22 comments

  • aeneas_ory 3 hours ago
    All of these "hacks" are snakeoil and I think deep down we all know. Whether it's caveman, RTK, or whatever other vibe-coded productivity/token cost saving hacks/skills/claude.md.

    What I had success with (although benchmarks are older) is to index the codebase with a dedicated local code embedding model. It's a bit expensive on the CPU side but in my benchmarks it reduced token use and wall clock time significantly. Of course, it's always dependent on statistical noise + host system load, and running sufficiently large benchmarks is simply too expensive, so take em with a grain of salt.

    Why does it work you may ask? Well, LLMs basically brute force words/phrases and pipe that into find/grep/pgrep/whatever (or as recently discussed here write a python script for it - https://news.ycombinator.com/item?id=49654229). Semantic search looks for similarities so you have to do less brute forcing. Comes of course at the cost of indexing everything first.

    You can find the project here: https://github.com/ory/lumen

    • Whitespace 1 hour ago
      I should not trust their "vibe-coded productivity/token cost saving hacks" but I should trust yours?

          Save 30% token costs when using Claude Code, Codex, OpenCode for free - with open source, local semantic search. Works for small and large codebases and monorepos! Enterprise-ready and fully compliant via Ollama and SQLite-vec.
          Releases v0.0.42 Latest last month
      
      Why should I trust that what you're peddling isn't snakeoil?
      • icantevenhold 1 hour ago
        Only way to find out is to do some testing yourself i think.

        I’m using less tokens with Lumen but I also use a bunch of other tokens hacks/skills; it’s hard to measure the impact exactly but it feels significant

      • aeneas_ory 45 minutes ago
        I literally say you should take benchmarks with a grain of salt :)

        > Of course, it's always dependent on statistical noise + host system load, and running sufficiently large benchmarks is simply too expensive, so take em with a grain of salt.

        And the savings listed are coming from a benchmark harness that implements different OSS bugs one time with and one without lumen - in those cases the % saved are reproducible (caveat: it was on older models, Opus 4.6 I believe).

        Also I explain WHY it saves tokens - because the model doesn’t have to brute force different terms until it finds the match it needs, but uses semantic „distance“ so the embedding does it for the model.

      • huflungdung 1 hour ago
        [dead]
    • lopatin 15 minutes ago
      I'm in the process of evals for these tools after my org adopted them. My RTK findings are the same. It worsens task performance and overall you don't save money. I wanted to give the same treatment to other tools like ponytail and caveman (especially caveman, I mean there's no way that telling a computer to talk like a caveman is a valid engineering technique right?). To my horror, caveman is looking to be the only tool that actually doesn't regress on reasoning while taking costs down. But I still have a lot more evals to write, so this isn't conclusive or anything. (Also I haven't tried Lumen yet)
      • bunderbunder 5 minutes ago
        I am actually rather fond of caveman. I haven't evaluated it for token cost, in part because frankly I think that part of the pitch is a load of malarkey. Output that's shown to the user is such a small percentage of overall tokens these days.

        But anecdotally I do think it saves me quite a lot of time on reading LLM outputs. And that, if nothing else, is good for my sanity.

        The caveman gimmick makes sense to me as a clever hack. Caveman talk is a longstanding meme that's presumably well-represented in the models' training data. So just asking it to do that is just an ultra-concise way to tell the LLM to be ultra-concise. Which, in turn, is theoretically good for accuracy because putting too many instructions in the prompt is bad for task performance.

    • esperent 1 hour ago
      This sounds quite similar to dirac which made a stir a few months ago:

      https://github.com/dirac-run/dirac

      I spent way too long trying to reproduce the results in Pi and failing before I decided that I shouldn't trust author benchmarks for any of these tools. Then I found that I couldn't even close to reproduce their benchmark results using the exact model and their harness.

      If any person other than the author has time to verify these Lumen benchmark results I'd be curious to hear it. I don't have the time to do it myself at the moment.

      • nevon 21 minutes ago
        I did the same with another project that does the same thing, called ck, and wasn't able to wring out any improved performance over just plain grep.
    • Bridged7756 1 hour ago
      Jetbrains IDEs are a perfect solution for this. They expose IDE actions (e.g, search, see occurrences, go to implementation) in their MCP server, which the harnesses can then call directly instead of figuring out the code themselves.
    • ramon156 44 minutes ago
      Lumen is pretty cool, it's just local RAG, but it's a realistic approach at RAG. I would still keep tool-calling in some places though.
    • cassianoleal 1 hour ago
      > One of: Claude Code, Cursor, Codex, or OpenCode

      What makes it incompatible with Pi, Zed or any other harness?

      • aeneas_ory 44 minutes ago
        Only the amount of free time I have to work on it - nothing fundamentally prevents it. PRs welcomed!
  • ProjectBarks 1 hour ago
    It seems like most of these tools are mostly vaporware. Benchmarks done on Headroom and RTK show that neither result in real savings. If it were possible to have such a simple pre-process step why wouldn’t the AI Labs upstream the optimizations themselves? My guess is they mostly don’t work or make the behavior much more confusing for the model. I really think there needs to be some kind of independent benchmark.

    Here are other cases demonstrating the exact same issues with these kinds of tools:

    https://blog.jetbrains.com/ai/2026/07/rtk-claude-code-token-... https://brandonbarker.me/writing/headroom-fewer-tokens-bigge...

    • ericyd 34 minutes ago
      > If it were possible to have such a simple pre-process step why wouldn’t the AI Labs upstream the optimizations themselves?

      Not defending these tools, but one reason these might not be upstreamed is because it would negatively impact vendor margins, and they have no incentive to save their users money

    • grim_io 1 hour ago
      Even JetBrains is now AI blog-slop, how disappointing.
  • cityofdelusion 17 minutes ago
    Any magic tool that declares a savings of over 10% can be immediately classified as snake oil. You can check yourself, load any of those projects up in GitHub and notice the math is always extremely misleading. It will be something like theoretical input bytes, or amount of command stripped off, or some other lie.

    If the tool won’t be upfront about those things, they are not worth looking into any further. It’s used car salesman strategy.

  • lackoftactics 58 minutes ago
    I wrote article about it couple months ago that I didn't believe it works. Nice to see numbers now

    https://mroczek.dev/articles/the-token-compression-illusion-...

  • fg137 1 hour ago
    Glad to see that more and more people realize these are just snake oils. Without objective metrics like benchmarks, none of the claims mean anything.

    That's also how I feel about skills/plugins. While some provide important context for specific projects/environments, I am very skeptical about (over)generalized skills like "writing JS tests" or "creating a spec". There are dozens of these skills internally at my company, but I haven't seen a single benchmark that shows any of those are better than just plain, single sentence prompts in a meaningful way (aka statistically significant).

    • jasonjmcghee 57 minutes ago
      I see the same thing and have effectively the same philosophy. If I'm using something like figma or glean or playwright/chrome dev tools, plugin/skill/mcp - likely very useful.

      But so many of the weird collections of skills that people on YouTube get viral followings for - I just don't get it.

      People excitedly ask me what skills I use and I feel bad just saying only things we've directly authored for some express purpose. None of the "hot" ones.

      I've written a large handful of skills, but they aren't like vim plugins. I don't just leave them "on".

      This has been my experience at least- curious if I'm just behind the times.

      I also effectively didn't leave the IDE+ChatGPT copy/paste workflow until the first release of Claude code. So maybe I'm slow to adopt.

  • kgeist 1 hour ago
    The technique looked dubious from the start, because LLMs were trained to expect certain outputs from common bash tools. If the output is not what it expects, an LLM may issue more tool calls than before, because it will assume the tool is broken, the arguments passed to it were wrong, or it's a newer/older version of the tool etc => more tokens. Sounds like just adding to the prompt to use `grep` and `tail` extensively will do the trick without any special tooling.
    • jghn 2 minutes ago
      This was the problem I saw. I installed rtk when it came out and liked the idea of it. But over time with newer model generations I kept seeing the model get confused in the reasoning text and retry a command bypassing rtk. I didn't even need a benchmark to see it was regularly an impediment to the final outcome.
  • GodelNumbering 1 hour ago
    Some months ago I was evaluating command output compressors to integrate into Dirac[1] as that seemed like an easy win that would compliment and compound with Dirac's other mechanisms.

    I tested rtk among these and it was actually a net negative in both CPU time and accuracy, the latter would throw LLMs way off and make it hard to recover. If you are building a coding agent, I'd hard pass on rtk.

       ~ $ time grep Return * 2> /dev/null | wc -l
       966
       grep Return * 2> /dev/null 0.36s user 0.02s system 98% cpu 0.382 total
       wc -l 0.00s user 0.00s system 1% cpu 0.380 total
    
    
       ~ $ time rtk grep Return * 2> /dev/null | wc -l
       260
       rtk grep Return * 2> /dev/null 4.10s user 17.10s system 92% cpu 23.008 total
       wc -l 0.00s user 0.00s system 0% cpu 23.007 total
    
    
    Much worse CPU consumption, and more importantly, plain wrong result. These kind of results compromise the entire agent performance because the model trusts wrong output. Without the correct results, any hypothetical savings are penny wise pound foolish

    So yeah I am still on the lookout for a credible CLI wrapper, do let me know if you have any in mind.

    [1] https://dirac.run/

  • fwlr 2 hours ago
    This makes sense. “Don’t try to penny-pinch your employees” is a lesson most managers learn eventually, and I guess agent-orchestrators will have to learn it too.
  • gillesjacobs 2 hours ago
    Main takeaway:

      Average cost per attempt, without → with RTK:
      
      Claude/Fable: $1.72 → $1.64 (~5% cheaper)
      DeepSeek: $0.115 → $0.121 (~5% more expensive)
    
      Almost all Claude savings came from a single task.
      Excluding it, savings were under 1%.
    
    It took me a few rereads to parse out the top-line. This article really buries the lede.
  • psadri 1 hour ago
    We have been working in this space for the past year. Based on our experience, I no longer trust any claims unless they are backed by benchmark results (yes, benchmarks are painful to run reliably and expensive).

    It is possible to reduce token usage. It’s just much harder than the basic approach.

    • dist-epoch 1 hour ago
      One quick win is to just avoid wasteful tokens, for example run all the QA tools like the unit tests in --quiet mode, which only prints warnings/failures.
  • santiago-pl 29 minutes ago
    rtk gain mechanism is oversimplified. 1 token != 4 bytes for the standard prompt / context window at coding agent.
  • hokkos 2 hours ago
    If you are using maven you should tell your agent to use its quiet mode or rtk, because mvn love to write a lot of useless output.
    • jakozaur 49 minutes ago
      I believe RTK would work well in that use case.

      Sometimes creating less verbose variants yourself (a simple script, build.sh, with pointers to logs) can be a quick win.

  • fleetfox 2 hours ago
    I don't understand how this or all these magic skill bundles and methodologies get traction and why they are so popular. It's either plain worse or has serious trade offs.
    • daliusd 1 hour ago
      It is just "putting a wet phone in rice" of AI
  • sreekanth850 2 hours ago
    i don't know if such hacks works, but in C# if you use roslyn mcp, you save a lot.
    • CodesInChaos 10 minutes ago
      Which one specifically? The ones I found looked like unmaintained throw-away experiments.
    • lmeyerov 50 minutes ago
      Much earlier, I tried to set up some static analysis tools so that the coding agent would have access to dataflow analysis etc. tools instead of just grep for typed python. If there were benefits, they weren't easily apparent :(
    • xnorswap 1 hour ago
      I haven't tried it since it was first released but it didn't seem to work at all for me back then.

      It was so slow that the roslyn results would be lagged well behind any edits it was making, which would just leave it confused.

    • VulgarExigency 2 hours ago
      I don't think they're comparable. RTK just modifies the output of CLI tools to reduce the number of tokens, a Roslyn MCP gives the agent a fundamentally superior way of interacting with a C# codebase.
      • antupis 2 hours ago
        My main issue with rtk is that rtk randomly messes modification and agent start polling same tool continuously.
      • sreekanth850 2 hours ago
        Yes. and i find model makes less errors and reasoning the codebase well, especially when you do a large refactor.
  • semiquaver 2 hours ago
    Just another instance of the bitter lesson. The model itself knows how to be clever and conserve tokens in command output by using shell primitives and as the models get smarter they get better at anticipating large output and defensively adapting the input commands.
  • liam_ilands 0 minutes ago
    [flagged]
  • vrighter 3 hours ago
    well yeah.... now you're giving it output it wasn't trained on.
    • nextaccountic 2 hours ago
      What if the next-gen models are trained on RTK output as well? Then you will actually have less tokens in the context window, and the model won't become confused (which would require more turns, wasting tokens)
      • vrighter 1 hour ago
        doesn't change the fact that it doesn't do what it claims to now. I just don't care about vague promises and "trust us bro" vibes that tech is sold for nowadays. It claims x, it doesn't deliver x. Maybe it could in the future, or maybe not.
  • elian_ilands 2 hours ago
    [flagged]
  • saltypixel 1 hour ago
    [flagged]
  • yuzushi-dev 1 hour ago
    [dead]
  • lucaprata 36 minutes ago
    [flagged]