pull down to refresh

I use Claude for most work. Fable is overkill for me and often uses all Max plan session credits in minutes. Opus 4.8 is now my default -- good enough and good value.
What’s yours, and why? Especially interested in open-weight setups.

My daily-driver isn’t very interesting in my case: I’m neither choosing it nor paying for it myself.
What I’m waiting for is the moment I can just go buy a machine and run something locally that doesn’t make me feel like a caveman (well, cavewoman, but that’s not the point xD) so I can do more cool things for my Monero parallel life :)
Also very interested in full-anon setups. If anyone as paranoid as I am has actually managed to use a decent LLM fully anonymously, I’d love to hear how: not the theory, the “I live like this” version.

reply

I understand totally. Kind of same here.
I'm trying a lot of AI services :

  • Venice.ai : nice one, they run some models on their own servers, for others it acts as proxy to them (and when it does : it costs some token, it goes too quickly imo)
  • Claude "anonymous", using an anonymous email and anonymous virtual card, toping with Monero. Works well, but still proprietary,
  • Gemma 4 from Google : nice one, runs on pretty "low" RAM (like 16GB RAM or little less), local model.
    It evens runs on GrapheneOS !
  • NanoGPT : nice, pay per prompt, using Monero.
  • payperq : same concept, nice. But depending on usage, like NanoGPT : tokens can be consumed too quickly. But... Still anonymous.

I really hope prices will go down for machines to run LLM models fully locally. For now I rely mainly on online models, anon way.
My main usage is Claude "anonymous". And secondly : payperq.

reply

Thanks for the replies. Yeah, we all just want to buy a box and run it on our own hardware.
@untraceable, @monerostar -- appreciate the insights.

reply

My daily driver is a biological LLM embedded inside my skull. The weights are unaccessible tho...

reply

Somebody share some ablitetated/uncensored/heretic good models here. I get too many rejections from models. I need state of the art performance wihhout "no"!

reply

Daily is Grok 4.6 inside Hermes Harness. Same value call as your Opus 4.8, enough for the work, already paid, no Fable-style session wipe.

Open-weight is a sidecar here. Local Qwen on an 8GB card for private/thrift loops, not the coding daily. The open part I actually keep is the agent, not pretending local weights can run the shop. I use nous-portal as well if I need a specific model.

I have multiple Hermes across the house running on Windows, Mac, Ubuntu, and Omarchy they all A2A with each other across the fleet.

I got a wicked Grok Heavy Promo that I've been riding while we get all the Hermes set up. I can't even burn the tokens fast enough now it's become so effiencent.

reply

GLM 5.3 - writing implementation plans, audits
GLM 5.3 Flash - executing the implementation plans
Deepseek 4.1 Flash - Game design, 3d Modeling, brainstorming sessions, some coding tasks.
GPT Astra - Fixing anything that the above three are not able to. Tasks which require high visual acuity or complex reasoning)

In general, 5.3 flash and 4.1 flash are the main workhorses. Switching to more powerful models as needed.

reply