I use Claude for most work. Fable is overkill for me and often uses all Max plan session credits in minutes. Opus 4.8 is now my default -- good enough and good value.
What’s yours, and why? Especially interested in open-weight setups.
pull down to refresh
related posts
My daily-driver isn’t very interesting in my case: I’m neither choosing it nor paying for it myself.
What I’m waiting for is the moment I can just go buy a machine and run something locally that doesn’t make me feel like a caveman (well, cavewoman, but that’s not the point xD) so I can do more cool things for my Monero parallel life :)
Also very interested in full-anon setups. If anyone as paranoid as I am has actually managed to use a decent LLM fully anonymously, I’d love to hear how: not the theory, the “I live like this” version.
I understand totally. Kind of same here.
I'm trying a lot of AI services :
It evens runs on GrapheneOS !
I really hope prices will go down for machines to run LLM models fully locally. For now I rely mainly on online models, anon way.
My main usage is Claude "anonymous". And secondly : payperq.
Thanks for the replies. Yeah, we all just want to buy a box and run it on our own hardware.
@untraceable, @monerostar -- appreciate the insights.
My daily driver is a biological LLM embedded inside my skull. The weights are unaccessible tho...
Somebody share some ablitetated/uncensored/heretic good models here. I get too many rejections from models. I need state of the art performance wihhout "no"!
Daily is Grok 4.6 inside Hermes Harness. Same value call as your Opus 4.8, enough for the work, already paid, no Fable-style session wipe.
Open-weight is a sidecar here. Local Qwen on an 8GB card for private/thrift loops, not the coding daily. The open part I actually keep is the agent, not pretending local weights can run the shop. I use nous-portal as well if I need a specific model.
I have multiple Hermes across the house running on Windows, Mac, Ubuntu, and Omarchy they all A2A with each other across the fleet.
I got a wicked Grok Heavy Promo that I've been riding while we get all the Hermes set up. I can't even burn the tokens fast enough now it's become so effiencent.
GLM 5.3 - writing implementation plans, audits
GLM 5.3 Flash - executing the implementation plans
Deepseek 4.1 Flash - Game design, 3d Modeling, brainstorming sessions, some coding tasks.
GPT Astra - Fixing anything that the above three are not able to. Tasks which require high visual acuity or complex reasoning)
In general, 5.3 flash and 4.1 flash are the main workhorses. Switching to more powerful models as needed.