best local alternative to nsfwcharacterai?

TLDR
I've built local character-chat setups on modest cards before, and the RTX 4060 with 8GB is genuinely enough - as long as you stop dreaming about 70B models. Here's the honest starter stack I'd recommend after watching too many people burn a weekend on benchmark rabbit holes.
What actually runs on an 8GB card?
The stack that works is boring: **KoboldCpp** as your backend, **SillyTavern** as your frontend, and a 7B-9B roleplay finetune in GGUF format, quantized at Q4_K_M or Q5. That's it. That's the whole answer.
KoboldCpp is the easiest path because it runs GGUF files out of the box with a single executable - no Python environments, no dependency hell. Ollama or exllamav2 work too, and exllamav2 gives slightly better speed if you enjoy fiddling, but start with KoboldCpp.
On the model side, the sweet spot for 8GB VRAM is the 7B-9B class, mostly Mistral and Llama-based finetunes. A Q4_K_M quant of an 8B model is roughly 4.5 to 5GB on disk, which leaves you room for context - and context is where people get burned. The KV cache (the memory holding your conversation) plus buffers eats whatever is left over. A model that "fits" in 8GB does not mean it fits at 8k context. Realistically, plan for **4k to 8k tokens of context**, which is plenty for character chat but means long roleplays will need memory summaries rather than keeping everything in the window forever. If you overshoot, KoboldCpp can offload some layers to CPU - it still runs, just slower.
One thing I won't do: name a single "best" roleplay model as permanent truth. The roster turns over every few weeks. Check current community benchmarks and the chatter on /r/LocalLLaMA, look at recent roleplay finetunes on Hugging Face, and read the model cards - someone always posts VRAM usage at specific context lengths, which is exactly the number you need.
Is local actually better than a hosted NSFW site?
Privacy-wise, it's not close. Everything stays on your machine - no accounts, no logs, no wondering what a server is keeping. For anyone awkward about their chat history living on someone else's servers, that alone is worth the setup.
But be honest with yourself about the trade-offs. A hosted site gives you polish out of the box: slick UI, ready-made characters, zero configuration. Locally, expect 30 to 60 minutes of setup to get SillyTavern installed, connected to KoboldCpp, and loaded with character cards. The web UIs are clunkier, and features like group chat and world info work well once configured but they don't hand themselves to you. The reward is total control - no filter, unlimited chat length, custom characters, and it's free beyond the electricity.
And if you do want hosted.on-demand variety without any hardware fuss at all, live webcam chat is a different experience entirely - for that, I'd point people at xlovecam over the AI-character sites, because real people beat simulated ones more often than not.
Where do you go from here?
Download KoboldCpp, grab a fresh 8B roleplay GGUF at Q4_K_M, install SillyTavern, load a character card, and set context to 6k as a starting point. Then the real question: once you've tasted full privacy and control, will the hosted sites ever feel the same - or will you start eyeing a 12GB GPU upgrade instead?