Noob experience using local LLM as a D&D style DM.

ThreeJawedChuck@sh.itjust.works · 12 hours ago

It loads the rpc machine’s part of the model across the network every time you start the server,

I have to correct myself. It appears newer versions of rpc-server have a cache option and you can point them to a locally stored version of the model to avoid the network cost.

ThreeJawedChuck@sh.itjust.works · 1 day ago

Mistral (24B) models are really bad at long context, but this is not always the case. I find that Qwen 32B and Gemma 27B are solid at 32K

It looks like the Harbinger RPG model I’m using (from Latitude Games) is based on Mistral 24B, so maybe it inherits that limitation? I like it in other ways. It was trained on RPG games, which seems to help it for my use case. I did try some general purpose / vanilla models and felt they were not as good at D&D type scenarios.

It looks like Latitude also has a 70B Wayfarer model. Maybe it would do better at bigger contexts. I have several networked machines with 40GB VRAM between all them, and I can just squeak I4Q_XS x 70B into that unholy assembly if I run 24000 context (before the SWA patch, so maybe more now). I will try it! The drawback is speed. 70B models are slow on my setup, about 8 t/s at startup.

ThreeJawedChuck@sh.itjust.works · 2 days ago

Ah, great idea about the low temp for rules and high for creativity. I guess I can easily change it in the front end, although I also set the temp when I start the server, and I’m not sure which one takes priority. Hopefully the frontend does, so I can tweak it easily.

Also your post just got me thinking about the DRY sampler, which I’m using, but might be causing troubles for cases where the model legit should repeat itself, like an !inventory or !spells command. I might try to either disable it or add a custom break token, like the ! mark.

I think ST can show token probabilities, so I’ll try that too, thanks. I have so much to learn! I really should try other frontends though. ST is powerful in a lot of ways like dynamic management of the context, but there are other things I don’t like as much. It attaches a lot of info to a character that I don’t feel should be a property of a character. And all my D&D scenarios so far have been just me + 1 AI char, because even though ST has a “group chat” feature, I feel like it’s cumbersome and kind of annoying. It feels like the frontend was first designed around one AI char only, and then something got glued on to work around that limitation.

ThreeJawedChuck@sh.itjust.works · 2 days ago

Will do, thanks for the tip. Your description does sound like a good fit for the idea. As long as it supports network inference between machines with heterogeneous cards, it would work for what I have in mind.

ThreeJawedChuck@sh.itjust.works · 5 days ago

deleted by creator

ThreeJawedChuck@sh.itjust.works · 5 days ago

Or be a 90s computer text adventure

Zork on steroids!

ThreeJawedChuck@sh.itjust.works · 5 days ago

Thanks for your comments and thoughts! I appreciate hearing from more experienced people.

I feel like a little bit of prompt engineering would go a long way.

Yah, probably so. I tried to write a system prompt to steer the model toward what I wanted, but it’s going to take a lot more refinement and experimenting to dial it in. I like your idea of asking it to be unforgiving about rules. I hadn’t put anything like that in.

That’s a great idea about putting a D&D manual, or at least the important parts, into a RAG system. I haven’t tried RAG yet but it’s on my queue of matters to learn. I know what it is, I just haven’t tried it yet.

I’ve for sure seen that the quality of output starts to decline about 16K context, even on models that claim to support 128K. Also, I feel like the system prompt seems more effective when there are only let’s say 4K context tokens so far. After the context grows, the model becomes less and less inclined to follow the system prompt. I’ve been guessing this is because as the context grows, any given piece of it becomes more dilute, but I don’t really know.

For those reasons, I’m trying to use summarization to keep the context size under control, but I haven’t found a good approach yet. SillyTavern has an auto summary injecting system, but either I’m misunderstanding it, or I don’t like how it works, and I end up doing it manually.

I tried a few CoT models, but not since I moved to ST as a front end. I was using them with the standard llama-server web interface, which is a rather simple affair. My problem was that the thinking output seemed to spam up the context, leaving me much less ctx space for my own use. Each think block was like 500-800 tokens. It looks like ST might have an ability to only keep the most recent think block in the context, so I need to do more experimenting. The other problem I had was that the thinking could just take a lot of time.

ThreeJawedChuck@sh.itjust.works · edit-2 7 days ago

What was your setup for this experiment?

I’m using llama.cpp + sillytavern. I’m very much in learning mode with ST however, so I’m confident I could be using it in a more effective manner than I know how to at the moment. It seems like koboldcpp + ST ought to be similar to what I’m doing.

ThreeJawedChuck@sh.itjust.works · 7 days ago

but it’s no fun when the LLM simply says “yeah, sure whatever.” I

I hear ya. LLMs tend to heavily tilt toward what the user wants, which is not ideal for an RPG.

Have you tried any of the specialized RPG models? The one I’m using now has, at least twice so far, put me into a situation where I felt my party (2 chars, me and the AI) were going to die unless we ran away. We just finished a very difficult fight, used everything at our disposal, and sustained several serious injuries in the process. Then an even more powerful foe appeared, and it really felt that was going to be the end unless we ran. Would it really have killed us? I can’t say, but I did get a genuine sense of that. It might help that in the system prompt, I had put this:

The story should be genuinely dangerous and frightening, but survivable if we use our wits.

I have the feeling the generalist models are much more tilted in the “yeah, sure, whatever” direction. I tried at least one RPG focused model (Dan’s dangerous winds, or something like that) which was downright brutal, and would kill me off right away with no opportunity for me to do anything about it. That wasn’t fun for the opposite reason. But like you say, it’s also not fun to have no risk and no boundaries to test one’s mettle. The sweet spot is can be elusive.

I’m thinking that a non-LLM rules system around an LLM for descriptive purposes could really help here too, to enforce a kind of rigor on the experience.

ThreeJawedChuck@sh.itjust.works · edit-2 7 days ago

To add to my lame noob answer, I found this, which has a better rundown of ollama vs llama.cpp. I don’t know if it’s considered bad form to link to ##ddit on lemmy, so ~~I’ll just put the title here and you can search for it on there if you want~~ link added per comment from mutual_ayed below. There are a couple informative posts which are upvoted. “There is a big difference between use LM-Studio, Ollama, LLama.cpp?”

ThreeJawedChuck@sh.itjust.works · 7 days ago

Noob experience using local LLM as a D&D style DM.

ThreeJawedChuck@sh.itjust.works · 7 days ago

What’s the advantage over Ollama?

I’m very new to this so someone more knowledgeable should probably answer this for real.

My impression was that ollama somehow uses the llama.cpp source internally, but wraps it up to provide features like auto-downloading of models. I didn’t care about that, but I liked the very tiny dependency footprint of llama.cpp. I haven’t tried ollama for network inference.

There are other backends too which support network inference, and some posts allege they are better for that than llama.cpp is. vllm and … exllama or something like that? I haven’t looked into either of them. I’m running on inertia so far with llama.cpp, since it was so easy to get going and I’m kinda lazy.

ThreeJawedChuck@sh.itjust.works · 7 days ago

I like this project. Very nice!

I haven’t tried RAG yet, nor the fancy vector space whatsit which looks like it requires a specialized model(?) to create. I’ve been wanting to do something similar in spirit to your project here, but for an online RPG, so I dig this.

ThreeJawedChuck@sh.itjust.works · 7 days ago

Don't overlook llama.cpp's rpc-server feature.