memories-for-llms
ossConsolidating what a model learns about its user (preferences, habits, how they want things done) into LoRA adapter weights, so the model carries them itself. First public result: Unwritten Rules, a real personal rules file trained into a rank-16 adapter on Qwen3.5-9B.
memories-for-llms is about a model that gets better at understanding its user over time. What it learns about one person, their preferences, habits and ways of working, is consolidated offline into a per-user LoRA adapter, so it becomes part of the model rather than something the user has to repeat. The first public result is Unwritten Rules: Moving a Personal Rules File from the Prompt into LoRA Weights.
A rules file is the most explicit record of a person’s ways: people who use coding assistants keep one (formatting, words to avoid, how commit messages should read) and send it with every prompt. Unwritten Rules distils nine mechanically checkable rules from a real one into a rank-16 LoRA adapter on a 4-bit Qwen3.5-9B. A teacher with the rules in its prompt samples four answers per training prompt, per-rule checkers discard the ones that break a rule, and the student learns the survivors from the bare prompt. Each placement of the rules is then measured on what it costs: rule adherence over 360 held-out prompts, reasoning on GSM8K, knowledge on MMLU.

Prompting buys adherence at the price of reasoning, training the rules into the adapter buys more adherence at no such price, and doing both is worse than the adapter alone.
Findings
- The adapter follows the rules better than the prompt: +13.9 to +14.5 points of paired adherence over the strongest prompt in three adapters (+6.1 to +6.8 without R4, whose checker is strict about headers).
- The prompt costs reasoning, and the wording decides how much: the rules file as a system prompt costs 23.7 GSM8K points; two edits bring that to 9.0; an adapter trained on the edited wording scores 94.3 against the base’s 88.0, with MMLU unchanged.
- Never both: the adapter with the rules also in its prompt loses 25.0 and 11.3 GSM8K points against the adapter alone.
- Worked solutions are not needed: dropping them from training keeps GSM8K at 93.3, above the pre-registered 90.3 floor.
The paper also documents an evaluation pitfall: last-number GSM8K grading marks wrong 10.7 to 13.0 percent of header-first answers that contain the right number, against 4.0 percent for the unprompted model. The adapters come from two different teachers, and the study uses one base model, one seed and nine rules; the limitations section says where each of these limits the claims.
Experiments
e1: adapter on the rules file as written, with the 9B base itself as teacher.e2: three wordings of the rules as system prompts, no training.e3: adapter on the edited wording, with 142 checked worked solutions, Qwen3.6-35B-A3B as teacher.e4:e3without the worked solutions.
Every model output ships under results/eN/generations/, so the reports and every statistic in the paper can be re-scored without a GPU. Code is MIT; the paper, rules texts, prompts, training data and results are CC BY 4.0.