# Anish Shrestha > Anish Shrestha — machine learning and software engineer working on LLM infrastructure, agent tooling, and developer tools. Creator of verdant, second-brain, EVOKE, incr, cognitive-cache, certus, and wardrowbe. Open to ML, SWE, and LLM-engineering roles. Anish Shrestha (handle: anyesh) is a Senior Software Engineer at AlayaCare in Sydney, Australia, and the Australian engineering team's AI champion. Originally from Kathmandu, Nepal. Open to ML, SWE, and LLM-engineering roles in Sydney AU and remote. ## Pages - [Home](https://www.anyesh.me): overview, work timeline, project index, speaking, credentials. - [Projects](https://www.anyesh.me/projects): all 14 projects grouped by kind. - [Resume](https://www.anyesh.me/resume): full resume as HTML, plus a PDF download. - [Writing](https://www.anyesh.me/writing): posts and technical write-ups. - [Now](https://www.anyesh.me/now): current focus. - [Archive](https://www.anyesh.me/archive): older projects, 2018-2020. ## Experience ### Senior Software Engineer, AlayaCare (2022-05 to present) Backend and product engineering on a cloud-based SaaS for home care. Architected the Support at Home reconciliation pipeline, built a third-party-API testing harness, and lead AI coding-agent adoption as the Australian engineering team's AI champion. - Support at Home (2025–2026) — architected and built the entire reconciliation engine and pipeline for Services Australia's new Support at Home funding model (the successor to Home Care Packages). - Testing harness (2025–2026) — Services Australia doesn't provide sandbox or test endpoints, so I architected and built our own: a mock / monkey-patching layer that suppresses third-party API calls in-app and lets us run the rest of the system end-to-end against fixtures. - AI champion, Australian engineering team (2026) — proposed, built, and adopted across the team an end-to-end Cypress test-generation framework on top of Cursor's hooks, skills, and commands. Given a PRD or Jira ticket, the LLM explores the web app through browser tools, finds useful selectors, maintains memories and navigation indexes, and writes Cypress specs nearly one-shot. - Engineered HCP invoicing, claiming, smart reconciliation, and budget management. 30% reduction in customer support tickets. - Shipped the auto travel time feature in the scheduling and payroll workflow. 25% efficiency gain for clients. - Pioneered a core product delivery team driving technical design and strategic value for Australian clients. 100% compliance with Australian standards. Stack: Python, TypeScript, Vue.js, PostgreSQL, Celery, Redis, MySQL, Cursor, Cypress, NewRelic ### Machine Learning Engineer, Fusemachines (2020-02 to 2021-11) ML engineer across three Fusemachines clients: TIME Magazine, Hospital for Special Surgery, and Fuse AI. Pipelines, OLAP consolidation, viral-content prediction, implant supply chain optimization, and ML curriculum. - TIME Magazine — streamlined data pipelines for brand audits, insights, and strategic recommendations. Consolidated scattered data into a single OLAP system. - TIME Magazine — collaborated with the Data VP and a Ph.D. on a viral-content prediction ML pipeline. 90% accuracy. - TIME Magazine — automated reporting systems. 75% reduction in report generation time. - Hospital for Special Surgery — ML-driven implant supply chain optimization. Analyzed 15,000+ cases, predicted patient implant needs at over 90% accuracy. - Fuse AI — restructured and rebuilt the best ML and deep learning academic courses taught in Fuse Classroom to tens of thousands of students worldwide. Implemented core ML and NLP algorithms at the maths level to explain mechanics. Stack: Python, PyTorch, TensorFlow, Sklearn, GCP, Vertex AI, Linear Algebra, Probability, Calculus ## Open source - [EVOKE](https://www.anyesh.me/projects/evoke): OS-like memory management for the LLM KV cache. status in-progress; started 2025; stack Python, C++, llama.cpp fork, CUDA, Qwen, Jacobian lens; source https://github.com/Anyesh/EVOKE. - [redraft](https://www.anyesh.me/projects/redraft): Reactive LLM inference. When you edit an LLM's context, redraft recomputes only what the edit actually changed on both sides of the call: the prompt (reused KV prefix) and the answer (salvaged via self-speculative replay), implemented as a streaming mode inside llama.cpp. status shipped; started 2026; stack Python, C++, llama.cpp, FastAPI, pytest; source https://github.com/Anyesh/redraft. - [incr](https://www.anyesh.me/projects/incr): Incremental computation engine for Rust. Tracks dependencies between computations automatically and only reruns what's actually affected. status shipped; started 2023; stack Rust, Cargo, proptest, criterion; source https://github.com/Anyesh/incr. - [second-brain](https://www.anyesh.me/projects/second-brain): A universal memory layer for AI coding agents. Ingests conversation history from Claude Code, Cursor, and other tools into a single SQLite-backed store with vector and full-text search, then serves it via MCP so any AI agent can recall what you've discussed, decided, and built across projects and machines. status active; started 2025; stack Rust, SQLite, usearch, tantivy, BGE-small, MCP; source https://github.com/Anyesh/second-brain. - [verdant](https://www.anyesh.me/projects/verdant): Content-addressed cache for AI agent loops. Returns exact bytes from prior executions instead of re-running tools or re-calling LLMs. status active; started 2025; stack Rust, blake3, MCP; source https://github.com/Anyesh/verdant. - [cognitive-cache](https://www.anyesh.me/projects/cognitive-cache): Algorithmic context-window selection for LLM coding tools. Treats context as a constrained optimization problem, not retrieval. status shipped; started 2024; stack Python, scikit-learn, networkx, Hypothesis; source https://github.com/Anyesh/cognitive-cache. - [wardrowbe](https://www.anyesh.me/projects/wardrowbe): Put your wardrobe in rows. Snap. Organize. Wear. status shipped; started 2024; stack Next.js, TypeScript, FastAPI, Python, PostgreSQL, Redis, Docker, Ollama; source https://github.com/Anyesh/wardrowbe. - [memories-for-llms](https://www.anyesh.me/projects/memories-for-llms): LoRA-as-memory experiment: per-user durable memory lives in a rank-16 LoRA adapter on a Qwen3 student rather than in the prompt window. status in-progress; started 2025; stack Python, SQLite, QLoRA, unsloth, Qwen; source https://github.com/Anyesh/memories-for-llms. - [Certus](https://www.anyesh.me/projects/certus): A standard for AI-generated code to carry machine-checkable certificates of correctness, plus a Python reference implementation and a finetuned model that generates them. status shipped; started 2024; stack Python, Qwen 2.5 Coder, QLoRA, Hypothesis, unsloth; source https://github.com/Anyesh/certus. - [skillprobe](https://www.anyesh.me/projects/skillprobe): Automated testing for LLM skills. Launches Claude Code or Cursor as subprocesses, runs scenarios in isolated workspaces, and reports what passed and what didn't. status active; started 2025; stack Python, Claude Code, Cursor; source https://github.com/Anyesh/skillprobe. ## Proprietary - [wardrowbe.com](https://www.anyesh.me/projects/wardrowbe-cloud): The hosted version of wardrowbe. Virtual try-on, iOS and Android apps, cloud sync, fashion subscription. status active; started 2024; stack Next.js, TypeScript, FastAPI, Python, PostgreSQL, iOS, Android, Stable Diffusion. - [dreamery](https://www.anyesh.me/projects/dreamery): Generative-AI image SaaS. 2,000+ paying customers, 7,000+ images transformed. AI headshots, glamour shots, conceptual art. status shipped; started 2023; stack SvelteKit, ComfyUI, RunPod, DeepFace, Stripe, PostgreSQL, microservices. ## Side projects - [eon](https://www.anyesh.me/projects/eon): Artificial life simulator. Agents start with random neural networks and figure out how to survive through natural selection. status active; started 2024; stack Rust, Python, NEAT, Godot, WebSocket; source https://github.com/Anyesh/eon. - [art_gan](https://www.anyesh.me/projects/art-gan): GAN that generates modern art. status archived; started 2020; stack Python, GAN; source https://github.com/Anyesh/art_gan. ## Writing - [Running a 26B MoE on an 8 GB Jetson by streaming experts from SSD](https://www.anyesh.me/writing/moe-stream-jetson-orin) (2026-08-01): My Jetson Orin Nano has 8 GB of memory and Gemma 4 26B-A4B needs 13.3 GiB at Q4. I patched llama.cpp to stream the routed experts off the SSD instead, with logits bit-for-bit identical to the stock path, and recorded the model decoding on device. - [My context selector beat grep. An agent with grep beat it.](https://www.anyesh.me/writing/beating-grep-was-the-wrong-bar) (2026-07-31): I built cognitive-cache to pick which files an LLM should see, then spent two months measuring it. It beat naive grep by a statistically real margin, lost to a coding agent holding nothing but grep and read, and one of its six signals turned out to be worth exactly zero. Here is the whole arc, including the part where I killed it. - [J-space in practice: using Anthropic's Jacobian lens to decide what an LLM can forget](https://www.anyesh.me/writing/j-space-jacobian-lens-kv-cache-eviction) (2026-07-10): Anthropic's global workspace paper introduced J-space and the Jacobian lens. I turned the workspace readout into a KV cache eviction signal that beats SnapKV and H2O across three Qwen models, and shipped it in EVOKE. ## Skills - Languages: Python, Rust, TypeScript, JavaScript, SQL - LLM & agent infrastructure: llama.cpp (fork-level), KV-cache management, MCP servers, Model Context Protocol, Cursor / Claude Code harnesses, content-addressed caching, RAG, vector embeddings, knowledge graphs, QLoRA fine-tuning, end-to-end agent test frameworks - Frameworks & tools: PyTorch, TensorFlow, FastAPI, Django, Flask, Next.js, React, Svelte, Astro, GraphQL, Cypress, Celery, Redis - Infrastructure & cloud: Docker, GCP, AWS, RunPod, Linux, CUDA, GitHub Actions, Terraform ## Elsewhere - Email: sir.anishshrestha@gmail.com - GitHub: https://github.com/Anyesh - LinkedIn: https://www.linkedin.com/in/anyesh - Medium: https://medium.com/@anyesh - Learning Lab (interactive explainers): https://anyesh.github.io - Sitemap: https://www.anyesh.me/sitemap-index.xml