Anish Shrestha

Senior Software Engineer at AlayaCare, Sydney. Building LLM infrastructure on the side. Open to ML / SWE / LLM-engineering roles.

skills

Languages
Python, Rust, TypeScript, JavaScript, SQL
LLM & agent infrastructure
llama.cpp (fork-level), KV-cache management, MCP servers, Model Context Protocol, Cursor / Claude Code harnesses, content-addressed caching, RAG, vector embeddings, knowledge graphs, QLoRA fine-tuning, end-to-end agent test frameworks
Frameworks & tools
PyTorch, TensorFlow, FastAPI, Django, Flask, Next.js, React, Svelte, Astro, GraphQL, Cypress, Celery, Redis
Infrastructure & cloud
Docker, GCP, AWS, RunPod, Linux, CUDA, GitHub Actions, Terraform

experience

Senior Software Engineer · AlayaCare ↗

May 2022 – present · Sydney, Australia

Backend and product engineering on a cloud-based SaaS for home care. Architected the Support at Home reconciliation pipeline, built a third-party-API testing harness, and lead AI coding-agent adoption as the Australian engineering team's AI champion.

  • Support at Home (2025–2026) — architected and built the entire reconciliation engine and pipeline for Services Australia's new Support at Home funding model (the successor to Home Care Packages).
  • Testing harness (2025–2026) — Services Australia doesn't provide sandbox or test endpoints, so I architected and built our own: a mock / monkey-patching layer that suppresses third-party API calls in-app and lets us run the rest of the system end-to-end against fixtures.
  • AI champion, Australian engineering team (2026) — proposed, built, and adopted across the team an end-to-end Cypress test-generation framework on top of Cursor's hooks, skills, and commands. Given a PRD or Jira ticket, the LLM explores the web app through browser tools, finds useful selectors, maintains memories and navigation indexes, and writes Cypress specs nearly one-shot.
  • Engineered HCP invoicing, claiming, smart reconciliation, and budget management. 30% reduction in customer support tickets.
  • Shipped the auto travel time feature in the scheduling and payroll workflow. 25% efficiency gain for clients.
  • Pioneered a core product delivery team driving technical design and strategic value for Australian clients. 100% compliance with Australian standards.

Python · TypeScript · Vue.js · PostgreSQL · Celery · Redis · MySQL · Cursor · Cypress · NewRelic

Machine Learning Engineer · Fusemachines ↗

Feb 2020 – Nov 2021 · Kathmandu, Nepal

ML engineer across three Fusemachines clients: TIME Magazine, Hospital for Special Surgery, and Fuse AI. Pipelines, OLAP consolidation, viral-content prediction, implant supply chain optimization, and ML curriculum.

  • TIME Magazine — streamlined data pipelines for brand audits, insights, and strategic recommendations. Consolidated scattered data into a single OLAP system.
  • TIME Magazine — collaborated with the Data VP and a Ph.D. on a viral-content prediction ML pipeline. 90% accuracy.
  • TIME Magazine — automated reporting systems. 75% reduction in report generation time.
  • Hospital for Special Surgery — ML-driven implant supply chain optimization. Analyzed 15,000+ cases, predicted patient implant needs at over 90% accuracy.
  • Fuse AI — restructured and rebuilt the best ML and deep learning academic courses taught in Fuse Classroom to tens of thousands of students worldwide. Implemented core ML and NLP algorithms at the maths level to explain mechanics.

Python · PyTorch · TensorFlow · Sklearn · GCP · Vertex AI · Linear Algebra · Probability · Calculus

selected open source

  • EVOKE

    OS-like memory management for the LLM KV cache.

    Python · C++ · llama.cpp fork · CUDA · Qwen · Jacobian lens · source ↗

  • redraft

    Reactive LLM inference. When you edit an LLM's context, redraft recomputes only what the edit actually changed on both sides of the call: the prompt (reused KV prefix) and the answer (salvaged via self-speculative replay), implemented as a streaming mode inside llama.cpp.

    Python · C++ · llama.cpp · FastAPI · pytest · source ↗

  • incr

    Incremental computation engine for Rust. Tracks dependencies between computations automatically and only reruns what's actually affected.

    Rust · Cargo · proptest · criterion · source ↗

  • second-brain

    A universal memory layer for AI coding agents. Ingests conversation history from Claude Code, Cursor, and other tools into a single SQLite-backed store with vector and full-text search, then serves it via MCP so any AI agent can recall what you've discussed, decided, and built across projects and machines.

    Rust · SQLite · usearch · tantivy · BGE-small · MCP · source ↗

  • verdant

    Content-addressed cache for AI agent loops. Returns exact bytes from prior executions instead of re-running tools or re-calling LLMs.

    Rust · blake3 · MCP · source ↗

  • wardrowbe

    Put your wardrobe in rows. Snap. Organize. Wear.

    Next.js · TypeScript · FastAPI · Python · PostgreSQL · Redis · Docker · Ollama · source ↗

All 13 projects →

speaking

  • 2024-06 Google I/O Extended "What does the future of AI hold for us" Panelist · NSW Teachers Federation Conference Centre, Sydney
  • 2024-05 Google Developer Student Clubs — Google Labs "Leveraging open-source text-to-image generative AI for practical applications" Speaker · Google HQ, Sydney
  • 2024-04 Google Developer Group — Google Cloud "Dreamery: generative-AI on GCP. Distributed services, serverless GPU, 90% cost reduction" Speaker · Google Developer Group, Sydney
  • 2020-11 Fuse AI Training "End-to-end ML pipeline development and deployment workshop" AI Instructor · Fusemachines

education

  • Bachelor of Science in Computer Science Lord Buddha Education Foundation (APU) · Kathmandu, Nepal 2016 – 2019

certifications