
Maya Xi Meh
AI/ML Engineer building production-grade AI systems, LLM applications & intelligent products
Senior AI/ML Engineer · at PiperKoi · Sydney, New South Wales, Australia
Not currently taking on new work
About
I’m a Senior AI/ML Engineer focused on turning machine learning research and emerging AI capabilities into reliable, production-grade systems. I work across the full AI lifecycle — from problem formulation, data and experimentation through to model development, evaluation, deployment, observability, and continuous improvement.
My work spans LLM applications, generative AI, retrieval-augmented generation, agentic systems, recommendation and prediction systems, model evaluation, and ML infrastructure. I’m particularly interested in the engineering problems that emerge when AI moves from a prototype into a product: reliability, latency, cost, evaluation, security, scalability, and creating systems that people can actually trust.
I enjoy working at the intersection of software engineering, machine learning, and product. I’m equally comfortable designing an ML architecture, debugging a model failure, building an evaluation pipeline, or working with a product team to determine whether AI is genuinely the right solution to a problem.
What are you working on at work right now?
I’m currently working on production AI systems that combine large language models, retrieval, structured data, and traditional software systems. A major focus is making these systems dependable at scale — building evaluation frameworks, improving retrieval and reasoning quality, reducing latency and inference costs, and putting the right monitoring and guardrails around model behaviour.
The interesting part isn’t simply getting an LLM to produce a good answer. It’s engineering the surrounding system so that it continues to perform reliably across real users, messy data, changing models, and ambiguous requirements.
What's an interesting project you've shipped lately?
Recently, I shipped an AI-powered workflow that automated a previously manual knowledge-intensive process. The system combined retrieval, LLM reasoning, structured outputs, deterministic business logic, and automated evaluation rather than relying on a model alone.
The biggest engineering challenge was reliability. We built evaluation datasets around real-world failure cases, measured the system at each stage of the pipeline, and iterated on retrieval, prompting, model selection, and post-processing independently. The result was a system that could be evaluated and improved systematically rather than simply judged by whether individual responses “looked good.”
What's a problem you've been thinking about a lot?
I’ve been thinking a lot about how we move from impressive AI demos to AI systems that deserve to be trusted in production.
As models become increasingly capable, the bottleneck is often no longer the model itself. It’s evaluation, system design, data quality, observability, security, cost, and understanding failure modes. I’m particularly interested in developing better ways to measure AI systems against the outcomes that actually matter to users and businesses.
Work
Senior AI/ML Engineer, PiperKoi (Jul 2026 – Present)
Designed and deployed production ML and generative-AI systems serving real-world users. Built LLM applications using retrieval, tool use, structured generation, evaluation pipelines, and model orchestration. Developed scalable ML pipelines spanning data preparation, training, experimentation, deployment, monitoring, and retraining. Established automated evaluation frameworks for measuring model quality, hallucination, retrieval performance, latency, and cost. Improved inference efficiency through model selection, caching, batching, quantization, and architecture optimisation. Partnered with product and engineering teams to identify high-value applications of AI and translate ambiguous problems into measurable ML objectives. Mentored engineers on ML engineering, experimentation, system design, and production AI practices.
Projects
Production RAG & AI Knowledge System — Designed an end-to-end retrieval-augmented generation system combining document ingestion, chunking, embeddings, hybrid retrieval, reranking, LLM generation, citations, evaluation, and observability.
LLM Evaluation Framework — Built an evaluation framework for systematically measuring AI application quality across factuality, relevance, groundedness, retrieval quality, latency, cost, and safety. Used both deterministic tests and model-assisted evaluation.
Intelligent Recommendation Platform — Developed a recommendation system combining behavioural signals, machine-learning ranking, offline evaluation, and online experimentation to personalise user experiences.
ML Platform & MLOps — Designed reusable infrastructure for model training, experiment tracking, deployment, monitoring, feature management, and automated model lifecycle management.
Topics
- Artificial Intelligence
- Machine Learning
- Generative AI
- Large Language Models
- LLM Applications
- Retrieval-Augmented Generation
- RAG
- AI Agents
- Agentic AI
- Natural Language Processing
- Deep Learning
- MLOps
- ML Infrastructure
- Model Evaluation
- AI Evaluation
- LLM Evaluation
- Production AI
- Machine Learning Systems
- Recommender Systems
- Computer Vision
- Python
- PyTorch
- TensorFlow
- Kubernetes
- Docker
- AWS
- GCP
- Vector Databases
- Embeddings
- Model Serving
- Distributed Systems