Projects
Things I've built, studies I've run, and ventures I've tried to get off the ground.
Production systems
AI Engineer Intern, Garuda Robotics
You can delete a bad answer. You can't un-fly a drone. That changes where the checks go.
- Agentic AI
- LLM Safety
- Systems Design
- Drones
Data Science Intern, SP Digital
My first system where someone was paid to break it. That changed how I build.
- LLM Safety
- Guardrails
- Evaluation
- Langfuse
- LangGraph
AI Architecture Strategy Engine
Personal project
A tool for deciding whether to use prompting, retrieval, or fine-tuning for a given job. Most teams pick by instinct or by whatever they read about last. But the costs are all knowable, so I built something that works it out: it scores each option against your budget and latency limits, runs a thousand simulations with the numbers nudged around, and tells you where the answer flips from one to another.
- Python
- Multi-Agent
- LLMs
- System Design
Socratic Digital Twin
Developer, NUS Faculty of Dentistry × School of Computing
A collaboration between the NUS Faculty of Dentistry and the School of Computing. It's an AI tutor for dental students that's built to refuse to answer. The subject is orthodontic clinical reasoning, where being handed the answer defeats the point, so the system asks questions back instead. That constraint drives everything: a multi-stage pipeline that decides what to ask next, retrieval over the faculty's own teaching material rather than the open web, and a review step where a clinician signs off on content before a student ever sees it. Currently in development.
- LangGraph
- RAG
- Postgres
- Clinical AI
Echolens — PII Redaction Evaluation
Data Science Intern, SP Digital
Echolens strips personal information out of customer call transcripts. I built the evaluation pipeline that measures how well it does that. Most of the work wasn't the measurement, it was deciding what counts: a receipt number isn't personal information, a partial email address probably isn't either, and those rules have to be written down and applied consistently before any score means anything. I scored it so that missing something counts as worse than being over-cautious, because those two errors aren't equally bad here.
- PII
- Evaluation
- NLP
LLM Evaluation Framework
Personal project
A tool for comparing language models on the same task and seeing what each one actually costs you in quality, money and speed. I ran it across three different kinds of work: summarising lectures, reasoning through business decisions, and ranking documents by relevance. The most useful thing it turned up was a measurement problem. The standard ways of scoring text similarity rate one model far worse than another purely because it wraps its answer in formatting, when both are ranking the documents equally well. For anything where the output has a structure, those metrics quietly mislead you, and you need one that measures the thing you actually care about.
- Evaluation
- Python
- Benchmarking
MarkBind — Open Source Contributions
Contributor
Contributions to MarkBind, an open-source documentation site generator maintained at NUS. I was picked for it off the back of the software engineering course. It was the first time I'd worked in a codebase I hadn't written any of, with a review process I had to satisfy.
- Open Source
- Java
- Documentation Tooling
TrackUp
Team project
A desktop contact and event manager for founders and small business owners, built as a team software engineering project. The deliberate choice in it is that everything is driven by typed commands with a graphical view alongside, rather than the other way round. That's the opposite of what most contact tools do, and it's right for the specific person who lives in a terminal and finds clicking through forms slower than typing what they want.
- Java
- JavaFX
- CLI
Research
Independent study
I wrote down the method before I looked at any data. Halfway through I found what I was testing for, then worked out the maths was producing it, not the market.
- Market Microstructure
- Pre-registration
- Statistics
- Python
Independent study
The model worked, but the metric was worthless and the monitoring was blind. A study in how machine learning projects fail even when the model itself is good.
- Machine Learning
- Evaluation
- Pre-registration
- Causal Inference
Singapore Society Simulation
Researcher, NUS Odyssey
Can AI agents stand in for real people when you want to know how a population thinks about something? I gave each agent a demographic profile and had them argue a policy question, then compared what they concluded against what real people had said online. They came out quite different. Changing who was in the simulation also moved the answer a long way, which means you'd be measuring your own choice of participants as much as anything else.
- LLM Agents
- Social Simulation
- Research
Ventures
Lecture AI
Case study
Trend Intelligence for Singapore F&B
Co-founder
We started building a trend tracker for people deciding where to eat, and stopped when it became clear that consumers won't pay for that and Google already owns food discovery. So we moved to the other side of the counter, where the pain is sharper and there's an actual budget: restaurant operators trying to work out what to put on a menu. The idea is still unproven and I'd rather say so than pretend otherwise. We're running customer interviews now, and one objection has already killed a business model we liked, which is that you can't sell a monthly subscription into an industry where most operators never turn a profit.
- Market Research
- F&B
- Product Discovery
Knocks
Solo build
A card game my friends and I have played for years, which existed only as a physical deck, so it needed everyone in the same room. I built the online version. It then spread by word of mouth to friends of friends I've never met, who kept playing it. Nothing about it was technically hard, and it's still the first thing I built that reached people I didn't tell about it.
- Real-time
- Multiplayer
- WebSockets
Sixer
Solo build
A cricket draft game we used to run over WhatsApp, badly. Group chats are a terrible place to hold a game with rules. Someone always has to arbitrate, someone always misses a turn, and the state of play lives in whoever scrolled back furthest. So I moved it somewhere that actually holds the rules and keeps everyone in sync.
- Real-time
- Multiplayer
- WebSockets
Earlier work
ChessPhere
Co-founder
I've played competitive chess since I was young, and when the pandemic shut down every over-the-board tournament, the thing that disappeared wasn't the game. Online chess was fine. What disappeared was the community around it. So a few friends and I started running virtual tournaments and workshops, and we were drawing over a hundred players to each one. It never made money and it was never going to. What it taught me is that a small niche can matter enormously to the people inside it, and that's a reasonable thing to build for.
- Community
- Events
- Chess
Donation Nation
Co-founder
This started with me giving away things from my own house to local NGOs during the pandemic. It worked, and it obviously didn't scale, because the bottleneck was one person with a car and a limited amount of stuff. So we built a platform connecting donors with communities that needed things, working through established NGOs and logistics partners rather than moving goods ourselves. The hard part was never the donating. It was that nobody could find each other.
- Social Impact
- Platform
- NGO Partnerships