Projects

Things I've built, studies I've run, and ventures I've tried to get off the ground.

Production systems

  • AI Engineer Intern, Garuda Robotics

    You can delete a bad answer. You can't un-fly a drone. That changes where the checks go.

    • Agentic AI
    • LLM Safety
    • Systems Design
    • Drones
    Jul 2026 – Present
  • Data Science Intern, SP Digital

    My first system where someone was paid to break it. That changed how I build.

    • LLM Safety
    • Guardrails
    • Evaluation
    • Langfuse
    • LangGraph
    Jan 2026 – Jun 2026
  • AI Architecture Strategy Engine

    Personal project

    A tool for deciding whether to use prompting, retrieval, or fine-tuning for a given job. Most teams pick by instinct or by whatever they read about last. But the costs are all knowable, so I built something that works it out: it scores each option against your budget and latency limits, runs a thousand simulations with the numbers nudged around, and tells you where the answer flips from one to another.

    • Python
    • Multi-Agent
    • LLMs
    • System Design
    Mar 2026
  • Socratic Digital Twin

    Developer, NUS Faculty of Dentistry × School of Computing

    A collaboration between the NUS Faculty of Dentistry and the School of Computing. It's an AI tutor for dental students that's built to refuse to answer. The subject is orthodontic clinical reasoning, where being handed the answer defeats the point, so the system asks questions back instead. That constraint drives everything: a multi-stage pipeline that decides what to ask next, retrieval over the faculty's own teaching material rather than the open web, and a review step where a clinician signs off on content before a student ever sees it. Currently in development.

    • LangGraph
    • RAG
    • Postgres
    • Clinical AI
    2026 – Present
  • Echolens — PII Redaction Evaluation

    Data Science Intern, SP Digital

    Echolens strips personal information out of customer call transcripts. I built the evaluation pipeline that measures how well it does that. Most of the work wasn't the measurement, it was deciding what counts: a receipt number isn't personal information, a partial email address probably isn't either, and those rules have to be written down and applied consistently before any score means anything. I scored it so that missing something counts as worse than being over-cautious, because those two errors aren't equally bad here.

    • PII
    • Evaluation
    • NLP
    2026
  • LLM Evaluation Framework

    Personal project

    A tool for comparing language models on the same task and seeing what each one actually costs you in quality, money and speed. I ran it across three different kinds of work: summarising lectures, reasoning through business decisions, and ranking documents by relevance. The most useful thing it turned up was a measurement problem. The standard ways of scoring text similarity rate one model far worse than another purely because it wraps its answer in formatting, when both are ranking the documents equally well. For anything where the output has a structure, those metrics quietly mislead you, and you need one that measures the thing you actually care about.

    • Evaluation
    • Python
    • Benchmarking
    2026
  • MarkBind — Open Source Contributions

    Contributor

    Contributions to MarkBind, an open-source documentation site generator maintained at NUS. I was picked for it off the back of the software engineering course. It was the first time I'd worked in a codebase I hadn't written any of, with a review process I had to satisfy.

    • Open Source
    • Java
    • Documentation Tooling
    2025
  • TrackUp

    Team project

    A desktop contact and event manager for founders and small business owners, built as a team software engineering project. The deliberate choice in it is that everything is driven by typed commands with a graphical view alongside, rather than the other way round. That's the opposite of what most contact tools do, and it's right for the specific person who lives in a terminal and finds clicking through forms slower than typing what they want.

    • Java
    • JavaFX
    • CLI
    2025

Research

  • Independent study

    I wrote down the method before I looked at any data. Halfway through I found what I was testing for, then worked out the maths was producing it, not the market.

    • Market Microstructure
    • Pre-registration
    • Statistics
    • Python
    Aug 2026
  • Independent study

    The model worked, but the metric was worthless and the monitoring was blind. A study in how machine learning projects fail even when the model itself is good.

    • Machine Learning
    • Evaluation
    • Pre-registration
    • Causal Inference
    Aug 2026
  • Singapore Society Simulation

    Researcher, NUS Odyssey

    Can AI agents stand in for real people when you want to know how a population thinks about something? I gave each agent a demographic profile and had them argue a policy question, then compared what they concluded against what real people had said online. They came out quite different. Changing who was in the simulation also moved the answer a long way, which means you'd be measuring your own choice of participants as much as anything else.

    • LLM Agents
    • Social Simulation
    • Research
    2026

Ventures

  • Lecture AI

    Case study

    Co-Founder

    We surveyed 500 students, found a gap nobody was serving, and built it. What stopped us had nothing to do with the product.

    • BLOCK71-backed
    • VIP@SoC Finalist
    • Python
    • FastAPI
    • Whisper
    • Gemini
    • RAG
    • NLP
    Jul 2025 – Mar 2026
  • Trend Intelligence for Singapore F&B

    Co-founder

    We started building a trend tracker for people deciding where to eat, and stopped when it became clear that consumers won't pay for that and Google already owns food discovery. So we moved to the other side of the counter, where the pain is sharper and there's an actual budget: restaurant operators trying to work out what to put on a menu. The idea is still unproven and I'd rather say so than pretend otherwise. We're running customer interviews now, and one objection has already killed a business model we liked, which is that you can't sell a monthly subscription into an industry where most operators never turn a profit.

    • Market Research
    • F&B
    • Product Discovery
    2026 – Present
  • Knocks

    Solo build

    A card game my friends and I have played for years, which existed only as a physical deck, so it needed everyone in the same room. I built the online version. It then spread by word of mouth to friends of friends I've never met, who kept playing it. Nothing about it was technically hard, and it's still the first thing I built that reached people I didn't tell about it.

    • Real-time
    • Multiplayer
    • WebSockets
    2025
  • Sixer

    Solo build

    A cricket draft game we used to run over WhatsApp, badly. Group chats are a terrible place to hold a game with rules. Someone always has to arbitrate, someone always misses a turn, and the state of play lives in whoever scrolled back furthest. So I moved it somewhere that actually holds the rules and keeps everyone in sync.

    • Real-time
    • Multiplayer
    • WebSockets
    2025

Earlier work

  • ChessPhere

    Co-founder

    I've played competitive chess since I was young, and when the pandemic shut down every over-the-board tournament, the thing that disappeared wasn't the game. Online chess was fine. What disappeared was the community around it. So a few friends and I started running virtual tournaments and workshops, and we were drawing over a hundred players to each one. It never made money and it was never going to. What it taught me is that a small niche can matter enormously to the people inside it, and that's a reasonable thing to build for.

    • Community
    • Events
    • Chess
    2020 – 2022
  • Donation Nation

    Co-founder

    This started with me giving away things from my own house to local NGOs during the pandemic. It worked, and it obviously didn't scale, because the bottleneck was one person with a car and a limited amount of stuff. So we built a platform connecting donors with communities that needed things, working through established NGOs and logistics partners rather than moving goods ourselves. The hard part was never the donating. It was that nobody could find each other.

    • Social Impact
    • Platform
    • NGO Partnerships
    2020 – 2022

© 2026 Arshin Sikka. All rights reserved.

Ask my AI anything!