Got something to have a conversation on? Ping me.

the timeline.

scroll down to move forward in time.

2022 · quarantine project

talking dinosaur.

I made this during quarantine, when I was first starting to learn coding and API integrations. The dinosaur talks in English, and I built its voice myself: I took open TTS models and changed the voice properties until it sounded like a dinosaur.

01 making the voice.

I lowered the pitch, increased the buffer and layered animal sounds on top to make the dinosaur voice effect.

dec 2025 · national science fair 1st place

quantum transmitter.

A device that transmits data through a laser. To demonstrate it, I sent music through light.

01 how it works.

The transmitter is a laser driven by dual ESP32s that break music down into notes: we programmed every musical note as a different pulse with a different interval. The receiver is an Arduino that decodes those pulses back into an analog signal, which goes through an AUX cable into a speaker.

02 the demo.

music, sent through a laser beam.

03 at the fair.

04 what came next.

We won first place at the National Science Fair as beginners. I went on to publish a paper at ICYS and was a finalist in the international research competition, with research on machine learning via light.

apr 2026 · scicon 2026 science fair

roylee & tugu.

RoyLee and Tugu are two sides of the same coin: they operate as one, but they're different. I built them for SciCon 2026. RoyLee is a device that holds a conversation and can spur discussion with people. It runs on a Mongolian LLM I trained, built when there were no natural Mongolian language models. Tugu is RoyLee's extension, the physical form of its agentic features: a device we built from scratch that takes commands from RoyLee and carries them out.

01 the model.

I started RoyLee from a small model in the 1B–8B parameter range (Llama 3 8B, Qwen, Phi-3). That size balances real language understanding with being able to iterate fast on a single 24–80 GB GPU. I fine-tuned it with QLoRA.

how I trained it
  • 4-bit base. I loaded the base model in 4-bit NormalFloat (NF4) with bitsandbytes. This cut training memory drastically, while the active adapters' forward and backward passes stayed in 16-bit precision.
  • Dataset. I curated 500–2,000 highly refined instruction–response pairs, covering standard inputs, complex edge cases, and cases where the model has to refuse to answer.
  • Formatting. I applied a strict ChatML template, so control tokens like <|im_start|> and <|im_end|> are identical across the whole corpus and the model doesn't learn confusion.
  • Sequence length. I truncated sequences to a fixed context window matched to the task. Unnecessarily long sequences scale memory quadratically and cause out-of-memory crashes.
  • Training loop. I ran the fine-tune with SFTTrainer from Hugging Face trl, using gradient accumulation to simulate batch sizes the GPU couldn't hold natively.
  • Adapter merging. I fused the trained LoRA weights back into the base model, removing the latency of separate adapter matrices at inference.
  • Benchmarking. I evaluated the final checkpoint against a prompt-only baseline on a held-out test set: schema adherence (valid JSON), hallucination rate and domain accuracy.

02 speaking mongolian.

Then I built a Mongolian LLM that talks in Mongolian naturally. I trained it on smaller models first, then optimised the data.

how I taught it mongolian
  • Groundwork. Open-source work like Dorjzodovsuren's Mongolian-Llama3-v0.1 and bayartsogt's speech and language models laid vital groundwork. To get a fluent small model, I had to fix the tokenizer first.
  • The tokenizer bottleneck. Base models like Llama 3 or Phi-3 have English-centric tokenizers, so a single Mongolian word splits into 5–6 subword tokens. That bloats the context window and kills inference speed.
  • New vocabulary. I trained a custom Byte-Pair Encoding tokenizer on a clean Cyrillic corpus (CC100-Mongolian) and injected the top k ≈ 15,000 Mongolian tokens. I resized the embedding matrix from E ∈ ℝV×d to E′ ∈ ℝ(V+k)×d and initialised each new token as the average of its subwords, to keep what the base model already knew.
  • Continued pre-training. Instruction tuning alone can't teach a new language, so I continued pre-training on Mongolian text (Wikipedia, legal corpora, news) with a next-token prediction objective. I unfroze the resized embeddings and the transformer blocks and used high-rank LoRA (r = 128) on every attention and feed-forward layer.
  • No forgetting. I mixed 5–10% English data into pre-training, so the model kept its logic, coding and reasoning.
  • Instruction tuning. I built a dataset of 20,000+ instruction–response pairs in fluent Mongolian. I used translated datasets like Alpaca as a starting point, then checked the pairs by hand to remove stiff, machine-translated syntax.
  • DPO. I trained on preference pairs with Direct Preference Optimization, so the model favours the native-sounding "chosen" answer over the translated-feeling "rejected" one.
  • Quantisation. I merged the LoRA adapters back into the base weights and quantised to 4-bit (GGUF / AWQ) so the model stays small enough for local servers and multi-agent setups.
  • Evaluation. English benchmarks translate poorly, so I built a custom test suite: Cyrillic grammar, Mongolian cultural knowledge and contextual comprehension.

03 the device.

For the demo, we set the model up on a server. A Raspberry Pi is the part you talk to: we added a touchscreen LCD, a microphone for your voice and a speaker module for RoyLee's. The Pi stores the inputs and responses. For speech we used Chimege, an open-source Mongolian STT and TTS that's still in development. Because it's open source, I could change its parameters and behaviour to work with my model.

04 the voice pipeline.

Talking to a model only feels natural if it answers quickly, so I built the pipeline to hide the wait.

how I built the pipeline

speech → model

  • Voice activity detection. I ran VAD on the hardware, so only real speech frames are sent to the STT server over WebSockets. That means less latency and less bandwidth.
  • Normalisation. Raw Mongolian STT output often lacks punctuation or has unformatted numbers, so I clean up the transcript (digits, acronyms, Cyrillic spacing) before it reaches the model.
  • Prompt formatting. I wrap the transcript in the same ChatML tokens the model was trained on, with hidden system instructions that keep answers short enough to speak.

model → speech

  • Sentence chunking. RoyLee doesn't wait for the full answer: I buffer the token stream until a complete Mongolian sentence or clause, and send each chunk straight to TTS.
  • Phonetic expansion. I wrote a normaliser that expands abbreviations and maths symbols into spoken Mongolian, so the TTS reads them with the right prosody.
  • Overlap. While chunk N is being spoken, chunk N+1 is already being generated. I downsample the audio to 16 kHz mono WAV to avoid buffer underruns on the speaker and display.

architecture

  • Decoupled services. Instead of one monolithic script, I connected STT, the LLM server and TTS through asynchronous message brokers (MQTT / gRPC) so they run in parallel. The model's thinking time hides behind the audio of the first sentence.

05 tugu: roylee's hands.

RoyLee talks; Tugu acts. We built Tugu from scratch on a Raspberry Pi, strapped on LEDs, motors and other systems, and trained it to receive commands from RoyLee and execute them. You can tell RoyLee to turn the servo to 90°, then to 180° after 10 minutes, then back to 90° after 5 more. Or ask it to turn on a light sequence for a Christmas party.

exampleyou   → "RoyLee, turn on a sequence for the Christmas party."
roylee → manifest_hardware({"device": "led_array", "pattern": "christmas", "duration_ms": 60000})
tugu  → {"status": "done"}  // telemetry back to roylee
how tugu works

tool calling

  • Structured output. Tugu can't act on conversational text, so I made RoyLee answer in deterministic function calls. I defined explicit tool schemas in the system prompt, like a manifest_hardware tool that takes {"device": "led_array", "pattern": "pulse", "duration_ms": 5000} or {"device": "dc_motor", "direction": "forward", "pwm_duty": 85}.
  • No parsing errors. I used a model fine-tuned for robust tool calling, plus grammar constraints that guarantee every output matches the JSON schema before it's sent.

messaging

  • MQTT pub/sub. Blocking HTTP requests don't work for physical agent loops, so I deployed a Mosquitto MQTT broker. RoyLee publishes its JSON tool calls to a command topic (agent/pi/command).
  • No waiting. RoyLee fires the command and goes straight to its next reasoning step, instead of pausing while a 10-second LED sequence plays out.

the hardware daemon

  • Always listening. Tugu runs a Python daemon that subscribes to the command topic, using paho-mqtt for the network and gpiozero for the hardware.
  • Non-blocking actuation. I run LED patterns (breathing effects, chase sequences) and motor ramping as asyncio tasks and worker threads, so Tugu stays responsive to an emergency STOP even while a motor is running.
  • Power isolation. Pi GPIO pins output 3.3 V at very little current, so I routed the PWM signals through an external motor driver (an L298N or an I²C stepper controller). It turns RoyLee's speed values into the right voltage for the motors without frying the board.

closing the loop

  • Telemetry. When an action finishes, or a sensor like a limit switch or encoder triggers, Tugu publishes a JSON update to agent/pi/telemetry.
  • Context injection. RoyLee's orchestrator subscribes to that topic and injects each update into the model's context as a tool observation, for example {"role": "tool", "content": "Motor reached position. Current draw spike detected."}. RoyLee then decides its next action from what actually happened. That feedback is what makes it an agent, not a remote control.

06 the build.

07 competition day.

may 2026 · startup mongolia 2026 1st place acquired

monbus.

We made an AI transportation app for Ulaanbaatar that estimates travel times from weekly traffic-jam data. We built it for elders and anyone in Mongolia who doesn't speak English: you talk to an AI in Mongolian, and it plots your bus route and how long the trip will take in traffic.

01 the app.

Bus. Tell the AI where you're going and it plots the bus routes in Mongolian, with how long the trip will take depending on the traffic.

Scooters. We worked with Tapatrip to bring the scooters around Ulaanbaatar into one unified location system. We developed the scooter location system specifically for MonBus, so people can see where the free scooters are, plus their battery level and condition. It made things much easier for riders and for Tapatrip.

02 first place.

MonBus won first place at Startup Mongolia 2026 and $6,000 in prize money. We put it into API costs and scaled MonBus to many more people.

03 acquired.

In August, MonBus was acquired by UBCard, a government-led company, for $30,000.

now · in progress

visi.

Coming soon…

Got something to have a conversation on? Ping me.