01 / Systems · AI · Embedded

Systems software, built to be measured.

Four and a half years on both sides of the same problem. I build the C++, C, Python and Linux systems underneath AI infrastructure, and I design the adversarial problems and harnesses that decide whether generated code is actually correct.

Arun KumarSystems / AI / Embedded
Thanjavur, Tamil Nadu

02 / Selected work

Four systems I own end to end

01

Cerberus AI

Python → Pybind11 → C++17 → C11 arena

Most agent frameworks inherit the host language’s memory behaviour: allocate freely, pause for GC, pay for it in long sessions. Cerberus moves the hot path below that line, so a long-running session holds a flat memory profile instead of a sawtooth.

Read the architecture
Python agent APIdefine tools · run(task) Pybind11 bridgeGIL release · marshalling C++17 orchestratorThink → Plan → Act FastMCPtool routing C11 arena · AVX2-aligned
A single agent step, from Python call to arena write.

02

Multithreaded TCP server

Accept → bounded queue → fixed worker pool

A thread per client means a thousand clients cost roughly eight gigabytes of stack before any work happens. A fixed pool behind a mutex-guarded queue keeps resource usage flat no matter how many connections arrive.

Read the case study
thread per client 1000 × 8 MB ≈ 8 GB fixed pool 4 × 8 MB = 32 MB workers sleep on a condition variable, they do not spin
What the two strategies actually cost.

03

Bare-metal game engine

STM32F429I · Cortex-M4 · no RTOS, no HAL

An FSM-driven engine on ARM Cortex-M4 with nothing between the game and the panel. Hand-written SPI and I2C drivers for the ILI9341 display and STMPE811 touch controller, and a hunt-and-target opponent seeded from the hardware RNG.

Read the case study
An STM32F429I-DISCO development board: Cortex-M4 microcontroller, 8 MHz crystal, and the 2.4 inch panel it drives.
STM32F429I-DISCO, the board this ran on. Stock product photograph of the same board, panel shown unlit.

04

CAN body control module

Multi-node · arbitration · RTOS timing

Four control modules on one differential pair, terminated at 120 ohm. Arbitration is decided bit by bit: the lower identifier holds the bus and the loser backs off, retries, and loses nothing.

Read the case study
HL 120Ω120Ω Door Lighting Dashboard Gateway
Four modules, one differential pair, 120Ω at each end.
03 / How I work

I don’t just build systems. I measure them.

“A number without a method is marketing.” Working principle
  1. 01

    Define

    Name the constraint before writing anything. Contention shape and failure mode decide the design, not the other way round.

  2. 02

    Build

    Smaller surfaces, explicit failure paths, and comments that explain the constraint rather than the syntax.

  3. 03

    Measure

    Benchmark against a named baseline on a stated machine, keep the harness in the repo, and say plainly when a result is single-machine rather than rigorous.

  4. 04

    Break

    Most latency tails I have chased came from allocation behaviour, not algorithms. Find what is holding the rail down, and prove it.

  5. 05

    Optimise

    Arenas, slabs and explicit lifetimes make performance predictable, and make ASan and Valgrind output meaningful instead of noisy.

04 / Proof

Cold-start memory, measured

Cerberus versus LangChain on the same task, on the same machine. Resident memory sampled after the first completed agent step, so lazy imports and model-client initialisation are included on both sides.

Single-machine figures from my own Linux workstation, not a controlled lab. Directionally reliable and reproducible on request. I would not present them as vendor benchmarks.

Method & full benchmarks
Cerberus AI208 MB
LangChain424 MB

~2 ns

Arena allocation

Bump pointer, AVX2-aligned

< 0.5 ms

Context prune

O(1), independent of history

272k+

Hot-loop ops/sec

Core state machine, one thread

05 / Technical range

Three domains, one stack

Languages C++17/20/23 C11 Python 3 Embedded C CUDA C++

Systems

  • C++ · C
  • Linux · POSIX
  • Memory management
  • Concurrency
  • Networking
  • IPC / RPC
  • Performance engineering

AI

  • Python 3
  • Pybind11
  • FastMCP · Ollama
  • Agent architectures
  • Evaluation
  • Benchmarking

Embedded

  • Embedded C
  • ARM Cortex-M4 · STM32
  • CAN
  • SPI / I2C
  • LTDC · SDRAM
  • Bare metal
Arun Kumar
06 / Identity

Arun Kumar

Systems / AI / Embedded engineer

I work at the boundary where software meets the machine it runs on. Before the systems work there were two years of board-level repair, which is where the habit came from: never replace a part you cannot first prove is wrong.

That habit is the same one I run on a profiler now, and it is why every performance number on this site arrives with the machine it was measured on.

07 / Experience

Track record

  1. Jun 2026 —

    Software EngineerHandshake AI

    Adversarial problem design and Docker-sandboxed evaluation harnesses for frontier model validation.

  2. 2025 – 2026

    Consultant & SMEEcademic Tube

    150+ C++, Python and Linux solutions reviewed; coding agents benchmarked head to head on systems tasks.

  3. Jan – May 2025

    AI Training ContributorOutlier · Scale AI

    120+ C++ submissions reviewed and debugged, feeding structured assessments into fine-tuning datasets.

  4. May – Dec 2024

    Embedded Systems EngineerVector India

    Full-time programme in Linux internals, RTOS, TCP/IP and ARM, ending in a multi-node CAN control project.

  5. 2022 – 2024

    Hardware & Firmware SpecialistSelf-employed

    Board-level diagnostics and BIOS recovery across Intel, AMD and Apple hardware. Schematic tracing, not part swapping.

Full experience & credentials