Texas Instruments
July 2025 – July 2026Software Engineer, Edge AI – Full-Time
- Designed and implemented 7 low-level-neural-network operator kernels — Add, Sub, Mul, Conv2D, Linear, AvgPool2D, Sigmoid — for a custom Neural Processing Unit (NPU), optimizing tiling, DMA scheduling, memory-bank access and schedule generation to support downstream model deployment.
- Implemented N-dimensional broadcasting for Add/Sub/Mul via compact DMA and vector-SRAM replay, enabling complete coverage across 8 broadcast categories and broader operator support.
- Built a reusable deep-learning test framework validating kernels against PyTorch reference across 10,000+ test cases and 5 numeric precisions, diagnosing accelerator defects including OOB DMA, memory corruption, NaN propagation, and BF16 drift ahead of hardware bring-up.
- Built a Claude skill to auto-generate kernel code from finalized design specifications, cutting new-operator development time by ~4× and accelerating the node-library roadmap.
- Deployed production object-detection models to INT8 via quantization-aware distillation and an ONNX-to-PyTorch pipeline, retaining 97–100% FP32 accuracy for edge deployment.
