NeuroSpark

Aug 02, 2026
NeuroSpark Santa Clara, CA
About NeuroSpark NeuroSpark builds and operates a high-performance AI inference platform that helps enterprises run large language models faster, cheaper, and at scale. Inference infrastructure is the foundation the entire AI application layer runs on — every AI product ultimately depends on how fast, how reliably, and how affordably models can serve their users. Our vision is to make that layer so efficient that compute is never the reason a good AI product fails. About the Role Serving inference at scale is a scheduling problem. Requests arrive with wildly different shapes and latency expectations, GPUs are heterogeneous and expensive, and the difference between a platform that's fast and one that's economical usually comes down to how well work gets placed. That system is what you'll own. You'll design and build the scheduling and routing layer of our platform: how requests get admitted, prioritized, batched, and placed across a heterogeneous multi-cloud...