AI infrastructure company Infinity on Monday announced $15 million in funding at a $100 million valuation from investors including Touring Capital, Principal VC, and researchers from companies such as OpenAI and Anthropic.
The startup builds software that makes it easier for AI chips to run AI models. One of the big reasons Nvidia has become a top player is not just its high-performance chips, but also its CUDA (Compute Unified Device Architecture) software, which allows its GPUs (originally designed to run graphics) to function as general-purpose processing CPUs. The largest AI development frameworks, PyTorch and TensorFlow, are built on CUDA. This allows developers to write apps in popular languages like Python, use these leading AI frameworks, and have apps run on Nvidia chips by default.
Most of these app-level startups will not have the resources or know-how to write their own kernels (the low-level software that powers the chip) and port their apps to other AI chips. That's why Infinity is trying to build CUDA replacement kernel software that will work on all types of chips, including SRAM, GPUs, phone chips, and systolic arrays. Infinity is part of a new wave of startups trying to chip away at Nvidia's market dominance, product by product.
Infinity is building a universal inference library that can run on all chips, allowing these chips to automate the replication of cutting-edge research results.
Infinity was launched last year by Jeremy Nixon, a former Google Brain researcher and founder of the hacker network community AGI House. Nixon told TechCrunch that he decided to start the company because he was obsessed with the idea of ”automated invention,” or the belief that “AI systems can actually become metatechnology.” He said he himself invented a machine learning algorithm called Omega. This essentially created a new machine learning algorithm and automatically evaluated it in a feedback loop.
This success led him to think about other cases where this approach might work, and he turned to hardware, believing that automated systems could also generate the low-level code, such as the kernel, needed to run the chip more efficiently.
Infinity's AI research agent Ignition aims to write the low-level code needed for AI inference on Nvidia's alternative chips. Test, debug, and measure how fast your code runs on your hardware, and automatically rewrite your code as needed to improve performance. The system is self-optimizing. This means that the system continuously learns and improves itself. It also adapts to different chip architectures, regardless of its unique design, Nixon said. The result is what Infinity claims is a CUDA-level software stack.
Customers include AI chip maker (and potential Nvidia challenger) D-Matrix, and Infinity is in talks with other major chip and cloud companies, Nixon said.
However, humans are always in the know and provide general instructions while agents do much of the tedious grunt work. In one case study, the startup found that agents worked much faster than a human alone, reducing processes that could take years or months to hours or days. Infinity does not charge upfront license fees. Instead, some of the performance gains and cost savings are taken by measuring token changes per second.
Infinity currently has 26 employees, including design, operations and engineering.
If you buy through links in our articles, we may earn a small commission. This does not affect editorial independence.

