Expect more from your chips.

There is usually more performance left in the machine, especially when you specialize to exact models, workloads, and hardware. We research methods for AI-driven performance engineering and build systems that automatically realize those gains across the accelerator stack, from application and runtime decisions down to kernels at the assembly level.

The accelerator software stack: application, runtime, operators, kernels and PTX/SASS, all feeding one optimization system. APPLICATION LLM SERVING SIMULATION ROBOTICS ... RUNTIME SCHEDULER MEMORY MGMT OP SELECTION OPERATORS ATTENTION MATMUL CONV ... KERNELS TILING MEMORY MOVEMENT SCHEDULING PTX / SASS WGMMA TMA CP.ASYNC ... WGMMA.MMA ... LDGSTS ... CP.ASYNC.BULK ... BAR.SYNC ... LDG.E ... STG.E ... DEPBAR.LE ...

Products

Inference by the token, or on a dedicated endpoint.

  • Shared

    Per token, no commitment

    Billed on what you use. Better economics and performance enabled by our infrastructure.

    Maison Token Platform
  • Dedicated

    An endpoint tailored to your workload

    Isolated capacity, with the model, routing and silicon chosen for your traffic shape.

    Contact us
  • OpenAI-compatible endpoint.
  • Zero data retention.

I — Multi-silicon

The right accelerator is a tradeoff between performance, cost, and availability, and that tradeoff shifts with the workload. Our automated software generation lets us bring up new architectures quickly and move across vendors, rather than being tied to a single chip, price curve, or pool of supply.

II — Application-tailored Deployment

The right configuration follows the traffic. For a coding agent, heavy context reuse makes caching central to cost. For a real-time voice application, latency shapes batching and scheduling. We tune the system around the constraints that actually matter for each application..

Research

We research the foundations of AI-driven performance engineering and build systems to advance it.

We study how AI can understand, generate, measure, and optimize high-performance software, then turn that research into systems and tools. The work spans profiling, low-level program understanding, kernel generation, verification, benchmarking, and post-training.

  • SKX

    Our kernel optimization engine. Given a target kernel, SKX searches over candidate implementations, checks them for correctness, and profiles them against the hardware.

  • PTX Understanding

    Program analysis and LLMs work together to transform and optimize PTX. By operating and reasoning at the shared PTX layer across DSLs (e.g. Triton, TileLang, ThunderKittens, CUTLASS), our system learns and combines optimizations to generate kernels that outperform all individual DSLs.

    Preview Blog Post
  • SK-Trace

    A custom profiler for finding where optimization matters and modeling workload dependencies. Tracer maps execution across a workload, identifies the kernels and regions dominating performance, and estimates the value of improving them.

  • SKL

    A statically checked language for GPU programming, designed for both engineers and models. It gives us a more structured target for generating low-level, architecture-specific code.

  • SK-Check

    Correctness infrastructure for generated GPU kernels. Beyond reference-output checks, SK-Check catches illegal stream usage, output reuse, race conditions, and common reward hacks that can make an incorrect kernel appear valid to a naive checker.

  • GPU Timing

    High-fidelity measurement under controlled clocks, power, and thermal state. We study the measurement problem itself because optimization is only as reliable as the benchmark beneath it.

    Read the write-up
  • Post-Training

    Fine-tuning and reinforcement learning for low-level performance reasoning.

Company

We work at the frontier of AI and computer systems.

Background of team members

  • From AI and systems groups at Stanford, MIT, CMU, NVIDIA, Meta, and other leading institutions; built and and optimized GPU kernels, compilers, cryptographic systems, and high-performance distributed infrastructure
  • Developed KernelBench, a widely used benchmark for LLM-generated GPU kernels that helped define the emerging field of AI-driven kernel optimization.
  • Published research at leading systems and machine learning venues including PLDI, ICLR, ICML, COLM, and MICRO.
  • Top competitors in kernel programming, cybersecurity (CTFs), and algorithmic programming.

Backed by

  • Jump Capital
  • General Catalyst
  • Felicis
  • Cowboy Ventures
  • Link Ventures
  • Essence VC
  • CoreWeave Ventures
  • Ericsson Ventures

Additional investors include

David M. Siegel · Jeff Dean · Jonathan Frankle · Michael Carbin · Sachin Katti · Walden Yan

We are hiring

Contact

Get in touch with us.

Maison Token is currently in private beta. Contact us to request access.

For a dedicated deployment, tell us about the model, traffic, latency requirements, and capacity you need.

For licensing or a custom engagement, send us the workload, hardware, and what you are trying to improve. We’ll work with you to understand the constraints and define the right scope.