Back to collection
Research

Training and live inference need different capacity contracts

Burst compute and continuous low-latency service should not be sold as one undifferentiated resource.

A large workstation and a small fanless computer rest apart on a workbench.
Technical illustration

Training consumes large bursts of compute for limited periods; inference serves continuing requests under tighter latency and cost limits. Treating them as one product can leave expensive accelerators idle between jobs or make a live application wait behind a long training run. A compute contract therefore needs duration, latency, location, availability and energy constraints that describe the job being bought. Separate pricing and measured provider performance help connect machine time to a maintained service. Agent launch counts and trading activity do not tell a buyer whether a request will finish within its latency budget.

References