1.5M requests / day
$1,097
/ month simulated profit
- Provider net
- $115.01/day
- GPU cost
- $78.96/day
- Daily profit
- $36.05
Provider economics
Start with requests per day and target GPU utilization, then compare GLM‑5.2, MiniMax‑M3, and the existing Qwen3‑Coder deployment against Lambda list prices. The retained Omnious book remains visible as a reality check, not the default scenario.
1 · Choose the market
Book prices are Jul 31 mainnet asks and remain editable. Demand is the complete retained 366-hour paid sample.
2 · Set your assumptions
Traffic + utilization simulator
Start with the traffic you expect, then test how hard the GPUs need to run.
Traffic sets revenue. Online hours set GPU cost. Utilization tests whether the selected stack can carry that traffic; changing it never invents additional demand.
Lambda on-demand · $3.29/h · measured preset
Optional throughput override
Leave blank to see the throughput your traffic requires, or enter a real benchmark to map traffic onto measured capacity.
Conservative result from the live endpoint at 128 synchronized requests and the observed 38.7:1 input/output mix.
3 · Scenario output
1.5M requests / day
$1,097
/ month simulated profit
Measured capacity used
50%
50% target utilization
Traffic-to-capacity check
This traffic uses 50% of the entered benchmark while the stack is online. The selected 50% target corresponds to about 1.5M requests per day at this request size.
Break-even traffic
1M/day
At this request size and quote
Infra / month
$2,402
24 online hours each day
Net / request
$0.000077
After the router fee
Historical mainnet reference
Secondary comparison; it does not change your traffic simulation.
Method and limits
Traffic. Revenue is requests per day × average tokens per request × your input/output quotes. The default request shape comes from paid mainnet usage; traffic itself is editable and hypothetical.
Capacity. Qwen maps simulated traffic onto the conservative result from three live synchronized runs. Without a benchmark, the calculator reports the full-load throughput required to serve that traffic at the selected utilization; it does not claim the hardware can achieve it.
Cost. Daily GPU cost is hourly run-rate × online hours. Lambda prices are pre-tax list prices; storage, egress, orchestration, and operator labor remain excluded unless added to the editable run-rate.
Book comparison. The 366 retained hourly mainnet buckets from 16–31 July 2026 include finalized paid usage only. They remain a secondary reference and do not silently set the simulated traffic level.
The stack works on paper. Prove it under load next.
Deploy the runtime, benchmark the real request mix, then set a firm capacity and floor in the Provider Node flow.