One agent
Discover, schedule, optimize,
and operate compute
Benefits
Benefits
Benefits
Benefits
Benefits
Benefits
Benefits
Benefits
Benefits
Multi-GPU node
MIG partition
Time-sliced
Shared pool
Premium & critical production
Reserved capacity
Capacity held by entitlement or reservation.
| Compute Agent | FinOps Agent | |
|---|---|---|
| Primary purpose | Run, schedule, and keep compute healthy and utilized | Optimize, allocate, and govern what compute costs |
| Core data | Nodes, GPUs, utilization, health, capacity, scheduling | Billing, usage, pricing, commitments, allocation |
| Typical action | Schedule a job, allocate a GPU, drain a node, self-heal | Rightsize, flag an anomaly, recommend a commitment, raise a cost ticket |
| Main users | Platform, infrastructure, SRE, HPC and ML platform teams | FinOps, finance, cloud, product and engineering leaders |
| Key outcome | Utilized, healthy, performant compute | Lower unit cost, accountable, forecastable spend |
FAQ
It’s a secure, topology-aware software component that discovers, monitors, allocates, schedules, optimizes, and governs compute across bare-metal, virtualized, containerized, and cloud environments — from CPU nodes to GPU clusters. It turns raw servers and accelerators into reliable, well-utilized, accountable capacity.
GPU is a dedicated capability domain. The agent discovers and monitors every card — utilization, memory, temperature, ECC and XID errors, NVLink and PCIe health, MIG slices — schedules GPU- and topology-aware, manages MIG, vGPU, and time-slicing, validates distributed-training readiness, and remediates failing GPUs automatically.
Several, because training, inference, notebooks, and VDI have different needs: a dedicated physical GPU, a dedicated multi-GPU node, hardware-isolated MIG partitions, vGPU for VMs, time-slicing for low-intensity work, a shared GPU pool, and reserved capacity for premium or critical workloads.
At scale, failures are routine — mean time to failure is measured in hours. The agent runs continuous health probes, detects node and GPU faults early, drains affected nodes, checkpoints and reschedules recoverable jobs onto healthy capacity, and opens an enriched incident — keeping useful work (‘goodput’) high.
It detects idle GPUs, abandoned notebooks, stranded MIG partitions, and runaway jobs; consolidates and scales to zero; recommends MIG, time-slicing, or a different GPU tier; and attributes GPU-hours and cost per token or inference — then hands that usage and cost data to the FinOps Agent.
Yes — one agent spans bare-metal (PXE, BIOS, RAID, imaging), virtual machines and hypervisors, containers and Kubernetes, and cloud instances, with multi-tenant quotas, isolation, policy, and reporting for private cloud, GPU cloud, and managed services.
Copyright © 2026 • All Rights Reserved