Compute Agent

Make every GPU,
count.

A secure, topology-aware agent that discovers, monitors, allocates, schedules, optimizes, and governs compute across bare-metal, virtualized, containerized, and cloud environments — from CPU nodes to GPU clusters.
Compute Agent
Runs your compute across every platform and scheduler
KubernetesSlurmKubeVirtBare-metal / PXE NVIDIA GPU OperatorDCGMTerraformAnsible PrometheusGrafanaAWSAzureGoogle Cloud
… and more via cloud APIs, hypervisors, and open telemetry.
GPUs are expensive.
Most sit idle or fail.
Accelerators are the costliest thing in the data center — yet utilization is low and, at scale, jobs fail constantly. The Compute Agent keeps them busy, healthy, and accounted for.
Waste
~ 0 %
Average GPU utilization
A 2026 study across 23,000 clusters put average enterprise GPU utilization at roughly 5% — expensive accelerators sitting idle or stranded. Utilization is the single biggest lever.
Reliability
~ 0 %
Of large training jobs fail before finishing
Large production clusters shed roughly 40% of jobs before completion — CUDA errors or NVLink/NCCL faults waste most of those GPU-hours. Health and self-healing are essential.
Scale
Hours
Between failures at scale
At thousand-GPU scale, mean time to a job’s first failure is measured in single-digit hours. Humans can’t respond fast enough — detection and remediation have to be automated.
Reach

One agent

CPU nodes to GPU clusters
A single topology-aware agent manages bare-metal, virtualized, containerized, and cloud compute — CPU hosts through multi-GPU AI nodes — under one control plane.
What it does

Discover, schedule, optimize,
and operate compute

Nine core capabilities of the Compute Agent — from discovery and telemetry to topology-aware scheduling,
autoscaling, self-healing, and governance — unified in CERNE. Hover any capability to see what it does.
Compute & Accelerator Discovery
Know every core and
card
Detect physical servers, VMs, cloud instances, and Kubernetes nodes (CPUs, GPUs, memory, NICs) and register them for managed workloads.
Servers / VMs / cloudKubernetes nodesHardware inventoryNode registration

Benefits

Real-Time Telemetry
See every resource,
live
Stream CPU, memory, GPU, disk, network, and power metrics, plus OS health, processes, and logs — real-time observability across the whole fleet.
Storage & networkPower & thermalOS healthLogs & eventsLive metrics

Benefits

opology-Aware Scheduling
Place jobs where they run best
Schedule workloads by NVLink, and RDMA locality, node labels, QoS, and priority, so latency-sensitive and GPU jobs land on the right hardware.
GPU-awareNUMA / NVLinkRDMA / InfiniBandLabels & taintsQoS & priorityAffinity

Benefits

Workload & Job Lifecycle
Run everything
reliably
Deploy and manage VMs, containers, pods, and batch or ML jobs — with queueing, checkpointing, retries, and rolling updates for long-running work.
Containers & VMsBatch / ML jobsRetriesQueue managementCheckpoint / resume

Benefits

Provisioning & Configuration
Bring compute up, keep it consistent
Automate bare-metal and VM bring-up, image and driver management, and configuration — BIOS/RAID, CUDA and runtime versions — at fleet scale.
Bare-metal / PXEVM lifecyclePatch & driversConfig managementCUDA / runtime

Benefits

Allocation, Quotas & Isolation
Fair sharing, hard
boundaries
Allocate CPU, memory, and GPU enforce limits via quotas and isolate tenants across namespaces, VMs, and GPU partitions.
Resource allocationGPU quotasQoS classesTenant isolationFair-share

Benefits

Autoscaling &
Rightsizing
Match capacity to
demand
Autoscale and scale-to-zero on demand, and continuously rightsize, detecting idle and stranded capacity, consolidating workloads, and using spot where it fits.
AutoscalingScale-to-zeroRightsizingIdle detectionConsolidationSpot / preemptible

Benefits

Health, Failure Detection & Self-Healing
Failures are routine — handle them
Run health probes, detect node, GPU, thermal, and resource failures, and self-heal, restart workloads, reschedule, and replace unhealthy nodes.
Health probesFailure detectionSelf-healingNode drainingAuto-rescheduleRecovery

Benefits

Security, Cost &
Governance
Compliant, accountable compute
Integrate IAM and RBAC, enforce policy, attribute cost, track power and carbon, and connect to ITSM and automation, with full auditability.
Security postureIAM / RBACEnergy / carbonPolicyCost attributionAudit & ITSM

Benefits

A dedicated GPU brain
Built for accelerators, not just servers
GPU is a first-class capability domain — because AI factories, GPU clouds, and HPC live or die on how well
their accelerators are scheduled, shared, and kept healthy.

01

GPU discovery & health
Detects GPU model, memory, driver, MIG, NVLink, ECC and XID errors, PCIe and thermal state — a live accelerator inventory.

02

Utilization &
memory
Tracks GPU and SM utilization, tensor-core activity, and memory per process, pod, VM, or tenant — finding idle and stranded capacity.

03

GPU-aware
scheduling
Places jobs by GPU model, free memory, count, NVLink/RDMA topology, health, quota, and priority — single- and multi-GPU.

04

MIG, vGPU & time-slicing
Manages hardware MIG partitions, virtual-GPU profiles, and time-sliced sharing — the right isolation for each workload.

05

Distributed-training readiness
Validates GPU count, memory, NCCL, networking, image, and CUDA before a run — and orchestrates checkpoints against interruptions.

06

Inference & model-aware placement
Tunes replicas, batching, and GPU tier by model size, precision, and throughput — holding response time down and cost per request low.

07

Utilization & cost optimization
Detects idle GPUs, abandoned notebooks, and runaway jobs, and computes cost per GPU-hour, token, and inference for showback.

08

Automated GPU remediation
On ECC/XID errors or overheating, drains the node, checkpoints and reschedules jobs to healthy GPUs, and opens an enriched ITSM incident.
GPU allocation modes
One size never fits every workload
Training, inference, notebooks, and VDI need different isolation and utilization — so
the agent supports the full range of allocation modes.
Large-model & high-perf inference
Dedicated GPU
An entire GPU to one workload or
tenant.
Distributed training, HPC

Multi-GPU node

Multiple GPUs with high-speed local interconnect.
Multi-tenant inference, fine-tuning

MIG partition

Hardware-isolated compute and memory slices.
VDI, virtual workstations
vGPU
GPU presented to VMs via profiles.
Dev, notebooks, experiments

Time-sliced

Shared GPU time; higher utilization.
Multi-tenant platforms, burst

Shared pool

Scheduler draws from a pooled resource set.

Premium & critical production

Reserved capacity

Capacity held by entitlement or reservation.

Five pillars
Discover. Observe. Schedule.
Optimize. Operate.
Every capability of the Compute Agent rolls up into five simple jobs — the full loop
from raw hardware to reliable, well-utilized capacity.

1

Discover
Inventory your infrastructure, every node and accelerator, and track total, allocatable and free capacity.

2

Observe
Real-time CPU, memory, GPU, storage, network, and power telemetry, alongside OS and hardware health, across the fleet.

3

Schedule
Place workloads topology- and GPU-aware, with quotas, QoS, priority, and autoscaling to match capacity to demand.

4

Optimize
Rightsize resources and detect idle or stranded capacity, consolidate, scale to zero, and manage both cost and sustainbility together.

5

Operate
Provision, patch, self-heal, isolate multi – tenants, enforce policy, and keep a full audit trail for governance, reliability, at scale.
Compute Agent vs. FinOps Agent
Two agents, two jobs —
one control plane
The Compute Agent runs the hardware and keeps it utilized and healthy; the FinOps Agent governs what
it costs. They hand off to each other — and both live in CERNE.
Compute Agent FinOps Agent
Primary purpose Run, schedule, and keep compute healthy and utilized Optimize, allocate, and govern what compute costs
Core data Nodes, GPUs, utilization, health, capacity, scheduling Billing, usage, pricing, commitments, allocation
Typical action Schedule a job, allocate a GPU, drain a node, self-heal Rightsize, flag an anomaly, recommend a commitment, raise a cost ticket
Main users Platform, infrastructure, SRE, HPC and ML platform teams FinOps, finance, cloud, product and engineering leaders
Key outcome Utilized, healthy, performant compute Lower unit cost, accountable, forecastable spend

FAQ

Questions teams ask us

What is a compute agent?

It’s a secure, topology-aware software component that discovers, monitors, allocates, schedules, optimizes, and governs compute across bare-metal, virtualized, containerized, and cloud environments — from CPU nodes to GPU clusters. It turns raw servers and accelerators into reliable, well-utilized, accountable capacity.

How does it handle GPUs specifically?

GPU is a dedicated capability domain. The agent discovers and monitors every card — utilization, memory, temperature, ECC and XID errors, NVLink and PCIe health, MIG slices — schedules GPU- and topology-aware, manages MIG, vGPU, and time-slicing, validates distributed-training readiness, and remediates failing GPUs automatically.

What GPU allocation modes does it support?

Several, because training, inference, notebooks, and VDI have different needs: a dedicated physical GPU, a dedicated multi-GPU node, hardware-isolated MIG partitions, vGPU for VMs, time-slicing for low-intensity work, a shared GPU pool, and reserved capacity for premium or critical workloads.

How does it keep workloads healthy at scale?

At scale, failures are routine — mean time to failure is measured in hours. The agent runs continuous health probes, detects node and GPU faults early, drains affected nodes, checkpoints and reschedules recoverable jobs onto healthy capacity, and opens an enriched incident — keeping useful work (‘goodput’) high.

How does it improve GPU utilization and cost?

It detects idle GPUs, abandoned notebooks, stranded MIG partitions, and runaway jobs; consolidates and scales to zero; recommends MIG, time-slicing, or a different GPU tier; and attributes GPU-hours and cost per token or inference — then hands that usage and cost data to the FinOps Agent.

Does it work across bare-metal, VMs, containers, and cloud?

Yes — one agent spans bare-metal (PXE, BIOS, RAID, imaging), virtual machines and hypervisors, containers and Kubernetes, and cloud instances, with multi-tenant quotas, isolation, policy, and reporting for private cloud, GPU cloud, and managed services.

See it on your fleet

Utilized, healthy, accountable
compute

Request a demo and see the Compute Agent discover your nodes and GPUs, place jobs topology-aware, reclaim idle capacity, self-heal failing hardware, and attribute every GPU-hour — across bare-metal, VMs, containers, and cloud.

Ready to get started? 

Talk to an expert.

Technical Support

Available 24/7 to assist you with your queries.

Playground

Experience UnityOne AI in action.

About UnityOne AI ™

UnityOne AI™ is an agentic intelligence platform for ITOps management, comprising CERNE™, LUMI™, and VEKTOR™. CERNE™ replaces dozens of cloud management tools by unifying DCIM, AIOps, HCMP, FinOps, and GreenOps within a single AI-driven control plane. LUMI™, the AI copilot, provides contextual intelligence, operational recommendations, and workflow automation, while VEKTOR™ enables enterprises to provision, orchestrate, and scale AI factories with the lowest cost-to-serve. The UnityOne AI™ suite enables enterprises to simplify hybrid/multicloud operations, strengthen governance, optimize resource utilization, and accelerate transformation to AI-driven ITOps.