Skip to content
15 years of IT experience
CybexsoftConsultancy Services
← Back to case studies

AI infrastructure startup

Taming Kubernetes scale for GPU-heavy AI workloads

An AI startup faced unpredictable GPU demand and noisy clusters. We re-architected their Kubernetes platform for efficient scaling, better scheduling, and cost controls aligned to growth.

Industry
AI/ML
Team
ML infrastructure
Engagement
12-week engagement
Focus
Kubernetes, FinOps

The situation

What was wrong

Rapid model experimentation created bursty GPU usage and unpredictable costs. The platform team lacked clear guardrails for scaling and had limited visibility into workload efficiency.

  • Underutilized GPU nodes with frequent scale oscillations.
  • Limited cost attribution across research teams.
  • Reliability incidents during high-throughput training runs.
Engagement focus
KubernetesFinOpsObservability
Outcome

Stabilized reliability while lowering cloud spend by ~20%.

How we worked

The approach, step by step

Stabilise delivery early, then build the foundation that keeps it stable once we hand it back.

    1

    Cluster redesign

    Segmented workloads into dedicated node pools with tailored scaling policies and instance mixes.

    2

    Scheduling improvements

    Introduced GPU-aware scheduling, bin packing, and quotas to stabilize utilization.

    3

    Cost governance

    Built FinOps dashboards and budget alerts tied to team ownership.

    4

    Reliability guardrails

    Added SLOs, alerting, and runbooks for GPU-intensive workloads.

Handover

What we delivered

  • Multi-pool Kubernetes architecture
  • GPU scheduling policies and quotas
  • FinOps reporting with team-level attribution
  • SLOs and reliability playbooks

Stack

What it runs on

KubernetesKarpenterPrometheusGrafanaOpenCost

Afterwards

What changed

Measured after handover, once the client's own team was running the system without us.

20%
Cloud spend reduction
99.9%
Training platform uptime
2x
Faster scale-up

Planning something like this?

Tell us what you are trying to move, migrate or automate and we will reply within one business day with an honest read on the work — including the parts we think you should not do.

More engagements