Infrastructure · In production

M-Core Autoscaler

An orchestration layer that scales application and GPU capacity from minimal TOML configuration, with a control loop driven by machine learning rather than static utilisation thresholds.

SystemML-driven compute and GPU orchestration
StatusIn production
Engineering areaInfrastructure

An in-house alternative to Kubernetes that scales compute and GPU capacity from minimal TOML rather than sprawling YAML, with scaling decisions driven by machine learning instead of static thresholds.

Core capabilities

  • Scale origin servers up and down against live traffic fluctuation
  • Provision application compute and GPU capacity on demand
  • Drive scaling decisions with machine learning rather than static thresholds
  • Describe deployments in minimal TOML instead of large YAML manifests
  • Provision the GPU tier behind the Senon AI model family
  • Handle cold starts, warm restarts, and idle-resource reduction unattended
  • Coordinate capacity state with Senon Cloud

Why it exists

Kubernetes solves orchestration, but it solves it with a configuration surface large enough to become its own operational burden, with every deployment carrying a sprawl of YAML before anything runs. M-Core Autoscaler keeps the orchestration and drops the surface: deployments are described in minimal TOML, and the platform is built to handle the rest without an operator standing over it.

A control loop that learns

Scaling decisions are driven by machine learning rather than static utilisation thresholds. The loop observes health, capacity, utilisation, and workload demand, then adjusts application and GPU resources against actual traffic movement, scaling origin servers through fluctuation and handling cold starts, warm restarts, and idle-resource reduction on its own.

What it carries

M-Core Autoscaler provisions the GPU capacity behind the Senon AI model family, working alongside M-Core Engine and SGLang to keep model serving elastic and to return idle capacity rather than hold it. It also adjusts the origin capacity behind Senon Cloud, so compute and traffic respond together rather than independently when demand moves or a path degrades.

Next conversation

Interested in building something great?

Start a project