Managed DevOps Pools: Scale Agents Without Burning Money

Managed DevOps Pools: Scale Agents Without Burning Money

Most teams that move to Managed DevOps Pools make the same two decisions in the first hour: pick a virtual machine (VM) size and click Create. Everything else stays at the defaults. Then, three weeks later, someone complains that pipelines wait forever for an agent, or someone from finance asks why the pool bill looks like a small server farm.

Both complaints usually trace back to the same place: the scaling configuration nobody looked at. Managed DevOps Pools (MDP) gives you three separate dials, and they influence each other. Agent state decides what happens after a job. Standby mode decides what is waiting before a job. Maximum agents decides how far any of it can go.

We run MDP in production for our own pipelines, and these dials are what let us steer how much load our builds and deployments put on our systems, while the monthly bill stays predictable. I will come back to how that looks in practice once the mechanics are clear.

Let’s go through the three dials one by one, with configuration you can paste into a template. The examples are written in Terraform’s HashiCorp Configuration Language (HCL), with raw Azure Resource Manager (ARM) JSON (JavaScript Object Notation) and Azure command-line interface (CLI) snippets where the underlying schema matters. Everything is based on the official scaling documentation and the pool settings reference. The exact versions are listed with the complete pool example further down.

What actually limits how many agents you get

The hard ceiling is maximumConcurrency, shown as Maximum agents in the portal. It is the maximum number of agents that can be provisioned at the same time in the pool, nothing more.

Two things are easy to confuse here:

  • Maximum agents is not parallelism. Your Azure DevOps organization has its own count of self-hosted parallel jobs, and that number decides how many jobs run concurrently. A pool with 20 agents and an organization with 4 self-hosted parallel jobs runs 4 jobs at a time. The other 16 agent slots are never used, and if your standby settings keep them alive, you pay for idle machines.
  • Maximum agents is bounded by quota. If your subscription cannot provide the cores for the VM stock keeping unit (SKU) multiplied by your maximum agents, pool creation fails with an error naming the cores required and your current limit. Not every SKU exists in every region either, so check regional availability before you standardize on one.
resource "azurerm_managed_devops_pool" "build" {
  name                = "build-pool"
  maximum_concurrency = 8
  # ...
}

I treat this number as a budget guardrail first and a capacity setting second. Size it for your realistic peak, request quota for it early (quota requests have their own lead time), and align it with the parallel jobs you actually bought.

Stateless or stateful: decide what happens after the job

The agent state is configured through agentProfile.kind in ARM, and through the mutually exclusive stateless_agent and stateful_agent blocks in Terraform.

Stateless is the default. A new agent is procured for every job and discarded afterwards. You get a clean machine each time, nothing leaks from one run into the next, and the documentation explicitly recommends stateless pools as a defense against supply chain attacks. For a pipeline that builds third-party dependencies and signs artifacts, this is the setting I would defend in any architecture review.

stateless_agent {}

Stateful lets agents serve multiple jobs. The typical reason is build performance: warm package caches, restored solutions, incremental build output. You configure two timers:

  • maxAgentLifetime (maximum_agent_lifetime in Terraform): the maximum time an agent lives before it is shut down and discarded. Format dd.hh:mm:ss, default and maximum is seven days (7.00:00:00).
  • gracePeriodTimeSpan (grace_period_time_span in Terraform): how long an idle agent waits for new jobs after all current and queued jobs finish. Same format, and the default is no grace period.
stateful_agent {
  maximum_agent_lifetime = "7.00:00:00"
  grace_period_time_span = "00:30:00"
}

There are a few behaviors here that surprise people:

  • Without a grace period, a stateful agent is not stateful for long. If no job is queued, no standby schedule is active and no grace period is configured, the agent is discarded right after the job, and its state with it.
  • maxAgentLifetime wins over everything. Even with a 24/7 standby schedule, an agent is recycled once it hits the lifetime. Configure three days and your “always on” agents restart every three days.
  • A running job is not killed by the lifetime timer, but individual jobs in Managed DevOps Pools can run for two days at most, regardless of the configured lifetime.

The documentation calls grace periods the most cost-effective way to run stateful pools for pipelines with consistent load, and notably they do not require standby mode at all. If your team commits steadily during the day, a 20 to 30 minute grace period keeps warm agents around between builds without a schedule to maintain.

The security trade-off is real, though. A stateful agent that ran a pull request validation from an untrusted branch is a different asset than a fresh VM. I keep stateful pools for trusted main-branch builds and leave everything else stateless. If you want to see how gates fit into that picture, my article on continuous deployment security gates covers the other half of the story.

Standby agents: the three modes

Without standby agents, every job waits for a machine. The documentation puts the on-demand wait at anywhere from a few moments up to about 15 minutes. With a standby agent in the ready state, startup, connection and job start typically take between 10 seconds and a minute, depending on SKU speed, image and networking.

A standby agent is capacity you pay for in advance, whether a job arrives or not. Every standby decision comes down to how much waiting that money should buy.

Standby mode lives in agentProfile.resourcePredictionsProfile in ARM. In Terraform you nest a manual_resource_prediction or automatic_resource_prediction block (at most one of them, none means off) inside the agent block. There are three states: off, manual and automatic. The standby count in any scheme cannot exceed Maximum agents.

Off: agents on demand

This is the default. You configure nothing and omit the resourcePredictionsProfile section. It is perfectly fine for nightly builds, rarely used pools and any workload where nobody stares at a pipeline waiting for it to start.

Manual: you know your usage pattern

Manual mode is for teams who can describe their workload with a calendar. You set a timeZone once for the whole pool and a daysData list. The list has either one item (all-week scheme) or seven items starting with Sunday. Each item maps times to standby counts, and a count holds until the next entry.

Here is a weekday scheme for a team in Germany: four standby agents from 07:00 to 18:00, Monday to Friday. In Terraform, the azurerm provider flattens daysData into one <weekday>_schedule block per entry (plus the integer attribute all_week_schedule for the all-week scheme) and timeZone into time_zone_name.

stateless_agent {
  manual_resource_prediction {
    time_zone_name = "W. Europe Standard Time"

    monday_schedule {
      time  = "07:00:00"
      count = 4
    }
    monday_schedule {
      time  = "18:00:00"
      count = 0
    }
    tuesday_schedule {
      time  = "07:00:00"
      count = 4
    }
    tuesday_schedule {
      time  = "18:00:00"
      count = 0
    }
    wednesday_schedule {
      time  = "07:00:00"
      count = 4
    }
    wednesday_schedule {
      time  = "18:00:00"
      count = 0
    }
    thursday_schedule {
      time  = "07:00:00"
      count = 4
    }
    thursday_schedule {
      time  = "18:00:00"
      count = 0
    }
    friday_schedule {
      time  = "07:00:00"
      count = 4
    }
    friday_schedule {
      time  = "18:00:00"
      count = 0
    }
  }
}

The same scheme as raw ARM template JSON, which is the schema the service actually consumes:

"agentProfile": {
    "kind": "Stateless",
    "resourcePredictionsProfile": {
        "kind": "Manual"
    },
    "resourcePredictions": {
        "timeZone": "W. Europe Standard Time",
        "daysData": [
            {},
            { "07:00:00": 4, "18:00:00": 0 },
            { "07:00:00": 4, "18:00:00": 0 },
            { "07:00:00": 4, "18:00:00": 0 },
            { "07:00:00": 4, "18:00:00": 0 },
            { "07:00:00": 4, "18:00:00": 0 },
            {}
        ]
    }
}

This is where most manual schedules go wrong: an empty daysData item does not mean zero agents. It means no change. Counts do not reset at midnight or at the end of the week. If you forget the "18:00:00": 0 entry on Friday, your four standby agents run through the weekend and into Monday. The schedule above is correct because every working day closes with an explicit zero.

The same mechanic lets you span midnight: put "09:00:00": 1 in one day and "17:00:00": 0 in the next. The documentation also shows a ramp such as one agent from midnight, ten from 09:00 and none after 17:00.

Two notes on the time zone. The documentation points to TimeZoneInfo.GetSystemTimeZones for valid names and uses Windows-style identifiers such as Eastern Standard Time in its examples, so verify your value against that list instead of guessing. And there is only one timeZone per pool.

Automatic: let the history decide

Automatic mode looks at the past three weeks of usage (if available), groups queued sessions of the pool into five-minute periods and assigns a percentile to each hour. You control the percentile with predictionPreference:

ValuePercentilePortal label
MostCostEffective10thMost cost effective
MoreCostEffective25thMore cost effective
Balanced (default)50thBalanced
MorePerformance75thMore performance
BestPerformance90thBest performance
stateless_agent {
  automatic_resource_prediction {
    prediction_preference = "Balanced"
  }
}

I like automatic mode for teams whose load is real but irregular, because it removes the schedule maintenance burden. The catch is that it is backward looking. A new project, a release crunch or a holiday week is not in the history, so the first Monday after a big change can still feel slow. It also works on percentiles, so BestPerformance does not mean “never wait”. It provisions for the 90th percentile of the demand observed in each hour, so spikes above that still wait.

Azure CLI: a different JSON shape

The Azure CLI takes the profile as a JSON file through --agent-profile on az mdp pool create and az mdp pool update. Note that the shape differs from the ARM template one: the kind becomes a property name.

az mdp pool update \
  --name build-pool \
  --resource-group rg-devops-pools \
  --agent-profile agent-profile.json

The following agent-profile.json contains the automatic profile:

{
    "Stateless": {},
    "resourcePredictionsProfile": {
        "Automatic": {
            "predictionPreference": "Balanced"
        }
    }
}

And this is the stateful variant from the documentation:

{
    "Stateful": {
        "maxAgentLifetime": "7.00:00:00",
        "gracePeriodTimeSpan": "00:30:00"
    }
}

If you copy a snippet from a template into a CLI file, or the other way around, the result is a validation error that does not point at the real cause. I have watched that cost people an afternoon.

Standby agents across multiple images

A pool can offer several images, for example Windows and Ubuntu. The standby capacity is then divided by the buffer property in the images section, which is a percentage of the standby agents for that image. You can use '*' to distribute equally, or integers that must total 100. The REST application programming interface (API) defines buffer as a string (default "*"), and the azurerm provider’s buffer is a string too ("*" or "0" to "100").

virtual_machine_scale_set_fabric {
  sku_name = "Standard_D4ads_v5"

  image {
    well_known_image_name = "ubuntu-22.04"
    buffer                = "70"
  }

  image {
    well_known_image_name = "windows-2022"
    buffer                = "30"
  }
}

With a total of ten standby agents this gives seven Ubuntu and three Windows machines. Two details from the documentation matter in practice:

  • Jobs without an ImageOverride demand target the first image. If your pipelines do not state the image, the three Windows standby agents wait for nobody, while every Ubuntu job beyond the seven ready agents falls back to on-demand provisioning.
  • ImageVersionOverride demands defeat standby agents. Standby agents are provisioned with the image versions from the pool configuration, so a pipeline that pins a different image version always starts a fresh agent and pays the on-demand wait. Forcing an older task agent (for example through the Agent.Version demand) is a smaller penalty, but still a delay: the agent has to be downloaded again after the machine is allocated.

A complete pool, assembled

For reference, this is a complete pool resource. It uses the azurerm_managed_devops_pool resource from the AzureRM provider, which exposes agent state, manual and automatic standby schedules and image buffers as first-class blocks, so there is no need for the azapi provider here. Compared with the ARM schema, the provider hides discriminators such as the organization kind and the Vmss fabric kind, and you attach the pool to a Dev Center project through dev_center_project_id. Check the provider documentation for the version you pin. Treat it as a starting point, not a drop-in: parameter values, project association and permissions depend on your environment.

On versions: the azurerm_managed_devops_pool resource is available since AzureRM provider 4.68.0, and all Terraform snippets in this article pass terraform validate on 4.68.0 and 5.7.0. OpenTofu resolves the same hashicorp/azurerm provider from its own registry, so the HCL carries over unchanged, although I only ran the validation with Terraform. The ARM snippets target API version 2026-06-02, the latest stable version of Microsoft.DevOpsInfrastructure/pools. The scaling documentation still shows 2025-09-20, but the agent profile, standby and image buffer properties used here are identical in both versions; 2026-06-02 adds features on top, such as mixing several VM sizes in one pool. The az mdp CLI extension calls its own pinned API version, but the --agent-profile shape shown above is the one from the documentation.

variable "dev_center_project_id" {
  type = string
}

variable "azure_devops_organization_url" {
  type = string
}

resource "azurerm_managed_devops_pool" "build" {
  name                  = "build-pool"
  resource_group_name   = "rg-devops-pools"
  location              = "westeurope"
  dev_center_project_id = var.dev_center_project_id
  maximum_concurrency   = 8

  azure_devops_organization {
    organization {
      url         = var.azure_devops_organization_url
      parallelism = 8
    }
  }

  stateless_agent {
    manual_resource_prediction {
      time_zone_name = "W. Europe Standard Time"

      monday_schedule {
        time  = "07:00:00"
        count = 4
      }
      monday_schedule {
        time  = "18:00:00"
        count = 0
      }
      tuesday_schedule {
        time  = "07:00:00"
        count = 4
      }
      tuesday_schedule {
        time  = "18:00:00"
        count = 0
      }
      wednesday_schedule {
        time  = "07:00:00"
        count = 4
      }
      wednesday_schedule {
        time  = "18:00:00"
        count = 0
      }
      thursday_schedule {
        time  = "07:00:00"
        count = 4
      }
      thursday_schedule {
        time  = "18:00:00"
        count = 0
      }
      friday_schedule {
        time  = "07:00:00"
        count = 4
      }
      friday_schedule {
        time  = "18:00:00"
        count = 0
      }
    }
  }

  virtual_machine_scale_set_fabric {
    sku_name = "Standard_D4ads_v5"

    image {
      well_known_image_name = "ubuntu-22.04"
      buffer                = "*"
    }
  }
}

Notice that parallelism in the organization block matches maximum_concurrency (maximumConcurrency in ARM). According to the documentation, this property is the number of agents an organization can use, and the sum across all organizations of a pool must match the maximum agents. It is a pool-internal split for shared pools, not your purchased parallel jobs. Those are a separate constraint, as described in the first section.

How we use it: cascading pipelines and fan-out deployments

Our own pipeline landscape is a good stress test for all of this. Two patterns dominate it. The first is cascading pipelines: a build completes, its pipeline completion trigger starts the next pipeline, which in turn kicks off further pipelines. The second is fan-out deployments, where a small change, sometimes a single line in a shared template or configuration, starts a large number of deployments at once.

Both patterns produce bursts, not steady load. On a static fleet of self-hosted agents, a burst like that means either long queues or a fleet sized for the worst case that sits idle most of the day. With MDP we handle it with two settings:

  • Maximum agents is our load valve. Every running agent is also a client of something: package feeds, artifact storage, and above all the target systems of a deployment. Capping maximumConcurrency caps how many of those clients exist at the same time. When a fan-out starts more deployments than the pool allows, the remaining jobs wait in the queue and run as agents free up, instead of hitting our environments all at once. The ceiling turns a deployment storm into a controlled wave.
  • Standby agents keep the cascade moving. In a chain of pipelines, every stage that waits for a fresh machine adds its provisioning time to the total. A small number of standby agents during working hours removes most of that delay, because the next stage usually finds a ready agent.

The effect on cost is the part I value most. The ceiling is a hard upper bound on what the pool can spend at any moment, and standby capacity is a number we chose deliberately rather than something that grows with demand. A burst makes the queue longer, not the bill. That combination gives these pipelines a lot of throughput while the costs stay consistently under control.

Recommendations by scenario

There is no universally correct setup. This is what I would start with and why.

ScenarioAgent stateStandby modeReasoning
Small team, a handful of builds per dayStatelessOff, or manual with 1 agent during working hoursWaiting a few minutes now and then costs less than idle machines. Add one standby agent only when the wait annoys people.
Single-site team with clear peak hoursStateless for pull requests, stateful with grace period for trusted buildsManual weekday scheme, explicit zero at closing timeThe schedule is predictable, so use it. Size the count for the typical burst, not the worst day.
Steady load, warm caches matterStateful, 20 to 30 minute grace period, sensible maxAgentLifetimeOffPer the documentation the cheapest stateful option, and no schedule to maintain.
Irregular but real loadStatelessAutomatic, start at BalancedLet history do the work, then move one step in either direction after a few weeks.
Global team across time zonesStatelessAutomatic, or one manual scheme in Coordinated Universal Time (UTC)A manual schedule has a single time zone per pool. Either accept a broad window that covers all regions or split into regional pools.
Nightly or scheduled buildsStatelessOffNobody is waiting. A few minutes of provisioning time at 02:00 is free.
Cascading pipelines or fan-out deploymentsStatelessManual, a few agents during working hoursUse Maximum agents as the throttle that protects target systems, and standby agents to cut the provisioning delay between stages.

For the global team, I would lean on automatic mode first. If your regions have very different peaks, separate pools per region also give you cleaner cost attribution, at the price of splitting your parallel job allocation.

Pitfalls I would check first

When a pool “scales badly”, this is my checklist, in order:

  1. Are there actually ready agents when the job is queued? Jobs fall back to on-demand provisioning when more jobs are queued than standby agents exist, when they are queued outside the schedule, and when the standby count is empty.
  2. Does the pipeline target the right image? Without ImageOverride, jobs go to the first image.
  3. Does the pipeline pin an image or agent version? A pinned image version bypasses standby agents, and a forced agent version adds a download.
  4. Is the Friday zero missing in a manual schedule?
  5. Is maximumConcurrency higher than your organization’s self-hosted parallel jobs, or lower than the standby count you configured?
  6. Are proxies, firewalls or virtual network settings slowing down the agent connection?

Before you add more standby agents, check whether utilization justifies the ones you have. The cost guidance in the documentation says the same thing: in manual mode, look at the historical utilization of standby agents and reduce the count if it is not used. Details are in Manage cost and performance.

Conclusion

Scaling in Managed DevOps Pools is a latency purchase, not a capacity purchase. Capacity is bounded by quota and parallelism. What you buy with standby agents is the absence of waiting.

My practical recommendation: start stateless with standby off, measure how long jobs actually wait, and add the cheapest mechanism that fixes the wait. For steady load that is a grace period on a stateful pool for trusted builds. For predictable peaks it is a manual weekday schedule with explicit zeros. For irregular load it is automatic mode at Balanced. And whatever you choose, put it in Terraform, review it like code and revisit it after a month of real data.

Comments