File Nº 0x6D / Subject: m0rgxn
Local --:--:--
← All notesExhibit K · Filed 18 Aug 2026 · Azure Security · 15 min read

Modeling Managed-Identity Privilege Escalation as an Attack Graph

Filed

A graph model for when a foothold in Azure is worth escalating


Exhibit K — cover
Scope and honesty note. This is a framework and systematization piece, not a measurement study. I have not run a controlled experiment across many tenants, and there are no benchmark numbers here that I did not compute by hand from a topology I describe. Everything operational was done in isolated lab tenants I built for research during a red-team internship, never against production or third-party systems. Where I lean on other people's work I cite it. Where the model has gaps I say so.
I have written before about why the Azure managed-identity attack surface does not shrink even though the design removes stored secrets, and about the tool I built to walk that surface, Fenrir. Both of those were engineering write-ups. This one steps back and asks the question underneath them: when I hold some level of access to a resource group, is managed-identity abuse worth attempting, and how far does it reach if it is? That question is annoying to answer by hand because the answer is not local. Whether a compromised user can escalate depends on a chain: which control-plane roles the user holds, which resources in the group carry an identity, which of those identities are user-assigned and therefore shared with other hosts, and what data-plane roles those identities were granted somewhere else in the tenant. Each link is easy to check in isolation. The chain is what takes time, and it is the chain that decides whether an engagement continues or stops. The natural representation for a chain of preconditions is a graph. That is not a new idea in security, and the AD side of the house has lived on it for years through BloodHound. What I found missing was a small, honest model of the Azure managed-identity path specifically: one that says yes only when it can point to the exact steps, and that treats the shared-identity blast radius as a structural fact rather than something an analyst has to notice. This post is that model, and an account of how Fenrir implements a deliberately narrow version of it. Attack graphs are old. Phillips and Swiler proposed a graph-based system for network-vulnerability analysis in 1998 (NSPW '98), where nodes were attacker states and edges were exploits with pre- and post-conditions. Sheyner and colleagues made the generation and analysis automatic in "Automated Generation and Analysis of Attack Graphs" (IEEE S&P, 2002). Ou, Govindavajhala, and Appel then reframed the whole thing as logic programming in MulVAL (USENIX Security, 2005), where reachability falls out of Datalog rules over host configuration and network facts. The shared shape across all three is the one I use below: an attacker holds a state, edges have preconditions, and escalation is the transitive closure of edges whose preconditions the current state satisfies. On the offensive tooling side, the identity graph is well trodden for Active Directory and, more recently, for Azure. AzureHound collects Entra ID and Azure Resource Manager objects and feeds them into BloodHound, which then answers path queries over roles, group membership, and ownership. MicroBurst, by Karl Fosaaen at NetSPI, is the PowerShell toolkit most people reach for when enumerating and abusing Azure services, managed identities included. Andy Robbins at SpecterOps documented the managed-identity escalation paths themselves in detail, starting with "Managed Identity Attack Paths, Part 1: Automation Accounts". For the hybrid side, where on-prem AD and Entra ID meet, Dirk-jan Mollema's research is the reference I keep going back to. So why write another model. Because AzureHound and BloodHound are built for completeness: collect everything, show every path, let the analyst decide what is real. That is the right call for a discovery tool, and I use it. It is the wrong call for the specific decision I care about, which is a go or no-go on one attack technique from one starting point. A completeness-first graph shows me paths that look green but fall apart when I try to walk them, because they depend on inherited assignments or on roles that grant nothing actionable at the scope I actually have. I wanted the opposite bias: a model that under-reports, so that when it says READY, the path is one I can execute. The rest of this post is what that bias looks like when you write it down. Take the graph G = (V, E). Vertices come in four kinds.
  • Principals: users, groups, service principals, and the attacker's current identity. A principal is something that can hold role assignments.
  • Resources: split into compute resources (a VM, an App Service, a Logic App workflow, a Container Group, an Automation Account) and data resources (Storage, Key Vault, a container registry). The split matters because only compute resources give you an execution surface.
  • Managed identities: system-assigned or user-assigned, per Microsoft's own distinction. I model these as their own nodes rather than as attributes of a resource, and the reason is the whole point of the exercise, which I get to below.
  • Directory boundaries: the on-prem AD forest and the Entra ID tenant, connected by whatever sync method is configured. These are context for now, not fully traversable in the current model.
Role assignments are not nodes. They are the preconditions attached to edges. This is a modeling choice with consequences, and I made it because an assignment is only interesting in terms of what edge it unlocks. A Contributor assignment at a resource group is worth nothing on its own; it is worth something because it lets you run code on the hosts in that group. An attacker's state is the set of principals and identities whose tokens you currently control. You start with one: the foothold. Escalation means growing that set. You grow it by traversing an edge whose precondition is satisfied by the role assignments held by anything currently in your state. Reachability is then a fixpoint. Start from the foothold, repeatedly add any node reachable by a satisfiable edge, and stop when a full pass adds nothing new. Because the state only ever grows, and the graph is finite, this terminates. The reachable set is every identity you could end up holding a token for, and the edges you used are the proof. There are five edge types. The first four are traversable steps. The fifth is context. Each one has a precondition expressed as a set of RBAC roles, and those role sets are exactly the ones Fenrir checks, not a cleaned-up version for the diagram.
The five edge types, each with its RBAC precondition
control-plane-execute: from a principal to a compute resource. The precondition is any role that lets you run code on that host: Owner, Contributor, Virtual Machine Contributor, Website Contributor, Logic App Contributor, Automation Contributor and Automation Operator, and the Azure Container Instances Contributor role, among the others Fenrir tracks. The concrete mechanism differs by host type, which is the part that makes manual assessment tedious: a VM exposes execution through RunCommand, an App Service through Kudu, an Automation Account through runbook artifacts, a Container Group through its exec endpoint. imds-token: from a compute resource to a managed identity. The precondition is code execution on that host, which is to say an inbound control-plane-execute edge you have already traversed. Once you are running on the box, you request a token from the Instance Metadata Service at 169.254.169.254. No secret is involved, because the identity never had one. This is the edge that turns access to a machine into access as an identity. identity-attach: from a principal to a user-assigned identity. The precondition is the pair Managed Identity Operator plus a role that lets you write to a host. If you can assign an existing user-assigned identity to a resource you control, you can then pull that identity's token from IMDS on your own terms. If there is no host in the group to attach it to, the model records that you would have to create one first, which is a real cost and worth surfacing rather than hiding. self-role-assignment: from a principal to a data resource. The precondition is roleAssignments/write, which you hold if you have Owner, User Access Administrator, or Role Based Access Control Administrator. This edge skips the identity entirely: instead of stealing a token, you grant yourself a data-plane role like Storage Blob Data Contributor or Key Vault Secrets User and read the data directly. I include it because from a reachability standpoint it lands you in the same place, and a model that only looked at IMDS would miss the quieter path. sync-boundary: from one directory to the other. In a hybrid tenant, on-prem AD synchronizes to Entra ID through Password Hash Synchronization, Pass-through Authentication, or Federation, per Microsoft Entra Connect. I draw this edge because the hybrid trust is where a lot of real escalation actually happens, but I am honest that my current model treats it as context rather than a step it will traverse. More on that in the limitations. The single design decision I would defend hardest is modeling a managed identity as its own node with its own in-edges, rather than as a property of the resource it sits on. It costs nothing for a system-assigned identity, which has exactly one host. It pays off for user-assigned identities, which are standalone objects that can be attached to many resources at once. When a user-assigned identity is a node, the fact that two different hosts both attach it shows up as two in-edges landing on the same vertex. The shared blast radius becomes a property you can read off the graph: compromise either host, and you hold the identity, and therefore you hold every role that identity was granted anywhere in the tenant. In a per-resource view, where each host carries its own copy of "an identity," that sharing is invisible until an analyst happens to cross-reference client IDs by hand. The graph makes it structural. That is most of why I bothered with a graph at all instead of a checklist. Fenrir does not build the full graph and hand you a query language. It builds a deliberately narrow slice and answers one question. I think that restraint is what makes it useful in front of a client, so I want to be precise about where it sits relative to the model above.
From tenant collection to a go/no-go verdict
Collection runs against two APIs. From Microsoft Graph, Fenrir pulls the principal's identity, directory roles, PIM-eligible roles, group memberships, owned apps and service principals, and managed devices. From Azure Resource Manager it walks subscriptions down through resource groups to resources, resolves the principal's direct role assignments at each scope, and flags every resource carrying an identity. The ARM walk runs on a small thread pool because listing resources is the slow part, and role-definition IDs get resolved to readable names through a cache. The graph it builds is scoped hard. It only includes role assignments that sit directly at the resource-group scope, with no inheritance expansion. That is the deliberate under-reporting I described earlier, encoded as a rule. An assignment inherited from the subscription might well be exploitable, but the exploit phase acts at RG scope, so including a subscription-level assignment as an RG edge would produce a green light the tool cannot honor. I would rather the model be quiet about a real path than confident about a fake one. Subscription-scope assignments are surfaced as a note, not an edge. The query is the reachability fixpoint, projected down to a coarse verdict. Rather than return a set of reachable identities, Fenrir returns one of four states, because the four states map cleanly onto what you do next.
  • READY — a satisfiable control-plane-execute edge at RG scope reaches at least one identity-bearing resource. Next step: run the exploit phase.
  • BLOCKED_NO_TARGETS — control-plane-execute edges exist, but no compute resource in reach carries an identity. Next step: wait for one to appear, or use the identity-attach edge.
  • BLOCKED_NO_RIGHTS — no satisfiable control-plane-execute edge at RG scope. Next step: go after non-MI data access instead.
  • UNVERIFIED — the resource layer was not enumerated, because it ran with --no-resources. Next step: re-run without that flag.
Beyond the verdict, Fenrir reports escalation openings by intersecting the roles you hold against the preconditions for the identity-attach and self-role-assignment edges. That logic lives in one small function so the same rule feeds the panel you read, the JSON state file, and the tests. Expressing the edge preconditions as sets of role names rather than a tangle of conditionals is what let me keep the model and the code in sync; adding an edge precondition is close to a one-line change. The exploit phase then walks only the paths the READY verdict named. For each identity-bearing host it tries the appropriate execution mechanism, requests the token from IMDS, and exchanges it against ARM to see what that identity actually reaches. For any standalone user-assigned identity it finds, it looks for a host with that identity attached and pulls a token for that specific client ID, which is the identity-attach edge and the shared-identity structure made real. Every extraction it reports corresponds to a path the graph justified, and when a host fails, because RunCommand was disabled by policy or a plan does not support Kudu, it says so instead of swallowing it. Here is a small topology from one of my lab tenants, reduced to the nodes that matter. It is contrived to be minimal, but every edge is one Fenrir evaluates for real.
A traced escalation path through the lab topology
The foothold is a synced user, j.doe@corp, whose password I am assuming is already compromised. In the resource group rg-app, the user holds Contributor. The group contains two compute resources: vm-jump01, which carries both a system-assigned identity and a user-assigned identity named uai-deploy, and an App Service that also has uai-deploy attached. Elsewhere in the tenant, the system identity on vm-jump01 holds Key Vault Secrets User on a vault, and uai-deploy holds Storage Blob Data Contributor on a storage account. Walk the reachability from j.doe.
  1. control-plane-execute to vm-jump01. The Contributor assignment at rg-app satisfies the precondition, so RunCommand is available. The state now includes code execution on the VM.
  2. imds-token from vm-jump01 to its system identity, and to uai-deploy. Execution on the host is the precondition, and it is met, so both tokens come down from 169.254.169.254. The state now holds two identities.
  3. The system identity's Key Vault Secrets User role is a data-plane reach to the vault's secrets. The graph draws this as an ordinary out-edge to a data resource, no further escalation required.
  4. Here is the part the graph earns its keep on. uai-deploy is a single node with in-edges from both vm-jump01 and the App Service. Holding its token from the VM means holding it everywhere it is attached, so the App Service is now within reach without ever attacking it directly, and the identity's Storage Blob Data Contributor role hands you the blob containers.
The verdict is READY, and the reachable set is {system-MI, uai-deploy}, which projects to Key Vault secrets plus Storage blobs, plus lateral presence on a second host you never touched. A manual assessment gets to the same place, but it has to notice on its own that the App Service and the VM share a client ID. In the graph that sharing is one vertex with two parents, so it is not something to notice, it is the shape of the data. What the model does not tell you here is also worth stating. It says nothing about whether j.doe's inherited assignments at the subscription would have offered a shorter path, because those are excluded by the RG-scope rule. It does not model whether the sync back to on-prem AD gives the storage data any further value. And the READY verdict is conditional on RunCommand not being disabled by policy on vm-jump01, which is a fact the graph assumes and the exploit phase verifies. I would rather list these than have someone find them. The RG-scope rule trades recall for precision on purpose. Real escalations do run through subscription-scope and management-group-scope assignments, and this model will stay silent on all of them. It is the correct tradeoff for a go/no-go tool and the wrong one for a completeness audit, which is why AzureHound and this model are complementary rather than competing. PIM-eligible roles are collected but not treated as held, so a path that requires activating an eligible role is a path the reachability will miss. That is a real gap, not a stylistic choice, and closing it means modeling activation as its own edge with its own cost. Custom roles are the soft underbelly. The precondition sets are named built-in roles. A custom role that grants the same dataActions under a different name will not match, so the model can under-report against a tenant that rolls its own RBAC. Doing this properly means parsing Actions and dataActions rather than matching role names, which is on the list below. The sync-boundary edge is drawn but not traversed. The most interesting hybrid escalations, the ones in Dirk-jan Mollema's research, cross that boundary, and my model currently treats it as a labeled context edge rather than something reachability will walk. Directory-plane escalations more generally, through Graph app-role grants or consent, are outside the model entirely right now. Everything is a static snapshot of a single tenant. There is no notion of an edge's OPSEC cost or detectability, so the reachability treats a noisy RunCommand and a quiet token read as equivalent steps, which they are not on a real engagement. The gaps above are more or less a roadmap. Parsing dataActions instead of matching role names would make the precondition checks sound against custom roles and is the change I would do first, because it removes a whole class of silent under-reporting. Modeling PIM activation and the sync boundary as real edges with costs would extend reachability into the two areas where it currently stops early. And weighting edges by detectability would turn the flat reachability into something closer to a shortest-plausible-path query, which is the version of this a red teamer actually wants: not every path, but the quietest one that works. The honest evaluation I owe this, and have not done, is a precision and recall measurement against a labeled corpus of lab tenants: build tenants with known escalation paths, run the model, and count what it finds and what it misses. That is the difference between the framework here and a claim that it works, and it is the piece I would need before calling any of this a result rather than a design. The takeaways survive the inversion cleanly, and they are the same ones I ended the Fenrir write-up on. A managed identity is only as safe as the control-plane roles that can run code on its hosts, so Virtual Machine Contributor and its relatives are a path to every identity in the group, not just the box. A user-assigned identity is only as trustworthy as the least-trusted host it has ever been attached to, and the graph is the clearest way I know to see which hosts those are. And roleAssignments/write at resource-group scope is a quieter escalation than it looks, because it never touches an identity at all. If you want to know your own blast radius, the reachability query is the same one; you just run it as the owner.
Related write-ups: Azure Managed Identities: An Attacker's View of a Credential-less Design and Building Fenrir. Corrections and pushback are welcome; the model is a work in progress and the limitations section is the part most likely to grow.
§ — Also in evidenceView all →

Other exhibits


Exhibit L · 04 Sept 2026

SQL Injection

PortSwigger Web Security Academy notes on detecting and exploiting SQL injection, from UNION attacks to blind and out-of-band techniques

Filed
§ Contents