Building reliable (and fast!) directory sync
“Who’s your User?”
— Master Control Program, TRON (1982)
If user identity is important to your application, you need a way to manage it. Usually this means user accounts, and you create them when someone signs up.
But allowing employees to sign up for whatever SaaS app strikes their fancy is a management nightmare. So organizations like ways to control which users exist (or do not exist) in your app.
That controls who can sign in, but how about what they can do?
For that you need roles, or groups, which, like the users, come from the organization's identity provider - Entra, Google, Okta, etc. Groups form the basis of Firezone's access model. They determine who can access what.
The process of getting users and groups into your app is called directory sync, and it's surprisingly tricky to do well. In this post we'll cover what directory sync is exactly, the leading standard for implementing it, and why we opted to forgo it entirely to build our own engine from scratch.
Directory sync: a gentle primer
An organization's directory is the list of who works there and what they're allowed to access. It lives in the identity provider and is made of three things:
- Users: the people in the organization. Each has a name, an email, and a status like active or suspended.
- Groups: named collections of users, like
EngineeringorProduct. Groups let you grant access to many people at once. - Members: who belongs to what. A member of a group is either a user or another group. A user's access comes from every group they belong to, directly or through other groups.
Directory sync is the process of copying that directory into your app and keeping it up to date. The identity provider is the source of truth, so what your app holds is a copy. Over time that copy has to pick up three kinds of changes: new users and groups, updates like a name change, and removals. When someone joins, leaves, or changes teams, ideally your app finds out quickly.
It's worth clarifying what directory sync is not: Single Sign-On (SSO). This may seem obvious but many folks conflate the two. Leading authentication standards like OpenID Connect say nothing about how to bring users into the application. In directory sync we're referring to precisely the mechanism of bringing users (and groups) in, but not how they're authenticated (for that, see this post).
A simple example
Say your organization has the following directory:
In your database it might look something like this:
To answer common access questions in your app like "is Bob a member of Product?", you'd have to look at Product's members, then the members of any groups in there, and so on until you find Bob.
To avoid that, we can flatten groups, giving each user a row for every group they belong to, directly or not:
Now the same question is a simple lookup.
Ok so that's the data model. Now we need to get it from the identity provider into our database, and keep it up to date.
SCIM
Of course, you don't have to invent all of this yourself. There's a standard for doing this directory sync thing. It's called SCIM and here's how it describes itself:
The System for Cross-domain Identity Management (SCIM) specification is an HTTP-based protocol that makes managing identities in multi-domain scenarios easier to support via a standardized service.
In practice this means you build a list of REST endpoints in your app, like:
GET /Users
GET /Users/{id}
POST /Users
PUT /Users/{id}
PATCH /Users/{id}
DELETE /Users/{id}
GET /Groups
GET /Groups/{id}
POST /Groups
PUT /Groups/{id}
PATCH /Groups/{id}
DELETE /Groups/{id}
And the identity provider calls them to keep the directory in sync.
The nice thing about SCIM is that the identity provider pushes changes to your app, so the lag time between when an update happens at the provider and when it lands in your app can be much lower than a pull-based approach that polls the provider on a schedule.
Sounds great on paper. But what are the problems here?
Well, for starters, your app has to be alive and ready to receive directory updates pretty much all the time. Going down or getting overloaded for even a few a seconds means you might miss a critical directory update. Triggering a full sync from your app is not possible, so you lose those updates until the identity provider decides to tell you about them again.
But the more annoying thing about SCIM is, while it does a great job at standardizing the protocol (i.e. the wire format), it says nothing about how those endpoints should be called, what lifecycle events they map to at the provider, or even what kind of data each request contains.
Identity providers all differ wildly in how they implement SCIM. Here are some examples:
- Group membership updates: Okta can send additions and removals together in one
PATCH, while Entra requires them to be split and only allows one member removal perPATCH. - Even basic types differ within the same provider: Entra historically sent
active: "False"as a string instead of the SCIM-defined boolean, and only later added a compatibility flag to switch to compliant behavior. - Optional features vary: Okta explicitly doesn't use several SCIM capabilities, including bulk operations,
POSTsearches,/ServiceProviderConfig, and filtering onmeta.lastModified.
And when it comes to user deprovisioning, the critical operation of cutting off a departing employee's access, even more inconsistencies come up:
- Okta: Order matters: unassign a user before removing their group memberships, and those memberships can remain in the downstream app.
- Entra: Group removal doesn't necessarily deactivate a user: they remain in scope if another assigned group still grants access.
- JumpCloud: Deprovisioning can follow app unbinding, suspension, or deletion – not just removal from a provisioning group.
- OneLogin: Deleting a user can trigger deletion, suspension, or nothing at all, depending on the app's provisioning settings.
All of this boils down to the reality that you end up with many different provider-specific SCIM code paths instead of the one consistent implementation you were hoping to write.
It's entirely possible you'll end up with a more complicated (and brittle) implementation going the SCIM route than if you built a pull-based system that hits each provider's API separately.
Which is exactly what we did.
Pull-based sync
With pull-based sync, you do the calling. Your servers hit the identity provider's API on a schedule (or at the behest of the user), walk the directory, and reconcile it with your database. Every provider's API is a little different, but the shape is always the same, so most of the work can be shared.
How it works
Identity providers have API endpoints for listing users, groups, and a group's members. After authenticating, a simple sync algorithm could be:
- Get all groups
- Get all the groups' members
- For user members, get the user records corresponding to those members
- For group members, go back to (2)
- Repeat every N minutes
A few details to keep in mind:
- Pagination. Each list call returns one page of results, so "get all" means following the next-page link until there are no more.
- Stable IDs. Every user and group has an ID that never changes. Match records on that instead of email or name, so a rename updates a record instead of creating a duplicate (see the Alice problem).
- Cycles. A group can end up inside itself (A in B in A). Remember which groups you've already visited so step 4 doesn't loop forever.
- Flattening. This is where the flattened table comes from. When you find Engineering inside Product, every member of Engineering also gets a row for Product. It also means a removal needs a recompute instead of a delete, since Bob might still reach Product another way.
Simple enough. For tiny directories like our example, you can get away with walking the entire directory, gathering all the data up front, and writing it to your database in one go.
Larger ones are a little trickier.
Syncing larger directories
Larger directories take more time to walk. And the longer the directory takes to walk, the greater the risk of losing the accumulated state before you manage to flush it to the database. Deploy at the wrong time and you have to start over, delaying the time until new directory data makes it into your database.
What you can do to alleviate this somewhat is to initialize each sync with an epoch, checkpointing the data as you go, then finally removing all records older than the epoch:
- Initialize a start timestamp.
- Fetch one page and write to database with timestamp.
- Repeat for remaining pages.
- At the end, delete all records older than the start timestamp.
Size isn't the only thing that makes this hard. A few other problems show up as directories grow:
- Rate limits. The provider's API limits might be shared with every other app the customer uses. Hit the API too hard and you can slow down their other tools. Go slower than you'd like, and back off when the provider tells you to.
- Bad responses. Sometimes an API returns a normal-looking
200with part of the directory missing. If your sync treats "missing" as "deleted", it can remove a lot of real users. Set a limit, or circuit breaker, on how much a single sync is allowed to delete. This can prevent nasty surprises. - Nested groups. Providers disagree on who does the flattening, and some return stale results when you ask them to.
- Overlapping jobs. If two syncs for the same directory run at once, they can overwrite each other's changes. The likelihood this happens grows with the size of the directory, since a previous sync might still be running when the next one kicks off. Make sure only one runs at a time (a job queue with uniqueness constraints or advisory locks works well), and that a crashed job doesn't block the next one forever.
- Temporary vs. permanent errors. A timeout or a
503should be retried. A revoked credential shouldn't, because retrying just wastes your quota and hides the real problem. Your sync should tell the two apart and let an admin know about the second kind.
Ask for less
The less you fetch, the faster the sync. Most directories have a lot in them that has nothing to do with your app, like contractors, service accounts, and hundreds of groups for mailing lists and office locations.
Where you can, let the provider do the filtering. Ask only for the users and groups assigned to your app instead of everything. That means fewer pages to read and fewer calls to make, which also eases the pressure on the rate limits.
You could stop here and walk away with a fairly robust directory sync engine. But if you use the directory for access rules (like we do in Firezone), you're at risk of lagging important updates between the sync schedules.
Consistency properties
This process is very much eventually consistent: eventually your database will reflect the organization's directory as it exists in the identity provider.
The above process is a full sync - walk the directory, grab every user, group, and member, then reconcile it with your database - insert what's missing and remove what's gone.
Depending on the provider's rate limits and the size of the directory, this sync could be as quick as a few seconds, or sometimes more than an hour. In that time, it's quite possible the directory has changed: a new employee onboarded, another left the company, someone else went on parental leave.
Lagging a new employee onboarding might leave the employee unable to sign in. Annoying, to be sure. But lagging an employee termination is risky - until you pull that directory change in, the employee still has access. Are they going to steal company secrets?
SCIM was supposed to help here, but we saw how that goes. What else can we do to reduce the lag?
What about delta syncs?
One way to reduce the lag is to stop walking the whole directory. Remember where you left off, and next time only ask the provider what changed. Some providers have a feature built for this (Entra calls it a delta query).
We looked at this closely and decided against it. The main reason is the groups we flattened earlier:
- No transitive groups. A delta tells you a group's direct members changed. It doesn't tell you what that means for the groups above it. If Bob is added to Engineering, you'd have to work out on your own that he's now also in every group that contains Engineering. Get that wrong and Bob ends up with access he shouldn't have, or without access he should.
There were other problems too:
- Deletes. With the epoch approach, removed records take care of themselves, because anything we didn't see this time is gone. A delta feed has to report removals explicitly, and providers do it differently. Miss one and a former employee keeps access.
- Expiring tokens. A delta only works from a saved token, and tokens can expire (Entra's last at most seven days) or get lost. When that happens you need a full sync anyway, so now you have two code paths to maintain instead of one.
- Mistakes stick around. A full sync fixes its own mistakes the next time it runs. A missed delta stays wrong until someone notices.
To us, delta syncs solve the wrong problem. The goal isn't to avoid the full sync. It's to make the full sync fast and reliable enough to trust, and solve the lag problem through another means.
The solution: a hybrid approach
Well, it turns out identity providers helpfully offer their own set of real-time APIs you can use to get pushed-based directory updates - no SCIM required! (Google, Entra, Okta, JumpCloud). As a bonus, these APIs often send you updates faster than they'd normally arrive over SCIM. Entra documents up to a 40-minute lag time for its SCIM updates, for example.
Firezone uses a hybrid approach that combines these APIs with an optimized full sync approach from above, giving customers the best of both worlds: periodic reconciliation of the entire directory to ensure updates aren't missed, and with real-time triggers that provide near-real-time updates for critical changes. Since the real-time triggers handle urgent changes, the full sync can run much less often as a safety net, saving API quotas as a result.
The last piece is making sure the two don't overwrite each other. A simple serialization queue in your job processing system handles that just fine. It also helps to treat each real-time notification as a signal to re-read the object from the provider, instead of blindly trusting what the notification says.
Conclusion
Directory sync has a lot of moving parts once you try it on a real organization. Large directories, flaky APIs, providers that disagree on what "deleted" means, and jobs that step on each other all need to be handled. SCIM leaves most of that to your app, once per provider, and takes away your ability to ask the provider what's true when something looks wrong.
Pulling from the provider's APIs keeps you in control of when to look, what to trust, and what to do when something seems off. Add the providers' real-time hooks on top and you get most of what SCIM promised, with far less baggage.
There are more edge cases we didn't cover here. For example, how do you sync two (or more) directories containing some of the same users into the same account? That happens when an organization moves to another identity provider, or when two companies merge. Firezone supports this, and the identity side of it is covered in the Alice problem.
Directory sync is available on our Enterprise tier today. If this is an important feature for your organization, book some time to get a closer look at how it works and try it out.