A CFO’s Guide to Tracking Digital Labor and Agentic AI

tracking digital labor

Your ROSE Metric may be improving for the wrong reason.

For years, the ROSE Metric has been a useful way to measure how much recurring revenue a SaaS company generates for every dollar invested in employees and contractors. It is one of the cleanest ways to connect labor investment to recurring revenue.

But AI is distorting the picture. Those $1M ARR per FTE claims? Somewhat true but a little skewed.

A company can now add labor output without adding the next SDR, support rep, or software engineer. Instead, it can spend on AI agents, usage-based inference, orchestration tools, internal AI productivity tools, and workflow automation.

So if work moves off payroll and into AI spend, what happens to your ROSE performance?

It goes up.

And unless finance adjusts the framework, it may go up for reasons the metric was never designed to capture.

ROSE Classic = Recurring Revenue / (Employee Spend + Contractor Spend)

That metric tells you how much recurring revenue the business is generating for every dollar of human labor investment.

That still matters.

If ROSE Classic improves, that may mean your team is getting more productive. It may mean your systems are better. It means your team is allocating resources more effectively.

It may also mean your humans are being leveraged by AI. Or replaced by AI.

But that last point is exactly why the metric can become misleading if used alone.

Download my guides at the bottom of this post.

ARR per FTE is Becoming Easier to Game

This is where a lot of operators are going to fool themselves.

If a company is explicitly trying to build toward something like $10M ARR per FTE, it can keep headcount lean while pushing more work into digital systems.

That company may look incredible on ARR per FTE. It may also look incredible in the ROSE Classic.

But finance should ask a harder question:

Did we truly become more efficient, or did we just replace wage expense with digital labor expense?

That is the real issue.

If AI is doing meaningful recurring work that would otherwise require people, then some of the improvement in ROSE is not coming from better human productivity alone. It is coming from a shift in the labor model.

There’s no free ride here. You can’t hide internal AI spend in your “internal use software” GL account.

If your company wants minimum headcount and maximum ARR per FTE, finance still has to account for the non-wage labor layer making that possible.

That is why I think SaaS finance teams now need a second version of ROSE.

The Next Version of the ROSE Metric

Keep the original metric. Then add a second one that reflects economic reality.

ROSE Classic

Recurring Revenue / (Employee Spend + Contractor Spend)

This measures human leverage.

ROSE + Digital Labor

Recurring Revenue / (Employee Spend + Contractor Spend + Agentic AI Spend)

This measures blended labor efficiency.

That second metric is the one that keeps you honest. Here is a simple example.

Let’s say a company has:

  • Recurring revenue = $10M
  • Employee + contractor spend = $4M
  • Agentic AI spend = $1.5M

Then:

ROSE Classic = $10M / $4M = $2.50

ROSE + Digital Labor = $10M / $5.5M = $1.82

That is a big difference.

If you only look at ROSE Classic, you might conclude the company has become dramatically more efficient.

If you also look at ROSE + Digital Labor, you get a more accurate answer:

Humans may be more leveraged, but part of that leverage was purchased through digital labor.

digital labor mix

Not All AI Spend Belongs in ROSE

This is where finance teams can make a mess of the analysis.

Not every AI dollar is labor-like.

If you throw all token spend, all AI software, and all experimentation into the ROSE denominator, the metric becomes noisy and much less useful.

The better approach is to classify AI spend based on its economic role.

A Practical Framework for AI Spend

1. Customer-facing or product inference

This should be managed as COGS in Dev Ops.

If AI is powering a product feature used by customers, that cost is tied to delivering the service. It usually scales with product usage, not with internal workforce planning.

This does not belong in the ROSE Metric.

2. Internal Copilot or Productivity AI

This should be managed as departmental OpEx.

Think vibe coding, AI copilots, note-taking tools, ChatGPT, and general productivity software used by our internal departments.

These tools may improve performance, but in most cases they behave more like software or compute spend than direct labor replacement.

I would generally keep this out of the ROSE Metric.

3. Explicit labor-substitution or labor-avoidance AI

This is the category that matters for the next version of ROSE.

If the company has deployed AI that is clearly replacing, delaying, or avoiding hires, I would track that spend in a distinct sub-account such as:

  • Agentic AI
  • Digital Labor
  • AI Labor Substitution

And I would place it inside the benefiting function.

Examples might include:

  • AI SDRs reducing the need for incremental sales development hires
  • AI support agents absorbing level 1 ticket volume
  • AI finance agents handling recurring close, reconciliation, or reporting tasks
  • AI coding agents reducing the need for planned engineering hires

This is the investment that belongs in your blended labor efficiency framework.

4. Experimentation, evals, tuning, and prototypes

This should typically be managed as R&D or innovation spend.

This is learning and development work. It is not recurring labor substitution.

Do not force it into the ROSE.

The one question finance should ask

Do not ask:

Is AI labor?

Ask:

Is this AI spend filling a box on the org chart that otherwise would be filled by a human?

If the answer is yes, it belongs in your digital labor view.

If the answer is no, it probably belongs somewhere else on the P&L.

That single question will keep the ROSE Metric clean and useful.

ai spend matrix

The Framework Only Works if You Fix the GL Structure

Here is the mistake I think many SaaS finance teams will make:

They will talk about digital labor, but they will not change the chart of accounts.

And if AI spend is buried across cloud hosting, software subscriptions, DevOps, sales tools, and miscellaneous OpEx accounts, you will lose spend visibility fast.

If you want to measure ROSE + Digital Labor and Digital Labor Mix correctly, the accounting structure must support it.

1. Give product inference its own COGS account

If AI is part of the product experience, I would not bury that spend inside generic AWS, Azure, or hosting accounts.

I would create a separate inference account inside COGS / DevOps so finance can clearly see what portion of service delivery cost is AI-driven.

If possible, I would also track it by major vendor or model family so you can monitor usage trends and gross margin impact over time. And tie inference costs back to the product using it!

2. Create a separate Agentic AI or Digital Labor account

If the spend is explicitly replacing, delaying, or avoiding hires, it should not live inside a generic internal use software account.

As mentioned above, create a dedicated GL account such as:

  • Agentic AI
  • Digital Labor
  • AI Labor Substitution

This is the spend that should feed the  ROSE + Digital Labor and Digital Labor Mix. More about digital labor mix below.

If you do not isolate this spend, you will not be able to tell whether ROSE improved because of true efficiency gains or because labor simply moved off payroll.

3. Code the spend to the benefiting department

Do not just track AI at the total company level. Track it where it creates the work output.

Examples:

  • product inference → COGS / DevOps
  • AI SDR tools → Sales
  • AI support agents → Support
  • AI finance agents → G&A / Finance
  • AI coding agents → R&D

That is what makes the analysis useful.

It lets you compare digital labor investment against function-level productivity, budget ownership, and hiring plans.

If you are using specific AI models (Anthropic, Gemini, OpenAI, etc.), you may consider creating an API key for each internal department that draws on those tokens.

4. Keep ambiguous AI spend out of digital labor until it is clearly labor substitution

This is where judgment matters.

A broad copilot, ad hoc prompting, or a set of internal AI skills may improve productivity. But unless there is a clear case that the spend is replacing or avoiding labor, I would generally keep it in productivity OpEx rather than digital labor.

Otherwise, you risk overstating agentic AI adoption. AI ROI is hard enough to measure.

That means finance needs a classification rule:

Productivity first. Digital labor only when the labor-substitution case is explicit.

5. Revisit the coding every quarter

The classification will evolve.

Some tools that start as broad productivity software may eventually become clear labor-substitution tools. Some experiments may move into production. Some product inference may get blended with internal usage.

That is why I would review these accounts every month during close and ask:

  • Is this serving customers?
  • Is this productivity software?
  • Is this explicit labor substitution?
  • Is this experimentation or R&D?

If you want reliable AI metrics, this cannot be a one-time setup exercise.

It has to become part of the finance operating rhythm. AI changes almost every week.

gl structure for ai spend

One Supporting Metric I Would Add: Digital Labor Mix

Alongside ROSE Classic and ROSE + Digital Labor, I would track one supporting ratio:

Digital Labor Mix = Agentic AI Spend / (Employee Spend + Contractor Spend + Agentic AI Spend)

This answers a simple question:

What share of my total labor-like spend is digital labor?

Let’s say a company has:

  • Employee spend = $3.5M
  • Contractor spend = $500K
  • Agentic AI spend = $1M

That gives you:

$1M / $5M = 20%

digital labor mix

How to interpret it

In this case, 20% of the company’s blended labor stack is now digital.

That means one-fifth of what the company is effectively spending on labor is no longer wage-based. It is AI-based. You may also want to track the wage rates of the replaced labor or delayed labor. This would provide a clear ROI. Unless your digital labor is not working!

That does not automatically mean the company is more efficient or less efficient.

It means the labor model has changed.

And that is exactly the context finance needs when looking at ROSE, ARR per FTE, or headcount productivity metrics.

If traditional labor metrics are improving while Digital Labor Mix is also rising, that may mean the company is becoming more efficient.

It may also mean some of the improvement is coming from moving work off payroll and into digital labor.

Why I like Digital Labor Mix

I like this ratio because it is simple, clean, and hard to misread.

It is useful because it is:

  • bounded between 0% and 100%
  • useful for benchmarking over time
  • better at showing labor cost mix than a raw AI spend number

A raw AI spend number on its own is not very helpful.

But if you say:

“15% of our blended labor stack is now digital”

or

“Digital Labor Mix increased from 4% to 18% over the last four quarters”

that tells the business something meaningful. It shows how quickly the labor model is shifting.

How to use digital labor mix

I would not use Digital Labor Mix by itself.

I would use it alongside:

  • ROSE Classic, to measure human leverage
  • ROSE + Digital Labor, to measure blended labor efficiency

Together, those three views answer three different questions:

  • Are our people becoming more productive?
  • Is the company actually becoming more efficient after including digital labor?
  • How much of our labor model has shifted from people to AI?

That is a much better operating dashboard than ARR per FTE alone.

What I Would Tell SaaS Finance Teams to Do Now

Here are the action items to put in place now.

1. Create a separate digital labor sub-account

Stop letting labor-substitution AI spend hide inside random software, cloud, or functional budgets.

Create a clear sub-account for Agentic AI or Digital Labor and assign ownership by function.

2. Keep ROSE Classic, but report ROSE + Digital Labor alongside it

Do not replace the original metric. Keep it as your measure of human leverage.

But add the second version so leadership can see the difference between human productivity gains and broader labor model shifts.

3. Start tracking Digital Labor Mix every month or quarter

Even if the number is small today, track it now.

Because once it starts moving, it will change how you interpret productivity, hiring plans, and efficiency metrics across the company.

The Bottom Line

The original ROSE Metric still matters.

But in the age of agentic AI, it is no longer enough.

If your company is increasingly relying on AI to replace, delay, or avoid hires, then labor has not disappeared. It has changed form.

Some of it is still human.

Some of it is now digital.

That is why the next version of ROSE should measure both.

Keep ROSE Classic to measure human leverage.

Add ROSE + Digital Labor to measure blended labor efficiency.

Track Digital Labor Mix to understand how much of the labor stack is no longer human.

And improve your GL structure so those metrics mean something.

Because if you do not isolate digital labor, your labor efficiency metrics will eventually tell you a story that is only half true.

Download the AI Spend Guides