Tag me in

In an effort to improve AI categorization, I experimented with restructuring the tag taxonomy of an inbox. Surprisingly, I increased the model’s tagging accuracy from 69% (nice) to 90%. Let’s talk about why that matters, how I did it, what it means, and what you could do next.

Categorization is a foundation… and it’s crap

Ticket tagging is a perennial topic among Support teams. Partly because tags are important. They’re used for reporting and decision making, but they’re also used to trigger automated workflows. Get the tags wrong and you’ll miss important data and screw up your automated resolution rates.

They’re also a common point of discussion because so many Support teams are dealing with a fucking mess. Hundreds of tags, a bunch of duplicates and misspellings, some inherited from a previous setup that doesn’t make sense after the last re-org, a couple that are specific to an outage that happened two years ago but somehow got repurposed for a new feature. I couldn’t tell you how many tagging schemes I’ve seen that made no sense to the team using them.

We all know this. We talk about it all the time. And we have annual (sometimes quarterly!) sessions to audit, correct, and improve our inbox’s tags, which promptly become a new mess in a couple of weeks.

Now that AI can purportedly handle triage, it feels like you should be able to step back from that mess and not worry about it. Instead, teams are finding that the tangled foundation is hurting their ability to use the cool new toys they’re reading about online. The automated categorization and reporting and the bots meant to resolve tickets aren’t working for them. Some of that (a lot, but definitely not all of it) is because of the rotten foundation it’s all sitting on. The new bots can’t understand how you categorize your tickets because you don’t really understand how you categorize tickets.

Define some definitions

Let’s explain taxonomy real quick before we go on. It’s just a fancy way of saying ‘tagging system’. You might have some of the data in other field types, like Custom Fields in Zendesk, but it’s all roughly the same for the purposes of this conversation. Generally speaking, your taxonomy is the way you describe the structure of all that data.

For a two-tier taxonomy, you could have a top-level set of categories, like “Billing” and “Database” and “Emails”. Each of those could have sub-categories to describe specific problems or questions that customers write in about. “Billing” could have “invoices,” “chargebacks,” etc. “Database” might have “replication,” “backups,” and others.

A lot of teams use tags to categorize their tickets, or a mix of tags and custom dropdown fields. Over the years, I’ve seen a lot of teams that would have tags like “Billing chargebacks” or “Database.backups.feature_request” or some other format that made sense to them. Or didn’t make sense to them, as it often turned out, but at least it made sense to the person who came up with it in the first place.

My hypothesis going into this was that we can improve the accuracy of LLM categorization and tagging by changing the way we structure the taxonomy. We’d keep the same tags, more or less, but change how they relate to each other and how they’re applied to tickets. Maybe instead of database.backups.feature_request we could have separate fields for the Issue type (bug, feature_request, docs) and for the Product Area (database, billing, emails).

Less talking, more doing

My basic process for this experiment was this:

  • Build a sample set of tickets with Qbort. This way I could have some control over the types of tickets I’m working with.
  • Import the tickets into Qval. I wrote a set of rules to follow and Qval drafted a schema for the LLM to use when triaging.
  • Manually triage the tickets. I went through all the tickets and selected tags, urgency, and a sentiment.
  • Have the LLM triage the tickets and compare the results. Run the process in Qval and see how the LLM did compared to the well-trained human.

The model I chose for this was Jev, which I guess technically isn’t an LLM. It’s a classifier, which means it doesn’t output freeform text. It’s more restricted in how it reads and writes data and it has fewer ways to steer it compared to a harness using Opus or Astra or some other LLM. But I ran a few tests with Jev alongside those LLMs and they performed similarly.

Back to my testing… I generated 250 tickets with Qbort. The imaginary product being supported is a to-do app called “Q-DO” and the customers are largely non-developer consumers. Qval takes a RULES.md file, which defines the system (product, customers, and taxonomy) and converts it into an input prompt and an output schema for LLMs, as well as a set of questions, descriptions, and instructions for a classifier model, like Jev.

You can see the rules file, which includes the taxonomy, for that run here. The set of tags was rewritten as a flat list of all possible combinations of the two-tier taxonomy defined in the rules. sync_bug and task_list_feature_request, for example.

The results of the first run looked like this…

Metric Value
Tickets 250
LLM evaluator Jev (jev-latest)
Human evaluator brianlevine
Labels (human vs. LLM) 69% (Jaccard 0.69, n = 250)
Urgency (human vs. LLM) 82% agreement (n = 250)
Sentiment (human vs. LLM) 81% agreement (n = 250)

Notes:

  • Tags were called “labels” in Qval, but they’re interchangeable terms here.
  • This data uses the Jaccard index as a measure of accuracy for groups of tags.

You can see that in addition to tags, I set an Urgency level (low/medium/high) and a Sentiment (unhappy/neutral/happy). The thing that immediately stood out was that the fields where the model only had to pick one option out of three had much higher accuracy than the selection of tags. It seems like tagging was difficult for the model.

Except tagging 250 tickets is not easy for a human, either. It’s been a while since I worked a large queue, so it took a few minutes to remember how difficult it is to be consistent with ticket triage. I wrote the rules myself, so it should not be hard for me to stick to them diligently. But let me tell you something. That shit is hard.

Looking at my tags the next day, I didn’t agree with all of them. Some fell in a grey area and I chose one side on the first day and the second side on the next day. That’s not great. So I ran through a review and came back again on the next day to see if I agreed then, which (thankfully) I did. That gave us our baseline.

To improve that score, I restructured the taxonomy. For the first run, I used a taxonomy like this (edited down here for brevity):

Tier one:

  • sync
  • account
  • task lists
  • archive/search/export
  • app platform

Tier two:

  • bug
  • documentation
  • feature request
  • how-to

in which any tier one category can be paired with a tier two category. So if someone writes in about a bug in the task list and a feature request in the iOS app, you can add the task_list_bug and app_platform_feature_request tags.

In the second run, I restructured it to look like this (edited down here for brevity):

Product area:

  • sync
  • account
  • task list
  • archive

Platform:

  • ios
  • android

Issue:

  • bug
  • feature request
  • how-to/docs

and included the rule that the classifier (human or otherwise) can only pick one from each category per ticket. This does mean we lose a little granularity compared to the first run. We wouldn’t put the feature_request tag on the ticket that also includes bug. But that’s a price I’m willing to pay in this exercise. We can talk about it more in a minute.

The other thing you might notice is that platform got moved from a “tier one” category to its own category type. One of the issues I found when tagging and reviewing the first round’s results was that the classifiers had a hard time deciding how to describe a ticket if the platform was specified. If a customer complained about a bug in task lists on Android, what do you do? app_platform_bug and task_list_bug are both correct, but that data is duplicative (double bugs) and often miscategorized. Both the human (me!) and the model (Jev) would miss one that the other added, because we were often not fully confident that the platform was critical to the conversation. For reporting, we want to know how many issues came up for each platform, so we want to track that for each ticket where possible. With all that in mind, I moved platform to a category type. Similarly, I combined how-to and documentation in the second run since those were mistakenly tagged a lot of the time. It’s often unclear when one should be used over the other, so maybe they weren’t actually different. As long as we’re restructuring these tags, we may as well fix that one.

So now we have a different type of taxonomy. This wouldn’t really be described as a two-tier taxonomy, strictly speaking. It’s more a group of three categories with mutually exclusive options. It’s more limiting than the flat set of tags used in the first run. The point here is to make it easier for the classifier to choose the most correct options, so adding some constraints seemed like a good direction.

And it was a good direction. Here are the results from run 2:

Metric Value
Tickets 250
LLM evaluator Jev (jev-latest)
Human evaluator brianlevine
Issue (human vs. LLM) 96% agreement (n = 250)
Product-Area (human vs. LLM) 88% agreement (n = 250)
Platform (human vs. LLM) 93% agreement (n = 250)
Urgency (human vs. LLM) 92% agreement (n = 250)
Sentiment (human vs. LLM) 92% agreement (n = 250)

We should look at the similarity of tags as a group, though, since the first run looked at all of them as a group. Here’s a side-by-side comparison of the results, including a “labels” calculation. I’m comparing them two ways:

  • Grouping the second run’s categories as if they were flat tags, so product area: sharing + issue: bug would be the equivalent of sharing_bug from the first run.
  • Not grouping the second run’s categories and treating them as separate entities.

You’ll notice that the Jaccard index changes a little between those two groupings, but they’re both noticeably better than the first run.

Metric Run 1 Run 2
Label structure One multi-select (area × issue type, 31 options) Three single-choice questions (Issue, Product-Area, Platform)
Labels, mean Jaccard (run-1 method) 0.69 0.85
Labels, mean Jaccard (three separate labels) — 0.90
Issue agreement — 96%
Product-Area agreement — 88%
Platform agreement — 93%
Urgency agreement 82% 92%
Sentiment agreement 81% 92%

Jev got up to 90% accuracy using the new taxonomy structure. It’s simpler and the choices are more clearly defined with less potential overlap per ticket. As the human classifier, it was also much easier for me to decide which options to choose. It turns out that when you make a set easier for a computer to classify, you also make it easier for humans. Amazing.

By restructuring the way we tag our tickets, we were able to increase the model’s accuracy by 15 to 20 percentage points. Very little taxonomy work was done otherwise. We used nearly the same tags and the same definitions, but applied them in different ways.

Now let’s have some cake. We did something great.

Some things to consider before you finish that cake

I’m including this section because I know some people will nitpick the data. This is the internet after all. So here are some caveats to consider:

  • Changes to the taxonomy created an imperfect comparison. We did a little more than restructure when we decided to disallow multiple tag sets in run 2. And the change to how the platform is tracked might skew the data some. Little changes could have bigger effects. I don’t think it’s a problem here. The data is still useful for reporting purposes and for routing/automation purposes. Some data will get missed, inevitably, but with an accuracy improvement of 20 points, I think we more than make up for it. To account for this, I’ve shown the results of the second run’s grouping in two ways.
  • This is a fake data set. This might not be what YOUR data looks like. The improvements you can get out of your data and your taxonomy might differ.
  • Human accuracy is not 100%. Even a perfectly trained domain expert will choose differently or incorrectly sometimes. The upper limit on model accuracy is below 100% because the human tags we measure against aren’t perfect.

Well done, class

Support teams overwhelmingly have a convoluted and outdated tag taxonomy. When that taxonomy is then used to generate reports and trigger AI workflows, the results are often a disaster of inconclusive reports and broken automations. Restructuring that taxonomy can make it easier for humans and LLMs to use, which improves your ability to understand what your customers are doing and makes it possible to achieve those incredible AI resolution rates you’ve been hearing about.

You should try this with a larger set of data, preferably real (anonymized) ticket data. I’d love to hear what you find and how you’ve been able to improve your classifier models. All the tools are freely available on GitHub.


Have a comment or a response to this post? Take it to Bluesky or LinkedIn.