Blog › Gainsight, Vitally, Catalyst

Gainsight, Vitally, Catalyst: What They Track and What They Miss

Priya Mehta · · 7 min read
A comparison view of customer success platform dashboards with health score indicators
Back to Blog

Customer success platforms have matured considerably over the past several years. Gainsight, Vitally, and Catalyst each represent a real and considered approach to helping CS teams manage renewal risk. This piece is not a takedown of any of them. It is an attempt to be precise about what their design philosophy prioritizes and where the resulting gaps live, because understanding those gaps is the only way to think clearly about what else a CS team needs.

We have built Sturdy specifically around one of those gaps, so we have a point of view here. We will be direct about it.

What these platforms are built for

All three of these tools operate from the same foundational assumption: the most tractable signal for customer health is product engagement. This makes sense. Product usage data is quantitative, consistent, and relatively easy to pipe in from your analytics layer. Login frequency, feature adoption, workflow completion rates, seat utilization: these metrics aggregate cleanly, they do not require interpretation the way communication data does, and they can be scored in ways that alert CSMs at scale.

Gainsight has built a comprehensive CS operations platform around this model, with health scoring, lifecycle management, playbooks, and CTAs that can fire based on usage thresholds. It is the most enterprise-oriented of the three, with deep configurability and a long history of integrations. Vitally positions itself as a more modern take on the same category, with a cleaner interface, tighter Salesforce integration, and a design that emphasizes the individual CSM's workflow rather than CS ops-level configuration. Catalyst is in a similar position to Vitally in terms of positioning but tends to emphasize its revenue-motion integration more prominently, sitting at the intersection of CS and account management.

All three also support NPS and CSAT collection, which gives them a structured customer voice signal alongside the usage data. Some customers respond to NPS surveys. Some do not. The ones who are about to churn often do not, which creates a selection bias in NPS data that is well known in the CS community but rarely discussed in the context of how health scores are built.

The manual layer

Each of these platforms also allows CSMs to enter manual health overrides, sentiment notes, and risk flags. This is where practitioner judgment enters the system. A CSM who knows an account is at risk because of a conversation they had on a call can flag it. The problem is that this requires the CSM to already know, which is circular when we are talking about early warning signals. Manual input is valuable as a supplement to automated signals. It is not a detection mechanism.

The manual layer also creates consistency problems across CSM books of business. One CSM might flag a red account that another CSM would have left yellow, because the standard for what constitutes "at risk" is partly a judgment call and partly a function of how much time the CSM had to review the account that week. CS ops teams spend real time trying to normalize these flags. The effort is not wasted, but it is a symptom of the limitation.

What is not in the data model

None of these platforms, as they are publicly positioned and as CS teams typically use them, read your support queue systematically as a health signal source. Zendesk, Intercom, and Help Scout can send data to these tools, typically in the form of ticket counts or CSAT scores from closed tickets. What they do not do is read the text of the tickets: what specific words customers are using, whether the language in tickets has shifted from operational to investigatory, whether a customer who used to ask "how do I" questions is now asking "what does it cost to" questions.

This distinction matters because ticket count is a lagging indicator. By the time ticket volume is elevated enough to trip a health score threshold, the customer has been expressing frustration in the text of those tickets for weeks. The signal was there. It was just stored in a field that no one was reading programmatically.

The same gap exists for email threads. Support email between customers and your team carries communication-pattern information that is separate from ticket counts. Thread participation, response latency, whether a champion who used to write effusively has shifted to clipped functional sentences: these patterns are visible to anyone who goes back and reads the thread, but they are not captured in any health score field by any of the major CS platforms.

The NPS problem in more detail

NPS is worth examining specifically because it is the most common structured voice-of-customer signal in these health scores. The mechanics of B2B NPS create a few specific failure modes that are underappreciated.

First, survey timing is typically set by the CS tool, not by the customer's sentiment journey. Most CS platforms send NPS surveys at fixed intervals: 90 days after onboarding, then annually or semi-annually. This means the survey can land during a period of relative satisfaction even if the customer is heading toward churn. The score reflects that moment, not the trajectory.

Second, B2B NPS response rates are typically lower than in B2C contexts, and the customers who respond tend to be those who are actively engaged. Disengaged customers, who represent a meaningful slice of churn risk, tend not to respond. Health scores built on NPS data are therefore partially self-selected for engagement, which creates exactly the blind spot you do not want.

Third, NPS in B2B is often answered by a champion or a primary contact, not by the broader user population at the account. If the champion is satisfied but the team using the product is frustrated, the NPS score reflects the champion's view. The frustration in the team shows up in tickets.

The usage data assumption

We want to be precise about where we think product usage data is genuinely predictive and where it is not. Usage data is a strong signal for one kind of churn: the disengagement churn, where customers simply stop using the product because it has not become a habit or because the use case was narrower than expected. If customers are not logging in, that is real information.

Usage data is less reliable as a signal for the dissatisfaction churn: cases where customers are actively using the product but are quietly frustrated with specific aspects of it, are finding workarounds for things they expected to work differently, or are being driven by a sponsor who is still engaged while the broader team has checked out. These accounts look healthy on a usage dashboard until the renewal conversation reveals otherwise.

In our experience building Sturdy, the accounts that present this pattern are the ones CS teams find most surprising at renewal. They did not look at-risk. The signals were there, but they were in the language of support tickets and the tone of email threads, not in the login frequency.

What this means for how you think about your CS stack

The right framing is not "Gainsight does not detect churn" but rather "Gainsight detects one kind of churn signal well, and the stack has a gap where communication-layer signals would complete the picture." CS teams that recognize this tend to ask one of two questions: can we get the CS platform to read our communication data, or do we need a different tool for that layer?

The answer to the first question is mostly no, in practice. These platforms are designed around structured data: numbers, scores, timestamps, CRM fields. Reading unstructured communication data at the level required to produce reliable signals requires a different kind of data pipeline and a different kind of NLP inference layer. Zapier integrations and webhook feeds can surface individual ticket metadata, but they do not close the gap.

The gap is real and it is consequential for renewal outcomes. We built Sturdy to fill it, which is why we are direct about where it sits relative to the CS platform category. We are not trying to replace health scores or playbooks or lifecycle management. We are trying to read the communication data that sits underneath all of those tools and surface the patterns that do not appear in any of their health score fields.

More from the team

Why QBR Data Is Always Looking Backward

Why QBR Data Is Always Looking Backward

Quarterly business reviews are built around the last 90 days. But the decisions being made right now are in the tickets your customers are opening today.

Priya Mehta