A Better Way to Check Regression Models That Predict Counts and Categories
Analysts who model yes-or-no outcomes, ratings, or counts have lacked good tools to check whether the model fits. A new framework replaces the usual single-number residual with a function, and yields plots that reveal missing variables, interactions, and other misspecifications.

Every analyst who fits a regression model faces the same follow-up question: does the model actually fit? For ordinary linear regression the answer is easy to check. You look at residuals, the gaps between predicted and actual values, and plot them. Many models predict discrete outcomes: whether a customer churns, how a product is rated, or how many claims are filed. For those models, that check has never worked well. A paper in the Journal of the American Statistical Association proposes a fix.
Zewei Lin, Assistant Professor of Information Systems and Analytics at McCoy College of Business, is a co-author. His co-authors are Dungang Liu of the Lindner College of Business at the University of Cincinnati and Heping Zhang of Yale University.
The problem with discrete outcomes
Models for discrete outcomes belong to a family called generalized linear models, or GLMs. Logistic regression for binary outcomes and Poisson regression for counts are common members. The standard residuals for these models, called Pearson and deviance residuals, compress each observation into a single number. When the outcome can only take a few values, that number carries little information. Plots of these residuals often look the same whether the model is right or wrong.
The study
The researchers take a different approach. Instead of summarizing each observation’s residual as one number, they keep it as a function. The idea is that a function can hold the randomness the model’s structure does not explain, even when the data are discrete. The paper works out the mathematical properties of this functional residual.
From those properties, the authors build two diagnostic tools. One is a functional-residual-vs-covariate plot, which shows how residual patterns change with an explanatory variable. The other is a Function-to-Function plot, which plays the role of the Quantile-Quantile plot that analysts already use to check distributional assumptions.
What the researchers found
In numerical studies, the new plots revealed several common mistakes. They flagged a missing higher-order term, such as a squared variable. They flagged an omitted explanatory variable and an omitted interaction between variables. They also flagged a missing dispersion parameter and a missing zero-inflation component. The latter matters when a count outcome has more zeros than the model expects.
The framework applies across a wide range of models. It covers binary, ordinal, and count regressions, as well as semiparametric models such as generalized additive models. It also connects two earlier proposals, the surrogate residual and the probability-scale residual, as special cases. Because the plots can be read the way linear-regression plots are read, the same diagnostic habits work for discrete and continuous data.
What it means for analysts and their managers
Models of discrete outcomes drive many business decisions: credit approval, churn prediction, demand for units, and customer satisfaction scores. A model that fits badly can look fine on the usual checks and still mislead. This framework gives analysts a way to see the problem before it reaches a decision.
For managers who review analytical work, the practical point is to ask whether model diagnostics were done, and with what tools. A team that standardizes on one diagnostic approach for all its regression models will find reviews easier and errors rarer.
This summary is based on the paper’s abstract. The full article reports the data, methods, and detailed results.
What it means for managers
- Standard residual checks are weak for discrete outcomes. If your model predicts churn, ratings, or counts, the Pearson and deviance residuals you may rely on can miss real problems.
- The new plots read like the familiar ones for linear regression. The functional-residual-vs-covariate plot and the Function-to-Function plot work the way analysts already expect.
- One framework covers many models. Logistic, ordinal, count, and generalized additive models can all be checked the same way, which simplifies review across a team.
Liu, D., Lin, Z., & Zhang, H. (2025). A unified framework for residual diagnostics in generalized linear models and beyond. Journal of the American Statistical Association, 120(551), 1840-1852. 10.1080/01621459.2025.2504037


