Clinical outcome assessments (COA), incorporating multiple individual items (i.e., tasks, tests,ย questionsย and ratings), are used to follow the progression of many rare diseases. The most common way ofย utilizingย the results of such assessments is to calculate the total score and then perform inferenceย regardingย disease progression rate and treatment effects using this. An alternative strategy is to analyse all the item level data under a joint model. Commonly this is done under the Item Response Theory (IRT) concept originally developed in the psychometrics field. In IRT models, it is assumed that the different items all reflect an underlying disease severity (the US FDA use the term โreflective indicatorโ model in its guidance for COAs). It is therefore possible to model all item level data as dependent on a common latent variable that is allowed to change withย time,ย treatmentย and other covariates. Item level data is typically categorical and IRT models have some distinct numerical advantages: (i) it treats data using the appropriate numerical assumption (total score analyses typically assume data to be continuous), (ii) IRT models naturally adhere to the bounded nature of the scale, which total score models struggle to do, (iii) it is robust to missing item level data, and, most importantly (iv) it distinguishes the differing information content between items. These properties result in higher power, or higher precision of effect size, for IRT than total score models. There are even more options for the analysis of COAs and these include (i) disease staging, (ii) time-to-deterioration, (iii) responder definitions, and (iv) individual item level response. Longitudinal IRT modelsย representingย information at the most granular level with both respect to item responses and time can be valuable tools in evaluating trial design and analysis options. Even the use of modified COAs can be assessed using such models and under different analysis models. The use of IRT models comes with assumptions and assessments of adequacy of these is a central aspect of IRT model development.ย