Comparing Dosing Models for Groups of Patients

A practical framework for deciding whether one pharmacokinetic dosing method performs better than another across a patient population.

Use the same test set: Every method must be evaluated on the same patients, drug concentrations, dose history, infusion times, sampling times, and therapeutic targets. Otherwise, an apparent difference may reflect different data rather than a better model.
What a group comparison should answer

A useful comparison should answer five separate questions:

  1. Accuracy: How close are predicted concentrations or exposure values to what was observed?
  2. Bias: Does a method systematically overpredict or underpredict?
  3. Precision: How much do individual prediction errors vary around the average?
  4. Calibration: Do predicted and observed values agree across the full clinically relevant range?
  5. Clinical usefulness: Does the method more reliably place patients in the desired therapeutic range without unacceptable toxicity risk?

No single statistic answers all five questions. The best method is the one with reliable performance across these dimensions, not necessarily the one with the smallest value in one table.

1. Define the comparison before looking at results

Use a common patient cohort

Specify inclusion criteria in advance: drug, age range, body-size range, kidney function, clinical setting, number and timing of samples, and whether repeat fits from the same patient are allowed.

If repeat fits are included, report whether the analysis is patient-based or fit-based. Treating many observations from one patient as independent can make confidence intervals and significance tests look more certain than they are.

Use a common clinical question

Decide whether the comparison concerns concentration prediction, clearance and volume estimation, steady-state exposure, peak and trough attainment, or dose selection. A model can perform well for one purpose and poorly for another.

For vancomycin, report exposure such as 24-hour area under the concentration-time curve and target attainment. For aminoglycosides, report peak and trough attainment according to the dosing strategy.

2. Measure prediction error

For each measured concentration, calculate the difference between the observed value and the model prediction. Keep the sign because it identifies systematic direction, and also calculate absolute or squared error because it describes magnitude.

MeasureWhat it tells youHow to interpret it
Mean Prediction ErrorAverage signed error.Near zero suggests little overall bias, but positive and negative errors can cancel.
Mean Absolute Prediction ErrorAverage absolute size of error.Lower values indicate better typical accuracy, without cancellation of opposite errors.
Root Mean Square ErrorTypical error with extra emphasis on large mistakes.Lower values are better; report the concentration unit so the result has clinical meaning.
Relative Root Mean Square ErrorRoot Mean Square Error divided by the reference value.Useful for comparing drugs or concentration ranges with different scales.
Sum of Squared ErrorsTotal squared discrepancy across observations.Lower is better only when the same observations and weighting are used.

For a method fitted with measurement weighting, report the weighted error that was optimized and an unweighted error in the original concentration units. The weighted result explains the fitting objective; the unweighted result helps clinicians understand actual prediction differences.

3. Separate population prediction from individual fitting

For a population model, first evaluate the prediction made from population equations before patient-specific fitting. Then evaluate the individualized fit after measured concentrations have updated the patient's parameters.

A method may have modest population prediction error but require large individual parameter adjustments. That pattern suggests the population model is broadly useful but does not fully represent certain patient subgroups. Conversely, a method may fit concentrations well while producing implausible parameters, which warrants investigation rather than automatic acceptance.

4. Evaluate calibration, not only average error

Calibration asks whether predicted values track observed values throughout the clinical range. Plot predicted values on the horizontal axis and observed values on the vertical axis, then add a line of identity where prediction equals observation.

Useful calibration checks
  • A regression slope near 1 suggests that predictions change at an appropriate rate as observations change.
  • An intercept near 0 suggests little constant offset.
  • A high correlation alone is insufficient: predictions can correlate strongly while consistently missing the identity line.
  • Inspect residual plots against predicted concentration, time after dose, dose interval, kidney function, body size, and number of samples.
  • Use a Bland-Altman-style plot of prediction minus observation against their average to reveal concentration-dependent bias.

5. Compare model complexity fairly

A two-compartment model can follow early distribution and later elimination more flexibly than a one-compartment model, but additional parameters can improve fit simply because the model has more freedom. Compare models with both an error measure and a complexity penalty.

Akaike Information Criterion = n × ln(Sum of Squared Errors ÷ n) + 2 × k

Here, n is the number of observations and k is the number of independently estimated parameters. Use the same error definition and observations for every candidate model. A lower Akaike Information Criterion indicates the better balance of fit and complexity. A difference below 2 is usually weak evidence; a difference above 2 supports the lower value, and a difference above 10 is strong evidence.

For Bayesian fits, keep the population-parameter penalty conceptually separate from the Akaike Information Criterion. The Bayesian objective may include penalties for moving fitted parameters away from population values, while the Akaike Information Criterion penalizes the number of estimated parameters. They answer related but different questions and should not be added together.

6. Examine performance across patient subgroups

Group averages can hide clinically important failures. Stratify performance by characteristics that affect pharmacokinetics and dosing decisions:

Body size

Body weight, body mass index, lean body weight, and body surface area. Look for increasing error or bias in obesity, underweight patients, or extreme body size.

Renal function

Creatinine clearance, estimated kidney disease stage, and changing renal function. Check whether errors differ in reduced, normal, or augmented renal clearance.

Clinical sampling

Number of concentrations, early versus late samples, sampling around infusion, and time from dose. A model can appear strong when tested only with convenient samples.

Report the number of patients in every subgroup. Small subgroup results are exploratory and should not be treated as definitive evidence. Consistent directional bias across a clinically important subgroup is more concerning than a single isolated outlier.

7. Measure clinical target attainment

Prediction accuracy matters because it affects dosing decisions. Translate model outputs into clinical categories using targets defined before analysis.

Clinical questionRecommended group result
Did the method reach the desired exposure?Percentage of patients below, within, and above the target range.
Did the method avoid clinically important excess?Percentage above the toxicity-risk threshold and the size of the excess.
Did the method improve classification?Patient movement between below-target, target, and above-target categories.
Are results stable?Confidence intervals, bootstrap intervals, or a prespecified validation cohort.

For vancomycin, exposure-based endpoints such as 24-hour area under the concentration-time curve may be more informative than trough alone when that is the stated dosing objective. For aminoglycosides, use endpoints appropriate to the dosing strategy, such as peak exposure and low trough exposure. Do not call a model superior based only on a surrogate endpoint that is not the clinical dosing goal.

8. Use paired analysis when methods are applied to the same patients

When every method produces a result for the same patient, compare methods within each patient first. Calculate the patient-level difference, then summarize the differences across patients. This controls for much of the variation between patients and is usually more informative than comparing two unrelated group averages.

9. A recommended reporting sequence

  1. Describe the cohort, data quality, sampling schedule, and missing data.
  2. Confirm that every model used the same observations and dosing information.
  3. Report convergence failures and exclude or investigate them transparently.
  4. Show population prediction error and individualized fit error separately.
  5. Report signed bias, absolute error, root mean square error, and normalized error.
  6. Show predicted-versus-observed and residual plots.
  7. Compare Akaike Information Criterion values using the same observations and likelihood convention.
  8. Show subgroup performance by body size, renal function, and sampling pattern.
  9. Report clinical target attainment and the number of patients in each category.
  10. Validate the preferred method on a separate or temporally later cohort whenever possible.
Bottom line

For a group of patients, prefer the dosing method that is accurate, minimally biased, well calibrated, clinically reliable across important subgroups, and appropriately simple. A lower average error is valuable, but it should be supported by:

  • similar or better performance in early and late concentration samples;
  • no important systematic bias by body size or renal function;
  • acceptable convergence and clinically plausible parameters;
  • better or equivalent therapeutic target attainment; and
  • independent validation in patients not used to develop the method.