A practical framework for deciding whether one pharmacokinetic dosing method performs better than another across a patient population.
A useful comparison should answer five separate questions:
No single statistic answers all five questions. The best method is the one with reliable performance across these dimensions, not necessarily the one with the smallest value in one table.
Specify inclusion criteria in advance: drug, age range, body-size range, kidney function, clinical setting, number and timing of samples, and whether repeat fits from the same patient are allowed.
If repeat fits are included, report whether the analysis is patient-based or fit-based. Treating many observations from one patient as independent can make confidence intervals and significance tests look more certain than they are.
Decide whether the comparison concerns concentration prediction, clearance and volume estimation, steady-state exposure, peak and trough attainment, or dose selection. A model can perform well for one purpose and poorly for another.
For vancomycin, report exposure such as 24-hour area under the concentration-time curve and target attainment. For aminoglycosides, report peak and trough attainment according to the dosing strategy.
For each measured concentration, calculate the difference between the observed value and the model prediction. Keep the sign because it identifies systematic direction, and also calculate absolute or squared error because it describes magnitude.
| Measure | What it tells you | How to interpret it |
|---|---|---|
| Mean Prediction Error | Average signed error. | Near zero suggests little overall bias, but positive and negative errors can cancel. |
| Mean Absolute Prediction Error | Average absolute size of error. | Lower values indicate better typical accuracy, without cancellation of opposite errors. |
| Root Mean Square Error | Typical error with extra emphasis on large mistakes. | Lower values are better; report the concentration unit so the result has clinical meaning. |
| Relative Root Mean Square Error | Root Mean Square Error divided by the reference value. | Useful for comparing drugs or concentration ranges with different scales. |
| Sum of Squared Errors | Total squared discrepancy across observations. | Lower is better only when the same observations and weighting are used. |
For a method fitted with measurement weighting, report the weighted error that was optimized and an unweighted error in the original concentration units. The weighted result explains the fitting objective; the unweighted result helps clinicians understand actual prediction differences.
For a population model, first evaluate the prediction made from population equations before patient-specific fitting. Then evaluate the individualized fit after measured concentrations have updated the patient's parameters.
A method may have modest population prediction error but require large individual parameter adjustments. That pattern suggests the population model is broadly useful but does not fully represent certain patient subgroups. Conversely, a method may fit concentrations well while producing implausible parameters, which warrants investigation rather than automatic acceptance.
Calibration asks whether predicted values track observed values throughout the clinical range. Plot predicted values on the horizontal axis and observed values on the vertical axis, then add a line of identity where prediction equals observation.
A two-compartment model can follow early distribution and later elimination more flexibly than a one-compartment model, but additional parameters can improve fit simply because the model has more freedom. Compare models with both an error measure and a complexity penalty.
Akaike Information Criterion = n × ln(Sum of Squared Errors ÷ n) + 2 × k
Here, n is the number of observations and k is the number of independently estimated parameters. Use the same error definition and observations for every candidate model. A lower Akaike Information Criterion indicates the better balance of fit and complexity. A difference below 2 is usually weak evidence; a difference above 2 supports the lower value, and a difference above 10 is strong evidence.
For Bayesian fits, keep the population-parameter penalty conceptually separate from the Akaike Information Criterion. The Bayesian objective may include penalties for moving fitted parameters away from population values, while the Akaike Information Criterion penalizes the number of estimated parameters. They answer related but different questions and should not be added together.
Group averages can hide clinically important failures. Stratify performance by characteristics that affect pharmacokinetics and dosing decisions:
Body weight, body mass index, lean body weight, and body surface area. Look for increasing error or bias in obesity, underweight patients, or extreme body size.
Creatinine clearance, estimated kidney disease stage, and changing renal function. Check whether errors differ in reduced, normal, or augmented renal clearance.
Number of concentrations, early versus late samples, sampling around infusion, and time from dose. A model can appear strong when tested only with convenient samples.
Report the number of patients in every subgroup. Small subgroup results are exploratory and should not be treated as definitive evidence. Consistent directional bias across a clinically important subgroup is more concerning than a single isolated outlier.
Prediction accuracy matters because it affects dosing decisions. Translate model outputs into clinical categories using targets defined before analysis.
| Clinical question | Recommended group result |
|---|---|
| Did the method reach the desired exposure? | Percentage of patients below, within, and above the target range. |
| Did the method avoid clinically important excess? | Percentage above the toxicity-risk threshold and the size of the excess. |
| Did the method improve classification? | Patient movement between below-target, target, and above-target categories. |
| Are results stable? | Confidence intervals, bootstrap intervals, or a prespecified validation cohort. |
For vancomycin, exposure-based endpoints such as 24-hour area under the concentration-time curve may be more informative than trough alone when that is the stated dosing objective. For aminoglycosides, use endpoints appropriate to the dosing strategy, such as peak exposure and low trough exposure. Do not call a model superior based only on a surrogate endpoint that is not the clinical dosing goal.
When every method produces a result for the same patient, compare methods within each patient first. Calculate the patient-level difference, then summarize the differences across patients. This controls for much of the variation between patients and is usually more informative than comparing two unrelated group averages.
For a group of patients, prefer the dosing method that is accurate, minimally biased, well calibrated, clinically reliable across important subgroups, and appropriately simple. A lower average error is valuable, but it should be supported by: