Sites Data Inconsistency Report: Purpose and Interpretation
The Sites Data Inconsistency Report helps to identify sites whose data distributions and/or time patterns differ from the study-wide reference more than expected by random variation. In other words, to find sites that “look different” (too high/low, unusual spread, odd visit timing trends), so you can check for process issues, measurement differences, or data problems.
To access the Sites Data Inconsistency Report navigate use the menu Analysis -> Unsupervised CSM.

Additionally, the Combined Data Inconsistency score (Z-score) per site is displayed in the Sites List Page and indicates standardized distance from the reference (the higher the score the more unusual the site behaves).

How to read the report (workflow)
- Start with the Combined Sites Inconsistency Score (Z-score) chart: pick the sites above the confidence threshold.
- Use the Individual Statistical Tests Overview heatmap table: see which test(s) drive the site’s overall score.
- Click a cell: the Statistical Test Details section updates to explain that specific signal (distribution or time trend).
- Interpret with context: confirm if differences are clinically plausible or suggest data capture, lab, operational issues.
Chart 1 — Combined Sites Inconsistency Score (Z-score)

What you see
- A bar chart by Site ID, sorted from highest to lowest.
- A horizontal “Confidence threshold” line (value 3).
- Background shading indicating “below vs above threshold”.
How to interpret
- Z-score = standardized distance from the reference (in SD units).
- Higher Z-score ⇒ more unusual site overall.
- Example intuition: Z = 3 means “about 3 standard deviations away” (rare under normal assumptions).
- Use this chart to prioritize:
- Above threshold: investigate first.
- Just below threshold: monitor, especially if trending up over time.
Table 1 — Individual Statistical Tests Overview (heatmap table)

What you see
- Rows = sites
- Columns grouped by domain/parameter (e.g., Lab Hemoglobin, Platelet count, WBC count, Neutrophils, etc.)
- Each data item column consists of multiple statistical test columns, typically represented by:
- P-Val, KS P-Val, Adj. P-Val, Percentile, Timeseries
- Categorical tests (Nominal / Ordinal / Binary) — one Final Z-score column per test. An ordinal test can be expanded to show the underlying detectors it combines (Final, Mann–Whitney, Pearson χ²). Categorical data contains a limited set of possible results, such as Positive/Negative, ordered severity grades, or named result categories. The analysis compares the proportion of records in each category at the selected site with the pooled proportions from all other sites in the study.
How to interpret the color + numbers
By default, the results of all statistical tests per are collapsed into a single cell per data item. This cell's color reflects the strongest inconsistency detected across all tests, but no numeric value is shown, since results from different statistical tests aren't directly comparable.
Expanded statistical test cells display values. Darker cells indicate a stronger signal (more inconsistent vs. reference) for that statistic. Cells with values below 3 are colored white, as they lie below the confidence threshold. Gray cells are shown for those sites and statistics that have too less values to be statistically meaningful.
Practical interpretation tips (quick rules)
-
High combined Z-score + multiple dark cells:
Broad inconsistency → investigate site operations.
-
High combined Z-score driven by one parameter:
Focused issue → investigate that domain (e.g., specific lab).
-
Strong signal + low sample size:
Verify with more data / monitor trend before escalation.
- Always combine statistics with clinical plausibility and operational context (country, lab, device, training date, vendor changes).
Click any cell with displayed value to open the exact supporting plot below.
Once a Z-score exceeds conventional significance thresholds (around 3), further increases primarily indicate greater statistical certainty rather than proportionally greater practical severity. In applied site inconsistency, Z-scores beyond this range are therefore better interpreted as confirming the presence of inconsistency rather than scaling its magnitude.
Column meanings
|
Column |
Statistical meaning |
Plain-language meaning |
|
P-Val |
Evidence against “site matches reference” for a selected test |
“How unlikely is this difference if the site were normal?” |
|
KS P-Val |
P-value from Kolmogorov–Smirnov test (distribution shape difference) |
“Does the whole distribution look different (not just the mean)?” |
|
Adj. P-Val |
P-value adjusted for multiple comparisons (controls false positives) |
“Significance after correcting for many tests/sites” |
|
Percentile |
A percentile-based extremeness comparison (site vs reference) |
“How different is this the two tails of site’s distribution compared with others/reference?” |
|
Timeseries (TS) |
Time-pattern anomaly score (based on model vs observed over visit time) |
“Does this site’s trend over time look unusual?” |
| Final [Z] | Final categorical inconsistency score for the test; for ordinal tests it is the strongest (maximum) of the test's detectors | How unusual the site's category pattern is overall for this test |
| MWU | Mann–Whitney U — tests for a shift toward higher or lower ordered categories (ordinal) | Do the site's ordered results tend higher or lower than the reference? |
| Pearson χ² | Permutation Pearson chi-square — tests the overall category pattern (nominal and ordinal) | Does the whole mix of categories differ from the reference? |
| Fisher | Fisher's exact test — compares the positive/negative rate (binary) | Does the positive/negative rate differ from the reference? |
Export Data
The table with all sites can be exported as csv file by clicking the download button in the header of the table.

Section — Statistical Test Details (updates when you click a cell)
A) Timeseries details — “GAM baseline vs actual … (TS)”

What you see
- Scatter points for the selected site across relative study day / data value.
- A smooth baseline curve labeled GAM baseline (Generalized Additive Model).
-
A shaded 95% confidence interval band around the baseline.
How to interpret
- If points mostly stay within the band: site follows the expected time pattern.
- If points systematically shift above/below the baseline: possible site-specific bias (e.g., consistently higher values).
- If deviations appear only in certain time windows: possible phase-specific issue (startup training, device change, local lab change, visit scheduling artifacts).
-
Wider confidence band typically means less information / more uncertainty in that time region.
Typical follow-ups
- Check lab vendor / calibration, collection timing, unit conversions, visit window handling, data entry delays, protocol deviations at that site.
B) Distribution details — “Kernel Density Estimation plot: Site vs Reference”
(Shown when selecting P-Val or Adj. P-Val)

What you see
- Overlaid density/histogram-like distributions:
- Reference (all other sites or pooled data)
- Selected site
- A small summary table with metrics such as:
-
Central Tendency (mean/median-like)
Mean < Median: left/negative skew; investigate low values.
- Variability (Standard Deviation, Mean Absolute Deviation)
- Standard deviation ≈ Mean Absolute Deviation: variability is evenly distributed; few extreme values.
- Standard deviation >> Mean Absolute Deviation: large number of outliers and heavy tails
- Quantile Difference (25% and 75%)
-
Sample Size, Missing/Unidentifiable
Number of data points with NA values
-
Outlier ratio (%)
Percentage of data source values which are 3*1.4826⋅MAD from median
-
How to interpret
- Shift left/right: site has lower/higher typical values.
- Narrower/wider shape: site has less/more variability than reference.
- Different shape (skew/multimodal): may indicate mixed populations, data mixing, measurement/process differences, or data quality issues.
- Small sample size: treat signals cautiously (higher noise / instability).
Plain-language checks
- “Are site values consistently higher/lower than everyone else?”
- “Is the spread suspiciously tight (too perfect) or too wide (unstable process)?”
C) Distribution details — “Cumulative Distribution Function (CDF): Site vs Reference”
(Shown when selecting KS P-Val and also in the Percentile example.)

What you see
- Two step-like curves:
- Reference CDF
- Site CDF
- The KS test is driven by the maximum vertical gap between these curves.
How to interpret
- Curves overlap closely: distributions are similar.
- Consistent separation across the range: systematic difference (location/scale).
- Separation mainly in tails: site differs in extremes/outliers (e.g., unusually high values).
- Largest vertical gap point: where the site diverges most from reference.
Plain-language meaning
-
“At any given value, does this site accumulate observations faster/slower than normal?”
(i.e., “Does the whole pattern of values look different?”)
D) Site versus pooled peer category distribution and Statistical evidence versus practical deviation
(Shown when selecting Ordinal (Final, MWU, Person χ²), Nominal (Final [Z]) or Binary (Final [Z]))
The method depends on the data type:
- Binary data: Fisher's exact test compares the proportions of two possible results.
- Nominal data: a permutation Pearson chi-square test compares the overall pattern across categories that have no natural order.
- Ordinal data: the Mann–Whitney U test identifies a tendency toward higher or lower ordered results, while the permutation Pearson chi-square test checks for other changes in the complete category pattern.
Cramér's V and Cliff's delta may be shown as supporting effect-size measures. Cramér's V describes the strength of an overall category-pattern difference. Cliff's delta describes the direction and size of an ordered shift. They help explain a result but do not independently determine whether a site is flagged.
- Site versus pooled peer category distribution

How to read the category comparison
The detail view shows the site's category percentages beside the pooled peer percentages, together with exact counts and denominators. For binary data, the percentage-point difference shows how much more or less frequently each result occurs at the site. For nominal data, review which categories contribute most to the overall difference. For ordinal data, read the categories in their defined order to identify a tendency toward lower or higher results.
A Rare-category alert indicates that the signal may be driven by categories that are uncommon in the pooled reference or occur only at the selected site. It is an independent review flag, not proof of a data-quality problem and not a separate contribution to the combined score.
Underpowered means that the comparison can be calculated, but the available record counts are not sufficient to detect the predefined meaningful difference reliably.
Insufficient means that the comparison cannot be calculated from the available category pattern. Neither status means that the site has been shown to match the reference.
Always interpret a categorical signal together with clinical and operational context. Follow-up checks may include category coding or mapping, local laboratory or vendor conventions, changes in equipment or procedures, missing or unidentifiable values, and differences in the site's subject population.
- Statistical evidence versus practical deviation

How to interpret Z-score versus Total Variation Distance (TVD)
Each point in the chart represents one site for the selected categorical test. The two axes answer different questions:
- The absolute Z-score shows the strength of the statistical evidence. Values at or above 3 indicate strong evidence that the site's category pattern differs from the pooled reference. A higher Z-score means greater statistical certainty, not necessarily a larger practical difference.
- Total Variation Distance (TVD) shows the size of the overall distribution difference. It compares all relevant category percentages at the site with those in the pooled reference and summarizes the difference as a value from 0% to 100%. A TVD of 0% means that the distributions are identical. A TVD of 10% means, approximately, that 10% of the distribution would need to be reassigned between categories for the site and reference patterns to match. Values at or above 10% indicate a large practical difference for review.
TVD describes the overall magnitude of the difference, but it does not show which categories are responsible or whether an ordered result shifted higher or lower. Use the category comparison bars to identify the direction and source of the difference. For binary data, TVD is equivalent to the absolute difference in the configured positive-category rate. For nominal and ordinal data, the chart emphasizes differences among the main categories; unusual rare categories are highlighted separately.
| Zone | Interpretation |
|---|---|
|
Actionable anomaly (Z-score ≥ 3; TVD ≥ 10%) |
Strong statistical evidence and a large practical difference. Review first. |
|
Statistically unusual, minor (Z-score ≥ 3; TVD < 10%) |
The difference is unlikely to be random, but its practical magnitude is small. Interpret in context. |
|
Large deviation, uncertain (Z-score < 3; TVD ≥ 10%) |
The observed difference is large, but the statistical evidence is not yet strong. Review supporting data and record counts. |
|
No strong signal (Z-score < 3; TVD < 10%) |
Neither statistical evidence nor practical magnitude crosses the review boundary. |
