The tool goes under pressure too

Auditing the Human Scale Audit

A self-reflection tool can be useful before it is scientifically validated. The dangerous moment is when a convenient score begins to look more objective than the questions underneath it.

If the total score creates more false certainty than insight, Human Scale should be willing to delete it.

Define what the tool is

The Audit can support reflection, conversation, experiment design, or formal measurement. The first three require far less evidence than the fourth. Until calibration exists, the public claim should remain simple: this is a reflection tool, not a diagnostic or scientifically validated index.

The total score is a hypothesis

Two people can have the same total with entirely different lives. Adding nature, money, time, relationships, security, movement, or meaning can hide value judgments and allow one strength to appear to cancel a serious problem elsewhere.

Test whether a total score adds value beyond a profile of dimensions. Do not assume the answer is yes.

Five immediate stress tests

Content validity

What exactly is each question supposed to measure, and why does that construct belong?

Cognitive interviews

Ask people to think aloud. Find ambiguity, double questions, hidden assumptions, and response options that do not fit real lives.

Accessibility

Make sure disability, assistive support, or nonstandard independence are not graded as personal failure.

Overlap

A commute can reduce sleep, time, movement, family, and community. Decide whether downstream effects are intentionally counted separately.

Preference vs opportunity

Separate “can you?”, “do you?”, “do you want to?”, and “is the current amount working for you?” where those differ.

Weighting must be honest

If a total survives, publish exactly how weights are chosen. Equal weighting is a design choice. Expert weighting is a judgment. Outcome prediction is empirical. Personal weighting reflects the user. None should masquerade as the others.

A future results page could show an unweighted profile, a clearly labeled Human Scale default, a personal weighting, and an option to ignore the total entirely.

No invented clinical thresholds

A reflection tool should avoid labels such as “normal,” “deficient,” “at risk,” or “optimal” unless a validated measure actually supports them. Safer language is descriptive: this area appears to create more friction; this dimension is lower than your others; this may be worth investigating.

Minimum calibration sequence

1 · Adversarial review Evidence, accessibility, survey-method, psychometric, and ideological critics.
2 · Cognitive interviews Small diverse sample; rewrite aggressively.
3 · Usability & accessibility Mobile, keyboard, screen reader, low vision, cognitive load, plain language, slow connection.
4 · Pilot Missingness, distributions, completion, test–retest, redundancy, dimension structure.
5 · Fairness Stress-test items across disability, age, culture, language, income, household form, work, caregiving, and geography.
6 · Sensitivity See whether real Human Scale Lab changes move the intended dimensions without random spillover.
7 · Total-score decision Keep it, relabel it, make it optional, or delete it.

Privacy stays part of validity

The current local/no-signup approach is a strength. Calibration should use explicit opt-in research modes rather than silently converting a private reflection tool into behavioral data collection.

What success looks like

People understand questions consistently enough; irrelevant items can be handled fairly; disability and life stage are not scored as failure; dimensions are meaningfully distinct; weights are transparent; relevant real-life changes move relevant dimensions; and users leave with at least one useful question or experiment.

The best success metric may not be the number. It may be: did the tool help someone notice a real source of friction and choose a sensible next experiment?

Result of Cross-Cutting Audit 05

The Human Scale Audit should be treated as a hypothesis-generating reflection tool until its questions, dimensions, weighting, accessibility, reliability, and fairness survive real testing.