Sextant · Areté Intelligence · Research Appendix

How We Built Sextant

A traceable account of the research pipeline behind the diagnostic — the literature it stands on, the rubric it was authored against, and the triage that decided which of its items survived.

Representative Sextant output — fictional company and modeled assumptions.

What this is. This appendix documents the process that produced the Sextant instrument: how the underlying research literature was assembled and made queryable, how the authoring rubric was derived from it, and how candidate diagnostic items were proposed and cut before anything shipped.

Who it's for. Readers doing diligence on the instrument itself — methodologists, technical buyers, or anyone who wants to know where a specific number came from rather than take it on faith.

Optional depth. This page is not required reading. The product story lives on the rest of the site; nothing here is load-bearing for understanding a report. Read it only if you want to see the work behind the work.

~10 min read 8 sections Updated 2026-05-20
Section 1

Why an off-the-shelf survey was insufficient

Sextant anchors the Discover phase of the engagement model — Discover → Empower → Build. A single Sextant pass has to do the work a separate discovery survey and a separate readiness assessment would otherwise split across two engagements: surface where a client sits against a seven-pillar AI-readiness model, identify which organizational constraints will block execution, and locate where value is quietly being lost. Folding assessment into discovery is a deliberate choice — it keeps the diagnostic a single instrument rather than a bundle of disconnected surveys, and it means the instrument's content quality is directly load-bearing on every recommendation that follows it.

A traditional survey-design process for an instrument of this consequence runs roughly eighteen months: literature review, item drafting, expert review, cognitive interviews, two pilot waves, factor analysis, reliability estimation, norming. The team did not have eighteen months, and did not have a standing working group of survey methodologists on retainer. What it had was deep experience building agentic research and authoring workflows, and a working session that ratified the seven-pillar framework the instrument would measure against.

The bar for the instrument was set high regardless of the compressed timeline: defensible to a sophisticated buyer asking "where did you get this number," defensible to an HR leader who reads survey-methodology papers, and defensible to an academic reviewer who has seen too many vendor frameworks dressed up in borrowed citations. Meeting that bar meant treating instrument design as an engineering problem with an academic-research substrate, not as a set of items drafted by gut and refined by opinion.

The pipeline that resulted ran in roughly ten days:

Brainstorm (batched human feedback) → Parallel literature acquisition (5 research agents) → Hybrid RAG build (vector + BM25 + reranker) → Parallel synthesis (5 synthesis agents querying the RAG) → Rubric merge (single human pass) → Parallel item proposal (6 proposer agents) → Parallel item critique (6 critic agents, 4-axis rubric) → Cross-pillar dedup (single human pass) → Engineering execution

Each stage has a defined contract — what it consumes, what it produces, how it degrades on partial failure — and each stage is reusable. That reusability is the actual point: the corpus, the rubric, and the specialized agents don't get thrown away once Sextant ships. They are the asset the next instrument runs against.

Section 2

Assembling and searching the research corpus

The working session that scoped Sextant named the validated instruments it wanted to embed — a team psychological-safety scale, an AI-attitudes scale, an AI-anxiety scale, a trust-in-AI scale, an AI-literacy scale, an AI-competency scale, a technology-acceptance model, an organizational-readiness-for-change scale, and a fluency rubric for demonstrated (not self-reported) AI skill. None of the underlying papers were in hand at the depth required to author items against them, and items were deliberately not drafted from a model's training-data recall — that is exactly how vendor frameworks ship subtle measurement errors nobody catches until a client pushes back.

So five parallel research agents were dispatched, one per literature sub-domain — foundations, scale design, psychometric validation, workforce surveys, administration and pretesting — each under the same contract: pull open-access sources, capture paywalled canonical works with full citation and DOI where the paper itself couldn't be retrieved, and write a per-subdomain index flagging the standout sources and cross-cutting findings. The pulls were strict about open access; a corpus an outside academic can't verify isn't worth citing.

68
Open-access papers
21
Paywalled stubs, cited
5
Sub-domains, indexed
3
Load-bearing papers via library proxy

Two days later the corpus stood at 68 open-access papers, one HTML capture, and 21 stubs, plus three load-bearing validated-instrument papers retrieved through library proxy access once the open-access pass was exhausted. A handful of the findings the research agents surfaced materially shaped the eventual rubric: agree-disagree anchors are measurably weaker than construct-specific anchors; the empirical literature's breakoff inflection point sits around thirty minutes, well inside a longer session; a thirty-person cohort spread across five management layers averages roughly six respondents per layer — barely above the accepted cell-suppression floor, which forced a pre-registered collapse rule rather than a post-hoc one; and executive-skewed industry sources cannot ground individual-contributor-level claims, which only one source in the corpus — Microsoft's 2024 Work Trend Index — was positioned to do.

A 68-paper corpus is unusable to a synthesis agent working with a limited context window, so the team built a hybrid retrieval index over it — vector search paired with a BM25 lexical index, combined by reciprocal-rank fusion, with an optional cross-encoder rerank pass. Every synthesis agent queried this index during authoring instead of relying on a model's prior knowledge, which is what makes the resulting rubric traceable: every claim in it cites a retrieved passage that cites a paper, rather than reading as plausible-sounding hearsay.

Section 3

From literature to constructs

Eight validated psychometric instruments anchor the culture and psychological-readiness portion of the battery, each administered close to its published form to preserve norm comparability, and each shaped by a decision-utility filter that a purely academic instrument-design process would not apply.

ConstructInstrumentDisposition
Team psychological safetyTPS-7Retained at full 7-item validated length
AI attitudesGAAIS-10 (positive + negative subscales)Positive subscale trimmed to top-3 items by item-total correlation; negative subscale retained in full
AI anxietyAIAS-JR (from AIAS-21)Trimmed from 21 items to the 9-item Job-Replacement subscale
Trust calibrationS-TIASRetained with its validated two-factor trust/distrust structure intact
AI literacy, self-reportWang ALS-12Retained at full validated length
AI competency, objectiveAICOSRetained as a multi-choice knowledge check, not a Likert self-report
Behavioral intentUTAUT2Trimmed from 7 subscales to 2 (Performance Expectancy, Behavioral Intention)
Demonstrated AI fluency4D Fluency rubricRetained as a skill-demonstration instrument, not a self-report scale

One additional instrument — a well-known organizational-readiness-for-change scale — was evaluated and dropped entirely. In a thirty-person mid-market cohort it produces a single number that doesn't differentiate any specific recommendation from what the retained instruments already produce; including it would have added a dozen items of respondent burden for no marginal decision signal. That's the same filter, applied in the opposite direction: every trim above kept an instrument's most decision-relevant subscale, and this one instrument's most decision-relevant subscale turned out to be none of it.

Section 4

The authoring rubric and its hard stops

Five synthesis agents, one per corpus sub-domain, each produced an intermediate synthesis document — structured by principle and rule, citing specific source papers, and explicitly naming tensions where the literature disagreed (for example, one line of research argues for a "don't know" filter option on attitude items; another finds that respondents use the scale midpoint as a de facto don't-know response, undermining the filter's purpose). A single merge pass collapsed the five syntheses into one authoring rubric.

15
Rubric sections
15
Hard stops
5
Synthesis documents merged

Every rule in the rubric cites the synthesis paragraph that justifies it, and every synthesis paragraph cites the source paper behind it — a chain intended to survive a skeptical reviewer asking "why is this a rule." A representative sample of the hard stops:

Anchor fidelityNever lift a validated instrument into a different scale length, anchor wording, or item count than its validation paper specifies — no 5-point rescaling of a 7-point instrument.
Two-factor trust structureNever collapse a trust scale's Trust and Distrust factors into a single composite; the validated structure is two-factor, and collapsing it destroys the signal the instrument was chosen for.
No reverse coding on custom itemsNever insert reverse-coded items into custom item banks; use dedicated attention-check items instead.
Confidential, never anonymousNever claim anonymity to respondents. Use the word "confidential" — layer, tenure, and function together can re-identify a respondent in a thirty-person cohort even without a name attached.
Session orderingNever sequence a conversation before the survey portion of a session; conversational priming contaminates the validated psychometric measurement that most needs protecting.

None of these were stylistic preferences. Each one survived a real argument with a specific finding in the literature before it became a rule the item bank had to obey.

Section 5

Proposer / critic triage

The validated instruments cover roughly half the battery. The remaining six pillars — strategy, data foundation, technology adoption, governance, operating model, and value capture — needed bespoke items no published instrument captures. Standard practice is a single survey methodologist spending weeks drafting these by hand. Instead, six proposer agents were dispatched in parallel, one per pillar, each instructed to draw on the synthesis documents (queried live against the research index), push into assessment modes beyond simple Likert items, prioritize variety and quantity over premature polish, and tag every item with the specific decision it was meant to inform.

136
Candidate items proposed
93
Items retained after triage
6 + 6
Proposer + critic agents

A second wave of six critic agents, also parallel, scored every candidate item on a four-axis rubric — decision utility (would the answer change the recommendation?), methodological rigor (does it honor the rubric's hard stops?), implementation feasibility (existing assessment mode, small extension, or a costly novel build?), and cognitive cost relative to signal (low-load, high-signal items score best). Each axis scored 0–2; items scoring under 5 were cut, 5–6 were kept with modification, and 7 or above were kept outright. The four axes were deliberately chosen to represent the four constituencies the instrument has to satisfy at once — the business operator, the academic reviewer, the engineering team, and the respondent sitting through the session. Optimizing for only one of them is how instruments end up rigorous but unusable, or usable but indefensible.

The first pass returned 63 KEEP, 33 KEEP-WITH-MODIFICATION, and 40 CUT — 96 items provisionally retained out of 136 proposed. A second, cross-pillar deduplication pass then caught something the per-pillar critics structurally couldn't see: several novel assessment modes — adversarial scenario response, diagnose-from-symptoms judgment, defend-a-position prompts — had each independently surfaced in four or five different pillars, which would have produced format fatigue in a battery already pushing the upper end of a reasonable session length. The dedup pass kept the strongest instance of each novel mode per pillar and cut the redundant copies, landing on 93 final retained items, each with an explicit ownership assignment to a specific pillar and a specific downstream decision.

Section 6

Demonstrated skill, not self-report

A recurring rubric principle is that self-report attitude data and demonstrated capability are different signals, and a defensible instrument has to keep them visibly separate rather than blend them into one composite. Three parts of the battery are deliberately built to measure something closer to demonstrated skill or observed behavior than stated opinion:

  • Objective knowledge, not self-assessed knowledge. The AI-competency instrument is scored as a multi-choice knowledge check rather than a Likert self-report, precisely because self-assessed AI knowledge and measured AI knowledge correlate weakly — pairing them, rather than substituting one for the other, is what makes the culture pillar's readout credible.
  • Demonstrated fluency, not claimed fluency. The fluency rubric evaluates what a respondent actually does with an AI tool in a structured task, not what they say they can do with one.
  • Structural signal alongside attitudinal signal. A contested thesis in the literature holds that middle management is where AI adoption stalls — but a rival body of research argues the blocker there is structural (budget authority, protected change-management time, sanctioned tool access), not attitudinal (motivation, willingness). The rubric requires the battery to measure both before a report is allowed to name the constraint, rather than inferring a structural claim from attitudinal data alone.

The same discipline governs an indirect-prevalence technique used to estimate unsanctioned AI tool use without asking respondents to self-incriminate. That technique requires pooling across a much larger sample than a single cohort provides to be statistically trustworthy — so single-cohort results from it are locked out of client-facing reporting as exploratory-only, a rule enforced at the deliverable level rather than left to reviewer discretion.

Section 7

Known limitations

The instrument's own literature review surfaced constraints it could not fully resolve, and the rubric encodes them as guardrails rather than pretending they don't exist:

  • Small-cohort statistics. A thirty-person cohort spread across five management layers produces thin per-layer samples — averaging around six respondents per layer, close to the accepted floor for reporting a layer-level number at all. The pre-registered collapse rule (merge adjacent layers in a fixed order when a layer falls below the floor, never decided after the fact) manages this but doesn't eliminate the underlying sample-size constraint.
  • Executive-skewed source literature. Much of the industry research the corpus draws on for organizational-change claims samples executives and senior leaders disproportionately, which does not ground claims about individual-contributor experience. Only one source in the corpus — a large-scale workplace survey — sampled deeply enough at the individual-contributor level to be trusted for that layer, which is a thinner evidentiary base than the executive-level claims enjoy.
  • Indirect-prevalence estimation needs scale. The item-count technique used for unsanctioned-tool-use estimation is only statistically sound pooled across many cohorts; any single engagement's result from it is exploratory by construction, not a number a client should act on alone.
  • Contested organizational theory. Where the literature has a live, unresolved disagreement — most visibly, whether middle-management resistance to AI adoption is motivational or structural — the instrument measures both sides rather than picking a winner, which adds session length without fully resolving the theoretical debate the sources are still having with each other.

What would most strengthen the instrument going forward: a larger pooled sample across multiple cohorts to make the indirect-prevalence technique reportable per engagement rather than only in aggregate; an external validation study correlating Sextant scores against independent outcome data; and sector-specific overlays (a regulated-industry variant is already partly scoped) that would let sector-specific literature — rather than general workforce literature — ground sector-specific claims.

Section 8

Citations & further reading

The source document behind this appendix has no formal bibliography or hyperlinked references — the corpus itself, not a reference list, is the citation trail. What follows is presented at the level of detail the underlying research actually specifies: instrument names and the concepts or researchers named alongside them in the working documents. Where only a surname and a finding are on record, that's what's listed.

  • TPS-7 — Team Psychological Safety scale, 7 validated items plus AI-specific extension items.
  • GAAIS-10 — General Attitudes toward AI Scale, positive and negative subscales.
  • AIAS-21 / AIAS-JR — AI Anxiety Scale (Wang & Wang); Sextant administers the 9-item Job-Replacement subscale.
  • TIAS-AI / S-TIAS — Trust in AI Scale; two-factor trust/distrust structure per Scharowski & Perrig (2024) and McGrath (2025).
  • Wang ALS-12 — AI Literacy Scale, self-report.
  • AICOS — AI Competency Scale, objective multi-choice knowledge assessment.
  • UTAUT2 — Unified Theory of Acceptance and Use of Technology 2; Sextant retains the Performance Expectancy and Behavioral Intention subscales.
  • 4D Fluency rubric — a demonstrated-skill AI-fluency rubric, evaluated rather than self-reported.
  • ORIC — Organizational Readiness for Implementing Change (Shea); evaluated and dropped from the battery.
  • Saris, Revilla, Krosnick & Shaeffer (2010) — European Social Survey MTMM-experimental evidence that construct-specific response anchors outperform generic agree-disagree anchors.
  • Galesic & Bošnjak — empirical work on survey breakoff and attention decay, cited for the session-length ceiling that shaped the rubric's save-and-resume requirement.
  • Krosnick, and Sturgis — opposing findings on whether an explicit "don't know" filter option improves or contaminates attitude-item response quality.
  • McChrystal Group — structural (rather than motivational) counter-argument to the "frozen middle management" adoption-resistance thesis.
  • Microsoft, 2024 Work Trend Index — the corpus's primary source with individual-contributor-level sampling depth, used to ground IC-level claims that executive-skewed sources cannot support.
Sextant · Areté Intelligence · research appendix ← Methodology & Validation