# Narrative Pro methodology: curated evidence summaries

Reviewed 2026-09-07. Article version: methodology-2026-09-07.

This publication summarizes internal records. It is not raw run data, an
independently reproduced benchmark, or a peer-reviewed study. The summaries were
deliberately authored for public reading; private source narratives and test
artifacts are not included. Missing fields remain **Not recorded**. A record's
date is not a production release date.

## length-control-2026-06

- Date: 2026-06-27. Category: writing. Status: reported-observation.
- Question: Can measured revision approach a character limit without inventing facts?
- Hypothesis: Not recorded.
- Method and input: retrospective summary of a generate, measure and revise spike
  in a historical design note. Original input privacy classification: Not recorded;
  no input narratives are reproduced.
- Sample: six reported sample-by-target runs, with targets of 250 and 350
  characters. Unique input count and pairing: Not recorded.
- Model: gemini-3-flash-preview. Prompt version: Not recorded.
- Counting convention in the original runs: Not recorded. A later specification
  of normalized Unicode codepoints is not proof of the spike's measurement method.
- Observation: the note describes reaching a near-limit band through revision,
  but also describes naive prompts inventing numeric details.
- Limitations: original run-level inputs, outputs, prompts and counting records
  were not located in this bounded review. This does not prove they never existed.
  No benchmark rate, average call count, quality rating or customer outcome is
  published. Length fit alone is not evidence fidelity.
- Decision: measure outputs; ask for missing evidence; reject unsupported additions.

## field-fit-2026-09

- Date: 2026-09-06. Category: writing. Status: reported-observation.
- Question: Does a live draft meet the near-limit band and remain reviewable?
- Hypothesis: Not recorded.
- Method and input: local workflow QA with fictional evidence and disposable
  accounts. No account identifiers or output narratives are published.
- Sample: two individual generation observations described in a QA record,
  not a complete run ledger.
- Model: Gemini; exact model version: Not recorded. Prompt: first observation
  pre-v2 (exact version Not recorded); later observation labeled v2.
- Counting convention: the record reports Unicode codepoints with a target of
  350 and a zero-to-five-under band. Original normalization and count artifacts
  were not independently reconstructed for publication.
- Observation: both described drafts were labeled short of the requested band.
  The record separately reports explicit Apply and Undo working.
- Limitations: original run-level artifacts were not reconstructed; exact output
  counts and quotations are withheld. The observations are not a paired A/B test
  and cannot establish a model success rate or general factual accuracy.
- Decision: show short, within-band and over-limit states separately from writing
  feedback and factual review.

## synthetic-walkthroughs-2026-09

- Date: 2026-09-05 (remediation ledger, referring to evaluation on 2026-09-04).
  Category: usability. Status: reported-observation.
- Question: Where can a career-writing journey lose evidence, edits or intent?
- Hypothesis: Not recorded.
- Method and input: synthetic consumer-persona walkthroughs summarized by a
  remediation ledger. Only generalized usability findings are included here.
- Sample: 15 synthetic walkthroughs, not 15 recruited users.
- Model and prompt versions: Not recorded. Character-count convention: not
  applicable to the walkthrough sample unit.
- Observation: the ledger identifies lost unsaved work, signup-handoff fragility,
  unclear field-fit feedback, target-role confusion and evidence-preservation
  problems. Numeric preservation alone did not resolve unsupported semantic claims.
- Limitations: these were not interviews or recruited participants. No real-user
  prevalence, satisfaction or customer outcome can be inferred. Local software
  closure checks are not a follow-up usability study or production delivery.
- Decision: retain source evidence and edits, ask neutral questions, and separate
  desired roles from actual employment history.

## resume-reliability-2026-09

- Date: 2026-09-07. Category: engineering. Status: reported-observation.
- Question: Can resume editing preserve intent across delays, save failures and navigation?
- Hypothesis: Not recorded.
- Method and input: local execution ledger reporting unit regressions, controlled
  browser fixtures, type checks, static audits and a production build. No raw
  fixtures or personal records are included. No checks were rerun for this summary.
- Sample: regression scenarios for stale responses, recovery, target-role
  persistence, cancellation and optional Air Force entry. Assertions are not
  participants or model runs.
- Model and prompt versions: not applicable to controlled provider fixtures.
  Character-count convention: not applicable to these software checks.
- Observation: the ledger reports passing checks for preserving newer edits,
  recovering unsaved work and retaining target-role changes. Browser checks also
  found authentication-handoff and route-initialization defects missed by focused
  unit mocks; the recorded revision corrected them.
- Limitations: this is recorded local verification, not an independent article
  rerun. Software checks do not establish writing quality, live-model accuracy,
  customer outcomes or production deployment. Recovery is a time-limited latest
  copy per tab, not permanent backup. Independent release review remains required.
- Decision: keep proposals reviewable, reject stale results and test whole journeys.

## scoring-history

Reviewed 2026-09-07. Method description, not an empirical study.

The historical C6 framework classified data maturity on six levels: basic metric
mention, comparison, multiple-point analysis, multi-faceted statistical analysis,
prediction, and causal analysis with attribution. Its military formula was:

`score = 1.25 * (impactOrgLevel - currentOrgLevel) + c6Level`

The organizational-level difference represented rank-relative reach. The 1.25
multiplier was a chosen scoring weight, not measured evidence of 25 percent more
impact. The historical corporate formula was `score = c6Level + impactLevel`.
These definitions are historical product rules, not current military regulations
or evidence that wording establishes causation.

The inspected det-v1 implementation uses five dimensions, each from 0 to 2:
specificity, quantification, impact scope, structure and language. Their sum is
from 0 to 10. The deterministic editor path uses text heuristics, including
keyword scope and clause structure. Server grading can instead use a language
model for impact scope and coaching, with a deterministic fallback. Neither path
verifies the underlying accomplishment. The old and new scores are not a
validated common scale or probabilities of employment, promotion or awards.

## fictional-illustration

Reviewed 2026-09-07. Illustration only; not part of the experiment register.

The original, clarified and refined example was authored as fiction. Clarification
supplies new fictional evidence; refinement changes wording without adding facts.
The observed time change in that example is not proof of causation. No experimental
grade or measured quality improvement is assigned to it.

## next-protocols

Recorded 2026-09-07. Status: planned. Neither protocol has been run.

1. Paired evidence and field-fit evaluation: freeze synthetic inputs, counting
   rules, prompts and models; retain all attempts and failure types; distinguish
   unique inputs from repeats. Sample size and versions: Not recorded.
2. Human reading and reuse: define consent, recruitment and rating criteria;
   blind and randomize comparisons where feasible; retain disagreements. Sample
   size and recruitment records: Not recorded. No hiring-outcome claim is proposed.

## change-history

2026-09-07: first curated article record. Separates reported writing observations,
synthetic usability evaluation, engineering verification and fiction. Withholds
unsupported rates and exact historical output counts. Documents missing evidence
and planned protocols without presenting them as completed results.
