01 / Research question
What makes a stronger statement trustworthy?
How can we turn real accomplishments into concise, persuasive writing without inventing evidence?
Our working method separates evidence gathering from wording changes. First establish what the writer did, who used it and what was observed. Then clarify the language, measure the requested field fit and return a proposal for the writer to review.
The records reviewed here point to an important distinction: a draft can fit a character limit and still invent a number. Conversely, a useful draft can preserve the evidence while falling short of a requested length band. Neither outcome should be hidden behind a single writing score.
Scope of the evidence. This is an account of reported development observations, synthetic walkthroughs and local software verification. It is not peer-reviewed research, a representative model benchmark, or proof of hiring, promotion or award outcomes. Publication scope and record [1]
02 / Method evolution
From data maturity to actionable feedback
The historical C6 framework
C6 described six levels of data maturity: mentioning a metric, comparing data, analyzing multiple points, using multi-faceted statistical analysis, forecasting, and causal analysis with attribution. The military method combined that level with organizational reach relative to the writer's rank.
score = 1.25 * (impactOrgLevel - currentOrgLevel) + c6Level
Here, impactOrgLevel represented who benefited; currentOrgLevel represented the writer's rank-based organizational level. The 1.25 multiplier was a chosen scoring weight, not measured evidence of
25 percent greater impact. The historical corporate method added data maturity to an
impact level. These are historical product definitions, not current military
regulations. Scoring history [2]
The current rubric: det-v1
The five-dimension rubric makes revision feedback more specific: what is unclear, what evidence is present, how far the described effect reaches, and how the statement reads. Each dimension contributes 0 to 2 points, summed to 0 to 10. This is a feedback instrument, not a probability of selection.
| Dimension | Signal and boundary |
|---|---|
| Specificity | Named-detail and vague-language heuristics; a recognized name is not verified evidence. |
| Quantification | Numbers and metric-like units; a detected number is not proof of its accuracy. |
| Impact scope | Keywords suggesting team or wider reach; mentioning a department does not prove a department-wide effect. |
| Structure | Supplied action, impact and result sections, or a clause-count proxy; clauses alone do not establish a coherent explanation. |
| Language | Opening verbs and stock phrases; these rules cannot capture every audience's judgment. |
The editor's deterministic path applies fixed text rules without a model call. Server grading can use a language model for impact scope and coaching, with keyword fallback for scope. Model-assisted interpretation and deterministic scoring are different paths; neither verifies the accomplishment. Implementation summary [2]
The old and new scores do not share a validated common scale. We do not translate a C6 score into det-v1 or claim that a higher score demonstrates better career outcomes.
03 / Experimental method
Separate the evidence from the edit
The following is our evaluation framework, not a claim that every historical observation followed a complete protocol. Each study below states what was actually recorded; missing fields remain Not recorded.
- 01Establish the source
Action, ownership, reach, measurements and their limits.
- 02Draft and measure
Preserve supplied facts. Inspect wording and count the output.
- 03Review or ask
Accept a supported proposal, reject it, or request missing evidence.
Define success before comparing outputs
- Evidence fidelity: preserve the supplied numbers, units, ownership and uncertainty. Unsupported employers, awards, wider scope or causal claims are failures even when every numeric token is retained.
- Field fit: report exact length, under-limit fit and near-limit fit separately. State the target, counting convention and handling of failed attempts.
- Readability and quality: specify an audience and rating method. Without actual human-rating records, do not present a quality improvement as measured.
- Missing evidence: asking the writer for a fact is a legitimate outcome. A refusal or incomplete draft belongs in the record, not outside its denominator.
A character is a measurement choice
The worked figure uses Unicode codepoints after the product's configured punctuation substitutions and trailing-whitespace trimming. A codepoint is a Unicode unit, not necessarily one visible symbol. This differs from UTF-16 string length, word counts and some form counters. Historical studies retain their own stated convention; we do not substitute today's counter for missing records.
A fair prompt comparison needs the same inputs and comparable conditions, with model and prompt versions retained. Repeated trials on one input are not independent people or new source examples. We publish no numerical comparison chart here because the inspected historical records do not supply the required source rows.
04 / Experiment register
What the records show
Writing observations, synthetic usability evaluation and engineering verification answer different questions. They are kept separate below, with no combined success counter. All entries are reported observations: a public summary is not a reproducible run artifact.
Writing observation
Length control exposed an evidence problem
Evidence status: Reported observation
Can measured revision bring a draft close to a character limit without inventing facts?
Reported findings
- The note reports reaching its length band through measured revision, but the original run-level artifacts were not located in this bounded review.
- Naive prompts reportedly invented quantitative details. A draft can fit a field and still fail evidence fidelity.
What this does not establish
- Original run-level inputs, outputs, counting records and complete prompt versions were not located. No success rate or average call count is published.
- Repeated runs are not independent participants. There is no recorded blinded quality rating or hiring-outcome measurement.
Product decision. Measure the output independently; treat unsupported additions as failures. Ask for missing evidence instead of manufacturing a metric.
Study record and method
- Study ID
- length-control-2026-06
- Hypothesis
- Not recorded
- Method
- A design note describes a generate, measure and revise spike at two character targets. It also records invented numbers from naive prompts. This is a retrospective summary, not a reconstructed experiment.
- Input source
- Historical design-note summary; original input set not available in the inspected evidence.
- Privacy
- Original input privacy classification: Not recorded. Only a qualitative summary is public; no source narratives are reproduced.
- Sample unit
- Six reported sample-by-target runs at targets of 250 and 350 characters. The number of unique inputs and their pairing are Not recorded in the available summary.
- Counting convention
- Not recorded for the original runs. The subsequent design specified normalized Unicode codepoints; that specification does not establish how the spike was counted.
- Model
- gemini-3-flash-preview
- Prompt version
- Not recorded
Writing observation
Under the limit is not the same as close to it
Evidence status: Reported observation
Does a live draft meet the requested near-limit band, and remain a reviewable proposal?
Reported findings
- Both described drafts were short of the requested band. Being under the maximum did not establish near-limit reliability.
- The record reports explicit Apply and Undo working. That is a workflow observation, not a quality score for the generated text.
What this does not establish
- Original run-level artifacts were not reconstructed for publication. Exact output counts and quoted outputs are therefore withheld.
- These are not paired observations under controlled conditions and are not an A/B test of prompt versions.
- The record cannot estimate general model reliability or factual accuracy.
Product decision. Report short, within-band and over-limit states honestly; keep field fit separate from writing feedback and factual review.
Study record and method
- Study ID
- field-fit-2026-09
- Hypothesis
- Not recorded
- Method
- Local workflow QA recorded individual live-generation observations before and after a prompt revision. Each reported draft was below the requested near-limit band and labeled short. Apply and Undo were checked separately.
- Input source
- QA notes describing fictional evidence in disposable local accounts.
- Privacy
- Synthetic input; sanitized observations only. No account identifiers or generated narratives are published.
- Sample unit
- Two individual generation observations described in the QA record, not a complete run ledger or a defined population sample.
- Counting convention
- The QA record reports Unicode codepoints against a 350-codepoint target and a zero-to-five-under band. Original normalization and per-run count artifacts were not independently reconstructed.
- Model
- Gemini; exact model version: Not recorded
- Prompt version
- First observation: pre-v2, exact version Not recorded. Later observation: v2 as labeled in the QA record.
Usability evaluation
The journey must preserve the writer's work
Evidence status: Reported observation
Where can an end-to-end career-writing journey lose evidence, edits or intent?
Reported findings
- The ledger identifies lost unsaved work, fragile signup handoffs, unclear field-fit feedback and confusion between a target role and a role already held.
- Numeric evidence preservation and unsupported semantic claims required separate attention; protecting numbers alone did not establish truth.
What this does not establish
- These are not recruited participants, interviews or measured customer outcomes. Synthetic walkthroughs cannot estimate real-user prevalence or satisfaction.
- The ledger reports local remediation. Its passing software checks are not a repeat usability study or evidence of production delivery.
Product decision. Preserve source material and unsaved edits, make missing-evidence questions neutral, and distinguish the desired next role from employment history.
Study record and method
- Study ID
- synthetic-walkthroughs-2026-09
- Hypothesis
- Not recorded
- Method
- A remediation ledger summarizes a September 4 synthetic persona evaluation and subsequent local repairs. Walkthrough findings and software closure checks are different forms of evidence.
- Input source
- Synthetic consumer-persona walkthroughs, summarized in a dated remediation ledger.
- Privacy
- Synthetic scenarios; only generalized usability findings are public, without individual narratives or internal security details.
- Sample unit
- 15 synthetic walkthroughs, not 15 recruited users.
- Counting convention
- Not applicable; the sample unit is a synthetic walkthrough, not a model generation or a character count.
- Model
- Not recorded
- Prompt version
- Not recorded
Engineering verification
A suggestion must not overwrite a newer edit
Evidence status: Reported observation
Can resume editing preserve user intent across delayed responses, failed saves and navigation?
Reported findings
- The ledger reports checks passing for rejecting stale suggestions, retaining unsaved changes and preserving the target role, including explicit clearing.
- Browser checks caught an authentication-handoff issue and a route-initialization issue that focused unit mocks had not exposed; both were corrected in the recorded revision.
What this does not establish
- This summarizes recorded local engineering verification, not an independent rerun for this article.
- Software checks do not establish writing quality, live-model accuracy, customer outcomes or production deployment.
- Recovery is a time-limited latest copy per tab, not a permanent backup. Different-family independent review remains a release requirement.
Product decision. Keep proposals reviewable, reject stale results and retain recoverable edits. Test the whole journey as well as isolated functions.
Study record and method
- Study ID
- resume-reliability-2026-09
- Hypothesis
- Not recorded
- Method
- The local execution ledger reports unit regressions, controlled browser fixtures, static audits, type checks and a production build. Provider responses were controlled; this article does not rerun those checks.
- Input source
- Local software-verification ledger for the resume-stability revision.
- Privacy
- Controlled test fixtures; no personal records or raw test artifacts are published.
- Sample unit
- Scenario-based regression checks of stale proposals, recovery, target-role persistence, cancellation and optional Air Force entry. Assertions are not participants or writing samples.
- Counting convention
- Not applicable; software assertions and browser scenarios are not character-count observations.
- Model
- Not applicable; controlled provider responses
- Prompt version
- Not applicable; no model comparison
05 / Worked example
New evidence first. Better wording second.
This supporting example is deliberately fictional and is not experimentally graded. The clarified version introduces supplied facts that were absent from the original; only the final step is a wording-only revision.
Original, clarified, refined
This is our writing methodology, not scientific validation or a promise of hiring or promotion.
Fictional illustration, not a customer story or model benchmark.
Clarifying question
What did you personally change, who used it, and what did you measure?
Helped the training team improve onboarding.
44 characters; not a completed form field.
Starting point: the contribution and measured change are not yet specified.
Five lenses on the clarified example
- Ownership
- The writer designed the checklist. Eight instructors used it.
- Evidence
- The reported average changed from 38 to 25 minutes over four weeks.
- Significance
- 60 new hires describes reach, not a proven wider organizational effect.
- Clarity
- The refined wording connects the contribution, its use and the observation.
- Precision
- No sole-leadership, financial or proven-causation claim is added.
How our method evolved
Historically, C6 described data maturity, from basic mentions of metrics through causal analysis. The military approach combined it with rank-relative organizational reach. Its 1.25 multiplier was a chosen scoring weight, not a measured 25% increase in actual impact.
Current version: det-v1. The current rubric considers specificity, quantification, scope, structure and language. Editor checks are deterministic, with partly heuristic measures. Server grading can use an LLM for impact scope, with a keyword fallback.
These methods organize feedback; they are not universal measures of writing quality. Their scores are not directly comparable, and this illustration is not scored.
Full example transcript
Original
Helped the training team improve onboarding.
Starting point: the contribution and measured change are not yet specified.
Clarified
I designed a checklist. Eight instructors used it to onboard 60 new hires. Average setup time fell from 38 to 25 minutes over four weeks.
New fictional facts supplied for this illustration: checklist ownership, users, reach, measured times and period. These details were not inferred from the original.
Refined
Designed a checklist used by eight instructors to onboard 60 new hires; average setup time fell from 38 to 25 minutes over four weeks.
Wording only: the same supplied facts, joined into a concise statement. The measured change does not establish that the checklist caused it.
06 / Product implications
Keep the writer in control
These decisions describe the implementation reviewed locally as of 7 September 2026, not a claim that every change has been released to production.
- Evidence questions
- Retain source material and ask neutral follow-ups. A stronger statement may need a fact only the writer can supply.
- Measured field fit
- Show the count and the fit state separately from the writing grade. A short draft must not be presented as meeting a near-limit target.
- Reviewable proposals
- Keep proposed changes distinct from accepted writing. Newer user edits must not be overwritten by a delayed response.
- Statement Bank and resume reuse
- Save reviewed accomplishments for later use. Adapting one to a new role should change its emphasis, not turn the desired role into past employment or add unsupported reach.
Air Force evaluation drafting is an optional application of this method. It is not a universal character-limit rule or the default journey for every veteran. Future connected-evidence workflows are outside this article's implementation claims. Local verification scope [1]
07 / Limits and next experiments
What remains unknown
We have not established a representative exact-length success rate, universal semantic fidelity, or a measured improvement in hiring outcomes. Historical summaries lack the run-level evidence needed to reproduce their quantitative claims. Synthetic walkthroughs do not substitute for recruited user research, and deterministic rules can reward superficial textual signals.
Model behavior may vary with the input, prompt and version. Software regression tests establish bounded behavior under their fixtures, not the truth of an accomplishment. Human factual review remains necessary; retaining source material makes that review possible but does not automate it.
Next protocols, not completed results
Planned / Not run
Paired evidence and field-fit evaluation
Question. Does measured revision improve field fit without increasing unsupported claims?
Freeze a synthetic input set, prompt text, model version and counting function before comparing conditions. Retain every attempt, refusal, unsupported addition, timeout and missing-evidence response. Report unique inputs separately from repeated trials, with row-level counts and matched conditions.
Not run. Sample size, model and prompt versions: Not recorded. No result contributes to the register above.
Planned / Not run
Human reading and resume reuse
Question. Can readers identify the contribution and its limits, and can writers retain control during reuse?
Recruit consenting participants under a separate protocol; define audience and rating criteria in advance. Compare source fidelity and readability using blinded, randomized presentation where feasible, retain disagreements, and evaluate whether people can inspect and reuse their own evidence.
Not run. Recruitment, sample size and rating records: Not recorded. This is not a claim about hiring, promotion or award outcomes.
08 / References and change history
Sources, with their limits attached
These are Narrative Pro's own curated primary development summaries, not external scientific validation. The public record summarizes internal records rather than exposing raw data. Study-specific links above identify the relevant entries.
- Narrative Pro. Methodology: curated evidence summaries. Reviewed 7 September 2026. Writing observations, synthetic walkthroughs, engineering verification and prospective protocols.
- Narrative Pro. Scoring history and det-v1 implementation summary. Reviewed 7 September 2026. Historical definitions and current heuristic boundaries, not a validation study.
- Narrative Pro. Fictional illustration provenance. Reviewed 7 September 2026. Authored teaching example; no customer or experimental outcome.
How to interpret the source record
The source file is a deliberately sanitized Markdown document. Original historical run artifacts were not located or reconstructed in the bounded publication review. Unknown fields remain unknown; no raw personal records, inferred denominators or retrospective model-version guesses are included.
Change history
/ First curated article version. Separates evidence categories, documents C6 and det-v1, retains reported failures, and labels the worked example as fiction. Unsupported benchmark rates and exact historical output counts are withheld.
This review date is not a production publication date. Independent review and release approval are separate from this account of implementation evidence.