# Breton sentence-witness-first recovery · V11.7 method report **Status:** strict EAF content-gated public teaching-review edition; native-speaker/community review required; no automatic continuity claim, etymology, or V12 lexical admission. ## July 22 V11.7 EAF realignment and visible-evidence correction The Atlas must present the sentences, not merely report that sentences exist. In the v0.8.2 interface, the default **Then & Now · Every Sentence** view renders every release-authorized card for the selected Breton record. No separate sentence tab, top-N slice, collapsed sentence body, or hidden remainder stands between the reader and the evidence. Each card always contains three explicit language slots: the full attested Garifuna sentence, source-supplied English when available, and source-supplied Spanish when available. A missing translation remains visible as **TRANSLATION UNAVAILABLE — NO SENTENCE INVENTED**. The required sentence estate remains exact: - Direct EAF/ELAR: **10,964** - Verified sentences: **3,765**, including the **624-record Living Dictionary subset** without double weighting - Exact Suazo/LILA A–Y examples: **11,838** - Required unique union: **26,567** - Breton records: **7,029** V11.7 begins every EAF decision with the immutable V2 realignment ledger: - **47,912** EAF token occurrences - **12,482** literal token types - **20,028** exact raw-source-headword occurrences in **8,152** units - **34** certified morphology overlays - **21** morphology-only occurrences in **12** units - **2,798** unresolved units and **2** numeric-only units - **34,839** comparison candidates, all comparison-weight zero The Atlas junction retains **48,231** typed EAF assertions in **45,960** EAF record/unit cards. It does not discard difficult or conflicting evidence. It separates six dimensions that earlier builds had conflated: 1. exact raw-source headword identity; 2. source lexeme-to-sentence sense support; 3. certified present-day morphology; 4. Breton record-to-modern-target relation; 5. present-day teaching authorization; and 6. lexical admission or historical continuity. An exact EAF token may therefore have form-identity weight without Breton-relation weight. A source dictionary lexeme and its translated sentence may agree without proving that a meaning-retrieved modern target belongs to a particular Breton record. A certified root-plus-operator analysis may carry morphology weight while its Breton relation stays zero. Historical continuity and lexical admission remain zero throughout this automated pass. ### Why V11.6 was blocked Independent audit found that V11.6 still allowed common words in explanatory Breton prose to create thousands of apparent EAF relations: **3,230** accepted assertions through *because*, **1,471** through *all*, and **1,344** through *but*. Of **8,893** formerly accepted EAF assertions, **8,889** depended only on `MEANING_NOMINATED_HEADWORD_REVIEW`; just **4** had a direct record-target relation. V11.6 was therefore marked **DO NOT PUBLISH OR DEPLOY** even though its structural counts and deterministic output were internally consistent. V11.7 retains every one of those exact-form nominations for review but gives a meaning-nominated target zero Breton-relation, teaching, and independent record-level weight. Only a `DIRECT_SOURCE_RELATION_REVIEW` target plus a content-bearing four-way intersection may carry Breton-relation weight. The intersection is: `Breton sense ∩ modern target sense ∩ raw source lexeme sense ∩ source sentence proposition` Function words, discourse markers, common quantifiers, and generic retrieval terms cannot satisfy the content gate alone. English and Spanish generic terms stay in separate `EN:` and `ES:` namespaces; only explicit reviewed `SEM:` groups can cross languages. Spanish subjunctive `sea` remains a stopword and cannot become English *sea*. The fresh-water/sea contradiction gate continues to prevent `barana` from entering Breton record 712. After this correction, **8,859** assertions retain source lexeme-to-sentence sense weight, representing **1,163** unique strict V2 source assertions. Only **4** EAF assertions receive Breton-relation and independent record-level weight: one direct `bira`/sail relation and three direct `iri`/name relations. The formerly dominant *because*, *all*, and *but* routes now contribute **zero** accepted EAF relations. ### The `barana mau wabugurate` correction Unit `eaf:Lily_Zapata_2/CAB_LilyZapata_2.eaf@a1068` preserves the archival EAF surface **barana mau wabugurate** and source-supplied English **the ocean has taken our rudder** as immutable source layers. It also presents MARAGAÑ's native-speaker correction **BARULA BARANA** with the gloss **rudder** directly beside the archival material. The token roles remain separate: - observed `barana` is linked only to its independently attested sea/ocean dictionary sense; - observed `wabugurate` remains a repeated contextual lexeme candidate with rudder/helm controls such as `abugura`, `abuguragülei`, `simunu`, and `yadúi`; - sentence-level translation is not treated as word-by-word glossing; - `barüla duna` is displayed only as a comparative water-taking predicate, not as independent dictionary proof that `BARULA` or `BARULA BARANA` is the rudder headword. Because the native-speaker decision marks the archival transcription **DISPUTED_NOT_VERIFIED_MODERN_FORM**, all **50** public `a1068` cards remain visible but carry zero source-sentence, Breton-relation, teaching, and independent record-level weight. These consist of **46** exact `barana` candidates and **4** contextual `wabugurate` cards. No card assigns “rudder” to `barana`. All six `wabugurate` occurrences remain in the research ledger. Five retain rudder/steering sentence signals; the sixth retains its raft conflict. Every contextual/root comparison carries zero form-identity, morphology, sentence-sense, Breton-relation, teaching, continuity, and admission weight. ### Certified morphology retained without historical inflation V11.7 restores **63** `nirahu → irahü` morphology assertions across all **21** observed `nirahu` units and the three exact `irahü` Breton comparison targets at records 4516, 4517, and 5815. These routes carry morphology weight only. The broader labels *child* and *son* no longer cause the certified modern morphology to disappear, but they also cannot manufacture a Breton sense relation. ### Public projection The rights-filtered V11.7 public review projection contains **15,605** complete sentence units and **84,308** record/unit cards: - **16,244** accepted record-level sentence witnesses; - **68,064** visible zero-weight review candidates; - **50** public EAF cards, **0** accepted; - all authorized untranslated Suazo examples remain visible and carry zero accepted same-sense weight. Record 17 remains a direct completeness control: **59** full sentence cards, **8** accepted routes, and **51** zero-weight candidates. The physical-pot/jackpot and financial-*broke* collisions remain rejected. Every public card exposes its full source sentence, language slots, observed form, source headword and dictionary senses, declared comparison target, target relation state, assertion method, separate weights, explanation, limitation, teaching note, attribution, rights, publication decision, lineage, and review status. Counts are navigation metadata only; they never substitute for the evidence. The earlier phrase-first research architecture remains below as preserved lineage. Its former **6,311 Suazo-example** and **27,681 sandbox-unit** counts describe that predecessor prototype and are superseded for the required four-source sentence-witness denominator by the V11.7 figures above. ## Governing rule Garifuna speakers and attributed Garifuna usage are the voice of this work. Historical dictionaries, Taylor's scholarship, Breton, missionary New Testaments, and Catholic liturgical texts may test or corroborate a community-led reading. They cannot originate, replace, outvote, or certify it. This is implemented as a computation boundary, not only a display warning: 1. Primary outcomes and form links are calculated only from the community/EAF, verified-sentence, eligible prose, and direct Suazo lanes. 2. Henderson Karif and Taylor controls are computed separately and are non-voting. 3. Historical Catholic/Christian witnesses are computed separately and are non-voting. 4. V12 identity and attestation records are read-only lexical grounding. 5. Sound or vector similarity expands hypotheses only. 6. No historical-only result may create a modern reading. This design follows community-led revitalization and Indigenous data-governance principles: Indigenous leadership and authority must govern language work, not merely be consulted after an outside system has made its decision. See [UNESCO on Indigenous leadership in revitalization](https://www.unesco.org/en/articles/international-year-indigenous-languages), [UNESCO on community-driven language technology](https://www.unesco.org/en/articles/ifap-advocates-community-driven-technology-indigenous-language-revitalisation), and the [CARE Principles for Indigenous Data Governance](https://www.gida-global.org/careprinciples). ## Evidence lanes | Lane | Current prototype role | May affect primary decision? | Public status | |---|---|---:|---| | Attributed native-speaker analysis | Governs the question and candidate pathways | Yes | Attribution required | | Direct EAF/community phrases | Primary phrase evidence | Yes | Item-specific rights review | | Verified attributed sentences | Primary phrase evidence | Yes | Source-specific review | | Eligible non-scripture prose | Primary phrase evidence | Yes | Source-specific review | | Direct Suazo examples | Headword-attached usage context; exact example sense unresolved | No accepted sense weight without an attributed example-to-sense decision | Permission/scope review retained | | V12 | Exact identity and attestation spine | Grounding only | Read-only here | | Henderson Karif 1873 | Bay-of-Honduras historical lexical control | No | Manuscript is historical; transcription QA retained | | Taylor 1951 British Honduras | Sound law, morphology, lexical and ethnographic control | No | Private research; body-text publication not cleared | | 1857 Matthew, 1901 Mark, 1902 John | Dependent historical Christian phrase corroboration | No | Open-domain candidate; OCR review required | | Genon 1871 Catholic Mass | Historical liturgical form-attestation control | No | Public-domain-age candidate; OCR review required | | Restricted JW/Wycliffe/modern NT | Excluded | No | Zero body-text rows admitted | | Nisamina oral/sentence geometry | Candidate expansion | No | Private research only | The archive identity of Henderson's source is independently described by the University of Pennsylvania: an English–Garifuna/Karif dictionary documenting the language around the Bay of Honduras, written in 1873 from earlier material. The archive also notes revised entries and usage examples. See [UPenn OPenn, Ms. Coll. 700 Item 127](https://openn.library.upenn.edu/Data/0002/html/mscoll700_item127.html). ## Why phrase-first is the right retrieval unit Breton frequently records clauses, inflected predicates, possessive forms, and compact constructions rather than isolated modern citation headwords. The engine therefore searches for the independently translated proposition first, returns the attributed Garifuna utterance, and only then asks which observed forms or predeclared candidate families are present. Comparable-corpus bilingual lexicon induction supports this two-stage approach: comparable contexts can generate candidates, but dictionary induction remains noisy and should be evaluated separately from exact bitext alignment. See [Laville et al. 2020](https://aclanthology.org/2020.coling-main.527/), [Irvine and Callison-Burch 2017](https://aclanthology.org/J17-2001/), and [Shi et al. 2021](https://aclanthology.org/2021.acl-long.67/). The output schema follows TEI's core distinctions even though this sandbox emits JSON: form, sense, cited example, translation, source, responsibility, and certainty remain separate. TEI explicitly supports separate homographs/senses, cited dictionary examples, and responsibility/certainty. See [TEI dictionary entry examples](https://www.tei-c.org/release/doc/tei-p5-doc/en/html/examples-entry.html), [`cit` for sourced examples](https://tei-c.org/release/doc/tei-p5-doc/en/html/ref-cit.html), and [`certainty`](https://www.tei-c.org/release/doc/tei-p5-doc/en/html/ref-certainty.html). ## Sound correspondence control Taylor's 1951 analysis at PDF pages 57–58 is directly relevant to the director's c/k→g observations. Taylor describes Breton/Island Carib `k` as corresponding in Central American Carib to: - `k` in a restricted set, especially word-initially; - `g`, especially word-initially; or - `b`, especially medially. The engine therefore records **conditioned candidate pathways**, not a universal character replacement. This supports testing forms such as `águra` and `gamalaliti` while preventing unsafe conversion of every historical `c` or `k`. ## Preserved predecessor sandbox corpus and gates · superseded for V11 completion - Predecessor total indexed units: **27,681** - Direct EAF: **10,964** - Verified sentences: **3,765** - Eligible non-scripture prose: **1,126** - Predecessor direct Suazo examples: **6,311** (superseded by the exact **11,838-example** A–Y estate) - Historical open-domain NT verse alignments: **2,167** - Genon 1871 Catholic Mass OCR lines: **344** - Henderson manually transcribed Karif controls: **117** - Henderson first-pass HTR controls: **2,887** - Restricted JW/Wycliffe/modern NT rows: **0** - Missing attribution/reference rows: **0** - SQLite `quick_check`: **ok** - Historical controls appearing in primary decision results: **false** The 117 manually transcribed Karif rows and 2,887 first-pass HTR rows remain separate. The latter are useful for recall but may not be presented as cleaned authority without manuscript review. ## What the six tests now show ### `Loubáagnem` The community reading “beside him / at his side” has current phrase evidence and a predeclared form link (`oubawagu` family). Taylor independently lists possessive `-auba` “side.” The punishment reading retrieves punishment semantics but no linked candidate form. The system therefore preserves both semantic lanes while recording that the community reading currently has the stronger form-and-phrase chain. ### `Machoucanrou naraoüani` Current EAF/verified evidence links the machete/cutting interpretation to `náwüri/nauri` and the cut family. Henderson gives “axe” with `hara` as a Karif variant; Taylor records `haraua` “axe,” `isubara` “machete,” and possessed `-áuori`. The open-domain Matthew witness also contains `harawa` “axe” in an axe/cutting verse. These controls validate the user's warning that axe and machete are closely related but must remain semantically separate source meanings. ### `Abíchata` The community alcohol/intoxication reading has current phrase and form-family evidence through `abácharuada`. Taylor separately records `binu` “rum/alcohol,” `bacarua` “drunk,” throat vocabulary, and burning vocabulary. The straw lane has only component evidence, not phrase proof. The lightning lane has lightning semantics but no `abicha` form link. Unrelated `bata` material is labeled as a homograph/unrelated context instead of being counted as support. ### `ácoura` Current EAF phrases explicitly express “throw away” with `águra`-family forms. Henderson's manually transcribed Karif entry independently gives `agurajaru` for “away, throw.” This is a strong historical control for the director's c/k→g pathway, but the current community phrases remain the primary proof. ### `camaláliti` Current evidence links `gamalaliti` to the `umalali` voice/sound family. Henderson gives `humalale/humalalegu` “voice”; Taylor gives possessed `-umalali` “voice”; and aligned open-domain John verses use `umalali/numalali` in phrases translated with “voice” or “sound.” This is the cleanest demonstration of the full architecture: community proposition first, then V12 identity, then regional historical controls, then dependent religious corroboration. ### `Abácharacauti` Current phrases link warm/hot semantics to `bachatu`, `süti`, and related warm forms. Taylor directly gives `abacaha` “to warm” and `bacá` “warm.” Henderson HTR supplies additional warm entries, but they remain first-pass review evidence. The broader distributed/attenuative morphology is still not proved by the current sentence set. ## Catholic and NT handling The 1871 Genon OCR contains a bounded Catholic Mass sequence beginning under `LIDAN LEMES`, including confession, Kyrie, Gloria, Credo, offering/consecration material, Agnus Dei, and Communion. Because the OCR is Garifuna/Latin/Spanish line material without a reviewed section-level bilingual alignment, it is presently used only to attest candidate forms at exact PDF page and OCR line. The NT lane admits only aligned rows with nonblank historical Garifuna, public-domain English, and public-domain Spanish from the enumerated complete-candidate sources: - Matthew 1857: 950 aligned verses - Mark 1901: 546 aligned verses - John 1902: 671 aligned verses Internet Archive identifies the 1857 Matthew and 1902 John files and states that the contributing organization is unaware of copyright restrictions. See the [Internet Archive source record](https://archive.org/details/CABPOR_DBS_HS). Google distributes the 1901 Mark edition free and identifies its date and publisher. See [Google Play Books](https://play.google.com/store/books/details/Uganu_buiditi_kaysi_St_Mark_lidan_garifuna?hl=en_US&id=FUZEAQAAMAAJ). These editions may corroborate phrase/form use, but they are institutionally and genealogically dependent historical witnesses. They do not count as independent modern community speech. ## Known limitations and continuing review 1. Taylor's local *Conversations and Letter from the Black Carib of British Honduras* extraction is glyph-corrupted and was inventoried but not admitted. It needs clean page-image recovery and bilingual segmentation. 2. Henderson's 2,887 HTR rows need manuscript-image review; only 117 rows currently have a distinct manual-transcription lane. 3. Taylor's 1951 OCR has page/line geometry but two-column line order can interleave neighboring entries. Context windows remain review evidence, not automatic lexical alignment. 4. Religious sources may share a missionary lineage with Breton. Their agreement cannot be treated as independent replication. 5. Phrase retrieval supplies evidence candidates; morphology, sense identity, and etymology still require speaker/elder adjudication. V11 now stores a sentence-witness graph per Breton record with separate source unit, headword/sense target, observed form, candidate frame, lineage, review, and publication decisions. Continuing work must add native-speaker decisions and recurring correspondence sets without flattening any node into one confidence percentage.