The Free Knowledge Library

The proposed knowledge-library function would build, maintain, and freely distribute a continuously updated resource for Ora and the broader public, using infrastructure that could remain replicable and publicly accessible rather than concentrated on one institution’s servers.

Current status. The Ora project publishes selected public-domain artifacts and framework material. A continuously updated Foundation knowledge library, its committees, review system, decentralized nodes, signing system, and intake process do not currently exist. The earlier seven-component proposal and the separate provisional Passion/Operation model remain unresolved.

The proposed library responds to two needs: Ora needs free knowledge to access, and the public benefits from a carefully sourced resource available to anyone. Replicability, distribution, and verifiability are design goals, not claims about a deployed Foundation system.

Constitutional principles

The proposed library would be guided by principles such as these:

  • Provenance-weighted, source-traceable output. Every claim is connected to its source. Source reliability is evaluated and weighted, not asserted.
  • Public-domain output. A future library would aim to keep published datasets freely available and in the public domain. It should not become a leverage point because it cannot be enclosed.
  • No editorial bias by processing. The algorithms execute specifications. They do not make editorial judgments. Where the specification does not clearly address content, the content is flagged for human review rather than processed by guess.
  • Public processing specifications. The specifications that govern each domain are public documents, open to inspection by anyone. Anyone can see how a piece of content came to be in the library, against what criteria, with what provenance weight.
  • Automation serves specification, not the reverse. If an algorithm cannot faithfully execute a specification, the algorithm is fixed or the specification is amended by its committee. The algorithm does not silently approximate.
  • No monetization of user interaction data. User retrievals, queries, and behavior are not monetized, sold, or shared. The library is infrastructure, not surveillance.

A proposed four-layer design

Domain-specific specifications authored by subject-matter experts, executed faithfully by automated processing, with review for edge cases. This is a design pattern, not a claim that the layers are staffed or operating.

  • Layer 1 — Constitutional principles (the six above) govern all library work.
  • Layer 2 — Expert committees author the specifications that govern each domain. What qualifies as a source, how provenance is weighted, how content is processed and atomized, what edge cases require human review.
  • Layer 3 — Judicial review handles edge cases flagged by users or algorithms. Rulings become precedent for future specification amendments by the relevant committee.
  • Layer 4 — Algorithms execute specifications faithfully, deterministically, and auditably. Every algorithmic decision is traceable to the specification that authorized it.

No formal committee or judicial-review process is operating at present. If the function is later established, staffing and review arrangements would be decided at that time.

The universal pipeline

The proposed knowledge work would use the same processing pattern across domains. The pattern is the infrastructure; only the specifications would differ.

Source identification → provenance verification → processing (any document → atomic notes → structured output) → cross-referencing → indexing → publication. Each domain’s specification document governs steps 1 through 3. Steps 4 through 6 are infrastructure-level and domain-agnostic.

Decentralized infrastructure

The proposed infrastructure could use independent nodes, content-addressed storage, and cryptographic provenance verification. No Foundation-operated nodes, partner hosting network, IPFS deployment, or signing authority is claimed here.

In the design, a future Foundation would author specifications rather than make itself the only host. The intended result is a library that could remain available beyond any one institution; that result has not yet been implemented.

Phasing

Phase 1 — easy public-domain material. Encyclopedia (Wikipedia equivalent), source library (Wikisource equivalent), dictionary/lexicon (Wiktionary equivalent). The Level 3 provenance base that makes Ora useful on day one. Source material already digitized, already in the public domain, already in formats that can be processed without negotiation.

Phase 2 — news and current events. The library pipeline applied to current events, solving the AI training-cutoff problem with continuously updated provenance-verified news. Initial scope: US national news. Journalistic standards apply: source verification, multi-source corroboration, provenance tracking on all claims.

Phase 3 and beyond. Textbooks (Wikibooks equivalent), courses (Wikiversity equivalent), government data (Federal Reserve papers, census data, regulatory filings, congressional records, agency reports).

Deferred domains. Legal, medical, museum/cultural heritage, patents, and literature/film. Each presents specific challenges beyond the standard pipeline. Deferred not because they are unimportant but because their processing frameworks have not yet been designed.

  • Main Street Independent — an existing, separately operated publication that uses Ora’s framework methodology on a daily news beat. It is evidence of the methodology in use, not a Foundation program or legal relationship.