π§ Indexa Knowledge Base β Strategy & Policy
What this document is. The constitution for Indexaβs content knowledge base. It defines how we mirror the 82 published articles into Obsidian, how we decompose them into a reusable peptide knowledge graph, how that graph is governed and ingested into Supabase, and how downstream channels (TikTok, Threads) borrow from it without ever polluting it. Read this before building anything. Once approved, every scraping run, note, and content piece must comply with it.
1. North Star
Build one canonical, interlinked body of peptide knowledge that is true, traceable, and reusable β so that:
- Customers and researchers can learn from it (surfaced via the site / Supabase).
- Every social asset (TikTok, Threads, captions, emails) is a derivative of it, never an original source.
- The knowledge compounds: each new paper, anecdote, or protocol enriches existing nodes instead of spawning duplicates.
The governing principle: The knowledge base states what is true. Channels decide how to say it. Truth lives in one place; tone, voice, and platform packaging are applied on the way out, never baked in.
2. Three-Layer Architecture
LAYER 1 β SOURCE MIRROR LAYER 2 β KNOWLEDGE GRAPH LAYER 3 β CHANNEL OUTPUT
(faithful copy of the 82 (atomic, deduplicated, (TikTok / Threads / email β
published articles) evidence-tagged truth) tone-manipulated derivatives)
Article notes βββextractβββΆ Peptide / Mechanism / Paper ββmanipulateβββΆ Content posts
(1:1 rebuild) Protocol / Concept notes (alias brands, hooks)
βββββββββββββ Obsidian = source of truth βββββββββββββ
β ingest
βΌ
Supabase (read model)
- Layer 1 β Source Mirror. A faithful 1:1 rebuild of each published article. This is the provenance anchor β proof of what we already say publicly, and the raw material for extraction. We do not editorialise here.
- Layer 2 β Knowledge Graph. Atomic notes (one peptide, one mechanism, one paper, one protocol = one note) that articles link into. This is the reusable, dedupeβd truth that gets enriched over time and ingested to Supabase.
- Layer 3 β Channel Output. Lives in the existing
Content/pipeline, not in the knowledge base. Pulls facts from Layer 2, then applies channel tone/policy. Covered by Content β Playbook.
Hard rule: Layers 1 and 2 are the source of truth. Layer 3 never writes back into them. A TikTok hook is not allowed to change what a peptide note says.
3. The Hybrid Note Model
Every piece of knowledge is one of six note types. The first two are Layer 1; the rest are Layer 2.
| Type | Tag | One note = | Purpose |
|---|---|---|---|
| Article | #kb-article | one published blog post (1:1 mirror) | Provenance + the assembled, customer-facing narrative |
| Resource | #kb-resource | one site resource/guide (dosage, reconstitution, storage, dictionary) | Mirror of evergreen tool/guide pages |
| Peptide | #kb-peptide | one compound (e.g. BPC-157) | The canonical entry for a molecule β the hub everything links to |
| Mechanism | #kb-mechanism | one biological pathway/action (e.g. angiogenesis, GH secretagogue) | Cross-functional concept reused across peptides |
| Paper | #kb-paper | one study / citation (PubMed/PMC) | Evidence node β the acute breakdown of a single paper |
| Concept | #kb-concept | one reusable idea (e.g. reconstitution, COA, half-life, stacking) | Glossary-grade building blocks |
| Transcript | #kb-transcript | one external source (podcast/talk/interview transcript) | Layer-1 anecdote source β provenance: anecdote; its claims thread into peptide notesβ Anecdotes/Protocols sections |
| Source | #kb-source | one reputable external reference (journal/conference/oncology network) | Higher-tier external evidence (provenance: literature) β e.g. ASCO/JAMA findings; cited like transcripts but weighted above anecdote |
| Supplement | #kb-supplement | one non-peptide nutritional adjunct (CoQ10, NAD, methylene blue, glycine) | Adjunct entity β dose/route/timing/intake/mechanism; linked from stacks |
| Compound | #kb-compound | one non-peptide therapeutic (MK-677, orforglipron, enclomiphene, clomiphene, cardarine) | Non-peptide drugs (SERMs, small-molecule GH/GLP-1 agonists, etc.) that act on the same mechanisms as peptides but arenβt peptides β kept honest rather than mislabeled kb-peptide |
| Behavior | #kb-behavior | one lifestyle/behavioral intervention (sauna, cold plunge, Zone 2, HIIT, fasting, sleep, sunlight) | Behavioral layer β the non-pharmacological inputs that compose into stacks alongside peptides/supplements |
| Stack | #kb-stack | one combination protocol (peptides + supplements + behaviors) | Layer 3 (applied) β synergy rationale, timing, intake, cycling, cautions, monitoring |
| Method | #kb-method | one cross-cutting methodology/reference (e.g. animalβhuman dose conversion) | Reference playbooks for working with the data β the convention for turning animal-study doses into traceable human-equivalent breakdowns |
Why hybrid (vs. 1:1-only or fully-atomic): the article mirror preserves SEO structure and provenance; the atomic layer is what makes the base a graph β reusable, enrichable, and clean to ingest. Articles become thin assemblies that link to fat, canonical concept notes. One fact (e.g. BPC-157βs angiogenesis mechanism) is written once in the Peptide/Mechanism note and referenced by every article that touches it.
Anti-duplication law: if a fact belongs to a peptide, mechanism, paper, or concept, it is written in that atomic note and linked β never re-typed inside an article. Articles narrate and contextualise; atomic notes hold the truth.
4. Folder Structure
New home: 01 - Distribution/Knowledge Base/. This is the βcompile all content pillarsβ folder.
01 - Distribution/
βββ Knowledge Base/
βββ Knowledge Base β Strategy & Policy.md β this file
βββ _Templates/
β βββ _Template β Article.md
β βββ _Template β Peptide.md
β βββ _Template β Mechanism.md
β βββ _Template β Paper.md
β βββ _Template β Concept.md
βββ Articles/ β Layer 1: 86 mirrored posts
βββ Resources/ β Layer 1: dosage, reconstitution, storage, dictionary
βββ Transcripts/ β Layer 1: external anecdote sources (podcasts/talks)
βββ Sources/ β Layer 1: reputable external references (journals, conferences, oncology networks)
βββ Peptides/ β Layer 2: one note per compound
βββ Mechanisms/ β Layer 2: pathways / actions
βββ Papers/ β Layer 2: cited studies
βββ Concepts/ β Layer 2: glossary building blocks
βββ Supplements/ β Layer 2: non-peptide nutritional adjuncts (CoQ10, NAD, β¦)
βββ Compounds/ β Layer 2: non-peptide therapeutics (MK-677, orforglipron, SERMs, β¦)
βββ Behaviors/ β Layer 2: lifestyle interventions (sauna, cold, Zone 2, HIIT, fasting, sleep, sunlight)
βββ Stacks/ β Layer 3: applied protocols (peptides + supplements + behaviors)
βββ Methods/ β cross-cutting reference methods (animalβhuman dose conversion, β¦)
βββ Bases/
βββ Articles.base
βββ Peptides.base
βββ Papers.base
Channel output stays where it already lives: 01 - Distribution/Content/ (Layer 3). The knowledge base feeds it; they do not merge.
5. Note Schemas (frontmatter)
Aligned to Data Model & Conventions golden rules: frontmatter is the data, filename = ID, dates are ISO, money is numeric, cross-links are wikilinks, tag drives the Base.
Article β #kb-article
type: kb-article
id: ART-bpc-157-complete-guide-2026
title: "BPC-157 Complete Guide 2026: Benefits & Dosing"
slug: bpc-157-complete-guide-2026
source_url: https://indexalabs.com/blog/bpc-157-complete-guide-2026-benefits-dosing
published: 2026-03-10
cohort: seo-2026 # or product-2025
pillar: education # education | product | social-proof | promo | research
primary_peptide: "[[BPC-157]]"
peptides: ["[[BPC-157]]"]
mechanisms: ["[[Angiogenesis]]", "[[Cytoprotection]]"]
papers: ["[[PMID-12345678]]"]
products: [bpc-157] # storefront slugs (Supabase), NOT vault notes β plain strings, never wikilinks
target_keyword: "bpc-157 dosing"
scrape_status: mirrored # pending | mirrored | extracted | verified
content_hash: "" # sha256 of source body, for re-scrape diffing
tags: [kb-article]Peptide β #kb-peptide (the hub note)
type: kb-peptide
id: BPC-157
title: "BPC-157"
aliases: ["Body Protection Compound-157", "PL 14736"]
category: tissue-repair # matches site category taxonomy
sequence: "GEPPPGKPADDAGLV"
molar_mass: 1419.5
half_life: "~4 hours (subcutaneous, est.)"
mechanisms: ["[[Angiogenesis]]", "[[Cytoprotection]]"]
stacks_with: ["[[TB-500]]"]
papers: ["[[PMID-12345678]]"]
product: bpc-157 # storefront slug (Supabase), NOT a vault note β plain string
forms: [] # extra product slugs that are delivery-forms of this compound
is_blend: false # true for blends; then list components as wikilinks
components: [] # for blends only: ["[[BPC-157]]", "[[TB-500]]"]
evidence_tier: B # S | A | B | C | D (Research-pillar tier list)
research_use_only: true
tags: [kb-peptide]Mechanism β #kb-mechanism
type: kb-mechanism
id: Angiogenesis
title: "Angiogenesis"
domain: vascular # vascular | metabolic | neuro | immune | hormonal | dermal
peptides: ["[[BPC-157]]", "[[GHK-Cu]]"]
papers: ["[[PMID-12345678]]"]
tags: [kb-mechanism]Paper β #kb-paper (evidence node)
type: kb-paper
id: PMID-12345678
title: "Pentadecapeptide BPC 157 and tendon healing"
authors: "Sikiric P, et al."
year: 2014
journal: "J Orthop Res"
doi: "10.1002/jor.xxxxx"
pmid: 12345678
url: https://pubmed.ncbi.nlm.nih.gov/12345678/
model: animal # in-vitro | animal | human-rct | human-obs | review
peptides: ["[[BPC-157]]"]
mechanisms: ["[[Angiogenesis]]"]
evidence_weight: moderate # strong | moderate | weak | anecdotal
tags: [kb-paper]Concept β #kb-concept
type: kb-concept
id: Reconstitution
title: "Reconstitution"
related: ["[[Bacteriostatic Water]]", "[[Storage & Handling]]"]
tags: [kb-concept]Enrichment fields (added over time, never overwriting source): atomic notes carry free-text body sections for ## Anecdotes, ## Physician protocols, ## Practical notes, each line provenance-tagged (see Β§9).
6. Naming & ID Conventions
Extends the existing ID table in Data Model & Conventions:
| Entity | Pattern | Example |
|---|---|---|
| KB Article | ART-<slug> (filename = article title) | ART-bpc-157-complete-guide-2026 |
| Peptide | <COMPOUND> (filename = common name) | BPC-157 |
| Mechanism | <Name> | Angiogenesis |
| Paper | PMID-<id> or DOI-<id> | PMID-12345678 |
| Concept | <Name> | Reconstitution |
Rules: filename matches the human-readable title so [[BPC-157]] resolves naturally; id in frontmatter is the stable machine key for Supabase. Slugs mirror the live URL exactly so we can round-trip to the site.
7. Interlinking & SEO Policy
The link graph is the SEO asset. Rules:
- Every article links to its primary peptide and to every mechanism/paper it cites. No orphan articles.
- Hub-and-spoke: Peptide notes are hubs. Articles, papers, and mechanisms spoke into them. A peptide hub should be reachable in β€2 hops from any related note.
- Reciprocal links: if an article links a peptide, the peptideβs
peptides/related list need not list every article (Bases derive that), but mechanismβpeptide and peptideβpaper links are maintained on both ends. - Cluster topology mirrors site categories (GLP-1/Metabolic, Growth Hormone, Tissue Repair, Nootropics, etc.) so internal-link clusters map to the siteβs category silos β this is what compounds ranking.
- Comparison articles (βX vs Yβ, βbest peptides forβ¦β) must link both/all peptides and the relevant mechanism β these are the highest-value link hubs.
- Anchor text = descriptive, never βclick hereβ. Wikilink display text should read as the target keyword where natural.
- No dead links. A
[[wikilink]]to a not-yet-created atomic note is allowed (it marks a stub to fill) but stubs are tracked and burned down before ingestion.
8. Scraping Methodology & Policy
Source of record: the live site sitemap (indexalabs.com/sitemap.xml) β already confirmed to list 82 /blog/ URLs plus resources and product pages. The sitemap is the canonical work list, not the rendered /blog index.
The 82 split into two cohorts (drives cohort field and review priority):
product-2025(~40 posts, lastmod 2025-01) β one-per-product research guides aligned to SKUs.seo-2026(~42 posts, lastmod 2026-03) β comparison, βbest-ofβ, dosing, and delivery-method articles built for search.
Method (per article):
- Pull the URL list from the sitemap; write a manifest (
Bases/Articles.baseseeded asscrape_status: pending). - Fetch each article. The site is partly client-rendered, so the pipeline must render JS (browser-based fetch) rather than raw HTML when the body is empty.
- Convert HTML body β clean Markdown (headings, lists, tables preserved; nav/footer/boilerplate stripped).
- Write the Article note (Layer 1) with full frontmatter,
source_url, and acontent_hashof the source body. - Set
scrape_status: mirrored.
Extraction (Layer 1 β Layer 2), per article:
- Identify the peptide(s), mechanism(s), and any cited papers. Create/append the atomic notes (donβt duplicate β link if they exist).
- Replace inline factual claims in the article with links to the atomic notes that own them.
- Set
scrape_status: extracted, thenverifiedafter the Β§9 review.
Re-scrape policy: on a scheduled cadence, re-fetch and compare content_hash. If changed, flag the article for re-mirror + re-extract so the base never drifts from the live site.
Politeness & integrity: scrape our own site only, throttle requests, preserve published dates, never invent content that isnβt on the page. Mirrors are faithful; interpretation happens only in Layer 2 with citations.
9. Provenance, Enrichment & Citation Policy
The baseβs value is trust, so every claim is traceable.
- Three provenance classes, tagged on each enrichment line/section:
[indexa]β already published by us (from a mirrored article).[literature]β backed by a[[PMID-β¦]]paper note (Research pillar).[anecdote]/[protocol]β community report or physician protocol; explicitly marked as such, never presented as established fact.
- No unsourced claims enter Layer 2. If it canβt be tagged to one of the above, it doesnβt get written.
- Evidence tiers (
evidence_tierSβD on peptides,evidence_weighton papers) power the βMost Researched Peptides Tier Listβ Research-pillar asset and let channels signal confidence honestly. - Stack backing (mandatory): every synergy, dosing, timing, intake, and cycling claim in a
kb-stackmust trace to a[[paper]](literature) or a named applied-protocol source ([[transcript]], e.g. a physician). The stackβsbacked_by:frontmatter lists them, and the## Evidence backingsection maps claims β sources. A stack with unbacked claims stays draft, neververified. Dosing/timing figures are recorded as cited research / applied-protocol information under research-use framing β never as personalised medical advice. - Compliance guardrail (non-negotiable): every peptide/article note carries
research_use_only: true, and no note makes medical-treatment or dosing-for-humans claims. Framing is research/educational throughout β consistent with the siteβs own βresearch and educational purposes onlyβ stance. This guardrail propagates to Layer 3.
10. Supabase Ingestion Model
Obsidian is the source of truth (write side); Supabase is the read model the site/app queries. The forward pipeline already exists β it does not need to be built.
Whatβs already live (project gqkhhfuafflcfkjczsxg):
- The
ingest-contentEdge Function (/functions/v1/ingest-content, auth via scopedx-ingest-key) maps note frontmatter β the live storefront tables:guides(+guide_sections,guide_products) andproducts(+ research/variants/prices/images/tags). Upserts by slug; everything landspublished = false(draft) until reviewed in Studio. - 86 published
guidesplusguide_sectionsare the canonical blog content. This KBβs Layer-1 Articles folder was backfilled one-time from those rows (Supabase β vault) so the vault becomes the forward author-source. From here, edits flow vault βingest-contentβ Supabase draft β publish.
The gap β the atomic graph has no DB home yet. The live schema stores articles (guides) and products, but not the Layer-2 atomic notes (peptides, mechanisms, papers, concepts). Decision needed (see Β§14): either (a) add kb_peptides / kb_mechanisms / kb_papers / kb_concepts + join tables and extend ingest-content to populate them, or (b) keep the atomic graph vault-side only and surface it to customers via RAG/embeddings rather than relational tables.
Direction stays one-way: vault β Supabase (via ingest-content). Supabase rows are never hand-authored as the source; the vault is. The one exception was this initial backfill, now complete.
11. Knowledge β Content Manipulation Layer (Layer 3)
This is the bridge to TikTok/Threads, governed by Content β Playbook. The knowledge base supplies whatβs true; the playbook supplies how each channel speaks.
- One-way borrow. A content post references the atomic notes it draws from (
source_notes: ["[[BPC-157]]"]) so we can trace any claim back to the base. Posts never edit the base. - Channel tone is applied on extraction, not stored in the base. The peptide note is neutral and evidence-led; the TikTok script is punchy; the Threads alias post is community-voiced. Same fact, three packagings.
- Threads alias rule (from memory, preserved): Threads posts run under disposable alias brands, never as INDEXA β but the facts still come from the INDEXA base. The base is brand-neutral truth; the alias is just a delivery vehicle.
- Research pillar pulls directly from
kb-papernotes and the tier list β this is the trust engine for the knowledge-minded Threads audience. - Guardrail inheritance:
research_use_onlyand the no-medical-claims rule carry into every derived asset automatically.
12. Governance β Source of Truth Rules
- Single source of truth: Layer 1 + Layer 2 in Obsidian. The live site and Supabase are projections; social posts are derivatives.
- One thing, one note (inherited golden rule). Facts are written once in the atomic note that owns them.
- Frontmatter is the data; the body is context + provenance-tagged enrichment.
- No fact without provenance (Β§9). No medical/treatment claims, ever.
- Mirrors stay faithful; interpretation stays cited. Editorialising happens only in Layer 2, against papers.
- Changes flow downhill: vault β Supabase β site/app; vault β playbook β channels. Never upward.
- Stubs are debt: unfilled
[[links]]are tracked and burned down before a cluster is markedverified.
13. Build Roadmap
| Phase | Deliverable | Status |
|---|---|---|
| 0. Strategy | This document, approved | β done |
| 1. Scaffold | Folder tree + 5 note templates + 3 Bases (manifest seeded from sitemap, 86 rows pending) | β done |
| 2. Mirror | All 86 articles backfilled from Supabase guides/guide_sections as Layer-1 notes (mirrored, 455 sections) | β done |
| 3. Extract | Peptide/Mechanism/Paper/Concept atomic notes created; articles re-linked (extracted). Seed papers from references_raw. | β¬ next |
| 4. Verify | Provenance + compliance pass; stubs burned down (verified) | |
| 5. Ingest (forward) | ingest-content Edge Function already live (vault β Supabase drafts). Extend to atomic graph only if Β§14.6 picks relational tables. | partial β forward pipeline exists |
| 6. Operate | Re-sync cadence + content-borrow workflow into Layer 3 |
14. Open Decisions β need your call before Phase 1
- Atomic depth for the first pass. Mirror all 82 first (Phase 2 complete) then extract, or do mirror+extract article-by-article? (Recommend: mirror all 82 first β faster to a complete Layer 1, cleaner dedupe in extraction.)
- Papers scope. Do articles already cite specific PubMed IDs we can harvest, or do we build
kb-papernotes from scratch during extraction? (Changes Phase 3 effort significantly.) - Supabase + GitHub access. Share the Supabase project (schema/keys) and the vaultβs git repo if you want CI-driven ingestion β otherwise Iβll spec a manual export script.
- Bases vs. backlinks for reverse edges. Confirm we rely on Obsidian Bases to derive βarticles mentioning peptide Xβ rather than maintaining reverse lists by hand (recommended β less drift).
- Resource pages. Include all six site resources (dosage, reconstitution, calculator, storage, payment, dictionary) as
kb-resource, or only the science-bearing ones (exclude payment)? (Recommend: exclude payment-guide β not knowledge.) - Atomic-graph DB home. The live schema has
guides+productsbut no home for peptide/mechanism/paper/concept notes. Either (a) addkb_*tables + join tables and extendingest-content, or (b) keep the atomic graph vault-side and serve it to customers via RAG/embeddings. (Recommend: start with (b) β the graph is for learning/retrieval, not storefront rows; relational tables can come later if a feature needs them.)
Phases 0β2 are complete: strategy approved, vault scaffolded, all 86 articles mirrored from Supabase. Next is Phase 3 β extracting the atomic peptide/mechanism/paper graph and re-linking the articles into it.