saif-docs Search Index Reference¶
Field contract and known limitations for the shared saif-docs Azure AI Search
index, for consumers querying it directly (primarily agent-based retrieval
consumers). The index is built from every app publishing docs to the shared
$web container — legends, iac-azure-modules, iac-okta-modules,
iac-aws-modules, cloud-foundations, and other publishing apps. Schema source:
infra/environments/search/index.json
and infra/environments/search/indexer.json in this repo.
Retrievable Fields¶
| Field | Type | Source | Meaning |
|---|---|---|---|
id |
Edm.String |
metadata_storage_path, base64-encoded |
Document key. Not human-readable. |
title |
Edm.String |
metadata_title |
Page title, searchable and semantic-weighted. |
description |
Edm.String |
metadata_description |
Page summary from the HTML <meta name="description"> tag, searchable and semantic-weighted. |
content |
Edm.String |
Extracted document body | Full text of the page. Markdown documents include leading YAML frontmatter as literal content. |
path |
Edm.String |
metadata_storage_path |
Full blob URI. See Canonical location. |
repo |
Edm.String |
metadata_storage_path token 4 |
Filterable/facetable app identifier. See app vs repo. |
app |
Edm.String |
metadata_storage_path token 4 (same mapping as repo) |
Searchable, semantic-weighted app identifier. See app vs repo. |
section |
Edm.String |
metadata_storage_path token 5 |
Filterable/facetable subfolder beneath the app. Filter-only — deliberately not searchable and not ranked. See section. |
tags |
Edm.String |
metadata_keywords |
Page keywords as a single joined string. Not usable for per-tag filtering. Searchable/semantic-weighted. See tags is a single string, not a tag list. |
format |
Edm.String |
metadata_storage_file_extension |
Filterable/facetable file type (e.g. .html, .md). See Deduplication contract for reliability caveats. |
content_type |
Edm.String |
metadata_storage_content_type |
Filterable/facetable MIME type as supplied by the uploader (not a content classification). Documented fallback if format proves unreliable. |
last_modified |
Edm.DateTimeOffset |
metadata_storage_last_modified |
Filterable/sortable blob last-modified timestamp. |
app vs repo¶
app and repo are populated by the same field mapping and always hold the
same value: the path segment immediately after $web in the blob's storage
path (position 4 when split on /). They differ only in index attributes:
appissearchableand included in the semantic configuration'sprioritizedKeywordsFields, so it carries relevance weight in both keyword scoring (docs-default) and semantic ranking.repoisfilterableandfacetablebut notsearchable, so it is for$filter/faceting, not free-text matching.
Every publishing app, including cloud-foundations, uploads under its own folder in $web, so token 4 is an app name and both fields derive correctly. cloud-foundations gets this from containerFolderName: cloud-foundations on the docs_content_cloud_foundations deployment in .azdo/azure-pipelines.yml, which puts its pages at $web/cloud-foundations/....
Stale documents from before that move: cloud-foundations previously published to the $web root with no containerFolderName, so its pages took whatever top-level folder sat at token 4 (guides, reference, assets) instead of an app name. Those documents are still in the index with those wrong values, and the old blobs are still live at the container root because the root-site deployment runs with skipDelete: true and prunes nothing. Reprocessing alone cannot fix this: an unchanged blob is never revisited on an incremental run, and even a forced reset would just re-derive the same wrong values from the same unchanged path. They will keep their wrong app/repo until the blobs themselves are deleted, which native soft-delete deletion detection (see Deletion Detection) picks up on the indexer's next run and removes the matching document; see #158 for that cleanup. A consumer filtering repo eq 'cloud-foundations' today gets the correctly-derived copies only; the pre-move duplicates surface under repo eq 'guides' and repo eq 'reference'.
section¶
section is the path segment one level below app — position 5 when the blob path is split on /. For https://saifdocs<env>.blob.core.windows.net/$web/legends/guides/getting-started.html the split yields position 3 $web, position 4 legends (app/repo), position 5 guides (section).
It is filterable and facetable but deliberately not searchable, not in the docs-default scoring profile, and not in the semantic configuration. That is the point of the field. It exists so a consumer can exclude a class of content outright rather than try to de-emphasize it through ranking:
A filter is the right instrument here because ranking cannot do this job for every consumer. Scoring-profile text weights are per-field, not per-document-class, and scoring functions are boost-only, so a profile cannot push a category down. Worse, for the consumers most likely to want this, the profile is inert: a weights-only profile like docs-default receives no post-semantic boost, and agentic retrieval ignores defaultScoringProfile entirely (see Ranking Caveats). A $filter is honored in every query mode, including agentic retrieval.
A derived field is required because Azure AI Search $filter cannot match on a path prefix. The OData expression grammar admits only geo.distance, geo.intersects, search.in, search.ismatch, and search.ismatchscoring — there are no string functions at all, so startswith(path, ...) is not expressible. search.ismatch is not a substitute either, because path is searchable: false and therefore cannot be targeted by a full-text function.
Known limitations¶
Root-level pages are dropped from the index. This is the serious one. extractTokenAtPosition returns an error when the requested position does not exist, and a field-mapping function error is a document-level failure, not a field-level one — the document is counted as a failed item and never reaches the index. Because the indexer sets maxFailedItems: -1, those failures are tolerated silently and the run still reports success. This applies to exactly two blobs, both built by mkdocs.root.yml and uploaded to the container root by the docs_content_root deployment: index.html, the docs.saif.com landing page, and 404.html. MkDocs also emits sitemap.xml and sitemap.xml.gz there, but both are excluded by excludedFileNameExtensions. The remaining two split to positions 0 through 4 only, so position 5 does not exist and they fail to reach the index rather than indexing with a null section.
"Fail to reach the index" assumes no document already exists at that key. It does not: index.html at the container root previously served this repo's own cloud-foundations landing page, before this split moved that content under cloud-foundations/, so it is likely already indexed with an app/repo/section derived from that old content. A mapping failure skips the attempted update, it does not delete the prior document, so that stale entry can persist under its old values and outdated preview text indefinitely. This is the same class of problem as the stale app/repo documents above, and the same fix applies: the blob has to be deleted (or the document removed directly) for native soft-delete deletion detection to clear it; see #158.
This is permanent and intended, not a pending defect. The landing page has to sit at the container root to serve https://docs.saif.com/, a 404 stub has no retrieval value, and neither page is worth a result slot, so neither will ever acquire a section. Do not read this as something a containerFolderName will fix: cloud-foundations' own docs already publish under cloud-foundations/ and index normally, and these two root-site pages are what is left over.
Note the timing, because it is counterintuitive: an ordinary incremental run does not trigger this, since unchanged blobs are never revisited. The loss materializes when those specific blobs are rewritten, or when the indexer is reset — which is to say the remediation for the backfill limitation is the same action that triggers this loss. Anyone running a reset should check itemsFailed and the errors array on GET /indexers/saif-docs-indexer/status afterwards, not just the run status, and should expect these two failures to be present every time.
A page published directly under the app folder gets a filename, not a section. $web/{app}/index.html yields section = 'index.html', because position 5 is the filename rather than a folder. section is only meaningful for pages published at least two levels deep.
Documents indexed before cloud-foundations moved under its own folder carry a misderived section. For those, app holds a top-level folder name (guides, reference) and section holds whatever sits one level below it, which is not a section name in the sense the field is meant to carry. This affects only that stale set, not pages published since the move. See app vs repo.
section is subject to the same backfill limitation as every other field added after initial indexing. It is null for any document not reprocessed since the field was added. Consumers writing an exclusion filter should prefer the null-safe form, since section ne 'release-notes' evaluates to true for null and is therefore already safe for exclusion — but an inclusion filter such as section eq 'guides' will silently exclude most of the corpus during the backfill window.
Canonical Location¶
path is the full blob storage URI (e.g.
https://saifdocs<env>.blob.core.windows.net/$web/legends/guides/getting-started.html),
not a public URL. The corresponding public URL follows the pattern:
where {app} is the same path segment app/repo derive from, and {...} is
the remainder of the blob path after that segment. This repo is a worked
example: $web/cloud-foundations/reference/saif-docs-search-index.html serves
as https://docs.saif.com/cloud-foundations/reference/saif-docs-search-index/,
with app/repo = cloud-foundations and section = reference.
The root site is the only exception. Its blobs sit directly at the container
root and serve from https://docs.saif.com/ with no {app} segment, which is
also why they are absent from the index (see
section's known limitations).
Excluded Content¶
The indexer's excludedFileNameExtensions skips assets and build artifacts that carry no search value: stylesheets, scripts, source maps, images, fonts, archives, .drawio, .py, .txt, and .xml.
.xml is excluded primarily to keep MkDocs' generated sitemap.xml out of the index. It has no retrievable content worth ranking, and for apps publishing under a folder it would otherwise index as a normal document with section = 'sitemap.xml'. If an app ever needs to publish meaningful XML documentation, this exclusion has to be revisited.
Exclusion only stops future crawling. It does not remove documents already in the index — the blobs themselves are still live, so deletion detection doesn't apply here regardless of soft-delete state; only an index rebuild (or the blob actually being deleted) clears them. Any sitemap.xml documents indexed before this exclusion was added will persist until then. The deduplication filter already screens them out once they carry format = '.xml', but stragglers indexed before format existed hold format = null and pass through it — the same null-coverage gap described there.
Deduplication Contract¶
Apps that publish both HTML and Markdown twins of a page (some apps do this
via a copy_markdown.py-style build step) get two documents indexed per page.
Consumers that want exactly one document per page must send this as a
request-level $filter:
Send it as $filter on the search request itself, not as a client-side
post-filter applied after retrieving the top-k results — post-filtering after
retrieval still lets duplicates consume result slots and can leave a query with
fewer (or zero) usable results than requested.
Why or format eq null is required, not optional: format is populated by
reprocessing a blob through the indexer, and an ordinary indexer run only
touches blobs that changed since the last run (see
Backfill limitation). Documents not yet reprocessed have
format = null. A bare format eq '.html' filter would exclude every one of
those unprocessed documents — which, during and shortly after rollout, is most
of the corpus — silently hiding the majority of the index rather than
deduplicating it. Consumers should tighten the filter to format eq '.html'
only once coverage is verified at 100%.
Reliability caveat: format's source field,
metadata_storage_file_extension, is not on Azure's documented list of
standard blob-property fields extracted by indexers (as of the current
Content metadata properties
reference). It is widely used in practice but unconfirmed for this index until
verified against real blobs in platformdev. Until that verification, treat
format as provisional and content_type (which is documented) as the
fallback signal for distinguishing HTML from Markdown documents.
Configuration Version¶
The index eTag, returned by GET /indexes/saif-docs, identifies the index
schema only — its fields, scoring profile, and semantic configuration. It
does not identify the state of the corpus. Backfill and reprocessing change
field population and therefore ranking behavior without changing the eTag, so
an eTag comparison alone cannot tell a consumer whether results have changed
since a prior evaluation. Evaluation or regression runs that need reproducible
results should also record the indexer's run state (e.g. last successful run
timestamp), not just the eTag.
A query-only consumer identity may not have permission to call
GET /indexes/saif-docs at all; do not assume the eTag is reachable from every
credential that can query the index.
Ranking Caveats¶
The docs-default scoring profile weights title 5, description 4, tags
3, app 2, content 1. What that weighting actually controls depends on the
query type:
- Plain keyword queries: the profile controls ranking end to end.
- Semantic queries (a common consumer query mode): the profile shapes
which documents get selected as L1 candidates before semantic reranking, but
it does not reorder the final results after semantic reranking runs.
Semantic ranking only applies post-semantic boosting to scoring profile
functions — a weights-only profile like
docs-defaultgets no post-semantic boost at all. - Agentic retrieval: the profile has no effect whatsoever. Azure explicitly
excludes
defaultScoringProfilefrom agentic retrieval.
Consumers evaluating "does changing the scoring profile change my results" need to account for which query mode they're using — a profile change can be invisible to a semantic or agentic consumer while still affecting a plain keyword one.
Backfill Limitation¶
Newly added fields (description, tags, app, section, format,
content_type, last_modified) are null for any document the indexer has not
reprocessed since the fields were added.
These fields do not fill in on their own. An earlier version of this document said coverage "fills in page by page, as content actually changes." That was optimistic to the point of being wrong, and it is corrected here. A blob indexer's change detection is driven by the blob's LastModified timestamp, which the indexer tracks through an internal high-water mark. The docs deploy uses azcopy sync --compare-hash=MD5, so an unedited page is never re-uploaded, its LastModified never moves, and the indexer never revisits it. Azure's own wording is blunt about the consequence: "If the underlying content is unchanged, a run operation has no effect… To reprocess all documents, you need to reset the indexer." Adding fields to the index and mappings to the indexer does not change this. Updating an indexer definition does kick off a run, but that run is still incremental and will process nothing.
So the realistic expectation is not incremental convergence. It is that these fields stay null indefinitely for the existing corpus, and are populated only for pages that are genuinely rewritten. Coverage does not approach 100% through ordinary publishing activity.
Remediation is a reset, and it is not part of the deploy pipeline. Full population requires clearing the high-water mark and re-crawling:
POST {searchEndpoint}/indexers/saif-docs-indexer/reset?api-version=<version>
POST {searchEndpoint}/indexers/saif-docs-indexer/run?api-version=<version>
Reset clears change-tracking state only. It does not delete documents from the index, and it does not clean up orphans. It is cheap for this index because there is no skillset, so nothing has to be re-enriched — the cost is one full re-crawl of the $web container.
The deploy pipeline does not do this. The configure-search-index activity in the SAIF/pipeline-templates Azure DevOps repo runs PUT /indexes, PUT /datasources, PUT /indexers, then POST /indexers/{name}/run, with no reset step. Until that activity grows an opt-in reset, backfill is a manual, out-of-band operation.
Before running a reset, read section's known limitations. A reset is what causes the root site's index.html and 404.html to drop out of the index, and the failure is silent under maxFailedItems: -1. Check itemsFailed and errors on GET /indexers/saif-docs-indexer/status after any reset-driven run.
Deletion Detection¶
infra/environments/search/datasource.json sets a
NativeBlobSoftDeleteDeletionDetectionPolicy, which depends on Azure Blob
Storage soft delete being enabled on the source account.
Historical state (verified 2026-09-14, platformdev): it was not enabled.
az storage account blob-service-properties show --account-name saifdocsplatformdev
returned deleteRetentionPolicy.enabled: false. The docs_storage module
invocation in infra/environments/environment/docs-site.tf passed no
soft-delete arguments at all, and its >= 4.0.0, < 5.0.0 version range
couldn't reach the module version (5.1.0) that added
blob_delete_retention_enabled / blob_delete_retention_days support, so the
Terraform alone confirmed nothing about the live account either way — only
the account check did.
Fix (cloud-foundations#144): as of this change,
infra/environments/environment/docs-site.tf pins module.docs_storage to
>= 5.1.0, < 6.0.0 and sets blob_delete_retention_enabled = true with
blob_delete_retention_days = 7 (7 days comfortably covers a missed PT24H
indexer run). blob_versioning_enabled is set explicitly to false — Azure
Storage allows soft delete and versioning together, but Azure AI Search's
native-soft-delete deletion detection specifically does not tolerate the
combination. This still needs the normal terraform plan/apply pipeline
run before it takes effect live; this doc will need a follow-up update once
that's confirmed applied.
Once applied, deleting a blob will be detected on the indexer's next run and the corresponding search document removed. An indexer reset alone still does not do this — reset "doesn't trigger deletion or clean up of orphaned documents in the search index."
Retroactive orphans: per Azure's own documentation on changed and deleted blobs,
the deletion detection policy only catches blobs deleted after it's in
place — documents for blobs deleted before soft delete was enabled remain in
the index indefinitely, even after enabling and resetting the policy. Those
have to be swept directly by diffing live $web blobs against indexed
path values and deleting the orphans via the index's batch delete API.
That capability doesn't exist yet as a repeatable, pipeline-driven operation
— it belongs in the SAIF/pipeline-templates Azure DevOps repo alongside the
existing configure-search-index activity (see Backfill limitation).
Until it lands, this is a manual, out-of-band cleanup.
tags Is a Single String, Not a Tag List¶
Two separate problems, and both matter before anyone builds on this field.
It is null everywhere today. tags is sourced from metadata_keywords, which the indexer extracts from an HTML document's <meta name="keywords"> tag. MkDocs Material — the theme used by every app publishing to docs.saif.com — only emits <meta name="description"> out of the box; it does not emit a keywords meta tag. Until a publishing app's theme is customized to emit one from frontmatter tags, tags stays null for that app's pages. No app does this today.
It is not a tag list, and filterable: true overstates what it can do. tags is Edm.String, a single scalar. When a keywords meta tag does arrive, its contents land as one comma-joined string, so tags eq 'terraform' is a whole-string exact match against the entire joined list and matches only a page whose complete keyword string is exactly terraform. It is not per-tag filtering. Real per-tag filtering requires Collection(Edm.String).
Decision: tags stays Edm.String, and filterable stays on. Recorded deliberately, because changing a field's type later requires dropping and recreating the whole index — the same wall that forced app to exist as a duplicate of repo. The reasoning:
- The obvious migration does not exist. There is no built-in skill that splits a string on an arbitrary delimiter.
SplitSkillsplits by length or sentence boundary and has nodelimiterparameter, andConditionalSkill's expression language has no string functions at all. Getting a collection out of a delimited string means either a customWebApiSkill— new hosted code, new infrastructure, new ownership, and a new document-level failure surface on a shared index — or thejsonArrayToStringCollectionfield mapping function. - The
jsonArrayToStringCollectionroute has an unacceptable failure mode here. It requires the source value to already be a JSON array string, and it errors on anything it cannot parse. Field-mapping errors are document-level, and this indexer runs withmaxFailedItems: -1, so a single app emitting an ordinary comma-joined keywords tag would silently drop every one of its pages from a shared index while the run reported success. That risk is not worth taking for a field no consumer has asked to filter on. - No consumer needs it. The requesting need was to de-emphasize a class of content, and
sectionanswers that properly. - Removing
filterableis not free either. Changing a field's index attributes requires the same full rebuild as changing its type, so dropping the flag costs exactly what keeping it costs, and keeping it preserves whole-string exact match andsearch.in(tags, '...')for the day a single-keyword convention emerges.
Consumers should treat tags as a free-text signal only: it is searchable and semantic-weighted, so keyword content contributes to relevance. Do not build per-tag filtering on it. If per-tag filtering is ever genuinely required, it is a full index rebuild plus a custom skill, and it should be scoped as such rather than bolted on.