Overview
NuGet.Services.Metadata.Catalog is the shared library that owns the NuGet V3 catalog protocol end-to-end. The catalog is an append-only, time-ordered log of every package lifecycle event (publish, edit, delete) stored as a hierarchy of JSON-LD documents in Azure Blob Storage. This library covers both sides of that protocol: a writer stack that produces the catalog and a collector stack that lets other services consume it.
On the write side, AppendOnlyCatalogWriter batches CatalogItem objects, serializes them as RDF graphs framed into JSON-LD (using dotNetRDF and json-ld.net, both net472-only), and commits them to storage. Each commit produces leaf documents for individual package events, updates the current page (page0.json, page1.json, …), and rewrites the root index.json. Pages are capped at a configurable MaxPageSize; when a page fills, the writer starts the next numbered page and sets an aggressive cache-control header on the now-finished previous page.
On the read side, CommitCollector and its subclasses (DnxCatalogCollector, IconsCollector, SortingCollector) walk the index, fetch pages, and call OnProcessBatchAsync for each time-stamped batch of CatalogCommitItem records. Progress is tracked by a cursor — either an in-memory MemoryCursor, a blob-backed DurableCursor, or an HTTP-readable HttpReadCursor — so that each collector can resume from exactly where it left off across restarts. Cursors record an exclusive lower bound (front) and an inclusive upper bound (back), and commits within a batch share a timestamp so the cursor never advances partially through an atomic commit.
Role in System
Append-Only Writer
AppendOnlyCatalogWriter maintains a numbered page series in blob storage. When the current page would exceed MaxPageSize, a new page is started. The finished page receives an aggressive cache-control value only after the root index is safely updated, preventing stale-cache issues on retry.Cursor-Driven Collectors
Every collector holds a
front cursor (exclusive minimum) and a back cursor (inclusive maximum). The cursor value is persisted to a blob after each successfully processed commit timestamp, enabling safe resume after failure without reprocessing already-handled events.JSON-LD / RDF Representation
Catalog leaves are RDF graphs serialized to JSON-LD using embedded context files (
Catalog.json, Container.json, PackageDetails.json). Nuspec XML is transformed to RDF triples via an embedded XSLT (nuspec.xslt). The schema URIs live in http://schema.nuget.org/schema# and http://schema.nuget.org/catalog#.Icon Pipeline
IconsCollector + IconProcessor handle both embedded icons (extracted from the .nupkg zip) and external icon URLs (fetched via HTTP, size-capped at 1 MB). Content type is determined by magic-byte inspection (PNG, JPEG, GIF, ICO, SVG) rather than file extension or HTTP headers.Flat-Container (DNX) Writer
DnxCatalogCollector drives DnxMaker to maintain the flat-container resource: for each PackageDetails leaf it copies the .nupkg and writes the .nuspec; for each PackageDelete leaf it removes the corresponding blobs and updates the per-package version list JSON.Gallery DB Bridge
GalleryDatabaseQueryService bridges the SQL Gallery database to the catalog write path. It queries Packages, PackageDeprecations, VulnerablePackageVersionRanges, and PackageVulnerabilities in a single parameterized query with control-break logic to accumulate multiple vulnerability rows per package version.Key Files and Classes
Dependencies
NuGet Package References
Internal Project References
Notable Patterns and Implementation Details
The library targets both
net472 and netstandard2.1. The catalog write path (JSON-LD serialization, RDF graph construction, AppendOnlyCatalogWriter, PackageCatalogItem, JsonLdWriter, AzureStorage) is compiled only for net472 because dotNetRDF and json-ld.net are not netstandard2.1-compatible. The read/collector path compiles for both targets.Cursor safety contract:
CommitCollector.FetchAsync advances the front cursor to a commit timestamp only after all items in that timestamp’s batch have been successfully processed. If two consecutive batches share the same timestamp, the cursor does not advance between them — it advances only when the timestamp changes or after the final batch. This prevents partial progress through an atomic catalog commit.Catalog page cache-control timing:
AppendOnlyCatalogWriter deliberately does not set the final (aggressive) Cache-Control value on a completed page until after the root index.json has been saved. This prevents CDN edge nodes from caching a “finished” page before the index acknowledges it, which would otherwise leave consumers unable to discover new pages if the index write fails.- Nuspec-to-RDF via XSLT: Rather than a hand-written nuspec parser,
PackageCatalogItemcallsUtils.CreateNuspecGraphwhich appliesnuspec.xslt(embedded resource) to convert nuspec XML into RDF/XML. dotNetRDF then loads that RDF/XML into an in-memory graph for further triple assertions. GetListedsentinel date: A package withPublished = 1900-01-01T00:00:00Z(Constants.UnpublishedDate) is treated as unlisted. Thelistedpredicate in the catalog leaf reflects this convention rather than any separate boolean column.- Icon magic-byte detection:
IconProcessor.DetermineContentTypechecks raw bytes — PNG (89 50 4E 47), JPEG (FF D8 FF), ICO (00 00 01 00), GIF87a/89a headers — in descending popularity order before falling back to SVG text heuristic. The HTTPContent-Typeheader from the external source is intentionally ignored. AggregateCursorminimum semantics: When multiple downstream collectors share a single catalog (e.g., DNX and search), anAggregateCursorwrapping all their individual cursors is used as thebackcursor for the writer, so the writer does not advance past what the slowest reader has consumed.StringInternerinCatalogIndexReader: When reading all pages in parallel to build a full snapshot, package ID and version strings are interned via a lock-freeConcurrentDictionary-basedStringInternerto reduce memory pressure from duplicate strings across thousands of catalog entries.