Overview
Catalog2Registration is a .NET Framework 4.7.2 console application that runs as a continuous background job. Its purpose is to keep the NuGet V3 Registration resource up to date. The registration resource is the structured JSON-LD data that NuGet clients query to discover all versions of a package, their metadata, and their dependencies. This job is the sole writer of those blobs; the blobs are served directly from Azure Blob Storage to NuGet clients without any intermediary service. The job operates as a catalog collector. It polls the NuGet V3 catalog index for new commit items, which represent package publishes (PackageDetails) and package deletions (PackageDelete). When new items are found, the job batches them and processes each affected package ID in parallel. For each ID it reads the existing registration index from blob storage, merges the incoming changes using an in-memory merge algorithm, and then writes back only the blobs that changed—leaves, pages, and the index—while deleting blobs that are no longer needed.
A key design decision is the concept of “hives.” The job writes the same data to three separate Azure Blob Storage containers with different characteristics: a Legacy hive (plain JSON, SemVer 1 only), a Gzipped hive (gzip-compressed JSON, SemVer 1 only), and a SemVer2 hive (gzip-compressed JSON, SemVer 1 and SemVer 2). The Gzipped hive is treated as a replica of the Legacy hive—writes to Legacy are automatically mirrored to Gzipped with URL rewriting. The SemVer2 hive is processed independently and also includes package deprecation and vulnerability metadata that are omitted from the other two hives.
Role in System
Three Registration Hives
Writes to Legacy (plain JSON, SemVer 1), Gzipped (gzip, SemVer 1), and SemVer2 (gzip, SemVer 1+2) containers. Gzipped is a replica of Legacy with URL rewriting applied automatically during writes.
Cursor-Based Progress
Uses a durable front cursor stored in blob storage to track how far through the catalog the job has processed. Optionally depends on upstream cursors via HTTP so it does not run ahead of upstream jobs.
Inlined vs. Non-Inlined Pages
For packages with 127 or fewer total versions, page items are inlined directly into the registration index JSON. Above that threshold, each page is written as a separate blob, keeping the index small for large packages.
Parallel Merge Algorithm
Uses a textbook merge algorithm to combine sorted catalog commits with the existing sorted registration pages and leaves, touching only the blobs that actually change rather than rewriting everything.
Key Files and Classes
Dependencies
NuGet Package References
The project itself declares no direct NuGet package references; all external packages are pulled in transitively through the two internal project references below. The most significant transitive dependencies include:Internal Project References
Notable Patterns and Implementation Details
The Gzipped hive is a replica of the Legacy hive. In
RegistrationUpdater, the HiveToReplicaHives dictionary maps HiveType.Legacy to [HiveType.Gzipped] and HiveType.SemVer2 to an empty list. When HiveStorage.WriteAsync is called, it iterates the primary hive plus its replicas, rewrites the entity’s internal URLs for each target hive using EntityBuilder.UpdateIndexUrls / UpdatePageUrls / UpdateLeafUrls, serializes and writes in parallel, then restores the original URLs so the caller sees no side effects.SemVer 2.0.0 packages are excluded from the Legacy and Gzipped hives.
HiveUpdater detects this by checking PackageDetailsCatalogLeaf.IsSemVer2() on incoming items. When a SemVer 2 package is found heading to a non-SemVer2 hive, its catalog commit item is replaced with a synthetic PackageDelete entry before the merge, ensuring that any previously stored version is removed without requiring a separate deletion flow.Every blob write optionally ensures at least one Azure snapshot exists for the blob (
EnsureSingleSnapshot = true by default). After uploading a blob, HiveStorage lists the blob with snapshot details and creates a snapshot if none is present. This provides a recovery point against accidental deletion via tooling such as Azure Storage Explorer without using soft-delete.