XBRL¶
The project's principal public input is the SEC's Financial Statement and Notes Data Sets: normalized bulk extracts of XBRL information from SEC filings.
The XBRL source documentation is split into two parts:
- SEC bulk XBRL data describes the original SEC tables, their keys, and their meaning.
- Pre-project Parquet processing documents the normalization layer that converts the bulk tables into the Parquet files mounted at
/dataand used as this repository's starting point.
The distinction matters. This repository begins from the Parquet inputs; it does not download or convert the SEC TSV releases itself. Its Julia pipeline builds the concept registry and final disclosure datasets from those normalized inputs.