Skip to content

Pre-project Parquet processing

Purpose and boundary

This repository starts from normalized Parquet files mounted at /data. A processing layer outside this repository converts SEC bulk TSV releases into those files before the marketing-disclosure pipeline begins.

This page records the observed input contract of that layer. It is deliberately separate from the in-project XBRL processing documented under Processed datasets. The pre-project conversion code is not currently versioned in this repository, so this documentation should not be read as a line-by-line implementation specification.

Transformation overview

SEC bulk TSV release
  sub / num / txt / tag / pre / cal / dim / ren
                 │
                 ▼
Pre-project normalization and form-specific partitioning
                 │
                 ▼
Parquet inputs in /data
                 │
                 ▼
This repository: registry and final disclosure datasets

The normalization layer preserves the long, filing-level structure. It carries filing attributes from submission data into the fact files and resolves dimension and presentation/report metadata so the project pipeline can read a compact set of Parquet inputs without repeatedly joining the raw TSV tables.

Source-to-Parquet mapping

Parquet input Principal SEC source tables Observed transformation/output role
filing_entities.parquet sub.tsv One filing-level record with adsh, issuer identity, form, reporting period, filing/acceptance times, SIC, location/country fields, and other submission metadata.
form_10K_notes_numeric_long.parquet num.tsv + sub.tsv + dim.tsv Annual-form numeric facts in long format. Filing identity/form fields and resolved segments/segt accompany the original numeric fact context.
form_10Q_notes_numeric_long.parquet num.tsv + sub.tsv + dim.tsv Quarterly-form numeric facts in the same long schema as the annual file.
form_10K_notes_text_long.parquet txt.tsv + sub.tsv + dim.tsv Annual-form text facts with their raw context, filing attributes, and resolved segment information.
form_10Q_notes_text_long.parquet txt.tsv + sub.tsv + dim.tsv Quarterly-form text facts in the same schema as the annual text file.
concepts.parquet tag.tsv Concept metadata: tag, taxonomy version, custom/standard flag, data type, label, and documentation.
presentation.parquet pre.tsv + ren.tsv Presentation arcs enriched with report/role names and labels, statement placement, and line information.
calculations.parquet cal.tsv Filing calculation arcs: parent and child tags/versions, group, arc weight, and sign.

dim.tsv and ren.tsv are therefore represented through enriched columns rather than being retained as separate project inputs. In particular, segments and segt in the fact Parquet files are the readable interpretation of the raw dimension hash (dimh).

Parquet fact schemas

The annual and quarterly numeric files share this core schema:

adsh, tag, version, ddate, qtrs, uom, dimh, iprx, value,
footnote, footlen, dimn, coreg, durp, datp, dcml,
cik, form, period, fy, fp, filed, fye, segments, segt

The text files preserve the analogous text context, including lang, context, escaped, srclen, txtlen, and value, plus the same filing and segment fields. The fact files remain long: each row is a reported fact rather than an issuer-period aggregate.

Form partitioning

The Parquet filenames separate annual and quarterly starting inputs. The downstream project currently uses 10-K/10-K/A in the annual build and 10-Q/10-Q/A in the quarterly build. The form field remains in each Parquet file and is retained by the final outputs, so later scope changes can be made without losing filing-form provenance.

The word notes in the current filenames is an input naming convention. Researchers should use the retained filing form, concept, presentation, duration, and dimension fields—not the filename alone—to determine the economic scope of a fact.

Data-quality implications

The pre-project layer should preserve, rather than resolve, most reporting ambiguities. The same issuer and ddate can have multiple filings; a tag can occur with several units, durations, or segment contexts; and standard versus custom concepts have different comparability properties. The in-project pipeline consequently keeps adsh, ddate, filed, and fact context fields instead of collapsing them to one firm-year number.