Schema coverage¶
This page is an honest map of how much of the DDI Lifecycle 3.3 schema
ddi-l supports, and at which level. "Support" means three different things,
so the table separates them into three tiers.
The three tiers¶
- High-level API: you can create and manage these with the simple
DocumentAPI:new_study(),add_question(),add_variable(),add_item(),items(),find(),remove(). This is the easy, one-line path most authoring needs. - Typed model: every type in the module is a typed Python class with
full
from_xml()/to_xml()support, so you can read, build, and edit it precisely fromddi_l.models.*even when there is no one-line helper. - Round-trip preserved: opening a file and saving it again keeps the content byte-for-byte faithful, even for parts the library does not model in detail. Unrecognized elements are preserved verbatim.
Nothing is lost
All of DDI 3.3 is modeled and validated: the model layer is generated
from the official XSD (about 500 types), and doc.validate() checks
against the real instance_3_3.xsd. Round-trip fidelity is complete: a
100,000-element production file re-saves with zero element loss. The only
thing that varies by module is ergonomics: how much typing you do.
Coverage by module¶
| DDI module | High-level API | Typed model | Round-trip |
|---|---|---|---|
| Study Unit | Yes | Yes | Yes |
| Conceptual Component (concepts, universes, unit types) | Yes | Yes | Yes |
| Data Collection (questions, instruments, questionnaire flow) | Yes | Yes | Yes |
| Logical Product (variables, code lists, data relationships, NCubes) | Yes | Yes | Yes |
| Reusable (shared types: references, names, dates…) | n/a | Yes | Yes |
| Archive (citations, coverage, provenance) | Yes | Yes | Yes |
| Methodology (data collection methodology) | Yes | Yes | Yes |
Process types (urn:ddi-l:extension:process:1) |
No | Yes | Yes |
| Comparative (harmonization maps) | Yes | Yes | Yes |
| Physical Data Product: record layouts | Yes | Yes | Yes |
| Physical Data Product: structures, NCube | Yes | Yes | Yes |
| Physical Instance (data files) | Yes | Yes | Yes |
| Group (study series) | Yes | Yes | Yes |
| Resource / Local Holding Packages | Yes | Yes | Yes |
| Translation information | Yes | Yes | Yes |
| DDI Profile (usage declaration) | Yes | Yes | Yes |
| Dataset (inline data) | Yes | Yes | Yes |
"No" in the High-level API column means there is no one-line helper; use the typed model instead. n/a means not applicable.
About the Process types
ddi_l.models.process models Process, ProcessStep, ProcessControl,
ProcessMethod and their schemes. DDI Lifecycle does not define these:
they appear in none of the 3.1, 3.2 or 3.3 schemas, neither as elements nor
as complexTypes. They are a ddi-l extension, and they serialize into
urn:ddi-l:extension:process:1 so that is unmistakable in the output. Other
DDI tools will not recognise them. Build them from ddi_l.models.process
when you want the modelling; do not expect them to interoperate.
DDI's own process vocabulary is ProcessingEvent, ProcessingInstruction
and ControlConstruct, which live in ddi:datacollection:3_3 and are
modeled separately. Questionnaire and data-capture
flow in DDI 3.3 instead lives in Data Collection as control constructs
(QuestionConstruct, Sequence, IfThenElse, StatementItem,
ComputationItem, Loop), which are first-class add_item types.
What "add_item" covers¶
add_item() / items() currently recognize 30 item types, including the
full questionnaire-flow set (QuestionConstruct, Sequence, IfThenElse,
StatementItem, ComputationItem, Loop; see
Module 8), PhysicalInstance,
RecordLayout, DataRelationship, and NCube. A physical instance describes
one data file; a record layout maps variables to positions within it.
The samples on this page all continue from this setup:
import ddi_l as ddi
doc = ddi.new_study(title="Health Survey", agency="example.org")
age = doc.add_variable(name="age")
income = doc.add_variable(name="income")
year = doc.add_variable(name="year")
region = doc.add_variable(name="region")
population = doc.add_variable(name="population")
footnote = doc.add_variable(name="footnote")
age_2020 = doc.add_variable(name="age_2020")
age_2021 = doc.add_variable(name="age_2021")
sex_codes = doc.add_code_list(name="Sex Codes")
from ddi_l.models.physical import PhysicalInstance
pi = doc.add_item(PhysicalInstance, name="2021 Microdata File")
pi.set_data_file("https://example.org/health-2021.csv")
pi.set_record_count(15000)
# Map variables to fixed-width columns
rl = doc.add_record_layout()
rl.add_data_item(age.to_reference(), start_position=1, width=2)
rl.add_data_item(income.to_reference(), start_position=3, width=8)
# Logical records (which variables make up one case)
dr = doc.add_data_relationship()
dr.add_logical_record() # all variables, one rectangular record
# Multidimensional (cube) data (with a coordinate region for an attribute)
cube = doc.add_ncube(name="Population by year and region")
cube.add_dimension(year.to_reference())
cube.add_dimension(region.to_reference())
cube.add_measure(population.to_reference())
region = cube.add_coordinate_region()
cube.add_attribute(footnote.to_reference(), attachment_region=region)
# Variable types: numeric / coded / text / datetime
# (numeric ranges take inclusivity flags and missing values)
doc.add_variable(name="age").set_numeric(
"Integer", low=0, high=120, low_inclusive=True, missing_values=["-9"]
)
doc.add_variable(name="sex").set_coded(sex_codes)
Reaching the model layer¶
When a module has no one-line helper, you still have full typed access. Build
the object from ddi_l.models.* and attach it, or read it back after
open_ddi(). For example, the physical data layer lives in
ddi_l.models.physical, the archive layer in ddi_l.models.archive, and so
on. Because every type round-trips, you can also open an existing file, reach
the part you need, edit it, and save; nothing else in the document changes.
Groups¶
A document can be organized as a Group (a study series / publication
package). doc.add_group() moves the study under a <g:Group>, and the
document stays fully editable. Group-organized files open and round-trip:
doc = ddi.new_study(title="Wave 1", agency="example.org")
doc.add_variable(name="age")
doc.add_group() # organize the study into a group
doc.add_variable(name="income") # still edits the (now grouped) study
Add more studies with doc.add_study(title=...), and relate items across
them with a Comparison:
doc.add_study(title="Wave 2") # a second study in the series
cmp = doc.add_comparison(name="2020 to 2021")
cmp.add_variable_map(age_2020.to_reference(), age_2021.to_reference())
The high-level add_* helpers edit the primary study by default. To edit a
different study in the series, use its cursor: doc.study(identifier) returns a
StudyCursor whose add_* helpers target that study:
wave2 = doc.add_study(title="Wave 2")
doc.study(wave2.identifier).add_variable(name="income") # edits wave 2
doc.study().add_variable(name="age") # edits the primary study
Instance-level packages, archive, and translation¶
Beyond the study, a DDIInstance can carry reusable and holding packages, an
archive, and translation information, each with a one-line helper:
doc.add_archive() # archive lifecycle metadata on the study
doc.add_resource_package() # reusable metadata shared across studies
doc.add_local_holding_package() # a local holding of a deposited study
doc.add_translation_information( # which languages the instance was translated between
languages=["en", "fr"], description="Translated from French."
)
Read them back with doc.archives, doc.resource_packages,
doc.local_holding_packages, and doc.translation_information. Packages and
translation information are serialized into their correct positions in the
DDIInstance content model on save.
Coverage summary¶
The high-level API spans every DDI Lifecycle 3.3 instance module: study
units, concepts, data collection and questionnaire flow, logical products
(variables with typed representations including inclusivity flags and missing
values, code lists, data relationships, NCubes with
dimensions/measures/attributes and coordinate regions), physical instances,
structures and record layouts (including segment keys and storage/decimal
details), inline datasets (item, record, and variable sets), study groups with
multi-study series (doc.study(id)) and the full set of comparison maps
(variable/concept/… maps plus managed-item and representation maps),
instance-level packages (resource, local-holding, archive, translation), and DDI
profiles. The one group that is model-layer-only is the Process types
(urn:ddi-l:extension:process:1), which DDI does not define at all. Everything
in every actual DDI module is modeled, validated, and round-tripped.