Release notes¶
0.1.0¶
The first release of ddi-l: a Python toolkit for creating, reading,
updating, and validating DDI Lifecycle 3.3
XML documents, with a simple CRUD API over an XSD-generated model layer.
Authoring¶
- Simple CRUD API:
ddi.new_study(),ddi.open_ddi(), and theDocumentclass withadd_question(),add_variable(),add_concept(),add_universe(),add_code_list(),add_item(),find(),remove(),save(), andvalidate(). Everyadd_*helper takes an optionallabel=(andlabel_lang=), so an item can be labelled where it is created rather than in a second pass. That is the difference between output that lints clean and output that warns about itself. - Generic
add_item()/items(): 30 DDI item types (QuestionItem, Variable, Category, Instrument, Concept, Universe, and more) through a single type registry. - Variable representations:
Variable.set_numeric(),set_coded(),set_text(), andset_datetime()declare what kind of values a variable holds, with optional numeric ranges (low_inclusive/high_inclusive),missing_values,blank_is_missing_value, or a code-list reference. The setters chain ontoadd_variable(). - Custom properties:
set_property(),get_property(),properties, andremove_property()on any DDI item. - Versioning:
increment_major_version(),increment_minor_version(),increment_subversion(), version rationales, and version responsibility. - Multilingual support: a
lang=parameter on everyadd_*method, andInternationalStringfor additional translations.
Study structure¶
- Multi-study series:
doc.add_study()adds further studies to the group.doc.study(identifier)returns aStudyCursorwhoseadd_variable(),add_question(), and otheradd_*helpers target that specific study, so non-primary studies are editable through the high-level API.doc.study()with no argument targets the primary study. - Study groups:
doc.add_group()organizes a study into a<g:Group>(a study series or publication package). The document stays fully editable, and group-organized files open and round-trip. - Instance-level packages:
doc.add_resource_package()anddoc.add_local_holding_package()attach aResourcePackage(reusable metadata) or aLocalHoldingPackage(a local holding of a deposited study, referencing the primary study by default) directly on theDDIInstance;doc.resource_packages/doc.local_holding_packagesread them back. - Archive:
doc.add_archive()attaches anArchivemodule (archive lifecycle metadata) to the primary study;doc.archivesreads them back. - Translation information:
doc.add_translation_information(languages=..., description=...)sets the instance'sTranslationInformation;doc.translation_informationreads it back. - Comparisons:
doc.add_comparison()records harmonization maps between items across studies or versions:add_variable_map(),add_concept_map(),add_managed_item_map(),add_representation_map(), acorrespondence()builder, andsource_scheme/target_scheme/correspondencearguments on the map helpers. - DDI profiles:
doc.add_ddi_profile()attaches aDDIProfiledeclaring which DDI elements a system uses, viaadd_used()/add_not_used()XPath statements.
Data description¶
- Data relationships:
doc.add_data_relationship()builds aDataRelationshipwhoseadd_logical_record()declares which variables make up one case (the rectangular-file default). - NCubes:
doc.add_ncube()builds a multidimensional cube;add_dimension(variable_ref)adds an axis,add_measure(variable_ref)adds a measured value, andadd_attribute()adds a qualifying variable.add_coordinate_region()andadd_dimension_value()describe a region of a cube, andadd_attribute(..., attachment_region=)attaches to one. - Physical instances:
PhysicalInstanceis a first-classadd_itemtype at the study level, describing a concrete data file withset_data_file(),set_record_count(), andset_citation_title(). - Record layouts:
doc.add_record_layout()maps variables to positions in a data file viaadd_data_item(variable_ref, start_position=, width=), which also acceptsstorage_format,delimiter, anddecimal_positions. Passlogical_record=to tie the backing structure to a modeled logical record.PhysicalStructure.link_logical_record(..., key_variable=)declares a segment key. - Inline datasets:
doc.add_dataset()stores data values directly in the document, inItemSet(add_item_value()),RecordSet(set_variable_order()/add_record()), orVariableSet(add_variable_item()) form.
Questionnaires¶
- Flow constructs:
QuestionConstruct,Sequence,IfThenElse,StatementItem,ComputationItem, andLoopare first-classadd_item()types, stored in aControlConstructSchemethat is built, serialized, and round-tripped automatically, and linked to one another with.to_reference().
Validation and quality¶
- Offline schema validation: the DDI 3.1, 3.2, and 3.3 XSDs ship inside
the package, so
doc.validate()andddi validatework with no network access and no separate download. - Reads DDI 3.1, 3.2 and 3.3:
read_ddi(),ddi validate,ddi lint,ddi to-jsonandddi roundtripdetect the version a document declares. TheDocumentauthoring API and the typed models target DDI 3.3; 3.1 and 3.2 documents are worked with as XML throughDDIDocument.root. - Lint engine:
ddi lintandDocument.lint()check reference integrity, missing labels and citation completeness. The label rule follows the schema: it only asks for a label where the element's content model has anr:Labelslot (174 of DDI 3.3's 1247 elements). The citation-language rule matches language ranges per RFC 4647, so a requiredenis satisfied byen,en-CAoren-Latn-CA, but not byeng. - Lint-clean output by default: the
DocumentAPI labels the modules, schemes and physical wrappers it creates, and everyadd_*helper acceptslabel=.set_scheme_label()labels any generated scheme wrapper. Documents you parse are never modified. - Lint defaults suit a published library: the agency allow-list is
opt-in (
configure_lint(allowed_agencies=[...])or--allowed-agency); a missing agency is always an error.ddi lintexits non-zero on errors only; pass--fail-severity warningfor a strict gate. - Labels only where the schema allows them: label support is derived
from the XSDs. Setting a label on a type with no
r:Labelslot warns instead of producing an invalid document. - Reference checks:
Document.save()reports references to items the study does not contain as a singleDDIReferenceWarning. - Clear errors: malformed XML raises
DDIParseErrorwith the parser's line and column; unresolvable references raiseDDIReferenceError(aLookupError); duplicate identifiers raiseDuplicateIdentifierError, and empty or non-string names are rejected when an item is added. Schema failures raiseSchemaValidationError, aDDIValidationError, whose.issuescarries the individual problems. - Examples that pass our own gates: the bundled
example_instance.xmlandQuality_of_Life.xmlare schema-valid and free of lint errors;example_instance.xmlhas no findings at any severity. Methodologyuses DDI's own namespace:<Methodology>is written inddi:datacollection:3_3, as DDI 3.3 declares it.- Extension types are visibly ours: the few
ProcessandMethodologyhelper types that no DDI Lifecycle release defines serialize intourn:ddi-l:extension:*, so nothing can be mistaken for an official DDI namespace. - Round-trip fidelity: unknown XML elements and attributes are preserved when opening and re-saving files.
- The
lxmlbackend is a speed choice, not a dialect: documentsddi-lauthors serialize to byte-identical output with and withoutddi-l[full]. A third-party file whose namespaces are laid out differently (a prefixed root plus per-elementxmlns=redeclarations, as inQuality_of_Life.xml) can differ in prefix spelling between backends; its content is identical.
Performance¶
import ddi_lloads names on first use, so importing the package and starting theddicommand take a few hundredths of a second.read_ddi()builds its resolver index lazily.- With
ddi-l[full], validation first checks the document with libxml2 and runs the detailed xmlschema validator only when that check fails: valid documents validate in milliseconds, and reported issues are identical on both backends. iter_variables(),iter_questions()anditerparse_ddi()stream large documents and yield fully populated objects.- See Performance for measured timings.
HTTP API¶
- Optional Litestar service:
pip install 'ddi-l[server]'thenddi serveexposes validate, lint and convert over HTTP, with OpenAPI docs at/schema. Every endpoint is a thin wrapper overddi_l.operations, which the CLI also calls, so the two cannot disagree about whether a document is valid. - Self-describing OpenAPI:
/schemaserves Swagger UI, and a browser opening the root is redirected there. Every endpoint declares its request body, query parameters and a typed response schema with a captured example; "Try it out" arrives prefilled with a valid DDI instance. - Service-appropriate defaults: a 32 MiB request body cap, a bounded
number of concurrent jobs (
--max-jobs,503beyond it), the default schema preloaded at startup, no disk writes, CORS off unless origins are named, and the same XML hardening as the library. - JSON-LD output:
POST /v1/convert/jsonldandddi to-jsonldrender a study as linked data using the DDI Alliance's own DDI-RDF Discovery vocabulary ("Disco"), with Dublin Core and SKOS for labels and DDI URNs as IRIs. Covers DDI's discovery subset and is one-way by design, as the Disco specification intends;to-jsonremains the lossless, round-trippable format (XML to JSON and back returns the bytes it was given, and so does/v1/roundtrip). ddi_l.operations: the transport-neutral core, usable directly when embedding ddi-l in another application.
Model layer¶
- XSD-driven model generation: every model class is generated from the official DDI 3.3 XSD, the single source of truth for fields, element ordering, namespaces, and documentation. All 508 XSD complex types are covered.
- Generic XML engine:
from_xml()/to_xml()driven by the generated_FIELD_XML_MAP,_ATTR_XML_MAP, and_ELEMENT_ORDERtables, emitting children in XSD-declared order. - Stable DDI namespace prefixes: output uses the canonical
r:,s:,d:,l:,c:,a:,p:,pr:(ddiprofile), andprc:(process) prefixes rather than generatedns0/p0names, and serialization prefers a namespace's registered prefix over an auto-generated one. - Namespace constants: exported for every DDI module, including
DDI_PROFILE_NSandDATASET_NS. - Typed: ships
py.typed; the public API is fully annotated.
Tooling and distribution¶
- CLI:
ddi validate,ddi lint,ddi to-json,ddi to-jsonld,ddi from-json,ddi roundtrip,ddi versions, andddi serve. Commands take-o/--outputand--ddi-version;ddi validate --format jsonprints a JSON array (empty when valid) for scripts. - Optional lxml backend:
pip install ddi-l[full]enables faster parsing of large documents. - Generated API reference built from the source docstrings, alongside the guides.
- Every documented code sample is executed in CI, including the R examples.
- Python 3.11, 3.12, 3.13, and 3.14 supported and tested in CI.
- PyPI publishing via GitHub OIDC trusted publishing.
- Bilingual documentation in English and French, including a guide to using ddi-l from R through reticulate.
- Training curriculum: a 16-module progressive course with exercises, quizzes, an instructor guide, answer keys, and a bilingual cheat sheet.