Data Dictionaries for Mobile Health Programs
A data dictionary is a structured list of the fields your program collects, with a clear definition for each one. It records what each field means, who enters it, which values are allowed, and where the data live.
It is one of the most useful documents a mobile health program can bring to an early conversation with a researcher. It shows what the program already collects and whether a study might use existing measures instead of adding forms or questions for staff and patients.
Why it changes the conversation
Without a dictionary, a researcher may have to make assumptions about what the program collects. A protocol built around measures the program does not have may require new forms, more time during visits, additional training, and more questions for patients.
With a dictionary, the conversation starts with the available data. If a program already records referral completion using a stated definition and knows the field’s completeness rate, the research team can decide whether that measure fits the question. Using an appropriate existing measure can protect the workflow, reduce burden, and shorten the time needed to begin the study.
A written definition also reduces the risk that a field will be interpreted differently later, especially when a finding depends on what counted as a completed referral or another outcome.
What goes in one
One row per field. At minimum:
- Field name as it appears in your system
- Plain-language meaning, written so someone outside the program understands it
- Where it lives, naming the system, form, or file
- Who enters it, by role and at what point in the visit
- Type and permitted values, including the full list for any dropdown
- Whether it is required
- Completeness, as the share of eligible records with a usable value, and from what date
- Known limitations, in plain words
Add these where they apply:
- The date the field was introduced or last changed
- Any earlier definition, and when it changed
- Whether the field contains identifiers or free text
- Which sites or services use it, if not all
- The source of a clinical measure, including version
- Who to ask about it
Naming a contact is important because knowledge about a field’s history or purpose may otherwise exist only with one staff member.
Build it from what you have
You do not need a system for this. A spreadsheet works.
- Start from the inventory of sources. See Data Inventory and Quality.
- Take one source at a time and list its fields. Export the field list where you can.
- Fill in meaning, entry, and permitted values by asking the people who use the field.
- Run a completeness check for the fields that matter and record the result.
- Write the known limitations honestly, including fields you no longer trust.
- Note the fields you collect but never use.
Fields that no one uses may be candidates for removal. Reducing unnecessary collection can improve completion of the measures that remain.
Begin with the fields closest to your priority question. A well-defined set of fifteen fields may be more useful than an incomplete dictionary of the entire system, and the document can grow over time.
Keep it current
An outdated dictionary can give a research partner an inaccurate picture of the program’s data.
Review it when a form changes, when a system is upgraded, when a site or service is added, and on a set schedule regardless. Record the date of each review on the document. When a definition changes, keep the old one and the date it changed, because any analysis crossing that date has to account for it.
Using it in a partnership conversation
Send the dictionary before the first substantive meeting so the research team can review it in advance.
Expect a researcher to ask which fields are reliable enough to build on, whether any existing measure is close to a validated instrument, what the completeness rate is for the specific population in question, and whether anything can be linked to another source.
Be direct about weak or incomplete fields. A researcher who learns about a limitation early may be able to adjust the question, measures, or design. A problem discovered after the protocol is written can require substantial revision and additional staff time.
Frequently asked questions
How is a data dictionary different from a measure set?
A measure set is what you decided to track and why. A data dictionary describes every field that exists, including ones you no longer use. The measure set points at the question. The dictionary describes the raw material. See the Mobile Clinic Outcomes and Measures Library.
We use a parent organization’s EHR. Can we still write one?
Yes. A data dictionary can be especially useful in that situation. You may not control the system, but you can document the fields your program relies on and how your staff use them, which may differ from how the system’s designers intended. Ask the data team for the field list. See Getting Your Data Out of a Health System EHR.
Who should own the dictionary?
Assign one person to coordinate updates and give that person the time and authority to do the work. Other staff can contribute subject-matter knowledge, but responsibility for maintaining the document should be clear.
Does the dictionary contain patient data?
A data dictionary should describe fields rather than include patient records or values. When it contains definitions only and no confidential system details, it can generally be shared with a prospective partner before patient-level data are disclosed. Review the document for sensitive information before sharing it.
Related resources
Have a research question or a program worth studying?
GMHRC helps researchers and mobile healthcare programs find each other and plan work that is useful to both.