Research Methods for Mobile Health Programs
Research begins with a plan. Before anything is collected, you should be able to say what question you are answering, what claim you need to support, what design can support that claim, and what you will do with the answer.
The design you choose should match the claim you need to make. Counting how many patients you served supports a claim about reach. Saying that your program caused an improvement in blood pressure control requires a design that addresses other plausible explanations, usually with a comparison and data collected over time.
Many questions asked by mobile health programs do not require a randomized trial. Programs do need to know where their design sits on the range between counting and causation, so they do not claim more than the design supports.
Evaluation, quality improvement, and research
These three overlap heavily. They often use the same data, the same measures, and the same methods. What separates them is purpose.
| Purpose | Typical question | Who it is for | |
|---|---|---|---|
| Quality improvement | Make this program work better here | Are we doing what we intended, and is it improving? | Your team |
| Program evaluation | Judge whether this program is working and worth continuing | Did the program achieve what it set out to do? | Leadership, funders, boards |
| Research | Produce knowledge that applies beyond this program | Does this approach work, and for whom? | The field |
The boundaries blur in practice. An evaluation designed carefully enough to inform other programs starts to look like research. A quality improvement project that tests an intervention with a research design may be research whatever it is called.
Two practical consequences follow.
The label changes the review pathway. Research involving human subjects generally requires review. Quality improvement often does not. That difference is not yours to decide informally. Describe your purpose, activities, data, and intended use accurately, and let the authorized office determine which it is. See IRB for Beginners.
The label does not change the quality of the design. A weak evaluation is not made stronger by calling it quality improvement, and a strong quality improvement project is not made weaker by not being called research. The design determines what you can claim.
A research plan is most of the work
Collecting data without a plan produces a spreadsheet nobody can interpret. Write these down before collection starts.
The decision. What will change depending on the answer. If nothing changes, reconsider whether the study is worth doing.
The question. Specific enough to answer. “Does our program help people” is not a question. “Among patients seen at our three rural sites, did the share completing a follow-up visit within 90 days change after we added appointment reminders” is.
The claim you need to support. Description, association, or cause. This determines the design more than anything else on the list.
The population and setting. Who is included, who is excluded, which sites, what period.
The measures. Exact numerator, denominator, and permitted values for each. See the Mobile Clinic Outcomes and Measures Library.
The data sources. Where each measure comes from and how complete it is. See Data Inventory and Quality, and Data Dictionaries.
The comparison, if any. What you are comparing against, and why that comparison is fair.
The analysis plan. How you will analyze the data, written before you look at it. This is the step most often skipped and the one that most protects you from finding a pattern that is not there.
How much data you need. Enough to detect a difference that would matter. A statistician can answer this quickly and it is a good first question for a research partner.
Timeline and responsibilities. Who does what, by when.
The review and privacy pathway. IRB determination, consent, and data protections. See Data Privacy and Data-Sharing Standards.
How results get back. To staff, to participants, to the community. See Engaging Communities Across the Whole Research Process.
The limitations you already know. Writing them down at the start is more credible than discovering them at the end.
Observational, experimental, and quasi-experimental
These three terms describe how a study assigns or observes an intervention.
Observational. You watch what happens without assigning anyone to anything. You describe who came, what they received, and what followed. Most program data are observational.
Experimental. The study team controls who receives the program, usually by randomizing. With an adequate sample, randomization helps make the compared groups similar at the start, including on factors the study did not measure.
Quasi-experimental. There is a comparison, or a structure that mimics an experiment, but assignment is not random. A new site opens and an old one has not yet. A policy starts in one county and not the neighboring one. You compare before and after across many time points.
Quasi-experimental designs can be practical when randomization is not feasible. The stronger designs can support cautious causal conclusions when their assumptions are reasonable and the analysis addresses plausible alternative explanations.
The range of designs
There is a range between counting people served and a randomized trial. The labels you may encounter for this are levels of evidence or a hierarchy of evidence. The useful way to think about it is simpler: what claim can this design support, and what else could explain the result?
| Design | What it answers | What it cannot rule out |
|---|---|---|
| Descriptive counts | Who came, how many, what they received | Anything about effect |
| Cross-sectional | What things look like at one point, and what is associated with what | Direction of the relationship, and other causes |
| Pre-post, single group | Whether the measure changed after the program started | That something else changed at the same time |
| Interrupted time series | Whether the trend changed at the point the program started | Another event at the same moment, though this is a much stronger design than pre-post |
| Comparison group, non-randomized | Whether the change differed from a group that did not receive the program | That the groups differed to begin with in ways you did not measure |
| Stepped wedge | Whether outcomes changed as sites received the program in randomized order | Time trends, contamination between sites, and changes during rollout |
| Randomized controlled trial | Whether the program caused the difference | Bias from loss to follow-up, nonadherence, measurement problems, or limited generalizability |
Two things are worth noticing.
Pre-post is weaker than most people assume. Measuring before and after a single time is vulnerable to seasonality, to other things happening at once, and to the tendency for unusual measurements to drift back toward average on their own. It is still useful. It is not evidence that your program caused the change.
Interrupted time series is stronger than most people assume. Many measurements before and many after let you see whether the existing trend changed at the point your program started. It is one of the stronger quasi-experimental designs and can be used without a separate control group. It may fit programs with reliable routine data collected at enough time points before and after an intervention.
Comparison groups
A comparison group is a set of people, sites, or periods that did not receive the program, or received it later or differently, whose outcomes you compare with those who did.
Without one, you can observe that something changed. You cannot say your program changed it. Other explanations remain open: the season, a policy change, another program, a change in who was coming through the door, or the ordinary tendency of extreme measurements to move back toward average.
This is why senior leaders, policymakers, and some funders ask for a comparison group. They want evidence that addresses alternative explanations before they act on the result.
| Type | What it is | Strengths and cautions |
|---|---|---|
| None, single group | Before and after in the same group | Simple and often sufficient for internal use. Cannot support a causal claim. |
| Historical | The same program in an earlier period | Easy from existing data. Anything else that changed over time is confounded with the program. |
| Concurrent, non-equivalent | Another site or community not yet served | Same period, so shared events affect both. The two places may differ in ways that matter. |
| Matched | Individuals or sites matched on measured characteristics | Reduces measured differences. Cannot match on what was never measured. |
| Waitlist or delayed start | People or sites who receive the program later | Can fit a planned expansion in which all sites are scheduled to receive the program. Time trends still matter. |
| Randomized | Assignment decided by chance | Supports causal inference when conducted and analyzed well. May be infeasible or inappropriate. |
The waitlist option deserves attention from mobile programs. If you are adding sites over the next two years, the sites not yet reached may provide a comparison group. Planning the order deliberately and collecting comparable data at all sites from the start can create a much stronger evaluation. The added data collection, coordination, and analysis still require staff time and should be included in the budget. Randomizing the order may create a stepped wedge design.
Match the design to the purpose
| Who is asking | What they usually need | Design that fits |
|---|---|---|
| Your own team | Whether the program is running as intended | Counts, quality improvement cycles, pre-post |
| Your board or a local funder | Whether the program is reaching people and delivering care | Descriptive measures with context, pre-post |
| A health system deciding whether to sustain the program | Whether it changes outcomes and what it costs | Comparison group, cost per outcome |
| A state agency or policymaker | Whether this would work elsewhere and at scale | Quasi-experimental design with an appropriate comparison or multiple observations over time |
| A peer-reviewed journal or national funder | Depends on the claim | Stronger quasi-experimental, stepped wedge, or randomized |
Start from who needs to be convinced and what they need to be convinced of. Then choose the least burdensome design that supports that claim.
The failure that matters
The common problem in this field is not that programs use simple designs. Simple designs are appropriate for most program questions and they are what the data support.
The problem is using a simple design and then making a claim it cannot carry. “Blood pressure control improved by 12 percent among our patients” is defensible from pre-post data. “Our program improved blood pressure control by 12 percent” is a causal claim, and without a comparison group it is not supported.
The second version is more persuasive right up until someone who knows the difference reads it. Then it costs you the credibility the first version would have kept.
Say what your design supports. Name the limitation before someone else does.
Frequently asked questions
Do we need a control group?
Not for most program questions. You need one when you want to say your program caused a change, and when the audience will act on that claim. Ask who needs convincing and what they need to believe.
Is a randomized trial always better?
It is a strong design for showing cause in the setting where it was conducted. It may be infeasible, unaffordable, or inappropriate when there is no ethical or practical way to assign the intervention. Choose the strongest feasible design that fits the question, setting, and available resources.
We only have a year of data. What can we do?
Descriptive or pre-post work may be possible. If you have reliable routine data going back far enough, an interrupted time series may also be feasible. If the program is expanding, ask a methodologist whether collecting comparable data at sites not yet served could support a future comparison. Include the added work in the project plan and budget.
Who decides whether this is research or quality improvement?
The authorized institutional office or IRB, based on an accurate description of what you intend to do. See IRB for Beginners.
Do we need a statistician?
For descriptive work, usually not. For anything involving a comparison, a trend over time, or a claim about effect, get advice early rather than after collection. A CTSA hub biostatistics service may offer an initial consultation or other support. Eligibility, cost, and capacity vary. See Getting Your Data Out of a Health System EHR.
Authoritative resources
- CDC Program Evaluation Framework
- Interrupted time series analysis, Kontopantelis and colleagues, BMJ
- The stepped wedge cluster randomised trial: rationale, design, analysis, and reporting, BMJ
- Quality Improvement Activities FAQs, HHS Office for Human Research Protections
Related resources
Have a research question or a program worth studying?
GMHRC helps researchers and mobile healthcare programs find each other and plan work that is useful to both.