Encyclopedia

How do you prove a project will work, and then prove it did?

Evidence, Evaluation, and Data

A project is shown likely to work by citing research behind each component and laying out a causal chain a reviewer can test. It is shown to have worked by measuring named indicators against baselines, on a schedule, using data sources secured before the work began.

Key takeaways

  • Evaluation is a spine through the proposal, not a closing section.
  • Cited evidence and generated evidence are separate obligations with separate costs.
  • Outputs prove delivery happened; outcomes prove anything changed.
  • A measure with no named data source is decoration.
  • The outcomes promised at submission become the rows of post-award reports.

What is a logic model, and why build one first?

A logic model is a one-page map of how a project turns resources into results, running from situation and inputs through activities, outputs, and tiered outcomes to impact. How to build a logic model covers the columns, the assumptions row most applicants omit, and the construction method.

Build it backward. Starting from the activities an organization already runs produces a model that justifies existing work; starting from the intended outcome forces the harder question of what would actually have to happen. Presented forward and built backward, a logic model rarely contains orphan activities — work with no outcome attached — which reviewers read as budget padding.

The test that keeps a logic model honest is the if-then chain between adjacent columns. The W.K. Kellogg Foundation states it plainly: “If you have access to them, then you can use them to accomplish your planned activities. If you accomplish your planned activities, then you will hopefully deliver the amount of product and/or service that you intended” (W.K. Kellogg Foundation, Logic Model Development Guide). Where a link reads as a leap, the model is missing a box. Federal reviewers score the artifact directly, weighing “the quality of the logic model or other conceptual framework underlying the proposed project, including how inputs are related to outcomes” (34 CFR 75.210).

Do you need a theory of change or a logic model?

A theory of change explains why an intervention should work and what must be true first; a logic model depicts what one funded project will do and produce. Theory of change versus logic model compares the two across scope, direction, and required contents, and covers the ceiling of accountability.

The Center for Theory of Change defines the instrument as “a comprehensive description and illustration of how and why a desired change is expected to happen in a particular context,” focused on “filling in what has been described as the ‘missing middle’ between what a program or change initiative does (its activities or interventions) and how these lead to desired goals being achieved” (Center for Theory of Change). A logic model is usually a subset of a larger theory of change — the operational slice one award pays for.

Which one a funder wants is decided by the notice rather than by the literature. Federal education regulations do not define a theory of change at all; they define a “logic model (also referred to as theory of action)” and score its quality (34 CFR 77.1). Large private and international funders more often ask for a theory of change by name and expect assumptions and a narrative. Retitling a logic model without adding causal rationales fools nobody who reads these for a living.

What separates an output from an outcome?

An output is a direct product of program activity and sits fully within staff control; an outcome is a change in the people or systems the project touched, which the project influences but cannot guarantee. Outputs, outcomes, and impact covers the controllability test, the three outcome tiers, and when a project may claim impact.

Delivering forty workshops is an output because the organization scheduled the room, recruited, and counted heads. Whether anyone learned anything is an outcome, and it can fail even when attendance is perfect. The Urban Institute states the boundary from the measurement side: end outcomes are “the consequences/results of what the program did, not what the program itself did” (Lampkin and Hatry, Urban Institute).

Impact is the word most often over-claimed. In federal usage it means a causally attributed effect, which requires a comparison condition — a control group, a matched comparison, a discontinuity, or a credible modeled counterfactual. A single-group before-and-after comparison is not one, however large the change. Program evaluation practice draws the same line: outcome evaluation “cannot determine what caused specific outcomes (causality), only whether they have been achieved” (CDC, Program Evaluation). Using a causal verb on a non-causal design is the fastest way to lose a reviewer who has run a study.

What belongs in a grant evaluation plan?

A grant evaluation plan states the questions, the indicators that answer them, the data sources, who collects and analyzes what, on what schedule, and what it costs. Writing an evaluation plan covers the eight components, the evaluation matrix, independence requirements, and how to price the work.

Two families of questions belong in every plan. Implementation questions ask whether the program was delivered as designed, at the intended dose, to the intended population. Outcome questions ask whether the intended changes occurred. Process work is what makes a disappointing outcome interpretable — without it, nobody can tell whether the theory was wrong or the service never reached anyone.

Sequencing is the underlying cause of most weak evaluation sections. A plan written last inherits whatever the narrative happened to promise and whatever money is left. The Office of Management and Budget states the ordering directly: “Rigorous evaluation planning must be grounded in a theory of change and take into account existing evidence and gaps in evidence” (OMB M-20-12). Timing also constrains which questions are answerable at all — the National Institute of Justice notes that outcome work belongs “at the end of the program’s development, when the program is stable and unlikely to change in fundamental ways” (NIJ, Plan a Program Evaluation from the Start).

What does evidence-based actually require?

“Evidence-based” means a named project component is supported by research at a defined level — not that the applicant has a track record and not that the idea sounds sensible. Levels of evidence in grant funding covers the four federal tiers, the study designs that reach each one, tiered-evidence grantmaking, and the clearinghouses that carry the ratings.

Two features of the definition do most of the work. The unit is the project component — “an activity, strategy, intervention, process, product, practice, or policy included in a project” — rather than the project as a whole, so a proposal earns a tier for a specific named practice (34 CFR 77.1). And all four tiers, including the lowest, count as evidence-based, so a notice that says only “evidence-based” is setting a low bar.

Design sets a ceiling that no amount of care can lift. A quasi-experimental study can meet What Works Clearinghouse standards with reservations but not without them, which means a matched-comparison study caps at the moderate tier regardless of execution quality; the handbook governing those ratings is published (What Works Clearinghouse Handbooks). The lowest tier is not a free pass either: it requires a logic model containing the component and citations of research supporting each causal link the model asserts.

Where does credible needs data come from?

Credible needs data comes from federal statistical agencies, program agencies, and state and local open-data systems, all free. Public data sources for needs statements is a directory organized by domain — population and income, labor market, health, education, housing, food, justice, environment, business, and nonprofit landscape — with the smallest published geography for each.

Geography, not topic, is the usual constraint. A county statistic used to describe three census tracts understates concentrated need, and a tract statistic used for a multi-county region is too noisy to defend. The Census Bureau frames the core tradeoff between survey products as one of precision versus currency, with population thresholds rather than preference deciding which product is published for a given place (Census Bureau, When to Use 1-year or 5-year Estimates).

Three habits separate defensible needs data from the rest. Label model-based small-area estimates as models rather than counts, because presenting a modeled tract figure as a headcount is the kind of error a statistically literate reviewer catches. Cite the publisher, product, table, geography, and vintage rather than “census data.” And always benchmark, because federal criteria ask reviewers to weigh “a comparison to local, State, regional, national, or international data” (34 CFR 75.210).

How do you choose performance measures you can sustain?

Choose performance measures that are valid, reliably collectable, feasible at the required cadence, sensitive enough to move inside the period of performance, and meaningful to both funder and community. Performance measurement and indicators covers the selection criteria, leading versus lagging indicators, targets and baselines, and the real cost of collection.

Performance measurement is not program evaluation, and the distinction decides what a proposal owes. Measurement is “the ongoing monitoring and reporting of program accomplishments, particularly progress toward pre-established goals” (CDC, Program Evaluation); evaluation explains why the numbers moved. A measures table alone is scored as an incomplete evaluation section.

Fewer measures collected well beat a long list collected badly, because every measure recurs in every reporting period for the life of the award. Two rules keep the set sustainable. Adopt any funder-prescribed common measures verbatim, since a locally reworded definition produces numbers the funder cannot aggregate. And pair every long-horizon outcome with a leading indicator that can move inside the grant period — a three-year project whose only measure is a five-year lagging indicator is structurally unable to show progress, no matter how well it is run.

How the pieces fit together

Five instruments do different jobs, and asking one to do another’s is the most common structural error in an evidence section. The table below distinguishes them.

InstrumentQuestion it answersWhen it is built
Needs dataHow large is the problem here?Before anything else
Theory of changeWhy should this work at all?Before the design
Logic modelWhat will we do and produce?With the design
Evaluation planHow will we know, and why?With the objectives
Performance measuresDid the numbers move?Every reporting period

The chain runs in one direction, and each link constrains the next. Needs data establishes the gap and, in doing so, supplies the baseline: federal regulation defines a baseline as “the starting point from which performance is measured and targets are set,” and judges whether a target is ambitious relative to that baseline (34 CFR 77.1). A statement of need written with a precise current condition therefore does double duty; one written vaguely damages two sections.

The logic model turns that gap into a causal claim, and the outcome boxes become the objectives. That ordering matters: goals, objectives, and activities written before the model produce commitments retrofitted to a need, which is exactly the misalignment reviewers are trained to notice. Each outcome box then earns an indicator, a named data source, and a measurement date, or it comes off the page. The Pennsylvania Coalition Against Rape compresses the whole relationship into one line: “a theory of change is what puts the logic in a logic model” (PCAR, Theory of Change and Logic Models).

The chain does not stop at submission. Federal performance reporting requires a comparison of accomplishments to the objectives established for the award, plus explanations of why goals were not met (2 CFR 200.329). The outcome boxes drafted during proposal week become the rows of a report filed years later, and the reporting obligations inherit whatever the model promised. That is the practical argument for treating evaluation as a spine rather than a section: a measure invented at the end of drafting is a measure someone else will have to produce, on a fixed deadline, from a data source nobody secured.

Frequently asked questions

When should evaluation planning start in a grant application?

Before the objectives are written and alongside the project design. Evaluation planning done last inherits promises it cannot measure and money that is already allocated. Baseline collection in particular has to be scheduled before implementation begins, or the design loses its comparison point and two sections contradict each other on the page.

What is the difference between program evaluation and research?

Program evaluation judges a specific program against its own goals; research seeks generalizable knowledge. The distinction has a regulatory consequence, because federal human-subjects rules attach to activities designed to contribute to generalizable knowledge. Work intended only to improve one local program often falls outside that definition, but get a written determination rather than assuming.

How do you report an evaluation that shows the program did not work?

Report it, with the implementation data that explains it. Federal performance reporting anticipates shortfalls and requires an explanation of why goals were not met, so under-performance is a reporting category rather than an admission. Process data showing dose, reach, and fidelity is what separates a wrong theory from a service that never arrived.

Do foundations expect the same evidence rigor as federal agencies?

Usually less formally and sometimes more consequentially, since a program officer’s judgment is not bounded by a published rubric. Foundation reporting typically asks for a short outcome set against stated goals with narrative explanation. The selection discipline is identical: fewer measures, defined precisely, drawn from sources that will still exist at report time.

Who should own performance measurement in a small organization?

One named person with authority to require data from program staff and protected time to close each reporting period. Splitting ownership between programs and finance without a designated owner is the most common structural cause of late reports and of accomplishments that cannot be reconciled with expenditures.

Can one evidence base support several proposals?

The citations transfer; the tier claim does not automatically. Each notice defines its own levels, sample requirements, and population-match rules, and tier vocabulary from one agency rarely maps onto another’s definitions. Reuse the research library and rewrite the claim against the definitions section of the notice you are answering.

Sources

  1. U.S. Department of Education. 34 CFR 77.1 — Definitions that apply to all Department programs. Cornell Legal Information Institute. https://www.law.cornell.edu/cfr/text/34/77.1 (accessed 2026-08-11)
  2. U.S. Department of Education. 34 CFR 75.210 — General selection criteria. Cornell Legal Information Institute. https://www.law.cornell.edu/cfr/text/34/75.210 (accessed 2026-08-11)
  3. W.K. Kellogg Foundation. Logic Model Development Guide (2004). Hosted by NACCHO. https://www.naccho.org/uploads/downloadable-resources/Programs/Public-Health-Infrastructure/KelloggLogicModelGuide_161122_162808.pdf (accessed 2026-08-11)
  4. Center for Theory of Change. What is Theory of Change? https://www.theoryofchange.org/what-is-theory-of-change/ (accessed 2026-08-11)
  5. Centers for Disease Control and Prevention. Program Evaluation. https://www.cdc.gov/evaluation/php/about/index.html (accessed 2026-08-11)
  6. U.S. Office of Management and Budget. M-20-12: Phase 4 Implementation of the Foundations for Evidence-Based Policymaking Act of 2018. https://www.whitehouse.gov/wp-content/uploads/2020/03/M-20-12.pdf (accessed 2026-08-11)
  7. National Institute of Justice. Plan a Program Evaluation from the Start. U.S. Department of Justice. https://nij.ojp.gov/topics/articles/plan-program-evaluation-start (accessed 2026-08-11)
  8. Lampkin, Linda M., and Harry P. Hatry. Key Steps in Outcome Management. Urban Institute. https://www.urban.org/sites/default/files/publication/42736/310776-Key-Steps-in-Outcome-Management.PDF (accessed 2026-08-11)
  9. Institute of Education Sciences. What Works Clearinghouse Handbooks. U.S. Department of Education. https://ies.ed.gov/ncee/wwc/handbooks (accessed 2026-08-11)
  10. U.S. Census Bureau. When to Use 1-year or 5-year Estimates. American Community Survey. https://www.census.gov/programs-surveys/acs/guidance/estimates.html (accessed 2026-08-11)
  11. Pennsylvania Coalition Against Rape. Theory of Change and Logic Models. https://pcar.org/sites/default/files/resource-pdfs/tab_2018_logic_models_508.pdf (accessed 2026-08-11)
  12. U.S. Office of Management and Budget. 2 CFR 200.329 — Monitoring and reporting program performance. Cornell Legal Information Institute. https://www.law.cornell.edu/cfr/text/2/200.329 (accessed 2026-08-11)

Articles in this section

  1. How to Build a Logic ModelWhat is a logic model and how do you build one?
  2. Theory of Change vs Logic ModelWhat is the difference between a theory of change and a logic model?
  3. Outputs, Outcomes, and ImpactWhat is the difference between an output and an outcome?
  4. Writing an Evaluation PlanHow do you write an evaluation plan for a grant?
  5. Levels of Evidence in Grant FundingWhat does evidence-based mean in a grant application?
  6. Public Data Sources for Needs StatementsWhere do you find data for a grant needs statement?
  7. Performance Measurement and IndicatorsHow do you choose performance measures for a grant?

Reading Is Research. Searching Is Progress.

Put the encyclopedia to work — search every open grant and get matched by eligibility.