How do you write an evaluation plan for a grant?
Writing an Evaluation Plan
An evaluation plan states what questions the project will answer, what indicators will answer them, where the data comes from, who collects and analyzes it, on what schedule, and what it costs. Build it from the objectives, not after them, and fund it as a budget line.
Current figures — verified 2026-08-11
Item Value Source Evaluation budget rule of thumb 10 percent of total program cost National Institute of Justice Typical evaluation budget range, education and training proposals 5 to 15 percent of total project budget NC State Research Development Office SIPPRA evaluation set-aside up to 15 percent of the project award, not contingent on achieving outcomes Urban Institute, Pay for Success Initiative These figures change. Verify against the linked source before relying on them. Report an outdated figure
Key takeaways
- Evaluation questions come from the objectives, numbered and answerable.
- Every indicator needs a named data source before it goes on the page.
- Process evaluation is what makes a null outcome result interpretable.
- An unfunded evaluation plan reads as a plan nobody intends to execute.
- Reviewers score what is written, not what an evaluator will later develop.
What is an evaluation plan in a grant proposal?
An evaluation plan is the written design for how a project will be measured and judged. Federal regulation defines evaluation as “an assessment using systematic data collection and analysis of one or more programs, policies, practices, and organizations intended to assess their implementation, outcomes, effectiveness, or efficiency” (34 CFR 77.1).
An evaluation plan is distinct from a performance measurement table, and conflating the two is a scored error. Performance measurement monitors accomplishments against pre-established goals; evaluation explains why the numbers moved. The Office of Management and Budget treats the two as separate components of evidence (OMB M-19-23), and a section containing only a measures table is marked incomplete against an evaluation criterion.
Reviewers score the plan against published criteria. The Department of Education’s evaluation criterion rewards methods that “provide guidance for quality assurance and continuous improvement” and that supply “formative, diagnostic, or interim data that is a periodic assessment of progress toward achieving intended outcomes” (34 CFR 75.210). Write the plan against the specific project — as the NC State Research Development Office puts it, “Reviewers can see a generic evaluation plan a mile away because they are likely reading numerous proposals with the same one-size-fits-all plan inserted” (NC State RDO, 2020). Everything else in the evidence, evaluation, and data hub feeds this section.
What are the components of a grant evaluation plan?
A complete evaluation plan has eight components, and a reviewer looking for a missing one will find it. There are eight parts:
- Evaluation questions, numbered, split into implementation and outcome questions, each traceable to a logic model element.
- Indicators and measures, one row per question, with construct, operational definition, and instrument.
- Data sources, named specifically — the district student information system, the state wage record file — not described generically.
- Collection methods and schedule, stating who collects what from whom, at which waves, at what assumed response rate.
- Sample and comparison strategy, including expected attrition and, for causal designs, the smallest detectable effect.
- Analysis plan, naming the estimator, unit of analysis, handling of missing data, and pre-specified subgroups.
- Timeline, aligned to budget periods and reporting deadlines, with baseline collection before implementation begins.
- Use and dissemination, covering how interim findings feed decisions and who receives results.
The CDC framework organizes the same work as six steps — assess context, describe the program, focus the evaluation questions and design, gather credible evidence, generate and support conclusions, and ensure use and share lessons learned (CDC Program Evaluation Framework, 2024). In CDC’s planning guidance, the program description “often includes a logic model,” and question prioritization rests on a shared understanding of the theory of change that the model expresses (CDC, Developing an Effective Evaluation Plan). Either structure covers the same ground; use the funder’s if the notice specifies one.
How do you turn objectives into evaluation questions?
Evaluation questions come from the objectives, restated as things a study could answer. Take each objective, ask what evidence would show it was achieved, and write that as a question. An objective committing to a change in employment produces a question about employment; an objective committing to a delivery volume produces an implementation question, not an outcome question.
Two families of questions belong in every plan. Implementation and process questions ask whether the program was delivered as designed, at the intended dose, to the intended population. Outcome questions ask whether the intended changes occurred. A process evaluation is the only thing that makes a disappointing outcome interpretable — without it, nobody can tell whether the theory was wrong or the service never reached anyone.
Funders use two overlapping vocabularies for the split. Formative evaluation runs while the program operates and feeds improvement; summative evaluation runs at the end and asks whether goals were met, requiring baseline and end-of-program data (NC State RDO, 2020). The CDC taxonomy separates process evaluation, which assesses “how well program implementation followed the original plan,” from outcome evaluation, which measures achievement of intended outcomes but “cannot determine what caused specific outcomes” (CDC Approach to Program Evaluation).
Timing constrains which questions are answerable. The National Institute of Justice is direct: “Conducting certain evaluations, like outcome evaluations, is difficult when a program is too new because program elements, strategies or procedures often still are being adjusted and finalized,” and impact work belongs “at the end of the program’s development, when the program is stable and unlikely to change in fundamental ways” (NIJ, 2015).
What makes an indicator, baseline, and target defensible?
A defensible indicator is specific, observable, understandable, relevant, time-bound, and reliable. Urban Institute indicator guidance names those criteria directly: unique and unambiguous, practical and cost-effective to collect, comprehensible, measuring an important and predictive dimension, covering a specified period, and producing accurate, unbiased, verifiable data (Urban Institute and The Center for What Works). Add one test the textbooks skip: an indicator needs a data source that already exists or that the budget pays to create.
Baselines make targets assessable. Federal regulation defines a baseline as “the starting point from which performance is measured and targets are set,” a performance target as “a level of performance that an applicant would seek to meet during the course of a project,” and judges whether a target is ambitious relative “to the context of the relevant performance measure and the baseline for that measure” (34 CFR 77.1). A target with no baseline behind it cannot be judged ambitious or timid, so it cannot be scored at all.
Set targets from three anchors: the effect sizes reported in the research cited for your evidence claim, your own historical trend on the same measure, and the gap between your value and an external benchmark. When no baseline exists, say so and budget a baseline collection period in the first year — the invented starting value will be compared against real data later. The choices here follow from how the project defines outputs, outcomes, and impact.
What data sources and methods should an evaluation plan name?
An evaluation plan should name every system, instrument, and record set it depends on, and state the access status of each. Prefer administrative data already collected by an operating system; use primary collection only where administrative data cannot answer the question. Administrative data is cheaper, arrives without participant burden, and does not depend on response rates.
Four access questions belong on the page, because reviewers penalize designs that assume data nobody has agreed to share. State whether a data-sharing agreement is executed, in negotiation, or not yet started. State the refresh lag of each source, since a system that closes quarterly cannot support a monthly feedback loop. State any statutory constraint — student records under FERPA, health records under HIPAA — and how the design satisfies it. State the retention and destruction plan for identifiable data.
Human-subjects review applies more often than applicants expect. Federal policy defines research as “a systematic investigation, including research development, testing, and evaluation, designed to develop or contribute to generalizable knowledge,” and a human subject as a living individual about whom an investigator obtains information through intervention or interaction, or obtains identifiable private information (45 CFR 46.102). Evaluation designed only to improve one local program is often outside that definition; the same work designed to produce publishable findings is usually inside it. Decide early, name the reviewing board, and budget the fee and the review weeks.
Address data quality across six dimensions: validity, reliability, completeness, timeliness, consistency across sites and years, and security. Naming expected missingness preempts the most common objection to a survey-based design.
Who should conduct the evaluation, and what does it cost?
Who conducts the evaluation is decided by what the notice of funding opportunity requires and by what the design can support. Internal staff can credibly run performance measurement, continuous improvement, and process evaluation. External evaluators are required whenever a notice uses the word “independent,” and are the norm for outcome and impact designs meant to produce citable evidence.
The regulatory bar is higher than hiring someone. An independent evaluation is “an evaluation of a project component that is designed and carried out independently of, but in coordination with, the entities that develop or implement the project component” (34 CFR 77.1). A subcontractor who takes design direction from the program team may not satisfy it. The government-wide analogue is OMB’s independence and objectivity standard, one of five federal evaluation standards alongside relevance and utility, rigor, transparency, and ethics (OMB M-20-12).
A hybrid split scores well and costs less than full outsourcing: internal staff own routine data collection and improvement cycles, while an external evaluator owns design, instrumentation, analysis, and reporting. Selecting that evaluator is a documented step, not a name in a sentence. Ask whether the evaluator specializes in the method the design requires, has evaluated this program type before, and will participate in proposal development (NC State RDO, 2020). Attach a scope of work, a capability statement, and a letter of commitment. The American Evaluation Association’s Guiding Principles — systematic inquiry, competence, integrity, respect for people, and common good and equity — are worth naming as the ethical baseline (American Evaluation Association).
Cost scales with design, not with organizational preference. The commonly cited benchmarks appear in the current figures above; treat them as heuristics that the notice’s own cap or the design’s requirements will override. Budget beyond the evaluator’s fee: collection burden on program staff, participant incentives, review board fees, data licenses, and the cost of the system that will hold the data. The Administration for Children and Families ties capacity to money directly, committing to ensure evaluators have appropriate expertise and to “allocating sufficient resources for evaluation activities” (ACF Evaluation Policy).
What does an evaluation matrix look like?
An evaluation matrix is a single table mapping each evaluation question to the indicator that answers it, the data source, the collection method, the frequency, and the person accountable. The matrix is the most useful artifact in the plan, because it exposes gaps that narrative hides — a question with no indicator, an indicator with no source, a source with no owner. The matrix below is illustrative, drawn from a hypothetical workforce training project.
| Evaluation question | Indicator | Data source | Method | Frequency | Responsible party |
|---|---|---|---|---|---|
| Was training delivered at the planned dose? | Contact hours delivered per cohort against 120-hour plan | Course management system | Administrative extract | Monthly | Program director |
| Did participants gain the intended competencies? | Share of enrollees passing the certification exam | State certification registry | Records match | Per cohort | External evaluator |
| Did navigation reach the intended caseload? | Contacts per participant per month against 1:40 caseload standard | Case management system | Administrative extract | Monthly | Navigation supervisor |
| Did completers obtain employment in the target occupation? | Share employed in target occupation within 90 days of exit | State wage records under data-sharing agreement | Records match | Quarterly, with lag | External evaluator |
| Did earnings improve relative to pre-enrollment? | Change in median quarterly earnings, pre- versus post-enrollment | State wage records | Pre-post comparison, reported as a performance measure | Semiannual | External evaluator |
Two design choices in that table are worth copying. The employment row names the legal instrument that makes the data reachable. The earnings row labels itself a performance measure rather than a causal estimate, which is how to report a pre-post comparison without over-claiming.
Set the reporting cadence against the award’s actual deadlines. Federal performance reporting requires a comparison of accomplishments to the objectives established for the award, on a frequency the awarding agency sets within the bounds of the Uniform Guidance (2 CFR 200.329). Work backward from each due date through data close, analysis, and internal review, and check that every measure can close in time. Then state which findings trigger which decisions, at which points, and who has authority to act — federal regulation defines continuous improvement as using implementation and outcome data “to inform necessary changes throughout the project,” which is a commitment to act, not only to observe.
What goes wrong in a grant evaluation plan?
Seven failures account for most of the points lost in evaluation sections, and each is visible on the page. There are seven:
- A plan that measures outputs only. Service counts under an outcomes heading show that delivery has not been separated from benefit.
- Indicators with no data source. Every measure needs a named system, instrument, or record set, plus that source’s access status.
- A design that cannot detect the effect it claims. A small sample with no stated minimum detectable effect produces an uninformative null, which damages the program’s future more than no study would.
- Causal language on a non-causal design. Reserve “increased” and “caused” for designs with a comparison condition.
- An evaluator named without a scope or a budget line. A name with no statement of work, no letter, and no cost reads as a placeholder.
- Data access assumed rather than secured. No agreement, no review board plan, no analysis of the privacy statute governing the records.
- Everything written in the future tense. “The evaluator will develop an evaluation plan” is scored as an absent plan.
The underlying cause of most of them is sequencing. A plan written last, after the narrative and budget are set, inherits whatever the narrative happened to promise and whatever money is left. A plan written from the logic model alongside the project design has a chance of being executable. OMB puts the ordering plainly: “Rigorous evaluation planning must be grounded in a theory of change and take into account existing evidence and gaps in evidence” (OMB M-20-12) — an argument for drafting the theory of change before anyone opens the evaluation template.
Frequently asked questions
Is an evaluation plan required if the notice does not ask for one?
Not required, but usually worth including. A short plan demonstrates that the applicant can distinguish success from activity and would notice failure. For small foundation grants, a half-page describing the questions, measures, sources, and reporting cadence is enough.
Do you need institutional review board approval for a program evaluation?
Institutional review board approval depends on whether the activity meets the regulatory definition of research involving human subjects (45 CFR 46.102). Evaluation intended solely to improve one local program often falls outside it; evaluation designed to produce generalizable, publishable findings usually does not. Get a written determination rather than assuming.
What is the difference between an evaluation plan and a performance measurement plan?
A performance measurement plan tracks whether the intended numbers moved. An evaluation plan explains why, and an impact design establishes whether the program caused the movement. Funders scoring an evaluation criterion expect both the measures and the explanatory design; a measures table alone is scored as partial.
Can evaluation activities continue after the period of performance ends?
Usually not on grant funds. Follow-up windows that extend past the period of performance need either a no-cost extension, a separate funding source, or a design change that pulls the measurement inside the award. Plan the final analysis against the closeout deadline rather than the calendar.
Who owns the evaluation data and the findings?
Settle it in the evaluator’s scope of work before submission. Federal evaluation standards treat transparency as an expectation that findings are released regardless of what they show (OMB M-20-12), so a contract that lets either party suppress unfavorable results conflicts with the standard the funder is applying.
Related topics
- Evidence, Evaluation, and Data — the hub for proving a project will work, and then proving it did
- Performance Measurement and Indicators
- Levels of Evidence in Grant Funding
- Public Data Sources for Needs Statements
- Writing the Project Design
Sources
- U.S. Department of Education. 34 CFR 77.1 — Definitions that apply to all Department programs. Legal Information Institute, Cornell Law School. https://www.law.cornell.edu/cfr/text/34/77.1 (accessed 2026-08-11)
- U.S. Department of Education. 34 CFR 75.210 — General selection criteria. Legal Information Institute, Cornell Law School. https://www.law.cornell.edu/cfr/text/34/75.210 (accessed 2026-08-11)
- Office of Management and Budget. M-20-12, Phase 4 Implementation of the Foundations for Evidence-Based Policymaking Act of 2018: Program Evaluation Standards and Practices. March 10, 2020. https://www.whitehouse.gov/wp-content/uploads/2020/03/M-20-12.pdf (accessed 2026-08-11)
- Office of Management and Budget. M-19-23, Phase 1 Implementation of the Foundations for Evidence-Based Policymaking Act of 2018: Learning Agendas, Personnel, and Planning Guidance. July 10, 2019. https://www.whitehouse.gov/wp-content/uploads/2019/07/M-19-23.pdf (accessed 2026-08-11)
- Kidder, D. P., Fierro, L. A., Luna, E., et al. CDC Program Evaluation Framework, 2024. MMWR Recommendations and Reports 73(RR-6):1–37. https://www.cdc.gov/mmwr/volumes/73/rr/rr7306a1.htm (accessed 2026-08-11)
- Centers for Disease Control and Prevention. CDC Approach to Program Evaluation. https://www.cdc.gov/evaluation/php/about/index.html (accessed 2026-08-11)
- Centers for Disease Control and Prevention. Developing an Effective Evaluation Plan. https://www.cdc.gov/tobacco/stateandcommunity/tobacco-control/pdfs/developing_eval_plan.pdf (accessed 2026-08-11)
- National Institute of Justice. Plan for Program Evaluation from the Start. March 1, 2015. https://nij.ojp.gov/topics/articles/plan-program-evaluation-start (accessed 2026-08-11)
- NC State University Research Development Office. Evaluation Plan: A Key Component of Education and Training Proposals. March 5, 2020. https://research.ncsu.edu/rdo/evaluation-plan-a-key-component-of-education-and-training-proposals/ (accessed 2026-08-11)
- Lampkin, L. M., Winkler, M. K., Kerlin, J., Hatry, H. P., et al. Building a Common Outcome Framework To Measure Nonprofit Performance. Urban Institute and The Center for What Works. https://www.urban.org/sites/default/files/publication/43036/411404-Building-a-Common-Outcome-Framework-To-Measure-Nonprofit-Performance.PDF (accessed 2026-08-11)
- U.S. Department of Health and Human Services. 45 CFR 46.102 — Definitions for purposes of this policy. Legal Information Institute, Cornell Law School. https://www.law.cornell.edu/cfr/text/45/46.102 (accessed 2026-08-11)
- Administration for Children and Families, U.S. Department of Health and Human Services. ACF Evaluation Policy. November 9, 2021. https://www.acf.hhs.gov/opre/report/acf-evaluation-policy (accessed 2026-08-11)
- American Evaluation Association. Guiding Principles for Evaluators. https://www.eval.org/About/Guiding-Principles (accessed 2026-08-11)
- Office of Management and Budget. 2 CFR 200.329 — Monitoring and reporting program performance. Legal Information Institute, Cornell Law School. https://www.law.cornell.edu/cfr/text/2/200.329 (accessed 2026-08-11)
- Urban Institute, Pay for Success Initiative. Social Impact Partnerships to Pay for Results Act (SIPPRA). https://pfs.urban.org/node/2502.html (accessed 2026-08-11)
Continue in this section
- How to Build a Logic ModelWhat is a logic model and how do you build one?
- Outputs, Outcomes, and ImpactWhat is the difference between an output and an outcome?
- Performance Measurement and IndicatorsHow do you choose performance measures for a grant?
- Levels of Evidence in Grant FundingWhat does evidence-based mean in a grant application?