Start with the worked examples to see complete reasoning, then use the shorter pattern library for variation. Level guidance and frameworks show how the same task changes as the evidence, audience, or assignment becomes more demanding.
A strong evaluation report makes the judgment process visible: readers can see what was evaluated, which criteria or questions controlled the assessment, where the evidence came from, what the results show, which limitations constrain the conclusions, and how any recommendation follows from the findings rather than from preference alone.
Define the evaluation purpose, intended users, questions, criteria, scope, and decision context before collecting or summarizing evidence.
Describe methods and data sources at enough depth for readers to understand how the findings were produced and what the evidence cannot establish.
Present results separately from interpretation when that distinction improves traceability.
Connect conclusions to explicit criteria and evidence rather than relying on general impressions such as successful or ineffective.
Tailor recommendations to the evaluator’s mandate, evidence strength, feasibility, stakeholder effects, and the decisions the report is actually meant to inform.
Worked format lab
See complete reasoning, not just isolated lines
Use these fuller examples to see what changes between a recognizable pattern and a finished piece of writing. The examples are original or explicitly illustrative, so they demonstrate structure without inventing real-world evidence.
Worked example 1Pilot program evaluation
Illustrative internal evaluation: a fictional six-week onboarding pilot for 48 new hires.
Executive summary\nThe pilot was evaluated against three questions: whether participants completed the core onboarding sequence, whether they could identify role-critical resources by the end of week two, and whether support requests decreased compared with the previous cohort. Forty-four of 48 participants completed the sequence. In a short end-of-week-two task check, 37 of 42 respondents located all three required resources without assistance. Support records show 0.9 onboarding-related contacts per participant compared with 1.4 in the prior cohort. The comparison is useful but not experimental: the cohorts differed in team mix and the follow-up period is short.\n\nConclusion\nThe available evidence supports continuing the pilot for one more cohort while improving the resource task and tracking role-specific outcomes. It does not yet establish that the pilot improves longer-term performance.\n\nRecommendation\nContinue for one cohort, retain the three current success measures, add a 60-day manager check, and review the decision after that evidence is available.
Why it works: The example shows questions, results, limitations, conclusion, and recommendation without converting a short operational comparison into a causal claim.
Worked example 2Service evaluation
Illustrative library digital-help service.
Purpose\nThis evaluation asks whether the digital-help desk is reaching the intended users and resolving common access problems without staff completing private transactions on patrons’ behalf.\n\nMethods\nThe fictional evaluation reviews four weeks of appointment records, issue categories, completion status, and an anonymous exit question. Records contain 126 visits; 112 have complete issue coding.\n\nResults\nAccount setup and document upload represent 61% of coded visits. Eighty-three percent of complete records end with the patron able to continue independently; 11% are referred to another service, and 6% remain unresolved. The exit question is optional and was answered by 58 patrons, so satisfaction results are not representative of all users.\n\nConclusion\nThe desk appears useful for the two most common tasks, but current data are not sufficient to judge equitable reach or longer-term independence.\n\nRecommendation\nKeep the service, improve reason-for-visit coding, and add a privacy-safe measure of repeat assistance before making a larger staffing decision.
Why it works: The report separates operational evidence from gaps in representativeness and future-outcome evidence.
Prompt → finished structure
See the decisions between the assignment and the final form
These transformations make the hidden planning step visible so the template does not become a fill-in-the-blanks substitute for judgment.
Transformation 1Activity data → evaluation report
Starting material: Source contains attendance totals, satisfaction comments, completion counts, and anecdotes but no evaluation questions.
Decisions Define what the program was meant to achieve, turn that into explicit questions and criteria, map each available source to those questions, separate outputs from outcomes, identify missing evidence, and write conclusions only where the data answer the question.
Starting material: Draft highlights strong participation and favorable comments and calls the program effective.
Decisions Add the intended outcome, include contrary or incomplete evidence, distinguish satisfaction from outcome evidence, state sampling/follow-up limits, and revise the conclusion to the strongest claim the design supports.
Consequential claims require stronger methods and governance than an ordinary internal review.
Reusable frameworks
Start from the decisions the format requires
Framework 1
Observation → function
1. What can the viewpoint actually perceive?
2. Which 1–2 details matter now?
3. What do those details change in image, pace, relationship, or action?
4. What interpretation remains uncertain?
Training-program evaluation: compare completion, knowledge checks, participant feedback, and supervisor observations with the program objectives while noting that short follow-up cannot establish long-term behavior change.
2
Pilot evaluation: assess adoption, task completion, support burden, and user experience against predeclared success criteria, then distinguish a promising operational result from evidence of causal impact.
3
Community-program evaluation: report reach, participant characteristics, implementation fidelity, outcomes, limitations, and stakeholder perspectives without treating the loudest feedback as representative of everyone.
4
Website redesign evaluation: compare task success, error rate, time on task, and qualitative feedback before and after the redesign while noting any sampling or instrumentation changes.
5
Policy evaluation: state the intended outcome, implementation conditions, observed indicators, comparison basis, unintended effects, and evidence limits before recommending continuation or revision.
6
Course evaluation report: synthesize assessment results, completion patterns, learner feedback, and instructional changes, keeping satisfaction distinct from demonstrated learning.
7
Supplier evaluation: score delivery, quality, responsiveness, and risk using the same criteria across suppliers, then explain any weighting before recommending action.
8
Inconclusive evaluation: explain that the available follow-up period is too short for the primary outcome and recommend continued measurement rather than forcing a positive or negative verdict.
Turn an example into your own writing
Keep the underlying decision or pattern, then replace the subject, evidence, relationship, constraints, and tone with details that belong to your situation. If your final line still works after swapping only one noun, it may be too close to the example.