Hackathon jury calibration: two practice projects before the first pitch
Agree on score anchors, assess two practice projects and record evidence, recusals and missing information. Three CSV templates for judges.
·7 min read
What you’ll learn
Score two practice projects independently before discussing differences through facts and published score anchors.
Separate the observed fact, interpretation and score; calibration does not change criteria, weights or tie-break rules.
Leave scores blank for missing evidence or recusal and record the status through the agreed procedure.
One judge gives a working prototype three points. A colleague gives it one: they watched the demonstration but could not repeat the scenario. If the panel discusses this only after the final, teams will have been assessed on different scales. Before the first pitch, review two practice projects and establish which observations each judge considers sufficient.
The meeting below provides a working procedure. Every project, record and score in the examples is fictional and created for training. They do not describe Stavleak customers, participants or event results.
Fix the rules and available materials
Give the panel the published rules, criteria, eligibility requirements and submission pack. Identify document versions and the point when scoring closes. Judges need to know whether they may run code, use a backup demonstration recording and ask clarifying questions. Check access to these materials beforehand.
The organizer sets weights, aggregation and the tie-break before competition work begins. Calibration before pitches clarifies how to apply the published rules. It does not authorize a new criterion, a changed weight or a retrospective requirement for user research. If a material procedure is missing, refer it to the responsible organizer and record the decision through the established process.
Assign someone separately to check submission completeness and eligibility. A missing required link calls for a decision under the rules rather than an improvised implementation penalty. Align the process with the judging criteria guide and hackathon rules.
Describe observable anchors on a 0-4 scale
The exercise uses three criteria: challenge fit C1, reproducibility of the main scenario C2, and testing evidence and limitations C3. This is an illustrative scale. The organizer prepares and publishes a matrix suited to the track in advance.
0: In the permitted test conditions, the main scenario demonstrably fails; the available fallback also fails to produce the claimed result.
1: One part works; the transition to the result is performed manually and shown.
2: The main path works on the supplied data; launch steps and limitations are recorded.
3: Another person repeats the path from a clean start; one named invalid input is handled.
4: An independent repeat and several negative scenarios include expected and observed results.
A zero here records an established test outcome. If a link failed to open and the permitted alternative material has not yet been reviewed, the score is unresolved. Level four does not establish production readiness either.
RUBRIC.csv describes all five levels for every criterion. C1 examines the connection to the challenge, C2 execution, and C3 the basis for conclusions. A successful test should not automatically earn maximum scores in all three columns.
Give judges two explicitly fictional practice projects
The practice challenge is to allocate limited mentor slots to teams without assigning two meetings to the same mentor at the same time. Every input request is synthetic. The exercise contains no real users or customer data.
Project A, “Slots,” accepts six requests. Following the instructions, a judge starts it again and produces a schedule without clashes. A duplicate booking is rejected with a message. The team supplies a log of that run and states limitations, but does not compare its approach with an alternative. The practice matrix uses these reference scores: C1 = 2, C2 = 3, C3 = 2.
Project B, “Dashboard,” presents a request form and an attractive schedule. The form accepts an entry; an operator then moves meetings manually. In the saved practice schedule, one mentor has two meetings at the same time. Usability testing is claimed, but observation records are absent. Reference scores are C1 = 0 because the demonstrated result violates the central challenge, and C2 = 1 for the working part of the scenario. C3 remains unscored pending clarification under the rules.
These scores apply only to the described materials and practice rubric anchors. A different track or permitted test method may produce a different outcome. Do not secretly use a recognizable current participant’s project as the reference.
Separate observation, interpretation and score
The score sheet contains three separate records. An observation states what happened and where it can be checked. An interpretation identifies the supported criterion anchor. The score is the value associated with that anchor. Add a material link or identifier, and a timestamp when referring to a demonstration recording.
For A, a judge could record: “A second request for an occupied slot was rejected in practice run A-NEG-1.” The interpretation is that the main path was reproduced and one invalid input handled. The C2 score is 3. Another judge can verify the basis for that decision.
For B, “clear interface” does not replace a usability test record. Note that the log was not supplied, select needs_clarification and write the question. Leave the score empty. Judges should not infer research from confident delivery or move an impression of visual design into a technical criterion.
Collect individual scores before discussion
The facilitator provides identical materials and asks each judge to complete the form independently. Save the first values before discussion. Judges do not see the chair’s score and are not required to choose the most popular number.
Then compare rows by criterion. If A receives 2 and 3 for C2, ask what each person reproduced and which level-three anchor was checked. One judge may have missed the negative test; another may have mistaken the team’s recording for an independent run. Restore access or explain the already published test method, preserving the original scores and the reason for clarification.
Review B through the same process. Agreement that C3 lacks sufficient information is a useful meeting outcome. It does not require a shared numeric score. In the calibration record, note the ambiguity, rule reference, responsible person and next action.
Proceed when judges can explain the boundaries and follow the same procedure. If time is limited, reduce the practice materials while preserving an independent review and discussion of one disputed case.
Handle missing information and conflicts of interest
Define statuses before the first project. scored means a score exists and is supported. needs_clarification records missing information. recused records a judge’s withdrawal. not_applicable is permitted only where the rules provide for it. An ordinary empty cell means pending, not zero.
A judge who advised a team, works with a member or has a personal interest reports that connection to the organizer through the agreed channel. Apply the published recusal procedure by assigning a replacement or using the specified independent-review process. Do not include that judge’s empty score in an average as zero.
Assign an owner and deadline to missing-data requests. Clarification must not become permission to improve the project after submission closes. Preserve the submitted version and decision explanation. The shared record needs an adequate description of the connection; personal details are unnecessary in participant feedback.
Check the result and use three templates
Before announcing results, check that required reviews are complete, recusals are resolved, the formula uses published weights, and the agreed tie-break has been applied. An unresolved required score leaves the result incomplete. Record any data-entry correction separately with its reason and responsible person.
Download RUBRIC.csv, score-sheet.csv and calibration-record.csv. The files contain labelled A/B practice rows. Copy the structure and remove those examples before real judging. Weights and the published-rule reference are left for the organizer to complete. The CSV files do not calculate a final score automatically.
Check your matrix against the judging criteria toolkit. The toolkit contains a separate 100-point template; publish one coherent scoring system for your event in advance. Store the approved rubric, individual reviews and calibration record alongside the final decision under the event rules.