AI hackathon: how to test a solution before the final pitch
Prepare a held-out test set, compare an AI prototype with a simple baseline and give judges a record they can use to repeat the evaluation.
·6 min read
A team shows a successful model response. A judge changes the input text and gets a different result: a date disappears, an invented city appears or the request hangs. To investigate before the presentation, the organizer needs an evaluation protocol: which examples to test, what to compare the answers against and what to save after each run.
Prepare it before development starts. Teams will know the evaluation format, and judges will receive material they can use to repeat a run and inspect an error. Below, we build that protocol for a single illustrative AI prototype.
1. Define the result you will test
Our illustrative task is a service that extracts a city, an event format and a registration deadline from a short announcement. All announcements are fictional. The expected response is JSON with the fields city, format and deadline. Missing values must be null; the format can be online, offline or hybrid; dates use YYYY-MM-DD.
An example input is “Almaty, online. Register by 18 October 2026.” The expected values are Almaty, online and 2026-10-18. “Register tomorrow” without a publication date does not establish a definite deadline. The expected value in that case is null, rather than a guess.
Before the event, agree on accepted city name variants and rules for incomplete dates and ambiguous wording. Record them in the challenge brief. If two reviewers disagree about the correct answer, fix the task description and labels first.
2. Separate development from the held-out evaluation
To demonstrate the protocol, we will use 30 artificial announcements: 18 are available to teams and 12 are held out. These are chosen illustrative conditions, not a recommended sample size or a universal ratio. This set is too small to establish performance on real incoming data.
Teams use the available examples to refine the prompt, text processing and approach. If a team trains a model, its development data needs separate training and validation subsets. Keep the held-out test set for the final evaluation. Google explains this separation: choosing changes based on test answers turns the test into part of the tuning process.
Appoint an owner for the held-out set. Until solutions are frozen, that person keeps inputs and expected answers separately from the teams. Group repeated announcements about the same event so their variants cannot land on opposite sides of the split. For our example, include Russian and Kazakh texts, missing fields and ambiguous dates chosen in advance.
If a preprocessing step learns from data, fit it only on the training subset. The scikit-learn documentation specifically warns about leakage through preprocessing. This also applies to steps that appear to be simple data preparation.
3. Prepare a simple baseline
Build a baseline using a city dictionary, explicit event format indicators and recognition of fully specified dates. It also returns null when information is insufficient. Freeze the dictionary version and rules before running the held-out evaluation.
Test the baseline and the AI solution on identical inputs with identical comparison rules. This shows where the model performs better than a simple approach and where it introduces errors. Google's ML guidance recommends starting with a simple solution and defining what to measure in advance.
Check fields separately: city, format, deadline and handling of missing data. Save the count of correct answers alongside the number of fields checked. Do not combine them into a single “accuracy” figure if the number hides what you counted. Agree how these checks contribute to the overall assessment through the judging criteria.
4. Save a record of every run
Before revealing the held-out inputs, freeze the solution commit and evaluation settings. Ask each team for a short run record containing these fields:
Run identifier, commit and held-out dataset version.
Model: name and available API version, or local checkpoint identifier.
Prompt version, response schema and preprocessing rules.
Runtime and dependency versions, generation settings and seed where applicable.
Each example's identifier, raw response, status and run time.
An API version does not always freeze an external service's internal state. Record what the provider allows you to pin and which conditions the team cannot control. Keep keys, tokens and real user data out of the record.
Save responses before repairing their format. Otherwise, you cannot tell whether the model returned valid JSON or the application recovered it after an error. Record separately the processed result shown to the user.
For a nondeterministic solution, agree the number of repeats in advance and save every result. Do not pick a successful run after evaluation. A fixed seed alone does not guarantee identical output: PyTorch warns about differences across releases, platforms and CPU/GPU execution. For an external API, also account for its provider's limitations.
5. Inspect errors example by example
A judge needs a log with one record per input per run. Use this template: run_id, example_id, input text, expected fields, raw response, processed response, error type, run status and reviewer's note.
In our example, distinguish a missing date, an invented city, invalid JSON, a correct null and a missing response caused by a failure. Correctly declining to guess differs from an extraction error. An API access problem also needs a separate status.
Do not remove stalled requests or missing responses from the report. Show the number of all planned examples, completed runs and failures. If a label is wrong, record the correction and recalculate both solutions using the same rule. Once a team tunes its solution against held-out errors, that set no longer provides an independent test of the next version.
6. Let a judge repeat the evaluation
First, check whether the submitted package runs according to its instructions. For access, a fixed version and a fallback demo, use the submission acceptance checklist. Run unreviewed code in a disposable isolated environment without work secrets, using a separate test account.
Then a judge selects examples from the agreed held-out set and repeats the run with the recorded settings. Compare the raw responses with the saved ones. If they differ, record the difference, the conditions and the team's explanation. A video documents the result shown, but does not replace this run.
Before the pitch, give the judges the solution version, run record, baseline comparison and error log. State the size and composition of the evaluation, untested input types, external service dependencies and where a human must check the result. Judges can refer to specific responses and failures when the team explains what the prototype can already do and which evaluation it needs next.