JAW 2026 · Hackathon
The Last Frontier of AI
Quantifying Trust & Reliability in Generative AI Output
“Let's conquer the last frontier of AI, let's build for AI reliability.”
In collaboration with
GikaGraph.ai
Now live
Validation dataset released
The validation set is now available on GitHub. Final submission closes at 12 PM on 13th August, so grab it and refine your solution.
The Challenge
The world is mesmerised by the creativity of AI. However, creativity and reliability are two opposing objectives.
Existing AI models fail to capture multi-hop reasoning, numerical dependencies, and strict logical constraints over complex data. How do we prove and trust AI outputs that require long contextual memory, stretching far beyond the model context window, numerical consistency, and multi-hop reasoning over heterogeneous data? Your goal is to develop a rigorous framework that ensures AI outputs are reliable, factual, and trustworthy for complex real-world workloads.
The Objective
As Generative AI transitions into complex workflows, the industry faces a critical gap: how do we trust AI-generated responses?
The AI community still lacks a rigorous framework to generate reliable model responses and validate their quality. Models continuously struggle with numerical alignment and long-horizon reasoning over complex datasets.
Hackathon Task
- Input A raw, complex multi-modal dataset (combining structured and unstructured data) and a real-world example set of questions.
- Output A framework that processes this multi-modal dataset to ensure AI outputs are accurate for workloads requiring long contextual memory, numerical consistency, and deep reasoning capabilities, materially improving model reliability and eliminating traditional hallucination vectors.
Evaluation Metric
For each question in the validation set, there will be one exact factual answer. Your score depends on how close your generated response is to the factual target. A live leaderboard shows your results immediately after each submission.
Important note
Answers will not be directly present in the dataset. There will be no simple text chunk containing a direct answer, your system must reason, aggregate, and compute across the data.
- 100% score is awarded for an exact match to the ground truth.
- Scoring basis: each question in the validation set is evaluated numerically with respect to the correct answer.
Timeline
- 5th August The dataset and problem statement are released. Participants begin building their solutions.
- 10th August, 3 PM The validation dataset is released. Participants get 69 hours to refine and improve their solutions.
- 13th August, 12 PM IST Final submission deadline.
- 13th August, by 3 PM Winners are announced on the JAW 2026 website.
- 15th August, afternoon Winners present their solutions in a 20-minute session, remotely or in person, with slides. Prizes are awarded at JAW 2026.
How to Register
- Register using only your BITS email ID.
- The team leader registers the team first and receives a registration code by email.
- Other team members join the team using that registration code.
- A team can have a maximum of four members.
Register and submit on the evaluation platform.
Key details
- Prizes: ₹25,000 total prize pool, ₹15,000 for 1st place and ₹10,000 for 2nd place.
- Eligibility: Open to all students.
- Team size: Up to 4 students per team.
- How to participate: Get the problem statement and dataset on GitHub, then register and submit your solution.
Questions?
For questions about the hackathon, write to dhruv.kumar@pilani.bits-pilani.ac.in or contact@gikagraph.ai.