The evidence

GAiDFLY is built on a specific published finding, and on a specific problem in strategy education. Both are worth stating plainly.

The finding

In 2026, researchers at Columbia Business School released a working paper analyzing nearly a thousand student debates with an AI adversary, comparing students who debated by voice with students who debated by text (Jain & Wang, 2026). Debating aloud was associated with significantly greater divergent ideation — students who argued out loud explored a measurably wider range of ideas than students who typed. The study is correlational, and we present it as such. But the direction of the finding matches what every seminar instructor already knows: something different happens when you have to say it.

The bottleneck

The most reliable way to verify that a student owns an argument — rather than having assembled one — is oral defense: sustained, adversarial questioning by someone who has read the work. It is also the least scalable thing a faculty does. One examiner, one student, one hour. GAiDFLY exists to scale the rehearsal for that moment, not the moment itself. The machine takes the thousand hours of sparring; the human keeps the verdict.

The method

A GAiDFLY session follows a very old shape. The student commits to a position and argues it first — the frame stays theirs. GAiDFLY then argues the other side, one short challenge at a time. The practice of arguing both sides of a hard question is older than the university: an anonymous Greek text called the Dissoi Logoi — ‘double arguments’ — laid out paired cases for and against, and Roman rhetorical schooling made arguing in utramque partem, on either side, a standard exercise.

GAiDFLY’s probing moves come from the Socratic elenchus, the cross-examination that tests whether the speaker’s own commitments hold together. Even the short turns are Socrates’ rule: in Plato’s Protagoras he refuses to continue unless the long speeches stop, because only short exchanges can be examined.

The session ends where the old dialogues did not: the student says where they land and restates their case as they would argue it tomorrow. GAiDFLY renders no verdict — in Plato’s terms it practices dialectic, never eristic: argument in service of testing, not winning. The resolution belongs to the student; the judgment belongs to the humans. Afterward it writes a debrief for student and instructor: ideas explored, pressure points and how they were handled, and what changed between the opening argument and the closing restatement.

What it doesn’t do

The debrief observes; it never scores. GAiDFLY assigns no grades, ranks no students, and renders no verdict on whether an argument was good — that judgment belongs to the student and their instructor. We think the absence of machine grading is not a missing feature but the point: the tool stays an instrument, and the standard of proof stays human.