Human-AI Collaboration in Research
Designing your own way of doing research with AI.
Sébastien Martin Associate Professor of Operations · Northwestern KelloggAI-Accelerated Research in OM/OR · July 24, 2026
Start normally. One line of expectation-setting: this talk is rapid-fire, lots of examples; I'd rather show too much and go deeper in the discussion. ~1 min.
Introduction
Sébastien Martin
Research
Large-scale optimization
Platform operations
AI in operations
AI in education
Teaching
AI Foundations for Managers (new)
Operations Management
Very fast — many in the room know me. Just: who I am and the topics I care about. Point at the QR for the curious (AIML 901 syllabus). ~30-45 sec.
Introduction
Today
1
AI can now do serious research
Fast signposting — rapid-fire talk, deep discussion after. (Roadmap will grow as the plan settles — later sections not yet reworked.) ~30 sec.
AI can now do serious research
Even basic ChatGPT can make important research contributions.
One super important thing before anything else — this is capability you can access today.
AI can now do serious research
Let’s start a research task together.
Live
Window switch: launch the computer agent on the rMC task, live, then walk away — "we'll come back to this at the end." Say verbally: a real task from one of my active papers; it works unattended while we talk. Keep it brisk; the payoff is the comeback slide ~20 min from now. ~2 min including the launch.
AI can now do serious research
This week
Mon Jul 20
The Jacobian conjecture falls
Levent Alpöge
Tue Jul 21
Terence Tao discusses it
Wed Jul 22
The DGG conjecture falls
Dmitry Rybin
Fri Jul 24
🎂
Sébastien’s birthday
Two long-open conjectures fell THIS WEEK, with AI in the loop — Jacobian with Claude (Alpöge; credit Akhil Mathew for the question, open since 1939), DGG with GPT-5.6 Pro (Rybin, PhD student at CUHK-Shenzhen; prompt ~58 words: "here is the conjecture, keep going"). Tue: Tao's viral session digesting the Jacobian counterexample — even the best use it as a thinking partner; he discusses AND verifies calculations. Community-verified within hours, generalized to an infinite family, not yet peer-reviewed — the verification theme. (Friday: yes, really my birthday — let the joke land itself.) I narrate everything; the conjecture is explained on the NEXT slide, then I open the conversations myself from the browser (links list in the plan note): Rybin chatgpt.com/share/6a60b2eb-0b64-83ee-9c76-7931ca1de063 (point at "worked for 52m/89m/93m/88m" — persistence, not magic) and Tao chatgpt.com/share/6a5fdc7a-d6f8-83e8-bbea-8deb42cfed56. Tao = deep collaboration; Rybin/Alpöge = near-autonomous. ~4-5 min total incl. next slide. Keep both tabs pre-loaded.
AI can now do serious research
The Dinitz–Garg–Goemans conjecture
A fractional flow
min-cost — demands may split across paths
s
t
demand 10
6
4
An unsplittable flow
each demand travels on one single path
s
t
demand 10
10
The conjecture (Goemans, open ~30 years): any fractional flow can be made unsplittable at no extra cost , adding at most dmax — the largest single demand — on any arc.
Quick verbal setup, pitched at rusty-but-fluent OR people: one source, terminals with demands, arcs have capacities AND costs. A fractional flow is a regular min-cost flow — demands may split. Unsplittable adds one combinatorial constraint: each demand on a single path — each customer's shipment can't be split across trucks. DGG 1999 (celebrated): you can always round, exceeding fractional flow by less than d_max per arc. Goemans conjectured rounding is also FREE — no cost premium. This week's counterexample: 7 nodes, three demands (15/10/15), fractional cost 58, every unsplittable with violation ≤ 15 costs ≥ 60. Finite object, checkable by direct computation — which is exactly the verification theme. It's OUR math — LP relaxations, rounding, flows — and it fell to a short "keep going" prompt. Then: switch to the browser and show the two conversations. ~2 min.
AI can now do serious research
Say verbally: the total human contribution there was asking. Should we stop trying to prove things? Is this just bad news? I don't believe so — at least not right now. But we each need to figure out where we fit — and research is competitive: people who master these tools will be far more competitive. Pivot: let me tell you two stories from my own research. ~1 min.
AI can now do serious research
Human-AI Interactions and Societal Pitfalls
Francisco Castro
UCLA Anderson
Jian Gao
UCLA Anderson Ph.D.
The spiral of homogenization
People create with AI
Output homogenizes
AI trains on that output
Where does this converge?
One example from my own research. The paper (major revision at M&SOM) with Francisco Castro (UCLA Anderson) and Jian Gao (UCLA Anderson PhD, now at Amazon). Broad idea: we study homogenization — how people's output, and eventually preferences, collapse toward the AI's style when everyone creates with AI; and the AI then trains on that output. Where does this converge? The key mathematical result — the variance-collapse theorem behind this spiral — is the one we could not crack ourselves. ~2 min.
AI can now do serious research
How we proved it
A fresh session starts
ChatGPT 5.5 Pro reads the notes
The research notes
the problem, leads, lemmas, partial results
The AI works
makes progress, cleans up the notes
We read and add intuition
staying at a high level
restart fresh
The proof
eventually
In my experience, the ChatGPT Pro models are the best models for theory. Unfortunately, they require an expensive subscription.
Tell the attempts arc verbally: just asking gave nothing; the elaborate multi-agent machine I engineered barely moved for months; then a smarter model came out and something much SIMPLER worked. The loop: the problem is given once, in full, into the notes. The AI works and makes progress. We read, discuss, and add intuition, staying at a high level. Then we RESTART from a fresh session: a new AI reads the notes of the previous one, with two jobs, make progress and clean up the notes. The notes are its scratchpad: leads, background, lemmas, partial results. Callout beat: the ChatGPT Pro models are the best for theory in my experience, but the subscription is expensive. [Window switch: show the actual conversation live from my machine.] ~3 min + demo.
AI can now do serious research
This was real, creative research
After the proof converged, Francisco read the whole document cold. The title carries the claim; the screenshot is the evidence, and it names Francisco itself. Let the quote land on its own: "astonished", "everything checks out", "several arguments highly non-trivial", "could have taken a long time without AI." The proof was truly beautiful, and honestly creative. ~45 sec.
AI can now do serious research
So what was our role?
We asked the right question.
We designed the process and advised the AI along the way.
We verified everything, and we own the result.
This WAS a form of core research, with the human as a kind of advisor: asking the right question, designing the (simple) process, maintaining intuition while the AI went very deep, and doing the verification. We own the result. (The creativity beat already landed on the previous slide.) ~1.5 min.
Computer agents
Why figuring out how to do research with AI is an operations management problem.
Transition: that was ChatGPT in a browser. The next capability level is agents that work on your computer.
Computer agents
A quick recap of AI progress
ChatGPT
AI can write and answer questions
GPT-4 & the AI wave
AI can be really smart
Agentic AI
AI can “reason” and use tools
Claude Code & coding agents
AI can work on its own on complex tasks
Claude Cowork & computer agents
AI can control a computer and do real work
Four years in one slide. Each step is a real jump in what AI can do: write, then be genuinely smart, then reason and use tools, then work on its own on complex tasks, and now control a whole computer. The line keeps going up and to the right, fast. If it keeps going, the end state is not a smarter chatbot — it is something that can take on open-ended work the way a capable collaborator would. ~1.5 min.
Computer agents
Computer agents
& Cowork
Anthropic
Codex
& Work
OpenAI
This is the same LLM, except that it can control your computer.
The capability layer for everything in this section: an agent that runs on your own machine, sees your files, and actually does the work. Two main options, interchangeable for our purposes: Claude Code and Cowork (Anthropic), Codex and Work (OpenAI). The point is to get into one: same LLM as the chat box, except it can control your computer. ~1.5 min.
Computer agents
Give it your research folder
paper.tex
simulations/
data/
research-notes/
ai-instructions.md
A computer agent
Claude Code · Codex
reads, edits, runs
Like co-working with a PhD student in your Dropbox folder.
The easiest way to do research with a computer agent: keep working the way you already work. Point it at the folder of one paper: the LaTeX, the simulations, the data, your research notes, organized like your Dropbox already is. You can also drop instruction files in the folder; the agent reads them every time it starts, so it works the way YOU want on this project. No new structure, no new tool. [Optionally show a real example folder.] ~1.5 min.
Computer agents
What does this unlock?
Compile and fix your LaTeX
Hunt down missing citations
Run simulations while you sleep
Write and edit the paper with you
Write math directly in LaTeX
Keep the folder organized
Write notes to its future self
Follow your instruction files
Answer any question about the paper
The only skill it takes is not an AI skill: keeping a folder organized . And the AI can help with that too.
Rapid-fire, the point is abundance: with nothing more than a computer agent and a shared folder, all of this comes essentially for free. It can fix compiling issues and missing citations (it handles the Overleaf pain for you), tweak parameters and iterate overnight, help write the paper directly if that makes sense to you, do the math straight into your file, keep things tidy, write its own research notes for the future, and answer any question about the paper or the proof while you work. The whole purpose of these two slides: something as simple as this unlocks a massive amount of opportunity. ~2 min.
Computer agents
How do you work with a PhD student?
out of the loop
in the loop
The grant provider
you fund it, they do it
A small army
feedback, spread thin
A few students
a true co-researcher
An equal collaborator
when 1 + 1 = 3
You drive
the AI just speeds you up
Apologies for the analogy: AI does not replace PhD students.
Walk the arrow left to right; everyone in the room has lived these modes with students. The grant provider: full automation, the paper machine. You are basically funding research that others do, and you learn little. A small army: you see more and give some feedback, but your attention is spread thin and you become the bottleneck. A few students: you will not master every proof detail or every sentence, but you are a true co-researcher, very involved. An equal collaborator: a real back and forth. Like the best coauthorships, they take what you do not like, you keep what you love, and one plus one equals three. You drive: the AI is just a tool that accelerates what took time before (LaTeX errors, beautiful visualizations to think with) without changing the research. The difference with AI: you get to design the collaborator, and where you sit on this arrow is your choice. ~2 min.
Computer agents
AI creates a new job: optimizing the research process
Creating extremely personal and powerful research tools
AI peer review
Voice-first writing
Reinventing how we work with PhD students and collaborators
A literature second brain
An automated paper machine
Increasing the scope
An automated research-lead machine
A digital research twin
Large research group coordination
The question is not whether to use AI in research, but how you want to use it . The sky is the limit, and there is room for innovation.
The reframe after the folder slides: those showed how easy it is to USE AI. This slide is the new job it creates: designing and optimizing your research process. Not enumeration, invention: reinventing the discovery process (interactive visualizations, seeing your problem in 3D), AI peer review, increasing the scope (numerical experiments that never made sense by hand), reinventing how we work with real PhD students and collaborators, a literature second brain, an automated paper machine (Rad's session), an automated research-lead machine, large research group coordination. The two highlighted ones, voice-first research and the digital research twin, are what I show next, briefly. ~2 min.
Computer agents
I talk, it writes Live
The draft
as if a PhD student sent it
I read and react out loud
recorded; hand-wavy or “move this figure to page 2”
The AI does the work
it knows how I like to write
~15 minutes later
I no longer do the writing directly, but I iterate far more on the writing than I used to. And it is how this slide deck was made.
The workflow, on an actual paper: a draft exists (from the AI, or a coauthor, like a PhD student sending you a chapter). I read it and record myself reacting out loud, every little sentence, sometimes high level and hand-wavy, sometimes super precise ("move this figure to page 2"). The AI has the recording, does the work, and ~15 minutes later the next iteration is back. Two points: it has a lot of saved information about how I like to write, and when it gets something wrong I improve that; and I iterate far MORE on the writing than I used to. Live: show a real recording-to-revision loop from a recent paper session. Nice beat: this deck was made exactly this way, this morning. The assistant on screen is quietly powered by something I'll reveal later — do NOT explain the knowledge graph yet. ~4 min with demo.
Computer agents
My approach: trying to build a second brain Live
email
meetings & recordings
AI work sessions
calendar, messages…
automated daily update
The knowledge graph
essentially a Wikipedia,
and the memory of my AI
My AI
edits its own notes
guides everything it does
Running quietly behind everything you saw today. Not sure I recommend it (it takes real care), but it works really well for me.
Framing: my own workflow, almost a research project of its own. The simplest way to explain it: essentially a Wikipedia of my work, maintained BY the AI. Every day it updates it from new information (email, meetings, recordings, work sessions), and it edits its own notes; this is the memory of my AI. In return, the graph guides everything it does. So I am no longer confined to one folder: more and more, I just talk to the AI and that is it. It powered everything you saw today. [Live: demo-mode assistant, Sébastien's track.] Happy to answer questions later. ~3 min.
Computer agents
My own path with AI
Finding the perfect research process is real work. But it is work that keeps on giving.
Close the section on the honest curve. Because what we are optimizing is a PROCESS, it does not work right away: expect the bumps. Finding the perfect research process is real work, but it is work that keeps on giving. ~1 min.
Reflections
Where is all of this going?
Final section: advice for both ends of the room, the open questions, the agent comeback, and the close.
Reflections
How to get started
1 · Get a computer agent
Claude Code or Codex — not the chat box. Ten minutes to install.
2 · Give it something real
A task from your actual research: a figure, a robustness check, a literature pass.
3 · Trust the process
The first sessions will be mediocre. Invest in the setup anyway — it compounds.
Already familiar? Keep building. The real work is finding the research process that works best for you: it takes time, and it needs to be updated as AI changes.
Both ends of the room on one slide. For the basic-ChatGPT users, no judgment: here is what I would do this week. For the experienced: keep building, there is no limit; the real work is finding the process that works best for you, it takes time, and it needs updating as AI changes. (The separate "no limit" oneliner was deleted; this box carries it.) ~2 min.
Reflections
Open questions
AI can already do research autonomously.
Which field is likely to change the most?
Currently the hardest task is verification. What does this mean for peer review?
What does authorship mean when you can run a project without deeply understanding it?
How do we handle all this and still enjoy our jobs?
To me, all of this shows that the future of research is uncertain , not that it will be bad. The more I integrate AI in my research, the more empowered I feel. I am optimistic.
These are for the discussion, not for me to answer. Stance in the gray box: uncertainty, not doom; there are many things to discover, just slightly different from before; the more I integrate AI, the more empowered I feel. Say verbally if useful: how do we know the humans verified it; if the reviewer does the verification, shouldn't they be an author? 10,000 specifications until one "works": p-hacking with no intent, and theory just as exposed. ~2 min.
Reflections
Let’s check in on our agent.
Window switch: open what the minute-zero agent produced. Context to say verbally (from the deleted "What has the agent been building?" slide): Relative Monte Carlo — an active paper with Audrey Bazerghi and Garrett van Ryzin; the prompt asked for an animated ride-hailing simulator (our trained rMC dispatch policy vs. the closest-car heuristic on the same request stream); built to develop intuition with a coauthor, not for show. Fallback if the live run disappointed: ~/gdrive/research/2023 - relative monte-carlo/202606 - ridesharing simulator/viz/index.html. ~2.5 min.
Reflections
Thank you!
sebastien.martin@kellogg.northwestern.edu
And please help us spread the word about this seminar series!
Say the close myself: this is just my opinion, but one thing I'm convinced of is that AI will have a big impact on research, and that it can make our job really fun. We have so many things to figure out together; that is why this seminar and these conversations exist, and why it matters to our field that we set them up. Point at the box: help us advertise the series. Thank you, hand off to Gad. ~1 min.