# I Built a Research OS for My Chat-First Workflow. Then I Lost My Chats.

> My research lived in a Google Sheet, my memory, and dozens of chats. I rebuilt that setup as a small Research OS with AI agents over a weekend, and days later it was the reason I could keep working after losing access to ChatGPT.

**Published:** 2026-09-29
**Canonical URL:** https://yihuisong.com/article/research-os-v1

Since I started my builder site in March, I have published nearly twenty long-form essays. Behind the pieces, I have been running research, interviewing industry experts, and working on several independent projects in parallel. Friends who follow the site sometimes ask how I keep all of it moving at once.

For months, the answer was a Google Sheet and my memory. By this summer, neither could keep up. Over one weekend I rebuilt the workflow with AI agents, and the new system misread my work several times before I could trust what it told me. Days after I finished it, it became the reason I could keep working when I lost access to ChatGPT.

## How the Work Actually Moved, and Where It Broke

When I only had a few ideas, the Google Sheet was the right tool. One row per topic, with columns for Topic, Category, Contact, Progress, and Post Time, was easy to scan and easy to change. Almost every idea started in a chat. I would open ChatGPT with a question about a product, a startup everyone was talking about, or something interesting I had read, and keep going until the shape of a topic became clear. A topic would grow into a research plan and eventually an essay draft in Google Docs, experiments went into their own repos, and the Sheet sat in the middle as the place where it all came back together.

![My old Google Sheet tracker, with columns for Topic, Category, Contact, Progress, and Post Time](/images/research-os-v1/sheet-tracker.jpg)

As the number of ideas grew, two things about them stopped fitting in rows.

The first was progress. Every idea followed the same path, from a first spark to accumulated context, finding people to interview, then research solid enough to write from. Each topic moved at its own speed and sometimes sat at one step for a long time. A Progress column reading "researching" could not say which step an idea had reached or what it was waiting on. The work itself happened in chats and repos, and the Sheet only knew what I remembered to manually carry back.

The second was connection. Some ideas turned out to be related to each other. Research for one would sharpen another, and a few started to look like a series that could add up to more systematic thinking than any single essay. The Sheet had no way to show that. Every row stood alone, and the relationships lived in my head or in whichever chat they had surfaced in.

## Before Research OS, There Was a Repo Called "chatgpt"

My first fixes came out of the connection problem. When several ideas started pointing at one larger research direction, big enough for a series, I asked ChatGPT to create a Markdown document to track it, and had later chats update that document whenever they found something relevant.

The product did not always cooperate. A new session often could not edit an artifact an earlier session had created, so it would write a separate Markdown note, labeled as an addition to the original. The tracker became a stack of appends, and stitching them together was once again my job.

Then I found out ChatGPT could read and write a GitHub repo. So I created one called "chatgpt" as a reservoir where GPT could keep the important artifacts it created, such as the research plan for a series, an ongoing experiment tracker, and a new pipeline document listing every idea we had discussed recently. From then on, every chat could read and update the same files.

That kept my artifacts in sync with what happened across chats. But it did not answer the questions I asked most often: what is P0 right now, what is active, what is waiting on something, and what comes next. Ideas also often arrived before they had a home. A small annoyance with one product's user experience could turn into a design pattern I started noticing across a whole category, and until I knew what it was, it stayed in the chat where it started.

What I wanted was an incubation pool for future essays, showing for each idea how far along it was, what was blocking it, and which other ideas it fed. Later chats had to be able to read that state, the reasoning had to stay in the documents where it already lived, and a new idea needed a way in before I knew where it belonged.

## The First Version Starts Making Decisions for Me

I shared my needs with ChatGPT, and it turned them into a plan for a small "Research OS." On a Saturday night I handed the first piece, one structured record per topic, to a coding agent, with a warning not to invent anything the source did not contain.

That warning lasted about nine minutes.

The agent's first report contained this line: "Priority was derived, not invented." It mapped the groups in my pipeline document straight onto priorities. The three topics I had ranked under "Now" all became P0, erasing the order. The nine topics under "Next," which I had deliberately left unranked, all became P1. The agent made planning decisions before I had.

I had every priority wiped, kept the old groups separately as history, and typed the rule: "Priority is a human planning decision, not a derived synonym for pipeline bucket." Then I opened the first dashboard and spent my first quarter-hour setting priorities by hand on all the topics listed. In the end, the Now topics kept their original order at P0, P1, and P3, and the Next candidates spread across three different levels. The agent had read my document accurately and still misread what I meant by it.

Using the dashboard also surfaced three things the records could not tell me:

* **What a topic was.** Some rows were too thin to remind me what the topic even was, so I asked that every topic page answer one question first: "What is this thing, and why did I keep it?"
* **Whether it was done.** Some of the topics were stale because I had kept working or published essays in places the repo could not see, so I needed a way to mark them "published."
* **Whether I still wanted it.** Some were ones I no longer wanted to write about, so they needed a terminal state called "discarded." The rule that came out of it held for the rest of the build: my old documents provide context, and the pipeline owns current status.

## The Dashboard Works, and I Still Cannot Read My Own Research

The first dashboard was a flat table, one row per topic. With summaries in place, I could recognize each topic on its own. The trouble was seeing them together. Every row looked as important and as unrelated as the next.

The dependency column made this obvious. A cell would say `3 links`. The number was accurate, but it told me nothing: which three, in which direction, and why? Was this topic blocked by those three, or feeding them? The count was also hiding a data problem. The agent had stored some relationships twice, once on each topic, and the two copies did not always agree, so A could claim to depend on B while B claimed to depend on A. In my next spec, I asked for every relationship to be spelled out as a complete phrase, so I would not have to guess what a count meant.

The shape was missing too. Some topics belonged to research lines that already spanned several essays, some stood alone, and some existed mainly as evidence for others. My pipeline document already grouped them this way. The records had dropped that grouping on the way in.

Therefore, I asked for a new default view: a Research Map that grouped topics by research line and spelled out relationships like "evidence source for." Underneath, each relationship was now stored once, in one direction, with the reverse worked out for display. The old table stayed as the place to edit.

I also wanted each topic page to show the full background. Copying that text into every record would have meant keeping two copies in sync by hand. Instead, each record just points to its section of the pipeline document, and the topic page pulls that section in when I open it. This way, the reasoning lives in one place, and I see it wherever I need it.

## Dogfooding Makes the Design Sharper

Once the Research Map existed, I kept using it through the night. Three problems came up, and fixing each one made every topic easier to read.

Two of them came from the same place. Every topic had a stage for how far the work had gotten, and it still carried its old planning group from my pipeline document. Those groups tracked where my attention was: `Now` for what I was focused on, `Next` for what was coming up, `Evidence-dependent` for topics blocked until some evidence came in, and `Parked` for things I had set aside.

First, the word "parked" appeared both as a stage and as a planning group. As a stage, pausing a topic erased how far its work had gotten, so picking it back up meant guessing where I had left off. Pausing is a decision about my attention, so it belonged with the planning groups. A quick rename to "deferred" left the duplicate in place, so I removed "parked" from the stages, and ten topics went back to the stage their work had actually reached.

Second, I found two rankings for the same topics. "Now" and "Next" implied an order, and once I had set every priority by hand, that order overlapped with priority and sometimes disagreed with it. A topic named "Agent Trust Part 2" still sat in "Now" at P3, and one of the "Next" topics was rated P1. I retired both labels and let priority carry the order, since that was the ranking I had actually made. The rest of the groups became a new field called activity: whether a topic is active, waiting on evidence, or deferred.

Third, a group label was simply wrong. The map had a group called "Older research," which totally confused me. It came from a section of my pipeline document, and its topics were not the oldest ones at all, so the group was dissolved. Topics that had been folded into bigger ones got their own group. The only thing the others had in common was a note on what would make each of them worth picking up again. That note now lives in each topic's own summary, so I see it whenever I open the topic.

By Sunday evening, one more field had gone the same way, and the agent named the pattern: "each for the same reason: it looked like structure but wasn't doing any." What survived were three questions I can answer for any topic: how far along it is, how much it matters right now, and whether I am actively working on it.

## Chat Can Read the State, and a New Idea Still Gets Lost

By then, the three places my work lived had clear jobs:

* Chat is where I think, often before I know which project an idea belongs to.
* Research OS holds the context and current state worth carrying between conversations.
* Each project's own repo holds its code and experiments.

Now any new chat can read one small repo and catch up on where everything stands.

Reading already worked. I had recently been working on a Notion deep dive as the first piece of my Enterprise AI research. Once the dashboard's changes were pushed to GitHub, a later chat read the Notion topic back as P0, drafting, active, under Enterprise AI, with its target date and next action. The priority and date were ones I had set by hand.

Writing was the real challenge. A chat about Plaid API constraints turned into a product experiment on household spending awareness. ChatGPT saved it to the repo as an evidence document and a new section in my pipeline document. As I hoped, the raw idea left the conversation and landed somewhere durable.

But that was only half of what I needed. When the coding agent later checked, the topic had no record. The dashboard and the Research Map could not see it, and nothing would have warned me. The new section also used a status word the dashboard had dropped earlier, because my pipeline document still listed it as valid. The chat had followed the document in front of it.

I passed the miss back to ChatGPT to find out where and why it had happened. Capturing an idea and registering it turned out to be two different steps. Writing it into Markdown preserved it, which the repo had always done well. Making it show up with a stage, a priority, and a place on the map was the second step, and nothing in the system required it.

## I Repair One Complete Loop and Call It V1

First, the spending experiment was registered as a proper topic in the pipeline, as evidence at a real stage. Then three changes made sure the next chat would not repeat the mistake. A new rule, written into the pipeline document and the pipeline's README, says anything a chat captures has to be classified and registered as a topic before the capture counts as done. The outdated status word came out of the pipeline document, so no chat would copy it again. And a health check now warns whenever a section in the pipeline document has no matching record. Priority stayed the one thing the rule never sets, since I set it myself.

Running the fix exposed two more gaps. The first was that write access and the ability to execute are different capabilities. ChatGPT could edit the repo but could not run the dashboard. The fix broke every page, and the new health check flagged the very topic it had just registered. The coding agent fixed both and confirmed the fix by running the dashboard in a sandbox. From then on, a change only counts as done after it has been checked somewhere the dashboard can actually run.

The other gap ran in the opposite direction. The dashboard could change my data, but only on my laptop. When I moved the spending experiment to P3 and deferred it, that change reached GitHub only because it got swept into the agent's fix commit. Commit and push became the explicit last step. I kept it manual on purpose, so I review every change before it goes out.

The full path now runs from a chat, through registration and a real check, to a push that any later chat can read. One idea has made it all the way through, but it entered through the broken version and was repaired in place. No new idea has taken the repaired path from start to finish yet, and I will not claim it works until one does.

![The Research Map, the default view: topics grouped by research line, each with its stage, priority, target date, next action, and relationships spelled out](/images/research-os-v1/research-map.jpg)

![The Pipeline view, the editable table where I set priority, stage, and activity by hand](/images/research-os-v1/pipeline-table.jpg)

## What Survived When My ChatGPT Workspace Went Dark

Research OS will keep changing, and it has already been tested harder than I planned. Right after I finished v1, the ChatGPT Business workspace I belonged to as a member was deactivated. The owner had received a policy warning about a request he made in Codex, and when his account went down, every member lost access at once. My project notes, brainstorm sessions, and months of career development context were in there. I never got a warning myself, and there was no window to back anything up. It was a hard hit.

![What ChatGPT showed me once the workspace was deactivated: a 402 error with the code "deactivated_workspace"](/images/research-os-v1/workspace-deactivated.jpg "narrow")

In my old setup, this would have been much worse. The Sheet only held rows, and much of the reasoning behind them lived in the chats that had just disappeared. Luckily, the current state of every project was in the repo. So were the important outputs of the chats that had done the heavy lifting, including series plans, experiment trackers, handoff docs, and every priority I had set by hand. The conversations are gone for now. What they produced is still mine, and I can pick up most projects where I left them.

No model company or agent product should be able to take my thinking away from me, without warning, without a way to export it, and because of something someone else did. This week one did, and the only reason I could keep working is that the most important part had already left the chat. Backing up the chats themselves is now the first thing I am adding to Research OS. Wherever I do my thinking next, I will keep a copy somewhere I control.
