Marble

The Integrated R&D Environment

Much of what looks like an R&D speed problem is a verification problem. This is a vision for solving it: a working environment where scientists and AI agents find prior work, run models the organization trusts, and trace every answer to its evidence.

Picture a Tuesday morning in a consumer-products R&D lab.

A product researcher opens six tabs before her coffee is cold. Yesterday's stability readouts live in a LIMS export. The claims tracker is an Excel file with a version number in its name. The model she actually needs runs on one workstation in modeling and simulation, and the two people who know how to run it have a queue. Somewhere on a shared drive sits last year's consumer study, four hundred pages of gold that nobody has five days to read. By noon she has touched six tools, asked three people for files, and hasn't yet done any science.

Talk to R&D teams across the industry and you hear the same sentence in different accents: we have the data, and we have the models. We just can't get to what we've learned.

The diagnosis

R&D teams are not usually waiting for ideas. They are waiting for enough evidence to act.

The job, stripped down, is to deliver product and packaging with superiority that is demonstrable and feasible, under real constraints of cost, capability, and consumer preference. Every step toward a launch is a question: will this formula stay stable for two years, does this claim survive testing, will consumers prefer it, can the plant run it. Every answer a team accepts is a decision about how much evidence is enough. Verification doesn't create truth; it creates evidence sufficient to make the next decision with confidence.

Hence the sentence this vision hangs on:

R&D moves as fast as it can verify answers.

Not as fast as it can generate ideas, staff projects, or schedule reviews.

That engine runs below its potential for a reason everyone in R&D already knows: silos. Studies get repeated because the first one can't be found or trusted. Models get built once, used once, and retired to the laptop of the person who wrote them. Insights stay buried in old experiments and outside activity.

Four small islands, one for each R&D function: a lab bench, a shopping cart, a factory, a clipboard. Each carries a filing cabinet and a padlocked chest. On two of them, two scientists are separately building the same small machine, neither aware of the other. A scroll in a bottle drifts unnoticed in the water between them.
Trapped knowledge · Every function keeps its own island, and the work that would cross between them stays locked on the shore.

The economics of evidence

Evidence can be bought at very different prices. No test at all is free and worth what it costs. External references are cheap and secondhand. An unvalidated model is cheap and unproven. A real experiment buys solid evidence and charges weeks for it. A market-scale test buys the most confidence and charges the most of everything.

The outlier on the curve is the validated model. Expensive to earn once, nearly free every time after: the only purchase on the spectrum that gets cheaper the more you use it.

cost + timeno testreferenceunvalidated modelexperimentbest dealvalidated modelmarket testcertainty
The evidence-for-cost spectrum · Confidence rises with what you spend, right up to the validated model, which sits off the curve entirely.

An organization that can build, trust, and reuse validated models answers most questions in minutes and saves physical experiments for the questions that deserve them. Run the necessary experiments. Never repeat a study because the organization couldn't find or trust the first one. Model every answer you can.

None of that is a new insight. The silos are why it doesn't happen.

Why now

AI changes the economics of finally fixing this. Agents can now do real work, but only as well as they can verify it. An agent with access to validated models and structured past work is a colleague. An agent without them is a confident guesser.

Two panels. On the left, a robot holds up a speech bubble reading 749 while standing on a loose pile of unconnected folders; the citation line hanging from the bubble is dashed and attaches to nothing, and the scientist beside it stands with her arms folded. On the right, the same robot with the same number stands on a connected graph of models, and an unbroken chain runs from the bubble down to one validated model, which the scientist is pointing at.
Colleague, or confident guesser · The same number, twice. What changes is whether anything underneath it can be traced.

It is tempting to think a general-purpose assistant over the company drive gets you there. It doesn't. For any single study you could, in principle, gather every relevant file into one conversation, but that burden would land on every scientist, for every question, forever. Retrieval bolted onto a document pile produces plausible summaries with quiet errors, and after the second wrong number, adoption dies. What compounds is different: a living record of which models are trusted for which conditions, which studies produced which evidence, and who validated what, maintained as people work rather than assembled per question. That record cannot be generic. It has to live where the R&D work happens.

What transfers to R&D is the pattern, not the software: one working environment; the organization's tools behind one interface; versioned, executable artifacts; traceable outputs; humans and agents sharing the same loop. What doesn't transfer matters just as much: R&D's evidence is physical, its experiments cannot all be automated, and half the work happens outside any lab, in consumer panels, on store shelves, on plant floors. The pattern holds; the environment has to be built for that.

What an Integrated R&D Environment is

An Integrated R&D Environment (IRE) is the working layer where scientists and agents run the R&D loop together, sitting above the systems of record. The ELN, LIMS, and PLM remain the record of truth; the IRE is where the day-to-day work happens against them.

What it lets a scientist do is concrete: find and reuse what the organization has already learned. Run any model the organization trusts, without owning a Python environment. Ask how any number was produced and get a chain of evidence instead of a meeting. And hand entire stretches of the loop to an agent that works under those same rules. Underneath: models run in a shared environment, and the connections between models, runs, studies, and people are captured from the work itself, with human review where it matters.

Here is the same Tuesday morning inside an IRE. She drops the study folder into the workspace, in whatever formats it arrived. The agent maps sixty-three files to thirty participants, asks one clarifying question instead of assuming, picks the validated version of the right model, runs it, and returns a read where every number carries its citation: which model, which version, which data. She asks how it knows something; the answer is a chain she can click through. And while she worked, the organization's record of what it knows grew by one study, without a separate filing step.

Two panels. On the left, a scientist holds her head at a desk ringed by six disconnected windows, a locked workstation with two people queued behind it, and a teetering stack of reports. On the right, the same scientist works calmly at one screen with a small friendly agent beside her, and a single continuous thread runs from the screen through a folder, a validated model, a chart, and a small constellation of connected nodes.
The same Tuesday, before and after · Same person, same desk. The six disconnected tabs become one thread she can follow from the folder to the answer.

Five working principles:

  • Every quantitative output is grounded in a traceable calculation, model run, or source. A claim without evidence doesn't ship.
  • Models are versioned, owned, and validated, and they carry the conditions they were built under. The environment says when a model shouldn't be trusted for a question, and proposes the experiment that would fix that.
  • Provenance is captured from the work itself, with selective human review, never as separate knowledge-management labor.
  • The systems of record stay the systems of record. The IRE references them; it never replaces them.
  • Start narrow, compound underneath. Each piece pays for itself; the connected record underneath is what compounds.
Five icons in a row: a number in a speech bubble chained to a document; a gear carrying a version tag and a ribbon seal; a walking figure leaving a trail of connected dots; a thin working layer resting above three database cylinders; a staircase of bricks whose first courses are solid blue.
The five principles · Grounded numbers, validated models, provenance as exhaust, a layer above the records, and a build that starts narrow.

What it is not

Not a new ELN, LIMS, or PLM: it feeds them and reads from them, it does not compete with them. Not another knowledge-management program: nobody's job becomes filing. And not a generic copilot: the value is the organization's own validated models and evidence, wired into how its R&D actually decides.

How it starts

Not with another system of record. Large organizations already run serious multi-year programs for ontologies, data models, and consolidation, and those matter. But a program whose value arrives in year three does not change how anyone works on Tuesday.

Value has to land on day one, on the first team that touches it, or culture never moves.

timevalue deliveredday one
Value from day one · A multi-year program asks for patience it may not get. A first piece lands on the first team that touches it, and compounds from there.

So it starts with models, for the reason the economics already gave: validated models are the fastest reusable verification an organization owns, they exist in every function today, trapped, and making them runnable pays the first team that touches them, the first week. That is the wedge. Not search, not workflow, not another integration project, because none of those change what a scientist can verify on day one.

Two pieces get built first. Make every model the organization already has runnable by anyone and callable by agents: upload once, run from anywhere. And let the record of connections build from that work, linking models, runs, studies, and people from the first week. From there, pieces get added where each team needs them: simulated consumer panels for concept screening, digital twins for process work, whatever the domain calls for. The foundations stay the same.

A wall mid-construction. Its two lowest courses are solid blue, the bottom one patterned with gears and the one above it with connected dots. The courses above are drawn in light outline, each carrying a faint pictogram: a consumer panel, a factory, a speech bubble, a shopping cart. Two figures are laying the next outlined piece together under a simple sun.
The construction · Two courses go down first, the shared runtime and the record of connections. Everything above them is added per domain, and each piece pays for itself.

What changes

Fewer days spent waiting for evidence. Fewer studies repeated because the first one couldn't be found or trusted. Models reused across teams instead of dying on the laptop where they were born, which matters more every month now that engineers everywhere have coding agents and are producing more tools than any organization can manually track. Answers that carry their own audit trail. And smaller teams able to answer bigger questions, because a single scientist with a good agent can traverse more of the loop without waiting on three other calendars.

That is the Integrated R&D Environment. We are building it, starting with models.

If you lead an R&D organization and this maps a little too well onto your Tuesday mornings, we'd love to compare notes.