纠偏成为下一次的基准。
纠偏变成回归用例,阈值不过回归就改不了。生产中已拦下一次错误阈值、发现两例漏采;下一步,让签核后的判断成为任务模型训练数据。
给申办方直接使用的 AI 原生临床操作系统。交付结果,也交付结果背后的完整证据链:从源数据到人类签核,每个结论都有出处,每次判断都留下依据。
从零到日常使用
注册临床获得支持
支持多项研究
其中一项已获 FDA IND 批准
据 2026 年 9 月公司材料,系统在一家细胞治疗(ATMP)企业生产运行;合作方按协议匿名。查看完整背景 ↗
从 IIT 与注册临床辅助切入,再扩展至细胞治疗和 ATMP 企业。申办方、研究中心与 CRO 可以在同一条证据链上协作,让项目知识留在申办方。
连接原始材料、分析结果与人类签核,为注册路径上的研究提供可追溯的上下文。Helix 已在生产中支撑一项注册临床。
从决策链更短的研究进入,让团队聚焦科学问题。Helix 在生产中同时支持多项 IIT,其中一项已获 FDA IND 批准。
让新项目、新中心加入同一张上下文图谱。以企业年度订阅扩展,后续通过 CRO / SMO 渠道服务更多研究团队。
执行、校准、记忆,三个侧面共同维持临床上下文。证据链经得起追问,靠的是每一步都与其他每一步对得上。
底层安全:沙箱执行、六维权限、标识与数据分离;AI 见不到患者是谁。
数据进入申办方自己的临床数据湖仓。连接器、OCR 与视觉语言模型解析原始材料,用本地授权的受控术语归一,再在质量门禁下发布。
DAG 编排与质量门禁:不过门禁,不发布。每次发布形成带字段级血缘的不可变版本快照。
MedDRA · CDISC SDTM · LOINC · UCUM基础大模型目前仍通过外部 API 调用,由可切换供应商的接入层连接。查询、交接与预警使用同一快照号,项目资产留在申办方。
随着任务模型、评测集和上下文累积,逐步降低对单一模型,尤其闭源模型的依赖。完全私有化是建设目标,尚不作为当前能力宣称。
Co-pilot 可以起草、引用、分析、提议。
签字的,永远是人。
今天,让草稿、答案与记录对齐同一份上下文。下一步,把剂量爬坡、观察窗口、项目去留和监管质疑等决策,连同完整证据链一起交到人面前。
决策支持是发展方向;临床判断与签核始终由人负责。
试验留下规则、阈值、评测用例与项目上下文。模型可以换,供应商可以换,申办方学到的一切都留下。CRO 仍可在 Helix 中工作,系统成为团队共同的记忆。
纠偏变成回归用例,阈值不过回归就改不了。生产中已拦下一次错误阈值、发现两例漏采;下一步,让签核后的判断成为任务模型训练数据。
新项目、新中心加入同一张上下文图谱,避免再次形成信息孤岛。
执行框架在每轮工作中积累规则、评测与项目记忆,下一项研究从已有经验起步。
从真实研究出发,把临床判断与工程能力放在同一支团队里。
生物医药 CMO / COO,20 年中、美、加、澳临床与注册经验。
百余项 I–III 期研究,5 项 BLA;具备 FDA 快速通道、孤儿药与突破性疗法认定经验。
前硅谷 GenAI 头部公司技术负责人、AI Lab 产品负责人,交付过千万级日活应用。
前线 AI 工程交付团队已在岗,将临床工作与系统开发直接连接。
目标约 80% 为系统收入。定价锚定申办方价值,收入不随交付人头线性增长;客户买到的是持续沉淀自己资产的系统。
以下为 2026 年 9 月保密材料中的公司测算与融资计划,保留原文口径,未经独立核验。
自下而上按 75 家细胞治疗申办方、180 个 IIT 中心与稳态客单价测算。
募集资金约 60% 计划用于临床执行框架、任务模型、评测集与专家知识蒸馏。具体条款在保密协议下单独沟通。
2026 年 9 月公司介绍,保密材料。按章节展开,阅读完整内容与原始口径。
Helix 已在一家细胞治疗(ATMP)领头羊 biotech 生产运行。支撑一项注册临床,预计未来半年内获批;同时支撑几项 IIT,其中一项已获 FDA IND 批准。数据进来,AI 干活,人签字,从数据到签字这条链是完整的。从零到日常在用,20 周。这是生产系统,不是 pilot。
Helix 是一套给 sponsor 直接使用的 AI-native clinical OS。一项试验,不管是 IIT 还是注册临床,最后交出去的只有一样东西:一条证据链。今天这条链靠八个职能的人手工拼起来,每一次交接都会断。行业的数字就是结果:进入 I 期的项目不到 8% 走到获批,耗时以年计,单药全周期成本约 $0.32B,算上失败分摊更高。Helix 交付的不只是结果,更是结果背后那条链:每个结论可追溯到出处、经得起监管稽查、随着人的每一次签核持续校准对齐。我们用 AI-native 的方式围绕这条链把临床流程重建了一遍,并且它在生产里跑起来了。这是通往 GCP 一直设定的那个目标的一条新路,而且走在合规的路上,不是绕开合规。临床开发每年仍花 $85–126B 买人手工搬运上下文,而其他信息密集行业早已围绕替人干活的软件重建过一遍。
前沿模型不会自己变成临床系统。真正难的不是让 AI 聪明,是让一整套东西彼此一致:模型看到的材料、它被允许做的事、人签过的决定、试验实际走到哪一步,四者随时对得上。我们把这叫 coherence,包在模型外面的那层 harness 就是为维持它而建的。六个构件,三个侧面。执行侧:Policy Engine 装着规则、阈值、门禁和受控生成插槽,模型只在划定的边界里生成;End-to-End Task Models 是按具体临床任务微调蒸馏的专用模型。校准侧:Eval Harness 由红线测试、回归用例和真实项目校准样本组成,任何改动都要先过它;Decision Ledger 记下谁在什么依据下改了什么、为什么。记忆侧:Context Graph 把受试者、样本、事件、访视、方案版本当一等对象,带时间轴、版本和血缘,而不是文本块;Work Grammar 让交接、验收、决定登记在所有项目里用同一种写法。底下是安全层:沙箱执行、六维权限、带字段级血缘的不可变快照、标识与数据分离,AI 见不到患者是谁。最上面一条规矩:AI 起草,人签核,不过门禁的东西不会成为记录。每一次签核都是一次校准,让整套系统重新对齐一次。一条证据链经得起稽查,靠的不是某一步聪明,而是每一步都和其他所有步对得上。这就是 coherence,其他一切都站在它上面。
harness 跑在申办方自己的临床数据湖仓之上,这是我们建的另一半。数据从 EDC、中心实验室、影像、PDF 和既往归档进来,由连接器、OCR 和视觉语言模型解析,用本地授权的受控术语(MedDRA、CDISC SDTM、LOINC、UCUM)归一。之后按四个状态投影:原始采集、SDTM、ADaM、TFL,由 DAG 编排并设质量门禁,不过门禁不发布,发布的是带字段级血缘的不可变版本化快照。查询、交接、预警读的都是同一个快照号,一个数字在哪里都是同一个数字。这里是申办方的数据变成复利资产的地方,也是专用任务模型训练的底子。
harness 和湖仓是申办方私有的。底层大模型目前还不是:今天仍通过外部 API 调用,中间有一层可以换供应商。这是这个阶段的有意安排,不是终态。随着湖仓越填越满、harness 累积起任务模型、评测集和上下文,依赖任何单一模型、尤其是闭源模型的那部分工作会持续缩小。终态是完全私有化的闭环部署。我们是往那里建,不提前宣称。
Helix 今天发挥作用的地方主要是辅助:人在干活,系统让整个试验保持 coherent,把每一份草稿、每一个答案、每一条记录对齐到同一份上下文。这是必要的地板,已经在生产里。但价值不止于此。每多用一轮,系统就离决策本身更近一步:剂量怎么爬、观察窗开多长、项目要不要继续、监管会从哪里质疑。这些决策才是申办方真正买单的东西,每一个都值几个月和一大笔钱,也是做决策的人一走就丢掉的东西。Helix 的方向是把决策连同完整的证据链一起摆到桌上,让人签的是一个决策,而不是自己去拼一个决策。
CRO 的工作随合同结束离开申办方,人、判断和上下文一起走。Helix 站在这条线的申办方一侧。每一项在它上面跑过的试验都留下规则、阈值、评测用例和项目上下文,归申办方所有,下一项从这里起步。三条曲线叠在一起。学习曲线:纠偏变成回归用例,阈值不过回归就改不了,生产里已经拦下过一次错误的阈值、抓出过两例漏采;下一步,签过的判断变成专用任务模型的训练数据。增长曲线:每个新项目、新中心都加进同一张上下文图谱,而不是再开一个孤岛。自生长曲线:harness 每转一圈厚一层,不需要有人专门去建。CRO 仍是申办方想要人手的地方的人手,而且可以在 Helix 里干活;Helix 是这些人手赖以工作的记忆。模型可以换,供应商可以换,申办方学到的一切都留下。
订阅为主,按项目加按席位,交付物计价叠加,目标约 80% 为系统收入。选订阅,因为收入可预测、不随人头增长;因为定价锚的是申办方的价值而模型算力持续降价;因为申办方买到的是一套会自己沉淀资产的系统,而不是每次重新买的服务。从决策链最短的进入:先是 IIT 项目与注册临床辅助,再是细胞治疗 / ATMP biotech 的企业年度订阅,再是 CRO / SMO 渠道。市场:eClinical 软件今天是 $11.5B 的市场(严口径 TAM);把系统能吃掉的 CRO 服务算进来是 $71–100B(宽口径 TAM)。SAM:中国加亚太可服务治疗领域的临床支出自上而下 $4.0–8.7B;自下而上按 75 家细胞治疗申办方加 180 个 IIT 中心的具名名单、稳态客单算,$0.21B。
联合创始人 · 临床:生物医药 CMO / COO,20 年中美加澳临床与注册,百余项 I / II / III 期研究、5 项 BLA,FDA 快速通道、孤儿药、突破性疗法认定。联合创始人 · 产研:前硅谷 GenAI 头部公司 tech lead、AI Lab product lead,千万级 DAU 应用。AI FDE 交付团队已在岗。
Susie Tan Co-founder | Jacky Xu Co-founder
锚定方按协议匿名。效率数字为内部试算目标区间,非已验证 SLA。
The AI-native clinical operating system sponsors run themselves. We deliver the evidence chain, not just the result.
Helix runs in production at a leading cell-therapy (ATMP) biotech. It supports one registrational trial, on track for approval within the next 6 months, and several IIT, one of which has cleared FDA IND. Data comes in, AI does the work, human in the loop, and the chain from data to signature stays intact. Zero to daily use took 20 weeks. This is a production system, not a pilot.
Helix is an AI-native clinical OS that sponsors use directly. Every trial, investigator-initiated or registrational, ultimately submits one thing: an evidence chain. Today that chain is assembled by hand across eight functions and breaks at every handoff. The industry's numbers show it: fewer than 8% of programs entering Phase I reach approval, after years of development and roughly $0.32B in full-cycle cost per drug, more once failures are counted. Helix delivers not just the result but the chain behind it: every conclusion traceable to source, auditable by a regulator, and continuously calibrated against what humans sign. We rebuilt the clinical process AI-natively around that chain, and it works in production. It is a new path to the goal GCP has always set, and it runs on the compliance road rather than around it. Clinical development still spends $85–126B a year on people carrying context by hand, while every other information-heavy industry has been rebuilt around software that does the work.
A frontier model does not become a clinical system on its own. The hard part is not making the AI clever; it is keeping everything coherent: what the model may read, what it may do, what humans have signed, and where the trial actually stands, all consistent with one another. We call that coherence, and the harness around the model exists to maintain it. Six components on three sides. Execution: a Policy Engine holds rules, thresholds, gates and controlled generation slots, so the model generates only inside a defined boundary; End-to-End Task Models are fine-tuned and distilled for specific clinical tasks. Calibration: an Eval Harness of red-line tests, regression cases and real-project calibration samples that every change must pass; a Decision Ledger that records who changed what, on what basis, and why. Memory: a Context Graph where subject, sample, event, visit and protocol version are first-class objects with timelines, versions and lineage, not text chunks; a Work Grammar that gives handoffs, acceptance and decision registration one shared form across projects. Beneath it all, the safety layer: sandboxed execution, six-dimension permissions, immutable snapshots with field-level lineage, and identity kept apart from data so the AI never sees who the patient is. One rule above all: the AI drafts, a human signs, and nothing becomes a record without passing a gate. Every signature is a calibration that brings the system back into alignment. An evidence chain survives an audit not because any one step is clever, but because every step agrees with every other. That coherence is the ground everything else stands on.
The harness runs on the sponsor's own clinical data lakehouse, the other half of what we built. Data arrives from EDC, central labs, imaging, PDFs and prior archives; connectors, OCR and vision-language parsing bring it in, and locally licensed controlled terminology (MedDRA, CDISC SDTM, LOINC, UCUM) normalizes it. It is projected through four states, raw capture, SDTM, ADaM and TFL, under DAG orchestration with quality gates: nothing ships until it passes, and what ships is an immutable, versioned snapshot with field-level lineage. Queries, handoffs and alerts all read the same snapshot ID, so a number is the same number everywhere. This is where the sponsor's data becomes a compounding asset, and the ground the task models are trained on.
The harness and the lakehouse are private to the sponsor. The foundation model is not yet: today it is called through external APIs behind a layer that lets us switch vendors. That is deliberate for this stage, not the end state. As the lakehouse fills and the harness accumulates task models, eval sets and context, the share of work that depends on any single model, closed models especially, keeps shrinking. The end state is a fully private, closed-loop deployment. We are building toward it, not claiming it early.
Today Helix earns its keep as assistance: keeping the trial coherent while people do the work, aligning every draft, answer and record to one context. That floor is in production. It is not where the value ends. Each cycle of use moves the system closer to the decisions themselves: how to escalate a dose, how long to hold an observation window, whether a program should proceed, where a regulator will push back. Those decisions are what a sponsor pays for. Each is worth months and real money, and each walks out the door with the person who made it. Helix's direction is to put the decision on the table with its full evidence chain attached, so the human signs a decision rather than assembles one.
A CRO's work leaves with the contract; people, judgment and context go with it. Helix sits on the sponsor's side of that line. Every trial run on it leaves rules, thresholds, eval cases and project context behind, owned by the sponsor, and the next trial starts from there. Three curves stack. Learning: corrections become regression cases, a threshold cannot change until it passes them, and in production this has already stopped a wrong threshold and caught two missed cases; next, signed judgments become training data for task models. Growth: each new program and site adds to the same context graph instead of opening a new silo. Self-growth: the harness gets thicker with every cycle without anyone being asked to build it. CROs remain the sponsor's hands wherever it wants hands, and they can work inside Helix; Helix is the memory those hands work from. Swap the model or the vendor, and the sponsor keeps everything it has learned.
Subscription first, by program and by seat, with priced deliverables on top; target mix around 80% system revenue. Subscription because revenue is predictable and does not scale with headcount, because price is anchored to sponsor value while model compute keeps getting cheaper, and because the sponsor is buying a system that accumulates its own assets rather than a service bought again every time. Entry through the shortest decision chain: IIT programs and registrational-trial support first, then cell-therapy and ATMP biotechs on annual enterprise subscription, then CRO/SMO channels. Market: eClinical software is an $11.5B market today (TAM, strict); $71–100B if the CRO services a system can absorb are counted (TAM, broad). SAM: $4.0–8.7B of clinical spend in serviceable therapy classes across China and APAC, top-down; $0.21B bottom-up from a named pool of 75 cell-therapy sponsors and 180 IIT centers at steady-state ACV.
Co-founder, Technology: former tech lead at a top Silicon Valley GenAI company and AI Lab product lead; shipped applications with 10M+ DAU. Co-founder, Clinical: biopharma CMO/COO with 20 years across China, the US, Canada and Australia; 100+ Phase I–III studies; five BLAs; FDA Fast Track, Orphan Drug and Breakthrough designations. A forward-deployed AI engineering team is already in place. Raising a Seed round; about 60% of proceeds go to the harness, task models, eval sets and expert knowledge distillation. Terms shared one-on-one under NDA.
GenPrime AI · Susie Co-founder · Jacky Xu, Co-founder
Anchor partner anonymised under agreement. Efficiency figures are internal target ranges, not validated SLAs. ALCOA+ and 21 CFR Part 11 are reference standards; no certification claimed.
“未来半年内获批”为 2026 年 9 月材料中的预期,不代表已获批。效率数字为内部目标,非已验证的服务承诺。ALCOA+ 与 21 CFR Part 11 为参考标准,不宣称认证。
本页系统关系为文字与流程示意。真实世界研究仍保留为拓展方向,不列入上述生产成果。
无论是正在推进的研究、技术集成,还是合作与投资交流,都可以从具体问题开始。填写右侧信息,生成写给 Jacky Xu 的邮件草稿。