RoleHunter

Job description

Overview

Physical Superintelligence is a startup with roots at Google, NVIDIA, Harvard, Meta, MIT, Oxford, Johns Hopkins, Cambridge, and the Perimeter Institute building AI systems to discover new physics at scale. We are seeking engineers to build platform infrastructure at the intersection of computational science, AI systems, and software engineering.

Our mission is to discover and commercialize transformative physics breakthroughs at scale with artificial superintelligence, safely, verifiably, and for broad public benefit.

The last century's golden age of physics gave us transistors, lasers, and nuclear energy. We believe artificial superintelligence will unlock the next one. We're creating the infrastructure to industrialize scientific discovery and usher in this new era.

We have one product: new physics, at scale.

We are seeking a Member of Technical Staff, Data Systems to own the data our results rest on: everything the platform's runs produce, and the shared corpora every team depends on. These systems have a property most infrastructure doesn't. Their failures cannot be fixed by a rewrite. A provenance field you didn't capture at collection time cannot be added later, at any price. We are hiring the engineer who designs so that it never happens.

 

Role and Responsibilities

  • Build the data plane for scientific results at scale: content addressing, lineage captured at write time, rights metadata carried on the object rather than bolted on after. Every result should be traceable to what produced it, mechanically, without anyone having to remember to do it

  • Scale the catalogs and query paths that make a billion-object corpus usable. Metadata operations break before byte volume does. You design for the scale before it arrives, and you own the database layer the catalogs live on.

  • Own the databases behind the platform, not just the object store. Transactional stores, analytical query paths, filesystems, schema evolution, and the judgment of which engine fits which job. The data plane is a system of systems, and you reason across all of it.

  • Own ingestion of the shared datasets every team depends on: scientific corpora, public data feeds, whatever comes next. Brought in once, licensed once, versioned, served to everyone. Not re-downloaded per team.

  • Own the replicate-versus-fetch economics across storage tiers and clouds. What lives where, what it costs to keep it there, what it costs to move it. Placement decisions get modeled, argued in dollars, and revisited as the corpus grows.

  • Understand data for AI, because that is what most of this data is for: training sets, eval corpora, and captured traces, each versioned, rights-tracked, and lineaged, so any model or eval can be traced to the exact data that built it and rebuilt from it.

What We're Looking For

  • Five or more years building data infrastructure in production at companies known for engineering rigor. A data generalist: databases, storage systems, filesystems, catalogs, and the data platforms other teams built on every day. An engineer who scales systems, not an administrator who operates them. Platform ownership, not pipeline authorship.

  • You think in data identity: immutable content, names as pointers to versions, provenance that cannot drift from the data it describes. You know what lineage costs when it is an afterthought, and how rights and retention get enforced in the data path rather than in a policy document.

  • You have watched a catalog, namespace, or metadata service break before the storage under it did. You can explain what you changed.

  • You can produce a replicate-versus-fetch break-even from memory, argue a tiering decision in dollars, and defend a placement policy to the person paying the bill.

Nice to Have

  • Content-addressed storage, data versioning, or lineage systems in production, whether built on object-store internals or from scratch.

  • Checkpointing at GPU scale: absorbing the burst when a training run dumps state from every card at once, and making restore fast enough to matter.

  • Distributed-systems depth alongside the data work: schedulers, workflow engines, high-volume services. Our data and execution work sit next to each other, and the engineer strong in both is the one we most want to meet.

  • Experience with scientific or research data: instrument output, simulation results, dataset licensing, or archives meant to be readable a decade later.

  • Table formats and catalogs at scale, and the economics of multi-cloud data placement.

How We Work

We hold a high technical bar and give people full ownership of their work, from spec to ship to on-call. We write contracts before logic, test against real systems instead of mocks, and favor simple designs that ship over clever ones that do not. Our development process is AI-native: we work with agentic coding tools daily, write specs that are legible to humans and agents alike, and lead with leverage.

Location and Compensation

This role is based in Boston. We will consider remote candidates on a case-by-case basis. We offer competitive compensation including salary, benefits, and meaningful early-stage equity. We evaluate on technical breadth, systems thinking, scientific curiosity, and shipping velocity. We are an equal opportunity employer and value diverse perspectives in building platforms for AI-driven discovery.

Member of Technical Staff, Data Systems at Psi | rolehunter