one platform for scientific data

Store, govern, share, and compute on multidimensional data — built on the open standards our team authors.

get started

One data model, top to bottom

Scientific data belongs in tensors, not tables — and the tools built for tables can't manage it. Earthmover is a full stack for multidimensional data: an open-source storage engine at the bottom, management and governance above it, and compute services on top — reached by people and by agents over the same interfaces. Each layer stands on its own, and they are designed to work together.

interface

Humans

Web appCLIPython

Agents

MCPSkillsCLIPython
platform

Scalable tensor compute services

TilesEDRDAPopenEOZax ServerSQL

Data management and governance

Data catalogCredential vendingAccess controlsData marketplace
open source

Tensor storage engine

ACID transactionsData versioningBest-in-class I/OVirtualization of legacy formats
cloud

Bring your own bucket

Or use storage we manage for you

AWSGoogle CloudAzureCloudflareJASMIN generic S3-compatible

Icechunk

Tensor storage engine

A Zarr-compatible transactional storage engine, written in Rust. ACID transactions, git-like version history, best-in-class cloud I/O, and zero-copy virtualization of NetCDF, HDF5, GRIB, and TIFF archives. Apache 2.0, serverless, and yours to run anywhere.

explore icechunk →

Arraylake

Data management and governance

A managed data lake platform for your Icechunk repositories: a catalog that natively understands arrays, single sign-on and fine-grained access control, credential vending so users never handle cloud keys, and a marketplace for sharing data across organizations.

explore arraylake →

Compute

Scalable tensor compute services

Compute that runs next to your data and serves it over the protocols your tools already speak — map tiles, EDR, DAP2, openEO. Turn a repository into a data product without building or operating a delivery stack.

explore compute →

Data you don't have to ingest

The same platform that manages your own data can deliver someone else's. Subscribe to analysis-ready weather, climate, and Earth-system datasets from trusted providers and they appear as repositories in your organization — versioned, updated incrementally, and queryable immediately. No download, no ingestion pipeline to maintain.

Design principles

Cloud native

Object storage as the primary data layer, storage and compute that scale independently. No filesystem assumptions anywhere in the stack.

Open standards

Built on Zarr, an Open Geospatial Consortium standard, and Icechunk, an Apache-2.0 project. No proprietary format stands between you and your data.

Bring your own bucket

Your data stays in your object storage, in your cloud account, under your policies — or use storage we manage for you.

Built for humans and agents

The same catalog, permissions, and APIs serve a scientist in a notebook, a production pipeline, and an AI agent exploring your archive.

Want to see the platform in action?