#Icechunk

35 posts

Science data comes in tables and tensors

Science data comes in tables and tensors

Apache Iceberg comes to Arraylake. Tables now live alongside Icechunk arrays in the same catalog, under the same permissions, in your own buckets — queryable from DuckDB, Spark, Trino, Snowflake, and Zax-SQL.

Joe Hamman
Joe Hamman

CTO & Co-founder

Matt Iannucci
Matt Iannucci

Staff Engineer

Our compute roadmap: introducing Zax SQL

Our compute roadmap: introducing Zax SQL

We're introducing Zax, a new query and compute engine for multidimensional scientific data, and Zax-SQL, the first interface to it. Zax plans your computation for you, so you are no longer responsible for tuning chunk sizes, worker counts, and memory limits.

Ryan Abernathey
Ryan Abernathey

CEO & Co-founder

Brian Davis
Brian Davis

Head of Product

Joe Hamman
Joe Hamman

CTO & Co-founder

Deepak Cherian
Deepak Cherian

Principal Research Engineer

Sebastian Galkin
Sebastian Galkin

Head of Engineering

Making open-source AI weather forecasting models easy to run

Making open-source AI weather forecasting models easy to run

A joint post from Earthmover and Hugging Face on running open-source AI weather forecasting models like Aurora using free ERA5 data from the Earthmover Marketplace — with a live demo space and a full local tutorial.

Emma Scharfmann
Emma Scharfmann

ML for Science, Hugging Face

Aaron Spring
Aaron Spring

ML, AI and geospatial data engineering

Ryan Abernathey
Ryan Abernathey

CEO & Co-founder

Kashif Rasul
Kashif Rasul

AI Research Science, Hugging Face

Having it all: virtual and materialized data products

Having it all: virtual and materialized data products

A guest post from dynamical.org on how virtual Icechunk Zarrs complement materialized products to deliver low-latency, complete, and fast-access weather data — plus the low-latency NOAA HRRR dataset it describes, now on the Earthmover Data Marketplace.

Alden Keefe Sampson
Alden Keefe Sampson

dynamical.org

Marshall Moutenot
Marshall Moutenot

dynamical.org

The University of Southern Mississippi makes 90 million ocean profiles conversational: AQUAVIEW's approach to the World Ocean Database

The University of Southern Mississippi makes 90 million ocean profiles conversational: AQUAVIEW's approach to the World Ocean Database

How AQUAVIEW reorganized NOAA's World Ocean Database around instrument types, powered it with Arraylake, and built a natural language interface that turns hours-long data workflows into seconds.

Henry Jones PhD
Henry Jones PhD

Director of Research Development and Scientific Entrepreneurship, University of Southern Mississippi

Joshua Hill PhD
Joshua Hill PhD

Director, Institute for Advanced Analytics and Society, University of Southern Mississippi

How Kettle Uses Earthmover to Power Wildfire Risk Modeling at Scale

How Kettle Uses Earthmover to Power Wildfire Risk Modeling at Scale

The Company Kettle is not a typical insurance company. Using AI to build smarter insurance products, Kettle provides insurance for property owners in areas affected by catastrophic climate events, with a particular focus on wildfire. Their AI models consume over 130 terabytes of satellite, weather,

Margaret Francis
Margaret Francis

COO

Filtered Subscriptions: Fine-Grained Cloud-Native Data Sharing

Filtered Subscriptions: Fine-Grained Cloud-Native Data Sharing

Earthmover's filtered subscriptions allow data providers to create secure, read-only views into multidimensional data cubes, enabling more granular cloud-native data exchange between provider and consumer.

Ryan Abernathey
Ryan Abernathey

CEO & Co-founder

Sebastian Galkin
Sebastian Galkin

Head of Engineering

Lindsey Nield
Lindsey Nield

Senior Engineer

I/O-Maxing Tensors in the Cloud

I/O-Maxing Tensors in the Cloud

Zarr Python with Icechunk or Obstore now fully saturates the network between EC2 and S3, achieving the physically maximum possible throughput for reading and writing tensor data in the cloud. Benchmarks compare Zarr, Tensorstore, TileDB, and Parquet stacks across a range of chunk sizes and instance types.

Ryan Abernathey
Ryan Abernathey

CEO & Co-founder

Multi-Player Mode: Why Teams That Use Zarr Need Icechunk

Multi-Player Mode: Why Teams That Use Zarr Need Icechunk

Zarr lacks built-in support for concurrent readers and writers, leading to inconsistent reads and conflicting writes in team settings. Icechunk solves this by adding atomic updates, consistent snapshots, and Git-like version control on top of Zarr.

Lindsey Nield
Lindsey Nield

Senior Engineer

Icechunk 1.0: Production-Grade Cloud-Native Array Storage Is Here

Icechunk 1.0: Production-Grade Cloud-Native Array Storage Is Here

Icechunk 1.0 is now stable and production-ready, bringing transactional safety, efficient versioning, high-performance Rust-based I/O, and virtual references for HDF5 and NetCDF to cloud-native array storage. The release includes manifest splitting, distributed writes, conflict resolution, and a 30 TB ERA5 sample dataset.

Ryan Abernathey
Ryan Abernathey

CEO & Co-founder

Icechunk: Efficient storage of versioned array data

Icechunk: Efficient storage of versioned array data

Icechunk stores versioned array data efficiently by never copying or rewriting existing chunks, so each new version only consumes storage for the data that actually changed. Older versions can be expired and garbage-collected when they are no longer needed.

Sebastian Galkin
Sebastian Galkin

Head of Engineering

Zarr takes Cloud-Native Geospatial by storm

Zarr takes Cloud-Native Geospatial by storm

At the 2025 Cloud-Native Geospatial conference, Zarr adoption was surging across the geospatial domain, with Copernicus Sentinel, USGS Landsat, Google Earth Engine, and ESRI ArcGIS all embracing the format for cloud-optimized array data.

Joe Hamman
Joe Hamman

CTO & Co-founder

Announcing Icechunk!

Announcing Icechunk!

Earthmover announces Icechunk, an open-source transactional storage engine for Zarr that brings ACID transactions, time travel, data versioning, and high-performance Rust-based I/O to multidimensional array data in cloud object storage.

Ryan Abernathey
Ryan Abernathey

CEO & Co-founder