About the Role
Massive is building the financial market data platform developers actually want to build on. This role owns the middle of it: taking raw market, reference, and alternative data and turning it into modeled, well-shaped, fast datasets while also designing the API surface customers query them through.
This is a deliberately full-stack data role, spanning what many companies split into "data engineer" and "analytics engineer." You'll do some ingestion - mostly consuming third-party APIs, which means rate limits, pagination, incremental loads, backfills, and schema drift - but the center of gravity is everything after the data lands: cleaning it, modeling it, and deciding how it should be physically laid out so that queries stay fast as volume grows.
Most of this work is zero-to-one. You'll be handed a dataset and a direction rather than a spec, and you'll be expected to decompose it, make the calls, and drive it to something in production that other people depend on. The team is small enough that you'll wear several hats in a given month - modeling on Monday, chasing a performance regression on Wednesday, debating about an API contract on Friday.
We're looking for someone genuinely interested in the data itself: what the fields mean, how entities relate, where the edge cases live, and how all of that maps to the question a customer is actually trying to answer.
Responsibilities
- Model the data. Design the entities, relationships, and grains that turn raw feeds into datasets people can reason about - and document why the model is shaped the way it is.
- Build and own the transformation layer in Python and SQL: cleaning, normalization, validation, and reconciliation across overlapping sources.
- Ingest from external APIs and feeds, handling rate limiting, pagination, incremental and backfill loads, schema-drift detection, and replay.
- Preserve the source. Treat raw provider data as immutable and complete, and build cleaned and curated layers on top of it
- Make it fast. Choose partitioning schemes, sort orders, file sizes, clustering, and indexes. Profile with EXPLAIN / query plans and engine metrics, and fix the slow path instead of adding hardware.
- Partner with Product to design the API surface customers use to consume these datasets - resource and query design, filtering and pagination semantics, contracts, and versioning.
- Partner with Data Science and AI Engineering, who are among your primary internal customers, on the datasets and features their work depends on.
- Set the quality bar: validation at ingest, anomaly and drift detection in production, freshness and completeness expectations, and clear documentation of assumptions and known limitations.
Skills & Qualifications
- Strong Python and SQL. Both, daily, and your SQL goes well past joins and aggregates - window functions, CTEs, and an instinct for what the planner is going to do with it.
- Real experience modeling analytical data. You can walk us through a schema you designed, the tradeoffs you made, and what you'd change today. You know when to normalize and when to denormalize, and you can explain why.
- A track record of making tables and queries faster - partitioning, file layout and sort order, indexing or clustering, and reading query plans to find the actual bottleneck rather than guessing.
- Hands-on experience with at least one modern analytical engine or lakehouse: DataFusion, DuckDB, Databricks, Snowflake, ClickHouse, Trino, BigQuery, or similar.
- Depth, not just exposure, in columnar formats and object storage. Parquet, Iceberg or Delta, S3-compatible storage - and you can explain why a columnar format wins for a given access pattern, not just that it does.
- You drive. You take an ambiguous problem, break it down, pick a path, and get to something working - then say clearly what you'd fix next. You know when to ship a workaround versus fix the root cause.
- You explain your thinking well. A lot of this job is making a modeling or storage decision legible to a researcher, an AI engineer, or the next person to touch the pipeline. If you can make a hard concept feel simple, that counts here.
- Fluency with AI coding tools as a daily companion. We expect you to use Claude Code, Cursor, or equivalents to move faster on parsing, scaffolding, refactors, and boilerplate - and to steer them deliberately: good context, verification against the actual data, and rejecting output that's confidently wrong.
- Curiosity about the domain. Prior financial data experience is a real advantage, but we would rather hire someone who asks sharp questions about a dataset than someone who has already seen this exact one.
- Rust is preferred. We reach for it where performance matters, including work in and around DataFusion.
- Arrow and Arrow Flight SQL for moving and serving result sets efficiently.
- Experience working alongside data science or ML teams - feature pipelines, dataset versioning, evaluation sets.
- Financial or market data is preferred: equities, options, corporate actions, reference data, fundamentals, or alternative data.
- dbt, SQLMesh, or similar transformation and semantic-layer tooling.
- Extracting structured data from irregular sources (XBRL, HTML, PDF).
About Massive
At Massive, we’re on a mission to modernize Wall Street by empowering developers with the tools to shape the future of finance. We’re reimagining financial market data for the 21st century, removing barriers, simplifying access, and creating frictionless, forward-thinking technologies.
Join us and become part of a passionate team that consistently sets new industry standards, creating a profound impact on the world of finance and technology, and leveling the playing field by providing fair access for all.

