Staff Software Engineer - Curated Data
Dune
Job Description
About Dune
Dune's mission is to make onchain finance observable. We're the industry standard for onchain data: a blockchain data and intelligence provider that institutions, protocols, and analysts trust to understand the onchain world and the frontiers of finance. We deliver structured datasets, spanning stablecoins, RWAs, tokens, lending, trading, and more, from 130+ chains and counting, to 1,000+ industry leaders including Visa, WisdomTree, FINRA, the IMF, Bloomberg, Standard Chartered, Coinbase, Forbes, and the Financial Times.
We're a tight knit team of 50 hard working people, spread across Europe and eastern US timezones, having an outsized impact on the industry. We take pride in being ambitious and doing world class work while staying humble and lighthearted. We believe in building open, verifiable data that lets individuals and institutions do deep research into ecosystems like Ethereum, Solana, Robinhood and many more.
We're backed by some of the world's best investors. In February 2022, we announced our Series B funding round led by Coatue and Union Square Ventures, an important milestone that let us double down on our mission.
About The Role
Data Products builds and owns datasets end to end: from raw chain data through decoding to the 3000+ models and 4 petabytes we curate, share directly with customers, and replicate into their warehouses.
The role will focus on the lifecycle of building high quality data: orchestrating thousands of interdependent models, propagating schema changes without breaking downstream consumers, and propagating corrections.
That is a software architecture problem in a data domain. This role is a hybrid: a backend engineer who thinks in systems and contracts, working on data.
You will be the engineer we hand ambiguous product requirements to, and will come back with a design, a sequence, and work the team can pick up, while building the hardest parts yourself.
Responsibilities
- Design and build the control plane for our curated data lifecycle: dependency-aware orchestration, backfills, restatements, retries, partial failure, and recovery
- Decide, dataset by dataset, whether the answer is a model, a service or a job, and own that architecture through production
- Design the contracts between ingestion and curation so a dataset can be reasoned about end to end
- Build alerting and data quality signals that catch real problems and stay quiet otherwise, so on-call is about incidents rather than noise
- Work across Go, Kotlin, Rust, Python and SQL, choosing the right tool rather than the familiar one
- Break large problems into work other engineers can own, and sequence it so we ship something useful early
Requirements
- You are a backend engineer who has gone deep on data systems, or a data engineer who became a strong software engineer. You ship production services, not only pipelines
- You have built or materially extended orchestration and scheduling systems, and can explain precisely what breaks at scale and why
- You have handled schema evolution and data correctness in a system with real consumers downstream, where a breaking change has a cost
- You have built or operated stateful stream processing in production (Flink, Kafka Streams, Spark Structured Streaming, RisingWave, Materialize, Feldera)
- You have strong SQL and modeling skills on large datasets, and an interest in how the query engine underneath actually executes your work
- You have solid computer science fundamentals and distributed systems understanding
- You debug independently and drive root cause analysis to a fix that holds
- You use AI tools well enough that they have changed how you work, you understand their failure modes and dislike ai-slop
- You communicate clearly in writing and get the best out of a distributed team
Nice to Have
- Deep experience with a transformation framework such as dbt or SQLMesh: specifically, having hit its limits and built beyond them
- Data lake formats such as Parquet, Iceberg or Deltalog
- Stateful stream processing in production (Flink, Kafka Streams, Spark Structured Streaming)
- Experience at a company where the data is the product
Benefits
- A competitive salary and equity package. Both salary and equity is top 25% of companies in the space
- Our employee equity scheme has world-class employee-friendly terms with a heavily discounted strike price (90%) and a 10-year exercise window
- 5 weeks PTO + local public holidays (that can be swapped to suit you)
- A fully remote-first approach within a distributed team with flexible working hours; you structure your own day
- A healthy mix of async and sync work, so you can focus on what truly matters—no more wasted time on endless meetings
- Private medical insurance, dental and vision as standard
- Paid parental leave to help you celebrate this important milestone, transition to your new life, and bond with your new baby. We offer 16 weeks to primary and 6 weeks to secondary caregivers, fully paid. Plus a 2-week part-time phased return at full pay to help you get used to your new (and slightly more complex!) schedule
- Quarterly offsites in various exciting locations as a company or team to connect, work together and have fun (so far in Tuscany, Berlin, Austria and Athens)
- Each person gets a yearly travel allowance to connect and co-work with someone or a team of people for a few days
- An allowance for your at-home setup, to ensure you are happy, comfortable and productive. If you prefer a local co-working space, we'll pay for your desk
- Work with some of the best people you'll ever get to meet
- Awesome Dune swag
We are dedicated to building a diverse, inclusive, and authentic workplace, so if you're excited about this role but your experience doesn't align perfectly with every qualification in the job description, we encourage you to apply anyway. You may be just the right candidate for this or other roles.
Unchain Data provides Web3 data job aggregation as a common good. Jobs are posted by third parties and are not individually verified. Always exercise caution: never download software requested during a hiring process, avoid clicking unfamiliar links in interviews, make sure to verify URLs are legit, and use trusted meeting tools like Google Meet or Zoom.
Want to land this role?
Practice SQL on real blockchain data. Build the query skills crypto data teams test for in interviews.
Explore WizardCampFurther reading
Similar Jobs
Senior Engineer: Real-Time AMM & Trading Data
Cielo · Remote - EU, US
Senior Analytics Engineer
Reap · Remote / Hong Kong
Mid/Senior-level GoLang Data Engineer (EU)
Inca Digital · Remote - Europe
Senior Software Engineer, Data & AI
Lightspark · Los Angeles, CA
Staff Software Engineer - Data
Dune · Oslo, Norway, Dominican Republic, New York, NY, USA, San Antonio la Isla, State of Mexico, Mexico
Want to land this role?
Practice SQL on real blockchain data. Build the query skills crypto data teams test for in interviews.
Explore WizardCamp
