I am an expert in database systems, and equally comfortable reading research papers, working through the mechanics of delivering production software, and everything in between. My goal is to help advance the state of the art in database processing in the industry as a whole, and to stay deeply technical while doing so. Throughout my career, I have helped shepherd technologies from research lab to industry multiple times, including Vertica (the commercialization of the C-Store research project), Apache DataFusion (state-of-the-art vectorized query processing), and the Eureqa machine learning algorithm.
Currently I am a Staff Engineer at InfluxData, working on the InfluxDB 3.0
time series database. I am part of the PMC (governance committees) of Apache
DataFusion, Apache Arrow, and Apache Parquet, and a Member of the Apache
Software Foundation. My time is split between working on InfluxData’s products,
and contributing to and maintaining the code and community of DataFusion and
arrow-rs, the Rust implementations of Arrow and Parquet. With Evan Kaplan,
I coined the term FDAP Stack to describe modern systems composed with these
reusable components.
I am a huge believer in the power of the open source model to create and maintain high quality software and the communities that sustain them over the long term. It is hard to overstate the enjoyment of working with great developers from around the world on software that powers thousands of companies and countless users each day.
Andrew Lamb is a Staff Engineer at InfluxData and a member of the Apache DataFusion, Apache Arrow, and Apache Parquet PMCs, with more than 22 years of experience in systems software and database internals. He enjoys helping bring new technologies from the research lab to industry, including Vertica (the commercialization of C-Store), Apache DataFusion, and the Eureqa machine learning algorithm. An expert in database systems, he is equally at home reading research papers and delivering production software, and has worked in environments ranging from two-person startups to large multinational corporations and worldwide distributed open source projects, in roles including architect, manager, and VP. He writes papers, gives talks, and blogs about software engineering and databases.
2024-09-23
Apache Arrow DataFusion: A Fast, Embeddable, Modular Analytic Query Engine (talk)
Carnegie Mellon University: Database Building Blocks Seminar Series - Fall 2024
(slides
recording
)
2024-06-19
Apache Arrow DataFusion: A Fast, Embeddable, Modular Analytic Query Engine.
Andrew Lamb, Yijie Shen, Daniƫl Heres, Jayjeet Chakraborty, Mehmet Ozan Kabak, Chao Sun, and Liang-Chi Hsieh
2024 International Conference on Management of Data (SIGMOD 2024), June 9-15, 2024, Santiago, Chile
(ACM DOI,
PDF
)
2012-08-27
The Vertica Analytic Database: C-Store 7 Years Later.
Andrew Lamb, Matt Fuller, Ramakrishna Varadarajan, Nga Tran, Ben Vandiver, Lyric Doshi, Chuck Bear.
38th International Conference on Very Large Data Bases, Proceedings of the VLDB Endowment, Vol. 5, No. 12 (VLDB 2012)
(PDF
,
PDF alternate
)