Тёмный

Fokko Driesprong - Tip of the Iceberg 

Plain Schwarz
Подписаться 2,9 тыс.
Просмотров 815
50% 1

Apache Iceberg is a high-performance format for huge analytic tables. Iceberg brings the reliability and simplicity of SQL tables to big data while making it possible for engines to work with the same tables, at the same time. Iceberg is a layer on top of your traditional Parquet tables with all the best practices from the database world. Using this you can do ACID operations on a table that solely lives in cloud storage.
In the talk, I'll first introduce Iceberg and its history, and the companies that are using and actively contributing to it. We'll take a peek under the hood and I'll explain the different concepts such as metadata, manifest lists, and manifest itself, and how it uses this to help the query engine, and maintain correctness. Next, I'll go through the schema, partition, and sorting evolution and how this is done in a lazy fashion so you don't have to rewrite your multi-petabyte table, and finally I'll do a quick demo using PyIceberg.
Speaker: Fokko Driesprong
More: 2023.berlinbuzzwords.de/sessi...
Web: 2023.berlinbuzzwords.de/
Fediverse: floss.social/@berlinbuzzwords
Linkedin: / 13978964
Twitter: / berlinbuzzwords

Развлечения

Опубликовано:

 

26 июл 2024

Поделиться:

Ссылка:

Скачать:

Готовим ссылку...

Добавить в:

Мой плейлист
Посмотреть позже
Комментарии    
Далее
How To Price For B2B | Startup School
17:46
Просмотров 25 тыс.
Да блин 😀
0:19
Просмотров 4,4 млн