Posts about DuckDB
Articles and tutorials about DuckDB: running it in AWS Lambda, Cloudflare Workers and the browser with DuckDB-Wasm, querying Iceberg and S3 Tables, and building lightweight data lakes.
14 posts
Filter by tag
Filter by tag:

Custom DuckDB Wasm builds for Cloudflare Workers
Run DuckDB SQL queries, including remote Parquet, S3 and R2 reads, directly inside Cloudflare Workers with Ducklings, a minimal async DuckDB Wasm build.

Using Iceberg Catalogs in the Browser with DuckDB-Wasm
Query Apache Iceberg catalogs like S3 Tables and Cloudflare R2 Data Catalog from the browser with DuckDB-Wasm and the Iceberg extension, no backend needed.

TypeScript scripts as DuckDB Table Functions
Query REST APIs, GraphQL endpoints and web pages with SQL by using Bun TypeScript scripts as DuckDB table functions via the shellfs and arrow extensions.

Using Amazon SageMaker Lakehouse with DuckDB
Query S3 Tables through Amazon SageMaker Lakehouse and the AWS Glue Iceberg REST catalog with DuckDB, including IAM role and Lake Formation setup.

Welcome to the age of $10/month Lakehouses
Compare Iceberg, Delta Lake and DuckLake, then deploy a serverless DuckLake lakehouse for about $10 a month on Cloudflare Containers, R2 and Neon Postgres.

Using DuckDB databases as lightweight Data Lake access layer
Share Parquet data in S3 or R2 via a small DuckDB database file of views: a lightweight, low-cost data lake access layer that is queryable with plain SQL.

Handling GTFS data with DuckDB
Load GTFS public transit schedule feeds into a DuckDB database, query stops, routes and trips with SQL, and export the data to Parquet.

Querying IP addresses and CIDR ranges with DuckDB
Look up whether an IP address belongs to a CIDR range in DuckDB with three reusable SQL macros that compute network and broadcast addresses.

Chat with a Duck
Query your data in plain English: use the DuckDB-NSQL LLM via Ollama to generate DuckDB SQL from natural language prompts in the browser-based SQL Workbench.

Using DuckDB-WASM for in-browser Data Engineering
Analyze, transform and visualize local and remote data in the browser with DuckDB-WASM and SQL Workbench, including a full example data engineering pipeline.

Gathering and analyzing public cloud provider IP address data with DuckDB & Observable
Unify, clean and analyze the published IP address ranges of AWS, Azure, Google Cloud, Cloudflare, Oracle, DigitalOcean and Fastly with DuckDB and Observable.

Casual data engineering, or: A poor man's Data Lake in the cloud - Part I
Build a cost-effective, near-realtime serverless data lake on AWS with CloudFront, Kinesis Firehose, Lambda, S3, Glue and DuckDB.

Using DuckDB to repartition parquet data in S3
Repartition Parquet files in S3 data lakes with DuckDB in AWS Lambda: a serverless alternative to Athena, which is limited to 100 partitions per query.

Using DuckDB in AWS Lambda
Run DuckDB serverless in AWS Lambda with a custom Lambda layer: how it is built, how to configure and deploy it, and example queries against S3 data.