Blog Posts
Insights and articles about DuckDB, data engineering, serverless and open source.
Filter by tag
Filter by tag:

Query S3 Tables with DuckDB
Step-by-step guide to set up an Amazon S3 Table bucket and query its Apache Iceberg tables with DuckDB, including IAM permissions, secrets and catalog attach.

Querying IP addresses and CIDR ranges with DuckDB
Look up whether an IP address belongs to a CIDR range in DuckDB with three reusable SQL macros that compute network and broadcast addresses.

Chat with a Duck
Query your data in plain English: use the DuckDB-NSQL LLM via Ollama to generate DuckDB SQL from natural language prompts in the browser-based SQL Workbench.

Using DuckDB-WASM for in-browser Data Engineering
Analyze, transform and visualize local and remote data in the browser with DuckDB-WASM and SQL Workbench, including a full example data engineering pipeline.

Retrieving Lambda@Edge CloudWatch Logs
Lambda@Edge writes CloudWatch logs to the region closest to the viewer. Learn how to find them and aggregate logs from all regions with Kinesis.

List of free AWS Knowledge Badges
The complete list of free AWS Knowledge badges you can earn on AWS Skill Builder, with direct links to each badge readiness learning path.

Serverless Maps for fun and profit
Host fast, low-cost web maps on AWS without a tile server: generate PMTiles from OpenStreetMap data and serve them from S3 and CloudFront with Lambda.

Gathering and analyzing public cloud provider IP address data with DuckDB & Observable
Unify, clean and analyze the published IP address ranges of AWS, Azure, Google Cloud, Cloudflare, Oracle, DigitalOcean and Fastly with DuckDB and Observable.

Casual data engineering, or: A poor man's Data Lake in the cloud - Part I
Build a cost-effective, near-realtime serverless data lake on AWS with CloudFront, Kinesis Firehose, Lambda, S3, Glue and DuckDB.

Using DuckDB to repartition parquet data in S3
Repartition Parquet files in S3 data lakes with DuckDB in AWS Lambda: a serverless alternative to Athena, which is limited to 100 partitions per query.