AWS Big Data Blog Official Big Data Blog of Amazon Web Services
- Deliver real-time data to streaming tables for Apache Iceberg with Amazon Kinesis Data Streamsby Nikit Pednekar on August 31, 2026 at 11:03 pm
Amazon Kinesis Data Streams now supports streaming tables, a fully managed capability that continuously delivers your streaming data as queryable Apache Iceberg tables on Amazon S3 Tables. Streaming tables cut delivery costs by up to 25% and query costs by up to 30% through inline compaction, with no infrastructure to operate.
- Measuring and improving search quality with Amazon OpenSearch Serviceby Aruna Govindaraju on August 31, 2026 at 9:13 pm
Most teams struggle to answer a deceptively simple question: is my search returning relevant results? This post shows how to capture User Behavior Insights (UBI) data on Amazon OpenSearch Service and use Search Relevance Workbench (SRW) to turn those signals into relevance judgments and evaluate search quality.
- Integrate Amazon Redshift and IAM Identity Center with enhanced VPC routingby Maneesh Sharma on August 31, 2026 at 5:22 pm
Amazon Redshift now supports AWS IAM Identity Center authentication on clusters and workgroups that use enhanced VPC routing. Create two interface VPC endpoints to give your users single sign-on with their corporate credentials while keeping all authentication traffic on the AWS private network.
- Build a real-time event pipeline with Spark Real-Time Mode on AWS Glue 6.0by Shoukat Ghouse on August 31, 2026 at 4:22 pm
With AWS Glue 6.0, you can build real-time, near-real-time, and batch data pipelines on a single platform. Using a financial market-risk example, learn how to flag high-risk trades with sub-second latency using Spark Real-Time Mode, store heterogeneous pricing vectors with Apache Iceberg v3 Variant columns, and run batch analytics with Arrow-native UDFs.
- Razor Group’s journey to a modern data lakehouse on AWSby Yaswanth Kothainti on August 28, 2026 at 4:25 pm
Razor Group, one of Europe’s leading ecommerce aggregators managing 250+ brands, migrated from always-on Amazon Redshift clusters to an open lakehouse on Apache Iceberg, Amazon S3 Tables, and Apache Spark. Learn the architectural decisions, the five-phase migration, and the results: 65% faster P95 queries and a 63% infrastructure cost reduction.
- How Picnic configured multiple OAuth providers for Amazon MQby Oscar Mapfumo Sibanda on August 27, 2026 at 4:07 pm
Picnic runs RabbitMQ as the messaging backbone for hundreds of microservices on Amazon MQ for RabbitMQ. This post shows how to configure one broker to trust multiple OAuth 2.0 identity providers, Keycloak for operators and AWS IAM for services, so you can eliminate static credentials while maintaining separate identity paths for people and workloads.
- Build with geospatial and variant types in Iceberg v3 on AWS Glue 6.0by Shoukat Ghouse on August 27, 2026 at 4:00 pm
AWS Glue 6.0 with Apache Spark 4.1 adds support for Apache Iceberg v3: native geospatial types, nanosecond-precision timestamps, the VARIANT type, and DEFAULT column values. This post builds a connected vehicle fleet telemetry pipeline that uses all four in a single Iceberg v3 table, from ingestion through spatial, nanosecond, and variant queries.
- Migrate an OAuth 2.0 authenticated Apache Kafka cluster to Amazon MSK with MSK Replicatorby Subham Rakshit on August 26, 2026 at 7:11 pm
MSK Replicator now supports OAuth 2.0 (SASL/OAUTHBEARER) authentication to external Apache Kafka clusters. This post walks through the three supported grant types, how to configure Replicator for each, the network and TLS prerequisites that are commonly missed, and how to handle identity providers behind an additional federation layer.
- Announcing in-place ZooKeeper-to-KRaft cluster upgrades for Amazon MSKby Austin Groeneveld on August 26, 2026 at 6:59 pm
Amazon MSK now supports in-place upgrades from ZooKeeper to KRaft metadata mode. You can modernize your existing cluster’s metadata management through the familiar version upgrade workflow, with no new cluster to provision and no data migration. This post covers the prerequisites and the step-by-step upgrade process.
- Amazon MSK Service 101: How many partitions does an Amazon MSK topic need?by Yashika Jain on August 26, 2026 at 3:56 pm
How many partitions does your Amazon MSK topic need? Choosing the right partition count affects throughput, scalability, and operational complexity. This post provides practical guidance for sizing partitions, covering per-partition throughput, consumer parallelism, partition keys, and Amazon MSK partition-per-broker guidelines.
- AWS and DuckLabs: Building the future of analytics togetherby Mai-Lan Tomsen Bukovec on August 26, 2026 at 1:28 pm
Today we are announcing that Amazon has signed a definitive agreement to acquire DuckLabs, the Amsterdam-based company behind the open-source analytical database DuckDB. We expect the transaction to close shortly, subject to customary closing conditions. Hannes Mühleisen and Mark Raasveldt, who created DuckDB and co-founded DuckLabs, will continue leading the team and the open-source project’s technical direction as part of AWS. The DuckDB open-source project will also continue to be driven by the DuckLabs team, remain open source under the independent Foundation (the non-profit that oversees DuckDB), and available under the MIT license as it does today.
- PythonOperator and BashOperator Now Available on Amazon Managed Workflows for Apache Airflow (Amazon MWAA) Serverlessby Pradeep Kumar Nalluri on August 25, 2026 at 7:28 pm
You can now use PythonOperator and BashOperator to run custom Python functions and shell scripts directly in the Amazon MWAA Serverless runtime, without provisioning additional infrastructure. This post walks through building a serverless pipeline that converts CSV files to JSON using a PythonOperator and verifies the output with a BashOperator.
- Enable cross-cloud analytics with Amazon S3 Tables and Google BigQuery, Part 1: IAM-based access controlby Lakshmi Nair on August 25, 2026 at 4:38 pm
Your Google BigQuery users need to query data that lives in Amazon S3 Tables on AWS without copying it across clouds. This post shows how to connect BigQuery to Amazon S3 Tables through the AWS Glue Iceberg REST Catalog using IAM-based access control, so you keep one governed dataset and query it live from BigQuery.
- Enable cross-cloud analytics with Amazon S3 Tables and Google BigQuery, Part 2: access control with Lake Formationby Lakshmi Nair on August 25, 2026 at 4:38 pm
In Part 2 of this series, connect Google BigQuery to Amazon S3 Tables using AWS Lake Formation credential vending. Lake Formation manages fine-grained permissions and issues short-lived, scoped credentials to external engines, so you can centrally govern which teams and query engines read your Iceberg tables on AWS without managing IAM policies for every consumer.
- GPU-accelerated Apache Spark with Amazon EMR and NVIDIA RTX PRO 4500 on Amazon EC2 G7 instances runs up to 3.7x fasterby McCall Peltier on August 25, 2026 at 4:17 pm
Amazon EMR on EKS now runs Apache Spark up to 3.7x faster on Amazon EC2 G7 instances with NVIDIA RTX PRO 4500 Blackwell GPUs than on comparable CPU instances, with no changes to existing Spark code. See the TPC-DS benchmark results, the cost comparison, and how to get started.
- Introducing AWS Glue 6.0 for faster and more cost-effective data integrationby Aarthi Srinivasan on August 24, 2026 at 7:06 pm
AWS Glue 6.0 is now available, lowering AWS Glue pricing by 30%, adding an AWS optimized build of Apache Spark 4.1, and introducing Apache Iceberg V3 capabilities suitable for enterprise adoption. This post covers the key capabilities and performance benefits, with code examples to help you get started.
- Upgrade AWS Glue jobs to Glue 6.0 with AI-powered Spark upgradesby Prasad Nadig on August 24, 2026 at 7:06 pm
Walk through upgrading a PySpark ETL job from AWS Glue 5.1 to AWS Glue 6.0 using the generative AI upgrades for Apache Spark. The upgrade analysis automatically detects incompatibilities, applies fixes, and validates results with data quality checks.
- Long-term system tables retention in Amazon Redshift with Amazon S3 Tablesby Nidhi Nayak on August 20, 2026 at 7:20 pm
Amazon Redshift system table integration with Amazon S3 Tables automatically delivers your system table logs to Amazon S3 Tables in Apache Iceberg format. You can retain this data well beyond the 7-day limit for compliance, auditing, and cross-warehouse observability, without custom ETL pipelines or cluster resource consumption.
- Track SageMaker Unified Studio project costs with custom tags and AWS CURby Nisha Gambhir on August 20, 2026 at 4:19 pm
Learn how to track Amazon SageMaker Unified Studio project costs by custom tags. This serverless solution enriches AWS Cost and Usage Report (CUR) data with custom project tags and visualizes cost by CostCenter, Team, or Environment in an Amazon Quick Sight dashboard.
- Secure SageMaker Unified Studio access with SAML and conditional policiesby Manos Samatas on August 19, 2026 at 8:35 pm
Learn how to secure Amazon SageMaker Unified Studio by integrating it with an external SAML identity provider such as Okta. This post shows you how to apply conditional access policies that enforce device compliance, IP-based restrictions, and multi-factor authentication for your data and AI workloads.
























