AWS Big Data Blog

Category: Analytics

Amazon MSK Service 101: How many partitions does an Amazon MSK topic need?

Amazon MSK Service 101: How many partitions does an Amazon MSK topic need?

How many partitions does your Amazon MSK topic need? Choosing the right partition count affects throughput, scalability, and operational complexity. This post provides practical guidance for sizing partitions, covering per-partition throughput, consumer parallelism, partition keys, and Amazon MSK partition-per-broker guidelines.

AWS and DuckLabs: Building the future of analytics together

Today we are announcing that Amazon has signed a definitive agreement to acquire DuckLabs, the Amsterdam-based company behind the open-source analytical database DuckDB. We expect the transaction to close shortly, subject to customary closing conditions. Hannes Mühleisen and Mark Raasveldt, who created DuckDB and co-founded DuckLabs, will continue leading the team and the open-source project’s technical direction as part of AWS. The DuckDB open-source project will also continue to be driven by the DuckLabs team, remain open source under the independent Foundation (the non-profit that oversees DuckDB), and available under the MIT license as it does today.

Enable cross-cloud analytics with Amazon S3 Tables and Google BigQuery, Part 2: access control with Lake Formation

Enable cross-cloud analytics with Amazon S3 Tables and Google BigQuery, Part 2: access control with Lake Formation

In Part 2 of this series, connect Google BigQuery to Amazon S3 Tables using AWS Lake Formation credential vending. Lake Formation manages fine-grained permissions and issues short-lived, scoped credentials to external engines, so you can centrally govern which teams and query engines read your Iceberg tables on AWS without managing IAM policies for every consumer.

GPU-accelerated Apache Spark with Amazon EMR and NVIDIA RTX PRO 4500 on Amazon EC2 G7 instances runs up to 3.7x faster

GPU-accelerated Apache Spark with Amazon EMR and NVIDIA RTX PRO 4500 on Amazon EC2 G7 instances runs up to 3.7x faster

Amazon EMR on EKS now runs Apache Spark up to 3.7x faster on Amazon EC2 G7 instances with NVIDIA RTX PRO 4500 Blackwell GPUs than on comparable CPU instances, with no changes to existing Spark code. See the TPC-DS benchmark results, the cost comparison, and how to get started.

Introducing AWS Glue 6.0 for faster and more cost-effective data integration

Introducing AWS Glue 6.0 for faster and more cost-effective data integration

AWS Glue 6.0 is now available, lowering AWS Glue pricing by 30%, adding an AWS optimized build of Apache Spark 4.1, and introducing Apache Iceberg V3 capabilities suitable for enterprise adoption. This post covers the key capabilities and performance benefits, with code examples to help you get started.

Track SageMaker Unified Studio project costs with custom tags and AWS CUR

Track SageMaker Unified Studio project costs with custom tags and AWS CUR

Learn how to track Amazon SageMaker Unified Studio project costs by custom tags. This serverless solution enriches AWS Cost and Usage Report (CUR) data with custom project tags and visualizes cost by CostCenter, Team, or Environment in an Amazon Quick Sight dashboard.

Secure SageMaker Unified Studio access with SAML and conditional policies

Secure SageMaker Unified Studio access with SAML and conditional policies

Learn how to secure Amazon SageMaker Unified Studio by integrating it with an external SAML identity provider such as Okta. This post shows you how to apply conditional access policies that enforce device compliance, IP-based restrictions, and multi-factor authentication for your data and AI workloads.

Querying raw log data with SQL and PPL with the optimized engine in Amazon OpenSearch Service

Querying raw log data using SQL and PPL with the optimized engine in Amazon OpenSearch Service

Learn how to run fast analytical queries directly against raw log and trace data in Amazon OpenSearch Service using PPL and SQL. Follow a single incident investigation, one query at a time, and see how the new optimized engine answers each question directly from raw spans.

Amazon MSK simplifies configuring custom domain names

Amazon MSK simplifies configuring custom domain names

With Amazon MSK, you can now configure custom domain names for provisioned clusters using a single configuration property that works identically on ZooKeeper and KRaft. Define the domain once and Amazon MSK applies it across every broker, so custom domain names keep working as the cluster scales.