AWS Glue is Amazon's serverless data-integration service for discovering, preparing, and combining data: Spark and Python ETL jobs, a central Data Catalog, crawlers, and the visual Glue Studio. It is the default ETL layer for AWS-centric data platforms feeding S3 data lakes, Redshift, and Athena. Billing is per DPU-hour with no idle cost, and the Data Catalog's first million objects and requests are free each month. It suits teams standardizing on AWS that want managed Spark-based pipelines without cluster operations.
awsserverlesssparkdata-catalogdata-lake
Overview
- Founded
- 2017
- Pricing tier
- Low
- Startup-friendly
- No
- Enterprise-ready
- Yes
ETL / data pipeline
- Free tier
- Yes: Data Catalog: first 1M objects + 1M requests/month free; jobs bill per DPU-hour.
- Open-source
- No
- Self-hostable
- No
- ELT
- Yes
- Change data capture
- No
- Prebuilt connectors
- Yes
