Explore / Careers & Certification
AWS Certified Data Engineer - Associate (DEA-C01) Exam Prep

Build every skill in the current DEA-C01 exam guide from first principles, service by service, with short SQL and configuration examples and original scenario questions whose distractors are explained. Independent of AWS and no pass is promised, but you will finish with a full 65-question timed mock
Expert · 60 levels · 2 free · Created Oct 2026 · Professionally curated by levelupwith.com
What's inside
- Level 1: Exam Briefing: Format, Domains and Question CraftFree
[All domains – exam orientation] Sets out how DEA-C01 works (65 questions of which 50 are scored and 15 unscored, 130 minutes, scaled score 100–1,000 with a minimum passing score of 720, compensatory scoring), maps the four weighted content domains and their 17 tasks, and states plainly that this course is independent of AWS, uses only original practice questions and promises no pass. - Level 2: Kinesis Data Streams from First PrinciplesFree
[Data Ingestion and Transformation] Task 1.1 (Skill 1.1.1): explains what a streaming source is and how Amazon Kinesis Data Streams stores ordered records in shards, covering partition keys, sequence numbers, retention, on-demand versus provisioned capacity and the per-shard throughput trade-offs. - Level 3: Stream Consumers, Fan-Out, Lambda and Firehose
[Data Ingestion and Transformation] Task 1.1 (Skills 1.1.7, 1.1.10): teaches how streaming data is distributed to many consumers and gathered from many producers, covering shared versus enhanced fan-out, calling a Lambda function from Kinesis through an event source mapping, and delivery with Amazon Data Firehose (listed in the exam guide as Amazon Kinesis Data Firehose). - Level 4: Amazon MSK and the Other Streaming Sources
[Data Ingestion and Transformation] Task 1.1 (Skill 1.1.1): introduces Apache Kafka concepts and Amazon Managed Streaming for Apache Kafka (Amazon MSK), then the remaining streaming sources the guide names: Amazon DynamoDB Streams, AWS DMS change data capture, AWS Glue streaming jobs and Amazon Redshift streaming ingestion. - Level 5: Replayability, State and Rate Limits
[Data Ingestion and Transformation] Task 1.1 (Skills 1.1.9, 1.1.11, 1.1.12): covers the reliability properties of ingestion pipelines: describing replayability, defining stateful and stateless data transactions, and implementing throttling and overcoming rate limits in DynamoDB, Amazon RDS and Kinesis. - Level 6: Batch Ingestion Sources and Configuration
[Data Ingestion and Transformation] Task 1.1 (Skills 1.1.2, 1.1.3): teaches reading from batch sources (Amazon S3, AWS Glue, Amazon EMR, AWS DMS, Amazon Redshift, AWS Lambda and Amazon AppFlow) and implementing appropriate configuration options for batch ingestion such as incremental loads, file sizing and parallelism. - Level 7: Moving Data In: Transfer and Migration Services
[Data Ingestion and Transformation] Task 1.1 (Skill 1.1.2, in-scope Migration and Transfer services): contrasts the services that move existing data sets and workloads into AWS: AWS DataSync, AWS Snow Family, AWS Transfer Family, AWS Application Discovery Service and AWS Application Migration Service, plus third-party data sets from AWS Data Exchange. - Level 8: Schedulers, Event Triggers, Data APIs and Allowlists
[Data Ingestion and Transformation] Task 1.1 (Skills 1.1.4, 1.1.5, 1.1.6, 1.1.8): shows how ingestion is started and connected: setting up schedulers with Amazon EventBridge, Apache Airflow or time-based schedules for jobs and crawlers, setting up event triggers with Amazon S3 Event Notifications and EventBridge, consuming data APIs, and creating allowlists for IP addresses. - Level 9: The Three Vs and Choosing a Transformation Service
[Data Ingestion and Transformation] Task 1.2 (Skills 1.2.5, 1.2.9): defines volume, velocity and variety of data (structured, semi-structured and unstructured) and uses them to implement the right transformation service for a requirement: Amazon EMR, AWS Glue, Lambda or Amazon Redshift, with AWS Batch and Amazon EC2 as compute alternatives. - Level 10: AWS Glue ETL: Connections and Multi-Source Integration
[Data Ingestion and Transformation] Task 1.2 (Skills 1.2.2, 1.2.3): explains AWS Glue ETL from first principles (Spark jobs, DynamicFrames, workers and DPUs) and how to connect to different data sources through JDBC and ODBC and integrate data from multiple sources in one job. - Level 11: Amazon EMR, Apache Flink and Processing Cost
[Data Ingestion and Transformation] Task 1.2 (Skills 1.2.4, 1.2.5): teaches Amazon EMR's architecture and deployment options and Amazon Managed Service for Apache Flink for stateful stream processing, with the levers for optimizing costs while processing data. - Level 12: File and Table Formats: CSV to Parquet and Iceberg
[Data Ingestion and Transformation] Task 1.2 (Skill 1.2.6): explains row versus columnar storage and how to transform data between formats, for example from .csv to Apache Parquet, covering compression, splittability and partitioning, and introduces Apache Iceberg as an open table format layered on Parquet files. - Level 13: Containers for Data Processing: ECR, ECS and EKS
[Data Ingestion and Transformation] Task 1.2 (Skill 1.2.1): introduces containers from first principles and how to optimize container usage for performance needs on Amazon ECS and Amazon EKS, with images stored in Amazon ECR and AWS Fargate or EC2 as the capacity underneath. - Level 14: Creating Data APIs and Integrating LLMs
[Data Ingestion and Transformation] Task 1.2 (Skills 1.2.8, 1.2.10): shows how to create data APIs that make data available to other systems, with Amazon API Gateway in front of Lambda and the Amazon Redshift Data API, and how to integrate large language models for data processing through Amazon Bedrock. - Level 15: Troubleshooting Transformation Failures and Slowness
[Data Ingestion and Transformation] Task 1.2 (Skill 1.2.7): builds a diagnostic method for troubleshooting and debugging common transformation failures and performance issues in AWS Glue, Amazon EMR, Lambda and Amazon Redshift, including data skew, out-of-memory errors, the small-files problem and schema mismatches. - Level 16: Orchestration Services and Serverless Workflows
[Data Ingestion and Transformation] Task 1.3 (Skills 1.3.1, 1.3.3): contrasts the orchestration services used to build workflows for data ETL pipelines (AWS Step Functions, Amazon Managed Workflows for Apache Airflow (Amazon MWAA), AWS Glue workflows, EventBridge and Lambda) and how to implement and maintain serverless workflows. - Level 17: Resilient Pipelines and Alerts with SNS and SQS
[Data Ingestion and Transformation] Task 1.3 (Skills 1.3.2, 1.3.4): teaches how to build data pipelines for performance, availability, scalability, resiliency and fault tolerance, and how to use notification services (Amazon SNS and Amazon SQS) to send alerts and decouple stages. - Level 18: Languages, Distributed Computing and Efficient Code
[Data Ingestion and Transformation] Task 1.4 (Skills 1.4.1, 1.4.3, 1.4.10, 1.4.11): covers the language-agnostic programming concepts the exam expects: which languages and frameworks suit which data-engineering job (Python, SQL, Scala, R, Java, Bash, PowerShell), defining distributed computing, describing data structures and algorithms such as graph and tree data structures, and optimizing code to reduce runtime. - Level 19: Configuring Lambda: Concurrency, Performance, Storage
[Data Ingestion and Transformation] Task 1.4 (Skills 1.4.2, 1.4.7): goes deep on AWS Lambda for data pipelines: configuring functions to meet concurrency and performance needs, and using and mounting storage volumes from within Lambda functions. - Level 20: Infrastructure as Code: CloudFormation, CDK and SAM
[Data Ingestion and Transformation] Task 1.4 (Skills 1.4.5, 1.4.6, 1.4.8): teaches Infrastructure as Code for deploying data engineering solutions repeatably with AWS CloudFormation and AWS CDK, and using AWS SAM to package and deploy serverless data pipelines such as Lambda functions, Step Functions state machines and DynamoDB tables. - Level 21: CI/CD for Data Pipelines: CodePipeline, CodeBuild, CodeDeploy
[Data Ingestion and Transformation] Task 1.4 (Skills 1.4.4, 1.4.9): how continuous integration and continuous delivery implement, test and deploy data pipelines, using version control, AWS CodePipeline, AWS CodeBuild, AWS CodeDeploy and container images in Amazon ECR. - Level 22: Choosing Storage: S3, EBS, EFS and Streams as Stores
[Data Store Management] Task 2.1 (Skills 2.1.1, 2.1.2): the first-principles difference between object, block and file storage in Amazon S3, Amazon EBS and Amazon EFS, and when Kinesis Data Streams or Amazon MSK retention acts as a short-term store, judged on cost, performance and access pattern. - Level 23: Amazon Redshift as a Data Store: Provisioned and Serverless
[Data Store Management] Task 2.1 (Skills 2.1.1, 2.1.2): how Amazon Redshift stores and serves data through columnar, massively parallel processing, and how to configure provisioned RA3 clusters or Redshift Serverless, workload management and concurrency scaling for cost and performance. - Level 24: Amazon RDS, Aurora and Managing Locks
[Data Store Management] Task 2.1 (Skills 2.1.1, 2.1.2, 2.1.6): how Amazon RDS and Amazon Aurora serve transactional workloads with read replicas and Multi-AZ, and how to manage locks that block access to data in Amazon RDS and Amazon Redshift. - Level 25: DynamoDB and Purpose-Built NoSQL Stores
[Data Store Management] Task 2.1 (Skills 2.1.1 to 2.1.3): configuring Amazon DynamoDB capacity modes and access paths, and applying Amazon MemoryDB for fast key/value pair access, Amazon Keyspaces, Amazon DocumentDB and Amazon Neptune to the use cases they were built for. - Level 26: Vector Stores, Vectorization and Index Types (HNSW, IVF)
[Data Store Management] Tasks 2.1 and 2.4 (Skills 2.1.3, 2.1.8, 2.4.6): what embeddings and vectorization are, how an Amazon Bedrock knowledge base chunks, embeds and stores data, and how the HNSW and IVF vector index types trade recall, speed, memory and build time in Amazon Aurora PostgreSQL and Amazon OpenSearch Service. - Level 27: Redshift Spectrum, Federated Queries and Materialized Views
[Data Store Management] Task 2.1 (Skills 2.1.4, 2.1.5): reaching data where it lives with Amazon Redshift Spectrum, Amazon Redshift federated queries and Amazon Redshift materialized views, and integrating AWS Transfer Family as a migration tool that lands partner files in Amazon S3. - Level 28: Managing Open Table Formats: Apache Iceberg and S3 Tables
[Data Store Management] Task 2.1 (Skill 2.1.7): how Apache Iceberg adds snapshots, ACID transactions, hidden partitioning and time travel on top of Parquet files, and how Amazon S3 Tables, AWS Glue, Amazon EMR and Athena create and maintain Iceberg tables. - Level 29: Technical Catalogs: Glue Data Catalog and Hive Metastore
[Data Store Management] Task 2.2 (Skills 2.2.1, 2.2.2): building and referencing a technical data catalog with the AWS Glue Data Catalog or an Apache Hive metastore, so that Athena, Amazon EMR, Redshift Spectrum and AWS Glue jobs consume data from its source through one shared definition. - Level 30: Glue Crawlers, Connections and Partition Synchronization
[Data Store Management] Task 2.2 (Skills 2.2.3 to 2.2.5): discovering schemas with AWS Glue crawlers and classifiers, creating AWS Glue connections to new sources and targets, and keeping partitions synchronized with the data catalog. - Level 31: Business Data Catalogs with Amazon SageMaker Catalog
[Data Store Management] Task 2.2 (Skill 2.2.6): creating and managing a business data catalog in Amazon SageMaker Catalog within Amazon SageMaker Unified Studio, where assets carry business metadata and glossary terms and consumers find and subscribe to published data. - Level 32: Loading and Unloading: Redshift COPY and UNLOAD
[Data Store Management] Task 2.3 (Skill 2.3.1): performing load and unload operations that move data between Amazon S3 and Amazon Redshift with the COPY and UNLOAD commands, tuned for parallelism, file format and cost. - Level 33: S3 Lifecycle Policies, Storage Classes and S3 Glacier
[Data Store Management] Task 2.3 (Skills 2.3.2, 2.3.3): managing S3 Lifecycle policies that change the storage tier of S3 data and expire data when it reaches a specific age, across the S3 storage classes including S3 Intelligent-Tiering and the S3 Glacier storage classes. - Level 34: Versioning, DynamoDB TTL, Deletion and Resiliency
[Data Store Management] Task 2.3 (Skills 2.3.4 to 2.3.6): managing S3 Versioning and DynamoDB TTL, deleting data to meet business and legal requirements, and protecting data with appropriate resiliency and availability by using AWS Backup, replication, snapshots and S3 Object Lock. - Level 35: Schema Design and Optimization: Redshift, DynamoDB, Lakes
[Data Store Management] Task 2.4 (Skills 2.4.1, 2.4.5): designing schemas for Amazon Redshift, DynamoDB and Lake Formation data lakes, and applying best practices for indexing, partitioning strategies, compression and other data optimization techniques. - Level 36: Schema Evolution, Schema Conversion and Data Lineage
[Data Store Management] Task 2.4 (Skills 2.4.2 to 2.4.4): addressing changes to the characteristics of data, performing schema conversion with AWS SCT and AWS DMS Schema Conversion, and establishing data lineage with Amazon SageMaker ML Lineage Tracking and Amazon SageMaker Catalog. - Level 37: Operating and Troubleshooting MWAA and Step Functions
[Data Operations and Support] Task 3.1 (Skills 3.1.1, 3.1.2): running orchestrated data pipelines in production on Amazon MWAA and AWS Step Functions, and troubleshooting Amazon managed workflows when DAGs, tasks or state machine executions fail or stall. - Level 38: Calling AWS SDKs and Consuming Data APIs
[Data Operations and Support] Task 3.1 (Skills 3.1.3, 3.1.5): calling SDKs and the AWS CLI to access Amazon features from code, and consuming and maintaining data APIs such as the Amazon Redshift Data API and endpoints published through Amazon API Gateway. - Level 39: Event-Driven Automation: Lambda, EventBridge, Service Features
[Data Operations and Support] Task 3.1 (Skills 3.1.4, 3.1.8, 3.1.9): automating data processing with AWS Lambda, managing events and schedulers in Amazon EventBridge, and using the built-in processing features of Amazon EMR, Amazon Redshift and AWS Glue such as EMR steps, scheduled queries, stored procedures, Glue triggers and job bookmarks. - Level 40: Preparing and Querying Data: DataBrew, Unified Studio, Athena
[Data Operations and Support] Task 3.1 (Skills 3.1.6, 3.1.7): preparing data for transformation with AWS Glue DataBrew and Amazon SageMaker Unified Studio, and querying data in Amazon S3 with Amazon Athena using workgroups, query result locations and CTAS. - Level 41: Visualising, Verifying and Cleaning Data
[Data Operations and Support] Serves Task 3.2 (Skills 3.2.1 and 3.2.2): how to visualise data with Amazon Quick (called Amazon QuickSight in the guide's skill text) and AWS Glue DataBrew, and how to verify and clean data with Lambda, Athena, Jupyter Notebooks and Amazon SageMaker Data Wrangler. - Level 42: Analytical SQL and Views in Redshift and Athena
[Data Operations and Support] Serves Task 3.2 (Skills 3.2.3 and 3.2.6): writing SQL in Amazon Redshift and Athena to query data and create views, and defining data aggregation, rolling average, grouping and pivoting with short worked examples. - Level 43: Athena Spark Notebooks; Provisioned vs Serverless
[Data Operations and Support] Serves Task 3.2 (Skills 3.2.4 and 3.2.5): exploring data with Athena notebooks that use Apache Spark, and describing the trade-offs between provisioned services and serverless services across the analytics stack. - Level 44: Monitoring, Logging and Alerts for Pipelines
[Data Operations and Support] Serves Task 3.3 (Skills 3.3.2, 3.3.3 and 3.3.7): deploying logging and monitoring for traceability with Amazon CloudWatch metrics and alarms, configuring and automating CloudWatch Logs for application data, sending alerts through Amazon SNS and EventBridge, and building dashboards in Amazon Managed Grafana. - Level 45: Tracking API Calls and Analysing Pipeline Logs
[Data Operations and Support] Serves Task 3.3 (Skills 3.3.1, 3.3.5 and 3.3.8): using AWS CloudTrail to track API calls, extracting logs for audits, and analysing logs with Athena, Amazon EMR, Amazon OpenSearch Service and CloudWatch Logs Insights, including big data application logs. - Level 46: Troubleshooting Glue and EMR Pipeline Performance
[Data Operations and Support] Serves Task 3.3 (Skills 3.3.4 and 3.3.6): troubleshooting performance issues and maintaining pipelines in AWS Glue and Amazon EMR by reading job metrics, the Spark UI and logs, and applying the fix that matches the symptom. - Level 47: Cost Visibility and Operational Review
[Data Operations and Support] Serves Task 3.3 and the guide's stated aim of optimising cost and performance, using in-scope services: AWS Budgets and AWS Cost Explorer for cost tracking, AWS Systems Manager for operating the compute behind pipelines, the AWS Well-Architected Tool for reviewing a workload, and Amazon Q as an assistant for troubleshooting. - Level 48: Data Quality Rules and Consistency Checks
[Data Operations and Support] Serves Task 3.4 (Skills 3.4.1, 3.4.2 and 3.4.3): running data quality checks while processing data (for example, checking for empty fields), defining data quality rules in AWS Glue DataBrew and AWS Glue Data Quality, and investigating data consistency. - Level 49: Data Sampling Techniques and Handling Data Skew
[Data Operations and Support] Serves Task 3.4 (Skills 3.4.4 and 3.4.5): describing data sampling techniques and implementing data skew mechanisms in distributed processing and in Amazon Redshift. - Level 50: IAM Roles, Managed Services and Unified Studio Domains
[Data Security and Governance] Serves Task 4.1 (Skills 4.1.2, 4.1.4, 4.1.6 and 4.1.7): creating and updating IAM groups and roles, setting up IAM roles for access by AWS Lambda, Amazon API Gateway, the AWS CLI and AWS CloudFormation, describing the differences between managed and unmanaged services, and using domains, domain units and projects in SageMaker Unified Studio. - Level 51: Security Groups, VPC Endpoints and Endpoint Policies
[Data Security and Governance] Serves Task 4.1 (Skills 4.1.1 and 4.1.5): updating VPC security groups and applying IAM policies to endpoints and services such as S3 Access Points and AWS PrivateLink, with AWS WAF and AWS Shield named as the protection for public-facing data APIs. - Level 52: Storing and Rotating Credentials
[Data Security and Governance] Serves Task 4.1 (Skill 4.1.3) and Task 4.2 (Skill 4.2.2): creating and rotating credentials with AWS Secrets Manager, and storing application and database credentials in Secrets Manager or AWS Systems Manager Parameter Store. - Level 53: Custom IAM Policies and Least Privilege
[Data Security and Governance] Serves Task 4.2 (Skills 4.2.1, 4.2.5 and 4.2.6): creating custom IAM policies when a managed policy does not meet the need, constructing policies that meet the principle of least privilege, and applying role-based, tag-based and attribute-based authorization methods. - Level 54: Lake Formation Permissions and Redshift Access Control
[Data Security and Governance] Serves Task 4.2 (Skills 4.2.3 and 4.2.4): managing permissions through AWS Lake Formation for Amazon Redshift, Amazon EMR, Amazon Athena and Amazon S3, and giving database users, groups and roles access and authority in a database such as Amazon Redshift. - Level 55: Encryption with AWS KMS, in Transit and Across Accounts
[Data Security and Governance] Serves Task 4.3 (Skills 4.3.2, 4.3.3 and 4.3.4): using encryption keys in AWS KMS to encrypt and decrypt data, configuring encryption across AWS account boundaries, and enabling encryption in transit or before transit. - Level 56: Data Masking and Anonymization
[Data Security and Governance] Serves Task 4.3 (Skill 4.3.1): applying data masking and anonymization according to compliance laws or company policies, using Amazon Redshift dynamic data masking, AWS Glue and Glue DataBrew transforms, and hashing or tokenization inside pipelines. - Level 57: Preparing Logs for Audit with CloudTrail Lake
[Data Security and Governance] Serves Task 4.4 (all skills): using AWS CloudTrail to track API calls for audit, storing application logs in Amazon CloudWatch Logs, using AWS CloudTrail Lake for centralised logging queries, analysing logs with Athena, CloudWatch Logs Insights and Amazon OpenSearch Service, and integrating services such as Amazon EMR for large volumes of log data. - Level 58: PII Identification, Data Sovereignty and AWS Config
[Data Security and Governance] Serves Task 4.5 (Skills 4.5.2, 4.5.3, 4.5.4 and 4.5.5): implementing PII identification with Amazon Macie and Lake Formation, preventing backups or replications of data to disallowed AWS Regions, maintaining data sovereignty, and viewing configuration changes in an account with AWS Config. - Level 59: Data Sharing, SageMaker Catalog and Governance Patterns
[Data Security and Governance] Serves Task 4.5 (Skills 4.5.1, 4.5.6 and 4.5.7): granting permissions for data sharing such as data sharing for Amazon Redshift, managing data access through Amazon SageMaker Catalog projects, and describing data governance frameworks and data sharing patterns, with AWS Data Exchange as the route for third-party data. - Level 60: Timed Mock Exam: 65 Questions in 130 Minutes Timed mock
[All domains] A timed mock of 65 original multiple-choice and multiple-response scenario questions in 130 minutes, drawn from Data Ingestion and Transformation (34%), Data Store Management (26%), Data Operations and Support (22%) and Data Security and Governance (18%) in proportion to their weightings, with a 72% pass mark; it is independent of AWS and does not guarantee a pass on the real exam.
Access
The first 2 levels are free with a free account. Every level, the podcast edition and the AI tutor come with All Access at £4.99/month or any Creator plan — see pricing.