In the lab, once we have the input data with the path above as at the beginning of the lab, we will configure AWS Glue and Crawler to run on a schedule once a day. Crawler will scan the path containing the input Parquet file, save it on S3, and then create a database along with accompanying tables. When there is a new version of the report, the data table will be automatically updated.
Amazon Athena helps us access and view the contents of the parquet file via SQL commands. Amazon Athena is a serverless solution that supports executing SQL query commands on large amounts of data. Athena is only charged for the data scanned, unlike traditional database solutions.
The detailed configuration steps so that Amazon Athena can access data files via AWS Glue are as follows:
Access AWS Management Console

In the AWS Glue interface

Cost_MasterCrawler. Then, select Next







costmaster. Select Create database









