Professional-Data-Engineer Braindumps PDF, Google Professional-Data-Engineer Exam Cram [Q186-Q208]

Share

Professional-Data-Engineer Braindumps PDF, Google Professional-Data-Engineer Exam Cram

New 2024 Professional-Data-Engineer Sample Questions Reliable Professional-Data-Engineer Test Engine


Google Professional-Data-Engineer Certification Exam is a highly prestigious certification program offered by Google for individuals who want to establish themselves as professional data engineers. Google Certified Professional Data Engineer Exam certification validates the skills and knowledge required to design, build, operationalize, secure, and monitor data processing systems. It is designed for individuals who have experience working with data processing systems, data warehousing, and data analysis technologies.

 

NEW QUESTION # 186
Your neural network model is taking days to train. You want to increase the training speed. What can you do?

  • A. Subsample your training dataset.
  • B. Subsample your test dataset.
  • C. Increase the number of layers in your neural network.
  • D. Increase the number of input features to your model.

Answer: C

Explanation:
Explanation/Reference: https://towardsdatascience.com/how-to-increase-the-accuracy-of-a-neural-network-9f5d1c6f407d


NEW QUESTION # 187
You have designed an Apache Beam processing pipeline that reads from a Pub/Sub topic. The topic has a message retention duration of one day, and writes to a Cloud Storage bucket. You need to select a bucket location and processing strategy to prevent data loss in case of a regional outage with an RPO of 15 minutes.
What should you do?

  • A. 1 Use a regional Cloud Storage bucket
    2 Monitor Dataflow metrics with Cloud Monitoring to determine when an outage occurs
    3 Seek the subscription back in time by one day to recover the acknowledged messages
    4 Start the Dataflow job in a secondary region and write in a bucket in the same region
  • B. 1. Use a dual-region Cloud Storage bucket.
    2. Monitor Dataflow metrics with Cloud Monitoring to determine when an outage occurs
    3 Seek the subscription back in time by 15 minutes to recover the acknowledged messages
    4 Start the Dataflow job in a secondary region
  • C. 1 Use a multi-regional Cloud Storage bucket
    2 Monitor Dataflow metrics with Cloud Monitoring to determine when an outage occurs
    3 Seek the subscription back in time by 60 minutes to recover the acknowledged messages
    4 Start the Dataflow job in a secondary region
  • D. 1. Use a dual-region Cloud Storage bucket with turbo replication enabled
    2 Monitor Dataflow metrics with Cloud Monitoring to determine when an outage occurs
    3 Seek the subscription back in time by 60 minutes to recover the acknowledged messages
    4 Start the Dataflow job in a secondary region.

Answer: B

Explanation:
A dual-region Cloud Storage bucket is a type of bucket that stores data redundantly across two regions within the same continent. This provides higher availability and durability than a regional bucket, which stores data in a single region. A dual-region bucket also provides lower latency and higher throughput than a multi-regional bucket, which stores data across multiple regions within a continent or across continents. A dual-region bucket with turbo replication enabled is a premium option that offers even faster replication across regions, but it is more expensive and not necessary for this scenario.
By using a dual-region Cloud Storage bucket, you can ensure that your data is protected from regional outages, and that you can access it from either region with low latency and high performance. You can also monitor the Dataflow metrics with Cloud Monitoring to determine when an outage occurs, and seek the subscription back in time by 15 minutes to recover the acknowledged messages. Seeking a subscription allows you to replay the messages from a Pub/Sub topic that were published within the message retention duration, which is one day in this case. By seeking the subscription back in time by 15 minutes, you can meet the RPO of 15 minutes, which means the maximum amount of data loss that is acceptable for your business. You can then start the Dataflow job in a secondary region and write to the same dual-region bucket, which will resume the processing of the messages and prevent data loss.
Option A is not a good solution, as using a regional Cloud Storage bucket does not provide any redundancy or protection from regional outages. If the region where the bucket is located experiences an outage, you will not be able to access your data or write new data to the bucket. Seeking the subscription back in time by one day is also unnecessary and inefficient, as it will replay all the messages from the past day, even though you only need to recover the messages from the past 15 minutes.
Option B is not a good solution, as using a multi-regional Cloud Storage bucket does not provide the best performance or cost-efficiency for this scenario. A multi-regional bucket stores data across multiple regions within a continent or across continents, which provides higher availability and durability than a dual-region bucket, but also higher latency and lower throughput. A multi-regional bucket is more suitable for serving data to a global audience, not for processing data with Dataflow within a single continent. Seeking the subscription back in time by 60 minutes is also unnecessary and inefficient, as it will replay more messages than needed to meet the RPO of 15 minutes.
Option D is not a good solution, as using a dual-region Cloud Storage bucket with turbo replication enabled does not provide any additional benefit for this scenario, but only increases the cost. Turbo replication is a premium option that offers faster replication across regions, but it is not required to meet the RPO of 15 minutes. Seeking the subscription back in time by 60 minutes is also unnecessary and inefficient, as it will replay more messages than needed to meet the RPO of 15 minutes. References: Storage locations | Cloud Storage | Google Cloud, Dataflow metrics | Cloud Dataflow | Google Cloud, Seeking a subscription | Cloud Pub/Sub | Google Cloud, Recovery point objective (RPO) | Acronis.


NEW QUESTION # 188
Your company has recently grown rapidly and now ingesting data at a significantly higher rate than it was
previously. You manage the daily batch MapReduce analytics jobs in Apache Hadoop. However, the
recent increase in data has meant the batch jobs are falling behind. You were asked to recommend ways
the development team could increase the responsiveness of the analytics without increasing costs. What
should you recommend they do?

  • A. Decrease the size of the Hadoop cluster but also rewrite the job in Hive.
  • B. Rewrite the job in Pig.
  • C. Increase the size of the Hadoop cluster.
  • D. Rewrite the job in Apache Spark.

Answer: B


NEW QUESTION # 189
You work for a large bank that operates in locations throughout North America. You are setting up a data storage system that will handle bank account transactions. You require ACID compliance and the ability to access data with SQL. Which solution is appropriate?

  • A. Store transaction in Cloud Spanner. Use locking read-write transactions.
  • B. Store transaction data in Cloud Spanner. Enable stale reads to reduce latency.
  • C. Store transaction data in Cloud SQL. Use a federated query BigQuery for analysis.
  • D. Store transaction data in BigQuery. Disabled the query cache to ensure consistency.

Answer: D


NEW QUESTION # 190
You are implementing several batch jobs that must be executed on a schedule. These jobs have many interdependent steps that must be executed in a specific order. Portions of the jobs involve executing shell scripts, running Hadoop jobs, and running queries in BigQuery. The jobs are expected to run for many minutes up to several hours. If the steps fail, they must be retried a fixed number of times. Which service should you use to manage the execution of these jobs?

  • A. Cloud Composer
  • B. Cloud Functions
  • C. Cloud Scheduler
  • D. Cloud Dataflow

Answer: A


NEW QUESTION # 191
You receive data files in CSV format monthly from a third party. You need to cleanse this data, but every third month the schema of the files changes. Your requirements for implementing these transformations include:
* Executing the transformations on a schedule
* Enabling non-developer analysts to modify transformations
* Providing a graphical tool for designing transformations
What should you do?

  • A. Load each month's CSV data into BigQuery, and write a SQL query to transform the data to a standard schema. Merge the transformed tables together with a SQL query
  • B. Use Apache Spark on Cloud Dataproc to infer the schema of the CSV file before creating a Dataframe.
    Then implement the transformations in Spark SQL before writing the data out to Cloud Storage and loading into BigQuery
  • C. Use Cloud Dataprep to build and maintain the transformation recipes, and execute them on a scheduled basis
  • D. Help the analysts write a Cloud Dataflow pipeline in Python to perform the transformation. The Python code should be stored in a revision control system and modified as the incoming data's schema changes

Answer: B


NEW QUESTION # 192
You need to migrate a 2TB relational database to Google Cloud Platform. You do not have the resources to significantly refactor the application that uses this database and cost to operate is of primary concern.
Which service do you select for storing and serving your data?

  • A. Cloud Bigtable
  • B. Cloud Firestore
  • C. Cloud SQL
  • D. Cloud Spanner

Answer: C


NEW QUESTION # 193
Your team is working on a binary classification problem. You have trained a support vector machine (SVM) classifier with default parameters, and received an area under the Curve (AUC) of 0.87 on the validation set.
You want to increase the AUC of the model. What should you do?

  • A. Train a classifier with deep neural networks, because neural networks would always beat SVMs
  • B. Perform hyperparameter tuning
  • C. Deploy the model and measure the real-world AUC; it's always higher because of generalization
  • D. Scale predictions you get out of the model (tune a scaling factor as a hyperparameter) in order to get the highest AUC

Answer: B

Explanation:
https://towardsdatascience.com/understanding-hyperparameters-and-its-optimisation-techniques-f0debba07568


NEW QUESTION # 194
You have a petabyte of analytics data and need to design a storage and processing platform for it. You must be able to perform data warehouse-style analytics on the data in Google Cloud and expose the dataset as files for batch analysis tools in other cloud providers. What should you do?

  • A. Store the full dataset in BigQuery, and store a compressed copy of the data in a Cloud Storage bucket.
  • B. Store and process the entire dataset in Cloud Bigtable.
  • C. Store the warm data as files in Cloud Storage, and store the active data in BigQuery. Keep this ratio as
    80% warm and 20% active.
  • D. Store and process the entire dataset in BigQuery.

Answer: A


NEW QUESTION # 195
Your startup has never implemented a formal security policy. Currently, everyone in the company has access to the datasets stored in Google BigQuery. Teams have freedom to use the service as they see fit, and they have not documented their use cases. You have been asked to secure the data warehouse. You need to discover what everyone is doing. What should you do first?

  • A. Use Stackdriver Monitoring to see the usage of BigQuery query slots.
  • B. Use the Google Cloud Billing API to see what account the warehouse is being billed to.
  • C. Use Google Stackdriver Audit Logs to review data access.
  • D. Get the identity and access management IIAM) policy of each table

Answer: C

Explanation:
First we need to know who is accessing what then we can create suitable policies. Stackdriver is used to track access logs for Bigquery.


NEW QUESTION # 196
All Google Cloud Bigtable client requests go through a front-end server ______ they are sent to a Cloud Bigtable node.

  • A. once
  • B. before
  • C. only if
  • D. after

Answer: B

Explanation:
In a Cloud Bigtable architecture all client requests go through a front-end server before they are sent to a Cloud Bigtable node.
The nodes are organized into a Cloud Bigtable cluster, which belongs to a Cloud Bigtable instance, which is a container for the cluster. Each node in the cluster handles a subset of the requests to the cluster.
When additional nodes are added to a cluster, you can increase the number of simultaneous requests that the cluster can handle, as well as the maximum throughput for the entire cluster.
Reference: https://cloud.google.com/bigtable/docs/overview


NEW QUESTION # 197
You have data located in BigQuery that is used to generate reports for your company. You have noticed some weekly executive report fields do not correspond to format according to company standards for example, report errors include different telephone formats and different country code identifiers. This is a frequent issue, so you need to create a recurring job to normalize the data. You want a quick solution that requires no coding What should you do?

  • A. Create a Spark job and submit it to Dataproc Serverless.
  • B. Use Dataflow SQL to create a job that normalizes the data, and that after the first run of the job, schedule the pipeline to execute recurrently.
  • C. Use Cloud Data Fusion and Wrangler to normalize the data, and set up a recurring job.
  • D. Use BigQuery and GoogleSQL to normalize the data, and schedule recurring quenes in BigQuery.

Answer: C

Explanation:
Cloud Data Fusion is a fully managed, cloud-native data integration service that allows you to build and manage data pipelines with a graphical interface. Wrangler is a feature of Cloud Data Fusion that enables you to interactively explore, clean, and transform data using a spreadsheet-like UI. You can use Wrangler to normalize the data in BigQuery by applying various directives, such as parsing, formatting, replacing, and validating data. You can also preview the results and export the wrangled data to BigQuery or other destinations. You can then set up a recurring job in Cloud Data Fusion to run the Wrangler pipeline on a schedule, such as weekly or daily. This way, you can create a quick and code-free solution to normalize the data for your reports. References:
* Cloud Data Fusion overview
* Wrangler overview
* Wrangle data from BigQuery
* [Scheduling pipelines]


NEW QUESTION # 198
Cloud Dataproc charges you only for what you really use with _____ billing.

  • A. month-by-month
  • B. minute-by-minute
  • C. week-by-week
  • D. hour-by-hour

Answer: B

Explanation:
Explanation
One of the advantages of Cloud Dataproc is its low cost. Dataproc charges for what you really use with minute-by-minute billing and a low, ten-minute-minimum billing period.
Reference: https://cloud.google.com/dataproc/docs/concepts/overview


NEW QUESTION # 199
You operate a logistics company, and you want to improve event delivery reliability for vehicle-based sensors.
You operate small data centers around the world to capture these events, but leased lines that provide connectivity from your event collection infrastructure to your event processing infrastructure are unreliable, with unpredictable latency. You want to address this issue in the most cost-effective way. What should you do?

  • A. Establish a Cloud Interconnect between all remote data centers and Google.
  • B. Write a Cloud Dataflow pipeline that aggregates all data in session windows.
  • C. Have the data acquisition devices publish data to Cloud Pub/Sub.
  • D. Deploy small Kafka clusters in your data centers to buffer events.

Answer: C


NEW QUESTION # 200
You have a data pipeline that writes data to Cloud Bigtable using well-designed row keys. You want to monitor your pipeline to determine when to increase the size of you Cloud Bigtable cluster. Which two actions can you take to accomplish this? Choose 2 answers.

  • A. Monitor storage utilization. Increase the size of the Cloud Bigtable cluster when utilization increases above 70% of max capacity.
  • B. Monitor the latency of write operations. Increase the size of the Cloud Bigtable cluster when there is a sustained increase in write latency.
  • C. Review Key Visualizer metrics. Increase the size of the Cloud Bigtable cluster when the Read pressure index is above 100.
  • D. Review Key Visualizer metrics. Increase the size of the Cloud Bigtable cluster when the Write pressure index is above 100.
  • E. Monitor latency of read operations. Increase the size of the Cloud Bigtable cluster of read operations take longer than 100 ms.

Answer: B,C


NEW QUESTION # 201
MJTelco Case Study
Company Overview
MJTelco is a startup that plans to build networks in rapidly growing, underserved markets around the world.
The company has patents for innovative optical communications hardware. Based on these patents, they can create many reliable, high-speed backbone links with inexpensive hardware.
Company Background
Founded by experienced telecom executives, MJTelco uses technologies originally developed to overcome communications challenges in space. Fundamental to their operation, they need to create a distributed data infrastructure that drives real-time analysis and incorporates machine learning to continuously optimize their topologies. Because their hardware is inexpensive, they plan to overdeploy the network allowing them to account for the impact of dynamic regional politics on location availability and cost.
Their management and operations teams are situated all around the globe creating many-to-many relationship between data consumers and provides in their system. After careful consideration, they decided public cloud is the perfect environment to support their needs.
Solution Concept
MJTelco is running a successful proof-of-concept (PoC) project in its labs. They have two primary needs:
* Scale and harden their PoC to support significantly more data flows generated when they ramp to more than 50,000 installations.
* Refine their machine-learning cycles to verify and improve the dynamic models they use to control topology definition.
MJTelco will also use three separate operating environments - development/test, staging, and production - to meet the needs of running experiments, deploying new features, and serving production customers.
Business Requirements
* Scale up their production environment with minimal cost, instantiating resources when and where needed in an unpredictable, distributed telecom user community.
* Ensure security of their proprietary data to protect their leading-edge machine learning and analysis.
* Provide reliable and timely access to data for analysis from distributed research workers
* Maintain isolated environments that support rapid iteration of their machine-learning models without affecting their customers.
Technical Requirements
Ensure secure and efficient transport and storage of telemetry data
Rapidly scale instances to support between 10,000 and 100,000 data providers with multiple flows each.
Allow analysis and presentation against data tables tracking up to 2 years of data storing approximately 100m records/day Support rapid iteration of monitoring infrastructure focused on awareness of data pipeline problems both in telemetry flows and in production learning cycles.
CEO Statement
Our business model relies on our patents, analytics and dynamic machine learning. Our inexpensive hardware is organized to be highly reliable, which gives us cost advantages. We need to quickly stabilize our large distributed data pipelines to meet our reliability and capacity commitments.
CTO Statement
Our public cloud services must operate as advertised. We need resources that scale and keep our data secure. We also need environments in which our data scientists can carefully study and quickly adapt our models. Because we rely on automation to process our data, we also need our development and test environments to work as we iterate.
CFO Statement
The project is too large for us to maintain the hardware and software required for the data and analysis. Also, we cannot afford to staff an operations team to monitor so many data feeds, so we will rely on automation and infrastructure. Google Cloud's machine learning will allow our quantitative researchers to work on our high- value problems instead of problems with our data pipelines.
Given the record streams MJTelco is interested in ingesting per day, they are concerned about the cost of Google BigQuery increasing. MJTelco asks you to provide a design solution. They require a single large data table called tracking_table. Additionally, they want to minimize the cost of daily queries while performing fine-grained analysis of each day's events. They also want to use streaming ingestion. What should you do?

  • A. Create a partitioned table called tracking_table and include a TIMESTAMP column.
  • B. Create a table called tracking_table with a TIMESTAMP column to represent the day.
  • C. Create a table called tracking_table and include a DATE column.
  • D. Create sharded tables for each day following the pattern tracking_table_YYYYMMDD.

Answer: A


NEW QUESTION # 202
Which is the preferred method to use to avoid hotspotting in time series data in Bigtable?

  • A. Randomization
  • B. Salting
  • C. Hashing
  • D. Field promotion

Answer: D

Explanation:
By default, prefer field promotion. Field promotion avoids hotspotting in almost all cases, and it tends to make it easier to design a row key that facilitates queries.
Reference: https://cloud.google.com/bigtable/docs/schema-design-time-
series#ensure_that_your_row_key_avoids_hotspotting


NEW QUESTION # 203
Which role must be assigned to a service account used by the virtual machines in a Dataproc cluster so they can execute jobs?

  • A. Dataproc Worker
  • B. Dataproc Runner
  • C. Dataproc Editor
  • D. Dataproc Viewer

Answer: A

Explanation:
Explanation
Service accounts used with Cloud Dataproc must have Dataproc/Dataproc Worker role (or have all the permissions granted by Dataproc Worker role).
Reference: https://cloud.google.com/dataproc/docs/concepts/service-accounts#important_notes


NEW QUESTION # 204
You designed a database for patient records as a pilot project to cover a few hundred patients in three clinics.
Your design used a single database table to represent all patients and their visits, and you used self-joins to generate reports. The server resource utilization was at 50%. Since then, the scope of the project has expanded.
The database must now store 100 times more patient records. You can no longer run the reports, because they either take too long or they encounter errors with insufficient compute resources. How should you adjust the database design?

  • A. Partition the table into smaller tables, with one for each clinic. Run queries against the smaller table pairs, and use unions for consolidated reports.
  • B. Add capacity (memory and disk space) to the database server by the order of 200.
  • C. Normalize the master patient-record table into the patient table and the visits table, and create other necessary tables to avoid self-join.
  • D. Shard the tables into smaller ones based on date ranges, and only generate reports with prespecified date ranges.

Answer: D


NEW QUESTION # 205
You are on the data governance team and are implementing security requirements to deploy resources. You need to ensure that resources are limited to only the europe-west 3 region You want to follow Google-recommended practices What should you do?

  • A. Set the constraints/gcp. resourceLocations organization policy constraint to in: europe-west3-locations.
  • B. Deploy resources with Terraform and implement a variable validation rule to ensure that the region is set to the europe-west3 region for all resources.
  • C. Create a Cloud Function to monitor all resources created and automatically destroy the ones created outside the europe-west3 region.
  • D. Set the constraints/gcp. resourceLocations organization policy constraint to in:eu-locations.

Answer: A

Explanation:
To ensure that resources are limited to only the europe-west3 region, you should set the organization policy constraint constraints/gcp.resourceLocations to in:europe-west3-locations. This policy restricts the deployment of resources to the specified locations, which in this case is the europe-west3 region. By setting this policy, you enforce location compliance across your Google Cloud resources, aligning with the best practices for data governance and regulatory compliance.
References:
* Professional Data Engineer Certification Exam Guide | Learn - Google Cloud1.
* Preparing for Google Cloud Certification: Cloud Data Engineer2.
* Professional Data Engineer Certification | Learn | Google Cloud3.
3: Professional Data Engineer Certification | Learn | Google Cloud 2: Preparing for Google Cloud Certification: Cloud Data Engineer 1: Professional Data Engineer Certification Exam Guide | Learn - Google Cloud


NEW QUESTION # 206
You are operating a Cloud Dataflow streaming pipeline. The pipeline aggregates events from a Cloud Pub/Sub subscription source, within a window, and sinks the resulting aggregation to a Cloud Storage bucket. The source has consistent throughput. You want to monitor an alert on behavior of the pipeline with Cloud Stackdriver to ensure that it is processing data. Which Stackdriver alerts should you create?

  • A. An alert based on an increase of instance/storage/used_bytes for the source and a rate of change decrease of subscription/num_undelivered_messages for the destination
  • B. An alert based on a decrease of subscription/num_undelivered_messages for the source and a rate of change increase of instance/storage/used_bytes for the destination
  • C. An alert based on an increase of subscription/num_undelivered_messages for the source and a rate of change decrease of instance/storage/used_bytes for the destination
  • D. An alert based on a decrease of instance/storage/used_bytes for the source and a rate of change increase of subscription/num_undelivered_messages for the destination

Answer: C


NEW QUESTION # 207
Your financial services company is moving to cloud technology and wants to store 50 TB of financial time-series data in the cloud. This data is updated frequently and new data will be streaming in all the time.
Your company also wants to move their existing Apache Hadoop jobs to the cloud to get insights into this data. Which product should they use to store the data?

  • A. Google Cloud Storage
  • B. Google BigQuery
  • C. Cloud Bigtable
  • D. Google Cloud Datastore

Answer: C


NEW QUESTION # 208
......


Google Professional-Data-Engineer certification is a highly respected and in-demand certification for data professionals. Google Certified Professional Data Engineer Exam certification is designed for individuals who possess the knowledge and skills to design, build, maintain, and troubleshoot data processing systems with a particular emphasis on the Google Cloud Platform. Google Certified Professional Data Engineer Exam certification is offered by Google and is recognized globally as a valuable credential for professionals in the data engineering field.

 

Feel Google Professional-Data-Engineer Dumps PDF Will likely be The best Option: https://www.real4prep.com/Professional-Data-Engineer-exam.html

Professional-Data-Engineer exam torrent Google study guide: https://drive.google.com/open?id=1n4dZXrKGNtLhjjIq1LYWzQQxpg7-Qkhd