OptimHire’s Job Auto-Applier that will automatically apply to jobs at 300,000+ Companies.
GCP Data Engineer
Job Summary: We are seeking a talented and experienced GCP Data Engineer to join our team for Pune location. The GCP Data Engineer will be designing, implementing, and maintaining the data infrastructure and pipelines that enable efficient and reliable data processing, storage, and analysis. Overall experience should be in the range of 4-7 years.
Key Responsibilities:
1. Data Pipeline Development: Design, build, and manage data pipelines for extracting, transforming, and loading (ETL) data from various sources into Google Cloud services like BigQuery, Cloud Storage, and more.
2. Data Modeling: Develop and maintain data models and schemas that facilitate efficient querying and analysis while adhering to best practices for data warehousing.
3. ETL Processes: Create and maintain ETL processes to cleanse, transform, and enrich data before loading it into data warehouses or other analytical systems.
4. Big Data Technologies: Utilize GCP big data technologies such as Google Dataproc, Dataflow, and Pub/Sub for processing and streaming large volumes of data.
5. Data Warehousing: Implement and manage data warehousing solutions using Google BigQuery, ensuring scalability, performance, and cost-efficiency.
6. Data Quality and Governance: Establish data quality standards, data lineage, and governance practices to ensure data accuracy, consistency, and security.
7. Serverless Computing: Develop data processing workflows using GCP serverless offerings like Cloud Functions, Cloud Run, and Cloud Scheduler.
Qualifications:
• BE/BTech/MCA in computer science from reputed college.
• Proven production level at least 4 years in Google Cloud Platform services, including BigQuery, Cloud Storage, Dataproc, Dataflow, Pub/Sub, etc.
• Project experience of at least 3 years programming languages commonly used in data engineering, such as Python or Java.
• Must have delivered E2E 1-2 projects using data modeling principles, relational databases, and data warehousing concepts.
• 3+ years of experience in designing and implementing ETL processes and data pipelines using technologies like AIRflow, DataFlow, DataFusion etc..
• At least 2 years of working proficiency in PySpark and SQL using tools like Data Bricks, DataProc, BigQuery etc.
• Expert with database query languages, such as SQL, for data manipulation and analysis.
• Understanding of data quality, data governance, and data security best practices.
Java (All Versions) - 3 years
Python - 3 years
SQL
ETL(Extract, Transform, Load)
PySpark
Cloud Storage
BigQuery
Google Cloud Platform - IAAS - 4 years
Anew-age AI-First Digital and Cloud Engineering services company
Healthcare Benefits
Professional Development