About Me
8+ Years(Total) of extensive IT experience in all phases of Software Development Life Cycle (SDLC) with skills in data analysis, design, development, testing and deployment of software systems for client/server multi-user business applications. 4+ Ye...testing and deployment of software systems for client/server multi-user business applications. 4+ Years of strong experience, working on Apache Hadoop ecosystem components like MapReduce, HDFS, HBase, Hive, Sqoop, Pig, Oozie, Zookeeper, Flume, Spark, Python with CDH4&5 distributions and EC2 cloud computing with AWS. Working closely with the stakeholders & solution architect, Ensuring architecture meets the business requirements, Building highly scalable, robust & fault-tolerant systems. Key participant in all phases of software development life cycle with Analysis, Design, Development, Integration, Implementation, Debugging, and Testing of Software Applications in client server environment Strong in Developing MapReduce applications, Configuring the Development Environment, Tuning Jobs and Creating MapReduce Workflows. Experience in performing data enrichment, cleansing, analytics, aggregations using Hive and Pig. Knowledge in Cloudera CDH4 and CDH5 distributions and Hortonworks (HDP Proficient in big data ingestion and streaming tools like Flume, Sqoop, Kafka, Storm and Kinesis Experience with different data formats like Json, Avro, parquet, RC and ORC and compressions like snappy & bzip. Experienced in analyzing data using HQL, PigLatin and extending HIVE and PIG core functionality by using custom UDFs. Good Knowledge/Understanding of NoSQL data bases and hands on work experience in writing applications on NoSQL databases like Hbase and MongoDB. Good knowledge on various scripting languages like Linux/Unixshell scripting and Python. Good knowledge of Dataware housing concepts and ETL processes. Importing / exporting data from RDBMS to HDFS for batch data process using SQOOP Configured Zookeeper to coordinate the servers in clusters to maintain the data consistency. Used Oozie and Control - M workflow engine for managing and scheduling Hadoop Jobs. Diverse experience in working with variety of Database like Oracle, MySql, Salesforce and Netezza. AWS provides a secure global infrastructure, plus a range of features that use to secure the data in the cloud Hands on experience on AWS cloud services (VPC, EC2, S3, RDS, Redshift, Data Pipeline, EMR, DynamoDB, WorkSpaces, Lambda, Kinesis, SNS, SQS) Good experience of AWS Elastic Block Storage (EBS), different volume types and use of various types of EBS volumes based on requirement. Ability to spin up different AWS Instances including EC2-Classing and EC2-VPC using cloud formation template Cognitive about designing, deploying and operating highly available, scalable and fault tolerant systems using Amazon Web Services (AWS). With the help of IAM created roles, users and groups and attached policies to provide minimum access to the resources, created topics in SNS to send notifications to subscribers as per the requirement. Implemented Amazon RDS multi-AZ for automatic failover and high availability at the database tier, created CloudFront distributions to serve content from edge locations to users so as to minimize the load on the frontend servers. Experienced in Performance Tuning and Query Optimization in AWS Redshift. Implemented POC to migrate map reduce programs into Spark transformations using Spark and Scala. Good knowledge in understanding Core Java and J2EE technologies such as Hibernate, JDBC, EJB, Servlets, JSP, JavaScript, Struts and spring. Experienced in using IDEs and Tools like Eclipse, Jenkins, Maven and IntelliJ. Familiar with data architecture including data ingestion pipeline design, Hadoop information architecture, data modeling and data mining, machine learning and advanced data processing. Experience optimizing ETL workflows. Experience on handling cluster when it is in Safe mode. Good knowledge of High-Availability, Fault Tolerance, Scalability, Database Concepts, System and Software Architecture, Security and IT Infrastructure. All the projects which I have worked for are Open Source Projects and has been tracked using JIRA. Load and transform large sets of structured, semi-structured and unstructured data using Hadoop ecosystems Good knowledge of single node and multi-node cluster setup Strong experience with good knowledge in SQL (Including Triggers, Stored Procedures). Experience in writing small to complex queries. Lead onshore & offshore service delivery functions to ensure end-to-end ownership of incidents and service requests. Getting in touch with the Junior developers and keeping them updated with the present cutting Edge technologies like Hadoop, Spark, SparkSQL. Highly motivated, subject oriented, has ability to work independently and as a part of the team with Excellent Technical, Analytical and Communication skills. Strong team player, ability to work independently and in a team as well, ability to adapt to a rapidly changing environment, commitment towards learning, Possess excellent communication, project management, documentation, interpersonal skills.
Show More