职位详情
国籍要求:马来西亚
职位描述
What You’ll Do:
Design and Develop Data Pipelines: Create and manage robust data pipelines using Apache Spark and Microsoft Fabric to process large volumes of data efficiently.
Data Integration: Integrate data from various sources, ensuring data quality, consistency, and reliability.
Performance Optimization: Optimize data processing workflows for performance and scalability.
Collaboration: Work closely with data scientists, analysts, and other stakeholders to understand data requirements and deliver solutions.
Data Governance: Implement and maintain data governance practices to ensure data security, privacy, and compliance.
Troubleshooting: Identify and resolve data-related issues and bottlenecks.
Documentation: Maintain comprehensive documentation of data engineering processes and workflows.
What You Need to Succeed:
Experience: 5 years in data engineering or related fields.
Technical Skills:
Proficiency in Databrick, Apache Spark or Microsoft Fabric.
Strong programming skills in languages such as Python, Scala, or Java.
Experience with cloud platforms (e.g., Azure, AWS, GCP).
Knowledge of SQL and NoSQL databases.
Familiarity with data warehousing solutions and ETL processes.
Analytical Skills: Strong problem-solving and analytical skills.
Communication: Excellent communication and collaboration skills.
Education: Bachelor’s degree in computer science, engineering, or a related field.
Additional Skills That Could Set You Apart:
Experience with big data technologies and frameworks.
Knowledge of machine learning and data science concepts.
Certification in relevant technologies (e.g., Apache Spark, Azure Data Engineer).