职位详情
国籍要求:马来西亚
职位描述
Responsibilities:
Responsible for monitoring and alerting implementation of core systems and applications, ensuring system stability and high availability
Participate in incident management, including troubleshooting, root cause analysis, resolution, and continuous improvement
Perform resource analysis, performance evaluation, and capacity planning for systems
Drive the implementation of DevOps practices, enhancing overall operational capabilities, including:
Continuous Integration (CI)
Application release and deployment
Continuous Deployment (CD)
Monitoring and alerting systems
Incident response and contingency planning
Intelligent operations (AIOps)
Promote standardization, automation, and intelligence in operations processes
Requirements:
Bachelor’s degree or above in Computer Science or a related field, with at least 1 year of relevant experience
Strong expertise in Linux operating systems with in-depth system-level understanding
Proficient in at least one scripting language and one compiled language; experience in large-scale system design is a plus
Solid understanding of TCP/IP and HTTP protocols, with hands-on troubleshooting experience preferred
Familiar with containerization and orchestration technologies; experience with Kubernetes (K8s) in production is preferred
Good understanding of database principles; familiarity with common database engines is a plus
Strong knowledge of distributed systems and commonly used open-source components, such as:
Nginx
Redis
Kafka
MySQL
HBase
Zookeeper
Hadoop
Experience in big data operations/development or machine learning algorithms is an advantage
Experience with CI/CD pipelines and large-scale cluster management is highly preferred
Strong sense of responsibility, proactive attitude, eagerness to learn, and good teamwork skills