The client is an organisation dedicated to creating a meaningful impact for its clients, employees, and the communities it serves.
With a strong commitment to being a force for positive change, its initiatives aim to address significant societal challenges and foster a better future. The organization's efforts focus on guiding clients toward purpose-driven growth while embedding practices that promote equity, inclusivity, and sustainability in business operations.
At Cloud Bridge Recruitment Services, we are assisting in hiring fifteen (15) Site Reliability Engineers (SREs) who will be instrumental in maintaining the performance and reliability of essential services. Your skills will help connect development and operations, ensuring a scalable, robust, and responsive infrastructure.
This position emphasizes solid system architecture and design, with a focus on core SRE practices such as Service Level Objectives (SLOs), Service Level Indicators (SLIs), and minimizing operational toil. You will work closely with diverse teams to enhance system reliability while promoting a culture of continuous learning and accountability.
This role will be based in Kuala Lumpur,Malaysia (work permit will be provided to successful candidates outside of Malaysia). Key Responsibilities Own the Managed Services ITSM ticketing queue, assigning and escalating tickets and resources accordingly. Analyse ticketing to create compliance reports as part of continual service improvements. Build and maintain relationships with customers, internal stakeholders, and third-party suppliers. Assist the Customer Success function with process improvements across the Managed Services spectrum. Work with Customer Success to ensure timely, risk-free onboarding of new customers. Build and maintain internal and external facing technical documentation. Track key actions and tasks per customer account. Help to build and improve the Managed Services teams, including driving and tracking certifications / skillsets to customer requirements and cloud roadmaps. Focus on the team individuals to align their aspirations with suitable roles in the team. Assist with team recruitment. Manage, lead and mentor your team members, including assigning resources to key customer accounts. Track resource utilisation within the team and manage accordingly. Investigate and resolve customer cloud solution issues. Help manage major incidents. Create and help improve process documentation. Qualifications Proficiency in programming languages such as Python, Golang, Java, or similar, focusing on operational efficiency. Demonstrated experience in system architecture and design, prioritizing reliability, and scalability. Strong understanding of SRE principles, including SLOs, SLIs, toil reduction, and incident post-mortems. Experience with cloud environments (e.g., AWS, Azure, Google Cloud) and their operational management. Strong expertise in Linux system administration. Proven experience in troubleshooting application support issues with a focus on performance and connectivity. Familiarity with networking concepts and effective troubleshooting techniques. Excellent problem-solving abilities and a proactive approach to operational challenges. Ability to work independently while effectively collaborating within a team environment. Essential Skills Familiarity with monitoring tools and performance optimization techniques. Experience in scripting or automation for system administration tasks. Knowledge of networking concepts and troubleshooting methodologies. Hands-on knowledge of cloud platforms (e.g., AWS, Azure, Google Cloud) and their services. Familiarity with DevOps practices and frameworks, including CI/CD, infrastructure as code, and containerization. Remuneration Salary range: 9K – 11K MYR (Malaysia Ringgit) per month for 2-4 years experience candidate. Salary range: 7K – 8K MYR (Malaysia Ringgit) per month for a 1-2 years experience candidate.
#J-18808-Ljbffr