Operations Support Engineer
About the job
About the Operations Support Engineer role
The Operations Support Engineer is responsible for monitoring and maintaining production and non-production environments, resolving technical and security incidents, and ensuring reliable system performance and service delivery. The role involves analysing runtime environments, developing data-driven improvement strategies, coordinating with application, architecture, security, and vendor teams, and supporting AWS environments for AI use cases. The engineer will manage incident and change processes, implement monitoring, access controls, automation, and preventive measures, while maintaining operational documentation and audit compliance. The role also includes managing a 24/7 operations team, preparing performance and status reports, and supporting cloud, DevOps, Agile, and information security practices. Experience with AWS, cloud-native monitoring, infrastructure automation, and AI applications using RAG, tool calling, workflow automation, and data lakes is advantageous.
Key Responsibilities:
- Responsibilities include monitoring system performance, resolving technical issues, coordinating with different teams for problem resolution, creating preventive measures(optional), and maintaining documentation related to system configuration, process, and service records
- Monitor and analyze the current state of various product runtime environment (production and non-production) to ensure optimum system performance, and work out data-based strategy for continuous improvement
- Work with application teams, solution architects, security consultants, and other teams to implement improvement plans
- Work closely with business users to setup AWS Quick environment to support AI use cases
- Manage application and security incidents, conduct problem determination, work with various internal teams and vendors to resolve issues on a timely basis to meet SLA, provides reporting and escalation to higher management or incident committee if necessary
- Develop operations and processes guide to ensure every aspect of operations is documented and complies with audit requirements
- Manage day-to-day operation activities, analyze statistics and write status and progress reports, and present findings to stakeholders and higher management
- Manage operations team consisting of staff and vendors, ensuring support is available on a 24/7 basis
Requirements:
- Bachelor’s degree in Computer Science, Information Technology, or a related field
- Proven experience as an Operations Engineer or similar role in an IT setting
- Implement change management and incident management workflows, using ITSM tools e.g. AWS Quick, Zendesk, ServiceDesk to automate workflows is advantageous
- Implement security and access control measures to control privileged access to test and production environment
- Implement full stack monitoring (i.e. application and infrastructure) using Application Performance Management (APM) tools. Familiarity with cloud native monitoring options (e.g. Cloudwatch) is preferred
- Identify and implement process automation to minimum downtime and human errors. Familiarity with scripting tools e.g. Terraform, Ansible is preferred
- Experienced in agile methodologies, DevOps pipelines, test-driven development, and info-security practices
- Able to work collaboratively with a high performance team and influence with positive energy
- Resourceful and able to work out solutions with innovative thinking and new tech
- Experienced with management cloud infrastructure and services / certification with GPC, GCC (i.e. AWS, Azure, Google Cloud) or equivalent cloud platforms will be preferred
- Experience in AWS ecosystem services and development of AI applications demonstrating proficiency in RAG, tool calling, workflow automation, data lakes are highly advantageous
- Excellent problem-solving skills
- Strong communication skills, with the ability to communicate complex technical issues to non-technical teams

