The Application Production Support Engineer is a detail-oriented and technically skilled professional responsible for managing and supporting critical business applications in a production environment, with a strong focus on delivering high-quality IT services.
This role is responsible for ensuring the stability, availability, performance, and reliability of applications within the team's scope. The professional will also manage incident resolution in a timely manner, collaborating closely with internal stakeholders (Development, Infrastructure, and Architecture teams) as well as external partners and service providers to drive sustainable and long-term solutions.
Main Responsibilities
Application Stability & Availability (Expert)
- Monitor, maintain, and support business-critical applications, ensuring high availability, stability, and optimal performance.
- Actively participate in Incident and Problem Management processes, including:
- Participation in Major Incident (P1/P2) Situation Rooms.
- Root Cause Analysis (RCA) activities.
- Identification of incident trends and recurring issues.
- Contribution to permanent corrective actions and preventive measures.
- Ensure compliance with ITIL governance practices, operational procedures, and established Service Level Agreements (SLAs).
- Execute application deployments, releases, and change requests following ITIL and DevOps methodologies.
- Proactively identify, analyze, and resolve technical issues to support uninterrupted business operations.
- Participate in on-call support rotations and provide 24/7 support coverage for critical applications when required.
Technical Support & Collaboration (Expert)
- Serve as a primary point of contact between Production Support and Development teams for troubleshooting and issue resolution.
- Collaborate closely with Scrum and DevOps teams to design, deploy, maintain, and continuously improve application services.
- Implement upgrades, patches, configuration changes, and new functionalities while minimizing business impact and ensuring service continuity.
- Contribute to the continuous improvement of operational processes, automation, and platform reliability.
Documentation & Knowledge Sharing (Expert)
- Create, maintain, and continuously update technical documentation, including operational procedures, system configurations, troubleshooting guides, and runbooks.
- Promote knowledge sharing and best practices across global support teams to improve operational efficiency and service quality.
- Support the development and maintenance of knowledge bases to facilitate issue resolution and onboarding activities.
Platform Monitoring & Observability (Expert)
- Implement, maintain, and optimize monitoring and observability solutions across production environments.
- Leverage platforms such as Dynatrace and other observability tools to ensure proactive monitoring and rapid incident detection.
- Collaborate with Development teams and Centers of Expertise to define and enhance observability standards and monitoring strategies.
- Promote an observability-first mindset, enabling early identification and resolution of potential service disruptions.
- Continuously improve monitoring dashboards, alerting mechanisms, and operational metrics.
General Responsibilities
- Complete all mandatory training required for the effective operation of the IT Production area and compliance with company policies.
- Perform additional activities, when required, to support business objectives and ensure the proper functioning of the Center of Expertise.
- Contribute to continuous improvement initiatives within the Production Support organization.