We are looking for a DevOps / Site Reliability Engineer (SRE) to support the development, operation and maintenance of large-scale distributed systems running on Linux and Open Source technologies. This role focuses on infrastructure reliability, Kubernetes operations, automation, observability and the continuous improvement of critical production environments.
Main Responsibilities
- Administer and operate Linux-based distributed systems and platforms.
- Manage and support Kubernetes clusters in on-premises environments.
- Implement and maintain DevOps and automation tools.
- Monitor, troubleshoot and optimize critical production systems.
- Support CI/CD implementation and operational practices.
- Administer relational and non-relational database platforms.
- Maintain observability, logging and monitoring solutions.
- Contribute to infrastructure scalability, performance and high availability initiatives.
- Provide operational support and participate in on-call activities.