DevOps / Infrastructure & Field Support Engineer

xBerry – we are an R&D House gaining experience in delivering custom solutions for international clients since 2016. We provide extensive expertise in embedded systems, machine learning, AR/VR technology, and image processing.
Incident Handling and System Maintenance
- Diagnosing and resolving issues related to:
- Kubernetes clusters,
- containers (Docker),
- Linux (Ubuntu) operating system,
- networking,
- storage (including NFS),
- Analyzing logs and service health across application and infrastructure layers,
- Restoring full system functionality in production environments,
- Performing system deployments and upgrades at customer sites,
- Participating in on-site interventions when issues cannot be resolved remotely.
Automation, Observability, and System Resilience
- Designing and developing automated troubleshooting mechanisms,
- Early detection of infrastructure and application-level issues,
- Automated validation of the health of key system components:
- OS,
- Kubernetes,
- containers,
- storage,
- networking,
- Building health checks and observability solutions (metrics, alerts, dashboards),
- Creating and maintaining:
- runbooks,
- standard recovery procedures,
- automated self-healing mechanisms,
- Documenting common incidents, root causes, and resolution methods.
Collaboration and Architecture Improvement
- Close cooperation with development and architecture teams,
- Contributing to architecture simplification and standardization,
- Improving overall system stability and reliability,
- Supporting long-term efforts to reduce operational overhead and manual interventions.
Technical Requirements
- Strong experience with Linux (Ubuntu) system administration and troubleshooting,
- Hands-on experience with Kubernetes, including cluster troubleshooting and container analysis,
- Practical knowledge of Docker,
- Solid understanding of networking and diagnosing network-related issues,
- Experience with NFS / storage troubleshooting,
- Operational knowledge of GPU / CUDA environments (compatibility, stability),
- Experience working with:
- RabbitMQ,
- PostgreSQL.
Additional Requirements
- Willingness to participate in an on-call / standby rotation,
- Readiness for business travel, including on-site customer visits,
- Ability to work independently in complex, distributed environments,
- Strong analytical and problem-solving skills.
XBERRY Sp. zoo. needs contact details and your personal data to be processed for the purposes of the current recruitment process. For information on how to withdraw consent, as well as our privacy practices and commitment to protecting your privacy, please see our Privacy Policy. You can contact our Human Resource Specialist by sending an email to: hr@xberry.tech
How the Recruitment process looks?
- 1CV Review
- 2HR call
- 3Technical interview
- 4Client meeting (depends on the project)
- 5Offer
