Site Reliability Engineer (SRE)

Ref. # 2249
Work type
На място
Place of work
гр. София

Published on:

21 July 2026

Отговорности

  • Design and implement reliable, scalable and efficient systems with cutting age technologies
  • You will drive reliability and observability improvements in the services within the engineering verticals
  • You will create and improve internal tools and automation software to make maintaining production services faster, easier and safer
  • Respond to resolve incidents, ensuring minimal impact on services
  • Develop and improve Product monitoring and logging systems based on Prometheus, Alertmanager, Grafana, fluent-bit and ELK stack and others
  • Support the release process via CI/CD pipelines
  • Automate the platform with infrastructure-as-code and configuration management tools like Ansible, Terraform and others
  • You will create Bash scripts and use Python
  • Maintain clear and comprehensive documentation for systems, processes and procedures. Share knowledge with team members and other colleagues from the development team
  • On-call rotation support for the production environment outages

Изисквания

  • 2+ years’ experience in an SRE, DevOps, Cloud Ops or Infrastructure role
  • Experience with monitoring systems like: Prometheus, Zabbix, Alertmanager, Grafana or others.
  • Experience with logging tools like: ELK stack, Fluentd/Fluent-Bit
  • Good knowledge in one or more CI/CD systems like Jenkins, Chef, ArgoCD, FluxCD,
  • Hands-on experience in one or more IaC/Configuration Management tools like Terraform, Pulumi, CloudFormation, Ansible, Puppet or others
  • You will need to use Linux and have some administration experience

Good to have:

  • We have microservices in Kubernetes and it’s important at least to understand Kubernetes basics.
  • Programming language skills – Python, Go, Bash preferred
Professional field
ИТ - Разработка / поддръжка на хардуер, ИТ - Разработка / поддръжка на софтуер