- Incident Management & Escalation:
Serve as the primary technical escalation point for production incidents and customer-impacting outages — responding to on-call alerts, triaging issues, and driving swift resolution in alignment with established incident management processes and SLA/SLO targets. - Cross-team Collaboration:
Work closely with development, platform, and other engineering teams to restore services quickly and safely, ensuring clear communication and coordinated response throughout the incident lifecycle. - Troubleshooting & Root Cause Analysis:
Investigate and diagnose complex system and application issues, perform thorough root cause analysis, and document findings with clear, actionable remediation plans validated across staging and production environments. - Post-Incident Reviews:
Participate in post-incident reviews, track follow-up actions through to completion, and ensure learnings are systematically applied to prevent recurrence. - Monitoring & Observability:
Continuously monitor system health, availability, latency, and error rates using observability tooling across metrics, logs, and traces, proactively identifying and addressing reliability risks before they impact users. - Reliability & Performance Optimization:
Assist with capacity planning, troubleshoot performance degradation under high load, and recommend improvements to architecture, configurations, and deployment practices to strengthen overall system resilience. - Toil Reduction & Automation:
Identify recurring operational issues and implement automation, configuration fixes, or preventive measures to reduce toil and improve long-term platform stability. Improve and develop new tools for visualisation, event correlation and log analysis including automated and AI assisted ones. - Documentation & Knowledge Sharing:
Maintain and continuously improve runbooks, troubleshooting guides, and incident response documentation, while providing technical guidance to L1/L2 support teams to elevate overall support quality.
Requirements:
- Strong knowledge of web application stack components and their roles (nginx, php, redis, rabbitmq, postgreSQL).
- Strong knowledge of monitoring and log analysis tools (Icinga2, Grafana, Elasticsearch, OpenSearch etc.).
- Strong understanding of linux and web application logging and event correlation.
- Strong communication skills, capable of translating complex technical concepts into clear, customer-friendly language.
- Good knowledge of PHP with experience with Web Frameworks (preferably Symfony).
- Good knowledge of SQL and PostgreSQL.
- Good knowledge of e-commerce related concepts (functional and technical level).
- Good knowledge of frontend technologies (CSS, JavaScript, HTML, frameworks).
- Experienced in troubleshooting of web applications, performing a root-cause analysis.
- Understanding of ITIL V4/V5.Our offer:
- Unlimited vacation, covered sick leaves
- The opportunity for professional growth
- Welcoming atmosphere (awesome team of professionals always ready to help)
- Onboarding program
- English Courses
About project:
Oro, Inc. is a leading enterprise software development company based in the U.S., offering B2B Commerce and CRM applications to companies around the world. As a product company, we focus on the development of Oro’s suite of software solutions, including: OroCRM, OroCommerce, OroPlatform and OroMarketplace. Oro was founded in 2012 and today has grown to 160+ employees with operations in France, Germany, Poland, Ukraine and the US.
At Oro, we are enthusiastic about what we do. And we do everything with passion.
This is not only because of the amazing products we create — but because our team is made up of a diverse group of talented technology professionals and industry leaders.
We are a fast-growing company and we’re hiring a Technical Project Manager to join our Professional Services team.