DevOps / Site Reliability Engineer specialising in building and operating production-grade platforms at scale – improving resilience, automating infrastructure, and enabling high-performing engineering teams.At Siemens Digital Industries Software I owned and operated CI/CD and platform infrastructure powering multiple engineering teams across hybrid cloud environments – AWS, VMware and OpenStack, running both Linux and Windows. I designed systems to withstand real production pressure, cutting change-to-production lead time by around 65% and MTTR by around 30%.Much of that work was governed self-service: reusable Terraform modules and CloudFormation templates letting teams provision compute, networking and access repeatably, without routing around security controls. Configuration and deployment across Linux and Windows fleets ran through Ansible/AWX, PowerShell and Python.I led end-to-end incident response in highavailability systems, driving RCA and long-term fixes for recurring
Add work experience to your profile. (optional)
Site Reliability Engineer supporting availability-critical 4G/5G
telecommunications platforms in large-scale distributed production
environments.
* Operated production services across RHEL, CentOS, Ubuntu and VMware,
maintaining high availability under real production load.
* Led time-sensitive incident response and root cause analysis for compute,
storage, network, database and operating-system failures, using Zabbix and
Splunk.
* Built Python, Bash and Ansible automation for patching, deployment and
remediation, reducing manual operations.
* Maintained CI/CD pipelines, Nexus/Artifactory repositories and rollback
processes.
* Supported high availability, backup/restore and disaster recovery, producing
runbooks, post-mortems and corrective actions with development and security
teams.
Primary owner of a business-critical platform across 3 AWS accounts and hybrid Linux, Windows and VMware environments, supporting a 10-19 engineer team.
Built reusable Terraform & CloudFormation modules for VPC, IAM, EC2, S3 — giving teams safe, self-service infrastructure provisioning
Automated deployments across Linux and Windows fleets with Ansible and Python
Ran containerized services on ECS/Fargate and EKS
Cut deployment lead time ~65% via CI/CD pipelines, automated testing and rollback automation
Cut incident recovery time (MTTR) ~30% through better monitoring, alerting and runbooks
Implemented least-privilege access controls and security scanning across the platform
Add work education to your profile. (optional)
We will review the reports from both freelancer and employer to give the best decision. It will take 3-5 business days for reviewing after receiving two reports.