Marcin Jasion
Senior Site Reliability Engineer
Senior Site Reliability Engineer based in Warsaw. I work where code meets infrastructure - diagnosing incidents, fixing the observability gaps that hide them, and shipping the boring-but-important changes that prevent the next outage. Currently helping engineering teams operate a system of 1000+ services at NatWest Boxed - mostly Java services in an asynchronous, Kafka-based architecture. 13 years in tech: started as a Java developer, found my home in the messy middle where applications meet production.
Experience
NatWest Boxed
Warsaw, PolandBanking-as-a-Service platform from NatWest Group. As an SRE consultant I support almost 30 engineering teams operating a system of 1000+ services - mostly Java services in an asynchronous, event-driven architecture built on Kafka. My work spans incident response, observability, and reliability improvements across the platform. Teams are organised using Team Topologies and mine is an enabling team, so it cuts across many different stream-aligned squads.
- Support almost 30 engineering teams resolving production incidents across a system of 1000+ services
- Act as an enabling team in a Team Topologies setup — bridging stream-aligned squads with shared SRE practice
- Improve observability - metrics, logs, traces - to shorten time-to-detect and time-to-diagnose
- Build agents for incident handling and streamline the resolution process for on-call engineers
- Drive reliability improvements that prevent recurring outages and reduce on-call toil
Datachain (former Iterative.ai)
Global (Remote)Datachain builds open-source tools for machine learning, offering team collaboration products as SaaS and self-hosted solutions.
- Designed a multi-cloud platform for running distributed ML experiments with billions of objects (AWS and GCP)
- Designed and implemented Kubernetes Pod autoscaling based on Celery queue depth
- Designed the on-premise, enterprise edition of DVC Studio to run reliably in arbitrary customer environments (ESXi, Hyper-V, KVM)
- Implemented new CI/CD pipeline reducing feature delivery time by 50%
- Coordinated large-scale engineering efforts for SOC2 certification
- Achieved 40% reduction in cloud expenses by optimizing system designs
- Expanded GitOps + IaC coverage to improve disaster recovery resilience
Equinix
Warsaw, PolandEquinix is the leader in the colocation data center market. Joined to implement CloudNative practices on Kubernetes and automate Istio Service Mesh for GDPR-compliant connectivity.
- Led migrations of 3 systems (40 microservices) to make infrastructure GDPR compliant
- Designed architecture and migration process from on-premise to cloud
- Delivered automation of cloud infrastructure (K8s, Istio Service Mesh, observability, CD pipelines)
- Supported other teams in solving AWS, Kubernetes, and networking issues
- Mentored teammates on infrastructure code quality and design
Codility
Warsaw, PolandTechnical interview platform for testing coding skills. Joined ahead of the migration from instance-based (EC2) infrastructure to CloudNative Kubernetes.
- Kept online services and infrastructure running in good health
- Ensured infrastructure security against unauthorized access and interruptions
- Educated engineers on modern cloud computing standards and good practices
- Designed architecture for migration from EC2 to Kubernetes
- Improved CI/CD pipelines and developer tooling
F5 Networks
Warsaw, PolandF5 specializes in application security, multi-cloud management, and network security. Joined to develop a secure gateway for CloudNative services; project concluded with the NGINX acquisition.
- Implemented CloudNative services in Go
- Designed pipelines for unit, integration, and end-to-end tests with multi-project execution
- Automated packaging into Amazon Machine Images (AMI)
TouK
Warsaw, PolandSoftware house building solutions for external clients. Moved from Java development to DevOps to gain experience as infrastructure administrator and cloud solutions architect.
- ELK Stack aggregating 50 TB of data for Play Mobile - 16 nodes, 2 TB memory
- Play Now - designed scalable Cloud Native delivery pipelines and third-party integrations for a new mobile operator
- Virginmobile MVNO - 2 years of infrastructure maintenance and architectural improvements
- Migrated parts of internal infrastructure to AWS (e.g. autoscaling GitLab runners)
- Developed application stack for managing product catalog for Play Mobile operator
Risco Software
Warsaw, PolandSoftware house delivering systems for financial institutions and private companies.
- Express ELIXIR Adapter - immediate transfer system for CitiBank S.A.
- Maintained back-office systems for FM Bank and BankBPS
- Developed components of the Paymax mobile payment system
Skills
- Claude
- Kubernetes
- Istio
- Go Development
- TypeScript
- Java
- AWS
- GCP
- Cloudflare
- Datadog
- Grafana
- Prometheus
- Argo CD
- GitHub
- Terraform
- Helm
- Prometheus / Grafana
Certifications
- AI_Devs 3: Agents - AI_Devs2024
- CKAD: Certified Kubernetes Application Developer - Linux FoundationMay 2020 - May 2023
- CKA: Certified Kubernetes Administrator - Linux FoundationNov 2019 - Nov 2022
Talks & Publications
- GitOps - Czyli konfigurowanie Kubernetesa Gitem - SysOps/DevOps Warszawa MeetUp #4827.02.2020
- GitOps - Czyli konfigurowanie Kubernetesa Gitem - Confitura 201929.06.2019
- Autoscaling GitLab CI - SysOps/DevOps Warszawa MeetUp #4227.06.2019
- DNS query metrics plugin for Telegraf - open-source contribution
Education
- M.Sc. in Engineering, Computer ScienceWarsaw University of Technology · 2013 – 2016
- Engineer's degree, Computer ScienceWarsaw University of Technology · 2009 – 2013
Languages
- Polish - Native
- English - Professional working