
Senior DevOps Engineer (On Premises Kubernetes) - MF
- Hybrid
- Melrose Arch, Gauteng, South Africa
- Midrand, Gauteng, South Africa
+1 more- Cloud
Join as a Senior DevOps Engineer (Azure) to build scalable cloud solutions, automate CI/CD, and work on impactful projects while mentoring teams and driving innovation.
Job description
DVT is one of the top software development companies on the continent. Our software engineers are consulting on cutting edge applications at top companies in South Africa, as well as consulting globally. You will have the opportunity to work alongside some of the most established developers in the country and globally with the latest technologies.
We are seeking an experienced DevOps Engineer with deep expertise in Kubernetes and on-premises infrastructure. This role is responsible for the end-to-end ownership of our Kubernetes platform, ensuring reliability, scalability, security, and operational excellence across development, test, and production environments.
The successful candidate will play a hands-on role in designing, implementing, automating, and supporting enterprise infrastructure running within our data centres. You will work closely with development teams to improve deployment processes, platform stability, observability, and developer productivity while building and maintaining Infrastructure as Code and CI/CD pipelines.
Job requirements
DUTIES AND RESPONSIBILITIES
Kubernetes Platform Ownership
Manage, maintain, and upgrade production Kubernetes clusters across on-premises environments.
Implement and enforce resource quotas, RBAC, network policies, and namespace governance.
Manage worker node lifecycle, capacity planning, and cluster scaling strategies.
Own Kubernetes cluster architecture, platform standards, and operational best practices.
Manage workloads and cluster lifecycle through Rancher or equivalent cluster management platforms.
Build, maintain, and optimise Helm charts and deployment templates across environments.
Ensure platform security, compliance, and resilience through proactive maintenance and governance.
Infrastructure & Automation
Design, implement, and maintain Infrastructure as Code using Terraform, Ansible, or similar technologies.
Automate server provisioning, configuration management, patching, and deployment processes.
Support Linux-based infrastructure, virtualisation platforms, storage systems, and networking components.
Work with infrastructure teams to ensure high availability and disaster recovery readiness.
Reliability & Operations
Define, monitor, and improve Service Level Indicators (SLIs) and Service Level Objectives (SLOs).
Participate in on-call support, incident response, root cause analysis, and post-incident reviews.
Identify and eliminate single points of failure across infrastructure and applications.
Drive continuous improvement initiatives to improve platform availability and performance.
Observability & Monitoring
Maintain and enhance monitoring, logging, and alerting platforms such as Prometheus, Grafana, Loki, ELK, or equivalent solutions.
Develop meaningful alerts and dashboards to improve operational visibility.
Support application performance monitoring and distributed tracing where appropriate.
Reduce alert fatigue through tuning and optimisation of monitoring systems.
CI/CD & Release Engineering
Build and maintain Azure DevOps or equivalent CI/CD pipelines for application deployment.
Implement automated testing, security scanning, and deployment controls within delivery pipelines.
Support development teams with deployment automation, troubleshooting, and platform enablement.
Promote DevOps best practices across development and operations teams.
Capacity, Performance & Optimisation
Monitor platform utilisation and forecast future infrastructure requirements.
Optimise workload placement, resource consumption, and infrastructure efficiency.
Provide recommendations to improve performance, scalability, and operational effectiveness.
Required Experience and Skills
Must-have
5+ years of experience in DevOps, Platform Engineering, Site Reliability Engineering, or Infrastructure Engineering roles.
Strong hands-on experience administering production Kubernetes environments.
Deep understanding of Kubernetes architecture, networking, storage, ingress, RBAC, and cluster operations.
Experience managing on-premises infrastructure and enterprise server environments.
Strong Linux administration skills.
Experience with Infrastructure as Code tools such as Terraform and Ansible.
Experience building and supporting CI/CD pipelines.
Proficiency in Bash, Python, Go, or similar scripting languages.
Strong networking knowledge including TCP/IP, DNS, routing, TLS, load balancing, and firewalls.
Experience supporting highly available and mission-critical environments.
Proven experience with incident management, root cause analysis, and operational support.
Advantageous
Certified Kubernetes Administrator (CKA) or equivalent certification.
Experience with Rancher for Kubernetes management.
Strong Helm chart development and release management experience.
Experience with VMware, Nutanix, Hyper-V, or other virtualisation platforms.
Experience with enterprise storage technologies and backup solutions.
Experience implementing observability platforms from the ground up.
Knowledge of GitOps tools such as ArgoCD or Flux.
Windows Server administration experience.
Exposure to cloud platforms such as Azure, AWS, or GCP, although this is not a primary requirement.
Who we are:
or
All done!
Your application has been successfully submitted!
You've already applied for this job
We appreciate your interest in this position. Unfortunately, you have already applied for this job.
