We need a senior Linux/Kubernetes platform engineer to build and document a secure, production-ready on-premises Kubernetes environment on three Dell PowerEdge servers. The platform will host a migration of our observability software.
We need an engineer with demonstrated experience designing, deploying, securing, and supporting bare-metal Kubernetes clusters running stateful production applications.
Current environment
Three Dell PowerEdge physical servers
RHEL 10 required
On-premises environment
Kubernetes-based hosting for Parlon observability software
Existing network/security team available for coordination on VLANs, firewall rules, DNS, certificates, and access requirements
Scope of work
Assess Dell server hardware, firmware, RAID/storage layout, NICs, and iDRAC configuration
Design the target architecture for a three-server production Kubernetes deployment
Install and harden RHEL 10 on all servers
Configure OS baseline: subscriptions/repos, SELinux, firewall, time synchronization, storage, logging, patching, and secure administrative access
Deploy and configure Kubernetes using a supportable approach; recommend kubeadm versus Red Hat OpenShift based on Parlon compatibility, licensing, operations, and supportability
Configure container runtime, CNI, ingress/load balancing, DNS, TLS, RBAC, NetworkPolicies, and cluster security
Design and implement persistent storage, backups, retention, capacity management, and restore testing for observability data
Deploy/migrate Parlon in coordination with the software vendor or our internal team
Validate application availability, data ingestion, dashboards/queries, alerting, performance, and failover/restart behavior
Produce complete documentation and conduct a knowledge-transfer session
Required experience
5+ years administering enterprise Linux, including Red Hat Enterprise Linux
Hands-on RHEL 9/10 experience in production
3+ years building and operating production Kubernetes clusters
Strong bare-metal/on-prem Kubernetes experience; cloud-only Kubernetes experience is not sufficient
Kubernetes administration: kubeadm and/or Red Hat OpenShift, CRI-O/containerd, Helm, RBAC, upgrades, backup/restore, and troubleshooting
CNI/networking expertise: Cilium or Calico preferred; NetworkPolicies, ingress, load balancing, DNS, MTU, routing, and firewall troubleshooting
Dell PowerEdge experience: iDRAC, RAID/HBA, BIOS/firmware, hardware monitoring, NIC bonding, VLANs
Stateful workload and persistent-storage experience: CSI, NFS/iSCSI/Ceph/Longhorn or equivalent
Security experience: SELinux, firewalld/nftables, TLS, secrets, image security, least privilege, hardening, and patch management
Automation using Ansible and Git; infrastructure documentation is mandatory
Excellent written English and ability to provide clear runbooks/diagrams
Strongly preferred
RHCE, RHCA, Red Hat OpenShift certification, CKA, and/or CKS
Experience deploying observability platforms such as Prometheus, Grafana, Loki, OpenTelemetry, Elasticsearch/OpenSearch, ClickHouse, VictoriaMetrics, or similar
Experience migrating a stateful application from an existing server/platform into Kubernetes
Experience integrating Kubernetes with enterprise DNS, PKI, load balancers, SIEM, and monitoring systems
Experience with Cilium/Hubble or advanced Kubernetes network observability
Deliverables
Architecture/design document, including cluster topology, network diagram, storage design, IP/port matrix, and capacity assumptions
Secure RHEL 10 build on all three Dell servers
Fully working Kubernetes platform with health checks and documented operational procedures
Persistent storage, backup policy, and a successful restore test
Parlon deployment/migration plan, validation plan, cutover plan, and rollback plan
Infrastructure-as-code/configuration artifacts in Git, preferably Ansible plus Helm/Kustomize manifests
Operations runbook: patching, upgrades, certificate renewal, node replacement, backup/restore, incident triage, and escalation paths
Recorded or live knowledge-transfer session with our technical team