DevOps Engineer – Infrastructure and Edge Operations
Location
North America | South America
Format
Remote, Hybrid
Tech Stack
DevOps
On behalf of our Client, Scalors is looking for a DevOps Engineer to join a remote team for a full-time position.
About Client: Our Client delivers cutting-edge software solutions for the cruise and hospitality industries, driving efficiency and reliability across mission-critical systems. Our Infrastructure & Edge Operations team ensures that our global platforms run smoothly, securely, and at scale.
Role summary:
Remote first, work from anywhere, with occasional travel to customer sites or ships as required for onboarding or incident response. Candidate must be available for on call rotations.
You will own build, release, operations, observability and deployment automation for cloud and shipboard
environments.
You will design and operate secure, resilient networks and runtime platforms that support microservices, event streaming and offline sync for guest facing systems such as POS, gangway, dining and guest services.
You will be customer facing, and able to translate shipboard constraints into reliable deployment and support practices.
You will also use AI tools and techniques to design, automate and improve our infrastructure, observability and incident operations.
Core responsibilities:
CI CD and release automation, design and maintain CI CD pipelines for multi environment deployments, including cloud, shipboard on premise and air gapped scenarios.
Infrastructure as code, author and maintain Terraform, Ansible or Pulumi templates for cloud and on premise platforms.
Kubernetes and container platforms, operate and tune Kubernetes clusters in cloud and shipboard edge nodes, manage Helm charts or Kustomize manifests and GitOps flows.
Shipboard network design and troubleshooting, configure IP networks, VLANs, routing, NAT, port forwarding, DNS, NTP and firewall rules for shipboard deployments, including satellite link and VSAT awareness.
Edge constraints and offline design, design resilient services for intermittent connectivity, offline first clients, local stores such as SQLite or Couchbase Lite, CRDT or other reconciliation strategies.
Black box microservice handling, integrate and operate against third party or opaque services using robust contract testing, timeouts and circuit breaker patterns, and observability wrappers to detect failures without requiring changes to the black box.
Observability and diagnostics, deploy and operate Prometheus, Grafana, OpenTelemetry, Jaeger, Elastic or OpenSearch, Loki and tracing to achieve end to end visibility across cloud and shipboard.
Messaging and data platforms, operate Kafka and topic management, ensure broker resiliency and correct retention and consumer behavior for shipboard replication scenarios.
Security and compliance, implement mTLS, certificate lifecycle, Vault secrets management, vulnerability scanning, container image signing, and operate according to PCI and GDPR requirements on payment and guest data flows.
Hardening and incident response, lead troubleshooting for production incidents, runbooks, on call rotations, post incident reviews and automated rollback strategies.
Automation and developer enablement, provide tools and templates that enable developers to produce production quality artifacts, and self service infrastructure for test and demo environments.
Customer engagement, collaborate with customer IT to validate IP designs, test failover scenarios and document network and deployment runbooks.
Backups and disaster recovery, implement backup and restore for stateful components, test DR in constrained network conditions.
Cost and capacity planning, monitor capacity, plan cluster sizing and network bandwidth usage, and recommend changes to reduce cost and risk.
AI driven design and operations, apply AI and machine learning tools to improve infrastructure design, incident management, observability and automation, while maintaining strict security and audit controls.
Required skills and experience:
3 plus years operating production cloud and container platforms, or equivalent experience.
Strong practical knowledge of Kubernetes, Docker and container runtimes.
Hands on experience with Terraform or other IaC tools, and configuration management such as Ansible.
Deep understanding of IP networking concepts, VLANs, routing, NAT, DNS, DHCP, bonding and firewall policies, with experience applying these on site in constrained WAN environments.
Experience operating Kafka, MySQL, Elasticsearch or OpenSearch, Redis or Hazelcast, and understanding replication and retention trade offs.
Observability and tracing experience, with Prometheus, Grafana, OpenTelemetry or Jaeger, and log pipelines in Elastic or OpenSearch or Loki.
Strong scripting and automation skills in shell, Python or Go.
Solid security fundamentals, including TLS, PKI, secrets management, container image security and familiarity with PCI compliance basics.
Excellent troubleshooting skills, including working in constrained network conditions, and customer facing communication skills.
Experience with Git based workflows, and at least one CI CD platform such as Jenkins, GitLab CI, ArgoCD or GitHub Actions.
Practical experience using AI tools and models in production workflows for infrastructure as code, observability summarization and automated runbook generation, including knowledge of model governance, data redaction and secure private deployments.