DALIL INFORMATION TECHNOLOGY jobs
4 Jobs Found
<p>bout the Platform</p><p>Project: Morocco Steady-State Operations</p><p>Environment: Government-grade, on-premises bare-metal infrastructure. Three security-zone Kubernetes clusters (DMZ, Core, Red Zone) underpinned by a Fortinet security fabric (FortiGate 901G HA pair, FortiSwitch, FortiManager, FortiPAM, FortiWeb WAF, FortiADC, FortiAnalyzer) and Dell PowerEdge R670 compute nodes with Ceph distributed storage. Backup is delivered via Dell PowerProtect.</p><p>Operations model: The platform has been deployed and accepted. These roles cover Day 2 steady-state operations — monitoring, incident response, change management, backup validation, and continuous improvement. No deployment or project delivery responsibilities.</p><p>Role Overview</p><p>The Infrastructure & Platform Operations Engineer is the primary Day 2 owner of the compute, Kubernetes, storage, and backup layer of the platform. You maintain cluster health, respond to node and workload incidents, manage the backup and recovery lifecycle, and ensure the observability pipeline is accurate and actionable.</p><p>You work alongside the Network & Security Operations Engineer, who owns the Fortinet fabric. Together, the two roles provide full-stack operational coverage. When a platform incident has a network/security dimension — for example, a pod failing to reach an external endpoint — you collaborate across the boundary and escalate correctly.</p><p>Day-to-Day Responsibilities</p><p>Kubernetes Cluster Operations</p><ul><li><p>Monitor health of all three Kubernetes clusters (DMZ, Core, Red Zone): node readiness, control plane component status, etcd health, and API server responsiveness.</p></li><li><p>Investigate and resolve pod scheduling failures, CrashLoopBackOff events, OOMKill incidents, and persistent volume claim binding issues.</p></li><li><p>Manage node maintenance: cordon, drain, perform maintenance (firmware, OS patch, hardware swap), and return nodes to service.</p></li><li><p>Monitor and enforce namespace resource quotas, limit ranges, and network policy posture — flag and remediate any drift from the accepted baseline.</p></li><li><p>Manage RBAC: onboard/offboard cluster users and service accounts, review permissions quarterly, maintain least-privilege posture.</p></li><li><p>Renew and rotate Kubernetes cluster certificates, ingress TLS certificates, and any expiring secrets before expiry.</p></li><li><p>Coordinate with the application team on deployment readiness: validate ingress routing, service endpoints, and resource availability before application releases.</p></li></ul><p>Ceph Storage Operations</p><ul><li><p>Monitor Ceph cluster health dashboard daily: OSD status, PG states, replication factor compliance, and capacity utilisation against defined thresholds.</p></li><li><p>Respond to OSD failure alerts: investigate disk health (SMART data), replace failed OSDs, and monitor rebalancing to completion before any second failure.</p></li><li><p>Monitor Ceph pool capacity and raise alerts when utilisation exceeds planned thresholds — coordinate capacity planning reviews with the client.</p></li><li><p>Validate PVC provisioning for new workloads: confirm correct pool assignment, access mode, and reclaim policy.</p></li></ul><p>Dell Compute & Node Operations</p><ul><li><p>Monitor Dell PowerEdge R670 nodes via iDRAC: hardware alerts (PSU, DIMM, NIC, disk, thermal), BIOS event logs, and remote console readiness.</p></li><li><p>Manage firmware lifecycle: track available BIOS, iDRAC, NIC, and PERC firmware updates; schedule and apply via approved maintenance windows.</p></li><li><p>Maintain Linux OS baseline on all Kubernetes nodes: security patches, kernel updates, containerd and kubelet version alignment with the cluster compatibility matrix.</p></li><li><p>Maintain hardware inventory register: track serial numbers, firmware versions, and warranty status per node.</p></li></ul><p>Backup & Recovery Operations</p><ul><li><p>Monitor Dell PowerProtect backup jobs daily: confirm job completion, investigate and remediate failures, and maintain the backup success rate SLA.</p></li><li><p>Execute and document periodic restore tests (minimum quarterly): restore at least one VM and one data scenario, record restore time, verify data integrity.</p></li><li><p>Manage backup policy changes through change control: scope adjustments, retention updates, and schedule changes.</p></li><li><p>Maintain the restore runbook with current, tested procedures — update after every restore test.</p></li></ul><p>Observability & Alerting</p><ul><li><p>Maintain the observability pipeline: ensure log collection, metrics scraping, and alerting are functioning across all nodes and clusters.</p></li><li><p>Triage platform-level alerts: distinguish noise from genuine incidents, tune alerting thresholds, and escalate per the defined severity model.</p></li><li><p>Produce weekly platform health summaries: cluster utilisation, backup status, storage capacity, and outstanding incidents.</p></li></ul><br><p><strong>Desired Candidate Profile</strong></p><p>Must-Have Requirements</p><ul><li><p>Kubernetes operations experience in a production environment: node management, workload troubleshooting, RBAC, PVC/storage, certificate rotation. CKA required.</p></li><li><p>Linux systems administration: OS patching, kernel management, systemd, journald, process and disk troubleshooting.</p></li><li><p>Ceph or comparable distributed storage operations: OSD monitoring, failure response, capacity management.</p></li><li><p>Backup platform operations (Dell PowerProtect, Veeam, or equivalent): job monitoring, failure remediation, restore testing and evidence production.</p></li><li><p>Bare-metal server operations: iDRAC or iLO monitoring, firmware management, hardware fault response.</p></li><li><p>Observability tooling: experience with Prometheus, Grafana, Loki, or equivalent for cluster-level monitoring.</p></li><li><p>On-site availability in Morocco on a permanent or long-term basis.</p></li></ul><p>Good-to-Have</p><ul><li><p>CKS (Certified Kubernetes Security Specialist) — strongly preferred given the government security classification of the platform.</p></li><li><p>Dell PowerEdge-specific experience (PowerEdge R670 or R-series equivalent, iDRAC 9/10, OMSA).</p></li><li><p>Dell PowerProtect-specific experience vs generic backup platform knowledge.</p></li><li><p>Kubernetes upgrade experience: minor version upgrades of control plane and worker nodes on a live cluster.</p></li><li><p>Familiarity with Fortinet/network layer concepts (enough to collaborate on cross-domain incidents with the Network & Security Ops Engineer).</p></li><li><p>French or Arabic language capability.</p></li></ul>
<p>Job Description: Network & Security Operations Engineer</p><p>Location</p><p>On-site — Kingdom of Morocco</p><p>Role Type</p><p>Permanent / Long-term — Day 2 Operations (BAU)</p><p>Certification</p><p>NSE 5 minimum — NSE 7 / FCSS preferred</p><p>Clearance</p><p>Government site security clearance required</p><p>About the Platform</p><p>Project: Morocco Steady-State Operations</p><p>Environment: Government-grade, on-premises bare-metal infrastructure. Three security-zone Kubernetes clusters (DMZ, Core, Red Zone) underpinned by a Fortinet security fabric (FortiGate 901G HA pair, FortiSwitch, FortiManager, FortiPAM, FortiWeb WAF, FortiADC, FortiAnalyzer) and Dell PowerEdge R670 compute nodes with Ceph distributed storage. Backup is delivered via Dell PowerProtect.</p><p>Operations model: The platform has been deployed and accepted. These roles cover Day 2 steady-state operations — monitoring, incident response, change management, backup validation, and continuous improvement. No deployment or project delivery responsibilities.</p><p>Role Overview</p><p>The Network & Security Operations Engineer is the primary Day 2 owner of the entire Fortinet security fabric protecting the platform. You maintain service continuity, respond to security events and network incidents, manage configuration changes through the change control process, and ensure the logging and audit trail infrastructure is operational at all times — a compliance-critical requirement for this government platform.</p><p>You work alongside the Infrastructure & Platform Operations Engineer, who owns the compute, Kubernetes, and backup layer. Together, the two roles provide full-stack operational coverage with clear domain ownership.</p><p>Day-to-Day Responsibilities</p><p>Monitoring & Incident Response</p><ul><li><p>Monitor FortiAnalyzer dashboards and alert feeds for security events, anomalous traffic, and policy violations across all zones.</p></li><li><p>Triage WAF alerts from FortiWeb: distinguish false positives from genuine threats, tune policies, and escalate confirmed incidents through the defined severity ladder.</p></li><li><p>Monitor FortiGate HA pair health: failover readiness, session table, CPU/memory baselines, and interface status.</p></li><li><p>Respond to FortiSwitch alerts: port flaps, uplink failures, spanning-tree events, and VLAN mismatches.</p></li><li><p>Own the first-response runbook for network and security incidents — document, contain, and escalate per the agreed SLA thresholds.</p></li></ul><p>Change & Configuration Management</p><ul><li><p>Process and implement firewall rule change requests through the change control board: validate business justification, implement on FortiManager, push to FortiGate, and produce evidence.</p></li><li><p>Manage VLAN and port changes on FortiSwitch: coordinate with the application and infrastructure teams to ensure correct zone placement.</p></li><li><p>Maintain FortiPAM user roster: onboard new privileged accounts, offboard leavers, enforce least-privilege RBAC, and produce monthly session audit reports.</p></li><li><p>Track and apply FortiOS, FortiSwitch, FortiWeb, and FortiManager firmware updates through the approved maintenance window process.</p></li><li><p>Renew and manage SSL/TLS certificates on FortiADC and FortiWeb ahead of expiry.</p></li></ul><p>Compliance & Audit</p><ul><li><p>Validate daily that log ingestion is functioning across all Fortinet components into FortiAnalyzer — any gap is a compliance incident.</p></li><li><p>Generate periodic access, session, and policy-change audit reports for security team consumption.</p></li><li><p>Conduct quarterly firewall rule reviews: identify stale, overly permissive, or undocumented rules and raise change requests for remediation.</p></li><li><p>Support any external security assessments or audits with evidence packs from FortiAnalyzer and FortiManager.</p></li></ul><p>Documentation & Continuous Improvement</p><ul><li><p>Maintain as-built accuracy: update network topology diagrams, VLAN registers, and IP plans after any change.</p></li><li><p>Update and improve operational runbooks based on incident learnings.</p></li><li><p>Identify recurring alert patterns and propose WAF tuning or firewall policy improvements to reduce noise.</p></li></ul><p>Must-Have Requirements</p><ul><li><p>FortiGate operations experience in a production environment: policy management, rule review, HA monitoring, firmware upgrades.</p></li><li><p>FortiManager operations: config push, managed-device monitoring, RBAC, audit trail review.</p></li><li><p>FortiAnalyzer or equivalent SIEM: log ingestion monitoring, security event investigation, report generation.</p></li><li><p>VLAN and switching operations: port changes, trunk management, VLAN troubleshooting.</p></li><li><p>NSE 5 minimum — NSE 7 / FCSS preferred. Certificate copy required at interview.</p></li><li><p>Experience operating network/security infrastructure in a government, critical infrastructure, or regulated environment.</p></li><li><p>Familiarity with PAM platforms: session monitoring, privileged account lifecycle.</p></li><li><p>On-site availability in Morocco on a permanent or long-term basis.</p></li></ul><br><p><strong>Desired Candidate Profile</strong></p><p>Good-to-Have</p><ul><li><p>FortiWeb WAF operations: signature updates, alert tuning, false-positive management.</p></li><li><p>FortiPAM-specific experience (or CyberArk / BeyondTrust in a comparable role).</p></li><li><p>FortiADC or equivalent ADC/load-balancer operations (F5 LTM, Citrix NetScaler).</p></li><li><p>SSL certificate lifecycle management across multiple services.</p></li><li><p>Familiarity with API/PNR data classification requirements or aviation/border security data handling.</p></li><li><p>French or Arabic language capability.</p></li></ul>