Senior Data Engineer

Translated
On-site
Morocco , Boulemane
--

Job Details

Job Description

Roles & Responsibilities

As part of strengthening our Data teams, we are looking for profiles capable of designing, industrializing and optimizing data platforms (batch & real-time) within distributed environments based on Cloudera.

Your responsibilities:

Development & Industrialization

  • Develop massive data processing pipelines in PySpark (Batch and Real-Time / Streaming modes).
  • Set up real-time streams via Kafka (topics, partitions, schemes, offsets) and ensure ingestion with NiFi.
  • Model and optimize NoSQL schemas, especially on Cassandra (tables, keys, clustering, replication) and Hive.
  • Integrate and transform data from multiple sources (APIs, databases, streams, files).

Quality, Performance & Reliability

  • Deploy Data Quality mechanisms (controls, monitoring, alerting).
  • Optimize Spark processing (partitioning, tuning, data formats) specifically for distributed architectures.
  • To ensure the supervision and resolution of incidents in a production environment.

CI/CD & Governance

  • Industrialize developments via CI/CD chains (automated testing, deployments).
  • Documenting flows, models and best practices.
  • Collaborating occasionally with third-party teams on Dataviz topics and collaborating with Data Scientists (particularly on customer segmentation algorithms).
  • Contribute to data governance (catalogue, traceability, security).

Desired Candidate Profile

Qualifications

Experience: ~4 years of experience in distributed environments and Big Data architectures.

Architecture: Strong interest and significant experience in On-Premise infrastructures.

Spark / PySpark: Essential mastery in Batch and Streaming processing.

NoSQL: Confirmed experience on Cassandra and/or Hive

Streaming & Ingestion: Very good mastery of Kafka and NiFi

Nice-to-have / Bonus:

  • Good understanding of the Hadoop/HDFS ecosystem. Mastery of Cloudera is a plus but remains optional.
  • Concepts in Data Science (models of customer clustering/segmentation) and Dataviz tools.
  • Tools & DevOps: Skills in Git, CI/CD (GitLab CI) and an orchestration tool (Airflow, Luigi, Prefect).

Similar Jobs