MLOps Engineer
- Location
- Abu Dhabi, Abu Dhabi Emirate, United Arab Emirates
- Country
- UAE
- Category
- Other
- Employment type
- Full-time
- Application
- External application
- Posted
- 22 Sept 2026
- Closes
- 15 Oct 2026
CV Match Score
Loading match details…
About this role
About AI71:
AI71 is an industry leader in artificial intelligence, delivering innovative solutions that empower developers, businesses and governments to solve complex challenges. AI71 builds secure, enterprise-ready applications powered by cutting-edge technology—tailored for knowledge workers and sector-specific needs. AI71 bridges the gap between advanced AI and real-world impact. Guided by a strong commitment to research and responsibility, we create transformative solutions that drive progress and empower communities.
The Role:
As an MLOps Engineer you set the ML infrastructure and reliability strategy across AI71's platform, including how LLMs and other deep learning models are deployed, fine-tuned, and served at scale. You own architecture decisions across both SaaS and on-prem operating models, mentor engineers across teams, and drive multi-quarter ML infrastructure strategy. You are a force multiplier.
What You'll Do:
* Define ML infrastructure architecture across the platform: model deployment strategy (vLLM, Triton, or TGI), pipeline engineering (MLflow or Kubeflow), and cloud-native infrastructure across major cloud platforms (AWS, Azure, or GCP)
* Set direction for ML system reliability: monitoring, latency / throughput / availability targets, and incident response across research and production environments.
* Mentor senior MLOps engineers; raise the operational bar across multiple teams.
* Drive cross-team initiatives that improve inference performance and cost-efficiency, including distributed training frameworks (DeepSpeed, FSDP, Accelerate).
* Partner with ML researchers, product, and engineering leadership on multi-quarter ML infrastructure strategy.
* Ensure ML infrastructure scales across managed SaaS and fully air-gapped on-prem deployments.
What You'll Bring:
* 10+ years of MLOps, ML infrastructure, or machine learning engineering with history of architectural ownership.
* Proven track record architecting large-scale model deployment (including LLMs) and ML infrastructure at scale.
* Deep cloud expertise across major cloud platforms (AWS, Azure, or GCP) and strong Python proficiency
* Mentorship record — engineers you have grown now operate independently at higher levels.
* Deep comfort architecting ML systems that run in both managed SaaS and on-premises / disconnected air-gapped environments.
* Kubernetes at architectural depth — GPU scheduling, multi-tenancy, operators, and the failure modes of distributed workloads on shared clusters.
* Strong communication, stakeholder management, and decision-making skills, with a passion for building diverse, inclusive engineering teams.
Strong Preference:
* Ownership of production reliability at platform level: SLO definition, incident command, postmortem practice, and driving reliability improvements across teams rather than services.
* Architecture-level experience with distributed training and fine-tuning at scale (DeepSpeed, FSDP, Megatron-LM), including cluster design, checkpointing strategy, and failure recovery.
* Deep GPU systems knowledge: CUDA, NCCL, interconnect topology (NVLink, InfiniBand/RoCE), and diagnosing performance and communication problems across nodes.
* Model optimization strategy at portfolio level: quantization (FP8, AWQ, GPTQ), speculative decoding, with measurable cost or latency outcomes across multiple systems.
* Experience in regulated or security-constrained environments — compliance-driven architecture, model governance, lineage, audit, and secrets management.
* On-prem / air-gap ML delivery architecture experience at scale.
* Track record maturing MLOps practice in a growing organization: standards, platform abstractions, and paved paths that outlived your involvement.
* Bare-metal GPU cluster architecture, including scheduling (Slurm or Kubernetes) and hardware lifecycle in customer or owned data centers.
Nice to Have:
* Conference speaking, technical writing, or industry thought leadership.
* Open-source contributions to inference, serving, or ML infrastructure projects, particularly maintainer-level involvement.
* C/C++ or CUDA kernel experience for performance-critical paths.
* Arabic language skills.
Why AI71:
* Mission-Driven Work: Work on cutting-edge AI applications with a talented and passionate team, solving real-world challenges in critical sectors.
* Unparalleled Opportunity: This is a chance to innovate and solve real-world challenges using AI at a company with unique access to world-leading models and resources.
* Career Growth: We offer competitive compensation, benefits, and significant career growth opportunities as a foundational member of the team.
* World-Class Environment: Enjoy a flexible working environment and the latest tools & technologies needed to do your best work.