MLOps & AI Platform Engineer

About the Role

We are looking for an MLOps & AI Platform Engineer to build and evolve the infrastructure, tooling, and automation that enable scalable AI solutions in production.

In this role, you will create the foundation that empowers AI and data teams to develop, deploy, and operate machine learning and generative AI applications efficiently and reliably. You will help establish best practices, automate workflows, and ensure platform stability, performance, and security.

Key Responsibilities

  • Design, implement, and maintain scalable platforms supporting the full machine learning lifecycle.
  • Build and optimize infrastructure for training, deploying, and serving machine learning and generative AI models.
  • Develop CI/CD pipelines and automation workflows for AI and ML deployments.
  • Establish standards for model versioning, reproducibility, packaging, and deployment.
  • Manage and optimize compute resources, including GPU-based environments and inference services.
  • Implement monitoring, logging, and observability solutions for AI systems.
  • Ensure platform reliability, scalability, security, and operational excellence.
  • Collaborate closely with data scientists, AI engineers, software developers, and infrastructure teams.
  • Evaluate and integrate emerging technologies within the MLOps and Generative AI ecosystem.
  • Contribute to internal documentation, knowledge sharing, and platform adoption across technical teams.

Requirements

  • Proven experience in MLOps, DevOps, Platform Engineering, Data Engineering, or related fields.
  • Strong programming skills in Python.
  • Hands-on experience with Docker, Kubernetes, and modern CI/CD tools such as GitHub Actions, Azure DevOps, or Jenkins.
  • Experience with ML lifecycle management platforms such as MLflow or equivalent solutions.
  • Familiarity with cloud-based data and AI platforms, including Databricks or similar technologies.
  • Experience supporting machine learning and generative AI workloads in production environments.
  • Knowledge of GPU infrastructure, model serving, and inference optimization techniques.
  • Understanding of monitoring, logging, tracing, and observability practices.
  • Knowledge of security, governance, and compliance principles in enterprise environments.
  • Ability to balance innovation, scalability, and operational stability.
  • Strong communication, collaboration, and problem-solving skills.

Nice to Have

  • Experience with LLM deployment frameworks such as vLLM, Triton Inference Server, Hugging Face TGI, or similar solutions.
  • Familiarity with vector databases and Retrieval-Augmented Generation (RAG) architectures.
  • Experience with distributed computing frameworks and large-scale AI infrastructure.
  • Exposure to agentic AI frameworks and orchestration platforms.

What We Offer

  • Opportunity to work on cutting-edge AI and platform engineering initiatives.
  • Access to modern technologies, cloud platforms, GPU infrastructure, and large-scale AI environments.
  • International collaboration with experienced AI, data, and engineering professionals.
  • Continuous learning and professional growth opportunities.
  • Flexible and supportive work environment.
  • Challenging projects with high technical impact and ownership.
  • A culture that encourages innovation, experimentation, and continuous improvement.