Asset #122762

AI Infrastructure & Operations Engineer

Employment type
Permanent (incl. probation period)

Description:

AI Laboratory Infrastructure & Operations Engineer About Us We are an AI and Data Mining company developing. Our AI infrastructure includes NVIDIA GPU clusters, enterprise AI platforms, large-scale model training and inference environments, and intelligent AI workflow automation. We are expanding our AI Laboratory and are looking for an experienced engineer to build and operate our AI computing environment. Responsibilities You will participate in the planning, deployment, operation, and maintenance of our AI laboratory infrastructure, including: Design, deploy, and maintain AI server and GPU infrastructure Install, rack, cable, commission, and maintain AI servers in data center environments Plan and execute AI infrastructure expansion, upgrades, and hardware lifecycle management Deploy and manage Linux servers Install, configure, and maintain NVIDIA CUDA, Docker, Kubernetes, and the AI software stack Deploy and optimize AI model training and inference environments Build and maintain AI agent platforms Develop AI workflows and automation pipelines Maintain AI computing clusters and GPU infrastructure Monitor system performance and optimize GPU utilization Monitor power distribution, cooling systems, and overall hardware health Troubleshoot AI server hardware, GPU, networking, storage, operating system, and software issues Support AI researchers and software engineers Requirements Required Strong Linux system administration experience Experience with AI server hardware deployment, installation, maintenance, and troubleshooting Experience with data center infrastructure, including power distribution, rack installation, cabling, cooling, and hardware commissioning Experience deploying and maintaining NVIDIA GPU servers Experience with Docker Experience with Python Experience with CUDA Experience with AI frameworks (PyTorch / TensorFlow) Familiar with LLM deployment Experience with Git Experience with enterprise networking (10/25/100Gb Ethernet or InfiniBand) Experience assembling, upgrading and repairing enterprise servers Preferred Kubernetes Slurm Ansible GPU cluster management AI Agent Platforms LangChain MCP Vector Database RAG Workflow Automation CI/CD Nice to Have Experience with one or more: Multi-GPU training Distributed AI computing NVIDIA DGX AI Data Center Enterprise AI Infrastructure InfiniBand or other high-speed networking Storage systems (NAS / SAN) OpenAI API DeepSeek Qwen Llama We Offer Competitive salary Flexible working environment Opportunity to work with cutting-edge AI technologies Career development in AI infrastructure and intelligent systems Language English required Chinese is a strong advantage (Chinese (Mandarin) proficiency is highly preferred, as the role involves collaboration with our Chinese engineering teams and technical documentation.)AI Laboratory Infrastructure & Operations Engineer About Us We are an AI and Data Mining company developing. Our AI infrastructure includes NVIDIA GPU clusters, enterprise AI platforms, large-scale model training and inference environments, and intelligent AI workflow automation. We are expanding our AI Laboratory and are looking for an experienced engineer to build and operate our AI computing environment. Responsibilities You will participate in the planning, deployment, operation, and maintenance of our AI laboratory infrastructure, including: Design, deploy, and maintain AI server and GPU…

Overview

Type
job
Status
active
Visibility
public
City
Sollentuna
GPS
59.4293164, 17.9500556
Email
Views
5
Published
17. 7. 2026
Edited
5. 8. 2026

Location

Loading map…

Specifications

Region
Stockholms län
Duration
Ongoing
Employer
KYLG AB
Open positions
1
Profession
Infrastrukturarkitekt
Salary Type
Fixed monthly, weekly or hourly pay
Scope Of Work
100–100 %
Working Hours
Full-time