ITQuarksBook a 25-min call
HomeCareersLLMOps Engineer

Open role

LLMOps Engineer

Location:Skopje, Macedonia
Type:Full-time
Posted:23 July 2026

About the role

ITQuarks is an AI-first engineering partner. We run LLMs in production - both managed and self-hosted - for our own products (EigenVox, DocuGenius) and for enterprise clients who need cost control, low latency, or strict data residency. Open-weight models are a growing part of that footprint, and we need an engineer who owns how they are hosted, served, and kept fast.

We are hiring an LLMOps Engineer to build and operate our model-hosting and inference infrastructure: deploying foundational and fine-tuned models from the Hugging Face ecosystem, and running self-hosted inference on bare-metal and GPU infrastructure with runtimes such as llama.cpp and vLLM. This role is available full-time or part-time - we care about ownership and depth, not hours on a clock.

Responsibilities

  • Deploy and operate foundational and fine-tuned LLMs from the Hugging Face ecosystem (Transformers, TGI, Inference Endpoints)
  • Run self-hosted inference on bare-metal and GPU infrastructure with llama.cpp, vLLM, and comparable runtimes
  • Quantise and optimise models (GGUF, AWQ/GPTQ) for the target hardware - throughput, latency, and memory budgets
  • Design serving architectures: batching, caching, tensor parallelism, GPU scheduling, and autoscaling
  • Build CI/CD pipelines for model rollout, versioning, and rollback
  • Monitor inference in production - latency, cost per token, quality regressions - and run evaluation benchmarks
  • Harden model endpoints: authentication, rate limiting, network isolation, and audit logging
  • Advise on hosting decisions: managed API vs self-hosted, hardware sizing, and cost/performance trade-offs

What we're looking for

Mandatory

  • Hands-on experience deploying and serving LLMs in production - self-hosted, not only via managed APIs
  • Practical experience with llama.cpp, vLLM, or Hugging Face TGI, including model quantisation
  • Solid Linux systems and GPU infrastructure skills - drivers, CUDA, memory management, bare-metal provisioning
  • Experience with containerisation and orchestration - Docker and Kubernetes
  • Scripting and automation proficiency in Python and Bash

Preferred

  • Experience benchmarking and evaluating open-weight models (Llama, Mistral, Qwen, and similar)
  • Familiarity with inference observability and cost tracking - OpenTelemetry, Langfuse, or Prometheus/Grafana
  • Experience with fine-tuning workflows (LoRA/QLoRA) and shipping the results to production
  • Infrastructure-as-code experience - Terraform or Ansible

Apply for this role

Send your CV and a short note about why this role fits — no cover letter required.