CHILI publish - Full Stack AI/ML Engineer
Skip to content

Full Stack AI Engineer

  • Hybrid
    • Aalst, Vlaams Gewest, Belgium
  • Delivery

Join us to turn AI ideas into real product features used by our customers worldwide. Work with a curious team building smarter tools for creative automation.

Job description

We're looking for someone who'd rather build the eval that proves an agent got better than argue about vibes.

At CHILI publish, we build a cloud platform that helps brands and agencies create, automate, and scale digital content with ease. AI is becoming a key driver on our platform, powering intelligent content automation, smart templating, and the agentic systems our customers will rely on next. We're not hiring for a checkbox: we're looking for someone who wants to shape how AI gets built into a real product, used by real people, at scale, working closely with your team and directly with product management.

What you'll be doing

You'll work at the intersection of AI systems, backend engineering, and product: turning ideas and sketch requirements into AI features our customers use every day, working the way we expect the whole industry to work soon: with coding agents doing a lot of the typing, and engineers doing the deciding and the verification.

  • Design and ship agents, not just endpoints. Multi-step, tool-calling systems that plan, act, pause for the browser, recover, and report back. You'll own the whole loop: how a request is interpreted, which tools it reaches for, what runs client-side versus server-side, and what happens when step four of seven fails.

  • Treat prompts, tool catalogs and skills as engineered product surfaces. They're versioned, reviewed, tested and tuned like any other code we ship.

  • Make quality measurable. Build the evals and benchmarks that prove a change made the agent better. Non-determinism is the normal case here; your job is to turn "this feels better" into evidence humans know "it is better."

  • Build the retrieval and data plumbing that keeps agents grounded. Vector search, embedding and ingest pipelines, and the deployment work that keeps them running.

  • Own it in production. Streaming APIs, stateless and highly available services, latency and token budgets, and the structured logging and tracing that backs it up.

  • Turn fuzzy intent into shipped behaviour. You'll work directly with Product and Design, argue about what the agent should do in the cases nobody specified, write the ticket yourself, and ship it.

  • Build with agents, not just for them. A large share of the code here is written by agents. That shifts where your day goes: decomposing a fuzzy goal into something an agent can actually execute, then verifying the result against the running system.

  • Stay current with a landscape that moves weekly, and bring well-considered opinions on where we should invest next, including where we shouldn't.

Who you are

  • You're product-minded. You judge the agent's output the way a user would, not just the way a test suite does. "Technically correct" and "actually good" are two different questions, and you hold both at once.

  • You're comfortable with ambiguity. Requirements reach you as a screenshot, a hunch, or a complaint. You're happy to make the call, but you make your assumptions explicit before you build, and you ask the one question that unblocks you rather than guessing quietly.

  • You think outside the box. When the obvious approach is brittle, you find the one that isn't.

  • You take ownership end to end. From how a request is interpreted through what happens when it fails in production, you don't wait to be told something's broken — you notice, you fix it, and you're accountable for what you ship, whether you wrote it yourself or an agent did.

  • You trust evidence over assumption. Docs, comments and tickets describe intent; the running system is the truth. You verify against it, and you say so plainly when the two disagree.

  • Less is more. You'd rather delete than add, and you'd rather ship something small and clear than something clever nobody can maintain.

  • You thrive in a collaborative, international team driven by curiosity where English is the working language and diverse perspectives are genuinely valued.

Job requirements

Must-haves

  • 5+ years of professional software engineering experience, including at least 1–2 years building and operating agentic or tool-calling production systems, with attention to their reliability, latency, observability and cost characteristics.

  • Demonstrated ability to evaluate non-deterministic systems: designing test sets, benchmarks and automated graders, and distinguishing a genuine regression from grader noise.

  • Strong proficiency in TypeScript or Python for backend services (NodeJS and/or FastAPI or equivalent) and in TypeScript (React) for browser- and edge-side work.

  • Full-stack capability across the request path: you can design the API, integrate it with a frontend, and apply enough product and UX judgement to make the result intuitive for the people using it.

  • A practical command of the LLM ecosystem: prompt and context engineering, model selection and routing across providers, embeddings and retrieval, caching and tiering, and an informed view on when fine-tuning is and is not warranted.

  • Experience operating services in production: containerised deployment, readiness and warmup semantics, structured logging and tracing, and CI/CD.

  • The ability to communicate clearly: with engineers and non-engineers alike, and to record decisions in writing so that others can act on them.

Nice-to-haves

  • Exposure to content generation, creative tooling, digital asset management or marketing technology.

  • Cloud-native infrastructure experience (Azure, AWS or GCP; Kubernetes; Cloudflare Workers).

  • Experience with retrieval systems backed by a vector database (pgvector, Qdrant, Weaviate, Pinecone or similar), including embedding pipelines and index maintenance.

  • Experience serving computer-vision or other non-LLM models in production — ONNX Runtime, YOLO, OpenCV or comparable — including artifact pinning, verification and warmup.

  • Some exposure to data science or applied ML research is a plus, but not something we're specifically screening for.

How we work

We're an international team headquartered in Aalst, Belgium. Our way of working is highly collaborative — short, fast increments, tight loops between engineering, product and even early-adopter clients. That rhythm works best in person, so we're looking for someone who can be in our Aalst HQ at least 2 days a week. For that reason, we're specifically looking for candidates based in Belgium or able to commit to that regular on-site presence.

You'll be joining a small, high-leverage team: a hands-on AI lead, an AI product manager, and a CPO who are directly and closely involved in shaping our agentic product features. There's no bureaucracy between a good idea and shipping it — you'll be in the room where these decisions get made, not several layers removed from them.

Salary and overall package are competitive and discussed transparently from the first screening call.

What to expect

An initial screening with HR on job and cultural fit, a technical/team interview with our AI lead, and a final conversation with our VP of Delivery.

Curious?

If this resonates, we'd love to connect. Send us your CV and a short cover letter — tell us about an AI/ML system you're proud of building, and where you see yourself making the biggest impact at CHILI publish.

Job requirements

Must-haves:

  • At least 2 years of production experience building and operating Vector Databases (e.g. PGVector, Pinecone, Weaviate, Qdrant) and RAG architectures at scale.

  • Hands-on experience with MLOps: model deployment, versioning, monitoring, CI/CD for ML, and infrastructure tooling (e.g. MLflow, Weights & Biases, SageMaker, or similar).

  • Strong full-stack development ability - you can build the API layer, hook it up to a frontend, and know enough about UX to make it intuitive for clients. You feel comfortable building API’s in Node (TypeScript) and train, evaluate, run scripts and models in Python.

  • A solid grasp of LLM ecosystems - prompt engineering, fine-tuning trade-offs, embedding models, and how to build reliable, observable AI features in production, considering cost and performance.

  • The ability to communicate clearly about functional and technical aspects to both engineers and non-engineers.

Nice-to-haves:

  • Background in data science or applied ML research - familiarity with model evaluation, experimentation design, and statistical thinking.

  • Experience with cloud-native ML infrastructure (Azure, AWS or GCP ML- and DevOps tooling).

  • Exposure to content generation, document understanding, or creative (mar) tech - the domain we operate in.

or