Service

AI + ML

We integrate AI into business software — RAG (retrieval-augmented generation) over your data, AI agents that take real actions, structured-output extraction from PDFs / forms / emails, classification + scoring pipelines. Our internal use cases include the Merot Finance AI bookkeeping assistant, lead scoring in Merot Leads, and invoice OCR.

Concrete deliverables

  • RAG over your knowledge base — internal docs / Notion / Confluence / Slack archives → searchable AI assistant.
  • Structured extraction from documents — invoices, contracts, forms, expense reports. JSON output you can store in your DB.
  • AI agents that take actions — book meetings, draft emails, run database queries, post to Slack.
  • Classification + scoring — lead scoring, fraud detection, sentiment, content moderation.
  • Embeddings + search — semantic search over your product catalog, support tickets, code repository.
  • On-premises model serving — when data privacy means cloud APIs are off-limits. Llama 3, Mixtral, fine-tuned smaller models.

What we work with

We pick what fits your team. We don't push our preferences when your stack works.

API providers

OpenAI · Anthropic Claude · Google Gemini · Mistral AI · Cohere

Frameworks

LangChain (sometimes) · LlamaIndex · Vercel AI SDK · Anthropic SDK · OpenAI SDK

Vector databases

pgvector (Postgres) · Pinecone · Qdrant · Weaviate · Chroma

Self-hosted models

Llama 3 (8B-70B) · Mistral 7B / Mixtral 8x7B · Whisper (speech-to-text) · Stable Diffusion (image)

Inference infra

AWS Bedrock · Replicate · Together AI · self-hosted via vLLM / TGI

Eval + observability

LangSmith · Helicone · OpenAI usage dashboards · custom eval harnesses

How we work

01

Discovery (1 week)

Define the use case crisply: what input → what output. Write evals up front so we know what 'working' means.

02

Prototype (1-2 weeks)

Smallest possible LLM call that produces the desired output. Measure on the eval set. Decide: API or self-hosted, which model, what prompt.

03

Production (3-6 weeks)

Wrap with retries, fallbacks, observability, cost controls (token budgets, rate limits). Wire into your product.

04

Iterate

AI features need ongoing eval as models change. Monthly retainer or scheduled quarterly tune-up.

From our own production

Merot Finance AI assistant

Anthropic Claude integrated for bank-statement matching, journal-entry suggestions, and month-close review.

Merot Leads scoring

Claude for product-fit scoring on enrichment + outreach drafts. Custom prompt + structured-output extraction.

Invoice OCR pipeline

Multi-stage extraction: OCR → LLM structured output → human-review queue for low-confidence items.

Nearshore delivery by buyer market

Country pages explain the trust, timezone, compliance and delivery questions buyers ask before choosing a software partner.

North Macedonia

Custom software development for Macedonian companies: web apps, mobile apps, cloud systems, AI automation and integrations from Merot's Skopje engineering teams.

Germany

Senior Balkan software teams for German companies: web, mobile, AI, cloud and integrations with EU timezone overlap, GDPR-aware delivery and DACH communication.

Austria

Nearshore web, mobile, cloud and AI engineering for Austrian SMEs and product teams. EU timezone, DACH communication, GDPR-aware delivery.

Netherlands

Senior React, mobile, backend, cloud and AI teams for Dutch SaaS and operations companies. English-first delivery, EU timezone and GDPR-aware workflows.

Switzerland

Senior Balkan software teams for Swiss companies: secure web, backend, cloud, AI and integrations with DACH communication and privacy-aware delivery.

Sweden

Nearshore web, mobile, backend, cloud and AI engineering for Swedish SaaS and operations teams. Senior delivery, EU timezone, clean handoff.

Denmark

Senior nearshore engineers for Danish product and operations teams: web, backend, mobile, cloud, AI and integrations with EU timezone overlap.

Ireland

Senior EU nearshore engineers for Irish SaaS, fintech and operations teams. Web, backend, mobile, AI and cloud delivery with English-first collaboration.

United States

Senior European software engineers for US companies: custom web apps, mobile apps, backend systems, cloud, AI workflows and integrations with practical US overlap.

Canada

Custom web, mobile, backend, cloud and AI engineering for Canadian companies. Senior European delivery, practical timezone overlap and clean handoff.

United Kingdom

Senior nearshore software engineers for UK companies: custom web apps, backend platforms, cloud modernization, AI workflows and integrations.

France

Nearshore product engineering for French companies: web apps, mobile apps, AI workflows, cloud systems and integrations with EU timezone delivery.

Belgium

Nearshore web, backend, cloud, mobile and AI engineering for Belgian companies operating across languages, regions and EU compliance expectations.

Engagement model

AI pricing depends on workflow complexity, data quality, model choice, guardrails, eval requirements, integrations, expected volume and whether the work is prototype or production. Book a scoping call; we will separate engineering effort from model/API spend so the quote is clear.
Book a scoping call

Frequently asked questions — AI + ML

Should I use OpenAI, Anthropic, or self-host?

Default: start with Anthropic (Claude 3.5 Sonnet / Claude 4) or OpenAI (GPT-4 family) for the prototype. Switch to self-hosted only when (a) data residency requires it, or (b) the per-call cost exceeds engineering+infra cost of running yourself. Most clients stay on the API providers for years.

Will my data train someone's model?

Not on the enterprise tiers of OpenAI / Anthropic / Google — they have explicit no-training-on-customer-data terms. We turn on those settings during onboarding.

What if the AI hallucinates / produces wrong output?

Two layers: (1) Eval harness — we measure correctness on a labelled test set before shipping and again on every prompt change. (2) Production — high-confidence outputs go straight through; low-confidence outputs go to a human-review queue.

Cost — won't this get expensive?

It can if scope is vague. We control that by defining the workflow, setting model budgets, adding usage alerts, and deciding what must be automated versus what should stay human-reviewed. The estimate depends on the actual use case.

Do you do fine-tuning?

Sometimes — usually only when the prompt approach truly can't get there. Fine-tuning has higher up-front cost (curating training data) and re-tuning maintenance every model upgrade. Typically we recommend better prompts + RAG first.

Privacy / on-premises only — can you do that?

Yes. We've deployed Llama 3 70B and Mixtral 8x22B on-premise (single-GPU H100 or 4xA100 setups) for clients in regulated industries. Higher up-front cost, lower per-call cost, full data residency.

AI agents — are these real yet?

Cautiously yes. Single-purpose agents (book a meeting, draft an email, run a SQL query) work well with proper guardrails. Generic 'do anything' agents are still flaky. We scope to single-purpose by default.

Voice / speech?

Whisper for speech-to-text, ElevenLabs / OpenAI TTS for synthesis. We've built call-summary + voice-note transcription features for clients in the legal + healthcare verticals.

Let's scope your ai + ml project

Free 60-min discovery call. 6-page written brief in 48 hours.

Contact

Tell us what you are building.

Reply during business hours.

No spam.Directly to the Merot team.