AI + ML
We integrate AI into business software — RAG (retrieval-augmented generation) over your data, AI agents that take real actions, structured-output extraction from PDFs / forms / emails, classification + scoring pipelines. Our internal use cases include the Merot Finance AI bookkeeping assistant, lead scoring in Merot Leads, and invoice OCR.
Concrete deliverables
- RAG over your knowledge base — internal docs / Notion / Confluence / Slack archives → searchable AI assistant.
- Structured extraction from documents — invoices, contracts, forms, expense reports. JSON output you can store in your DB.
- AI agents that take actions — book meetings, draft emails, run database queries, post to Slack.
- Classification + scoring — lead scoring, fraud detection, sentiment, content moderation.
- Embeddings + search — semantic search over your product catalog, support tickets, code repository.
- On-premises model serving — when data privacy means cloud APIs are off-limits. Llama 3, Mixtral, fine-tuned smaller models.
What we work with
We pick what fits your team. We don't push our preferences when your stack works.
API providers
OpenAI · Anthropic Claude · Google Gemini · Mistral AI · Cohere
Frameworks
LangChain (sometimes) · LlamaIndex · Vercel AI SDK · Anthropic SDK · OpenAI SDK
Vector databases
pgvector (Postgres) · Pinecone · Qdrant · Weaviate · Chroma
Self-hosted models
Llama 3 (8B-70B) · Mistral 7B / Mixtral 8x7B · Whisper (speech-to-text) · Stable Diffusion (image)
Inference infra
AWS Bedrock · Replicate · Together AI · self-hosted via vLLM / TGI
Eval + observability
LangSmith · Helicone · OpenAI usage dashboards · custom eval harnesses
How we work
Discovery (1 week)
Define the use case crisply: what input → what output. Write evals up front so we know what 'working' means.
Prototype (1-2 weeks)
Smallest possible LLM call that produces the desired output. Measure on the eval set. Decide: API or self-hosted, which model, what prompt.
Production (3-6 weeks)
Wrap with retries, fallbacks, observability, cost controls (token budgets, rate limits). Wire into your product.
Iterate
AI features need ongoing eval as models change. Monthly retainer or scheduled quarterly tune-up.
From our own production
Merot Finance AI assistant
Anthropic Claude integrated for bank-statement matching, journal-entry suggestions, and month-close review.
Merot Leads scoring
Claude for product-fit scoring on enrichment + outreach drafts. Custom prompt + structured-output extraction.
Invoice OCR pipeline
Multi-stage extraction: OCR → LLM structured output → human-review queue for low-confidence items.
Nearshore delivery by buyer market
Country pages explain the trust, timezone, compliance and delivery questions buyers ask before choosing a software partner.
North Macedonia
Custom software development for Macedonian companies: web apps, mobile apps, cloud systems, AI automation and integrations from Merot's Skopje engineering teams.
Germany
Senior Balkan software teams for German companies: web, mobile, AI, cloud and integrations with EU timezone overlap, GDPR-aware delivery and DACH communication.
Austria
Nearshore web, mobile, cloud and AI engineering for Austrian SMEs and product teams. EU timezone, DACH communication, GDPR-aware delivery.
Netherlands
Senior React, mobile, backend, cloud and AI teams for Dutch SaaS and operations companies. English-first delivery, EU timezone and GDPR-aware workflows.
Switzerland
Senior Balkan software teams for Swiss companies: secure web, backend, cloud, AI and integrations with DACH communication and privacy-aware delivery.
Sweden
Nearshore web, mobile, backend, cloud and AI engineering for Swedish SaaS and operations teams. Senior delivery, EU timezone, clean handoff.
Denmark
Senior nearshore engineers for Danish product and operations teams: web, backend, mobile, cloud, AI and integrations with EU timezone overlap.
Ireland
Senior EU nearshore engineers for Irish SaaS, fintech and operations teams. Web, backend, mobile, AI and cloud delivery with English-first collaboration.
United States
Senior European software engineers for US companies: custom web apps, mobile apps, backend systems, cloud, AI workflows and integrations with practical US overlap.
Canada
Custom web, mobile, backend, cloud and AI engineering for Canadian companies. Senior European delivery, practical timezone overlap and clean handoff.
United Kingdom
Senior nearshore software engineers for UK companies: custom web apps, backend platforms, cloud modernization, AI workflows and integrations.
France
Nearshore product engineering for French companies: web apps, mobile apps, AI workflows, cloud systems and integrations with EU timezone delivery.
Belgium
Nearshore web, backend, cloud, mobile and AI engineering for Belgian companies operating across languages, regions and EU compliance expectations.
Where these engineers come from
Direct EOR employment in two markets, hiring advisory across four more.
Senior engineers from North Macedonia
Merot's home market — deepest pool. ~30,000 IT professionals. Direct EOR via MEROT DOOEL Skopje.
Labour-law + payroll detail →Engineers from Kosovo
Youngest population in Europe, EUR currency (no FX risk). Direct EOR via MEROT L.L.C. Pristina.
Labour-law + payroll detail →Plus 4 advisory markets
Albania, Serbia, Bulgaria, Montenegro — hiring advisory + vetted local payroll partners. See the full outsourcing landing for the trade-offs.
Outsourcing landing →Engagement model
Frequently asked questions — AI + ML
Should I use OpenAI, Anthropic, or self-host?
Default: start with Anthropic (Claude 3.5 Sonnet / Claude 4) or OpenAI (GPT-4 family) for the prototype. Switch to self-hosted only when (a) data residency requires it, or (b) the per-call cost exceeds engineering+infra cost of running yourself. Most clients stay on the API providers for years.
Will my data train someone's model?
Not on the enterprise tiers of OpenAI / Anthropic / Google — they have explicit no-training-on-customer-data terms. We turn on those settings during onboarding.
What if the AI hallucinates / produces wrong output?
Two layers: (1) Eval harness — we measure correctness on a labelled test set before shipping and again on every prompt change. (2) Production — high-confidence outputs go straight through; low-confidence outputs go to a human-review queue.
Cost — won't this get expensive?
It can if scope is vague. We control that by defining the workflow, setting model budgets, adding usage alerts, and deciding what must be automated versus what should stay human-reviewed. The estimate depends on the actual use case.
Do you do fine-tuning?
Sometimes — usually only when the prompt approach truly can't get there. Fine-tuning has higher up-front cost (curating training data) and re-tuning maintenance every model upgrade. Typically we recommend better prompts + RAG first.
Privacy / on-premises only — can you do that?
Yes. We've deployed Llama 3 70B and Mixtral 8x22B on-premise (single-GPU H100 or 4xA100 setups) for clients in regulated industries. Higher up-front cost, lower per-call cost, full data residency.
AI agents — are these real yet?
Cautiously yes. Single-purpose agents (book a meeting, draft an email, run a SQL query) work well with proper guardrails. Generic 'do anything' agents are still flaky. We scope to single-purpose by default.
Voice / speech?
Whisper for speech-to-text, ElevenLabs / OpenAI TTS for synthesis. We've built call-summary + voice-note transcription features for clients in the legal + healthcare verticals.
Let's scope your ai + ml project
Free 60-min discovery call. 6-page written brief in 48 hours.