PAULO VILA AI VERSION
McKinsey → BCG → Carrefour → Mastercard → Imaginos
I Build Systems. Technology executive and hands-on platform architect. 30 years designing, selling and delivering large multi-country technology programs across McKinsey, BCG, Carrefour and Mastercard — and today CEO and lead architect of a local-first, multi-agent AI platform written in Rust, running five production applications on one shared core. Over a decade at Mastercard scaling an innovation practice across Latin America: 10X growth in strategic engagement value, 5X revenue for Labs as a Service, 60+ client projects at 90%+ satisfaction, $100M+ in platform revenue, and 100+ C-suite design sprints that turned architectures into signed programs. The rare combination of the executive who owns the client relationship and the engineer who writes the runtime, in one person. M.S. Industrial Engineering, Universidad de los Andes.
Self-funded platform: five applications running in production, not yet commercialized.
- Own the end-to-end architecture of a multi-agent AI platform serving 5 production applications from one shared Rust core — orchestration, memory, routing, identity and runtime governance as shared services instead of per-application logic.
- Cut AI inference cost to $0 per query, against $0.01–0.05 per request on cloud APIs, by deploying an on-premises GPU cluster (2x RTX 3090) for LLM, image generation and multilingual embeddings — cloud demoted to failover only.
- Deployed 13 autonomous AI agents in 9 months with persistent scoped memory, semantic recall and multi-channel delivery (web, Telegram, WhatsApp) over one conversation context layer.
- Engineered sub-millisecond IPC (1.019 ms average, 1.136 ms p95 — 300x+ over the prior Next.js stack) with automatic PostgreSQL high-availability failover (~2 min RTO, RPO ~ 0) via streaming replication and quorum witness.
- Built a zero-shot NLP pipeline (GLiNER, GLiREL, GLiClass on GPU) for knowledge graph construction — entity recognition, relation extraction and sentiment classification with no labeled training data.
- Implemented Model Context Protocol (MCP) bridges and a self-correcting agentic coding loop: the runtime edits its own Rust source, runs the compiler and iterates until the build is clean.
Single-operator at platform scale: design the core abstractions, build them in Rust, ship to real users across web + Telegram + WhatsApp. Five apps (paulovila.org, vetra.trade, latinos.paulovila.org, movilo.club, elgarcero.com) share one runtime today.
- $0/query AI inference cost — on-prem GPU (2x RTX 3090) vs. ~$0.01–0.05/request on cloud APIs
- Rust warm-path latency: 1.019 ms avg, 1.136 ms p95 (Vetra-rust); 1.1 ms TTFB (Movilo) — 300x+ over prior Next.js stack
- 13 autonomous agents deployed in 9 months with persistent memory and semantic recall
- Zero-shot NLP pipeline: GLiNER + GLiREL + GLiClass (GPU) for knowledge graph extraction without labeled data
- Automatic Postgres HA: ~2 min RTO, RPO ~ 0 via streaming replication + quorum witness
- 5 production verticals (compliance, trading, healthcare, ecotourism, portfolio) on one shared runtime
- 10X growth in strategic engagement value as VP Digital Labs, Latin America
- $100M+ B2B credit card sourcing platform at Mastercard (18 months)
- Supported launch of Colombian neobank reaching 1M+ users
- 5x revenue growth for Mastercard Labs-as-a-Service in LAC
- 60+ innovation projects delivered, 90% client satisfaction
- 100+ C-suite Design Thinking sprints facilitated across Latin America
- 3 Mastercard President's Awards
Stanford University
The Wharton School
Mastercard, 2022 — women-led fintech LAC
Mastercard, 2020 — Digital Datathon LAC (best in company)
Mastercard — selected from ~30K global employees
Mastercard, 2016
Mastercard, 2016