Careers / AI Engineer
AI Engineer
AI Engineer, Scoring & Signal Intelligence Location: LatAm remote, strong preference for Costa Rica / Central America. Requires travel to the client's office in New York: 2 to 3 weeks onsite at kickoff, then roughly every other month. A valid passport and a current US travel visa (B1/B2) are required. 📌 Summary You will own the scoring and signal layer of a data-intensive internal platform in the financial services space. The system brings together licensed data providers, public web signals, and notes captured by the team, then turns all of it into a ranked, explainable view of which companies deserve attention and why. You will define how that score is built, prove it holds up, and keep sharpening it with direct input from the people who use it every day. ✅ Requirements (Must-Haves) Excellent English communication skills. You will work directly with the client's leadership and business team every day, in their meetings, explaining and defending your scoring decisions in live conversation. 3+ years building production LLM systems (agents, RAG, structured extraction) and 5+ years total in data science, analytics, or ML. We need both halves: the LLM half and the statistical rigor half. Strong Python and SQL. You write your own queries, build your own features, and do not wait for someone else to prepare the data. Demonstrated work on scoring, ranking, propensity, or lead prioritization systems. You have built something that assigned a number to an entity and had to defend that number to a skeptical business stakeholder. Rigorous LLM evaluation practice: eval sets, regression testing on prompt changes, measuring agent output quality. You can explain how you knew a prompt change made things better rather than just different. Experience turning messy, conflicting, multi-source data into a single usable signal, including entity matching and handling contradictory records. Comfort defining and defending business metrics with non-technical stakeholders. A background in analytics is an asset in this role, not a detour. Experience with feature stores, vector databases, or embedding-based similarity search on company or people data. A valid passport and a current US travel visa (B1/B2). ✅ Requirements (Nice-to-Haves) Exposure to private markets data: funding rounds, revenue estimation, ownership structures, company relationship graphs. Hands-on with commercial company data providers and their API quirks: coverage gaps, stale records, rate limits, and inconsistent identifiers. A product-driven mentality. The score is not a deliverable you hand over and forget, it is a product with daily users who will not file tickets when something is off. They will just trust the output less. The people who succeed in this role notice that a signal has gone stale, that a ranking stopped making sense, or that the team has started working around the tool, and they act on it before anyone asks. Time series or cohort analysis on entity-level signals, for example detecting acceleration in a metric rather than reading a static snapshot. Experience working in an Agentic Development environment (spec definition, AI agent delegation for code creation, human-in-the-loop review). Familiarity with Claude Code, Codex, or similar harnesses and AI Skills. 🌟 Bonus Points You have shipped an agentic system where the agent's judgment was the product, not a chatbot wrapper. Published work, open source, or talks on LLM evaluation, agent design, or scoring methodology. You have worked closely with investment or dealmaking teams and can already speak their language. Experience designing the human-in-the-loop feedback mechanism that makes a model improve from expert input over time.