$ whoami_
James Tavita
Software engineer, machine-learning researcher, and technical founder building systems that explain not only what will happen, but why.
I develop custom software, machine-learning systems, causal research, and AI evaluation infrastructure that help organizations make better decisions.
software engineering • causal inference • machine learning • AI evaluation
Projects
Deep-dive write-ups of real models — architecture, evaluation, error analysis, and limitations.
NBA Player Performance Forecasting: CatBoost vs. Random Forest
Team analysts, media, and award voters all want a credible answer to "how will this player perform next season" before the season happens — but naive year-over-year projections ignore usage changes, aging curves, and role shifts already latent in the box score.
BDS Media Engagement Analytics: What Predicts Views Across a Two-Channel Content Brand
A multi-host YouTube/Instagram media brand — a flagship golf-content channel plus a companion podcast — needed to know which levers actually move engagement (who's talking and for how long, what a title says, how long an episode runs, when it publishes, whether one channel's performance affects the other) instead of relying on editorial instinct alone.
Premier League Match Forecasting: Baselines, Challengers, and a Shipped iOS App
Predicting a Premier League match outcome (home win / draw / away win) is a decades-old actuarial problem with a low ceiling — bookmakers and simple rating systems already capture most of the signal. The real question wasn't whether a model could predict outcomes, but whether a more sophisticated model (gradient-boosted trees on rich player and team features) actually beats well-understood baselines like Elo and Dixon-Coles, and whether the whole system could be shipped as a product fans could query directly.
Loan Default Prediction: CatBoost on Imbalanced Tabular Credit Data
Given a loan application's financial and demographic profile, predict whether the loan will ultimately be paid back — a standard credit-risk screening problem where the cost of a false negative (approving a loan that defaults) and a false positive (declining a loan that would have been repaid) are asymmetric and business-defined, not something the model itself can resolve.
JIVE: Correcting Many-Instruments Bias in a Movie Box-Office Natural Experiment
Does a movie's opening-weekend ticket sales causally drive its total ticket sales, or does the naive correlation just reflect that both are driven by the same unobserved quality/demand? Opening sales is endogenous — correlated with unobserved quality — so a plain regression of total sales on opening sales conflates the causal effect with that confound.
House Prices: A Stacked Ensemble Across CatBoost, XGBoost, and LightGBM
Predict a home's sale price from 79 property features (Kaggle's Ames, Iowa House Prices competition) — a benchmark regression problem where the real difficulty isn't fitting any single model, it's handling a long tail of correlated, mixed-type features (dozens of quality/condition ratings, areas, and categorical construction details) without one encoding scheme handicapping one model family over another.
Who Leads a Mission: 504 Leadership Appointments, Mapped and Tested
Across the 504 missions of The Church of Jesus Christ of Latter-day Saints, is the distribution of leadership appointments anything other than membership spread across a map — and do the leaders drawn from different parts of the world differ?
Publications
Peer-reviewed research and business case studies I've co-authored.
Oversight Risk: How Committees Shape Portfolios
Scott Condie, Gabriel Lehnardt, James Tavita
From Payments to Power: How the PayPal Mafia Shaped Silicon Valley's Venture Landscape
Tom Hunsaker, Abdulaziz Alakeel, James Tavita
Ways of Working Together
Custom Software & Data Products
Internal tools, web applications, APIs, dashboards, data pipelines, and decision-support systems.
Machine Learning & Predictive Modeling
Forecasting, classification, ranking, optimization, simulation, model comparison, and production integration.
Causal Inference & Experimentation
Treatment-effect estimation, experiment design, observational studies, difference-in-differences, synthetic controls, instrumental variables, regression discontinuity, matching, and causal machine learning.
AI-Agent Evaluation
Role-specific evaluation, simulation environments, failure analysis, human-versus-agent comparisons, compliance testing, and production monitoring.
Technical Strategy & Prototyping
Product definition, architecture, technical due diligence, rapid prototyping, and translating executive requirements into technical systems.
Have a difficult technical or analytical problem?
Describe the decision, system, or research question you're working on. I can help determine whether it needs custom software, machine learning, causal analysis, AI evaluation, or a combination of them.