Mathew Stevens

Full-Stack Software Engineer | AI Agent Infra | Real-Time Systems | SDKs
Full-stack software engineer who ships production software solo and operates it under real load: two live apps across web, iOS, and Android with 2,250+ users, 66 reviews, and app-store releases. Also builds the layer underneath, from deterministic execution and exactly-once external actions to evaluation tooling that catches silent agent failures.

Experience

Independent Full-Stack Engineer / Product & Systems

2025 - Present
  • Built and deployed Agentic, an AI-agent authorization layer across web, mobile, CLI, and desktop, reaching 1000+ users in under 1 month with 40 reviews and a 4.2-star rating.
  • Built and launched SolPulse, a production trading platform across web, Android, and iOS, with an execution engine sustaining 10ms monitoring across up to 1,200 concurrent transactions.
  • Built AI-agent reliability/eval systems for exactly-once external actions and crash/recovery conformance; authored the original Terminal-Bench durable-outbox task where 6/6 frontier-agent trials failed clean reward-0.
  • Invited by AI Fund (Andrew Ng) as one of 12 engineers to build against a B2B sales problem posed by a Fortune 100 OEM; delivered ConfigPilot, a decision engine gating LLM recommendations behind deterministic margin, inventory, and policy checks.
  • Built Capability Record Replay, a model-to-deterministic-replay runtime where one LLM discovery run becomes a typed, approved capability that later executes with no model in the decision path; conformance suite killed 9/9 weakened replay engines.
  • Built cross-platform SDKs with native Android/iOS bridges and cryptographic signing across Unreal, Cocos, Godot, Unity, and Capacitor; contributed four merged upstream PRs to the Godot and Unity adapters.

Analytics Specialist, Socket Mobile, Inc.

Apr 2020 - Jul 2023
Fremont, CA
  • Owned web analytics end to end in Google Analytics and Google Tag Manager: data collection, tagging, QA, A/B and multivariate testing, dashboards, and reporting for marketing and leadership.
  • Partnered with the web team on the analysis behind a 25% year-over-year sales increase, turning traffic and conversion data into channel and campaign decisions leadership acted on.

Selected Projects

Agentic - AI Agent Signing Platform

Web, CLI, Desktop, App Store
  • Built an authorization layer where AI agents propose actions and users approve them on their own devices; private keys never leave user custody.
  • Shipped across web, mobile, CLI, desktop, MCP tools, SDK surfaces, and app-store releases behind one non-custodial approval boundary.
  • Implemented shared WalletBackend plus A2A AgentCard, AP2/ACP/MPP payment-adapter flows, MCP/Vercel AI tools, approval polling, simulation, and sign-and-send UX.

SolPulse - Trading Platform

React, Node, PostgreSQL, Redis
  • Built solo full-stack trading product with live website, App Store, Google Play, and browser-visible release proof.
  • Implemented execution engine with DCA, limit orders, conditional exits, multi-venue routing, dynamic fee optimization, and real-time strategy logs.
  • Built risk and fraud detection across 35 scoring signals: counterparty identity checks, anomaly and large-holder detection, max daily loss limits, and per-trade exposure caps.

Agent Reliability & Eval Integrity Systems

Python, TypeScript, Docker, Terminal-Bench
  • Built Agent Eval Foundry, a coding-agent benchmark production system with 25 task packages and validated graders; nine selected finalists recorded 52/54 required-deliverable failures across Codex and Claude trials (seven at 6/6, two at 5/6).
  • Built the original Klavis Terminal-Bench durable-outbox task plus its reference implementation, which passes 267/267 conformance checks.
  • Built Durable Agent Outbox and Crashpoint to measure exactly-once external side effects, UNKNOWN outcomes, receipt reconciliation, revocation, and crash/recovery across LangGraph, Temporal, and DBOS.
  • Built Cheat-Oracle and Judge-Artifact audits showing detector undercount and grader artifact in public agent benchmarks, with tamper-evident receipts, controls, and disclosure-ready reports.

Capability Record Replay - Agent Automation Runtime

TypeScript, Playwright, Anthropic, CLI
  • Built system where an LLM drives a legacy back-office UI once, then synthesis emits a typed, versioned capability artifact replayed without a model in the decision path.
  • Implemented browser and terminal surfaces, deterministic interpreter, policy chokepoint, signed approvals, typed business outcomes, recovery/failure taxonomy, evidence journals, and redaction canaries.
  • Verified a live Claude Opus discovery run and model-free replay across 25 hostile fixture scenarios.

Cross-Platform SDKs and Engine Integrations

Unreal, Cocos, Godot, Unity, iOS
  • Built Unreal Engine 5 mobile wallet plugin with Kotlin clientlib, C++ JNI bridge, Blueprint UFUNCTION surface, auth caching, SIWS, and sign-and-send flows.
  • Built the first Cocos Creator Mobile Wallet Adapter SDK from scratch: TypeScript API, native Android bridge, ECDH/AES-GCM wallet association, auth caching, and Token Duel demo game.
  • Delivered Godot MWA reconnect architecture, SIWS wallet-proof APIs, persistent auth-token cache, Unity WebGL UX fix, and iOS-native Swift/CryptoKit adapter flows.

Technical Skills

Languages: TypeScript, JavaScript, Node.js, Swift, C#, C++, GDScript, Kotlin, SQL, Python
Frontend: React, Vite, Zustand, Tailwind, Radix UI, Capacitor, SwiftUI
Backend: Express, FastAPI, Pydantic, PostgreSQL, Prisma, Redis, WebSockets, SSE, REST APIs
Security / Transaction Systems: web3.js, Mobile Wallet Adapter, Wallet Standard, Seed Vault, Jupiter, Anchor, SPL tokens, transaction signing
Platforms: Android, iOS, Unity, Unreal Engine, Godot, Cocos Creator
Systems / AI Infra: Docker, CI/CD, Terminal-Bench, Harbor, conformance suites, mutant tests, deterministic replay, policy gates, eval-integrity audits

Launch Metrics

1000+SolPulse users
4.3 / 26SolPulse rating/reviews
1250+Agentic users
4.2 / 40Agentic rating/reviews
2Store-listed apps
4Merged upstream PRs

Open-Source Proof

  • GitHub portfolio: github.com/mstevens843
  • Agent Eval Foundry: 25 task packages, screening evidence, verifier controls, and per-trial failure analysis.
  • Godot Mobile Wallet Adapter PRs: #449 reconnect/capabilities/sign-send bridge; #453 SIWS wallet-proof APIs; #454 auth-token cache.
  • Unity Wallet Adapter merged PR: #275 WebGL single-wallet auto-select UX fix.
  • Project proof: capability-record-replay, durable-agent-outbox, crashpoint, cheat-oracle, judge-artifact, agent-context-containment, model-regression-sentinel, toolcall-risk-classifier, ConfigPilot, Klavis TB3, TxShield.

Upstream Impact

  • Crashpoint / DBOS: Supplied Postgres crash/recovery evidence supporting merged workflow-identity fix #838.
  • Crashpoint / LangGraph: Reported pre-checkpoint gap #8764; reproduction adopted in AgentCI's merged fixture #160.
  • Judge-Artifact / Inspect: Contributed 3,986-transcript grader analysis and parser regressions to #2108 / #2310.

Education

San Francisco State University
B.S., Business Information Systems, 2018

Springboard
Software Engineering Certificate, 2025
Full-stack curriculum covering JavaScript, React, Node.js, Express, PostgreSQL, REST APIs, auth, and testing.

Focus Areas

Full-Stack Products AI Agent Infra Eval Integrity Real-Time Systems Mobile Apps SDK Engineering
Best-fit roles: full-stack product engineering, backend/API systems, AI-enabled systems, reliability/eval tooling, SDKs, mobile apps, and real-time infrastructure.