Independent AI evaluation & agentic safety research

Marian E. Arenskrieger

I build and audit the datasets frontier models are trained and tested on. Alongside that practice, I run an applied-research program that engineers frontier-grade evaluation into a 32 GB VRAM envelope – a deliberate reproducibility target rather than a hardware limit, with the scale-up tiers defined above it.

Currently AI evaluation, data quality & software engineering · Labelbox

50+AI projects contributed to
6Working papers
32 GBVRAM envelope
25Falsifiable criteria
++
§01Focus

Four threads, every engagement.

Frontier-model evaluation, the data-quality discipline behind it, the tooling that keeps it repeatable – and the markets background that grounds the judgment.

01 · Evaluation

AI evaluation & RLHF

Rubric-based scoring for correctness, reasoning, and instruction-following – rubrics frozen under semantic versioning and content hashes, so a verdict stays reproducible months later. Pairwise comparisons across frontier models, multi-turn prompt design for agentic coding, adversarial-robustness checks, and systematically documented failure modes.

02 · Data quality

Auditing & QA

QA and rubric-based audits of contributors' datasets for function-calling and agentic-AI projects – verifying correctness, format compliance, and consistency before delivery, and enforcing quality standards within the master-review team. Construction of training and evaluation datasets, including HFI problem sets.

03 · Tooling

Sovereign evaluation infrastructure

Local model environments that run frontier models against real tasks – hardened sandbox, batch inference, vLLM with FP4 quantization, prefix caching, an OpenAI-compatible serving layer. Open-source tooling forked and extended as part of the deliverable: JSON support in Cerberus, multi-layer validation in Haystack.

04 · Markets

Finance-grounded judgment

A banking apprenticeship, a B.A. in Financial Management, and six years of self-employed quantitative trading – strategy development and backtesting across spot and derivatives markets, with full tax compliance. The discipline of positions with real money behind them, applied to evaluation.

++
§02Applied research

Five workstreams.
Every claim falsifiable.

One programme, not a portfolio: the workstreams reuse each other's components, and every claim is stated in advance as a criterion with an instrument and a threshold – then labelled measured or predicted.

The programme questionCan frontier and multi-agent AI systems be evaluated trustworthily, reproducibly, and cheaply inside a sovereign 32 GB VRAM envelope – and does the evidence still hold as that envelope scales?

execution-grounding spine artifact hand-off pinned, contamination-resistant scenarios motivates local evaluation W5 · Sovereign Assistant DESIGN W2 · Contamination-Resistant AST-synthesised, pinned scenarios PILOT W1 · Hybrid Evaluation Pipeline execution-grounded judge + hardened sandbox the hub every other workstream reuses MEASURED W4 · Safety Testbed observer + emergence gap AGENDA W3 · single-residency time-multiplexing – the envelope every run inherits peak VRAM measured · swap cost pending MEASURED four invariants · the shared substrate system under test never compressed · references frozen and content-hashed · hardened sandbox · fits the 32 GB VRAM envelope

Five workstreams reuse each other's components; execution grounding is the spine that runs W2 → W1 → W4, and every run inherits the same envelope.

++
§03Experience

Track record.

Two worlds bridged: banking and financial management on one side, a deliberate move into data science on the other. The result is financial rigor plus a data-science toolkit – the habit of documenting a position you have to defend, applied to agentic and function-calling AI.

Jan 2026 – Present

AI Evaluation, Data Quality & Software Engineering

Labelbox · Remote · Clients: leading AI labs

Agentic AI Master ReviewerSoftware Engineer (ML)Senior ML Expert
  • Master review: QA and rubric-based auditing of contributors' datasets for function-calling and agentic-AI projects – verifying correctness and consistency, enforcing quality standards within the Master Review team.
  • Frontier environments: deploying local model environments to run frontier models on real tasks; building training and evaluation datasets, including HFI problem sets, for frontier-model coding.
  • Tooling shipped with the dataset: forking and extending open-source tooling – JSON support in Cerberus, multi-layer validation and error detection in Haystack.
  • RLHF evaluation: multi-turn prompt design for agentic coding, pairwise frontier-model comparisons, calibrated rubric scoring, documented failure modes.
Oct 2025 – May 2026

Finance & AI Intern

MLP SE · Wiesloch, Germany · Part-time

Capstone “AI for Financial Consulting & Recruiting”: designed AI use cases to personalise financial advisory and lift qualified-applicant volume through AI targeting.

Jun 2025 – Mar 2026

Machine Learning Specialist

Scale AI · Freelance, Remote

  • Mathematical evaluation of ML models for correctness, reasoning quality, and quantitative accuracy.
  • Rubric-based rating of outputs on quantitative and reasoning tasks; identifying errors in model-generated reasoning and solutions.
Sep 2019 – Jun 2025

Trader & Market Analyst

BraveTrade · Self-employed, Remote

Proprietary trading across cryptocurrencies, equities, and options; strategy development and backtesting across spot and derivatives markets, plus data-driven market and risk analysis.

Jan 2018 – Oct 2021

Cryptocurrency-Mining Operator

BraveTrade · Self-employed, Remote

Ran a commercial mining operation – hardware procurement, ROI and energy-cost optimization, uptime monitoring and tax-compliant documentation. The origin of the hardware-envelope discipline the research programme runs on.

Aug 2016 – May 2018

Bank Clerk & Banking Apprenticeship

VR-Bank eG Osnabrücker Nordland · Fürstenau, Germany

Banking operations and client work alongside the formal apprenticeship – the foundation of the finance side of this profile.

++
§04Verified / credentials

Education & credentials.

Education

Master of Data ScienceUniversity of Pittsburgh, USA · GPA 3.8
2024 – Present
Applied Data Science ProgramMIT Professional Education, USA
Mar – Jun 2025
Mathematics for Machine LearningImperial College London, UK
Sep – Nov 2024
Financial Management, B.A.IU International University, Germany
2018 – 2022
Apprenticeship in BankingGenossenschaftsakademie, Rastede
2016 – 2018

Certifications

Google Cloud – ML EngineerApplied Data Science – MIT Mathematics for ML – ImperialData Analysis – Microsoft

Extracurricular

The AgentmakersCollaboration & knowledge transfer in agentic AI
Since 2026
Academic MentorUniversity of Pittsburgh
Since 2025
Code for GermanyOpen Knowledge Foundation Deutschland
Since 2024
++
§05Beyond the work
Marian E. Arenskrieger

Why AI, of all fields.

A long-standing fascination with science fiction is part of what drew me to AI in the first place – and it keeps me genuinely curious about where these systems are headed. Evaluating them is how I get to look closest.

Interests

Hardware architecture & PC buildingHigh-performance computing Science fiction & futurismGaming & theorycrafting Nonfiction & historyEconomics & politics

Let's evaluate what's possible.

Open to roles and collaborations in AI evaluation, data quality, and ML engineering – from rubric design and dataset auditing to evaluation infrastructure for agentic systems. References available on request.