Research

The research behind SafeFlow.

SafeFlow's AI Assurance methodology is grounded in original research on operational AI reliability and repeated-inference evaluation — a depth-oriented approach to measuring how production AI systems behave under realistic use.

Publications

Research that informs our methodology.

The methodology is documented in public preprints, open to scrutiny.

arXiv:2602.11786 · Preprint

Evaluating LLM Safety Under Repeated Inference via Accelerated Prompt Stress Testing

A framework for estimating empirical failure probability under repeated inference across production-representative prompt distributions. It shows why single-shot benchmarking can understate the operational risk of deployed language models.

Read on arXiv
arXiv:2604.09606 · Preprint

Evaluating Reliability Gaps in Large Language Model Safety via Repeated Prompt Sampling

An analysis of reliability gaps in safety-critical LLM deployments, demonstrating how repeated prompt sampling surfaces failures that static evaluations miss and how those failures cluster in predictable semantic regions.

Read on arXiv