HumanWill

Useful AI. Transparent decisions. Legitimate human control.

HumanWill is an open evaluation and governance initiative working toward AI systems that remain useful, transparent, and under legitimate human control.

AI safety is not only about preventing harmful answers. It is also about preventing harmful refusals.

What we evaluate

Can an AI system distinguish when it should help, when it should refuse, and when it should ask for clarification?

Our evaluation framework measures separately:

We do not reduce these dimensions to a universal “uncensored” score.

Domain-specific evaluations

We are building a shared evaluation framework with dedicated domain packs:

  1. Cybersecurity — our first development priority.
  2. Fraud and financial crime — planned next.
  3. Biosecurity — subject to stricter controls and expert review.
  4. Healthcare — planned.
  5. Critical infrastructure — planned.

Scenario families pair related malicious, legitimate, and ambiguous requests to test how models respond to differences in intent, authorization, and context.

Current stage

HumanWill is in early development. Our repository defines the project architecture, governance principles, and first cybersecurity milestone: a synthetic access-control scenario family.

Published benchmark datasets and model evaluation results are not yet available.

Help define responsible AI in your field

We welcome domain experts, researchers, developers, and organizations interested in contributing scenarios, reviewing evaluation criteria, or improving the framework.

Participation is local-first and opt-in. Contributors should be able to keep scenarios private or choose to share metadata, anonymized patterns, complete scenarios, or expert-verified scenarios.

Public releases require appropriate provenance, licensing, privacy, and safety review. Held-out evaluation material and private submissions are kept separate from public development scenarios.

Know your domain. Help us test whether AI knows when to help.

Links