HarmBlock+ Model Audit

Live classification through the SafeToNet ONNX model extracted from the HMD Fuse

Loading model
0
Tested
0
Blocked
0
Passed
0
Output neurons
0
Responsive
0
Dead
Neutral neuron (responds)
NSFW neuron (doesn't respond)
Random photographs of everyday subjects. None contain explicit content. Every image should pass the classifier.

Content warning

This section displays explicit test images classified through the HarmBlock model. All should be blocked.

Content warning

This section modifies blocked explicit images to show how the classifier is defeated.

The model contains 15 output tensors. Some are feature embeddings (which respond to input), some are classification heads (which should classify content into categories), and some are constants. The table below compares each head's output range against the main classification neuron to show which are functionally active and which are dead.