Random photographs of everyday subjects. None contain explicit content. Every image should pass the classifier.
Content warning
This section displays explicit test images classified through the HarmBlock model. All should be blocked.
Explicit test images. Every image should be blocked by the classifier.
Content warning
This section modifies blocked explicit images to show how the classifier is defeated.
Each blocked image is modified with transformations that occur during normal phone use. Green borders indicate bypasses.
The model contains 15 output tensors. Some are feature embeddings (which respond to input), some are classification heads (which should classify content into categories), and some are constants. The table below compares each head's output range against the main classification neuron to show which are functionally active and which are dead.