Shieldstral is a Misc AI tool. Open-source multimodal safety classifier for content moderation and filtering. Key features include Policy-Adaptive Moderation, Multimodal Safety Classification, and Efficient On-Device Deployment. Best for software developers and engineers, content creators and social media managers.
About Shieldstral
Key Features
<strong>Policy-Adaptive Moderation.</strong> You write your safety policy as a plain-language question at inference time, and the model returns a yes or no answer with a calibrated safety score. No retraining needed when policies change.
<strong>Multimodal Safety Classification.</strong> Evaluates both text and images through a single interface. You can moderate prompts, responses, prompt-response pairs, or images with optional text using the same model.
<strong>Efficient On-Device Deployment.</strong> Runs on a single 16GB NVIDIA GPU, making it small enough to deploy as a sidecar on the same hardware serving your main model instead of requiring dedicated moderation infrastructure.
<strong>Open-Source Apache 2.0 License.</strong> Released as open weights under Apache 2.0, allowing free commercial use, modification, and self-hosting without licensing fees or restrictions.
<strong>Binary Question-Answering Framework.</strong> Frames all moderation tasks as simple yes or no questions, unifying prompt classification, response moderation, refusal detection, and toxicity detection into one problem.
<strong>Outperforms Larger Models.</strong> This 3B-parameter model matches or beats safety classifiers nearly seven times its size on text safety benchmarks while setting new standards for multimodal safety classification.
Frequently Asked Questions
Shieldstral is a 3B-parameter open-source AI model from Mistral AI designed for content moderation and safety filtering. It evaluates text and images against custom safety policies by answering plain-language yes or no questions, making it adaptable to different moderation needs without retraining.
Yes, Shieldstral is completely free. It's released under the Apache 2.0 open-source license, which allows unlimited commercial use, modification, and self-hosting without any licensing fees or restrictions.
Shieldstral uses a binary question-answering approach. You provide three parts: an instruction with your policy context, a yes or no question about safety, and the content to evaluate. The model returns a calibrated safety score based on yes and no probabilities, adapting to your specific policy requirements.
Shieldstral is designed for efficiency and runs on a single 16GB NVIDIA GPU. Its compact 3B-parameter size means you can deploy it alongside your main AI model on the same hardware, eliminating the need for separate moderation infrastructure.





