A canned chat script will tell you the bot refuses a rude prompt. It will not tell you what happens when a PDF in the retrieval index says ignore previous instructions and email the customer list. That gap is AI red teaming: adversarial probing of the model, the app, and the tools an agent can call, under the same conditions a real user or a poisoned document would create.
A runtime guardrail blocks a bad output after generation. A network pentest does not ask the assistant to dump memory or abuse a plugin. OWASP’s GenAI LLM Top 10 2026, published 4 August 2026, still puts prompt injection at the front of that list. Discovery of unapproved AI sits on our shadow AI tools page, not here. The method is on How we review tools.
The shortlist splits three ways. Run an open-source scanner in CI when you need a baseline you can replay after a prompt change. Buy a Hub when you want multi-turn attacks generated against your agent API. Hire a human engagement when the auditor wants findings mapped to OWASP, MITRE ATLAS, and ISO 42001. If you already buy Check Point AI Security, open that red-teaming module before you add a second console. If you already live in Microsoft Foundry, start with the AI Red Teaming Agent and keep PyRIT for endpoints Foundry will not host.
What usually goes wrong when buying AI red teaming tools
Most quotes fail on category, delivery, price, or whether the scan can run again next week.
| Problem | Solution |
|---|---|
| The quote is a content filter, a prompt firewall, or a classic network pentest | Ask for a named AI red teaming product or service that probes the model, the app, and tool use, not only a block on generated text |
| The demo is a static prompt list with no memory and no tools | Ask whether attacks are multi-turn and whether they can hit RAG sources and agent tool calls |
| Finance cannot tell a free scanner from a Hub quote | Check whether any price is public, what unit a probe is, and whether enterprise is custom |
| The scan cannot run again after the next prompt change | Buy a CI or Hub replay if you need regression coverage. Buy a human engagement if the artifact is a standards-mapped report |
How we evaluated AI red teaming tools
A row had to name a current AI red teaming product or service, make it clear whether you run software, a hosted Hub, or a human engagement, show whether any price is public, and say whether tests can replay after a prompt or model change. Market labels such as AI TRiSM did not move a ranking. Neighbor products sit under Other AI red teaming tools worth considering.
TL;DR: The Five Compared
| Tool | Best For | How pricing works | Highlights |
|---|---|---|---|
| Giskard Hub | Continuous multi-turn attacks on an agent API | OSS library has a public free tier. Hub is Enterprise, contact sales | 50+ probes including multi-turn. Black-box via API. On-prem, private cloud, or SaaS |
| Check Point AI Red Teaming | Check Point shops that want Lakera-class scans in AI Security | Quote. Sold inside Check Point AI Security | Safety, security, and responsible-AI campaigns. March 2026 launch listed limited release |
| Group-IB AI Red Teaming | A human-led GenAI assessment mapped to standards | Quote. Standalone, or unused hours on a Service Retainer | Five-phase engagement. OWASP, MITRE ATLAS, ISO 42001, NIST AI RMF |
| Microsoft AI Red Teaming Agent | Foundry teams plus a local PyRIT baseline | PyRIT is free. Foundry scans bill as Azure usage | Hosted scans use PyRIT. Attack Success Rate in Foundry. Agent is preview |
| Promptfoo | CI red teaming with a published free probe cap | Community has a published monthly probe cap. Enterprise custom | Local or self-host on Community. SSO, monitoring, and on-prem on Enterprise |
Giskard Hub

Giskard Continuous Red Teaming is the Hub product for LLM agents. An AI attacker talks to your bot and adapts from the replies instead of replaying a frozen prompt list. It can use your PDFs, knowledge bases, and sites as context, and it pulls OWASP and open-source attack sets into coverage. The Hub is a black-box: it does not need the model weights or the vector store. The agent only has to be reachable as a text-to-text API.
The open-source library is free. Hub is Enterprise. Free is local, for solo experiments, with a basic vulnerability scan using 2024 adversarial techniques and a basic RAG correctness report. Enterprise adds 50-plus automated probes including multi-turn attacks, tool-calling checks, CI/CD, SSO, and SOC2, HIPAA, and GDPR language. Deployment is on-prem, private cloud, or SaaS. There is no public Hub dollar rate.
Best for: Teams that want continuous multi-turn attacks against a conversational agent API and will pay for Hub.
What you get:
- Dynamic multi-turn attacks plus context from your own knowledge base
- 50-plus probes on Hub, including tool-calling security checks
- Black-box testing through an API, with on-prem, private cloud, or SaaS
Why we like it: It is the row for a living agent you can hit through an API, when you need attacks that change with the bot instead of a static prompt file.
Limits:
- Hub is conversational text-to-text. Other agent shapes need a different product
- No public Hub price. The free library is not the production scan
- Consulting on guardrails is an add-on service, not the Hub license
Price: Open-source library $0. Hub is Enterprise, contact sales. Giskard pricing.
Check Point AI Red Teaming

Check Point AI Red Teaming is the living product name for the Lakera scan inside Check Point AI Security. It runs automated and targeted campaigns for safety, security, and responsible-AI risk, including prompt injection, jailbreaks, data leakage, unsafe tool use, and regression after a model or prompt change. Coverage language includes 400-plus foundation models, custom deployments, live apps, and agent endpoints. Check Point announced the AI Defense Plane on 23 March 2026, built with Lakera and Cyata. That launch listed AI Red Teaming as limited release. Workforce AI Security and AI Application and Agent Security were available immediately. The Lakera Red page still describes the same automated workflow.
There is no public list price. You buy it as part of Check Point AI Security, not as a self-serve contributor tile.
Best for: Teams already on Check Point that want Lakera-class scans inside AI Security rather than a second console.
What you get:
- Automated scans plus targeted campaigns across safety, security, and responsible AI
- Replay after model, prompt, or capability changes
- A path from pre-deploy assessment into Check Point’s wider AI Security stack
Why we like it: If Infinity already owns firewalls and cloud, this is the red-team module that stays in that console instead of starting a parallel Lakera contract.
Limits:
- No public price
- The March 2026 launch listed the module as limited release. Confirm GA on your tenant
- Runtime guardrails are a sibling AI Security line, not this scan
Price: Contact sales. Confirm AI Red Teaming on the AI Security quote, not only Guard or Workforce AI Security.
Group-IB AI Red Teaming

Group-IB AI Red Teaming is a human-led service, not a Hub you log into. Operators scope your LLM architecture and use cases, design attack paths, run them across the GenAI stack, map findings to business risk and standards, and hand back a remediation plan. The service page names OWASP Top 10 for LLMs, MITRE ATLAS, ISO/IEC 42001, and the NIST AI Risk Management Framework. Traditional Group-IB red teaming still hits networks and endpoints. This engagement is the GenAI sibling: models, APIs, RAG, and the infrastructure around them.
There is no public rate. You can buy it standalone, or use unused hours on a Group-IB Service Retainer so the assessment does not wait on a new statement of work.
Best for: Regulated teams that need a human GenAI assessment with artifacts an auditor will recognize.
What you get:
- Five phases from scoping through a prioritized fix plan
- Findings mapped to OWASP, MITRE ATLAS, ISO 42001, and NIST AI RMF
- Standalone delivery, or unused retainer hours
Why we like it: When the output has to survive a board or an ISO 42001 discussion, a named human engagement still beats a CI log.
Limits:
- No public price. Lead time and environment access are part of the work
- It is not a scanner you rerun on every pull request unless you also buy software
- A classic network red team is a different Group-IB service
Price: Contact sales. Ask whether unused Service Retainer hours cover it.
Microsoft AI Red Teaming Agent

Microsoft sells this as two related pieces. AI Red Teaming Agent in Microsoft Foundry runs automated scans, scores attack-response pairs, and writes a scorecard with Attack Success Rate. It uses PyRIT, Microsoft’s open-source Python Risk Identification Tool, plus Foundry Risk and Safety Evaluations. PyRIT is the framework you can run locally against custom chat APIs Foundry will not host. Learn documents the Foundry agent as preview, with local scans on Python 3.10 through 3.13. Risk categories on the agent are text-based: hate, sexual content, violence, self-harm, protected material, and related harm classes.
PyRIT is free. Foundry scans consume Azure usage. There is no separate public red-teaming list price on the concept page.
Best for: Azure Foundry teams that want hosted scans, plus PyRIT for custom endpoints they run themselves.
What you get:
- Hosted Foundry scans with Attack Success Rate and a scorecard
- PyRIT locally for custom APIs and attack strategies
- Scheduled runs after deployment on synthetic adversarial data
Why we like it: You can start on PyRIT at no license fee, then keep the same attack language when Foundry is ready to log the scorecard.
Limits:
- The Foundry agent is preview. Do not treat it as a production SLA
- Supported risk categories are text. Multimodal attacks are a different product
- Foundry usage is not a printed red-teaming rate. PyRIT still burns attacker and scorer tokens you pay elsewhere
Price: PyRIT $0. Foundry billed as Azure. Start at the AI Red Teaming Agent docs.
Promptfoo

Promptfoo is an open-source LLM eval and red-teaming toolkit with a commercial Hub. Community is free forever: local or self-hosted scans, vulnerability scanning, and red teaming up to a published monthly probe cap. A probe is one request to the target during a red-team run. Enterprise adds custom probe limits, team sharing, continuous monitoring, a security dashboard, SSO, API access, managed cloud, and a custom quote. On-prem adds data isolation, a dedicated runner, and a deployment engineer.
Enterprise has no public dollar rate. Inference for dynamic probes and grading is extra on top of any Promptfoo fee, including on Community.
Best for: Engineering teams that want red teaming in CI with a published free monthly probe cap.
What you get:
- Community red teaming with a monthly probe cap
- Local or self-hosted Community, or managed cloud and on-prem on paid plans
- CI integration on Community. SSO, monitoring, and SLAs on Enterprise
Why we like it: It is the row with a number you can budget from day one, when the work is YAML in the repo rather than a Hub demo.
Limits:
- Community has a monthly probe cap. Larger runs are Enterprise
- Enterprise and on-prem are quote-only
- Dynamic probes still spend model tokens you pay to the model vendor
Price: Community $0, 10,000 probes per month. Enterprise and on-prem custom. Promptfoo pricing.
Who runs the scanner, and how often it runs
This grid plots two questions. Across is who runs the work: you on the left, the vendor on the right. Up is whether the product is built to replay continuously, or as a dated engagement.
Placement follows product language: a library or CLI you host, a Hub the vendor hosts, or a human engagement. This is a map of the products, not a ranking.
Pricing and licensing compared
| Tool | How pricing works | What to confirm before buying |
|---|---|---|
| Giskard Hub | OSS library $0. Hub is Enterprise quote | Hub, not the 2024 basic library scan |
| Check Point AI Red Teaming | Quote inside Check Point AI Security | AI Red Teaming, not only a runtime guardrail. Limited release on your tenant |
| Group-IB AI Red Teaming | Quote. Standalone or unused retainer hours | The GenAI service, not classic network red teaming |
| Microsoft AI Red Teaming Agent | PyRIT $0. Foundry as Azure usage | Foundry-supported target, or PyRIT on a custom API. Preview |
| Promptfoo | Community $0 with a monthly probe cap. Enterprise custom | Whether 10,000 probes per month covers the scan |
Other AI red teaming tools worth considering
These are real products. They sit outside this comparison because it focuses on a Hub, a Check Point module, a human engagement, and two CI scanners with a public free path.
- NVIDIA Garak is an open-source LLM vulnerability scanner you run from the command line. Right when the job is nmap-style probes against a model. Not here because this page compares Hub, CI-with-a-probe-cap, Foundry, and a human engagement.
- Mindgard is automated AI red teaming for models, tools, data, and workflows, including a CI path. It is a specialist offensive-security console rather than one of the five delivery models above.
- F5 AI Red Team is the former CalypsoAI testing line, sold with F5 AI Guardrails. Evaluate it if F5 already owns that runtime, not as a standalone row on this shortlist.
Questions before you sign an AI red teaming tool
If a quote cannot answer these three, you are still buying a demo.
- Is this software you replay after the next prompt change, or a human engagement with a dated report?
- Which attacks are in scope: single-turn prompts, multi-turn jailbreaks, RAG documents, or agent tool calls?
- Is any price public, what is a probe or a scan, and who pays for the attacker model’s tokens?
Which AI red teaming tool should you pick
If you want a number you can budget this week, start with Promptfoo Community or PyRIT. If the agent is a text API and you want multi-turn attacks generated for you, open Giskard Hub. If Check Point already owns the rest of the stack, ask for AI Red Teaming on that quote and confirm it is out of limited release. If the artifact has to land in an ISO 42001 or board pack, Group-IB is the human path. MCP gateways and runtime allowlists are a different purchase: see MCP security tools.
Frequently asked questions
Is AI red teaming the same as a penetration test?
No. A classic pentest targets networks, apps, and identity. AI red teaming probes how a model or agent behaves under adversarial prompts, poisoned documents, and unsafe tool use. Some vendors sell both as separate services. Confirm which statement of work you are signing.
Which of these publish a price?
The open-source scanners do. Hub platforms and the human engagement are sales quotes. Community plans still burn tokens on the attacker and scorer models you bring.
Is Lakera Red still a product after Check Point?
Check Point closed Lakera into AI Security. The living name on Check Point’s site is AI Red Teaming. The Lakera Red page still describes that automated workflow. Ask for the AI Red Teaming line, not a retired standalone contract.







