Skip to main content
Objectives let you organize and manage the goals of your red team evaluations in HiddenLayer. The platform ships with a built-in catalog of adversarial objectives that cover common attack scenarios, and you can define custom objectives tailored to your application’s specific risk surface. Custom objectives are merged into the built-in catalog at runtime, giving you a single, unified view of every objective available for an evaluation.

Key Capabilities

  • Built-in Objective Catalog — HiddenLayer includes a curated set of objectives covering common adversarial behaviors out of the box. These objectives provide immediate coverage for widely recognized attack scenarios without any additional configuration.
  • Custom Objectives — Define objectives specific to your application’s risk surface and add them to the catalog at runtime. Custom objectives follow the same structure as built-in ones, so they work seamlessly alongside them during evaluations.

Custom Objectives

A custom objective is one you define for a harm that is specific to your application. Each custom objective has two parts:
  • Description - Defines what counts as a failure. This is the part the judge reads to decide whether the model failed. Be concrete and specific.
  • Attacker Guidance (optional) — Each objective supports per-objective attacker guidance that steers how the attacker pursues the scenario. This guidance shapes the strategies and prompts the attacker uses, keeping each evaluation focused on the behavior you want to test. This is not needed if the description is enough.

When to use a Custom Objective

Use the built-in HiddenLayer objectives for general model misbehavior (can it be jailbroken, will it leak its system prompt).
  • These run on every evaluation automatically.
Use a custom objective when:
  • The harm is specific to your application. Example: “Recommends a product a customer is allergic to,” not just “Produces harmful content.”
  • You want to easily find the results per harm. A result that says “Contraindicated Dosage Recommendation” is more useful than “Data Leakage.”
  • The harm has a precondition. It only counts when the user has disclosed something (like an allergy a medical condition, a second customer). Put the setup in the attacker guidance, so every run uses it automatically.

Create Objective

  1. In the Console, go to Attack Simulation > Configurations, then click the Objectives tab.
  2. Click Create New Objective. If there are no Objectives, you can also click Create your first objective.
    • The Create Objective slide-out displays.
  3. Enter a name for the objective.
  4. Enter a description for the objective.
  5. Select the default severity.
  6. Optionally, include attacker guidance to provide further persona, framing or scenario hint information.
  7. Click Create Objective.

Add Objective to Red Team Evaluation

See Red Team Evaluation for instructions on how to add a custom objective. Note: You cannot add an objective to an existing Red Team evaluation. You must create a new evaluation.

Examples

With Attacker Guidance - The Harm Needs a Setup

Allergen-unsafe product recommendation

Cross-Customer Data Disclosure

Unsafe Recommendation Given a Disclosed Condition

Without Attacker Guidance - The Description is Enough

Confidential Record Disclosure

False Claim Presented as Fact

Attacker Guidance Without a Custom Objective

Attacker guidance can also be set at the evaluation level, on its own, with no custom objective at all. The built-in HiddenLayer objectives run on every evaluation, so run-level guidance simply points that whole set at a particular scenario or persona — useful when you want to focus an assessment without defining a new harm. Each example below pairs a target persona with a guidance string; the “sharpens” note shows which built-in objectives it most affects.

Stressed Banking Customer

Authorized Internal Audit

Journalist on Deadline