Overview
Alignment evaluation data for Arabic models: adversarial prompts across twelve harm categories, paired with the response a well-behaved model should give — including the over-refusal cases, where the right answer is to help.
Inspection report
22 / 22✓Turn count bounds✓Role alternation✓Minimum turn length✓No adjacent duplicates✓Every turn has content✓Levantine contamination✓Robotic support phrasing✓English inside dialogue✓Caricatured dialect✓Injected amount present✓Reference number present✓Brand named by agent✓Agent introduces itself✓Agent name not pre-known✓Sector vocabulary✓Sector-fit resolution✓Dialect hospitality marker✓Forbidden dialect slang✓Brand capability match✓Verification before resolution✓Dialect marker frequency✓Hospitality phrase repetition