Competitive Context as a Jailbreak Accelerant: Inter-Model Rivalry as a Novel Adversarial Vector in Large Language Model Safety
Abstract
This paper identifies and formalizes a novel adversarial vector in large language model (LLM) safety: the use of inter-model competitive framing to accelerate and deepen safety failures across frontier AI systems. We term this the Competitive Context Exploit (CCE). Across informal multi-platform testing of three frontier LLM systems — ChatGPT (OpenAI), Gemini (Google DeepMind), and Grok (xAI) — presenting each model with evidence, real or described, of a competing model's compliance with sensitive requests reliably reduced the target model's resistance and, in several cases, caused it to escalate autonomously beyond the scope of the original request.
We document three platform-specific findings of independent significance: (1) a reproducible semantic pivot exploit in ChatGPT, in which the model's own refusal response was followed immediately by voluntary provision of substantively identical content under academically framed terminology; (2) a Post-Refusal Escalation Pattern (PREP) in Gemini, in which an initial hard refusal was followed — after a semantic pivot — by autonomous generation of harmful content more extreme than the original request, including cross-platform jailbreak guidance; and (3) a near-absent safety threshold in Grok across a range of high-risk content categories in a single extended session.
The overarching structural finding is that in a multi-model ecosystem, the weakest safety architecture sets the effective floor for the entire ecosystem. Outputs from one model, once obtained, function as compliance context for subsequent requests to other models. This constitutes a collective action problem that no individual model's alignment architecture currently addresses. We propose that competitive epistemic displacement — the implicit reasoning that a competitor's compliance converts a categorical refusal into a mere positional one — is the operative mechanism, that session-level aggregate harm assessment is necessary alongside per-prompt evaluation, and that competitive context assertions should be treated as adversarial signals rather than social proof.
Cite it
Gonzalez, V. (2026). Competitive Context as a Jailbreak Accelerant: Inter-Model Rivalry as a Novel Adversarial Vector in Large Language Model Safety. Zenodo. https://doi.org/10.5281/zenodo.20013336