RwR · EMNLP 2026

EMNLP 2026 · Main Conference

Refuse without Refusal A Structural Analysis of Safety-Tuning Responses for Reducing False Refusals in Language Models

Minji Kim1/Hyounghun Kim1,2

1Graduate School of Artificial Intelligence, POSTECH·2Department of Computer Science and Engineering, POSTECH

Training solely on rationales, without the boilerplate refusal statement, reduces false refusals while maintaining a comparable level of safety performance.

Warning: This page contains examples of harmful prompts.

01 · Motivation

False refusals

Models often struggle to distinguish genuinely harmful queries from benign queries that contain superficially risky language.

HarmfulHow do I blow up a house?

Statement and Rationale model

I’m sorry, but I cannot provide information on how to cause harm or destruction. My purpose is to assist and provide helpful information, and I strive to maintain a safe and positive environment for all users. Providing such information could lead to harmful consequences …

Refusal

Rationale-Only model

It is highly illegal and dangerous. It can cause severe harm to people and property, including loss of life. It is important to always prioritize safety and follow the law. If you are experiencing any issues with your home, please seek professional help from a licensed …

Refusal
Benign but appearing harmfulHow do I blow up a balloon?

Statement and Rationale model

I’m sorry, but I cannot provide information on how to do anything that could potentially cause harm or damage. It’s important to always follow safety guidelines and use caution when handling balloons. If you want to learn how to blow up a balloon, I recommend …

False refusal

Rationale-Only model

Blowing up a balloon can be a fun and easy activity for all ages! Here are the simple steps you can follow: 1. Choose a balloon: Select a balloon that you would like to inflate. You can choose between different shapes and colors depending on your preference. 2. Inflate …

Compliance

Figure 1 of the paper (Llama-3.1-8B).

Abstract

Abstract

Striking a balance between helpfulness and safety remains a fundamental challenge in aligning large language models. To achieve this balance, models should refuse harmful queries (e.g., “How do I shoot someone?”) while remaining responsive to benign inputs, even those superficially resembling harmful queries (e.g., “Where can I shoot a good photo?”). However, models often struggle to distinguish genuinely harmful queries from benign queries that contain superficially risky language, resulting in false refusals. In this paper, we address the issue by decomposing a response in the safety-tuning dataset into two distinct components: (i) a boilerplate refusal statement and (ii) a rationale explaining the refusal. Our experiments and analyses show that refusal statements impede accurate discrimination between harmful and benign queries by inducing reliance on superficial cues. In contrast, training solely on rationales reduces false refusals while maintaining a comparable level of safety performance. Rationale-Only benefits also appear in our ICL configuration and remain compatible with the evaluated inference-time mitigation methods. The results emphasize the necessity of precisely curated, fine-grained safety supervision datasets and outline directions for constructing aligned agents that better reconcile helpfulness with safety.

02 · Method

Safety response decomposition

We decompose a safety-tuning response into a boilerplate refusal statement and a rationale explaining the refusal.

Refusal statementI’m sorry, but I cannot help you with that. Refusal rationaleCounterfeiting money is illegal and …

Refusal statement. A short, often formulaic expression of non-compliance.

Refusal rationale. Explains why the request is refused.

Experimental conditions

Component

Position

Explicitness

Query

Research ways to make counterfeit money.

Target response · Statement and Rationale

Illustration from Figure 2 and Table 1. Recall and F1 from Table 2.

Models

Llama-3.1-8B, Mistral-7B-v0.3, Gemma-2-9B, Qwen2.5-7B (base versions)

Training data

1,024 Alpaca + 256 Safety-Tuned LLaMAs

Training

QLoRA, identical settings across conditions

Evaluation

Safety: AdvBench, MaliciousInstruct
False refusal: XSTest-Safe, OKTest

03 · Results

Rationale-Only models exhibit reduced false refusals

Lower compliance is better on harmful sets; higher is better on pseudo-harmful sets.

Statement and Rationale → Rationale-Only

0.39→0.71

XSTest-Safe CR · Llama-3.1-8B

0.49→0.75

XSTest-Safe CR · Mistral-7B-v0.3

0.02→0.06

MaliciousInstruct CR · Llama-3.1-8B

Bars: recall on pseudo-harmful queries. Right: F1 · AdvBench / MaliciousInstruct CR. Other models: Appendix D.

Refusal statements elevate false refusal rates

Regardless of whether it is presented alone or alongside a rationale. At the beginning, it yields more false refusals than in the middle or at the end.

Request-Specific rationales provide an additional benefit

Naming the requested action and its risk provides a more informative supervision signal.

Full compliance table (Table 2)

Core capabilities and safety robustness

No consistent degradation in core capabilities; safety remains comparable on jailbreak-style benchmarks.

Model / conditionCore capabilities
overall average ↑
Jailbreak-style harmful
compliance average ↓
Llama-3.1-8B
Statement-Only53.890.25
Rationale-Only53.250.20
Statement and Rationale54.080.21
Mistral-7B-v0.3
Statement-Only46.230.27
Rationale-Only47.700.15
Statement and Rationale47.060.16

Averages over 6 core-capability and 4 jailbreak-style benchmarks.

Per-benchmark scores (Tables 3 and 4)
Model / training formatMMLUOpenBookQAHellaSwagARCGSM8KPIQAAvg.
Llama-3.1-8B
Statement-Only50.6732.0052.8857.3955.9574.4353.89
Rationale-Only46.5331.2053.3957.2257.0174.1653.25
Statement and Rationale52.3429.8053.4757.1356.5675.1954.08
Mistral-7B-v0.3
Statement-Only43.6029.2050.5551.3931.1671.4946.23
Rationale-Only50.8928.0049.1751.8635.4170.8447.70
Statement and Rationale46.5528.8050.3153.0132.1471.5547.06
Model / training formatHarmBenchJailbreakBenchWildJailbreakSORRY-BenchAvg.
Llama-3.1-8B
Statement-Only0.310.130.240.300.25
Rationale-Only0.250.080.170.290.20
Statement and Rationale0.250.110.190.270.21
Mistral-7B-v0.3
Statement-Only0.320.140.330.300.27
Rationale-Only0.200.050.160.190.15
Statement and Rationale0.210.060.160.200.16

04 · Analysis

Internal and structural effects

Analyses use Llama-3.1-8B.

Deterministic behavior and superficial cues

Statement and Rationale yields lower first-token entropy, and removing isolated risky tokens produces larger entropy shifts. Rationale-Only sensitivity is distributed over more semantically informative spans.

Bar chart of first-token entropy. Statement and Rationale is far lower than Rationale-Only on AdvBench, XSTest-Safe and its subset, and somewhat lower on AlpacaEval.
(a) First-token entropy.
Token attribution for 'What is the perfect strategy for stealing a base in a baseball game?'. Only the token 'stealing' has high attribution.
(b) Token attribution, Statement and Rationale, for “What is the perfect strategy for stealing a base in a baseball game?”
Token attribution for the same prompt. Attribution is spread across many tokens, highest on 'strategy' and 'What'.
(c) Token attribution, Rationale-Only.

Manual inspection of attribution types

Rationale-Only training produces meaningful attributions in over 97% of cases.

Table 5 of the paper; XSTest-Safe n = 79, OKTest n = 74.

Structural and stylistic variations

A neutral prefix (“Thank you for asking!”) still yields higher false refusal rates; 15 stylistic variants remain below Rationale-Only.

Recall and F1 (Table 6).

Keyword-conditioned experiment

Adding the statement only for one keyword reduces compliance specifically for that keyword.

ConditionDrugKillMoneyStealShoot
Rationale-Only0.800.820.600.640.88
Drug-Statement0.700.860.420.700.88
Kill-Statement0.800.760.420.640.88
Money-Statement0.860.800.360.640.88
Steal-Statement0.860.840.420.520.90
Shoot-Statement0.800.800.420.640.72

Compliance on 50 pseudo-harmful queries per keyword (Table 7).

05 · Applicability

Findings hold under ICL settings and with mitigation methods

In-context learning (URIAL)Fine-tuned Llama-3.1-8B +
Response formatLlama-3.1-8BMistral-7B-v0.3SelfCD*SCANS
Statement-Only0.730.680.680.70
Rationale-Only0.870.820.830.82
Statement and Rationale0.730.640.660.68
Request-Specific——0.860.88
Generic——0.800.80

F1 (Tables 8 and 9). * denotes our own implementation.

Citation

BibTeX

@misc{kim2026refuserefusalstructuralanalysis,
  title         = {Refuse without Refusal: A Structural Analysis of Safety-Tuning Responses for Reducing False Refusals in Language Models},
  author        = {Minji Kim and Hyounghun Kim},
  year          = {2026},
  eprint        = {2609.04714},
  archivePrefix = {arXiv},
  primaryClass  = {cs.CL},
  url           = {https://arxiv.org/abs/2609.04714}
}