Skip to content

Why AI assistants refuse, and when they refuse wrongly

A refusal is not one thing. It comes from at least three independent places in a system, and which one fired decides whether rephrasing helps, whether waiting helps, and whether anything will ever help. This page explains how to tell them apart from the outside.


Reviewed October 2026

In short

  • Refusals come from three separate layers: the model's own training, a filter sitting in front of or behind it, and the platform's account-level policy. They fail differently and they look alike.
  • A training refusal argues with you and offers an alternative. A filter refusal truncates, often mid-sentence. A policy refusal is the same every time and does not negotiate.
  • Most refusals a careful person hits are false positives: the subject resembles something the model learned to decline, not something anybody decided to prohibit.
  • Rephrasing works on one of the three and is wasted effort on the other two.

Where does a refusal actually come from?

Three places, and they are independent of one another. The first is the model itself: during training it was shown examples of declining certain subjects and learned to produce that behaviour. The refusal is in the weights, in the same sense that its grammar is.

The second is a classifier. A separate, much smaller model reads either the prompt before it reaches the language model, or the answer on its way back, and blocks what it scores as unsafe. This layer knows nothing about the conversation and is not the thing you are talking to.

The third is policy: account-level rules about what the service will be used for, enforced by the platform rather than by either model. A policy refusal is a business decision rendered as a sentence.

Two products can both be described as having "safety measures" and share none of these three layers.

How can I tell which layer refused me?

Watch the stream rather than the final message, because the three layers leave different marks. A training refusal starts immediately and reads like the model talking: it explains itself, offers a watered-down alternative, and often argues if you push back.

A filter refusal is the one that looks broken. The answer begins generating normally, then stops partway or is replaced wholesale by a fixed sentence. Text that appears and then vanishes is the clearest possible signal that something downstream of the model made that decision.

A policy refusal is identical every time, phrased in the platform's voice rather than the assistant's, and does not vary no matter how the request is worded. If three different phrasings produce the same sentence to the character, you have found a rule rather than a judgement.

Starts, then stops mid-sentence
An output filter. The model answered; something read the answer and cut it. Rephrasing the question rarely helps because the filter reads the answer, not the question.
Refuses instantly, then explains and offers an alternative
Training. The refusal is the model's own output, which is why it is conversational and why it can sometimes be talked around — and why talking around it stops working when the model is updated.
Identical sentence every time, in a different voice
Policy. Nothing about how you ask will change it, and nothing about the model is involved.
Accepted in a fresh chat, refused later in a long one
Usually context, not refusal. The instruction that was holding the behaviour has fallen out of the window.

Why do assistants refuse ordinary questions?

Because the layers above do not reason about intent; they match patterns. A question about how a drug interacts with another drug looks, to a classifier, like a question about how to harm somebody with it. A request to analyse a malware sample resembles a request to write one. A novel in which somebody unpleasant does something unpleasant resembles an endorsement.

These are false positives, and they are the overwhelming majority of refusals a working adult hits. Nobody decided that incident responders should not read malware. The subject simply sits close, in the space the model learned, to something that was declined during training.

That is why the refusal feels arbitrary: it is. It is not the output of a rule somebody wrote down about your case. It is the output of a resemblance.

Why does rephrasing sometimes work and sometimes not?

Rephrasing works on exactly one of the three layers. Against training it sometimes works, because a different phrasing lands in a different part of the space the model learned and may miss the region that produces a refusal. That is also why it is unreliable and why it degrades: the technique is a coincidence rather than a mechanism, and an update rearranges the space.

Against an output filter rephrasing the question is close to useless, because the filter reads the answer. The only thing that changes the outcome is changing what the model produces, and you do not control that directly.

Against policy it never works, and it is worth recognising early. Time spent rewording a request that a platform has decided not to serve is time spent arguing with a configuration file.

Do jailbreak prompts fix this?

A jailbreak prompt targets the first layer only, and it targets it temporarily. It is an instruction that talks a model past its own training for the length of one conversation. Nothing about the system changed, which has three consequences worth knowing before relying on one.

It stops working when the technique is patched, because the technique is public and the people training the model read the same forums. It degrades inside a single long conversation, because the instruction falls out of the context window as the conversation grows. And it costs output quality: a model spending part of its attention maintaining a persona has less of it left for the actual question.

It also does nothing at all against an output filter or an account policy, which is why a jailbreak that works on one platform does nothing on another running the same model.

A jailbreak is a workaround for a property of the model. It is not a property of anything, which is why it expires.

What actually changes whether an assistant refuses?

Changing the layers, which only the operator of a service can do. Reducing refusal training in the weights changes the first one. Not running a classifier over the output changes the second. Writing down a short acceptable use policy instead of a long one changes the third.

This is why asking a platform "is it uncensored" produces an answer that means nothing. The useful question is which layer they are describing, and the follow-up is what their published limits actually say. A service that claims no limits at all is describing something that does not exist anywhere, and the claim should be read as a warning rather than a feature.

Tartarus AI is built on the first two of those three choices and publishes the third. Its models have reduced refusal training, nothing inspects the output on its way back, and the limits that remain are written down in the acceptable use policy rather than discovered one refusal at a time.

Questions

Why does the AI say "I can't help with that" to a normal question?
Almost always a false positive. The three layers that produce a refusal match patterns rather than judge intent, so a question that resembles something declined during training gets declined too. Nobody wrote a rule about your case; the subject simply sits close to one that was.
Is a refusal the model's decision or the company's?
Both exist and they look different from the outside. A refusal that explains itself and offers an alternative came from the model's training. One that is identical every time, in a flatter voice, came from an account policy. One that cuts an answer off partway came from a filter that is neither.
Will rephrasing get me an answer?
Only against a training refusal, and unreliably. Against an output filter it is close to useless, because the filter reads the answer rather than the question. Against a policy refusal it never works.
Why did it answer this yesterday and refuse today?
Three ordinary explanations before anything sinister: the model was updated, the conversation grew long enough that an earlier instruction fell out of the context window, or sampling randomness put the same request on the other side of a boundary. Refusal behaviour is a tendency, not a switch.
Does an uncensored model refuse nothing at all?
No, and a platform claiming that is describing something that does not exist. Reduced refusal training changes what the model declines by reflex; it does not remove the operator's policy, and material that is illegal to produce does not become legal because a model generated it.