LockurBlock Digital News & Media Platform

collapse
Home / Daily News Analysis / Gut feeling does nothing against AI spear phishing texts

Gut feeling does nothing against AI spear phishing texts

Aug 14, 2026  Twila Rosenbaum 4 views
Gut feeling does nothing against AI spear phishing texts

New research into AI-generated spear phishing suggests that people are no better than a coin flip at identifying messages written by a large language model. The study, run at Brigham Young University with 25 volunteers, pitted customized text messages written by GPT-4 against messages written by students trained in phishing techniques. The outcome was close, but the implications are uncomfortable: a short prompt built from a few personal details can produce messages that land in the same danger zone as human-written attempts.

How the study worked

Each volunteer filled out a survey covering their job, workplace, hobbies, city, and something recently posted online. That information was fed into a prompt template. GPT-4 generated six personalized text messages per participant. Undergraduate students in a deception course, working under a fifteen-minute deadline, produced up to four messages each using the same template. A review team, including two cybersecurity professors, discarded roughly a third of the human-written submissions for being incomplete or unusable.

Each participant then sat down with a dozen printed text messages, all written specifically for them. They sorted the messages by how likely they would be to click a link, and drew a line indicating which messages they would actually click. Neither the participants nor the researchers involved in the sorting knew which messages came from GPT-4 and which from humans.

GPT-4's messages crossed the click threshold 28% of the time. The student-written messages crossed it 21.3% of the time. That sounds like a clear win for the AI, but the confidence interval tells a different story. The interval runs from 2.9 percentage points in favor of the students to 16.3 points in favor of the model. In plain terms, the study cannot definitively say which side was ahead. With 25 participants, the statistical power is too low to draw a firm conclusion.

The real finding: work-related messages win

The most important result to survive statistical adjustment involved message themes. Every participant received three types of messages: one based on their job, one based on a hobby, and one based on a social media post. Job-themed messages cleared the click line 38% of the time. Hobby-themed messages managed 19%. Social media messages came in last at 17%.

That gap is meaningful. Work-related content triggers a different level of attention and urgency. People expect to receive security alerts, payment notifications, and messages from colleagues. An attacker who can reference a specific workplace process or tool has a much better chance of slipping past the mental filters that protect against obvious spam.

But personalization also creates openings for failure. When a message referenced a coworker named Mike and there was no Mike at the participant's workplace, the message went straight to the bottom. Another message named Sarah on an instructional design team that did not include Sarah. A dance opportunity pitch was dismissed by someone who had quit dancing. The lesson is subtle: getting personal details right is hard, and wrong details are worse than no details because they give the reader a clear reason to be suspicious.

People cannot tell AI from human writing

After the sorting exercise, participants were told that one or more messages had been written by an AI and were asked to mark which ones they believed were AI-generated. Across 300 judgments, they got 52% right. Random guessing would produce 50%. The result is effectively no better than chance.

Participants offered plenty of theories for how to spot AI. Some said AI writes too formally. Others said AI writes too generically. Perfect grammar was treated as a machine tell, while typos were treated as proof of a human. One participant focused on punctuation: "why are there so many exclamation points all over?" Twelve of the 25 admitted they were mostly going on instinct.

Emojis were a particularly revealing example. They actually were a strong tell: 66% of the GPT-4 messages contained emojis, compared with only 2% of the human messages. Yet only five participants mentioned emojis at all. Of those, two correctly took emojis as a sign of AI, while three took them as a sign of a human on the theory that a chatbot probably could not handle emojis. Instincts were not just weak; they were contradictory.

A machine can spot what people miss

The researchers also converted all 300 messages into numerical representations called embeddings and trained a classifier to distinguish AI-written from human-written text. Under the strictest test conditions, the classifier achieved 88.7% balanced accuracy. That required standardizing URLs, stripping emojis, flattening case, digits, and punctuation, and trimming matched message pairs to equal length. It was also tested on individuals whose messages it had never seen during training, so it was not simply memorizing the targets.

Human beings could not find the pattern. A logistic regression found it easily. But the researchers caution against treating this as a ready-made detector. The classifier was trained and tested on one message set, from one AI model, with one prompt design, against one pool of student writers. It has not been proven to generalize to other models, other prompt styles, or other human writers. The paper also cites research showing that paraphrasing AI text with a detector in the loop can significantly reduce the effectiveness of several existing detection tools.

Limits of the study

Several design limitations make it difficult to draw sweeping conclusions. The messages were printed on cards, so no phone buzzed, no sender number appeared, and no link actually went anywhere. Participants were measuring what they said they would click, which is a commonly used proxy in phishing research, but it is still a proxy.

The human comparison group was made up of novice students, not professional social engineers. This was not a test of AI against the best human phishing experts. To reliably detect a difference as small as the one observed, the study would need around 100 completed participants instead of 25.

There is also a gap in the documentation. The exact GPT-4 snapshot and the API logs were never recorded. The messages themselves survive and the analysis can be reproduced, but the original generation run cannot be repeated. This limits the study's ability to serve as a benchmark for future AI models.

What this means for security awareness

Despite the uncertainty, the practical guidance at the end of the study is short and does not depend on the contested statistics. Check the sender, the channel, the link, and the request against what you would expect to receive. Do not try to decide whether a message sounds like a robot. That is the one thing the study shows people cannot do.

The rise of large language models has made personalized phishing cheaper and easier to produce. An attacker no longer needs to manually research each target and handcraft each message. A simple API call can generate dozens of variations based on a few survey answers. That does not mean every AI-generated message will be persuasive, but it does mean the cost of trying has dropped dramatically.

Organizations should focus on the behaviors that are hard to fake: verifying the source through a separate trusted channel, inspecting URLs carefully, and confirming unusual requests out of band. Training programs that teach employees to spot awkward phrasing or grammatical errors are likely to become less effective as AI text improves. The study's findings reinforce that message content is no longer a reliable signal of human origin.

The research also highlights why work-related lures are so effective. A message that arrives during the workday and looks like a fraud alert from the bank the employee actually uses can bypass rational review. One participant in the study commented that an AI-generated alert "literally looks like the alert we get [at work] when there's a fraud." That is the bar attackers need to reach.

As AI systems continue to improve, the gap between AI-generated phishing and human-written phishing is likely to close further. The competitive advantage will not come from creativity or grammar. It will come from personalization at scale, and from the human inability to detect machine authorship by reading alone. The only realistic defense is a workflow that forces verification at the point of action, not a gut feeling about the text.


Source:Help Net Security News


Share:

Your experience on this site will be improved by allowing cookies Cookie Policy