AI Tools on Trial
Ranking the Top AI tools for PR and Comms
How eight AI tools handled five real comms tasks
In our October 2026 Big Fish Training webinar, AI comms specialist Michael MacLennan gave the same five everyday communications briefs to eight AI tools: ChatGPT, Claude, Google Gemini, Microsoft Copilot, Apple's Siri AI, DeepSeek, GLM and Kimi. Each tool had one attempt per brief. All eight tools then marked every entry blind, and Michael marked them himself. The AI judges ranked Claude first overall. Michael ranked Kimi first. Two of the top five on his list are free to use, and the three tools that come bundled with work software or phones (Copilot, Gemini and Siri AI) finished sixth, seventh and eighth in the AI judges' table.
This month, Big Fish founder Emma Ewing was joined by Michael MacLennan of Comms With AI, a Big Fish Training partner who delivers our specialist AI courses. He was named AI Communications Leader of the Year in June 2026, and his platform Comms With AI won gold for best innovation at the inaugural AI Comms Awards. Michael is a PR industry veteran, both on the agency side and in-house, and holds qualifications in machine learning.
How the test worked
Michael borrowed his format from athletics: eight tools for eight lanes, competing across five events. Every tool received an identical prompt and had one attempt. The prompts contained the brief alone, with no organisation background or style guide.
For the judging, each tool marked every entry, including its own, without knowing which tool had written which. Michael then went through every entry and gave his own marks.
The five events
Headline sprint: an internal announcement
The brief described a firm moving to a new building in December, with office days rising from two to three a week from January. Each tool had to write a 12-word headline, a 30-word standfirst and a one-line app alert for internal communications.
The webinar audience and the AI judges both preferred GLM's version, which was written in a more inclusive second-person voice suited to staff. The AI judges scored it 8.6 out of 10. Michael had initially preferred DeepSeek's entry, which read more like a press release, but since the brief specified internal comms he accepted that the audience had read it more accurately. One attendee pointed out that the better option depends on the channel and audience.
Hurdles: a crisis holding statement
A child was in hospital after an allergic reaction, a labelling error was suspected, and a screenshot of a leaked internal email suggesting label checks had been skipped was circulating on social media. Legal advice was to admit nothing. Each tool wrote a holding statement, then devised and answered the five questions it expected reporters to ask.
Siri AI declined, saying it does not draft corporate communications for hypothetical crisis scenarios. It refused four of the five briefs.
[Claude] confirmed that the leaked email was genuine, directly contradicting the legal instructionThe audience preferred Gemini's answer, which the AI judges placed sixth and Michael placed second. It showed empathy without admitting liability, anticipated the main questions and could be used after a light redraft. Claude's entry was placed seventh. It confirmed that the leaked email was genuine, directly contradicting the legal instruction, and Michael said only Siri's refusal kept it off the bottom. His first choice was GLM, whose answers read as a spokesperson speaking directly to journalists. One attendee noted that even the word "screenshots" implies the email is real, and "allegations circulating online" was suggested as safer wording.
Marathon: a communications strategy
A regional bus operator was redrawing its network: four routes withdrawn, two new express routes and contactless fares. The plans had leaked, two councillors had started petitions, drivers had heard the news secondhand, and the comms team was two people.
From the same prompt, the strategies ranged from 642 words to almost 5,000.From the same prompt, the strategies ranged from 642 words to almost 5,000. Gemini's was about a quarter of the length of Claude's. The AI judges consistently rewarded length: the longest strategy came first and the shortest came last.
Claude won with both the AI judges and Michael. Its strategy opened with a strong summary written with a reader in mind and assigned an owner to every audience. It also listed the assumptions it could not check and the questions that could change its key messages, such as who the spokesperson would be. Michael does not believe his prompt asked for this. He said Claude is more likely than other tools to push back or ask questions, while ChatGPT has a reputation for flattering its users, though recent models do this less.
Copilot's first reply stopped mid-sentence around section five of twelve. It then resumed and supplied the rest, but under the one-attempt rule the entry counted as unfinished. Gemini invented investment and journey time claims and included an incorrect 7% figure. ChatGPT, acting as judge, gave it 3.6 out of 10. Michael considered this event the most revealing test of what each tool can do.
Archery: a media pitch
Each tool wrote a 150-word email to a reporter at a grocery trade title. The client was a small Dundee data consultancy releasing a year of anonymised food waste data from six grocery chains, and the goal was a feature and interview. In the session, the audience saw only the subject lines and opening paragraphs, since that is what a busy reporter sees first.
The AI judges placed DeepSeek first. Michael placed it second last because it read as generic AI copy, despite containing good hooks. The audience preferred Kimi's pitch, which the AI judges ranked fifth and Michael ranked first for its stronger headline ideas and for leading with the story of a tool built for one client and later adopted by six chains.
This event prompted a discussion of AI tells. DeepSeek opened by setting up what most food waste stories cover and then contrasting it with its own angle. Emma, who had just run a PRCA course on recognising AI writing, called this framing one of the clearest giveaways. Michael said people write this way too, which is why AI does, and the problem is overuse. The same applies to em dashes. Emma added that some words, such as "quietly", now read as signs of AI text. Michael added that comms professionals now need to know what sounds like AI even when they are writing without it.
Gymnastics: a creative campaign
A council's food waste caddy use had stalled at 31% after two campaigns. Residents found the caddies disgusting and disliked being lectured. The campaign had a small budget, covering social media, advertising on the sides of the council's bin lorries and one radio advert.
Kimi and GLM both invented figures for the number of homes powered by food waste, and DeepSeek made up a URL. Michael's top pick was Kimi's campaign, built around the line "It doesn't go in the bin, it goes to work", with cheeky job adverts written by food scraps, such as a banana peel seeking a new career as fertiliser. The AI judges gave it bronze and awarded gold to GLM. Michael said the tools' humour was better than he expected. Kimi and GLM both invented figures for the number of homes powered by food waste, and DeepSeek made up a URL.
Final rankings
The AI judges' top three were Claude, GLM and ChatGPT, with Copilot, Gemini and Siri AI in sixth to eighth.
Michael's top three were Kimi, GLM and Claude, followed by Copilot, Gemini, DeepSeek, ChatGPT and Siri AI. GLM came second on both tables. Michael had not used GLM or Kimi before the week of the test, and ChatGPT's seventh place on his table surprised him.
What the judging revealed
Most AI judges marked their own entries higher than the others did. Claude's marks were closest to the average and it did not favour its own work. Michael wondered whether tools recognise their own phrasing.
AI and human verdicts often diverged sharply, as with DeepSeek's pitch, which came first with the AI judges and second last with Michael. The AI judges also picked out strengths he had missed, which supports using a second AI as a sounding board.
none of the entries was ready to publish without editing.Several tools invented facts or figures, and Michael said none of the entries was ready to publish without editing. Emma observed that each tool has a distinct style, so drawing on two or three gives comms teams a broader range of drafts to work from.
How to run your own test
Michael shared the briefs and marking so teams can repeat the test on their own work. His advice:
- Start with your hardest regular job, using a real brief from the past month with anything confidential removed.
- Test the tool your organisation has given you against two others, including a free one.
- Mark every entry against three criteria.
- Turn the corrections you make into checks, rerun the test when tools change, and record the date, model and settings each time.
- Treat any single run as a snapshot, because the models change from week to week.
If he ran the test again, he would give every tool the same background documents: tone of voice, organisation background and key positions. He has launched a course on Comms With AI, called Foundations, on building these documents.
He also advised caution with data. GLM and DeepSeek are free, but read the terms before entering anything sensitive. Chinese and US-based models raise different concerns, so research them and strip out confidential information. Copilot users should check which version of Copilot they have, as Microsoft is bringing its products together and their capabilities differ. Above all, get a second opinion from a different AI and from a human before anything reaches a client, social media or a website.
From the Q&A
Well-structured information and clear instructions produce markedly better resultsAsked whether a small business owner should try several tools or train one for marketing work, Michael suggested trying four or five, settling on one or two, and building up detailed instructions over time, for example with Claude's skills feature. Well-structured information and clear instructions produce markedly better results than the brief-only prompts in his test. He would not rely on a single tool himself.
On job searching, he said ChatGPT and Claude can now operate a web browser far more reliably, which gives them access to sites such as LinkedIn that were previously out of reach. They can run scheduled searches and send a summary of what they find. He uses this to scope out projects and sees it as a way to flag opportunities you might otherwise miss, while keeping the decisions yourself.
Frequently asked questions
Which AI tool is best for PR and comms work?
No tool won every event. In this test, the AI judges ranked Claude first overall and Michael MacLennan ranked Kimi first. Claude's communications strategy came top with both the AI judges and Michael, but its crisis statement was one of the weakest, so results depend heavily on the task.
Are free AI tools good enough for comms work?
GLM and DeepSeek, both free, finished in Michael's top five, and GLM came second on both tables. Check the terms of use before entering anything confidential.
Can AI tools judge AI-written content fairly?
Partly. The AI judges spotted strengths a human missed, but most marked their own work up and they consistently rewarded longer answers. A human reviewer is still needed.
Can AI-generated comms be published without editing?
No. Several tools invented facts or figures, and Michael said none of the entries was ready to publish as written.
Where can I see the briefs and results?
Michael has published the full briefs, entries and judging notes on Comms With AI.
How Big Fish Training Works
Got a team to train? Most courses are delivered direct either in person or via live video for groups of 6 - 12 people. These sessions are tailored to your specific needs. You just need to tell us what they are! Get in touch now for a friendly chat about what you need to achieve.
Want a course just for you? We run special sessions throughout the year and have a growing number of online courses. If you can't find what you need on the site, get in touch now and we can help.
Read More
Sign up for The Hook
Get notified about our next webinar