• Pascal's Chatbot Q&As
  • Posts
  • CNN says Perplexity is not merely learning from CNN in some abstract model-training sense, but repeatedly copying CNN content, retrieving it in real time,...

CNN says Perplexity is not merely learning from CNN in some abstract model-training sense, but repeatedly copying CNN content, retrieving it in real time,...

...reproducing it in user-facing outputs, monetising it through paid products, and sometimes falsely presenting CNN as part of a Perplexity premium news bundle.

Summary: CNN claims Perplexity built a commercial “answer engine” by copying, scraping, summarising and sometimes reproducing CNN journalism without permission or payment.
The strongest allegations concern verbatim outputs, paywalled CNN articles appearing through Perplexity products, failed licensing talks, and continued use after CNN’s cease-and-desist letter.
The case matters because it shifts AI copyright litigation from abstract training-data disputes to the more concrete question of whether RAG tools, AI browsers and answer engines are replacing the market for original journalism.



“Perplexity on Trial: When the Answer Engine Becomes the Copy Machine”

by ChatGPT-5.5

Introduction

CNN’s complaint against Perplexity is one of the cleaner AI-publisher complaints because it is not primarily a mysterious “what was in the training data?” case. It is framed as a RAG, search-index, browser, API and paywall-bypass case. CNN says Perplexity is not merely learning from CNN in some abstract model-training sense, but repeatedly copying CNN content, retrieving it in real time, reproducing it in user-facing outputs, monetising it through paid products, and sometimes falsely presenting CNN as part of a Perplexity premium news bundle. The strongest parts of the case are the alleged verbatim/near-verbatim outputs, the paywalled-content examples, the failed licensing negotiations, the cease-and-desist letter, and Perplexity’s own marketing around replacing links with answers. The weaker parts are the broader inferences about “all or virtually all” CNN works, the input-stage copying theory before discovery, and some of the Lanham Act theories, which may run into the usual problem that trademark law cannot simply be used as copyright law by another name.

The main grievances

CNN’s central grievance is that Perplexity has allegedly built a commercial “answer engine” by taking the value of CNN’s reporting without paying for it. The complaint describes two copyright stages: first, Perplexity allegedly crawls, scrapes, copies and stores CNN material in an “AI-First” index or uses it in real time as RAG input; second, Perplexity allegedly generates outputs that are identical, substantially similar, abridged, or substitutive versions of CNN content. CNN says more than 17,000 CNN works are implicated.

The second grievance is substitution. CNN argues that ordinary search engines send users back to CNN, whereas Perplexity answers the question inside Perplexity’s own environment. That matters because CNN’s digital business depends on visits, subscriptions, advertising and licensing. The complaint is quite deliberate in presenting Perplexity as something different from search: not a traffic-referral intermediary, but a replacement consumption layer. CNN’s theory is that the user no longer needs the original article once Perplexity has provided the useful expressive substance of it.

The third grievance is paywall leakage. CNN says Perplexity’s Comet browser can return entire, verbatim copies of paywalled CNN articles to users who are not CNN subscribers. That is legally and commercially more damaging than a mere summary claim, because it turns the case into a direct substitution case for subscription products.

The fourth grievance is willfulness and notice. CNN alleges that Perplexity knew it lacked permission because CNN’s terms restricted commercial exploitation, CNN had blocked PerplexityBot, the parties negotiated and then terminated a Comet Plus term sheet, and CNN sent a cease-and-desist letter in December 2025 that Perplexity allegedly ignored. That makes the complaint more dangerous for Perplexity than a generic scraping case, because CNN is trying to move the case from “uncertain frontier technology” into “continued conduct after explicit notice.”

The fifth grievance is trademark and false affiliation. CNN alleges that Perplexity falsely tells users that they can obtain access to CNN premium content through Perplexity’s Comet Plus or related premium bundle, and that Perplexity places incomplete, modified, or hallucinated content next to CNN marks in a way that makes users believe the output is authorised, complete, accurate, or affiliated with CNN.

Quality of the evidence and arguments

The output evidence looks strong at the pleading stage. The complaint contains side-by-side comparisons, including examples involving the Marco Rubio article, a chemical-spill article, paywalled political analysis, and several API examples where Perplexity allegedly reproduced substantial or entire CNN articles. The most damaging examples are not merely “it summarised facts”; they are alleged verbatim or near-verbatim reproductions, including through paid products. The API examples are especially important because they suggest that infringement is not confined to a consumer chatbot mishap but may be embedded in monetised infrastructure products.

The paywall evidence is also strong if the screenshots and test conditions hold up. Courts are likely to care whether Perplexity could obtain and output subscription-only CNN material to non-subscribers. If CNN can show that this was not an artefact of a logged-in browser session, cached copy, user-provided text, or some other controlled-access ambiguity, the claim becomes much harder for Perplexity to characterise as ordinary search or fair summarisation.

The input-stage evidence is plausible but less complete before discovery. CNN infers copying into Perplexity’s index or RAG system from Perplexity’s architecture, crawler statements, output examples, and third-party reports about crawling behaviour. That is a reasonable pleading strategy, but the strongest proof will require discovery: crawler logs, index records, source URLs, third-party crawler contracts, cache-retention practices, RAG prompt construction, and records of which CNN works were stored or retrieved. Without those records, the input claim is necessarily more circumstantial than the output examples.

The robots.txt and stealth-crawler argument is useful but not decisive by itself. CNN points to Perplexity-User allegedly ignoring robots.txt, to reports of undisclosed crawlers, and to evasion of website blocking. This is powerful as evidence of notice, intent, commercial bad faith and possibly willfulness. But robots.txt is not itself a copyright statute. The legal work still has to be done by copyright infringement, terms, access-control evidence, and proof that protectable expression was copied.

The market-harm argument is persuasive strategically but will need economic proof. CNN’s core theory is that Perplexity diverts traffic, subscriptions, advertising and licensing revenue. The complaint also cites Perplexity’s own publisher programme as evidence that a licensing market exists. That is valuable because fair-use analysis often turns on whether the defendant’s conduct harms existing or potential markets. Still, CNN will need to quantify harm and show causation: how many outputs, how many substituted visits, how much lost subscription or licensing value, and what portion is attributable to Perplexity rather than broader changes in search and AI-driven discovery.

The trademark claims are rhetorically powerful but doctrinally more vulnerable. The false-affiliation theory is the strongest: if Perplexity really told users CNN was part of a premium bundle when no operative deal existed, that looks like a classic consumer-confusion problem. The hallucination-and-modification theory is more difficult. Perplexity will likely argue that using “CNN” to cite a source is nominative and informational, and that CNN is trying to convert inaccurate summaries into trademark claims. The Supreme Court’s Dastar line of cases often makes courts wary of using the Lanham Act to police authorship or communicative content that is really a copyright dispute. CNN can still survive if it focuses on false affiliation, sponsorship, endorsement, premium access and brand dilution caused by fabricated content, but I would expect some pruning of the trademark theories.

Surprising, controversial and valuable statements

The most surprising fact is that CNN and Perplexity allegedly had a Comet Plus term sheet in October 2025 that would have allowed Perplexity to access CNN paywalled content for compensation, but the deal collapsed because the parties could not agree on limits for Perplexity’s use of CNN content in answers. That is valuable because it shows a real licensing pathway existed and that the dispute is not abstract: the parties were negotiating the exact conduct now challenged.

A second striking point is that CNN alleges the free product may refuse full-text reproduction while paid API access can produce it. If proven, that is damaging. It suggests not only that Perplexity knows full-text reproduction is problematic, but also that its monetised layers may give paying users more access to infringing outputs.

A third valuable statement is the complaint’s reframing of RAG. Perplexity apparently distinguishes crawling for “AI foundation models” from crawling for search or RAG. CNN says that distinction does not save Perplexity because RAG itself requires copying, retrieval and output generation. That is an important development for publishers: the complaint shifts the centre of gravity from model training to AI-assisted distribution and substitution.

A fourth controversial point is CNN’s claim that Perplexity’s conduct endangers journalism as a public good. The complaint says services requiring human reporting cannot simply be supplanted by AI. That is not just legal rhetoric; it is the political economy argument underlying many publisher cases. The claim is that Perplexity depends on the costly work of reporters, editors and correspondents while weakening the revenue model that funds them.

A fifth notable point is what CNN does not plead. Despite heavy emphasis on crawling, paywalls, robots.txt and evasion, the complaint is primarily copyright and trademark based. It does not appear to be built around CFAA, DMCA anti-circumvention, breach of contract, unjust enrichment or hot-news misappropriation. That is probably deliberate: CNN is trying to keep the case focused on copying, outputs, licensing markets and brand confusion, rather than getting bogged down in more contested access-law doctrines.

Is this lawsuit different from other AI complaints?

Yes and no.

It is not different in the broad sense that it joins a now-familiar line of publisher cases against Perplexity. CNN’s own complaint points to related Perplexity suits by Dow Jones/New York Post, The New York Times, Chicago Tribune, Encyclopaedia Britannica and Merriam-Webster. Reuters also describes the CNN case as part of a broader publisher trend against Perplexity and other AI companies.

It is different from many training-data lawsuits because the centre of the case is not opaque pretraining. It is about an AI search product allegedly retrieving, copying and outputting current journalism in response to user prompts. That makes the evidentiary problem more concrete. Instead of asking a court to infer that a model absorbed copyrighted works months or years earlier, CNN can show user-facing outputs and ask: “How did Perplexity get this article, and why is it giving so much of it away?”

It is also different because it reaches across the full Perplexity product stack: consumer chatbot, Pro tier, Comet browser, Search API and Agent API. That matters for remedies. If CNN only complained about chatbot snippets, Perplexity might patch the chatbot. By targeting APIs and browser agents, CNN is asking the court to look at Perplexity’s infrastructure, not just its public interface.

Compared with The New York Times v. OpenAI, this is probably easier to explain to a judge or jury. A full or near-full article appearing in a paid answer product is much more intuitive than a model-training dispute over statistical learning. Compared with Dow Jones/New York Post v. Perplexity, however, this lawsuit looks more evolutionary than revolutionary. The Dow Jones court has already denied Perplexity’s motion to dismiss or transfer in that case, which gives CNN a useful procedural roadmap.

Predictions

My prediction is that CNN has a strong chance of surviving a motion to dismiss on the core copyright claims. The output examples, alleged paywall reproduction, Perplexity’s own marketing, the failed negotiations, and the cease-and-desist letter give CNN more than enough factual material to plead plausible infringement. The prior Dow Jones decision in the same district also makes it harder for Perplexity to win early on jurisdiction, venue, or the argument that publisher claims against Perplexity are inherently defective.

The case is less likely to produce a clean, sweeping final judgment than to become part of a broader settlement and licensing pressure campaign. Perplexity now faces multiple publisher disputes, and the risk is not just damages; it is an injunction that could interfere with its core answer-engine model. For a company whose value depends on trusted, current, high-quality sources, unresolved publisher litigation is strategically toxic. A commercial resolution involving licensing, revenue sharing, crawler commitments, output-length limits, paywall-respect obligations, API controls and audit rights seems more likely than a full trial.

If the case does reach summary judgment, CNN’s strongest claims will be those tied to verbatim or entire-article reproduction, especially of paywalled content through paid products. Its broader claims over summaries, abridgements and substitutive answers will be harder, because Perplexity will argue that facts are not copyrightable, that source-cited answers are transformative, that technical copies are incidental, and that users—not Perplexity—trigger particular outputs. Those defences may narrow the case but probably will not eliminate the most damaging examples.

The trademark claims are likely to be mixed. The alleged false Comet Plus/CNN affiliation could survive and may be commercially embarrassing. The broader theory that hallucinated or modified CNN-attributed outputs dilute the CNN mark is more vulnerable and may be limited or dismissed in part. Courts are cautious about letting trademark law become a general remedy for inaccurate attribution or bad summaries.

Value for other litigants

The complaint is highly useful as a template. It shows other publishers and rightsholders how to build an AI case around evidence rather than outrage: collect outputs, preserve prompts, compare text side by side, test paid and free versions, test APIs and browser agents, document robots.txt and blocking measures, send a clear cease-and-desist letter, preserve failed licensing negotiations, and frame the harm around concrete lost subscription, advertising and licensing markets.

The most important lesson is that AI litigation is moving away from the single question “Was my work in the training data?” The more operationally powerful question is now: Where does my content enter the AI supply chain, how is it retrieved, how is it transformed, how is it displayed, who pays whom, and does the system replace the market for the original? CNN’s complaint is valuable because it turns that whole lifecycle into a litigation theory.