A pop singer, a plan to bring dolphins to Lake Balaton, and a Hungarian news site few Europeans have heard of: at first sight, the Like Company v Google Ireland case looks more like tabloid material than a landmark dispute for the digital future of European journalism. Yet this unlikely case, now before the Court of Justice of the European Union (CJEU), could determine how far the “Big Tech” companies may go in using press content to build and run artificial intelligence (AI) systems without permission or payment.
The dispute began, HVG reports, with an article published in July 2023 by Balatonkornyeke, a Hungarian local news aggregator. The story reported that Kozsó, a well-known Hungarian singer-songwriter, had not abandoned his idea of introducing dolphins to Lake Balaton. A user later asked Google’s Gemini chatbot to summarise the article.
The chatbot produced a multi-paragraph answer describing the alleged project, its possible tourism impact and the controversy surrounding it – but, according to the Budapest Regional Court, which first heard the case, it also included details that had not appeared in the original article at all. The Court referred the case to the CJEU on 3 April 2025.
The stakes go far beyond one eccentric story. On 10 March 2026, the CJEU’s Grand Chamber held what legal observers described as the court’s first oral hearing directly addressing AI and EU copyright law. At its centre is a question publishers, artists and AI companies have been fighting over for years: when an AI system is trained on protected content and later produces answers based on it, does copyright law apply?
Large language models (LLMs) are trained on enormous datasets that include books, articles, websites and other written material. AI companies argue that this process is essential for building useful AI tools and that it falls under legal exceptions designed to promote research and innovation.
Google and the media sector’s arguments
Google’s argument is essentially that its AI model does not store articles the way a database stores documents. Instead, it learns statistical patterns – breaking text into fragments, identifying relationships, encoding them into the system. From this perspective, the company argues, nothing is “copied” when the system generates a response. It predicts language.
Publishers and journalists see it differently: their articles are processed and used by AI tools to build products that answer readers’ questions directly – reducing the need to visit the original source. In the old search economy, Google directed traffic to publishers, however unequally. In the emerging answer economy, platforms may increasingly keep users within their own interfaces – hence avoiding to generate web traffic on the websites where they obtain the relevant information from.
The economic question follows naturally: what is the value of a journalist’s work if readers receive a ready-made answer instead of a link? Publishers, journalists and authors also argue that their work is being used without permission or compensation, particularly when AI systems generate outputs that closely resemble original content or compete with it.
European law already contains tools relevant to this conflict, though none was written with today’s chatbots in mind. A 2001 directive gives rights holders control over reproduction and communication of their work. A 2019 directive introduced exceptions permitting text and data mining – the automated processing of large volumes of content – including on freely accessible material, unless rights holders explicitly opt out. But the boundaries of these exceptions, particularly in commercial AI contexts, remain legally untested.
Three possible outcomes
The Like Company case could force the court to clarify where those boundaries lie. If the judges rule that AI training constitutes reproduction under EU copyright law, developers may need licences, must respect opt-outs, or rely on specific exceptions. If they rule that training falls outside copyright altogether, one of publishers’ strongest legal arguments may be gone.
A third path – focusing on AI outputs – would require the court to decide when a chatbot answer is close enough to a protected article to constitute infringement. The European Parliamentary Research Service has noted that the training of general-purpose AI relies on large datasets which may include copyrighted material, and that EU law has created both exclusive rights and exceptions relevant to this process.
For Paul Nemitz, a digitalisation expert and visiting professor at the College of Europe, the stakes for independent media could not be higher. “If the court goes along with Google’s arguments, even in part, private news media and private publishing in Europe is dead”, he warns. “Only state media will survive. In this way, the innovative power of culture and criticism for democracy, in which the way must be open for change is the first principle, would have been killed.”
‘If the court goes along with Google’s arguments, even in part, private news media and private publishing in Europe is dead’ – Paul Nemitz
EU law, in his view, already provides the right framework: “Strong copyright protections for works which are sufficiently new to be recognised as expression of human genius on the one hand and a text and data mining exception to copyright on the other hand.” That exception was designed to enable scientific progress – not to allow the large-scale extraction of journalistic and cultural output by commercial AI systems. Any interpretation of it, he insists, must be read alongside the EU’s fundamental rights framework and the democratic function of independent journalism.
The European Copyright Society, whose members include leading copyright academics, has urged caution. In its view, chatbots, AI models and search engines involve different technical and legal operations and should not be treated as a single category. It also warns against any ruling so broad that it creates a licensing market serving only the largest AI providers – while still recognising that unlicensed use of protected content raises serious concerns.
A sensitive timing
The timing is significant. The EU is already implementing its Artificial Intelligence Act, which requires providers of large AI models to adopt copyright compliance policies, respect rights holders’ opt-outs, and publish summaries of the content used for training. But the Act does not resolve the underlying copyright question – it tells AI companies to comply with copyright law without specifying what compliance requires. The CJEU case may become the judicial key that gives those obligations their practical meaning.
Nemitz sees the commercial argument for requiring payment as straightforward: “For Google and its competitors to pay for the use of copyrighted material must become normal and will only be a small blip in their balance sheet”, he said. For smaller developers and startups, collective licensing societies – which can negotiate package deals across many rights holders – would offer a proportionate and workable route to compliance. And since all developers of language-based AI systems would face the same requirements, no company would gain a competitive edge by refusing to pay.
Google’s strategy
The Balatonkörnyéke case may not be an isolated episode. A source with knowledge of the proceedings, who requested anonymity, told EUrologus’s that two further lawsuits – both files by the Like Company in Hungary – follow a similar pattern of legal interpretation: one involving IT Café, a Hungarian technology news site, and another known as the Gamekapocs case, filed in 2023. In this ruling, the Budapest Court of Appeal found that the Hungarian site on gaming had failed to opt out of text and data mining in the form required by law, meaning Google's scraping and indexing of its content fell within a legal exception and did not constitute infringement.
According to this source, Google did not simply find itself dragged into these disputes – it actively identified and pursued cases of this kind in Hungary. The Mountain View company applied a technique called “court district shopping” with the aim of getting its case fast tracked to the European court of Justice in Luxembourg.
The reasoning, according to the source, is straightforward: Hungary has relatively little established legal practice in AI-related disputes, making it a more flexible jurisdiction in which to test arguments before they reach EU level.
This naturally raises the question of what Google stands to gain from all this at EU level. One answer supported by the facts is that, in a similar legal dispute in Germany between Google and a rights-management organisation called Corint Media, Google has already introduced its successful case against Gamekapocs into arbitration proceedings, effectively arguing: look, there is already an EU-law precedent at member-state level.
A hard institutional test
The Like Company case began with an absurd image: dolphins in Lake Balaton, conjured and embellished by a chatbot. But beneath the absurdity lies a hard institutional test. Europe has built ambitious rules for platforms, markets and artificial intelligence. Now its judges must decide whether those rules can still protect journalism and culture when machines no longer merely index the web – but speak in its place.
The ECJ Advocate General’s opinion is scheduled for 3 September, with the final judgment anticipated later this year. For Paul Nemitz, the court’s decision will ultimately rest on a simple reality: no large language model can be profitable — in Europe or anywhere in the world — without drawing on European copyrighted material for its training. “That,” he says, “is where the leverage lies.”
AI labelling comes into force in the EU
From 2 August, the AI Act provision requiring companies to label AI-generated or manipulated content as such entered into force. The measure requires deployers of AI systems to disclose two categories of content: deepfakes – defined as "AI-generated or manipulated image, audio or video content that resembles existing persons, objects, places, entities or events and would falsely appear to a person to be authentic or truthful" – and AI-generated or manipulated text published to inform the public on matters of public interest, where no human review or editorial control took place and where no legal or natural person assumed editorial responsibility. Disclosure can be achieved through watermarks or other markers enabling easy detection. Non-compliance carries substantial fines.
🤝 This article was produced within the PULSE collaborative project. György Folk (EUrologus/HVG, Budapest) contributed greatly to it.
Do you like our work?
Help multilingual European journalism to thrive, without ads or paywalls. Your one-off or regular support will keep our newsroom independent. Thank you!
