NEW YORK / RankWire.AI / – Hachette Book Group, Cengage Learning, and Elsevier have initiated a lawsuit against Google regarding its Gemini artificial intelligence platform. Author Scott Turow and his firm, S.C.R.I.B.E., have joined the proposed class-action suit. The complaint was filed on July 10 in the U.S. District Court for the Southern District of New York. The plaintiffs accuse Google of copying millions of copyrighted books and journal articles without authorization during the development and training of Gemini. As of July 15, the court had yet to rule on the allegations or certify the class.

According to the complaint, Google acquired content through Google Books, Google Play Books, and Google Scholar. Publishers and authors provided works for specific functions such as search, sales, and research purposes. The plaintiffs argue these arrangements did not permit extensive commercial AI training. They further contend Google downloaded large datasets from web scraping that contained copyrighted works, some sourced from known pirate sites and paywalled services.
The 57-page lawsuit presents four claims under federal law. Three focus on alleged reproduction through Google’s services, web scraping activities, and Gemini’s development or training. The fourth invokes the Digital Millennium Copyright Act, asserting Google removed or altered copyright management information from training data. The complaint also references internal discussions about using publisher-provided books, with one assessment estimating potential fines between $10 billion and $100 billion. These allegations have not yet been tested in court.
Class Definition Encompasses Registered Works
The proposed class includes owners of registered U.S. copyrights in eligible books and journal articles. To qualify, books must have an International Standard Book Number (ISBN), and articles must have a Digital Object Identifier (DOI) or an International Standard Serial Number (ISSN). The class covers works allegedly copied from Google services or obtained through web scraping, as well as those reproduced during Gemini’s development or training.
Eligibility also depends on registration timing. One criterion requires registration within five years of publication and prior to Google’s alleged reproduction or distribution. Another allows registration within three months of publication. The lawsuit excludes government entities, Google affiliates, certain court participants, and individuals who properly opt out of the class. The court must approve the class status before proceeding with broader representation.
Damages and Court-Ordered Accounting Sought
The plaintiffs are seeking either statutory damages or actual damages for proven infringements. They also request that Google account for profits attributable to any copyright violations. Their proposed remedies include an injunction, legal costs, and a jury trial. The complaint does not specify a total damages amount but asks Google to disclose Gemini training data, collection methods, and model capabilities via a court-ordered accounting.
This accounting would identify copyrighted works used in Gemini’s training, detailing how Google collected, copied, processed, and encoded those materials. The plaintiffs also seek court-supervised destruction of unauthorized copies under Google’s control. Earlier, Hachette and Cengage aimed to join separate AI litigation against Google in California. The New York case expands the scope to include Elsevier, Turow, and S.C.R.I.B.E., focusing on claims related to Google’s services, web scraping activities, and Gemini training processes.
