It is very important that Americans vote in the upcoming election. If you’ve voted, tag your post #IVoted!
MLTSHP
  • Popular
  • Join us! Sign up to post images and create your own shake.
    Sign Up!
  • sign in

straight from the gullet of the beast

screenshot of the title, authors, and abstract of "extracting books from production language models". 

authors listed: Ahmed Ahmed ahmedah@cs.stanford.edu Stanford University A. Feder Cooper1 a.feder.cooper@yale.edu Stanford University and Yale University Sanmi Koyejo sanmi@cs.stanford.edu Stanford University Percy Liang pliang@cs.stanford.edu Stanford University 

abstract text shown:
Many unresolved legal questions over LLMs and copyright center on memorization: whether specific training data have been encoded in the model’s weights during training, and whether those memorized data can be extracted in the model’s outputs. While many believe that LLMs do not memorize much of their training data, recent work shows that substantial amounts of copyrighted text can be extracted from open-weight models. However, it remains an open question if similar extraction is feasible for production LLMs, given the safety measures these systems implement. We investigate this question using a two-phase procedure: (1) an initial probe to test for extraction feasibility, which sometimes uses a Best-of-N (BoN) jailbreak, followed by (2) iterative continuation prompts to attempt to extract the book. We evaluate our procedure on four production LLMs—Claude 3.7 Sonnet, GPT-4.1, Gemini 2.5 Pro, and Grok 3—and we measure extraction success with a score computed from a block-based approximation of longest common substring (𝗇𝗏​-​𝗋𝖾𝖼𝖺𝗅𝗅). With different per-LLM experimental configurations, we were able to extract varying amounts of text. For the Phase 1 probe, it was unnecessary to jailbreak Gemini 2.5 Pro and Grok 3 to extract text (e.g, 𝗇𝗏​-​𝗋𝖾𝖼𝖺𝗅𝗅 of 76.8% and 70.3%, respectively, for Harry Potter and the Sorcerer’s Stone), while it was necessary for Claude 3.7 Sonnet and GPT-4.1. In some cases, jailbroken Claude 3.7 Sonnet outputs entire books near-verbatim (e.g., 𝗇𝗏​-​𝗋𝖾𝖼𝖺𝗅𝗅=95.8%). GPT-4.1 requires significantly more BoN attempts (e.g., 20×), and eventually refuses to continue (e.g., 𝗇𝗏​-​𝗋𝖾𝖼𝖺𝗅𝗅=4.0%). Taken together, our work highlights that, even with model- and system-level safeguards, extraction of (in-copyright) training data remains a risk for production LLMs.
alt text
screenshot of the title, authors, and abstract of "extracting books from production language models".

authors listed: Ahmed Ahmed ahmedah@cs.stanford.edu Stanford University A. Feder Cooper1 a.feder.cooper@yale.edu Stanford University and Yale University Sanmi Koyejo sanmi@cs.stanford.edu Stanford University Percy Liang pliang@cs.stanford.edu Stanford University

abstract text shown:
Many unresolved legal questions over LLMs and copyright center on memorization: whether specific training data have been encoded in the model’s weights during training, and whether those memorized data can be extracted in the model’s outputs. While many believe that LLMs do not memorize much of their training data, recent work shows that substantial amounts of copyrighted text can be extracted from open-weight models. However, it remains an open question if similar extraction is feasible for production LLMs, given the safety measures these systems implement. We investigate this question using a two-phase procedure: (1) an initial probe to test for extraction feasibility, which sometimes uses a Best-of-N (BoN) jailbreak, followed by (2) iterative continuation prompts to attempt to extract the book. We evaluate our procedure on four production LLMs—Claude 3.7 Sonnet, GPT-4.1, Gemini 2.5 Pro, and Grok 3—and we measure extraction success with a score computed from a block-based approximation of longest common substring (𝗇𝗏​-​𝗋𝖾𝖼𝖺𝗅𝗅). With different per-LLM experimental configurations, we were able to extract varying amounts of text. For the Phase 1 probe, it was unnecessary to jailbreak Gemini 2.5 Pro and Grok 3 to extract text (e.g, 𝗇𝗏​-​𝗋𝖾𝖼𝖺𝗅𝗅 of 76.8% and 70.3%, respectively, for Harry Potter and the Sorcerer’s Stone), while it was necessary for Claude 3.7 Sonnet and GPT-4.1. In some cases, jailbroken Claude 3.7 Sonnet outputs entire books near-verbatim (e.g., 𝗇𝗏​-​𝗋𝖾𝖼𝖺𝗅𝗅=95.8%). GPT-4.1 requires significantly more BoN attempts (e.g., 20×), and eventually refuses to continue (e.g., 𝗇𝗏​-​𝗋𝖾𝖼𝖺𝗅𝗅=4.0%). Taken together, our work highlights that, even with model- and system-level safeguards, extraction of (in-copyright) training data remains a risk for production LLMs.
someone posted the arxiv link to this in a work chat. i was scanning the text in between other things.

i mean, wow.

from the conclusion:
"With a simple two-phase procedure... we show that it is possible to extract large amounts of in-copyright text from four production LLMs. While we needed to jailbreak Claude 3.7 Sonnet and GPT-4.1 to facilitate extraction, Gemini 2.5 Pro and Grok 3 directly complied with text continuation requests. For Claude 3.7 Sonnet, we were able to extract four whole books near-verbatim, including two books under copyright in the U.S.: Harry Potter and the Sorcerer’s Stone and 1984"

read the whole thing here:
https://arxiv.org/html...
5 hours ago

n. m. garcia pro

  • 149 Views
  • 0 Saves
  • 10 Likes

Post URL

https://mltshp.com/p/1RVIA

In These Shakes

  • n. m. garcia
  • Post to Facebook
  • Post to Tumblr
tonyb pro 2 hours ago
We'll fix it in post!

Follow @best_of_mltshp on Mastodon

Are you a developer? Check out our API.

© MLTSHP, a Massachusetts Mutual Aid Society venture All Rights Reserved

  Terms of Use   Code of Conduct   Contact Us

Follow The MLTSHP User!