How to Protect Your Artwork from AI Scraping: A 2026 Playbook

Design tipsDesign tools •

How to Protect Your Artwork from AI Scraping: A 2026 Playbook

In August 2026, an artist-focused social and portfolio platform got scraped three times that month. Cara reported the attacks on its own blog: starting August 13th, three scrapes took 12M images, then 9M URLs, then 123,000 images plus members' names, locations, comments, and text posts (Cara scraping FAQ). The first scraper, Heft, apologized, deleted the dataset, and later helped build a detection tool with Cara's founder. Cara had NoAI tags. Cara had Glaze built in; its integration support was discontinued in September 2026. Those measures did not prevent the reported scrapes.

This is the honest starting point for any guide on how to protect your artwork from AI scraping: nothing is a shield. What exists instead are layers, and each layer defends against a different threat. Some stop lazy crawlers and nothing else. Some raise the cost of a specific abuse. Registration and other evidence can serve different roles in a copyright dispute. The listicles that rank your options by tool name miss this, so this playbook ranks them by what they actually defend, then gives you the order to do them in.

Key Takeaways

  • AI scraping protection in 2026 is layered defense, not one tool: access control, style cloaking, dataset opt-out, provenance proof, and detection each cover a different gap.

  • Glaze targets style mimicry and Nightshade targets training-data poisoning; their effectiveness depends on the models and defenses tested. But LightShed, published at USENIX Security 2025, demonstrated automated removal of their perturbations, so treat cloaking as one step in a process, not a badge.

  • Voluntary signals like NoAI tags and robots.txt blocks depend on compliance; Cara reported scrapes that did not honor its restrictions, though crawler blocking at the network layer (Cloudflare) has teeth that HTML tags do not.

  • In the United States, timely copyright registration generally preserves eligibility to seek statutory damages and attorney’s fees, subject to statutory exceptions; those awards are not guaranteed.

Before you begin: what are you actually defending against

"AI scraping" bundles four different events, and tools that claim to stop all of them are selling you a sticker.

  • Scraping itself: someone downloads your images. Tags, watermarks, and cloaking do not guarantee that publicly visible images cannot be downloaded. Cara warns that publicly visible images cannot be fully protected from determined scrapers.

  • Training use: your scraped images feed a model's training set. This is where opt-out registries and crawler blocking do real, partial work.

  • Style mimicry: a fine-tuned model learns to draw "in the style of you." This is the specific gap Glaze targets.

  • Attribution loss: your work circulates with your name detached. Provenance metadata can document a work’s history; registration serves a separate legal role.

Write your four threat levels down before touching a single tool, because the effort differs by audience. A studio illustrator with 200 posted pieces has different math than a hobbyist with a single Instagram grid. If your career depends on a recognizable style, layer 3 matters more than layer 4. If you license work commercially, layer 4 with registration is your whole game.

Layer 1: control access before a model arrives

Start by making it harder for crawlers to harvest you in bulk.

Start at the network layer. If your portfolio runs on Cloudflare, one setting blocks known AI crawlers, and the newer Pay Per Crawl system, currently in beta, offers participating publishers a way to charge eligible crawlers; access and setup depend on Cloudflare’s current program. The enforcement is real because the block sits in the pipe rather than in the page.

Then put robots.txt and meta tags in their proper place. A robots.txt listing (GPTBot, CCBot, ClaudeBot, and friends) plus noai meta tags express preferences, but noai is not a universal standard honored by all AI trainers. The Cara case is the counterexample: tags are voluntary, and the scrapers who mattered did not care. Do the work anyway, it costs minutes, but do not rely on them to stop noncompliant crawlers.

Finally, post preview sizes and keep the full-resolution files off public pages. Print-quality work belongs behind client delivery or a store. This does not stop training use, but smaller previews can limit the detail exposed, but an 800-pixel size does not guarantee protection against training or style mimicry, and commission-quality images sitting in a public folder are the actual inventory worth guarding. If a client pipeline runs through public links, the portfolio hosting itself deserves attention: the hosting mistakes that quietly sink design projects covers directory exposure that makes scraping trivial.

Layer 2: cloak and poison, with the caveats stated

Cloaking has important limits, so here is the version with the tested attack results attached.

Glaze applies image perturbations intended to disrupt style mimicry in tested fine-tuning setups; results depend on the model and any countermeasures. Nightshade is designed to poison associations in training data; the lab gives the example of a model trained on enough shaded cow images producing handbags instead. Results depend on the training setup and countermeasures. The University of Chicago lab behind both recommends applying Glaze to everything you post and treating Nightshade as an optional deterrent.

The caveats. A USENIX Security 2025 team built LightShed, an autoencoder that reconstructs and subtracts these perturbations, reporting a 99.98 percent true-positive rate detecting NightShade samples, with Glaze named among the defeated perturbation schemes. The LightShed study demonstrates limits in its tested setups; those results do not establish that every model or attack defeats Glaze. Cloaking is a moving target, and the lab says so: review the lab’s release notes for protection changes; a hardware-support update does not necessarily require re-cloaking every image.

Practical constraints before you plan an evening of it: Processing time depends on hardware and settings; Glaze’s guide says GPUs can greatly reduce it, while its highest render quality can take around an hour per image on a personal laptop. Nightshade can introduce visible changes, especially on flat colors and smooth backgrounds. For high-value series, keep full-resolution originals private and prepare the public version before running Glaze. The lab’s guide recommends resizing and watermarking before Glaze, while its FAQ permits converting or compressing the Glazed PNG afterward. The Glaze FAQ recommends PNG input and permits subsequent conversion and compression; it does not prescribe re-glazing after every platform recompression. The same logic applies to any AI-assisted pipeline touching your files: what AI vectorization does to clean line art is a useful case study in how quietly tools rewrite your actual work.

Layer 3: opt out where a registry exists

Data has already been scraped; the retroactive play is dataset-level exclusion, and it works only where a trainer honors the list.

Have I Been Trained, run by Spawning, has offered dataset search and opt-out registration through Spawning’s Do-Not-Train registry. The site currently displays a maintenance notice, so check access before planning a search or registration. Registry signals affect future use only where trainers honor them. The same registry feeds opt-out signals compliant vendors consume. It does not undo a completed scrape or constrain trainers who ignore the registry. It can signal an opt-out from future training use, including for images already listed in a dataset, where a trainer honors the registry; completion time and benefit vary.

A registry expresses your preference to participating trainers; it does not establish which crawlers negotiate with rights holders or what share of the market complies.

Layer 4: prove it is yours

Keep the C2PA claim and the invisible-watermark claim separate in your head.

C2PA provides tamper-evident provenance metadata; compatible tools can include it when Content Credentials are enabled on export. It records signed provenance assertions, rather than proving authorship by itself. Platforms may remove embedded metadata, but C2PA supports durable credentials using soft bindings such as fingerprinting or watermarks to help rediscover a credential. Preservation varies by tool and distribution path. Invisible watermarks are not access controls. Their robustness depends on the scheme and attack; LightShed’s perturbation-removal results do not establish that most provenance watermarks fail.

What remains genuinely enforceable is the oldest step: register the copyright. In the United States, registration before infringement or, for published works, within three months after first publication generally preserves eligibility to seek statutory damages and attorney’s fees under 17 U.S.C. § 412, subject to its exceptions. Registration does not guarantee an award. It costs a filing fee per work or batch, preparing an application and the Office’s processing time vary; the availability of remedies depends on the facts and legal requirements. For designers whose client contracts assume ownership, the practical companion is getting the paperwork right before the work ships: the graphic design hiring guide discusses revision terms and questions to ask before signing a contract.

Layer 5: detect, because prevention leaks

The final layer assumes you will get scraped anyway and asks what you do next. This is where the strangest 2026 story in this space lives: after the Cara attacks, founder Jingna Zhang partnered with the first scraper, Heft, on Lantern, an open-source tool described by WIRED whose project page describes one-way fingerprints and checks of publicly posted datasets within its coverage. Lantern is a work-in-progress project that aims to send alerts when it finds matching works. It cannot check companies’ non-public training datasets. Detection instead of prevention, built by the person who attacked.

You can do a manual version of the same idea today: run periodic searches of your distinctive pieces against Hugging Face datasets and the public scrape announcements that circulate on artist forums, and keep timestamped copies of originals (layered PSDs, process files, dated exports) as evidence that can help document the history of a work. This layer is easy to overlook. If imitation affects your income, make periodic checks part of your workflow before deciding whether legal action is appropriate.

The order to do this in

  • Block crawlers at the network layer (Cloudflare setting or equivalent host control) and resize public files to preview resolution.

  • Check whether Have I Been Trained registration is available, then use opt-out signals where participating trainers honor them.

  • Glaze the pieces that carry your signature style; review protection-related release notes before deciding whether to re-cloak.

  • File copyright registrations on the work you license or sell.

  • Set a quarterly detection sweep using Lantern if its current coverage fits your needs, alongside manual dataset checks. Cara’s FAQ describes Lantern as an independent, publicly available project; it is not part of Cara.

The time required for steps 1, 2, and 4 varies. Network blocking requires effective enforcement, opt-outs depend on trainers honoring them, and copyright remedies depend on legal requirements. Steps 3 and 5 are ongoing labor, so budget them the way you would budget backups: silently, on a schedule, because the one you forgot is the one that mattered.

Is design still a defensible career when the imitation is this cheap? We argued our position in is graphic design still a good career, and the short answer starts with owning what you make.


Linh Nguyen

Graphic Designer

Passionate Graphic Designer | Specializing in Illustration Design | Bringing Captivating Visuals to Life

Related Posts