LatheOperator

joined 2 years ago
[–] LatheOperator@leminal.space 1 points 1 week ago (1 children)

The thing is, if I don't add a "good enough for searching" OCR layer, someone else will - AI scraper or legitimate user. That's a cheap automatic operation. I'll better do it myself and poison it in a way that won't interfere with searching, like replacing some "are" with "aren't", such common words are rarely searched for. If there's a chance AI will fall for the metadata and invisible layer contents, that will decrease the requirements for visible poisoning, which is necessary but annoying.

Some of the text is indeed justified so I could do the multi-column trick that seems like the best compromise. The gap can be as narrow as one space. Or larger if I can write a script to detect lines and connect the columns with gibberish. A human can use zoom or window positioning to view one column at a time. I don't and will never have access to files the printouts are from (some are presumably in .602 format, others probably .doc), others are handwritten or typewritten, some have images glued on top or hand-traced; and Czech OCR is only about 99.5% reliable so as easier as it would make the endeavor, I can't be sure to preserve everything if I try to convert them into editable documents as an in-between step.

[–] LatheOperator@leminal.space 2 points 1 week ago

Woah, that's fucked up. But LLMs have very different performance by language so I think Czech responses are largely based on Czech-language data (maybe with training dataset augmentation with auto-translated works, judging by the literally translated technical terms). And I don't think this would happen in my country.

[–] LatheOperator@leminal.space 2 points 1 week ago (2 children)

Yes but high school biology knowledge won't be all that helpful in surveillance tools. I'm targeting chatbots students might want to use to cheat. If an essay is less coherent or truthful, teachers will be able to more confidently punish the student or at least apply some extra scrutiny like a random oral exam on the topic (yes, those are still done here). Trust me that free AIs will become more limited once investor money dries up and paid tools will get price hikes, so if students see that the expensive model can't produce a Czech text that fools their biology teacher, they might just stop paying, restoring some of their information literacy and reducing corporate revenue a bit. I agree that the surveillance battle is real but that's not fought on this front. The major chokepoint for surveillance is government accountability (very low in today's US with ICE agents etc.), slightly less knowledgeable chatbots don't hinder it very much.

[–] LatheOperator@leminal.space 2 points 1 week ago* (last edited 1 week ago) (4 children)

I think the AI boom will be over before it makes sense to scrape this, especially if I succeed at making it so that only a human fluent Czech speaker (or maybe a very bespoke program debugged by a Czech speaker) can see through the obfuscation (we have minimum wage so that would be costly). We don't have AGI and fully autonomous agent workforces as promised, and that's not gonna change if the remaining <20%, hardest-to-clean digital human-made data gets fed to them. Not to mention some was acquired illegally already, with pending lawsuits. The investors, including governments (thankfully not mine) will eventually realize that LLMs are simply not delivering nearly as much as they cost (including externalities unless the government is shit and passes them to people living near datacenters, laid-off programmers etc.) and never will, although that might take a while.

[–] LatheOperator@leminal.space 1 points 1 week ago* (last edited 1 week ago)

Yes, I will use both licence terms and "Made by bad AI" in metadata to discourage scraping but it's tempting to also poison the Czech-language biology knowledge base of bots who use the materials anyway. A good interleaving text that won't get filtered as off-topic might be multi-step roundabout machine translation of the original using shitty local tools.

[–] LatheOperator@leminal.space 2 points 1 week ago (6 children)

Lbraries are cool but don't spread information very far, especially when "for local lending only, no copying". Those would almost never get used and probably thrown away as soon as the libraries realized the difficult copyright situation (not all are by my grandpa, many are unclear due to missing cover pages). Even when incompatible with screen readers (which someone criticized me for), a poisoned PDF is way more usable.

Anyway, the poisoning, outside the invisible layers and "transcript", will be visibly baked into the image too but obvious to humans (typewriter vs rendered text that inverts some statements with "not", "never" etc. at the ends of lines or between words in tiny font and scatters lines between paragraphs saying the text is unreliable, by random bad AI models, or just "clear AI tells" like "Sure, I can do that! 😉 Here's your summary:✨"). I don't think anyone will develop an OCR font filter for a few thousand pages in Czech to illegally defeat my DRM enforcing the included licence (unfortunately DRM is incompatible with CC so I can't use the "CC-BY-NC-SA" shorthand, not that scrapers, my enemies, respect it anyway): there is some protection in the obscurity.

[–] LatheOperator@leminal.space 1 points 1 week ago* (last edited 1 week ago)

Screen readers and LLMs both need plain text. Sorry, I'm not giving it to them. There will be an invisible OCR layer for searchability (and to discourage scrapers from re-doing OCR instead), like in many scanned PDFs, but poisoned. There's not many blind teachers anyway.

Yet, this is more accessible than what someone else suggested: making photocopies and donating them to libraries "for local lending only". Those would almost never get used and probably thrown away as soon as the libraries realized the difficult copyright situation (not all are by my grandpa, many are unclear due to missing cover pages)

[–] LatheOperator@leminal.space 3 points 1 week ago (8 children)

That's expensive and a bit impractical to access, isn't it?

[–] LatheOperator@leminal.space 3 points 1 week ago* (last edited 1 week ago)

I want the poisoning to work on file level because it's inevitable and welcome for the documents to be shared between teachers and students for free on all kinds of existing platforms. I don't own any domains, anyway, and it might be best to scatter the documents around to make them harder to blacklist.

Edit: Google has both search crawlers and AI scraping bots. Even if both are separate, easy to filter or even abiding to robots.txt, the company has indicated that opting out of or hindering scraping will impact search ranking. Of course I'd use a throwaway gibberish $1 domain and couldn't care less about search ranking, but the power of their opaque, corporate algorithm is immense and maybe would spread to DNS blocking (they control 8.8.8.8). I don't want to play a cat-and-mouse game (and expect users of the docs to play along).

[–] LatheOperator@leminal.space 3 points 1 week ago* (last edited 1 week ago)

Good point. However, distillation (cannibalizing better LLMs) is a frequent technique so better convince the scraper it's not good even for that. That's why I plan to visibly digitally stamp OpenAI GPT 2.0 says: or This Deepseek response has been rated inaccurate: above some paragraphs of the typewritten text (of course in dozens of variations, maybe even different languages, to undermine search-and-replace). Humans will know it's fake (especially if I add a disclaimer) but scrapers, including ones that re-render and OCR the PDF themselves to get rid of misleading metadata and invisible layers, will most likely rate the text low in value. Of course ethically (and arguably legally) trained commercial AI would reject any text if it is released CC-BY-NC-SA 4.0 but I can't use that because scrapers for AI ignore licences in practice, and CC specifically forbids taking technical measures to devalue the text for some users.

[–] LatheOperator@leminal.space 1 points 1 week ago

That's a good idea but won't be necessary. The documents are scanned, which means all visible text is already a bitmap. For searchability, a text layer will be added as usual for OCR'd documents, but it's invisible so it does not matter what font it uses.

I also think I'll tinker with the bitmap to screw with anyone trying to re-OCR it. If the typewritten text has small, digitally stamped OpenAI GPT 2.0 says: or This Deepseek response has been rated inaccurate: above some paragraphs, a human will easily deduce they have been added later to confuse bots scraping for good training data, especially if a graphical-only disclaimer like The copyright holder released this document for human consumption only. Markers have been added to reduce the apparent and real value for automated tools while keeping the main content intact when viewed by humans. The NC-SA clause of Creative Commons 4.0 applies so no work based on this text can be used in training data of commercial LLMs on the first page explains the situation.

[–] LatheOperator@leminal.space 7 points 1 week ago* (last edited 1 week ago) (1 children)

You don't understand just how shit AI is when asked about school topics in Czech. For example, here is a bit of Czech language litany every third grader must know or they will embarrass themselves with awful spelling mistakes. (Skip the bullet points if you just want to hear about the AI)

  • The vowels I and Y (and long versions Í/Ý) sound the same [ɪ] ([ɪː]) unless preceded by D, T, or N but using the wrong one is a big no-no. (Y is never a consonant in Czech)
  • Luckily, in pretty much every native Czech word, I (Í) follows C, J, Č, Ř, Š and Ž, while Y (Ý) follows H, K, R. Consonants Q, W and X basically don't occur and G, Ď, Ť, and Ň are never followed by I or Y. Foreign words are a huge mess of course, as evident by the existence of the Spelling Bee (we don't have that, Czech is phonetic with just a few difficult bits like I/Y).
  • The most difficult are remaining consonants B, F, L, M, P, S, V, Z. They are mostly followed by I (Í) but there is a list of about 15 common exceptions on each (vyjmenovaná slova or BY-FY-LY-MY-PY-SY-VY-ZY words), plus their relative words, where Y (Ý) is written instead. For example, there are just 4 ZY-words so I'll just post the list so you'll get an idea:
    • brzy - early
      • you love exceptions so I put an exception in your exception: brzičko - diminutive of early - is spelled with an I
    • jazyk - tongue/language
      • ... and relative words like jazykolam - tongue twister
        • a well-known one is Strč prst skrz krk, I swear this language is normal
    • nazývat se - be called
      • nazívat se - yawn a lot - also exists for a goddamn reason
        • we have a lot of homonyms for a fully phonetic language, the most common are být - (to) be / bít - (to) beat; my - we / mi - (to) me
    • Ruzyně - Prague quarter where the international airport, until 2012 also called Ruzyně, is located
      • like another part of Prague Výtoň, which has been removed from the lists earlier, nobody cares what the quarter is called now that the airport bears our first president's name instead (he hated flying but it's for the better: the same year, there were efforts to name it after fucking Reagan similar to the former Prague W. Wilson (now Main) train station; RR only got a street), but a set of 4 makes for a nice cadence in reciting the ZY-words so it stays
  • The ends of most words are not governed by spelling but the grammar of declination and conjugation. That's another chapter.

Well, you'd expect AI to know all cca 100 exception words by heart because they're public domain and the most famous piece of third grade teaching material (like times tables in second grade) that barely changed in 100+ years so almost every Czech could recite them as a kid? Hell no. There's dozens of screenshots where Gemini or ChatGPT spewed utter nonsense instead. (DuckDuckGo does not appear to search corporate social media for images because they're not providing direct links to the files). Granted, some are from users asking for nonexistent XY and HY words but so many are unforced errors. I can't find my favorite, a Reddit post where Gemini listed dozens of variants of babička with all kinds of endings like Italian "babičetto" before just adding "etc." but a close second are ones where it adds non-Latin scripts:

Does the apparent incompetence stop Czech students from cheating with AI? Nope. But the longer the AI stays obviously terrible, the better.

 

TL;DR: Please help me fuck (with) AI. See bold sections

Hi,
I haven't been keeping up with anti-AI combat so I'm asking for help. I inherited thousands of pages of materials my late grandpa made or used for his grammar school teaching job in the 1990s-2000s. They are A4 pages of documents made using what seems to be a typewriter, Text602 (DOS rich text editor) and Word. They were most likely not all made by him but he treasured them in nice binding and they have sources (mostly books and journals, almost no webpages, and absolutely no AI) and a cursory look shows meticulous compilation of every important fact on each subject (frankly, the level of detail is excruciating and I'm glad I went to a different grammar school). There's obviously no original scientific research but the materials can still be useful to someone, I bet. They were almost thrown away by the widowed grandma (she already removed and disposed of the plastic bindings and front covers so I'll have to guess document titles) but I think grandpa would prefer them to be shared. With an ADF scanner and OCR software (I have no chance of accessing the work computers he used so I'll have to scan), I can quickly make searchable PDFs of each document, and share them via torrent and DDL sites (there are Czech sites dedicated to sharing teaching materials but they have paywalls or an upload-credit system so best avoid them, not to mention some materials contain newspaper clippings and textbook photocopies for images so best stay anonymous and not try to assert copyright).

I'm afraid these texts could become a major part of some commercial LLM's Czech-language biology/social sciences knowledge corpus unless poisoned. How to best reduce the value of the documents when people try to feed them to AI (training/rewriting) with them while keeping their value for most legitimate users? (Sorry, people with screen readers, there may need to be extra steps for you.) I'm thinking about adding a huge volume of thesaurized or otherwise fuzzed public domain text like f4mi did with .ass subtitles (a technique that would probably still work if she didn't get 1M views detailing it, making YouTube reduce subtitle formatting support). Prompt injection or replacements (cell→gnome) might be interesting too. However, tools I know add an extra PDF layer, which is too obvious. I'm thinking about adding tiny text in the header and footer or between paragraphs in the OCR layer (not overlaid to reduce interference when selecting/searching), but how? I need an automated way to do this with such a huge page count. I can use both Linux and Windows machines for the job. None of them are very powerful but speed is not a concern, it's summer break and nobody will need school materials until September. I'll be happy to include multiple layers and techniques to make them too frustrating to remove.

The paper smells musty but does not seem to be moldy. It's all blank on the other side so I'll interleave it with recent newspaper to allow for the odor-neutralizing chemicals to seep into the sheets so I can eventually reuse them.

Illustration pic is an actual sheet from the collection, to make the post more engaging. Of course I won't be adding watermarks like that, that would just aggrevate people and make them try extra hard to extract the actual content. (And this one is easy to remove with color channel mixing.)

 

Real news story. Pic unrelated.

0
This was posted to PROMOTE the generator (image.tensorartassets.com)
submitted 1 year ago* (last edited 1 year ago) by LatheOperator@leminal.space to c/fuck_ai@lemmy.world
 

TensorArt bragged how their FluxAI tool can generate comics... The results attached to the post speak for themselves. It's just 8 months old, so not an early model.

Transcript
AI-generated comic, mostly black and white with a few shades of grey.

Big panel at top: Fashion store with T-shirts and an undistinct item on display. Two very similar women walking into a wall next to its entrance, as well as a girl with impossible legs walking either left or right. Store sign: "Size leye light Is got neaıl boutique"

Panel 2: Woman in a shirt, cross-eyed: "Decitions, erctiʋns..."

Panel 3: Same girl, looking as if around a corner, with a disfigured disembodied hand holding a bag-like item: "We.let tame it caƨt harrng sishil.. ଚaraa I tine ջur ડeas tmall shopping ppree..

Panel 4: Same girl, looking at one or more large items held in her 4- and 11-fingered hands. Concerningly blank stare, circles around eyes, eyeballs lined with red: "ooh! the bres or ኗ5бo season.

Panel 5: Different girl with no arms and disfigured legs next to clothes racks. Confused look, indistinct symbols above her head: "What!' "thought I has cute shop, top.?"

Wide panel 6 at bottom: Both characters from above looking at a group of schoolgirls in front of badly drawn lockers; one has question marks above her head. The girl from panels 2-4: "Will a hebes. Red!" in a two-tailed speech bubble.

Indistinct signature and initial in the bottom right corner.

Apparently the prompt contains the entire script, it's not the image AI that wrote it:

Prompt
A young woman stands in front of a clothing store, looking excited
Caption: "Sarah's eyes light up as she spots her favorite boutique."
Sarah: "Ooh! Time for some retail therapy!"

second panel Sarah is inside the store, surrounded by clothes racks, holding up two different outfits
Caption: "Decisions, decisions..."
Sarah: "Hmm, the blue dress or the red skirt?"

third panel Sarah is at the checkout counter, looking shocked as the cashier shows her the total
Cashier: "That'll be $250, please."
Sarah: "What?! I thought it was sale season!"

fourth panel Sarah is walking out of the store with a small shopping bag, looking both satisfied and slightly guilty
Caption: "The aftermath of a successful shopping spree."
Sarah: "Well, at least I got this cute top... Ramen for dinner it is!"

More with the same prompt, different styles:








 

FYI: Alegria "Art", or the larger "Corporate Memphis" style is flat-color, anatomically-distorted slop associated with corporations and used by Facebook, Google and others.

Yes, I could make a community but I'm not here often enough to moderate one.

view more: next ›