this post was submitted on 10 Oct 2026
71 points (87.4% liked)

Technology

88681 readers
3111 users here now

This is a most excellent place for technology news and articles.


Our Rules


  1. Follow the lemmy.world rules.
  2. Only tech related news or articles.
  3. Be excellent to each other!
  4. Mod approved content bots can post up to 10 articles per day.
  5. Threads asking for personal tech support may be deleted.
  6. Politics threads may be removed.
  7. No memes allowed as posts, OK to post as comments.
  8. Only approved bots from the list below, this includes using AI responses and summaries. To ask if your bot can be added please contact a mod.
  9. Check for duplicates before posting, duplicates may be removed
  10. Accounts 7 days and younger will have their posts automatically removed.

Approved Bots


founded 3 years ago
MODERATORS
top 26 comments
sorted by: hot top controversial new old
[–] voronaam@lemmy.world 4 points 1 hour ago

I am sorry, to enroll a primary maintainers need to open a Pull Request with a Yaml file containing their email address?

And people do that? Here is an example PR: https://github.com/anthropics/oss-scanner/pull/139/changes

Did Anthropic just created a largest ever publicly available collection of email addresses of primary maintainers of OSS projects matching this criteria

established projects that have a critical impact on infrastructure and user security

(Quote from Anthropic)

Is it an invitation for every bad actor to scrape pull requests of this single repo, extract email addresses and target those with every fishing/hacking/account takeover attack imaginable?

Is not that insane?

[–] Wildmimic@anarchist.nexus 19 points 11 hours ago (1 children)

One of the uses i can fully stand behind - we all profit from secure software, and since it's open source anyways, there's nothing to lose by using this.

[–] skibidi@lemmy.world 10 points 8 hours ago (1 children)

The thing to lose is human expertise in finding vulnerabilities, which is needed to feed the machines. It seems likely fewer people will go into the field if it is viewed as 'solved' by a robot.

[–] Wildmimic@anarchist.nexus 6 points 8 hours ago* (last edited 8 hours ago)

That depends. If an coding LLM finds a vulnerability, would you have been able to find it yourself? And the more important part is: do i blindly merge whatever i get thrown at me or do i critically engage with it, analyze why this is a security issue and how to prevent it in the future?

The thing is: an LLM can NEVER replace secure design. The best LLM can't do shit if you have designed your software without security in mind from the start. And this also means this area will never be fully solved by machines, outside of them becoming able to code the thing securely from start to finish (and even then the design which is the basis of that software has to take security into account, or whatever you get will be insecure by default). This is one of the areas that can be massively enhanced by AI, but it can't replace people that are able to design secure systems.

[–] Zedstrian@sopuli.xyz 63 points 14 hours ago (2 children)

It's not free; doing so provides Anthropic with free access to training data for its models.

[–] DrCake@lemmy.world 69 points 13 hours ago (1 children)

I think it’s safe to assume that any public code on the internet is already in their training dataset.

At this point anything that is public in general, probably all the posts in the “fediverse” have been taken already

[–] dan@upvote.au 17 points 13 hours ago* (last edited 4 hours ago)

probably all the posts in the “fediverse” have been taken already

I don't doubt this, as it's trivial to do (by design). You don't even have to scrape anything.

New posts and comments in a Lemmy community (or Mastodon profile or any other ActivityPub-based Fediverse app) get pushed to every server where federation is allowed and at least one person is subscribed to it. You can spin up your own server, subscribe to a bunch of communities, wait a while, and over time your PostgreSQL database will be populated with profiles, posts, comments, etc.

[–] dan@upvote.au 27 points 13 hours ago* (last edited 13 hours ago) (1 children)

The code they're analyzing is free and open-source though, so they can already use it this way. One of the freedoms of free software is that people can study it and use it for whatever purpose they want.

Isn't it good that they're sharing data with the projects instead of just keeping it for themselves?

[–] Zedstrian@sopuli.xyz 7 points 9 hours ago (1 children)

One of the freedoms of free software is that people can study it and use it however they want, for whatever purpose they want.

Given that Anthropic and OpenAI derive their profits from their models being closed source rather than open source, I don't think they should be assisted in devouring freely-accessible data in their endeavor.

Not all developers may see it this way, but those contributing code to LLMs are training them for free at the expense of career opportunities for fellow developers, all the while Anthropic, OpenAI, and other AI companies rake in money for maintaining access to content that isn't theirs.

[–] dan@upvote.au 2 points 4 hours ago* (last edited 3 hours ago)

I don't think they should be assisted in devouring freely-accessible data in their endeavor

How is this assisting them, though? I see it as two different options:

  1. They use the open-source code without giving anything back; or
  2. They use the open-source code and help improve the project by reporting security issues.

In any case, a fundamental feature of both free and open-source software is that anyone can use it, regardless of if you like them or agree with them or not, without discrimination. You could instead use a license that forbids AI training, but then your code would no longer be free or open-source.

[–] pcouy@lemmy.pierre-couy.fr 22 points 14 hours ago

Loling at the previous comments here... As if AI labs don't already vacuum all open source code they can find

[–] NoLemurs@lemmy.world 2 points 10 hours ago* (last edited 10 hours ago) (3 children)

"The outputs of this opt-in vulnerability scanner will be fully model-generated, without human review or triage," Anthropic explained. "This will enable faster and more frequent scanning, but means that it is possible reports will be incorrect or invalid."

When I saw the headline, I was wondering about this specifically. This may make this service not super useful.

My experience with AI security reviews is that they're fantastic at finding faults, but they always find a list of things to complain about. If there are no real/serious faults they'll start finding things that kind of have the same shape as a security issue, but really aren't if you dig into them. I've regularly had an LLM generate a list of 10-15 issues ranging in severity from "nits" to "critical" where none of them were actual issues.

Periodic reviews seem like they could get annoying really quickly, becoming more of a maintenance burden than a help.

[–] shortwavesurfer@lemmy.zip 1 points 2 hours ago

Yeah. I did use an open weight model LLM to scan an app that I use and it came up with a critical severity issue and I checked on it first to see if I could reproduce the problem before ever submitting it to the developer as a problem.

Turns out it was actually a legitimate security problem that has now been fixed.

[–] Wildmimic@anarchist.nexus 2 points 8 hours ago* (last edited 8 hours ago)

Yeah, there is some fatigue to be expected if you run this thing on the regular.

But the library ecosystem is changing constantly, and that might mean that while it doesn't find a real issue at one point that there won't be one some time in the future. So i would say: Scan it semi-regulary for issues and review what was found, but don't go overboard. Somewhat similar to getting an MRT every year instead of once a decade.

[–] scytale@piefed.zip 2 points 8 hours ago

AI security reviews is that they're fantastic at finding faults, but they always find a list of things to complain about.

Same experience for me. At work, I always review the output because a lot of times it finds something that can be noteable but not relevant to the scope of what is being reviewed. Even with guardrails like instructions to not go beyond the stated scope, a human absolutely still needs to read and check the findings at the end.

[–] psx_crab@lemmy.zip 1 points 14 hours ago (1 children)

In return your code is theirs.

[–] BlueEther@no.lastname.nz 22 points 14 hours ago (1 children)

and they haven't already scraped the whole of OSS already?

[–] psx_crab@lemmy.zip 1 points 13 hours ago (1 children)

Assuming yes, it would've been outdated, and this give them access to fresh new code for a foreseeable future in the case of Microsoft blocking access from other AI to make it exclusive for Co-Pilot.

[–] Womble@piefed.world 17 points 13 hours ago (1 children)

They don't have to offer this to get new versions, that's kinda the point of foss software, it's there for anyone to download and use however they like whenever they want.

[–] BlueEther@no.lastname.nz 1 points 3 hours ago (1 children)

Hot how ever they like though, that would be code under MIT ant the like, not all FOSS. Where the other licenses actually fall?

[–] Womble@piefed.world 1 points 3 hours ago* (last edited 3 hours ago)

No, the GPL doesnt put any restrictions on what you can do with code, just on that code and derivatives of it have to be released (source code included) under the GPL as well.

The courts have tended to find that training LLMs on data you have legal access to is fair use and transformative.