this post was submitted on 25 Aug 2026
43 points (85.2% liked)

Technology

87652 readers
3359 users here now

This is a most excellent place for technology news and articles.


Our Rules


  1. Follow the lemmy.world rules.
  2. Only tech related news or articles.
  3. Be excellent to each other!
  4. Mod approved content bots can post up to 10 articles per day.
  5. Threads asking for personal tech support may be deleted.
  6. Politics threads may be removed.
  7. No memes allowed as posts, OK to post as comments.
  8. Only approved bots from the list below, this includes using AI responses and summaries. To ask if your bot can be added please contact a mod.
  9. Check for duplicates before posting, duplicates may be removed
  10. Accounts 7 days and younger will have their posts automatically removed.

Approved Bots


founded 3 years ago
MODERATORS
you are viewing a single comment's thread
view the rest of the comments
[–] unpossum@sh.itjust.works 1 points 9 hours ago (1 children)

Kaiser is one of the coauthors of Attention is all you need, the paper that introduced the Transformer architecture, the basis for all major LLMs.

3b1b has a writeup: https://www.3blue1brown.com/lessons/attention/

For example, imagine that the text we input was most of an entire mystery novel, all the way up to a point near the end, which reads:

Therefore the murderer was...

If the model is going to accurately predict the next word, that final vector in the sequence which began its life simply embedding the word was will have to have been updated by all of the attention blocks to represent much more than the individual word.

It will have to have somehow encoded all of the information from the full context window that's relevant to predicting the next word

Attention is the mechanism that lets an LLM use the (correct parts of the) entire context to predict the next word.

[–] h0tbeef@lemmy.zip 1 points 8 hours ago* (last edited 8 hours ago)

Oh yeah, I was able to figure out why the article said that with a couple search queries, but I still think that the phrasing in the article is misleading (and likely intentionally so).

Edit: Also, I think part of what’s confusing people is that they keep using terms like “attention” and “inference” to describe computer processes that may or may not have some kind of underlying similarity to the corresponding human capabilities. I also believe this to be deliberate obfuscation.