Peek LLM
/
Attention
Studio · Computed from GPT-2
Fig. 1
Computed from GPT-2 117M
stronger attention
masked · cannot look ahead
each row looks at the columns · row weights sum to 1