chart-attention-heads
A small-multiples grid of causal attention maps — one miniature token-by-token heat matrix per (layer, head), sortable by row entropy or previous-token mass, with the picked head enlarged beside it under token labels on both axes and its attention mass split into four disjoint shares that add to exactly 100.
chart-attention-heads
A small-multiples grid of causal attention maps — one miniature token-by-token heat matrix per (layer, head), sortable by row entropy or previous-token mass, with the picked head enlarged beside it under token labels on both axes and its attention mass split into four disjoint shares that add to exactly 100.