An M5 Max using Samsung's LPDDR5X-PIM would generate 650 tokens/second on an 8B model
Replace the M5 Max's LPDDR5X with eight Samsung LPDDR5X-PIM packages and the memory could deliver 4.9 TB/s where the weights live. The right parallel mapping could push Llama 3.1 8B toward 650 tokens per second, or a 27B dense model toward 190.




