CUDA Micro-benchmarks in Practice (Part 2): Peak Compute, Memory Bandwidth, and Hopper Asynchronous Pipelines