ocl: speedup host-memory allocation on GH200 #846

hfp · 2024-09-13T10:53:02Z

No description provided.

hfp · 2024-09-13T10:58:46Z

This change reduces the duration to allocate a single buffer by approximately 3.5s (tested on GH200 system). No other system has ever shown such high allocation cost for this particular implementation of c_dbcsr_acc_host_mem_allocate.

alazzaro · 2024-09-13T11:59:33Z

out of curiosity, is this something related to #841 ?
I'm checking an idea on how to make the memory pool more effective by introducing a configuration parameter to set a "minimal" allocation size. I will also add some statistics on memory allocations, although we can check the size of the buffers via the MPI size messages...

hfp · 2024-09-13T12:14:23Z

out of curiosity, is this something related to #841 ? I'm checking an idea on how to make the memory pool more effective by introducing a configuration parameter to set a "minimal" allocation size. I will also add some statistics on memory allocations, although we can check the size of the buffers via the MPI size messages...

This is not directly related (or I did not try to address #841), and I think #841 used the CUDA-based implementation. I have yet to try my reproducer (acc_bench_smm) on an Alps-like node using the CUDA backend (I will report back). It's possible the CUDA based host-memory allocation suffers a similar slowness. Though, #841 states "tested on H100" but we also call the H100 tuned parameters "H100" even when collected on GH200.

This is further speeding up host memory allocation. So far, an OpenCL cl_mem object was created with no host-pointer and it was up to the OpenCL runtime to allocate the host memory. It turns out, this is horribly slow for GH200 stack (for unknown reasons). It was only helpful to mark such OpenCL allocated host memory as "not worth to transfer initially" (CL_MAP_WRITE_INVALIDATE_REGION); a bug in itself. Normally, allocating host memory by relying on the OpenCL runtime yields best performance at least when this memory is the origin/destination of a (PCIe-)transfer. However, GH200 SW stack seems to struggle with this idea. Introduced code path (covered by XHINTS) specific to Nvidia, which simply wraps a malloc'ed pointer (host memory). * Implemented malloc based c_dbcsr_acc_host_mem_allocate. * Introduced compile-time ACC_OPENCL_XHINTS. * Updated tuned parameters (Mi250 and GH200). * Removed unused function (m_getuid).

ocl: speedup host-memory allocation on GH200

6613090

hfp merged commit cef934d into cp2k:develop Sep 13, 2024
22 checks passed

Provide feedback

Saved searches

Use saved searches to filter your results more quickly

ocl: speedup host-memory allocation on GH200 #846

ocl: speedup host-memory allocation on GH200 #846

hfp commented Sep 13, 2024

hfp commented Sep 13, 2024

alazzaro commented Sep 13, 2024

hfp commented Sep 13, 2024

ocl: speedup host-memory allocation on GH200 #846

ocl: speedup host-memory allocation on GH200 #846

Conversation

hfp commented Sep 13, 2024

hfp commented Sep 13, 2024

alazzaro commented Sep 13, 2024

hfp commented Sep 13, 2024