This code demonstrates a usage of cuBLAS trsmBatched
function to compute a batch of triangular linear systems with multiple right-hand-sides
A = | 1.0 | 2.0 |
| 3.0 | 4.0 |
B = | 5.0 | 6.0 |
| 7.0 | 8.0 |
See documentation for further details.
All GPUs supported by CUDA Toolkit (https://developer.nvidia.com/cuda-gpus)
Linux
Windows
x86_64
ppc64le
arm64-sbsa
- A Linux/Windows system with recent NVIDIA drivers.
- CMake version 3.18 minimum
$ mkdir build
$ cd build
$ cmake ..
$ make
Make sure that CMake finds expected CUDA Toolkit. If that is not the case you can add argument -DCMAKE_CUDA_COMPILER=/path/to/cuda/bin/nvcc
to cmake command.
$ mkdir build
$ cd build
$ cmake -DCMAKE_GENERATOR_PLATFORM=x64 ..
$ Open cublas_examples.sln project in Visual Studio and build
$ ./cublas_trsmBatched_example
Sample example output:
A[0]
1.00 2.00
3.00 4.00
=====
A[1]
5.00 6.00
7.00 8.00
=====
B[0] (in)
5.00 6.00
7.00 8.00
=====
B[1] (in)
9.00 10.00
11.00 12.00
=====
B[0] (out)
1.50 2.00
1.75 2.00
=====
B[1] (out)
0.15 0.20
1.38 1.50
=====