DeviceArray 参考

简介

DeviceArray 是指向一块存储在 GPU 内存(而非 CPU 内存)中的数据的句柄。它采用引用计数机制,因此只要至少存在一个指向该数据的句柄,数据就会保持存活。

本页描述 数据布局:即对于每种受支持的格式,DeviceArray 的每个像素或点在内存中是如何存储的。

如果您针对该缓冲区运行自己的 GPU 代码(CUDA 或 OpenCL 内核,或 PyTorch、CuPy、OpenCV 等框架),请参阅 DeviceArray — Stream、Queue 和 Context,了解 SDK 如何将其工作与您的工作进行排序,以及何时需要共享的 GPU 上下文。有关包含代码示例的任务导向说明,请参阅 GPU 访问教程。

数据布局

每个 DeviceArray<Format> 都以 行主序 存储其元素,并采用 通道交错 方式。通道数由 Format 决定,且在运行时永不改变。缓冲区是 1D、2D 还是 3D,取决于该格式的通道数以及数据源是有序(organized)还是无序(unorganized)。

格式目录

格式

元素类型

通道

字节数 / 元素

典型来源

PointXYZ

float

3

12

PointCloud。

PointXYZW

float

4

16

PointCloud (GPU 对齐)。

PointZ

float

1

4

PointCloud (深度图)。

SNR

float

1

4

PointCloud。

NormalXYZ

float

3

12

PointCloud。

ColorRGBA / BGRA (+ SRGB)

uint8_t

4

4

Frame2D / PointCloud。

ColorRGBAf

float

4

16

Frame2D (线性浮点)。

ColorRGB / BGR (+ SRGB)

uint8_t

3

3

Frame2D。

形状与步长规则

形状取决于通道数和数据组织方式。步长以 元素数量 (而非字节数)表示,并遵循行主序 C 排列。

情况

维度

形状

步长(元素)

有序,多通道

3

[H, W, C]

[W*C, C, 1]。

有序,单通道

2

[H, W]

[W, 1]。

无序,多通道

2

[N, C]

[C, 1]。

无序,单通道

1

[N]

[1]。

Use DeviceArray::stridesInBytes() when the consumer framework wants byte strides.

交错多通道格式

An organized multi-channel buffer such as DeviceArray<PointXYZ> from an image of height H and width W is one tightly packed block. Each pixel contributes C consecutive elements before the next pixel begins. There are no per-plane offsets and no padding between rows.

下图展示了一个 2x3 的 PointXYZ 缓冲区。每个像素 (r, c) 保存三个浮点数 X、Y、Z,它们在内存中连续排列。

Pixel grid (shape = [2, 3, 3]):

  col 0            col 1            col 2
+----------------+----------------+----------------+
| X00 Y00 Z00    | X01 Y01 Z01    | X02 Y02 Z02    |  row 0
+----------------+----------------+----------------+
| X10 Y10 Z10    | X11 Y11 Z11    | X12 Y12 Z12    |  row 1
+----------------+----------------+----------------+

Flat memory (row-major, channel-interleaved):

offset (floats):  0   1   2   3   4   5   6   7   8   9  10  11 ...
                +---+---+---+---+---+---+---+---+---+---+---+---+
                |X00|Y00|Z00|X01|Y01|Z01|X02|Y02|Z02|X10|Y10|Z10|...
                +---+---+---+---+---+---+---+---+---+---+---+---+
                 \_________/ \_________/ \_________/
                 pixel (0,0) pixel (0,1) pixel (0,2)

strides (elements): [W*C=9, C=3, 1]
strides (bytes)  : [36,     12,   4]

相同的布局也适用于 PointXYZW (4 个通道)、NormalXYZ (3 个)、ColorRGBA / BGRA 及其 SRGB 变体(4x uint8)、ColorRGBAf (4x float),以及 ColorRGB / BGR (3x uint8)。仅通道数和元素类型有所不同。

小技巧

PointXYZW 的存在是为了让每个点按 16 字节对齐,这是大多数 GPU API 和向量加载操作所偏好的方式。ColorRGB / BGR 每像素 3 字节,并非天然对齐;当消费者能从对齐中获益时,建议优先使用 4 通道变体。

平面单通道格式

Formats with one channel (PointZ, SNR) have no channel axis. An organized buffer is a plain 2D matrix.

Pixel grid (shape = [2, 3]):

  col 0     col 1     col 2
+---------+---------+---------+
| Z00     | Z01     | Z02     |  row 0
+---------+---------+---------+
| Z10     | Z11     | Z12     |  row 1
+---------+---------+---------+

Flat memory (row-major):

offset:  0    1    2    3    4    5
       +----+----+----+----+----+----+
       |Z00 |Z01 |Z02 |Z10 |Z11 |Z12 |
       +----+----+----+----+----+----+

strides (elements): [W=3, 1]
strides (bytes)  : [12,   4]     (float)

无序点云

无序的 PointCloud 会生成一个没有高度的、由 N 个元素组成的扁平列表。多通道格式会折叠行维度,并保留交错的通道;单通道格式则折叠为普通的一维数组。

Unorganized PointXYZ (shape = [N, 3]):

offset (floats): 0   1   2   3   4   5   6   7   8   ...
               +---+---+---+---+---+---+---+---+---+
               |X0 |Y0 |Z0 |X1 |Y1 |Z1 |X2 |Y2 |Z2 |...
               +---+---+---+---+---+---+---+---+---+
                \_________/ \_________/ \_________/
                point 0     point 1     point 2

strides (elements): [3, 1]

Unorganized SNR (shape = [N]):

offset (floats): 0    1    2    3    ...
               +----+----+----+----+
               |s0  |s1  |s2  |s3  |...
               +----+----+----+----+
strides (elements): [1]

在运行时读取布局信息

元数据访问器与上述模型是一致的。

  • DeviceArray::shape() returns a std::vector<int> of length 1, 2, or 3.

  • DeviceArray::strides() returns strides in element counts.

  • DeviceArray::stridesInBytes() returns strides in bytes.

  • DeviceArray::size() returns the number of logical elements (points or pixels, not channels).

  • DeviceArray::sizeInBytes() returns the total byte size of the buffer.

  • DeviceArray::backend() returns ComputeBackend::cuda or ComputeBackend::opencl.

Together with DeviceArray::devicePointer(), these are exactly the fields needed to construct a DLPack tensor around the buffer; see GPU 访问教程 for zero-copy interop with frameworks such as PyTorch and CuPy.

另请参阅