DeviceArray 参考
简介
DeviceArray 是指向一块存储在 GPU 内存(而非 CPU 内存)中的数据的句柄。它采用引用计数机制,因此只要至少存在一个指向该数据的句柄,数据就会保持存活。
本页描述 数据布局:即对于每种受支持的格式,DeviceArray 的每个像素或点在内存中是如何存储的。
如果您针对该缓冲区运行自己的 GPU 代码(CUDA 或 OpenCL 内核,或 PyTorch、CuPy、OpenCV 等框架),请参阅 DeviceArray — Stream、Queue 和 Context,了解 SDK 如何将其工作与您的工作进行排序,以及何时需要共享的 GPU 上下文。有关包含代码示例的任务导向说明,请参阅 GPU 访问教程。
数据布局
每个 DeviceArray<Format> 都以 行主序 存储其元素,并采用 通道交错 方式。通道数由 Format 决定,且在运行时永不改变。缓冲区是 1D、2D 还是 3D,取决于该格式的通道数以及数据源是有序(organized)还是无序(unorganized)。
格式目录
格式 |
元素类型 |
通道 |
字节数 / 元素 |
典型来源 |
|---|---|---|---|---|
|
|
3 |
12 |
|
|
|
4 |
16 |
|
|
|
1 |
4 |
|
|
|
1 |
4 |
|
|
|
3 |
12 |
|
|
|
4 |
4 |
|
|
|
4 |
16 |
|
|
|
3 |
3 |
|
形状与步长规则
形状取决于通道数和数据组织方式。步长以 元素数量 (而非字节数)表示,并遵循行主序 C 排列。
情况 |
维度 |
形状 |
步长(元素) |
|---|---|---|---|
有序,多通道 |
3 |
|
|
有序,单通道 |
2 |
|
|
无序,多通道 |
2 |
|
|
无序,单通道 |
1 |
|
|
Use DeviceArray::stridesInBytes() when the consumer framework wants byte strides.
交错多通道格式
An organized multi-channel buffer such as DeviceArray<PointXYZ> from an image of height H and width W is one tightly packed block.
Each pixel contributes C consecutive elements before the next pixel begins.
There are no per-plane offsets and no padding between rows.
下图展示了一个 2x3 的 PointXYZ 缓冲区。每个像素 (r, c) 保存三个浮点数 X、Y、Z,它们在内存中连续排列。
Pixel grid (shape = [2, 3, 3]):
col 0 col 1 col 2
+----------------+----------------+----------------+
| X00 Y00 Z00 | X01 Y01 Z01 | X02 Y02 Z02 | row 0
+----------------+----------------+----------------+
| X10 Y10 Z10 | X11 Y11 Z11 | X12 Y12 Z12 | row 1
+----------------+----------------+----------------+
Flat memory (row-major, channel-interleaved):
offset (floats): 0 1 2 3 4 5 6 7 8 9 10 11 ...
+---+---+---+---+---+---+---+---+---+---+---+---+
|X00|Y00|Z00|X01|Y01|Z01|X02|Y02|Z02|X10|Y10|Z10|...
+---+---+---+---+---+---+---+---+---+---+---+---+
\_________/ \_________/ \_________/
pixel (0,0) pixel (0,1) pixel (0,2)
strides (elements): [W*C=9, C=3, 1]
strides (bytes) : [36, 12, 4]
相同的布局也适用于 PointXYZW (4 个通道)、NormalXYZ (3 个)、ColorRGBA / BGRA 及其 SRGB 变体(4x uint8)、ColorRGBAf (4x float),以及 ColorRGB / BGR (3x uint8)。仅通道数和元素类型有所不同。
小技巧
PointXYZW 的存在是为了让每个点按 16 字节对齐,这是大多数 GPU API 和向量加载操作所偏好的方式。ColorRGB / BGR 每像素 3 字节,并非天然对齐;当消费者能从对齐中获益时,建议优先使用 4 通道变体。
平面单通道格式
Formats with one channel (PointZ, SNR) have no channel axis.
An organized buffer is a plain 2D matrix.
Pixel grid (shape = [2, 3]):
col 0 col 1 col 2
+---------+---------+---------+
| Z00 | Z01 | Z02 | row 0
+---------+---------+---------+
| Z10 | Z11 | Z12 | row 1
+---------+---------+---------+
Flat memory (row-major):
offset: 0 1 2 3 4 5
+----+----+----+----+----+----+
|Z00 |Z01 |Z02 |Z10 |Z11 |Z12 |
+----+----+----+----+----+----+
strides (elements): [W=3, 1]
strides (bytes) : [12, 4] (float)
无序点云
无序的 PointCloud 会生成一个没有高度的、由 N 个元素组成的扁平列表。多通道格式会折叠行维度,并保留交错的通道;单通道格式则折叠为普通的一维数组。
Unorganized PointXYZ (shape = [N, 3]):
offset (floats): 0 1 2 3 4 5 6 7 8 ...
+---+---+---+---+---+---+---+---+---+
|X0 |Y0 |Z0 |X1 |Y1 |Z1 |X2 |Y2 |Z2 |...
+---+---+---+---+---+---+---+---+---+
\_________/ \_________/ \_________/
point 0 point 1 point 2
strides (elements): [3, 1]
Unorganized SNR (shape = [N]):
offset (floats): 0 1 2 3 ...
+----+----+----+----+
|s0 |s1 |s2 |s3 |...
+----+----+----+----+
strides (elements): [1]
在运行时读取布局信息
元数据访问器与上述模型是一致的。
DeviceArray::shape()returns astd::vector<int>of length 1, 2, or 3.DeviceArray::strides()returns strides in element counts.DeviceArray::stridesInBytes()returns strides in bytes.DeviceArray::size()returns the number of logical elements (points or pixels, not channels).DeviceArray::sizeInBytes()returns the total byte size of the buffer.DeviceArray::backend()returnsComputeBackend::cudaorComputeBackend::opencl.
Together with DeviceArray::devicePointer(), these are exactly the fields needed to construct a DLPack tensor around the buffer; see GPU 访问教程 for zero-copy interop with frameworks such as PyTorch and CuPy.
另请参阅
DeviceArray — Stream、Queue 和 Context,了解流 / 队列排序和 GPU 上下文共享模型,在运行自己的 GPU 代码时需要用到。
GPU 访问教程,了解端到端使用方法和框架集成示例。
进阶:数据访问,了解
DeviceArray在更广泛的相机 API 体系中的作用。