Image recognition method, vehicle control method and electronic equipment
By fusing point clouds from multiple frames of images into a voxel space, constructing voxel structures and extracting features, and then selecting target point clouds for image recognition, the problem of low image recognition efficiency and accuracy in automated delivery vehicles is solved, thus improving driving safety.
Patent Information
- Application Number
- CN202410544505.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-04-30
- Publication Date
- 2025-10-31
- Estimated Expiration
- 2044-04-30
AI Technical Summary
The image recognition computing efficiency and accuracy of automated delivery vehicles are low during operation, resulting in insufficient driving safety.
Point clouds from multiple frames of images are fused into voxel space to construct voxel structures. Voxel features are extracted through attribute representations on feature maps, and feature masks are determined based on voxel coordinates and indices to select target point clouds for image recognition.
This improves the accuracy and efficiency of image recognition, enhancing the driving safety of automated delivery vehicles.
Smart Images

Figure CN120877265A_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of autonomous driving technology, specifically to an image recognition method, a vehicle control method, and an electronic device. Background Technology
[0002] Automated delivery vehicles can transport orders from delivery stations to designated delivery points, where human delivery personnel then deliver the orders to users, thus improving delivery efficiency. During their round trips, these delivery vehicles need to perform image recognition of their driving environment to avoid collisions or accidents. However, image recognition in these scenarios suffers from low computational efficiency and accuracy, leading to lower safety standards for delivery vehicles during the delivery process. Summary of the Invention
[0003] The purpose of this disclosure is to provide an image recognition method, a vehicle control method, and an electronic device.
[0004] To achieve the above objectives, a first aspect of this disclosure provides an image recognition method, comprising: Point clouds from multiple frames of images are fused into voxel space to construct a voxel structure on a single-frame representation. Based on the attribute representation on the feature map, attribute features of voxels in each voxel grid within the voxel structure are extracted according to the attribute representation. Based on the voxel coordinates and voxel indices of different voxel types within the voxel structure, feature masks corresponding to different coordinate axes of the voxels under the window structure are determined. These different coordinate axes are determined based on the voxels in a bird's-eye view. Voxels with different indexes in the feature masks are used to identify target point clouds corresponding to the point clouds in the multiple frames of the images. Based on the attribute features corresponding to the target point clouds, image recognition is performed on the multiple frames of the images to obtain the image recognition result.
[0005] A second aspect of this disclosure provides a vehicle control method, comprising: The system acquires environmental images of the vehicle during its operation; identifies the environmental images using the image recognition method described in any one of the first aspects to obtain an image recognition result; and controls the vehicle's operation based on the image recognition result.
[0006] A third aspect of this disclosure provides an electronic device comprising: A memory having a computer program stored thereon; a processor for executing the computer program in the memory to implement the steps of the method of any one of the first aspects.
[0007] The above technical solution achieves at least the following beneficial effects: It fuses point clouds from multiple frames of images into voxel space, constructing a voxel structure on a single-frame representation; based on the attribute representation on the feature map, it extracts the attribute features of voxels in each voxel grid within the voxel structure; based on the voxel coordinates and voxel indices of different voxel types within the voxel structure, it determines feature masks corresponding to different coordinate axes under the window structure, where different coordinate axes are determined based on the voxels in a bird's-eye view; it identifies the target point clouds from the point clouds corresponding to voxels with different representation indices in the feature masks across multiple frames; and it performs image recognition on the multiple frames based on the attribute features corresponding to the target point clouds, obtaining the image recognition results. This approach maintains the scale of the feature map on the coordinate axes, improving the accuracy of image recognition, while introducing a window for feature mask determination improves the efficiency of image recognition.
[0008] Other features and advantages of this disclosure will be described in detail in the following detailed description section. Attached Figure Description
[0009] The accompanying drawings are provided to further illustrate the present disclosure and form part of the specification. They are used together with the following detailed description to explain the present disclosure, but do not constitute a limitation thereof. In the drawings: Figure 1 This is a flowchart illustrating an image recognition method according to an exemplary embodiment.
[0010] Figure 2 This is an implementation illustrated according to an exemplary embodiment. Figure 1 The flowchart of S3 in the middle.
[0011] Figure 3 This is an implementation illustrated according to an exemplary embodiment. Figure 2 The flowchart of S34 in the middle.
[0012] Figure 4 This is an implementation illustrated according to an exemplary embodiment. Figure 3 The flowchart of S342 in the middle.
[0013] Figure 5 This is an implementation illustrated according to an exemplary embodiment. Figure 2 The flowchart of S33.
[0014] Figure 6 This is an implementation illustrated according to an exemplary embodiment. Figure 2 The flowchart of S35 in the middle.
[0015] Figure 7 This is an implementation illustrated according to an exemplary embodiment. Figure 1 The flowchart of S1.
[0016] Figure 8 This is an implementation illustrated according to an exemplary embodiment. Figure 7 The flowchart of S13.
[0017] Figure 9 This is an implementation illustrated according to an exemplary embodiment. Figure 8 The flowchart of S131.
[0018] Figure 10 This is an implementation illustrated according to an exemplary embodiment. Figure 8 The flowchart of S132.
[0019] Figure 11 This is an implementation illustrated according to an exemplary embodiment. Figure 1 The flowchart of S2.
[0020] Figure 12 This is a flowchart illustrating a vehicle control method according to an exemplary embodiment.
[0021] Figure 13 This is a block diagram illustrating an electronic device according to an exemplary embodiment. Detailed Implementation
[0022] The specific embodiments of this disclosure will be described in detail below with reference to the accompanying drawings. It should be understood that the specific embodiments described herein are for illustration and explanation only and are not intended to limit this disclosure.
[0023] It should be noted that all actions involving the acquisition of signals, information, or data in this disclosure are carried out in compliance with the relevant data protection laws and policies of the country where the location is situated, and with authorization from the owner of the relevant device.
[0024] Before introducing the image recognition method, vehicle control method, and electronic device provided in this disclosure, the techniques used in related scenarios are described. Point cloud data is typically generated by LiDAR or other sensors, containing a large number of points. However, these points are usually unevenly distributed in 3D space, resulting in sparsity in the point cloud. Convolutional Neural Networks (CNNs) waste significant computational resources when convolving sparse regions of the point cloud during convolution operations. Therefore, in related scenarios, Transformers (attention-based deep learning models) are used to process point clouds in natural language processing. Their self-attention mechanism allows the model to flexibly model long-distance relationships in the input data, making Transformers potentially effective in processing sparse point clouds and similar data. For example, it inherits all the benefits of window-based attention, reducing computational cost, the size of a single matrix, and improving computational parallelism.
[0025] However, directly applying the standard Transformer to sparse point clouds results in low computational efficiency and low accuracy in effective region partitioning. This is because optimizing and deploying the network structure only for the Set Attention-based attention mechanism neglects the deployment of data preprocessing and dynamic set recombination, leading to lower accuracy in effective region partitioning. Furthermore, the cost of switching between Transformer sequence features and voxel (voxel or volume pixel) point cloud feature spaces is far higher than visual feature extraction, resulting in low computational efficiency. Additionally, random data shuffling (scatter / gatter) in the dynamic feature encoding part is much more time-consuming than data rearrangement, also contributing to low computational efficiency. Moreover, the operators in the preprocessing and recombination processes are only supported on the PyTorch platform, which does not fully support the engineering deployment process of PyTorch-ONNX-TensorRT, resulting in low flexibility.
[0026] In view of this, this disclosure first provides an image recognition method, see [link to relevant documentation]. Figure 1 As shown, Figure 1 This is a flowchart illustrating an image recognition method according to an exemplary embodiment. It includes the following steps: S1. Fuse the point clouds from multiple frames of images into the voxel space to construct a voxel structure on a single-frame representation.
[0027] The voxel space is a data structure composed of voxels used to represent three-dimensional information. A voxel structure is a collection of voxels built on a single frame representation and used to store voxels.
[0028] In this embodiment, point clouds from multiple frames of images can be fused into a voxel space, allowing the point clouds to be converted into voxels and represented in a voxel structure, which can then be stored in a single-frame representation. For example, during data reading and preprocessing, multiple frames of the original image `range_image` from the sample image are input to the function `voxelize_concat_point_cloud`. This function calls the underlying registered `CustomOp` type `Voxelize ConcatPointCloud` to implement multi-frame fusion. The kernel function converts the point cloud data into a dense voxel structure `DenseVoxel`, and its internal steps include: first, allocating memory space `DeviceBuffer` for the parameters of the dense voxel structure; then, saving subsequently calculated parameters such as the index and starting position to the corresponding memory space.
[0029] Furthermore, based on the point indices (global index and thread index) of the point cloud in the image, the voxel indices corresponding to the voxels in the voxel space are calculated. Since the voxels in the voxel space are unordered, the CUDA CUB library function Sort Pairs can be used to sort the voxels in pairs based on their voxel indices and corresponding point indices, thus obtaining the sorted voxels.
[0030] Furthermore, the Unique function of the CUDA CUB library is used to remove duplicate values in the sorted voxel indices. One voxel index is retained for each voxel grid. Parameter checks are performed on the dense voxel structure, and the unique voxel indices are copied from device memory and synchronized to host memory. The Histogram Range function of the CUDA CUB library is used to solve for the number of voxels corresponding to each unique voxel index, and the Exclusive Sum function of the CUDA CUB library is used to solve for the starting position (offset) of each voxel in the sorted point index.
[0031] In this embodiment, the voxel structure built on a single-frame representation can be a point cloud from multiple frames of images. The point cloud data is dynamically divided into pillars according to spatial location using a Dynamic Pillar VFE 3D Layer (a network layer for dynamically processing 3D point clouds) in the model, and feature extraction is performed within each pillar. This captures the spatial distribution features of the point cloud data, thereby improving the accuracy of target detection. Dynamic pillaring divides the point cloud into multiple pillars based on its spatial distribution. Each pillar contains point cloud data within a certain range, ensuring a relatively balanced number of points within each pillar. Feature extraction is performed on the point cloud data within each pillar using a specific feature extraction method (such as a multilayer perceptron or convolutional neural network) to extract the feature representation of each pillar. These feature representations reflect the spatial distribution, density, and shape of the point cloud within each pillar. Feature aggregation combines the feature representations of all pillars to form a global feature map, thereby converting the input point cloud into sparse voxels.
[0032] For example, automated delivery vehicles can acquire multiple frames of images collected at the same time from multiple sensors, or obtain 3D point cloud data from multiple frames of images collected by a single sensor at different times, in order to more accurately understand the environment. By fusing these point cloud data into a unified voxel space, a voxel structure represented by a single frame can be constructed, which can more comprehensively reflect the surrounding environmental information.
[0033] S2. Based on the attribute representation on the feature map, extract the attribute features of the voxels in each voxel grid of the voxel structure on the attribute representation; The attribute representation can be a numerical representation of the attribute information of each point or region on the feature map. For example, the attribute can be the intensity or height of a voxel. A voxel mesh is a subdivided region in voxel space, and each mesh can include a certain number of voxels. Of course, there may be voxel meshes in a voxel structure that contain no voxels, i.e., there may be invalid voxel meshes.
[0034] In this embodiment of the disclosure, the attribute features of each voxel in the voxel grid can be extracted by the attribute representation (such as color, texture, height, etc.) on the feature map, so that pedestrians or vehicles in the image can be identified in subsequent steps.
[0035] S3. Based on the voxel coordinates and voxel indices of different types of voxels in the voxel structure, determine the feature masks of the voxels corresponding to different coordinate axes under the window structure. The different coordinate axes are determined based on the voxels in the bird's-eye view. In this context, a feature mask is used for multi-attention focusing on point clouds or voxels in specific regions or directions across multiple frames of images. This allows for image recognition based on the attribute features of the point clouds or voxels in those specific regions or directions. Specifically, a window structure can be defined based on the different types and coordinates of the voxels. This window structure can include multiple windows, and it defines the voxel region of interest. Then, based on the voxel indices and coordinates, a feature mask can be generated. This mask highlights the voxel features within a specified range while suppressing features in other regions, thus enabling the focus on voxels of interest through the feature mask.
[0036] Specifically, each element in the feature mask corresponds to a voxel. If the voxel is located within the region of interest, the element at the corresponding position in the feature mask is 1 (or a high value); otherwise, the element at the corresponding position in the feature mask is 0 (or a low value).
[0037] In this embodiment, to establish connections between sparse voxels, computation can be performed using rotating sets and blended windows, introducing feature propagation within and between windows while maintaining efficient computation. The extracted voxel features are then projected onto a bird's-eye view (BEV) feature map. Finally, various sensing heads can be attached to various 3D perception tasks.
[0038] In S4, the voxels with different indices in the feature mask are used to determine the point clouds corresponding to the images in multiple frames as the target point clouds.
[0039] In this embodiment, voxels with different indices can be selected using a feature mask; that is, voxels that exhibit significant variations or uniqueness in their feature representations. These voxels may represent important objects or scene structures in the point cloud data corresponding to multiple image frames. Identifying these voxels as target point clouds helps subsequent image recognition tasks to focus more on key information.
[0040] In S5, based on the attribute features corresponding to the target point cloud, image recognition is performed on multiple frames of the image to obtain the image recognition result.
[0041] In this embodiment of the disclosure, after determining the target point cloud, image recognition is performed using the attribute features (such as voxel height, reflection intensity, etc.) corresponding to these point clouds. For example, machine learning models (such as convolutional neural networks, support vector machines, etc.) are used to learn and classify the features. Through model learning and inference, recognition results for multiple frames of images can be obtained, such as object classification labels, detection boxes, or segmentation masks.
[0042] In this embodiment of the disclosure, image recognition is used to identify objects or scenes in multiple frames of images. In point cloud processing, image recognition can be used to classify, detect, or segment objects or scenes in point cloud data. For example, it can be used to implement 3D object detection and point cloud segmentation.
[0043] In image recognition, the model can focus only on the attribute features of voxels highlighted by the feature mask, ignoring the attribute features of other regions. This helps reduce computational load, improve recognition efficiency, and allows the model to focus more on the region of interest, thus improving image recognition accuracy. For example, in autonomous vehicle delivery scenarios, the model may only be concerned with objects and pedestrians on the road, without needing to consider buildings or trees on either side. By generating a feature mask that highlights only the voxel features of the road area, the recognition model can focus more on recognizing the attribute features of these key elements, thereby improving the driving safety of autonomous delivery vehicles.
[0044] In this embodiment, sparse features corresponding to multiple frames of images acquired in the same batch can be concatenated based on the effective voxel count of each sample and then fed into a network model. The network model can extract attribute features corresponding to the voxels, and then, based on a Set Attention Transformer model, extract network features of high-order information of the voxels annotated by the feature mask. These high-order network features can then be fed into downstream branches processing perception tasks to achieve specific tasks such as 3D object detection and point cloud segmentation, thus realizing image recognition.
[0045] The above technical solution fuses point clouds from multiple frames of images into a voxel space, constructing a voxel structure on a single-frame representation. Based on the attribute representation on the feature map, it extracts the attribute features of voxels in each voxel grid within the voxel structure. Based on the voxel coordinates and voxel indices of different voxel types within the voxel structure, it determines feature masks corresponding to different coordinate axes under a window structure. These different coordinate axes are determined based on the voxels in a bird's-eye view. Voxels with different representation indices in the feature masks are used to determine the point clouds corresponding to the multiple frames of the images as target point clouds. Based on the attribute features corresponding to the target point clouds, it performs image recognition on the multiple frames of the images to obtain the image recognition result. This approach maintains the scale of the feature map on the coordinate axes, improving the accuracy of image recognition, while introducing a window for feature mask determination improves the efficiency of image recognition.
[0046] Optionally, see Figure 2 As shown, in S3, determining the feature mask corresponding to different coordinate axes of the voxel under the window structure based on the voxel coordinates and voxel indices of different voxel types in the voxel structure includes: In S31, the unique inverse index of each voxel in the window structure is determined based on the sorted voxel index of each voxel in the voxel structure, the number of windows in the window structure, and the number of first voxels in the window structure.
[0047] Among them, the unique inverted index assigns a unique, reversed index value to each voxel.
[0048] In this embodiment of the disclosure, the starting point of the inverse index is determined based on the sorted voxel index of the first voxel of the first window in the window structure. For example, if the sorted voxel index of the first voxel of the first window in the voxel structure is start_index, then the starting point of the inverse index can be start_index.
[0049] Furthermore, determine the step size for each increment of the inverted index. For example, this can be determined based on the total number of voxels in the window structure and the number of windows. If each window has the same number of voxels (i.e., a fixed-size window), then the increment of the inverted index will be the reciprocal of the window size.
[0050] Furthermore, starting from the initial point, a new inverse index is assigned to each voxel in the window structure based on the calculated inverse index increment. This inverse index can be ordered from largest to smallest, thus allowing inverse index allocation in reverse order. When reaching the end of a window and moving to the next, the continuity of the inverse index needs to be ensured. This typically means that at the beginning of each new window, the inverse index will be based on the inverse index of the last voxel of the previous window, minus a value related to the window size. The generated inverse index is then associated with the corresponding voxel and stored in an appropriate data structure.
[0051] For example, suppose there are three windows, each containing 50 voxels. The sorted voxel indices in the voxel structure start from 0, so the voxel indices for the first window are 0 to 49, for the second window they are 50 to 99, and for the third window they are 100 to 149. We can start with the last voxel in the last window (index 149) and assign it an inverse index of 0 (because it's in reverse order). Then, the inverse indices decrease in increments of -1, assigning a unique inverse index to each voxel. For the voxels in the second window, the inverse index will start from -50 and decrease in the same increments. For the voxels in the first window, the inverse index will start from -100. Ultimately, each voxel will have a unique inverse index associated with it, and these inverse indices are consecutive and unique within the window structure.
[0052] In S32, based on the unique inverse index of each voxel in each window structure, the window set partitioning parameter, and the voxels in each window, the window voxel index in the bird's-eye view after the voxels in each window are partitioned into different sets is determined.
[0053] The window set partitioning parameter is used to determine the number of voxels in each set.
[0054] The window voxel index is an index assigned to each voxel in the bird's-eye view after voxels are divided into different sets according to the window set partitioning parameters. The window set partitioning parameters determine the number of voxels that should be included in each set, thus determining which sets the voxels are assigned to.
[0055] In this embodiment, voxels are divided into different sets according to window set partitioning parameters (e.g., each set includes 36 voxels). Then, a window voxel index is assigned to each set. This allows for quick location of voxels in a specific set. Continuing with the previous embodiment, since each window can include 50 voxels, indexing each window will form two sets, where the first set includes 36 voxels and the second set includes 14 voxels.
[0056] In a bird's-eye view, voxels lose features along the Z-axis (height direction). For tasks like obstacle recognition, more attention is paid to the size of obstacles along the X and Y axes to avoid collisions. Therefore, window voxel indexing in a bird's-eye view can be scaled along the X and Y axes while avoiding the introduction of features along the Z-axis, thus reducing computational load.
[0057] In S33, multiple different index sorting sequences are determined based on the unique inverse index of each voxel, the voxel coordinates of each window in the voxel structure of different types in the voxel structure, and the voxel index.
[0058] Specifically, the original coordinate voxel index can be determined based on the voxel coordinates and voxel indices before sorting, and the sorted coordinate voxel index can be determined based on the voxel coordinates and voxel indices after sorting. Then, a first index is constructed based on the unique inverse index of each voxel and the original coordinate voxel index, and a second index is constructed based on the unique inverse index of each voxel and the sorted coordinate voxel index. The first and second indices are then combined to obtain a combined index sequence. Finally, based on the Sort Pairs operator, the combined index sequence is sorted according to the first and second indices respectively, resulting in multiple different index sorting sequences, including a first index sorting sequence based on the first index and a second index sorting sequence based on the second index.
[0059] In S34, different global indexes of the windows are determined based on the unique inverse index of each voxel, the window voxel index in the corresponding window, and multiple different index sorting sequences.
[0060] The window global index is used to represent the index that uniquely identifies the position of each voxel within the corresponding window in the entire voxel structure. By constructing the window global index, voxel data from different angles can be easily aligned and fused, resulting in more complete and accurate 3D reconstruction results.
[0061] In this embodiment, a unique inverted index can be matched with the window voxel index. This can be achieved using a lookup table or similar data structure, allowing for rapid local location of each voxel within a specific window. Further, one or more index sorting sequences are selected, and the voxel order in the global index is determined based on these sequences. For example, the matched inverted index and window voxel index are sorted or rearranged according to the selected sequences. Based on the sorted inverted index and window voxel index, a global window index is generated. This global index is unique and accurately identifies the position of each voxel within a specific window. It combines the global uniqueness of the inverted index with the local position information of the window voxel index.
[0062] In S35, a feature mask corresponding to the coordinate axis is determined based on the index movement step size, index movement direction, coordinate axis under the bird's-eye view, and different global indexes of the window.
[0063] In this context, a feature mask is a binarized image or matrix used to highlight or suppress features in a specific region. In voxel processing, the feature mask is used to identify the voxel region of interest. The index movement step size and index movement direction are used to determine the parameters for traversing the voxel index when generating the feature mask. The size of the index movement step size determines the number of voxels traversed in each traversal, while the index movement direction determines the direction of traversal (e.g., from left to right, from top to bottom, etc.). The coordinate axes in a bird's-eye view are the coordinate axes used to describe the voxel positions when viewing 3D space from a top-down perspective; these are the X and Y axes.
[0064] In this embodiment, a feature mask matrix of the same size as the voxel structure is created, and all its elements are initialized to 0 (or other values representing non-features). This matrix will be used to store the final feature mask result. All window global indices are traversed according to the index movement step size and index movement direction. For each index, the position and attributes of its corresponding voxel within the voxel structure are found.
[0065] Furthermore, based on the voxel's attributes (such as type and size), it is determined whether it meets the conditions of a specific feature. For example, if the focus is on vehicles on a road, vehicle voxels can be identified based on their size and type. If a voxel meets the feature conditions, its value at the corresponding position in the feature mask matrix is updated to 1 (or another value representing the feature). In this way, the feature mask matrix gradually accumulates the location information of voxels that meet the conditions.
[0066] Furthermore, during the traversal, attention is paid to the voxel positions corresponding to the selected coordinate axes. Based on the direction and extent of the coordinate axes, the feature mask is adjusted or cropped accordingly to ensure it accurately reflects the feature distribution on that coordinate axis.
[0067] Optionally, see Figure 3 As shown, in S34, determining the window voxel index in the bird's-eye view after the voxels in each window are divided into different sets, based on the unique inverse index of each voxel in each window structure, the window set partitioning parameter, and the voxels in each window, includes: In S341, the voxels in each window are transformed by perspective to obtain the bird's-eye view voxels of each voxel in the bird's-eye view.
[0068] Among them, bird's-eye view voxels can be the projection representation of voxels on a two-dimensional plane under a bird's-eye view.
[0069] In this embodiment of the disclosure, voxels are projected from the original three-dimensional spatial perspective onto a bird's-eye view to obtain bird's-eye voxels. For example, voxels can be mapped onto a two-dimensional plane based on their three-dimensional coordinates by ignoring information in one dimension (e.g., height) to obtain bird's-eye voxels in the bird's-eye view for each of the windows.
[0070] In S342, based on the window set partitioning parameters and the voxels in each window, the bird's-eye view index of each bird's-eye view voxel in each window is determined under the bird's-eye view.
[0071] Among them, the bird's-eye view index can be a unique identifier or index assigned to a bird's-eye view voxel under the bird's-eye view.
[0072] In this embodiment of the disclosure, a unique bird's-eye view index is assigned to each bird's-eye view voxel based on the window set partitioning parameters and voxel attributes. This process may involve classifying, sorting, or grouping voxels in order to assign them appropriate indexes in the bird's-eye view.
[0073] In S343, based on the unique inverse index of each voxel in each window structure and the sorted voxel index of each voxel in the voxel structure, the bird's-eye index of each bird's-eye voxel in each window is converted into the window voxel index of the bird's-eye voxel in the corresponding window.
[0074] In this embodiment, a bird's-eye view index can be converted into a window voxel index by combining a unique inverse index and a sorted voxel index. This conversion is to associate the position of a voxel in the global structure with its position within a specific window, thereby facilitating subsequent data processing and analysis at the window level.
[0075] Optionally, see Figure 4 As shown, in S342, determining the bird's-eye view index of each bird's-eye view voxel in each window under the bird's-eye view based on the window set partitioning parameters and the voxels in each window includes: In S3421, the number of sets in each window is determined based on the window set partitioning parameters and the number of second voxels in each window.
[0076] Here, the number of sets refers to the number of sets into which the voxels in a window are divided according to the window set partitioning parameter. Following the previous embodiment, each window includes 50 voxels, and the window set partitioning parameter is 36. Therefore, the number of sets in the window is 2. Specifically, the first window contains 36 voxels, and the second window contains 14 voxels.
[0077] In S3422, the window index corresponding to each window and the window set index of each set in each window are determined according to the number of sets in each window and the window set partitioning parameters.
[0078] The window index uniquely identifies a window and is used to quickly locate and process data within that window in a data structure or algorithm. The window set index uniquely identifies each set within a window and is used to distinguish different voxel sets within the same window.
[0079] In this embodiment of the disclosure, a unique window index and window set index are assigned to each window and its internal sets based on the determined number of sets and window set partitioning parameters. These index values are used to quickly locate and manage voxel data within the window and its sets in subsequent processing. For example, suppose there are two windows (W1 and W2), each with a different voxel set. Based on the number of sets and partitioning parameters, window index 1 can be assigned to W1, and window set index A can be assigned to the first set within it, and window set index B can be assigned to the second set. Similarly, window index 2 is assigned to W2, and corresponding window set indices are assigned based on the sets within it.
[0080] In S3423, the bird's-eye view index of each bird's-eye view voxel in each window is determined based on the number of sets in each window, the corresponding window index, and the window set index of the set in each window.
[0081] Specifically, a unique bird's-eye view index is assigned to each bird's-eye view voxel by combining the determined number of sets, window index, and window set index. This process ensures that each voxel has a clear and unique identifier in the bird's-eye view, facilitating subsequent data processing and analysis.
[0082] For example, for each bird's-eye view voxel in a window, a unique bird's-eye index can be assigned by combining the window index of its own window, the window set index of its own set, and its relative position within the set. In this way, each vehicle has a clear identifier in the bird's-eye view.
[0083] Optionally, in S3422, determining the window index corresponding to each window and the window set index of each set in each window based on the number of sets in each window and the window set partitioning parameter includes: Based on the number of sets and the window set partitioning parameters, the voxels in each window are divided into sets, and set numbers are added to the sets in each window in sequence.
[0084] Among them, the set number is a unique number used to identify and distinguish different sets. Usually, each set is assigned a unique number or character identifier according to a certain rule (such as sequential numbering).
[0085] In this embodiment of the disclosure, voxels in the window are divided into different sets according to the determined number of sets and window set partitioning parameters. After the partitioning is completed, a unique set number is added to each set in a certain order (such as the order in which the sets appear or according to a certain priority).
[0086] Based on the number of sets corresponding to each window, calculate the cumulative sum of the set numbers in each window to obtain the cumulative sum sequence corresponding to each window.
[0087] The sum of the set numbers in each window can be calculated using the Inclusive Sum operator. The sum is calculated by adding the first set number to the second set number, then adding the first sum to the third set number, and so on, until the sum of all set numbers is calculated, resulting in a final sum. The sums corresponding to these sets form a sum sequence. For example, suppose there are two windows. The first window has three sets, numbered 1, 2, and 3. The first sum is 3 (1 + 2), and the second sum is 6 (3 + 3). Therefore, the sum sequence for this window is [3 6]. The second window has four sets, numbered 4, 5, 6, and 7. Similarly, the sum sequence is [9 15 22].
[0088] Based on the accumulated sum sequence, determine the window index corresponding to each window and the window set index of the set of each window.
[0089] The window index can be obtained directly from the window's order or number, while the window set index can be obtained by accumulating and transforming or calculating the values in the sequence. This process ensures that each window and each set within a window has a unique index value.
[0090] Optionally, determining the window index corresponding to each window and the window set index of the set of windows based on the accumulated sum sequence includes: Add the same marker value to each of the accumulated sum sequences corresponding to each window to obtain a marker value sequence corresponding to each window; The marker value is a fixed value used to distinguish the sums of different windows and generate a unique window index. In this embodiment, the length of the marker value sequence can be determined based on the longest sum sequence in the window. Then, a fixed value is added to the positions in the marker value sequence where there is a sum, and another fixed value is added to the positions in the marker value sequence where there is no sum. Taking the example of the aforementioned embodiment, the sum sequence of the first window is [3 6], and the sum sequence of the second window is [9 15 22]. Then, the marker value sequence corresponding to the first window is [1 10], and the marker value sequence corresponding to the second window is [1 1 1].
[0091] The sum of the marker values in the sequence of marker values corresponding to each window is used as the window index corresponding to that window. Similarly, the sum of the marked values can be calculated based on the Inclusive Sum operator. Following the previous example, the marked value sequence corresponding to the first window is [1 1 0], and the window index is [2 2]. The marked value sequence corresponding to the second window is [1 1 1], and the window index is [2 3].
[0092] Based on the cumulative sum sequence corresponding to each window and the window index, determine the window set index of each set in each window.
[0093] In this embodiment of the disclosure, the accumulated sum sequence and the window index can be combined to obtain the window set index of each set in each window, or the accumulated sum sequence and the window index can be added together at the same position to obtain the window set index of each set in each window.
[0094] Optionally, see Figure 5 As shown, in S33, determining multiple different index sorting sequences based on the unique inverse index of each voxel, the voxel coordinates of different types of voxels in the voxel structure of each window in the window structure, and the voxel index includes: In S331, a reference index corresponding to each voxel and its type is constructed based on the unique inverse index of each voxel, the voxel coordinates of each window in the window structure of different types in the voxel structure, and the voxel index.
[0095] In this embodiment of the disclosure, based on the unique inverse index of each voxel, the voxel coordinates of each window in the voxel structure of different types in the voxel structure, and the voxel index, an ordered reference index and an original reference index are assigned to each voxel in parallel.
[0096] In S332, based on the reference indexes corresponding to each voxel and different types, different index sorting sequences are obtained by sorting with the reference indexes of each type as the sorting target.
[0097] In this embodiment, based on the Sort Pairs function, the reference indices corresponding to each type are sorted using the reference index as the sorting target, resulting in different index sorting sequences. For example, voxels are sorted using ordered reference indices as the target; if the ordered reference indices are the same, they are then sorted according to the original reference index. Furthermore, voxels are sorted using the original reference index as the target; if the original reference indices are the same, they are then sorted according to the ordered reference index. This results in two different index sorting sequences.
[0098] Optionally, in S331, constructing a reference index corresponding to the type of each voxel in the window structure based on the unique inverse index of each voxel, the voxel coordinates of each window in the window structure of different types in the voxel structure, and the voxel index, includes: Based on the voxel coordinates and voxel indexes of the voxels of each window in the voxel structure for different types in the voxel structure, determine the window voxel coordinates of each voxel in the corresponding window corresponding to each type. Based on the window voxel coordinates of each voxel in the corresponding window and the window voxel coordinates of each type, construct an array of window voxel coordinates of each voxel in the window and the window voxel coordinates of each type. Based on the unique inverse index of each voxel and the window voxel coordinate array corresponding to each type, construct the reference index of each voxel in the window structure corresponding to the type.
[0099] In this embodiment of the disclosure, the sorted window coordinates of each voxel in the corresponding window can be determined according to the sorted voxel index of each voxel; the first window coordinate array of the window can be determined according to the sorted window coordinates of each voxel in the window; the original window coordinates of each voxel in the corresponding window can be determined according to the original voxel index of each voxel; and the second window coordinate array of the window can be determined according to the original window coordinates of each voxel in the window.
[0100] In this embodiment of the disclosure, the CUDA kernel function Get Vox Coord In Win Kernel can be used to calculate the coordinates of each voxel in its window in parallel based on the input sorted voxel index, and return an array of the coordinates of each sorted voxel in its window, as well as an array of the coordinates of each original voxel in its window.
[0101] In this embodiment of the disclosure, an index sequence can be constructed based on the ordered reference index and the corresponding original reference index assigned to each voxel; based on the Sort Pairs function, voxels are sorted with any index in the index sequence as the target, to obtain the first sorting sequence and the second sorting sequence corresponding to each window.
[0102] Optionally, see Figure 6 As shown, in S35, determining the feature mask corresponding to the coordinate axis based on the index movement step size, index movement direction, the coordinate axis under the bird's-eye view, and different global indices of the window includes: In S351, based on the coordinate axes under the bird's-eye view, coordinate axis set indices corresponding to the coordinate axes are assigned to different window global indices.
[0103] In this embodiment of the disclosure, the coordinate axes in the bird's-eye view are the X-axis and the Y-axis. Therefore, a set index in the Y-axis direction and a set index in the X-axis direction can be assigned to each voxel in parallel, thereby obtaining the X-axis set index and the Y-axis set index.
[0104] In S352, each coordinate axis set index is moved according to the index movement step size and the index movement direction to obtain the corresponding movement set index.
[0105] For example, if the index movement step is 1 and the index movement direction is to the right, then each element in the coordinate axis set index is moved 1 position to the right to obtain the corresponding movement set index.
[0106] In this embodiment of the disclosure, the coordinate axis set index can be kept at the same length, so a filling index can be assigned in parallel to each set whose internal voxel count is not full. For example, a filling index can be assigned in parallel to each set whose internal voxel count is not full using a fixed value (e.g., a fixed value of 0). For instance, if the coordinate axis set index is 36 digits, that is, the window set partitioning parameter can be used as the standard length of the coordinate axis set index, if the coordinate axis set index is less than 36 digits, 0s can be added to the end until the coordinate axis set index reaches the standard length.
[0107] In S353, the coordinate axis set index is compared with the corresponding movement set index.
[0108] In this embodiment of the disclosure, it is possible to compare whether the elements at the same position of the coordinate axis set index and the corresponding moving set index are the same.
[0109] In S354, the feature mask corresponding to the coordinate axis is determined based on whether the indices at each position in the index comparison result are the same.
[0110] In this embodiment of the disclosure, different masks are obtained based on whether the indices at each position in the index comparison result are the same. For example, if the indices at the same position in the index comparison result are the same, the mask is 1; if the indices at the same position in the index comparison result are different, the mask is 0. Thus, a feature mask is constructed based on the position of the index and the mask.
[0111] The bird's-eye view contains X-axis set indexes and Y-axis set indexes, thus, a feature mask corresponding to the X-axis and a feature mask corresponding to the Y-axis can be obtained respectively.
[0112] Optionally, the different types of voxel coordinates and voxel indices include: voxel coordinates and voxel indices before sorting, and voxel coordinates and voxel indices after sorting.
[0113] Optionally, see Figure 7 As shown, in S1, the process of fusing point clouds from multiple frames of images into voxel space to construct a voxel structure on a single-frame representation includes: In S11, a global index of the point cloud is constructed for each point cloud in the multiple frames of the image. As is understandable, a point cloud is a collection of numerous points in three-dimensional space, typically acquired using devices such as LiDAR or depth cameras. To uniquely identify each point cloud within the entire dataset or a series of data frames, a unique identifier, known as a global point cloud index, can be added. This global point cloud index allows for querying specific point clouds. Since point clouds are distributed unordered in space, directly processing and analyzing point cloud data is often inefficient; therefore, constructing a global point cloud index is necessary.
[0114] In this embodiment, the entire point cloud data space is divided into a series of small regions or units. The size and shape of these regions or units can be determined according to specific application requirements. Through spatial division, unordered point cloud data can be transformed into ordered structured data. Furthermore, each divided region or unit is assigned a unique identifier. This identifier can be in the form of numbers, strings, etc., but the key is to ensure that it is unique throughout the entire point cloud dataset. Simultaneously, the number of points contained in each region or unit and other relevant information are recorded.
[0115] Furthermore, after completing the division of regions or cells and assigning unique identifiers, the global index of the point cloud is constructed. In practice, the entire point cloud dataset is traversed. For each point, its spatial location determines its region or cell, and the point's information (such as coordinates, attributes, etc.) is associated with the unique identifier of that region or cell. This forms a unique mapping from points to regions or cells, i.e., the global index of the point cloud.
[0116] In S12, each point cloud is mapped onto a voxel grid in voxel space; In this context, a voxel is the smallest unit in three-dimensional space, similar to a pixel, used to represent three-dimensional data. A voxel grid is a three-dimensional mesh structure composed of voxels, used to divide three-dimensional space into a series of discrete units. Mapping refers to assigning points in a point cloud to corresponding voxel grids according to their spatial location.
[0117] In this embodiment, point cloud data, originally discretely distributed in three-dimensional space, is organized according to specific rules and structures. For each point in the point cloud, its position in voxel space is determined. For example, this can be achieved by calculating the position of the point's three-dimensional coordinates (X, Y, Z) relative to the voxel grid. Specifically, the voxel to which the point belongs can be determined by comparing the point's coordinates with the boundary of the voxel grid. This process may involve coordinate normalization, rounding, and other operations to ensure that the point can be accurately assigned to the corresponding voxel.
[0118] Furthermore, after determining the voxel to which each point belongs, these points can be mapped to the corresponding voxel mesh. This typically means associating the point's information (such as coordinates, color, normals, etc.) with the corresponding voxel in the voxel mesh. In this way, the originally discrete point cloud data can be organized into structured voxel mesh data.
[0119] Furthermore, during the mapping process, point cloud data may exhibit uneven distribution; some voxels may contain a large number of points, while others may contain only a few points or even none. To effectively represent this distribution, statistical methods can be employed, such as calculating the number, density, or average coordinates of points within a voxel, and storing this statistical information as voxel attributes. For example, this information can be stored in pre-allocated storage space for the voxels.
[0120] In this embodiment of the disclosure, data structures such as spatial hash tables can be used to perform the voxel allocation process, and parallel computing technology can be used to process the mapping operation of multiple points at the same time, thereby improving the processing speed.
[0121] In S13, based on the global index of each point cloud, a unique voxel index is determined after sorting the voxels in each voxel grid.
[0122] Sorting is the process of rearranging data, usually according to certain rules (such as numerical value, spatial position, etc.). A unique voxel index is a unique identifier assigned to each voxel after sorting, used to locate that voxel in the voxel structure.
[0123] It is understandable that when a point cloud is mapped to a voxel grid in voxel space, the voxel grid may contain a certain number of points. That is, different point clouds are mapped to the same voxel grid. Determining the unique voxel index after sorting the voxels in each voxel grid is actually organizing the points inside the voxel grid, so that a unique index can be assigned to each voxel.
[0124] In this embodiment, the sorting criteria can be determined based on specific application requirements. For example, the global index can be sorted based on attributes such as point coordinates, color, and density. This organizes the points within the voxel grid in a certain order. After sorting, a unique index can be assigned to each voxel in the voxel grid. This index is typically generated based on the sorted point order, ensuring that each voxel in the entire voxel grid has a corresponding unique index. This unique voxel index not only identifies the voxel's position in the grid but also implicitly contains the sorting information of the points within the voxel.
[0125] In S14, the starting position of each voxel in the voxel grid after sorting is determined in the global index of the point cloud.
[0126] The starting position can be the initial position of a voxel within a voxel grid in the sorted global index of the point cloud. Determining the starting position of a voxel in the global index establishes a relationship between the voxel grid and the global index. By using the starting position of each voxel in the global index, the corresponding data segment of that voxel grid can be quickly located.
[0127] In S15, the voxel structure on the single-frame representation is constructed based on the voxel point number of the voxels in the voxel grid, the corresponding starting position, and the unique voxel index.
[0128] The Voxel Count can be the number of voxels contained in a voxel grid.
[0129] In this embodiment of the disclosure, during the construction of the voxel structure, the basic framework of the voxel structure can be determined based on the size and resolution of the voxel mesh. This framework is typically a multidimensional array or a similar data structure used to store and manage voxel data. The dimensions of the array can be determined based on the dimensions of the voxel mesh.
[0130] Furthermore, after traversing and counting the voxel points of each voxel grid, data structures such as hash tables or index trees are used during the construction of the voxel structure to accelerate voxel lookup and location operations. Voxel data can also be compressed or encoded to reduce storage space usage and improve data transmission efficiency. An ordered, structured dataset is constructed, containing information such as the voxel point count, starting position, and unique voxel index for each voxel. This structure facilitates the querying, access, and manipulation of voxel data.
[0131] Optionally, see Figure 8 As shown, in S13, determining the unique voxel index after sorting the voxels in each voxel grid based on the global index of each point cloud includes: In S131, the original voxel index of each voxel in each voxel mesh is determined based on the global index of each point cloud.
[0132] In this embodiment, the voxel grid to which each point belongs is determined based on the global index of the point cloud. For example, the three-dimensional coordinates of the point can be compared with the spatial extent of the voxel grid to determine the grid cell to which the point belongs. Furthermore, for each voxel grid, the number of points it contains is counted, and the points are marked according to their global index. This marking information will be used for subsequent allocation of the original voxel index. Based on the point count and marking information within the voxel, an original index value is assigned to each voxel. This index value can be determined based on the voxel's position in the grid, the number of points, or other relevant attributes. At this stage, the index value may not be unique, as multiple voxels may have the same index value. Moreover, the original voxel index is in a state of disorder.
[0133] In S132, the voxels in each voxel grid are sorted according to the original voxel index of each voxel and the point cloud global index, so as to obtain the sorted voxel index of each voxel grid.
[0134] In this embodiment of the disclosure, voxels can be arranged in an orderly manner according to certain rules or standards to obtain a sorted voxel index. The sorting can be based on voxel attributes (such as the number, density, and spatial location of points within a voxel) or on the order of the point cloud global index.
[0135] In this embodiment of the disclosure, during the sorting process, based on the adjacent or overlapping nature of voxels in the voxel grid, the voxels can be sorted according to their spatial positional relationship and a certain traversal order (such as depth-first, breadth-first, etc.).
[0136] In this embodiment of the disclosure, algorithms such as quicksort, mergesort, and heapsort can be selected for voxel sorting. After sorting the voxels, the sorted voxel indices in each voxel grid can be obtained. These index values will be arranged in order according to the sorting rules.
[0137] In S133, duplicate voxel indices in each voxel grid after sorting are removed to obtain a unique voxel index corresponding to the voxel in each voxel grid.
[0138] In this embodiment of the disclosure, if there are multiple reread voxel indices in the same voxel grid, one voxel index is retained and the others are deleted. In this way, voxels in each voxel grid can be represented by the same voxel index.
[0139] Optionally, see Figure 9 As shown, in S131, determining the original voxel index of each voxel in the voxel mesh based on the global index of each point cloud includes: In S1311, taking any coordinate axis of the voxel space as the first target, the first quotient and the first remainder of the number of voxel grids of the global index of each point cloud on that coordinate axis are determined. In S1312, taking any remaining coordinate axis of the voxel space as the second target, the second quotient and the second remainder of the first quotient corresponding to each point cloud on the value of the number of voxel grids on that coordinate axis are determined. In S1313, the remaining coordinate axis of the voxel space is used as the third target to determine the third remainder of the second quotient corresponding to each point cloud on the value of the number of voxel grids on the coordinate axis. In S1314, the voxel index of each voxel in the voxel grid is determined based on the first remainder, the second remainder, and the third remainder.
[0140] In this embodiment of the disclosure, it is assumed that there is a three-dimensional point cloud dataset, which is obtained by LiDAR scanning, and each point is assigned a unique global point cloud index. Now, these points can be divided into a voxel grid, and an original voxel index can be assigned to each voxel.
[0141] For example, suppose the size of the voxel space is 10x10x10, meaning that 10 voxel grids are divided in each dimension. Taking the X-axis as the first target, suppose the global index of the point cloud at a certain point is idx = 555. This index can be divided by the number of voxel grids on the X-axis, 10, to obtain the first quotient quotient1 = 55 and the first remainder remainder1 = 5.
[0142] Furthermore, the Y-axis is selected as the second target. The first quotient quotient1 = 55 obtained in the previous step is divided by the number of voxel grids on the Y-axis, 10, to obtain the second quotient quotient2 = 5 and the second remainder remainder2 = 5.
[0143] Finally, the Z-axis is selected as the third target. Using the second quotient quotient2 = 5 obtained in the previous step, and dividing it by the number of voxel grids on the Z-axis, 10, since 5 is less than 10, the third remainder 3 = 5 is obtained directly. Based on the above calculation, three remainders are obtained: remainder1 = 5, remainder2 = 5, and remainder3 = 5. These three remainders actually represent the voxel grid positions of the point on the three coordinate axes. Therefore, these three remainders can be combined to form a triple (5, 5, 5), which serves as the original voxel index of the voxel where the point is located.
[0144] In this way, the voxel grid to which each point in voxel space belongs can be determined, and a unique original voxel index can be assigned to each voxel.
[0145] Optionally, see Figure 10 As shown, in S132, the step of sorting the voxels in each voxel grid according to the original voxel index of each voxel and the point cloud global index to obtain the sorted voxel index in each voxel grid includes: In S1321, an index sequence for each voxel is constructed based on the original voxel index corresponding to each voxel and the point cloud global index.
[0146] In this embodiment, each voxel has two key pieces of information: an initial voxel index and a point cloud global index. The initial voxel index is the identifier initially assigned to the voxel, while the point cloud global index is the unique identifier for each point in the point cloud. By combining these two indices, a unique index sequence can be constructed for each voxel. This index sequence contains both the voxel's position information in the mesh and the global information of points within the voxel. The initial voxel index or the point cloud global index can be placed first, followed by the remaining index, to construct the voxel index sequence.
[0147] In S1322, based on the index at any same position in each of the index sequences, the voxels in each voxel grid are sorted to obtain the sorted voxel index in each voxel grid.
[0148] In this embodiment of the disclosure, after obtaining the index sequence of each voxel, the voxels can be sorted according to the index value at a certain position in the index sequence. For example, the first position of the index sequence can be selected as the sorting key, and all voxels can be sorted according to this key. In this way, voxels with the same key value will be grouped together. For example, the voxels in each voxel grid can be sorted based on the Sort Pairs operator to obtain the sorted voxel indexes in each voxel grid.
[0149] Optionally, see Figure 11 As shown, in S2, the step of extracting the attribute features of voxels in each voxel grid of the voxel structure based on the attribute representation on the feature map includes: In S21, a valid voxel grid containing voxels is determined in the voxel structure.
[0150] Understandably, during voxelization, due to the sparsity or non-uniform spatial distribution of the point cloud, some voxel grids may not contain any point cloud data. These grids are ineffective for subsequent feature extraction and analysis. Therefore, it is necessary to first determine which grids are valid so that feature extraction can be performed only on these grids. This is achieved by checking whether each grid contains a voxel. For example, all grids can be traversed, and if a grid contains at least one voxel, it can be considered valid.
[0151] In S22, for the effective voxel grid, the voxel features of the voxels in the effective voxel grid are extracted according to the attribute representation on the feature map.
[0152] The process involves extracting voxel features from these grids. These features typically originate from a feature map, which may be learned from raw point cloud data using a deep learning model (such as a convolutional neural network). Each location on the feature map corresponds to a voxel grid and contains attribute representations of the voxels within that grid. Voxel features can then be extracted based on these attribute representations.
[0153] In this embodiment of the disclosure, the feature map is labeled with the attribute representations (such as color, height, density, etc.) of the voxels in the voxel grid. For each valid voxel grid, the corresponding attribute representation can be extracted from the feature map. These attribute representations constitute the voxel features of the voxels in that grid.
[0154] In S23, based on the voxel characteristics of the voxels in the effective voxel grid, the attribute characteristics of the voxels in each voxel grid of the voxel structure in the attribute representation are obtained.
[0155] In this embodiment, the voxel features extracted in the previous step are integrated to form the overall features of each voxel grid in the voxel structure in terms of attribute representation. This typically involves aggregation or statistical operations on voxel features to obtain representative features for each grid. Voxel features are extracted from each valid voxel grid, and these features are integrated to form the representative features of each grid. This can be achieved by performing statistical operations such as averaging, maximizing, and minimizing the voxel features of all voxels within the grid. Finally, the overall features of each voxel grid in the voxel structure in terms of attribute representation can be obtained.
[0156] Through the above steps, the attribute features of voxels in each voxel grid of the voxel structure can be extracted based on the feature map, providing strong support for subsequent point cloud data processing and scene understanding tasks.
[0157] Optionally, the voxel features include at least one of the following: voxel height, laser reflection intensity, origin distance with the starting point of the corresponding effective voxel grid as the origin, and origin angle with the starting point of the corresponding effective voxel grid as the origin. In this context, voxel height refers to the height information provided by the LiDAR scanning data for each point. During voxelization, the height of each point within the voxel grid is used to calculate the voxel's height characteristics. For example, voxel height can be the average or maximum height of all points within the grid.
[0158] The intensity of laser reflection is a key element in lidar's ability to sense the environment by emitting laser light and measuring the intensity of its reflected rays. Reflection intensity can provide information about the material of an object's surface; for example, a metallic surface may reflect a stronger laser beam, while vegetation may reflect a weaker one.
[0159] The origin can be the starting point or center point of each voxel grid. The origin distance is the distance from each point within the voxel to the origin. The origin angle can be the angle between the line connecting each point within the voxel to the origin and a reference axis (such as the X-axis). This feature can represent the directionality of points within the voxel.
[0160] In S23, obtaining the attribute features of voxels in each voxel grid of the voxel structure in the attribute representation based on the voxel features of the voxels in the effective voxel grid includes: Based on the voxel height of the voxel in the voxel features, determine the maximum height and average height of the voxels in each effective voxel grid; In this embodiment of the disclosure, the maximum height value and the average height can be found by traversing the voxel heights within each effective voxel grid.
[0161] Based on the reflection intensity of the voxels in the voxel features, the average radiation intensity and maximum reflection intensity of the voxels in each effective voxel grid are determined.
[0162] Similarly, the average and maximum reflection intensities of voxels within each effective voxel grid can be calculated by iterating through the reflection intensities of voxels to the laser within each effective voxel grid.
[0163] Based on the origin distance corresponding to the voxel in the voxel feature, determine the average origin distance of the voxels in each effective voxel grid; The average origin angle of the voxels in each effective voxel grid is determined based on the origin angle corresponding to the voxels in the voxel features.
[0164] Similarly, calculate the average distance from all points within each voxel grid to the origin. The average distance to the origin reflects the distribution of points within the voxel, helping to determine the shape and size of an object. Calculate the average angle between the lines connecting all points within each voxel grid to the origin. The average angle to the origin represents the directional distribution of points within the voxel, which is very useful for analyzing road direction, vehicle orientation, etc.
[0165] These steps allow for the extraction of meaningful attribute features from voxel structures, providing strong support for subsequent tasks such as autonomous driving decision-making and environmental perception.
[0166] This disclosure also provides a vehicle control method, see [link to relevant documentation] Figure 12 As shown, it includes: In S101, environmental images acquired by the vehicle during its driving process are obtained.
[0167] The vehicle can be a delivery vehicle used for order fulfillment, or a pure electric vehicle, a range-extended electric vehicle, or a gasoline-powered vehicle for passenger use. Therefore, for example, the driving image could be real-time image data captured by the delivery vehicle's onboard camera or other visual sensors while it is in motion. This data typically includes visual information about the vehicle's surrounding environment, such as roads, vehicles, pedestrians, and traffic signs.
[0168] For example, a delivery vehicle is making deliveries along a predetermined route on city streets. To ensure driving safety, the vehicle is equipped with advanced camera and sensor systems to capture images of the surrounding environment in real time. These images include information such as road conditions ahead, traffic signals, pedestrian and vehicle movements. For instance, as the delivery vehicle approaches an intersection, the camera captures the traffic conditions ahead, including the status of traffic lights, pedestrians crossing the road, and the position and speed of other vehicles at the intersection. This image data is acquired and processed in real time through the vehicle's internal systems.
[0169] In S102, the environmental image is identified using the image recognition method described in any of the foregoing embodiments to obtain an image recognition result.
[0170] In this embodiment, after acquiring the driving image, the delivery vehicle's computer system immediately performs image recognition using the image recognition method provided in the foregoing embodiments. This image recognition method can be applied to a model, such as a network model based on a multi-attention mechanism. This network model, trained and optimized with a large amount of data, can accurately identify key information in the image. It analyzes road signs, traffic signals, pedestrian dynamics, etc., to determine the safety status of the current driving environment.
[0171] For example, when captured images of an intersection are fed into a network model, the model can quickly identify that the traffic light is currently red, detect pedestrians crossing the road, and identify other vehicles waiting at the intersection. These identification results are extracted as key information.
[0172] In S103, the vehicle is controlled to move based on the image recognition result.
[0173] Driving control, in particular, involves adjusting the vehicle's driving status, including speed, direction, and braking, based on image recognition results through the delivery vehicle's control system. Driving control aims to ensure the vehicle drives safely and efficiently in complex environments while meeting the demands of delivery tasks.
[0174] For example, after obtaining the image recognition results, the delivery vehicle's control system makes corresponding decisions based on this information. If a dangerous situation is detected or an instruction to slow down or stop is needed, the control system immediately adjusts the vehicle's speed and direction to ensure driving safety. If the environment is safe and driving conditions are met, the vehicle continues to travel along the predetermined route. For instance, if the traffic light is detected to be red and pedestrians are crossing the road, the delivery vehicle's control system immediately makes a decision to slow down and prepare to stop. When the red light turns green and the pedestrians have crossed, the control system again determines the intersection is safe based on the image recognition results and then controls the vehicle to continue driving.
[0175] In this embodiment of the disclosure, in order to optimize the model structure and avoid modifying each plugin individually for each operator that does not support ONNX / TensorRT, it is necessary to migrate the pre-computation part of non-network weights to the data preprocessing process, and package the parameter processing required for the dynamic sparse partitioning part as a whole into CUDA kernel functions.
[0176] The above technical solution, based on the image recognition method provided in the foregoing embodiments, can quickly and accurately identify images collected by the delivery vehicle during its operation, thereby improving image recognition efficiency and accuracy, and thus enhancing the safety of the delivery vehicle.
[0177] This disclosure also provides an electronic device, including: A memory on which computer programs are stored; A processor for executing the computer program in the memory to implement the steps of any of the methods described in the foregoing embodiments.
[0178] Figure 13 This is a block diagram illustrating an electronic device 700 according to an exemplary embodiment. Figure 13 As shown, the electronic device 700 may include a processor 701 and a memory 702. The electronic device 700 may also include one or more of a multimedia component 703, an input / output (I / O) interface 704, and a communication component 705.
[0179] The processor 701 controls the overall operation of the electronic device 700 to complete all or part of the steps in the image recognition method or delivery vehicle control method described above. The memory 702 stores various types of data to support the operation of the electronic device 700. This data may include, for example, instructions for any application or method operating on the electronic device 700, and application-related data such as contact data, sent and received messages, pictures, audio, video, etc. The memory 702 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as Static Random Access Memory (SRAM), Electrically Erasable Programmable Read-Only Memory (EEPROM), Erasable Programmable Read-Only Memory (EPROM), Programmable Read-Only Memory (PROM), Read-Only Memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk. The multimedia component 703 may include a screen and audio components. The screen may be, for example, a touchscreen, and the audio component is used to output and / or input audio signals. For example, the audio component may include a microphone for receiving external audio signals. The received audio signals may be further stored in memory 702 or transmitted via communication component 705. The audio component also includes at least one speaker for outputting audio signals. I / O interface 704 provides an interface between processor 701 and other interface modules, such as a keyboard, mouse, buttons, etc. These buttons may be virtual or physical buttons. Communication component 705 is used for wired or wireless communication between the electronic device 700 and other devices. Wireless communication, such as Wi-Fi, Bluetooth, Near Field Communication (NFC), 2G, 3G, 4G, NB-IoT, eMTC, or other 5G technologies, or combinations thereof, is not limited here. Therefore, the corresponding communication component 705 may include: a Wi-Fi module, a Bluetooth module, an NFC module, etc.
[0180] In an exemplary embodiment, the electronic device 700 may be implemented by one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field-programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components to perform the image recognition method or delivery vehicle control method described above.
[0181] In another exemplary embodiment, a computer-readable storage medium including program instructions is also provided, which, when executed by a processor, implement the steps of the image recognition method or delivery vehicle control method described above. For example, the computer-readable storage medium may be the memory 702 including the program instructions described above, which may be executed by the processor 701 of the electronic device 700 to complete the image recognition method or delivery vehicle control method described above.
[0182] The preferred embodiments of this disclosure have been described in detail above with reference to the accompanying drawings. However, this disclosure is not limited to the specific details of the above embodiments. Within the scope of the technical concept of this disclosure, various simple modifications can be made to the technical solutions of this disclosure, and these simple modifications all fall within the protection scope of this disclosure.
[0183] It should also be noted that the various specific technical features described in the above specific embodiments can be combined in any suitable manner without contradiction. In order to avoid unnecessary repetition, this disclosure will not describe the various possible combinations separately.
[0184] Furthermore, various different embodiments of this disclosure can be combined in any way, as long as they do not violate the spirit of this disclosure, they should also be regarded as the content disclosed in this disclosure.
Claims
1. An image recognition method, characterized in that, include: Point clouds from multiple frames of images are fused into voxel space to construct a voxel structure on a single-frame representation; Based on the attribute representation on the feature map, extract the attribute features of voxels in each voxel grid of the voxel structure on the attribute representation; Based on the voxel coordinates and voxel indices of different voxel types in the voxel structure, the feature masks corresponding to different coordinate axes of the voxels under the window structure are determined. The different coordinate axes are determined based on the voxels in the bird's-eye view. The voxels with different indices in the feature mask are used to determine the point clouds corresponding to the images in multiple frames as the target point clouds. Based on the attribute features corresponding to the target point cloud, image recognition is performed on multiple frames of the image to obtain the image recognition result.
2. The method according to claim 1, characterized in that, The step of determining the feature mask corresponding to different coordinate axes of the voxels under the window structure based on the voxel coordinates and voxel indices of different voxel types in the voxel structure includes: Based on the sorted voxel index of each voxel in the voxel structure, the number of windows in the window structure, and the number of first voxels in the window structure, determine the unique inverse index of each voxel in the window structure; Based on the unique inverse index of each voxel in each window structure, the window set partitioning parameter, and the voxels in each window, the window voxel index in the bird's-eye view after the voxels in each window are partitioned into different sets is determined. The window set partitioning parameter is used to determine the number of voxels in each set. Based on the unique inverse index of each voxel, the voxel coordinates of each window in the window structure of different types in the voxel structure, and the voxel index, multiple different index sorting sequences are determined; Based on the unique inverse index of each voxel, the window voxel index in the corresponding window, and multiple different index sorting sequences, the different window global indexes in the window are determined; Based on the index movement step size, index movement direction, coordinate axes under the bird's-eye view, and different global indexes of the window, determine the feature mask corresponding to the coordinate axis.
3. The method according to claim 2, characterized in that, The step of determining the window voxel index in the bird's-eye view after the voxels in each window are divided into different sets, based on the unique inverse index of each voxel in each window structure, the window set partitioning parameter, and the voxels in each window, includes: The voxels in each window are transformed by perspective to obtain the bird's-eye view voxels of each voxel in the bird's-eye view. The number of sets in each window is determined based on the window set partitioning parameters and the number of second voxels in each window. Based on the number of sets and the window set partitioning parameters, the voxels in each window are divided into sets, and set numbers are added to the sets in each window in sequence. Based on the number of sets corresponding to each window, calculate the cumulative sum of the set numbers in each window to obtain the cumulative sum sequence corresponding to each window; Add the same marker value to each of the accumulated sum sequences corresponding to each window to obtain a marker value sequence corresponding to each window; The sum of the marker values in the sequence of marker values corresponding to each window is used as the window index corresponding to that window. Based on the cumulative sum sequence corresponding to each window and the window index, determine the window set index of each set in each window; Based on the number of sets in each window, the corresponding window index, and the window set index of each set in each window, determine the bird's-eye view index of each bird's-eye view voxel in the bird's-eye view. Based on the unique inverse index of each voxel in each window structure and the sorted voxel index of each voxel in each window in the voxel structure, the bird's-eye index of each bird's-eye voxel in each window is converted into the window voxel index of the bird's-eye voxel in the corresponding window.
4. The method according to claim 2, characterized in that, The step of determining multiple different index sorting sequences based on the unique inverse index of each voxel, the voxel coordinates of each window in the window structure of different types in the voxel structure, and the voxel index, includes: Based on the voxel coordinates and voxel indexes of the voxels of each window in the voxel structure for different types in the voxel structure, determine the window voxel coordinates of each voxel in the corresponding window corresponding to each type. Based on the window voxel coordinates of each voxel in the corresponding window and the window voxel coordinates of each type, construct an array of window voxel coordinates of each voxel in the window and the window voxel coordinates of each type. Based on the unique inverse index of each voxel and the window voxel coordinate array corresponding to each type, construct the reference index of each voxel and its type in the window structure; Based on the reference indices corresponding to each voxel and different types, different index sorting sequences are obtained by sorting with the reference indices of each type as the sorting target.
5. The method according to claim 2, characterized in that, The step of determining the feature mask corresponding to the coordinate axis based on the index movement step size, index movement direction, the coordinate axis under the bird's-eye view, and different window global indices includes: Based on the coordinate axes in the bird's-eye view, assign coordinate axis set indices corresponding to the coordinate axes to different window global indices; Based on the index movement step size and the index movement direction, each of the coordinate axis set indices is moved to obtain the corresponding movement set index; Compare the coordinate axis set index with the corresponding movement set index; Based on whether the indices at each position are the same in the index comparison results, the feature mask corresponding to the coordinate axis is determined.
6. The method according to any one of claims 1-5, characterized in that, The process of fusing point clouds from multiple frames of images into a voxel space and constructing a voxel structure on a single-frame representation includes: For each point cloud in the multiple frames of the image, a global index of the point cloud is constructed. Each point cloud is mapped onto a voxel grid in voxel space; Based on the global index of each point cloud, determine the unique voxel index after sorting the voxels in each voxel grid; Determine the starting position of each voxel in the global index of the point cloud after sorting; The voxel structure is constructed on a single-frame representation based on the voxel number of voxels in the voxel grid, the corresponding starting position, and the unique voxel index.
7. The method according to claim 6, characterized in that, The step of determining the unique voxel index after sorting the voxels in each voxel grid based on the global index of each point cloud includes: Using any coordinate axis of the voxel space as the first target, determine the first quotient and the first remainder of the number of voxel grids of the global index of each point cloud on that coordinate axis. Using any remaining coordinate axis of the voxel space as the second target, determine the second quotient and second remainder of the first quotient corresponding to each point cloud on the value of the number of voxel grids on that coordinate axis; Using the remaining coordinate axes of the voxel space as the third target, determine the third remainder of the second quotient corresponding to each point cloud on the value of the number of voxel grids on that coordinate axis; Based on the first remainder, the second remainder, and the third remainder, determine the voxel index corresponding to each voxel in the voxel grid; Based on the original voxel index corresponding to each voxel and the point cloud global index, construct the index sequence of the voxel; Based on the index at any same position in each of the index sequences, the voxels in each voxel grid are sorted to obtain the sorted voxel index in each voxel grid. Duplicate voxel indices are removed from each voxel grid after sorting, resulting in a unique voxel index corresponding to each voxel in each voxel grid.
8. The method according to any one of claims 1-5, characterized in that, The step of extracting attribute features of voxels in each voxel grid of the voxel structure based on the attribute representation on the feature map includes: Determine the effective voxel grid containing voxels within the voxel structure; For the effective voxel grid, the voxel features of the voxels in the effective voxel grid are extracted based on the attribute representation on the feature map; Based on the voxel characteristics of the voxels in the effective voxel grid, the attribute characteristics of the voxels in each voxel grid of the voxel structure in the attribute representation are obtained; and / or, The voxel features include at least one of the following: voxel height, laser reflection intensity, origin distance with the starting point of the corresponding effective voxel grid as the origin, and origin angle with the starting point of the corresponding effective voxel grid as the origin. The step of obtaining the attribute features of voxels in each voxel grid of the voxel structure in the attribute representation based on the voxel features of the voxels in the effective voxel grid includes: Based on the voxel height of the voxel in the voxel features, determine the maximum height and average height of the voxels in each effective voxel grid; Based on the reflection intensity of the voxels in the voxel features, determine the average radiation intensity and maximum reflection intensity of the voxels in each effective voxel grid; Based on the origin distance corresponding to the voxel in the voxel feature, determine the average origin distance of the voxels in each effective voxel grid; The average origin angle of the voxels in each effective voxel grid is determined based on the origin angle corresponding to the voxels in the voxel features.
9. A vehicle control method, characterized in that, include: Acquire environmental images during vehicle operation; The environmental image is identified using the image recognition method according to any one of claims 1-8 to obtain an image recognition result; The vehicle is controlled to move based on the image recognition results.
10. An electronic device, characterized in that, include: A memory on which computer programs are stored; A processor for executing the computer program in the memory to implement the steps of the method according to any one of claims 1-8.
Citation Information
Patent Citations
Target detection method and device, electronic equipment and computer readable storage medium
CN115170769A
Three-dimensional target detection method and system based on point cloud-image multi-cross mixing and storage medium
CN116664856A
Data processing method and device, computer equipment and storage medium
CN117011820A
Point cloud three-dimensional target detection method based on voxel context perception
CN117671360A
Target detection method and related device
WO2022178895A1