Multi-agent collaborative semantic grid prediction method and system based on uncertainty
By adopting a multi-agent collaboration method based on uncertainty in semantic occupancy raster prediction, using three-dimensional features and uncertainty distribution maps for information fusion, the problems of depth prediction difficulties and perspective limitations in semantic occupancy raster prediction are solved, and more accurate and comprehensive prediction results are achieved.
Patent Information
- Application Number
- CN202510098833.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-22
- Publication Date
- 2025-05-16
AI Technical Summary
The prior art has difficulties in semantic occupancy raster prediction, insufficient geometric accuracy, and inevitable viewing angle limitations and occlusion problems of single-view angles, resulting in blurred or errors in the boundaries of the prediction results.
The multi-agent collaborative semantic raster prediction method based on uncertainty is adopted. By extracting three-dimensional features, generating geometric and semantic uncertainty distribution maps, building message packages and performing information fusion features, decoding the fusion features to obtain the cooperative occupation raster prediction results.
Improve the accuracy and comprehensiveness of semantic occupancy raster prediction, reduce the adverse effects of inaccurate depth estimation on prediction, explicitly process image boundaries and make full use of the advantages of multi-view observation.
Smart Images

Figure CN120014626A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of image processing technology, and in particular to an uncertainty-based multi-agent collaborative semantic grid prediction method and system. Background Art
[0002] The emergence of multi-agent collaborative perception technology enables multiple agents to share and utilize each other's perception capabilities, thereby achieving a more comprehensive and complementary perception effect. This technology provides a new solution to the inevitable limitations of single-agent perception, such as occlusion problems and long-distance perception problems. Related methods of multi-agent collaborative perception have broad application prospects in many fields, such as multi-vehicle network autonomous driving systems, multi-robot warehouse automation systems, and multi-drone search and rescue missions.
[0003] Semantic occupancy grid prediction is a relatively new task in the field of single-agent perception. The goal of this task is to classify each voxel point in space as unoccupied or belonging to a specific category. Compared with the common object detection task in the traditional multi-agent collaborative perception field, semantic occupancy grid prediction provides a more sophisticated and detailed representation of the entire environment, thereby providing more sufficient and detailed information for decision-making and planning. However, compared with the object detection task, semantic occupancy grid prediction has two significant characteristics: first, the detection results of semantic occupancy grid prediction are dense, while the results of object detection tasks are usually highly sparse; second, object detection tasks usually require post-processing of the results output by deep neural networks, while semantic occupancy grid prediction directly outputs the perception results from deep neural networks. Due to these two objective problems, many collaborative methods originally applied to object detection tasks cannot be directly applied to semantic occupancy grid prediction.
[0004] In order to expand the collaborative perception technology to more complex and comprehensive perception tasks, a multi-agent collaborative semantic occupancy grid prediction method is proposed. Generally, both LiDAR point cloud and two-dimensional image can be used as input for semantic occupancy grid prediction. Among them, the two-dimensional image as input contains rich semantic information, which makes it advantageous in some aspects, but its depth prediction is more difficult, resulting in inferior geometric accuracy to LiDAR point cloud input. Among them, the LiDAR as input contains complete three-dimensional information and accurate geometric information, but the semantic certainty of the extracted features is insufficient. In addition, single-view inevitably has the problem of perspective limitations and occlusion.
[0005] It is worth noting that the currently commonly used pure visual perception method based on depth estimation requires projecting the features in the two-dimensional image plane to several different positions in the viewing cone in proportion to the probability of depth prediction. This method will cause corresponding features to appear in the unoccupied areas near the occupied areas, which will make the occupied grid instance category boundaries finally decoded blurred or even wrong. However, in previous object detection tasks, post-processing, especially the non-maximum suppression process, can solve the problem of erroneous detection results near the correct detection target to a certain extent. In order to explicitly handle instance boundaries and take full advantage of the additional advantages brought by multi-view observation, collaborative semantic occupancy grid prediction is not suitable for sharing feature maps that have been fully processed by the backbone network. Summary of the invention
[0006] In view of the deficiencies in the prior art, the purpose of the present invention is to provide a multi-agent collaborative semantic grid prediction method and system based on uncertainty.
[0007] According to one aspect of the present invention, a multi-agent collaborative semantic grid prediction method based on uncertainty is provided, comprising:
[0008] Extract the three-dimensional features of the input data and predict whether each spatial position is occupied and the category of the object occupying the spatial position as the occupancy grid prediction result;
[0009] Based on the input data, generating a geometric uncertainty distribution map;
[0010] generating a semantic uncertainty distribution map based on the three-dimensional features and the occupancy grid prediction result;
[0011] Constructing a message package based on the three-dimensional feature, the geometric uncertainty distribution map and the semantic uncertainty distribution map;
[0012] Fusing message packets from multiple agents and the three-dimensional features extracted by their own perception to obtain fused features;
[0013] The fused features are decoded to obtain a collaborative occupancy grid prediction result.
[0014] Preferably, extracting the three-dimensional features of the input data and predicting whether each spatial position is occupied and the category of the object occupying the spatial position as an occupancy grid prediction result includes:
[0015] With a coding network Φ enc Input collected from the ith agent Extract sparse features from:
[0016]
[0017] in, C s Indicates the number of channels of sparse features, H, W, and Z are the length, width, and height of the feature map respectively; for lidar point cloud input, For image input,
[0018] Using a backbone network Φ backbone Process sparse features into dense features:
[0019]
[0020] Among them, the above table o indicates that the feature only contains the self-agent perception information;
[0021] Using two different decoding networks Φ occ and Φ sem Decode the semantic occupancy grid prediction results from the dense features respectively and
[0022]
[0023] Preferably, generating a geometric uncertainty distribution map based on the input data comprises:
[0024] For image input, during the encoding and decoding process of image input, there is a projection process from the image plane to the space, and the projection process depends on the depth estimation:
[0025]
[0026] in
[0027] Using a geometric uncertainty generator Φ geounc , the depth estimation uncertainty of the position is generated from the original image input and the depth estimation and projected to all positions of the viewing cone to obtain the geometric uncertainty U geo :
[0028]
[0029] For the LiDAR point cloud input, the Sigmoid function of the number of laser point clouds contained in the observation area is defined to obtain the geometric uncertainty.
[0030] Preferably, generating a semantic uncertainty distribution map based on the three-dimensional features and the occupancy grid prediction result includes:
[0031] Using a semantic uncertainty generator Φ semunc, generate semantic uncertainty U from the occupancy grid prediction result and the dense feature map sem :
[0032]
[0033] Preferably, the message package is constructed based on the three-dimensional feature, the geometric uncertainty distribution map and the semantic uncertainty distribution map.
[0034]
[0035] The superscript s indicates that the feature map is information sent by the agent for collaboration.
[0036] Preferably, the step of fusing the message packets from multiple agents and the three-dimensional features extracted by the agents' own perception to obtain the fused features includes:
[0037] Where 1≤j≤N,j≠i
[0038]
[0039] The superscript C indicates the feature that integrates collaborative information, concat indicates the fusion module, and Φ fusion represents the fusion function related to the spatial distribution, and N represents the number of agents. represents the fusion feature, Indicates a fusion message packet.
[0040] Preferably, decoding the fused features to obtain a collaborative occupancy grid prediction result includes:
[0041] The sparse features are subjected to the process described in claim 2 to obtain the collaborative semantic occupancy grid prediction result. and
[0042] According to a second aspect of the present invention, there is provided a multi-agent collaborative semantic grid prediction system based on uncertainty, comprising:
[0043] Single-machine grid prediction module: extracts the three-dimensional features of the input data and predicts whether each spatial position is occupied and the category of the object occupying the spatial position as the occupancy grid prediction result;
[0044] Geometric uncertainty module: generating a geometric uncertainty distribution map based on the input data;
[0045] Semantic uncertainty module: generating a semantic uncertainty distribution map based on the three-dimensional features and the occupancy grid prediction result;
[0046] Message package module: constructing a message package based on the three-dimensional features, the geometric uncertainty distribution map and the semantic uncertainty distribution map;
[0047] Fusion module: fuses the message packets from multiple agents and the three-dimensional features extracted by their own perception to obtain fusion features;
[0048] Collaborative grid prediction module: decodes the fusion features to obtain collaborative occupancy grid prediction results.
[0049] According to a third aspect of the present invention, there is provided a terminal, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor can be used to execute any one of the methods described, or to run the system described, when executing the program.
[0050] According to a fourth aspect of the present invention, there is provided a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, can be used to execute any one of the methods described, or to execute the system described.
[0051] Compared with the prior art, the embodiments of the present invention have at least one of the following beneficial effects:
[0052] An embodiment of the present invention relates to a method and system for multi-agent collaborative semantic grid prediction based on uncertainty. On the one hand, through uncertainty-based information integration and fusion, it is possible to fully utilize the perception information from different sensors and different agents, thereby improving the benefits brought by collaboration; on the other hand, a fusion strategy that depends on spatial relationships is adopted to fuse sparse feature maps, effectively reducing the adverse effects of inaccurate depth estimation on semantic occupancy grid prediction.
[0053] An embodiment of the present invention relates to an uncertainty-based multi-agent collaborative semantic grid prediction method and system, which helps to obtain more accurate depth estimation by observing the same spatial area from different perspectives, thereby further improving the gain of collaboration among multiple agents.
[0054] An embodiment of the present invention relates to an uncertainty-based multi-agent collaborative semantic grid prediction method and system, in which the features used for multi-view fusion are features that have not been processed by the backbone network, and will be processed by the backbone network after fusion, effectively solving the problem in the prior art that collaborative semantic occupancy grid prediction is not suitable for sharing feature maps that have been fully processed by the backbone network, explicitly processing image boundaries and making full use of the additional advantages brought by multi-view observation, solving the problem of erroneous detection results occurring near the correct detection target. BRIEF DESCRIPTION OF THE DRAWINGS
[0055] Other features, objects and advantages of the present invention will become more apparent from the detailed description of non-limiting embodiments made with reference to the following drawings:
[0056] Figure 1 It is a framework flow chart of the uncertainty-based multi-agent collaborative semantic grid prediction method in one embodiment of the present invention. DETAILED DESCRIPTION
[0057] The present invention is described in detail below in conjunction with specific embodiments. The following embodiments will help those skilled in the art to further understand the present invention, but are not intended to limit the present invention in any form. It should be noted that, for those of ordinary skill in the art, several variations and improvements may be made without departing from the concept of the present invention. These all belong to the protection scope of the present invention.
[0058] In one embodiment of the present invention, a multi-agent collaborative semantic grid prediction method based on uncertainty is provided. Figure 1 As shown, it mainly includes the following steps:
[0059] S100, extracting three-dimensional features of input data, and predicting whether each spatial position is occupied and the category of the object occupying the spatial position as an occupancy grid prediction result;
[0060] S200, generating a geometric uncertainty distribution map based on the input data;
[0061] S300, generating a semantic uncertainty distribution map based on the three-dimensional features and occupancy grid prediction results of S100;
[0062] S400, constructing a message package based on the three-dimensional features of S100, the geometric uncertainty distribution map of S200 and the semantic uncertainty distribution map of S300;
[0063] S500, fusing the message packets from multiple intelligent agents and the three-dimensional features extracted by their own perception to obtain fused features;
[0064] S600, decoding the fused features of S500 to obtain a collaborative occupancy grid prediction result.
[0065] In the aforementioned embodiment, each agent provides uncertainty assessments of different locations in both geometric and semantic dimensions while making semantic occupancy grid predictions. On the one hand, it ensures that the most critical information can be screened out for transmission under conditions of limited communication, thereby providing high-value shared information; on the other hand, it provides sufficient perceptual quality information for multi-agent information fusion. By integrating multi-perspective observations of different agents, richer spatial features can be constructed to make up for their own perceptual disadvantages in two dimensions, thereby achieving efficient multi-agent collaborative semantic occupancy grid predictions.
[0066] In order to obtain semantic grid prediction information more accurately, in a preferred embodiment of the present invention, an improved solution of step S100 is proposed, that is, a deep neural network is used to extract the features of the input data, and predict whether each spatial position is occupied and the object category occupying the spatial position. The specific operation process is as follows:
[0067] With a coding network Φ enc Input collected from the ith agent Extract sparse features from Among them C s Indicates the number of channels of sparse features, H, W, Z are the length, width and height of the feature map respectively. For LiDAR point cloud input (after voxelization), For image input,
[0068]
[0069] Using a backbone network Φ backbone The sparse features are processed into dense features, and the superscript o indicates that the feature only contains the self-agent perception information:
[0070]
[0071] Using two different decoding networks Φ occ and Φ sem Decode semantic occupancy grid predictions from dense spatial features and
[0072]
[0073] Indicates whether the grid corresponding to each feature contains an object; Indicates the object category contained in the grid corresponding to each feature.
[0074] It should be noted that in the above embodiments The features used for multi-view fusion are features that have not been processed by the backbone network, and they will be processed by the backbone network again after fusion. This effectively solves the problem in the prior art that collaborative semantic occupancy grid prediction is not suitable for sharing feature maps that have been fully processed by the backbone network, explicitly handles image boundaries and fully utilizes the additional advantages brought by multi-view observation to solve the problem of erroneous detection results near the correct detection target.
[0075] The geometric uncertainty can indicate the quality of predicting whether a specific space is occupied. In a preferred embodiment of the present invention, a preferred solution of step S200 is provided, that is, a geometric uncertainty generator is used to generate a geometric uncertainty distribution map. Specifically, there are two cases:
[0076] In the first case, images are used as input.
[0077] In the decoding process based on image input, there is a projection process from the image plane to space, which depends on the estimation of depth, where
[0078]
[0079] Using a geometric uncertainty generator Φ geounc , the depth estimation uncertainty of the position is generated from the original input and the depth estimation and projected to all positions of the frustum to obtain the geometric uncertainty U geo :
[0080]
[0081] In the second case, the lidar point cloud is used as input.
[0082] For an agent that uses a LiDAR point cloud as input, the geometric uncertainty is obtained by defining the Sigmoid function of the number of laser point clouds contained in the observation area:
[0083]
[0084] Generally, the fewer the number of laser point clouds collected in the area, the higher the geometric uncertainty. Conversely, the more points, the lower the geometric uncertainty.
[0085] The geometric uncertainty distribution diagram in the above embodiment reflects the geometric uncertainty of the features in each spatial region. The smaller the value, the smaller the uncertainty, that is, the greater the possibility that the feature is located in the spatial region. This uncertainty feature map can effectively guide multi-perspective information fusion. On the receiving side, for multiple information packets from multiple perspectives, the features with low uncertainty are selected and used. At the same time, the uncertainty maps of multiple perspectives are cross-checked based on the principle of multi-perspective geometry, which can effectively eliminate the feature uncertainty of each perspective and determine the correct spatial position of each feature, thereby obtaining more accurate and comprehensive features, that is, obtaining accurate features in each spatial region.
[0086] Semantic uncertainty can reflect the accuracy of the prediction of the object category in a specific space. In order to maximize the advantages of different agents in the two dimensions of geometry and semantics, in another preferred embodiment of the present invention, a preferred solution of step S300 is proposed, that is, a semantic uncertainty generator is used to construct a semantic uncertainty distribution map. Specifically:
[0087] Using a semantic uncertainty generator Φ semunc , generate semantic uncertainty U from the semantic occupancy grid prediction results and dense feature maps sem :
[0088]
[0089] In the above embodiment, the semantic uncertainty distribution map reflects the semantic uncertainty of the features in each spatial region. The smaller the value, the smaller the uncertainty, that is, the greater the possibility that the feature belongs to the category. This uncertainty feature map can effectively guide multi-perspective information fusion. On the receiving side, for multiple information packets from multiple perspectives, the features with low uncertainty are selected and used. At the same time, the uncertainty maps of multiple perspectives are cross-checked based on the principle of multi-perspective geometry, which can effectively eliminate the feature uncertainty of each perspective and determine the accurate semantic features in each spatial region, that is, the semantic category of each spatial region.
[0090] In order to provide high-value shared information under the condition of limited communication resources, it is necessary to output and share the geometric uncertainty and semantic uncertainty as a whole. In a preferred embodiment of the present invention, a preferred solution of step S400 is provided, that is, using a collaborative information packaging module based on uncertainty to construct a high-value message package. Specific:
[0091] Message packets based on sparse feature graphs It is constructed by geometric and semantic uncertainty graphs; the superscript s indicates that the feature graph is the information sent by the agent for collaboration.
[0092]
[0093] It can be seen from the above embodiments that the information packet contains three parts: i) the first part is sparse but effective feature information within a single perspective. The effective information contained in this part can supplement the missing parts in other perspectives. The feature is sparse and has the advantage of small communication volume, thereby achieving an effective trade-off between communication volume consumption and collaboration effect. ii) The corresponding geometric uncertainty map can distinguish the accuracy of the geometric information in the features of the first part, thereby guiding the receiving side of the information packet to reasonably use the accepted features and select features with high certainty in the information packets from multiple perspectives on the receiving side; iii) The corresponding semantic uncertainty map can distinguish the accuracy of the semantic information in the features of the first part, thereby guiding the receiving side of the information packet to reasonably use the accepted features and select features with high certainty in the information packets from multiple perspectives on the receiving side.
[0094] Observations from multiple perspectives of different agents can integrate enhanced spatial features, thereby achieving efficient multi-agent collaborative semantic occupancy grid prediction. Therefore, in a preferred embodiment of the present invention, a preferred solution for step S500 is provided, that is, using a spatial feature fusion module based on geometric uncertainty to fuse message packets from multiple agents and data extracted by their own perception, so as to improve the effect of spatial occupancy prediction. For the i-th agent among the N collaborative agents:
[0095] Where 1≤j≤N,j≠i
[0096]
[0097] The superscript C represents the feature that integrates collaborative information, Φ fusion is the fusion function related to the spatial distribution.
[0098] In order to make full use of the advantages of different agents in the two dimensions of geometry and semantics, the above embodiment constructs two different degrees of certainty to quantify the perception effect of the agent, and uses these quantitative indicators to guide selective information fusion, thereby avoiding the interference of low-quality perception information and achieving better collaborative performance. In addition, a single perspective inevitably has perspective limitations and occlusion problems. The above embodiment fuses and aggregates information from multiple perspectives through multi-perspective fusion, and the occluded area in a single perspective may be supplemented in other perspectives, thereby obtaining more comprehensive and complete environmental information.
[0099] Based on the same inventive concept, in another embodiment of the present invention, a multi-agent collaborative semantic grid prediction system based on uncertainty is provided, comprising:
[0100] Single-machine grid prediction module: extracts the three-dimensional features of the input data and predicts whether each spatial position is occupied and the category of the object occupying the spatial position as the occupancy grid prediction result;
[0101] Geometric uncertainty module: Generates geometric uncertainty distribution map based on uncertainty-based multi-agent collaborative semantic grid prediction input data;
[0102] Semantic uncertainty module: Generates a semantic uncertainty distribution map based on uncertainty-based multi-agent collaborative semantic grid prediction three-dimensional features and uncertainty-based multi-agent collaborative semantic grid prediction occupancy grid prediction results;
[0103] Message package module: Construct message packages based on uncertainty-based multi-agent collaborative semantic grid prediction three-dimensional features, uncertainty-based multi-agent collaborative semantic grid prediction geometric degree distribution map, and uncertainty-based multi-agent collaborative semantic grid prediction semantic uncertainty distribution map;
[0104] Fusion module: fuses the message packets from multiple agents and the three-dimensional features extracted by their own perception to obtain fusion features;
[0105] Collaborative grid prediction module: decodes the uncertainty-based multi-agent collaborative semantic grid prediction fusion features to obtain collaborative occupancy grid prediction results.
[0106] The various modules / units in the above examples of the present invention may specifically refer to the implementation techniques of the corresponding steps of the uncertainty-based multi-agent collaborative semantic grid prediction method in the above embodiments, which will not be repeated here.
[0107] In order to obtain sufficient spatial and semantic information required for the grid prediction task, the above-mentioned embodiment utilizes the information of various types of sensors from multiple perspectives for complementation, and proposes an uncertainty-based multi-agent collaborative semantic occupancy grid prediction system. The geometric uncertainty and semantic uncertainty of each feature map are effectively reflected by introducing a geometric uncertainty map and a semantic uncertainty map, and they are used to guide multi-perspective information fusion. On the receiving side, for multiple information packets from multiple perspectives, the features with low uncertainty are selected and utilized. At the same time, the uncertainty maps of multiple perspectives are cross-checked based on the principle of multi-perspective geometry, which can effectively eliminate the feature uncertainty of each perspective, thereby obtaining more accurate and comprehensive features, that is, determining the correct geometric information and semantic information in each spatial area.
[0108] Based on the same inventive concept, in other embodiments of the present invention, a terminal is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor can be used to execute the above-mentioned method or to run the above-mentioned system when executing the program.
[0109] Optionally, the memory is used to store programs; the memory may include volatile memory (English: volatile memory), such as random-access memory (English: random-access memory, abbreviated: RAM), such as static random-access memory (English: static random-access memory, abbreviated: SRAM), double data rate synchronous dynamic random access memory (English: Double Data Rate Synchronous Dynamic RandomAccess Memory, abbreviated: DDRSDRAM), etc.; the memory may also include non-volatile memory (English: non-volatile memory), such as flash memory (English: flash memory). The memory is used to store computer programs (such as applications, functional modules, etc. that implement the above method), computer instructions, etc., and the above-mentioned computer programs, computer instructions, etc. can be partitioned and stored in one or more memories.
[0110] The processor is used to execute the computer program stored in the memory to implement the various steps in the method involved in the above embodiment. For details, please refer to the relevant description in the above method embodiment.
[0111] The processor and the memory may be independent structures or integrated structures. When the processor and the memory are independent structures, the memory and the processor may be coupled and connected via a bus.
[0112] Based on the same inventive concept, in other embodiments of the present invention, a computer-readable storage medium is provided, on which a computer program is stored. When the program is executed by a processor, it can be used to execute the above method, or run the above system.
[0113] Among them, computer-readable media include computer storage media and communication media, wherein the communication media include any media that facilitates the transmission of computer programs from one place to another. The storage medium can be any available medium that can be accessed by a general or special-purpose computer. An exemplary storage medium is coupled to the processor so that the processor can read information from the storage medium and write information to the storage medium. Of course, the storage medium can also be an integral part of the processor. The processor and the storage medium can be located in an ASIC. In addition, the ASIC can be located in a user device. Of course, the processor and the storage medium can also be present in a communication device as discrete components.
[0114] Those skilled in the art will appreciate that embodiments of the present invention may be provided as methods, systems, or computer program products. Therefore, the present invention may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Moreover, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0115] The present invention is described with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0116] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.
[0117] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.
[0118] Obviously, those skilled in the art can make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the claims of the present invention and their equivalents, the present invention is also intended to include these modifications and variations.
Claims
1. A multi-agent collaborative semantic grid prediction method based on uncertainty, characterized in that: include: Extract the three-dimensional features of the input data and predict whether each spatial position is occupied and the category of the object occupying the spatial position as the occupancy grid prediction result; Based on the input data, generating a geometric uncertainty distribution map; generating a semantic uncertainty distribution map based on the three-dimensional features and the occupancy grid prediction result; Constructing a message package based on the three-dimensional feature, the geometric uncertainty distribution map and the semantic uncertainty distribution map; Fusing the message packets from multiple agents and the three-dimensional features extracted by their own perception to obtain fused features; The fused features are decoded to obtain a collaborative occupancy grid prediction result.
2. The uncertainty-based multi-agent collaborative semantic grid prediction method according to claim 1 is characterized in that: The extracting of the three-dimensional features of the input data and predicting whether each spatial position is occupied and the category of the object occupying the spatial position as an occupancy grid prediction result includes: With a coding network Φ enc Input x collected from the ith agent i Extract sparse features from: in, C s Indicates the number of channels of sparse features, H, W, and Z are the length, width, and height of the feature map respectively; for lidar point cloud input, For image input, Using a backbone network Φ backbone Process the sparse features into dense features: Among them, the superscript o indicates that the feature only contains the self-agent perception information; Using two different decoding networks Φ occ and Φ sem Decode the semantic occupancy grid prediction results from the dense features respectively and Indicates whether the grid corresponding to each feature contains an object; Indicates the object category contained in the grid corresponding to each feature.
3. The uncertainty-based multi-agent collaborative semantic grid prediction method according to claim 2 is characterized in that: The step of generating a geometric uncertainty distribution diagram based on the input data comprises: For image input, during the encoding and decoding process of image input, there is a projection process from the image plane to the space, and the projection process depends on depth estimation: F depth,i =Φ depth (x i ) in represents depth estimation; Using a geometric uncertainty generator Φ geounc , the depth estimation uncertainty of the position is generated from the original image input and the depth estimation and projected to all positions of the viewing cone to obtain the geometric uncertainty U geo : For the LiDAR point cloud input, the Sigmoid function of the number of laser point clouds contained in the observation area is defined to obtain the geometric uncertainty.
4. The uncertainty-based multi-agent collaborative semantic grid prediction method according to claim 3 is characterized in that: The generating a semantic uncertainty distribution map based on the three-dimensional features and the occupancy grid prediction result includes: Using a semantic uncertainty generator Φ semunc , generate semantic uncertainty U from the occupancy grid prediction result and the dense feature map sem :
5. The uncertainty-based multi-agent collaborative semantic grid prediction method according to claim 4 is characterized in that: The message packet is constructed based on the three-dimensional feature, the geometric uncertainty distribution map and the semantic uncertainty distribution map. The superscript s indicates that the feature map is information sent by the agent for collaboration.
6. The uncertainty-based multi-agent collaborative semantic grid prediction method according to claim 5, characterized in that: The method of fusing the message packets from multiple agents and the three-dimensional features extracted by the agents' own perception to obtain the fused features includes: Where 1≤j≤N, j≠i The superscript C indicates the feature that integrates collaborative information, concat indicates the fusion module, and Φ fusion represents the fusion function related to the spatial distribution, N represents the number of agents, represents the fusion feature, Indicates a fusion message packet.
7. The uncertainty-based multi-agent collaborative semantic grid prediction method according to claim 6 is characterized in that: The decoding of the fused features to obtain a collaborative occupancy grid prediction result includes: The fusion feature F C sparse,i Processed as dense fusion features, F C dense,i =Φ backbone (F C sparse,i ); Using two different decoding networks Φ occ and Φ sem Decode the semantic occupancy grid prediction results from the dense fusion features respectively and 8. A multi-agent collaborative semantic grid prediction system based on uncertainty, characterized in that: include: Single-machine grid prediction module: extracts the three-dimensional features of the input data and predicts whether each spatial position is occupied and the category of the object occupying the spatial position as the occupancy grid prediction result; Geometric uncertainty module: generating a geometric uncertainty distribution map based on the input data; Semantic uncertainty module: generating a semantic uncertainty distribution map based on the three-dimensional features and the occupancy grid prediction result; Message package module: constructing a message package based on the three-dimensional features, the geometric uncertainty distribution map and the semantic uncertainty distribution map; Fusion module: fuses the message packets from multiple agents and the three-dimensional features extracted by their own perception to obtain fusion features; Collaborative grid prediction module: decodes the fusion features to obtain collaborative occupancy grid prediction results.
9. A terminal comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the program, it can be used to execute the method described in any one of claims 1 to 7, or to run the system described in claim 8.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, it can be used to execute the method described in any one of claims 1 to 7, or to execute the system described in claim 8.
Citation Information
Cited By
Enhanced cooperative positioning method and system for uncertainty perception gating network
CN121829513A