Heterogeneous cooperative sensing 3D target detection method and device based on laser radar

By adopting a two-stage training framework and global view-assisted reinforcement method in the lidar heterogeneous collaborative perception system, the problem of retraining the agent backbone network in the existing technology is solved, and efficient and robust heterogeneous collaborative perception 3D object detection is achieved.

CN120107927AActive Publication Date: 2025-06-06UNIV OF SCI & TECH BEIJING

Patent Information

Application Number
CN202510119210.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-24
Publication Date
2025-06-06
Estimated Expiration
2045-01-24

AI Technical Summary

Technical Problem

In the prior art, when the converged network is updated, the backbone network of all other types of agents must be retrained, which results in additional training costs for other types of agents from different companies and increases the complexity of the collaboration process, affecting the perceived performance of a single agent.

Method used

A heterogeneous collaborative perception 3D object detection method based on lidar is proposed, and a two-stage training framework is adopted: local training stage and collaborative training stage. The agent is independently trained and updated through the local training stage, and feature alignment and minimize differences between agents in the collaborative training stage. A global view-assisted reinforcement method is designed to enhance feature representation capabilities.

Benefits of technology

Each agent type is realized to achieve the best bicycle perception performance under its corresponding model, while ensuring data privacy, not affected by feature fusion network updates, improving the robustness and efficiency of collaborative perception.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120107927A_ABST
    Figure CN120107927A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of cooperative sensing, in particular to a heterogeneous cooperative sensing 3D target detection method and device based on a laser radar. The method comprises a two-stage heterogeneous collaborative perception training framework, wherein the two stages comprise a local training stage and a collaborative training stage; designing a difference method between the minimized agents, and performing feature fusion on the agents participating in collaboration; a global view auxiliary enhancement method is designed, a global feature map and a single-view-angle feature map pass through a gradient inversion layer and are combined with a domain classification head to carry out domain adversarial training, the feature representation capability of each agent is enhanced by using global information, and a unified feature space is constructed. According to the method provided by the invention, in a complex driving scene, the fusion network of the collaborative awareness task can be independently updated under the condition that each backbone network is not retrained, and meanwhile, the performance of bicycle awareness is not influenced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of collaborative sensing technology, and in particular to a laser radar-based heterogeneous collaborative sensing 3D target detection method and device. Background Art

[0002] Autonomous vehicles are widely regarded as an effective means to improve road safety. However, the inherent limitations of single-vehicle perception systems, such as susceptibility to occlusion, limited sensor coverage, and challenges in long-range perception, make them face many problems. In recent years, multi-agent collaborative perception technologies have been developed to address these problems in single-vehicle perception, such as vehicle-to-vehicle (V2V) and vehicle-to-everything (V2X) collaboration. In these systems, multiple agents located at different locations in the same environment cooperate and exchange information through communication to build a unified global perception map and improve the performance of single vehicle perception. However, most existing research focuses on collaborative systems based on homogeneous agents. These methods require the same type of agents, which limits their applicability and flexibility in real-world scenarios.

[0003] To promote real-world applications, several methods have studied heterogeneous collaborative perception by relaxing the same data type constraint and using data from different sensors (such as mechanical lidars or semi-solid lidars with different numbers of beams). However, collaborative training sharing of raw data may undermine the privacy of raw data from each company. In addition, these methods use the same backbone network to process all collaborative perception agents. This unified backbone network may not achieve optimal perception of different types of lidar data.

[0004] Recently, methods such as HEAL [Yifan Lu, Yue Hu, Yiqi Zhong, Dequan Wang, Yanfeng Wang, and Siheng Chen. An extensible framework for open heterogeneous collaborative perception.] not only integrate data from different lidar sensors, but also set different backbone network types for each agent and adopt a two-stage training strategy with reverse alignment. The HEAL method first trains a homogeneous collaborative perception network using data from a single sensor, then fixes the fusion sub-network and trains different backbone network types for different agents separately. When the fusion network is updated, the backbone networks of all other types of agents must be retrained, which will bring additional training costs to other types of agents (possibly from different companies), increase the complexity of the collaboration process, and affect the perception performance of a single agent. Summary of the invention

[0005] In order to solve the technical problem in the prior art that when the fusion network is updated, the backbone networks of all other types of agents must be retrained, which will bring additional training costs to other types of agents from different companies, the embodiment of the present invention provides a heterogeneous collaborative perception 3D target detection method and device based on laser radar. The technical solution is as follows:

[0006] On the one hand, a laser radar-based heterogeneous collaborative perception 3D target detection method is provided, characterized in that the method includes:

[0007] S1. Construct a two-stage heterogeneous collaborative perception training framework. The two stages include a local training stage and a collaborative training stage. The local training stage is used to independently train and update the agents participating in the collaboration. The agents participating in the collaboration use different laser radar sensors with different line counts and different network backbone structures.

[0008] S2. Align the features of the agents participating in the collaboration through the collaborative training phase; design a method to minimize the differences between agents and reduce the differences in feature maps between agents;

[0009] S3. Design a global view assisted enhancement method, pass the global feature map and the single-view feature map through the gradient inversion layer and combine them with the domain classification head for domain adversarial training, use global information to enhance the feature representation ability of each intelligent agent, construct a unified feature space, and complete heterogeneous collaborative perception 3D target detection based on lidar.

[0010] Optionally, the agents participating in the collaboration are independently trained and updated through a local training phase, including:

[0011] The backbone corresponding to each agent type independently trains its single-vehicle perception model through the local training phase of the heterogeneous collaborative perception training framework;

[0012] Among them, the backbone corresponding to the agent type only uses its own data when performing independent training and updating, and does not share the original data during collaborative training;

[0013] When a new agent type needs to be added, the feature map F extracted by the backbone network corresponding to the new agent type will be i =Φ i (X i )∈R (H×W×C) , where Φ i Represents the corresponding backbone network, H, W and C represent the height, width and number of channels of the feature map respectively; through the local training phase of the two-stage heterogeneous collaborative perception training framework, the agent is independently trained to update its single-vehicle perception model.

[0014] Optionally, in S2, feature alignment is performed on the agents participating in the collaboration through a collaborative training phase, including:

[0015] Through the collaborative training phase of the heterogeneous collaborative perception training framework, the models of all participating collaborative agents trained in the first phase are loaded, and the data required for collaborative training is input to obtain the BEV feature map of each agent;

[0016] Then, after the preliminary scale alignment layer, the feature maps from different agents are preliminarily aligned using 1×1 convolution;

[0017] Through the collaborative training phase of the heterogeneous collaborative perception training framework, the parameter information of all participating collaborative agents is aligned with the perspective of the ego vehicle.

[0018] Feature alignment includes: the participating cooperative vehicles communicate coordinate information to the ego vehicle, the ego vehicle calculates the coordinate transformation matrix based on the participating cooperative vehicles' coordinate information, and transforms the participating cooperative vehicles' feature maps into the ego vehicle's coordinate system;

[0019] The transformed features are expressed as in Represents a feature transformation operation.

[0020] Optionally, in S2, a method for minimizing the difference between agents is designed to reduce the difference in feature maps between agents, including:

[0021] The feature graph {F′ of each agent 1 , F′ 2 , ... F′ i , F′ j F′ k , ... F ′ N};

[0022] Construct a multi-scale fusion network to transform the feature map {F′ 1 , F′ 2 , ... F′ i , F′ j F′ k , ... F′ N}, input into the multi-scale fusion network; get the feature map representation extracted at each scale Where α represents the αth scale, i, j, k correspond to the agents participating in collaborative sensing;

[0023] Calculate the prospect estimation map S of the two agents participating in collaborative perception; apply 1x1 convolution to the feature map of each agent at each scale respectively, and the prospect estimation map of each agent at scale α is calculated as follows:

[0024]

[0025] Based on the prospect estimation graph of each agent at scale α, the weight matrix between agents i and j is calculated:

[0026]

[0027] At each scale α, calculate the KL divergence m between the feature maps of agents i and j α :

[0028]

[0029] Where N is the total number of collaborative agents, i, j ∈ {1, 2, ..., N}; A is the total number of scales, α ∈ {1, 2, ..., A};

[0030] Calculate the total difference across all scales for:

[0031]

[0032] Where M represents the total difference loss;

[0033] By minimizing The feature maps of different agents are aligned at each scale, and the features of all participating collaborative agents are fused.

[0034] Optionally, a global view assisted enhancement method is designed to pass the global feature map and the single view feature map through the gradient reversal layer and combine them with the domain classification head for domain adversarial training, using global information to enhance the feature representation ability of each agent and construct a unified feature space, including:

[0035] Get the feature maps of different agents at scale α

[0036] Input the feature maps of different agents at scale α into the foreground generator to obtain the foreground estimation map;

[0037] Use the softmax function to normalize the foreground estimation map and generate the weight matrix W for feature fusion;

[0038] Calculate the weighted feature map of different agents at scale α:

[0039]

[0040] The feature map after fusion at scale α Merge with the feature map of a single view to form a combined feature set

[0041] The combined feature set Fall is input into the gradient reversal layer to obtain the inverted features, and the inverted features are input into the domain classifier. The domain classifier is used to determine whether the input features come from the global fusion features or from the individual view features. The label is set to 1, and the individual view features The label is set to 0; output the field classification result;

[0042] Among them, the domain classifier network structure is:

[0043] Two 3*3 convolutional layers with padding 1:

[0044] Y1=relu(Conv 3×3 (F all ))

[0045] Y2=relu(Conv 3×3 (Y1)

[0046] Perform maximum pooling on the input feature map:

[0047] Y3=maxpool(Y2)

[0048] Flatten the feature map:

[0049] Y3=flatten(Y2)

[0050] After the transformation of three fully connected layers FC, the original output logits without activation function are finally obtained;

[0051] Y5=dropout(relu(fc(Y)))Y6=dropout(relu(fc(Y5)))

[0052] logits=fc(Y6)

[0053] Based on the domain classification results, the domain classification cross entropy loss at each scale α and the total domain classification loss at all scales are calculated.

[0054] Optionally, calculate the domain classification cross entropy loss at each scale α and the total domain classification loss at all scales, including:

[0055] Calculate the domain classification cross entropy loss at each scale α:

[0056]

[0057] Among them, y refers to the logits obtained by the domain classifier;

[0058] Calculate the domain classification cross entropy loss at each scale α:

[0059]

[0060] Optionally, S3 also includes:

[0061] Calculate the overall loss function of the trained model:

[0062]

[0063] Among them, β, γ and ∈ are hyperparameters.

[0064] On the other hand, a laser radar-based heterogeneous collaborative sensing 3D target detection device is provided, and the device is applied to a laser radar-based heterogeneous collaborative sensing 3D target detection method, and the device includes: a data preprocessing module, which is used to input a preset input program into a compiler front end and output an intermediate representation;

[0065] The heterogeneous collaborative perception framework design module is used to build a heterogeneous collaborative perception training framework. The backbones of different vendors independently update the intelligent agents through the heterogeneous collaborative perception training framework.

[0066] The consistency enhancement module is used to obtain the feature map of the BEV participating in the cooperation and then perform feature fusion through the heterogeneous cooperative perception training framework; the consistency enhancement module includes: a submodule for minimizing the feature differences between intelligent agents and a submodule for global perspective auxiliary enhancement;

[0067] The submodule for minimizing the feature differences between agents is used to design a method to minimize the differences between agents and reduce the differences in feature maps between agents.

[0068] The global view auxiliary enhancement submodule is used to design a global view auxiliary enhancement method. The global feature map and the single view feature map are passed through a gradient inversion layer and combined with a domain classification head for domain adversarial training. The global information is used to enhance the feature representation ability of each intelligent agent, construct a unified feature space, and complete heterogeneous collaborative perception 3D target detection based on lidar.

[0069] On the other hand, a laser radar-based heterogeneous collaborative perception 3D target detection device is provided, and the laser radar-based heterogeneous collaborative perception 3D target detection device includes: a processor; a memory, and the memory stores computer-readable instructions. When the computer-readable instructions are executed by the processor, any one of the above-mentioned laser radar-based heterogeneous collaborative perception 3D target detection methods is implemented.

[0070] On the other hand, a computer-readable storage medium is provided, in which at least one instruction is stored. The at least one instruction is loaded and executed by a processor to implement any one of the above-mentioned heterogeneous collaborative perception 3D target detection methods based on laser radar.

[0071] The beneficial effects brought about by the technical solution provided by the embodiment of the present invention include at least:

[0072] In the embodiment of the present invention, the core method of the framework mainly includes three parts: 1) a two-stage training strategy for heterogeneous collaborative perception tasks based on lidar is proposed, so that the intelligent agents of each manufacturer in the collaborative perception system remain relatively independent, their respective data have a certain degree of confidentiality, and their own single-vehicle perception performance is not affected by the collaborative perception task, and is not affected by the update of the feature fusion network; 2) in the feature fusion module part, the design of minimizing the difference between intelligent agents is introduced to reduce the difference in feature distribution between heterogeneous intelligent agents with different perspectives in the same scene, and improve the feature fusion effect; 3) in the feature fusion module part, the global view auxiliary reinforcement design is introduced, the global feature map and the single-view feature map are passed through the gradient inversion layer and combined with the domain classification head for domain adversarial training, and the global information is used to enhance the feature representation ability of each intelligent agent, and a unified feature space is constructed, reducing the domain difference. This design ensures the effectiveness and stability of multi-view feature fusion and enhances the robustness of collaborative perception. A large number of experiments show that for V2X and V2V heterogeneous collaborative perception, the perception effect of the present invention is better than the existing state-of-the-art methods. BRIEF DESCRIPTION OF THE DRAWINGS

[0073] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0074] Figure 1 A schematic diagram of a process flow of a laser radar-based heterogeneous collaborative perception 3D target detection method provided in an embodiment of the present invention;

[0075] Figure 2 A laser radar-based heterogeneous collaborative perception 3D target detection diagram provided in an embodiment of the present invention;

[0076] Figure 3 A block diagram of a laser radar-based heterogeneous collaborative perception 3D target detection device provided in an embodiment of the present invention;

[0077] Figure 4 A schematic diagram of the structure of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0078] The technical solution of the present invention is described below in conjunction with the accompanying drawings.

[0079] In the embodiments of the present invention, words such as "exemplarily" and "for example" are used to indicate examples, illustrations or explanations. Any embodiment or design described as "example" in the present invention should not be interpreted as being more preferred or more advantageous than other embodiments or designs. Specifically, the use of the word "example" is intended to present the concept in a specific way. In addition, in the embodiments of the present invention, the meaning expressed by "and / or" can be both, or it can be either of the two.

[0080] In the embodiments of the present invention, sometimes the subscripts such as W 1 It may be written in non-subscript form such as W1. When the difference is not emphasized, the meaning is the same.

[0081] In order to make the technical problems, technical solutions and advantages to be solved by the present invention more clear, a detailed description will be given below with reference to the accompanying drawings and specific embodiments.

[0082] The embodiment of the present invention provides a laser radar-based heterogeneous collaborative perception 3D target detection method, which can be implemented by a laser radar-based heterogeneous collaborative perception 3D target detection device, which can be a terminal or a server. Figure 1 The flowchart of the heterogeneous collaborative perception 3D target detection method based on lidar is shown in Figure 1 As shown, the present invention proposes a laser radar-based heterogeneous collaborative perception 3D target detection method, and the processing flow of the method may include the following steps:

[0083] S1. Construct a two-stage heterogeneous collaborative perception training framework. The two stages include a local training stage and a collaborative training stage. The local training stage is used to independently train and update the agents participating in the collaboration. Among them, the agents participating in the collaboration use lidar sensors with different line counts and different network backbone structures.

[0084] In a feasible implementation, the present invention proposes a heterogeneous collaborative perception training framework, which is a two-stage training strategy for heterogeneous collaborative perception tasks based on LiDAR. The original data of each type of intelligent agent is privacy-protected, and the integration of different types of collaborative agents can be achieved. When a new agent is added, the collaborative network can be updated independently without affecting other participating agents, without retraining their private backbone networks, to maximize the adaptability of heterogeneous features extracted from LiDAR sensors of different agents. In practical applications, each manufacturer's best single-vehicle perception model is independently trained and confidential, so it is impractical to require each manufacturer to contribute its data for collaborative perception task training. Therefore, the present invention fixes the backbone network for each manufacturer, retaining its best single-vehicle perception performance, while not affecting the original decision-making module of each vehicle.

[0085] In a feasible implementation, in S1, the agents participating in the collaboration are independently trained and updated through the local training phase, including:

[0086] The backbone corresponding to each agent type independently trains its single-vehicle perception model through the local training phase of the heterogeneous collaborative perception training framework;

[0087] Among them, the backbone corresponding to the agent type only uses its own data during independent training and updating, and does not share the original data during collaborative training.

[0088] When a new agent type needs to be added, the feature map F extracted by the backbone network corresponding to the new agent type will be i =Φ i (X i )∈R (H×W×C) , where Φ i Represents the corresponding backbone network, H, W and C represent the height, width and number of channels of the feature map respectively; through the local training phase of the two-stage heterogeneous collaborative perception training framework, the agent is independently trained to update its single-vehicle perception model.

[0089] In one feasible implementation, during the local training phase, each type of agent independently trains its single-vehicle perception model. Figure 2 As shown on the left, each agent converts the original 3D point cloud into a bird’s-eye view (BEV) map X i The feature map F extracted by its corresponding backbone network i =Φ i (X i )∈R (H×W×C) , where Φ i represents the corresponding backbone network, H, W and C represent the height, width and number of channels of the feature map respectively.

[0090] In the embodiment of the present invention, a novel heterogeneous training framework is proposed, which has the characteristic of protecting the privacy backbone network, so that in complex driving scenarios, the fusion network of collaborative perception tasks can be independently updated without retraining each backbone network, while maintaining the optimal performance of single-vehicle perception without being affected.

[0091] S2. Align the features of the agents participating in the collaboration through the collaborative training phase; design a method to minimize the differences between the agents and reduce the differences in feature maps between the agents.

[0092] In a feasible implementation, feature alignment is performed on the agents participating in the collaboration through the collaborative training phase, including:

[0093] Through the collaborative training phase of the heterogeneous collaborative perception training framework, the models of all participating collaborative agents trained in the first phase are loaded, the data required for collaborative training is input, and the BEV feature map of each agent is obtained through the backbone of each agent type;

[0094] Then, after the preliminary scale alignment layer, the feature maps from different agents are preliminarily aligned using 1×1 convolution;

[0095] Through the collaborative training phase of the heterogeneous collaborative perception training framework, the parameter information of all participating collaborative agents is aligned with the perspective conversion of the ego vehicle.

[0096] Feature alignment includes: the participating cooperative vehicles communicate coordinate information to the self-vehicle, the self-vehicle calculates the coordinate transformation matrix based on the coordinate information of the participating cooperative vehicles, and transforms the feature map of the participating cooperative vehicles into the self-vehicle coordinate system.

[0097] The transformed features are expressed as in Represents a feature transformation operation.

[0098] In the collaborative training phase, the backbone network parameters of each agent type obtained in the local training are fixed. Figure 2 As shown in Figure 1, one of the vehicles is called the ego vehicle, which broadcasts its own location information to all other cooperative vehicles. After receiving the location information of the ego vehicle, each cooperative vehicle will i Transform the view to align it with the features of the vehicle. The transformed features are represented as in Represents a feature transformation operation.

[0099] In a feasible implementation, in S2, a method for minimizing the difference between agents is designed to reduce the difference in feature graphs between agents, including:

[0100] The feature graph {F′ of each agent 1 , F′ 2 , ... F′ i , F′ j F′ k , ... F′ N};

[0101] Construct a multi-scale fusion network to transform the feature map {F′ 1 , F′ 2 , ... F′ i , F′ j F′ k , ... F′ N}, input into the multi-scale fusion network; get the feature map representation extracted at each scale Where α represents the αth scale, i, j, k correspond to the agents participating in collaborative sensing;

[0102] Calculate the prospect estimation map S of the two agents participating in collaborative perception; apply 1x1 convolution to the feature map of each agent at each scale respectively, and the prospect estimation map of each agent at scale α is calculated as follows:

[0103]

[0104] Based on the prospect estimation graph of each agent at scale α, the weight matrix between agents i and j is calculated:

[0105]

[0106] At each scale α, calculate the KL divergence m between the feature maps of agents i and j α :

[0107]

[0108] Where N is the total number of collaborative agents, i, j ∈ {1, 2, ..., N}; A is the total number of scales, α ∈ {1, 2, ..., A};

[0109] Calculate the total difference across all scales for:

[0110]

[0111] Where M represents the total difference loss;

[0112] By minimizing The feature maps of different agents are aligned at each scale, and the features of all participating collaborative agents are fused.

[0113] In one possible implementation, the parameters of the backbone network learned during training may vary. These differences directly lead to the feature maps {F′ 1 , F′ 2 , ... F′ i , F′ j F′ k , ... F′ N}, which in turn causes changes in information distribution. These differences make subsequent feature fusion more complicated and increase the difficulty of feature alignment and sharing. In order to minimize the differences in feature maps between agents, a simple and effective method is to use KL divergence to measure the differences in feature distribution. Therefore, the present invention calculates the KL divergence between feature maps of different agents, and uses the foreground estimation map of the agent to calculate the loss weight, reduce meaningless alignment, and increase the weight of effective alignment.

[0114] In one possible implementation, in this general form, The overall difference between all scales and agents is quantified by minimizing The present invention can align the feature maps of different agents at each scale, ensuring that the feature distribution at each spatial position is similar. This operation improves the quality of feature alignment across multiple spatial scales, produces a more robust multi-scale feature representation, and ensures that these features remain consistent across different agents.

[0115] S3. Design a global view assisted enhancement method, pass the global feature map and the single-view feature map through the gradient inversion layer and combine them with the domain classification head for domain adversarial training, use global information to enhance the feature representation ability of each intelligent agent, construct a unified feature space, and complete heterogeneous collaborative perception 3D target detection based on lidar.

[0116] In a feasible implementation, in S3, a global view assisted enhancement method is designed, the global feature map and the single view feature map are subjected to a gradient reversal layer and combined with a domain classification head for domain adversarial training, the global information is used to enhance the feature representation ability of each agent, and a unified feature space is constructed, including:

[0117] Get the feature maps of different agents at scale α

[0118] Input the feature maps of different agents at scale α into the foreground generator to obtain the foreground estimation map;

[0119] Use the softmax function to normalize the foreground estimation map and generate the weight matrix W for feature fusion;

[0120] Calculate the weighted feature map of different agents at scale α:

[0121]

[0122] The feature map after fusion at scale α Merge with the feature map of a single view to form a combined feature set

[0123] The combined feature set is input into the gradient reversal layer to obtain the inverted features, and the inverted features are input into the domain classifier. The domain classifier is used to determine whether the input features come from the global fusion features or from the individual view features. The label is set to 1, and the individual view features The label is set to 0; the field classification result is output;

[0124] Among them, the domain classifier network structure is:

[0125] Two 3*3 convolutional layers with padding 1:

[0126] Y1=relu(Conv 3×3 (F all ))

[0127] Y2=relu(Conv 3×3 (Y1))

[0128] Perform maximum pooling on the input feature map:

[0129] Y3=maxpool(Y2)

[0130] Flatten the feature map:

[0131] Y3=flatten(Y2)

[0132] After the transformation of three fully connected layers FC, the original output logits without activation function are finally obtained;

[0133] Y5 = dropout(relu(fc(Y4)))

[0134] Y6 = dropout(relu(fc(Y5)))

[0135] logits=fc(Y6)

[0136] Based on the domain classification results, the domain classification cross entropy loss at each scale α and the total domain classification loss at all scales are calculated.

[0137] In a feasible implementation, although the differences between agents have been minimized, it is still necessary to use the overall information to enhance the features of each agent. To solve this problem, the present invention can use the fused features (F u ) to form overall information and guide the alignment of distributions between different agents. However, due to occlusion problems, the feature map contains many empty value areas, and different agents observe different areas. When using KL divergence for alignment, empty features may produce meaningless alignment with non-empty features, which will reduce the expressive power of the network. Therefore, the present invention selects domain adversarial training, where adversarial training is used as a soft alignment technique.

[0138] In one feasible implementation, the domain classification cross entropy loss at each scale α and the total domain classification loss at all scales are calculated, including:

[0139] Calculate the domain classification cross entropy loss at each scale α:

[0140]

[0141] Among them, y refers to the logits obtained by the domain classifier;

[0142] Calculate the domain classification cross entropy loss at each scale α:

[0143]

[0144] In a feasible implementation, the main goal of the domain classifier is to distinguish whether the input feature comes from the global fusion feature or from the individual view feature, so that the model can effectively distinguish features from different sources. The gradient reversal layer L promotes the adversarial learning process between the feature extractor and the domain classifier, so that the feature extractor learns a way that makes it difficult to distinguish features from different domains. This encourages the model to suppress domain-specific feature differences during the learning process and promotes the extraction of cross-domain adaptive features.

[0145] Therefore, the model gradually balances the relationship between global features and local features, produces consistent feature representations under different viewpoints, and reduces domain differences. This design ensures the rationality and stability of multi-view feature fusion and enhances the robustness of collaborative perception.

[0146] In a feasible implementation manner, S3 further includes:

[0147] The fused feature map F u It is input to the classification head (for foreground / background classification) and the regression head (for generating bounding boxes). The perception result on the ego vehicle is represented as Y = head (F u), where head refers to the classification head and regression head. In addition to these two detection losses L cls and L reg , the present invention also uses the cross entropy loss L da To supervise the domain adversarial training, and use the difference between agents to minimize the loss L M To supervise the multi-scale feature extraction process. Calculate the overall loss function of the training model:

[0148]

[0149] Among them, β, γ and ∈ are super parameters. In the first stage (local training stage), the present invention adopts L cls +βL reg In the second stage (co-training stage), the present invention uses the total loss in the formula to train the model.

[0150] In an embodiment of the present invention, a novel two-stage method - dual-consistency multi-agent collaborative perception is proposed. A heterogeneous training framework with a privacy-preserving backbone network is adopted, and includes a local training phase and a collaborative training phase. In the local training phase, a single-vehicle perception model is trained for each collaborative vehicle type. In order to ensure that each agent achieves the best single-vehicle perception performance under its corresponding model and to ensure the privacy of the data, the present invention strictly assigns different types of data to corresponding models for training. In the collaborative training phase, the backbone network of the agent remains fixed, and only the fusion network and subsequent detection head are trained.

[0151] Figure 3 1 is a block diagram of a laser radar-based heterogeneous collaborative perception 3D target detection device 300 according to an exemplary embodiment. The device 300 is used for a laser radar-based heterogeneous collaborative perception 3D target detection method. Figure 3 The device includes a heterogeneous collaborative perception framework design module 310, a consistency enhancement module 320, a submodule 321 for minimizing feature differences between agents, and a global perspective auxiliary enhancement submodule 322. Among them:

[0152] The heterogeneous collaborative perception framework design module 310 is used to construct a heterogeneous collaborative perception training framework; each type of backbone of a manufacturer independently updates the intelligent agent through the heterogeneous collaborative perception training framework;

[0153] The consistency enhancement module 320 is used to obtain the feature map of the BEV participating in the cooperation, and then perform feature fusion through the heterogeneous cooperative perception training framework; the consistency enhancement module 320 includes: a submodule 321 for minimizing feature differences between agents and a submodule 322 for assisting global perspective enhancement;

[0154] The submodule 321 for minimizing the feature difference between agents is used to design a method for minimizing the difference between agents and reduce the feature map difference between agents;

[0155] The global view auxiliary enhancement submodule 322 is used to design a global view auxiliary enhancement method, which passes the global feature map and the single view feature map through a gradient inversion layer and combines them with a domain classification head for domain adversarial training, uses global information to enhance the feature representation ability of each intelligent agent, constructs a unified feature space, and completes heterogeneous collaborative perception 3D target detection based on lidar.

[0156] Optionally, the agents participating in the collaboration are independently trained and updated through a local training phase, including:

[0157] The backbone corresponding to each agent type independently trains its single-vehicle perception model through the local training phase of the heterogeneous collaborative perception training framework;

[0158] Among them, the backbone corresponding to the agent type only uses its own data during independent training and updating, and does not share the original data during collaborative training.

[0159] When a new agent type needs to be added, the feature map F extracted by the backbone network corresponding to the new agent type will be i =Φ i (X i )∈R (H×W×C) , where Φ i Represents the corresponding backbone network, H, W and C represent the height, width and number of channels of the feature map respectively; through the local training phase of the two-stage heterogeneous collaborative perception training framework, the agent is independently trained to update its single-vehicle perception model.

[0160] Optionally, the minimizing inter-agent feature difference submodule 321 is used to:

[0161] Through the collaborative training phase of the heterogeneous collaborative perception training framework, the models of all participating collaborative agents trained in the first phase are loaded, and the data required for collaborative training is input to obtain the BEV feature map of each agent;

[0162] Then, after the preliminary scale alignment layer, the feature maps from different agents are preliminarily aligned using 1×1 convolution;

[0163] Through the collaborative training phase of the heterogeneous collaborative perception training framework, the parameter information of all participating collaborative agents is aligned with the perspective of the ego vehicle.

[0164] Feature alignment includes: the participating cooperative vehicles communicate coordinate information to the ego vehicle, the ego vehicle calculates the coordinate transformation matrix based on the participating cooperative vehicles' coordinate information, and transforms the participating cooperative vehicles' feature maps into the ego vehicle's coordinate system;

[0165] The transformed features are expressed as in Represents a feature transformation operation.

[0166] Optionally, the minimizing inter-agent feature difference submodule 321 is used to:

[0167] The feature graph {F′ of each agent 1 , F′ 2 , ... F′ i , F′ j F′ k , ... F′ N};

[0168] Construct a multi-scale fusion network to transform the feature map {F′ 1 , F′ 2 , ... F′ i , F′ j F′ k , ... F′ N}, input into the multi-scale fusion network; get the feature map representation extracted at each scale Where α represents the αth scale, i, j, k correspond to the agents participating in collaborative sensing;

[0169] Calculate the prospect estimation map S of the two agents participating in collaborative perception; apply 1x1 convolution to the feature map of each agent at each scale respectively, and the prospect estimation map of each agent at scale α is calculated as follows:

[0170]

[0171] Based on the prospect estimation graph of each agent at scale α, the weight matrix between agents i and j is calculated:

[0172]

[0173] At each scale α, calculate the KL divergence m between the feature maps of agents i and j α :

[0174]

[0175] Where N is the total number of collaborative agents, i, j ∈ {1, 2, ..., N}; A is the total number of scales, α ∈ {1, 2, ..., A};

[0176] Calculate the total difference across all scales for:

[0177]

[0178] Where M represents the total difference loss;

[0179] By minimizing The feature maps of different agents are aligned at each scale, and the features of all participating collaborative agents are fused.

[0180] Optionally, the global perspective auxiliary enhancement submodule 322 is used to:

[0181] Get the feature maps of different agents at scale α

[0182] Input the feature maps of different agents at scale α into the foreground generator to obtain the foreground estimation map;

[0183] Use the softmax function to normalize the foreground estimation map and generate the weight matrix W for feature fusion;

[0184] Calculate the weighted feature map of different agents at scale α:

[0185]

[0186] The feature map after fusion at scale α Merge with the feature map of a single view to form a combined feature set

[0187] The combined feature set Fall is input into the gradient reversal layer to obtain the inverted features, and the inverted features are input into the domain classifier. The domain classifier is used to determine whether the input features come from the global fusion features or from the individual view features. The label is set to 1, and the individual view features The label is set to 0; output the field classification result;

[0188] Among them, the domain classifier network structure is:

[0189] Two 3*3 convolutional layers with padding 1:

[0190] Y1=relu(Conv 3×3 (F all ))

[0191] Y2=relu(Conv 3×3 (Y1))

[0192] Perform maximum pooling on the input feature map:

[0193] Y3=maxpool(Y2)

[0194] Flatten the feature map:

[0195] Y3=flatten(Y2)

[0196] After the transformation of three fully connected layers FC, the original output logits without activation function are finally obtained;

[0197] Y5 = dropout(relu(fc(Y4)))

[0198] Y6 = dropout(relu(fc(Y5)))

[0199] logits=fc(Y6)

[0200] Based on the domain classification results, the domain classification cross entropy loss at each scale α and the total domain classification loss at all scales are calculated.

[0201] Optionally, the global perspective auxiliary enhancement submodule 322 is used to:

[0202] Calculate the domain classification cross entropy loss at each scale α:

[0203]

[0204] Among them, y refers to the logits obtained by the domain classifier;

[0205] Calculate the domain classification cross entropy loss at each scale α:

[0206]

[0207] Optionally, the global perspective auxiliary enhancement submodule 322 is further used to:

[0208] Calculate the overall loss function of the trained model:

[0209]

[0210] Among them, β, γ and ∈ are hyperparameters.

[0211] In an embodiment of the present invention, a novel two-stage method - dual consistency multi-agent collaborative perception (DCon) is proposed. DCon adopts a heterogeneous training framework with a privacy-preserving backbone network, and includes a local training phase and a collaborative training phase. In the local training phase, DCon trains a single-vehicle perception model for each collaborative vehicle type. In order to ensure that each agent achieves the best single-vehicle perception performance under its corresponding model and to ensure the privacy of the data, we strictly assign different types of data to the corresponding models for training. In the collaborative training phase, the backbone network of the agent remains fixed, and only the fusion network and subsequent detection head are trained.

[0212] Figure 4 is a schematic diagram of the structure of a laser radar-based heterogeneous collaborative perception 3D target detection device provided by an embodiment of the present invention, such as Figure 4 As shown, the heterogeneous collaborative sensing 3D target detection device based on laser radar may include the above Figure 3 The heterogeneous collaborative sensing 3D target detection device based on laser radar is shown. Optionally, the heterogeneous collaborative sensing 3D target detection device based on laser radar 410 may include a first processor 2001.

[0213] Optionally, the laser radar-based heterogeneous collaborative perception 3D target detection device 410 may also include a memory 2002 and a transceiver 2003 .

[0214] The first processor 2001, the memory 2002 and the transceiver 2003 may be connected via a communication bus.

[0215] Combine the following Figure 4 The components of the laser radar-based heterogeneous collaborative perception 3D target detection device 410 are specifically introduced:

[0216] The first processor 2001 is the control center of the laser radar-based heterogeneous collaborative perception 3D target detection device 410, which can be a processor or a general term for multiple processing elements. For example, the first processor 2001 is one or more central processing units (CPUs), or an application specific integrated circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of the present invention, such as one or more microprocessors (digital signal processors, DSPs), or one or more field programmable gate arrays (field programmable gate arrays, FPGAs).

[0217] Optionally, the first processor 2001 can perform various functions of the laser radar-based heterogeneous collaborative perception 3D target detection device 410 by running or executing a software program stored in the memory 2002 and calling data stored in the memory 2002.

[0218] In a specific implementation, as an embodiment, the first processor 2001 may include one or more CPUs, such as Figure 4 CPU0 and CPU1 are shown in FIG.

[0219] In a specific implementation, as an embodiment, the laser radar-based heterogeneous collaborative perception 3D target detection device 410 may also include multiple processors, such as Figure 4 The first processor 2001 and the second processor 2004 are shown in FIG. Each of these processors can be a single-core processor (single-CPU) or a multi-core processor (multi-CPU). The processor here can refer to one or more devices, circuits, and / or processing cores for processing data (e.g., computer program instructions).

[0220] The memory 2002 is used to store the software program for executing the solution of the present invention, and is controlled to be executed by the first processor 2001. The specific implementation method can refer to the above method embodiment, which will not be repeated here.

[0221] Optionally, the memory 2002 may be a read-only memory (ROM) or other types of static storage devices that can store static information and instructions, a random access memory (RAM) or other types of dynamic storage devices that can store information and instructions, or an electrically erasable programmable read-only memory (EEPROM), a compact disc read-only memory (CD-ROM) or other optical disc storage, optical disc storage (including compressed optical disc, laser disc, optical disc, digital versatile disc, Blu-ray disc, etc.), a magnetic disk storage medium or other magnetic storage device, or any other medium that can be used to carry or store the desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto. The memory 2002 may be integrated with the first processor 2001, or may exist independently, and may be accessed through the interface circuit ( Figure 4 (not shown) is coupled to the first processor 2001, which is not specifically limited in this embodiment of the present invention.

[0222] The transceiver 2003 is used to communicate with a network device or a terminal device.

[0223] Optionally, the transceiver 2003 may include a receiver and a transmitter ( Figure 4 The receiver is used to implement a receiving function, and the transmitter is used to implement a sending function.

[0224] Optionally, the transceiver 2003 may be integrated with the first processor 2001, or may exist independently, and may be connected to the first processor 2001 through the interface circuit ( Figure 4 (not shown) is coupled to the first processor 2001, which is not specifically limited in this embodiment of the present invention.

[0225] It should be noted that Figure 4 The structure of the laser radar-based heterogeneous collaborative perception 3D target detection device 410 shown in the figure does not constitute a limitation on the router. The actual knowledge structure recognition device may include more or fewer components than shown in the figure, or combine certain components, or arrange the components differently.

[0226] In addition, the technical effects of the laser radar-based heterogeneous collaborative perception 3D target detection device 410 can refer to the technical effects of the laser radar-based heterogeneous collaborative perception 3D target detection method described in the above method embodiment, and will not be repeated here.

[0227] It should be understood that the first processor 2001 in the embodiment of the present invention may be a central processing unit (CPU), and the processor may also be other general-purpose processors, digital signal processors (DSP), application specific integrated circuits (ASIC), field programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.

[0228] It should also be understood that the memory in the embodiments of the present invention may be a volatile memory or a non-volatile memory, or may include both volatile and non-volatile memories. Among them, the non-volatile memory may be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory may be a random access memory (RAM), which is used as an external cache. By way of example and not limitation, many forms of random access memory (RAM) are available, such as static RAM (SRAM), dynamic random access memory (DRAM), synchronous DRAM (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link DRAM (SLDRAM), and direct rambus RAM (DR RAM).

[0229] The above embodiments can be implemented in whole or in part by software, hardware (such as circuits), firmware or any other combination. When implemented by software, the above embodiments can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded or executed on a computer, the process or function described in the embodiment of the present invention is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable sensors. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from one website, computer, server or data center to another website, computer, server or data center by wired (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or data center that contains one or more available media sets. The available medium can be a magnetic medium (for example, a floppy disk, a hard disk, a tape), an optical medium (for example, a DVD), or a semiconductor medium. The semiconductor medium can be a solid-state hard disk.

[0230] It should be understood that the term "and / or" in this article is only a description of the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist at the same time, and B exists alone. A and B can be singular or plural. In addition, the character " / " in this article generally indicates that the associated objects before and after are in an "or" relationship, but it may also indicate an "and / or" relationship. Please refer to the context for specific understanding.

[0231] It should be understood that in various embodiments of the present invention, the size of the serial numbers of the above-mentioned processes does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.

[0232] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of the present invention.

[0233] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0234] In addition, each functional unit in each embodiment of the present invention may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.

[0235] If the functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or the part of the technical solution, can be embodied in the form of a software product, which is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the methods described in various embodiments of the present invention.

[0236] The above is only a specific embodiment of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art can easily think of changes or substitutions within the technical scope disclosed by the present invention, which should be included in the protection scope of the present invention. Therefore, the protection scope of the present invention should be based on the protection scope of the claims.

Claims

1. A laser radar-based heterogeneous collaborative perception 3D target detection method, characterized in that: The method comprises: S1. Construct a two-stage heterogeneous collaborative perception training framework. The two stages include a local training stage and a collaborative training stage. The local training stage is used to independently train and update the agents participating in the collaboration. The agents participating in the collaboration use different laser radar sensors with different line counts and different network backbone structures. S2. Align the features of the agents participating in the collaboration through the collaborative training phase; design a method to minimize the differences between agents and reduce the differences in feature maps between agents; S3. Design a global view assisted enhancement method, pass the global feature map and the single-view feature map through the gradient inversion layer and combine them with the domain classification head for domain adversarial training, use global information to enhance the feature representation ability of each intelligent agent, construct a unified feature space, and complete heterogeneous collaborative perception 3D target detection based on lidar.

2. The laser radar-based heterogeneous collaborative perception 3D target detection method according to claim 1 is characterized in that: The local training phase is used to independently train and update the agents involved in the collaboration, including: The backbone corresponding to each agent type independently trains its single-vehicle perception model through the local training phase of the heterogeneous collaborative perception training framework; The backbone corresponding to the agent type only uses its own data when performing independent training and updating, and does not share the original data during collaborative training; When a new agent type needs to be added, the feature graph F extracted by the backbone network corresponding to the new agent type is i =Φ i (X i )∈R (H×W×C) , where Φ i represents the corresponding backbone network, H, W and C represent the height, width and number of channels of the feature map respectively; through the local training phase of the two-stage heterogeneous collaborative perception training framework, the agent is independently trained to update its single-vehicle perception model.

3. The laser radar-based heterogeneous collaborative perception 3D target detection method according to claim 2 is characterized in that: In S2, feature alignment is performed on the agents participating in the collaboration through the collaborative training phase, including: Through the collaborative training phase of the heterogeneous collaborative perception training framework, the models of all participating collaborative agents trained in the first phase are loaded, and the data required for collaborative training is input to obtain the BEV feature map of each agent; Then, after the preliminary scale alignment layer, the feature maps from different agents are preliminarily aligned using 1×1 convolution; Through the collaborative training phase of the heterogeneous collaborative perception training framework, the parameter information of all the collaborative agents is aligned with the perspective of the ego vehicle; The feature alignment includes: the participating cooperative vehicle communicates coordinate information to the self-vehicle, the self-vehicle calculates a coordinate conversion matrix according to the coordinate information of the participating cooperative vehicle, and converts the feature map of the participating cooperative vehicle into the self-vehicle coordinate system; The transformed features are expressed as in Represents a feature transformation operation.

4. The laser radar-based heterogeneous collaborative perception 3D target detection method according to claim 3 is characterized in that: In S2, a method for minimizing the differences between agents is designed to reduce the differences in feature maps between agents, including: The feature graph of each agent {F′1, F′2, ... F′ i , F′ j F′ k , ... F′ N }; Construct a multi-scale fusion network and transform the feature graphs {F′1, F′2, ... F′ i , F′ j F′ k , ... F′ N }, input into the multi-scale fusion network; get the feature map representation extracted at each scale Where α represents the αth scale, i, j, k correspond to the agents participating in collaborative sensing; Calculate the prospect estimation map S of the two agents participating in collaborative perception; apply 1x1 convolution to the feature map of each agent at each scale respectively, and the prospect estimation map of each agent at scale α is calculated as follows: Based on the prospect estimation graph of each agent at the scale α, the weight matrix between agents i and j is calculated: At each scale α, calculate the KL divergence m between the feature maps of agents i and j α : Where N is the total number of collaborative agents, i, j ∈ {1, 2, ..., N}; A is the total number of scales, α ∈ {1, 2, ..., A}; Calculate the total difference across all scales for: Where M represents the total difference loss; By minimizing The feature maps of different agents are aligned at each scale, and the features of all the participating collaborative agents are fused.

5. The laser radar-based heterogeneous collaborative perception 3D target detection method according to claim 6 is characterized in that: The global view assisted enhancement method is designed to perform domain adversarial training on the global feature map and the single view feature map through a gradient reversal layer and combined with a domain classification head, and use global information to enhance the feature representation ability of each agent and construct a unified feature space, including: Get the feature maps of different agents at scale α Input the feature maps of different agents at scale α into the foreground generator to obtain the foreground estimation map; Use the softmax function to normalize the foreground estimation map and generate the weight matrix W for feature fusion; Calculate the weighted feature map of different agents at scale α: The feature map after fusion at scale α Merge with the feature map of a single view to form a combined feature set The combined feature set Fall is input to the gradient reversal layer to obtain the inverted features, and the inverted features are input to the domain classifier. The domain classifier is used to determine whether the input features come from the global fusion features or from the individual view features. The label is set to 1, and the individual view features The label is set to 0; output the field classification result; Among them, the domain classifier network structure is: Two 3*3 convolutional layers with padding 1: Y1=relu(Conv 3×3 (F all )) Y2=relu(Conv 3×3 (Y1)) Perform maximum pooling on the input feature map: Y3=maxpool(Y2) Flatten the feature map: Y3=flatten(Y2) After the transformation of three fully connected layers FC, the original output logits without activation function are finally obtained; Y5 = dropout(relu(fc(Y4))) Y6 = dropout(relu(fc(Y5))) logits=fc(Y6) Based on the domain classification results, the domain classification cross entropy loss at each scale α and the total domain classification loss at all scales are calculated.

6. The laser radar-based heterogeneous collaborative perception 3D target detection method according to claim 5, characterized in that: Calculate the domain classification cross entropy loss at each scale α and the total domain classification loss at all scales, including: Calculate the domain classification cross entropy loss at each scale α: Among them, y refers to the logits obtained by the domain classifier; Calculate the domain classification cross entropy loss at each scale α:

7. The laser radar-based heterogeneous collaborative perception 3D target detection method according to claim 6, characterized in that: The S3 further includes: Calculate the overall loss function of the trained model: Among them, β, γ and ∈ are hyperparameters.

8. A laser radar-based heterogeneous collaborative perception 3D target detection device, the laser radar-based heterogeneous collaborative perception 3D target detection device is used to implement the laser radar-based heterogeneous collaborative perception 3D target detection method according to any one of claims 1 to 7, characterized in that: The device comprises: A heterogeneous collaborative perception framework design module is used to build a heterogeneous collaborative perception training framework; backbones of different vendor types independently update the intelligent agent through the heterogeneous collaborative perception training framework; A consistency enhancement module is used to obtain the feature map of the BEV participating in the cooperation and then perform feature fusion through the heterogeneous cooperative perception training framework; the consistency enhancement module includes: a submodule for minimizing feature differences between intelligent agents and a submodule for global perspective auxiliary enhancement; The submodule for minimizing the feature differences between agents is used to design a method to minimize the differences between agents and reduce the differences in feature maps between agents. The global view auxiliary enhancement submodule is used to design a global view auxiliary enhancement method. The global feature map and the single view feature map are passed through a gradient inversion layer and combined with a domain classification head for domain adversarial training. The global information is used to enhance the feature representation ability of each intelligent agent, construct a unified feature space, and complete heterogeneous collaborative perception 3D target detection based on lidar.

9. A laser radar-based heterogeneous collaborative perception 3D target detection device, the laser radar-based heterogeneous collaborative perception 3D target detection device comprising: processor; A memory having computer-readable instructions stored thereon, wherein when the computer-readable instructions are executed by the processor, any one of the heterogeneous collaborative perception 3D target detection methods based on laser radar as described in any one of claims 1 to 7 is implemented.

10. A computer-readable storage medium, wherein at least one instruction is stored in the storage medium, and the at least one instruction is loaded and executed by a processor to implement any one of the heterogeneous collaborative perception 3D target detection methods based on lidar as described in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Object sensing method and device, vehicle, electronic equipment and storage medium

    CN115019283A

  • Heterogeneous graph network-based multi-modal cooperative detection method and system

    CN115512319A

  • Multi-source data collaboration and fusion perception method for network automatic driving

    CN117237772A

  • Extensible cooperation method and system supporting modal model isomerism

    CN117315396A

  • Intelligent network connection automobile cooperative perception and data fusion method

    CN117956507A

Cited By

  • Target ranging method, device and equipment integrating monocular camera and high-precision map

    CN121702338A