Collaborative Target Detection Method and System Based on Collaboration Graph Fusion
By filtering blind spots and fusing local features in a collaborative graph, the method addresses resource inefficiencies in existing collaborative object detection, enhancing detection accuracy and reducing computational load.
Patent Information
- Application Number
- CN202210485437.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-05-06
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2042-05-06
AI Technical Summary
The existing graph-based collaborative object detection method has the problem of transmission global features leading to large computing resource occupation, information redundancy and overlapping area weights, which affects the detection accuracy.
Coarse-grained blind spot screening and fine-grained local feature collaborative map fusion method are adopted. By selecting the detection blind spots of the central vehicle and spreading local features, the attention map is used to fuse local features, reducing computing resource consumption and improving detection accuracy.
It effectively improves detection accuracy, reduces communication resource overhead, alleviates the pressure on computing resources, and achieves more accurate blind spot collaborative detection.
Smart Images

Figure CN114913495B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of autonomous driving, and particularly relates to a collaborative object detection method and system based on collaborative graph fusion. Background Art
[0002] Object detection is a fundamental task in computer vision, aiming to identify the position and category of objects in space. According to the dimension of the predicted object, object detection methods can be divided into 2D object detection and 3D object detection, and most of what is required in the field of autonomous driving is 3D object detection. According to whether candidate boxes are generated, object detection methods can be divided into single-stage object detection and two-stage object detection. Single-stage object detection directly predicts the position and category of an object, characterized by a simple model and short time consumption, but low accuracy; two-stage object detection first generates a series of candidate boxes and predicts their confidence levels, and then optimizes the final position based on these candidate boxes, characterized by a larger model and longer time consumption, but high accuracy.
[0003] Object detection is an important research direction in the field of autonomous driving vision. Vehicles in the field of autonomous driving are also called agents in the autonomous driving scenario. Traditional object detection is all single-agent object detection based on in-vehicle sensors. However, due to the occlusion of objects and the limitations of in-vehicle sensors themselves, single-vehicle detection has blind spots and often cannot achieve good detection results. To address the challenges faced by single-vehicle object detection, collaborative object detection has emerged. Collaborative object detection is a detection method based on multi-agent information fusion, which is achieved by inserting a multi-agent collaboration module into a traditional object detection framework. In the autonomous driving scenario, there are multiple vehicles on the road. The blind spot of one vehicle may be in the detection area of other vehicles. By transmitting the object information observed by other vehicles to the central vehicle, the central vehicle can obtain a more comprehensive view and thus complete more accurate object detection. During the collaborative object detection process, each vehicle can be either the central vehicle or a neighbor vehicle of other vehicles.
[0004] Collaborative Object Detection is a key vision technology in the field of autonomous driving. It refers to assisting a single agent to complete a more accurate object detection task through information exchange and data fusion among multiple agents in a scene, thereby alleviating problems such as object occlusion and abnormal sensor capture in autonomous driving scenarios. Collaborative object detection methods can be discussed from two perspectives: the collaboration stage and the fusion strategy. The collaboration stage refers to at which stage of object detection the collaboration module is inserted. According to the different collaboration stages, collaborative object detection methods can be divided into three categories: data-level collaboration, feature-level collaboration, and decision-level collaboration. Among them, data-level collaboration refers to fusing the original observation data of vehicles, feature-level collaboration refers to fusing the object features of vehicles, and decision-level collaboration refers to fusing the final detection data of vehicles. The fusion strategy refers to the specific fusion calculation process of the collaboration module, which can be divided into simple fusion, feature-based fusion, and graph-based fusion. Simple fusion adopts strategies such as taking the mean, maximum value, and concatenation. Feature-based fusion selects the vehicle with the greatest relevance. Graph-based fusion refers to constructing the multi-vehicle collaboration process into a graph and fusing the information of multiple vehicles through the process of graph learning.
[0005] Existing graph-based collaborative object detection methods mainly include V2VNet and DiscoNet. V2VNet adopts a spatial-aware Graph Neural Network (GNN) to complete multi-vehicle information fusion. V2VNet first compensates for the transmission delays of different vehicles, and then uses GNN to aggregate the features of surrounding vehicles to the central vehicle and determine the vehicles within the neighborhood range according to the global position. This method effectively expands the vehicle's field of view, thereby detecting occluded objects. DiscoNet also adopts a Graph Attention Networks (GAT) to achieve multi-vehicle collaboration. Different from V2VNet, the edges of the fusion graph in DiscoNet are not scalars but a matrix, which can reflect the contribution degree of each pixel feature. In addition, DiscoNet introduces a teacher-student network. The teacher network is a data-level collaborative object detection, and the student network is a feature-level collaborative object detection. The features of the teacher network are used as the supervision of the student network to improve the performance of feature-level collaborative object detection.
[0006] The problem with current mainstream graph-based collaborative object detection methods lies in directly propagating global features, that is, the object features from the global perspectives of adjacent vehicles. Since the central vehicle can itself accurately detect part of the perspective area and does not require all the perspective information of adjacent vehicles. The transmitted global features are characterized by large volume and information redundancy, which not only consume a large amount of computing resources but also increase the weights of overlapping parts, making the network unable to focus more attention on the areas that need collaboration. Summary of the Invention
[0007] The object of the present invention is to provide a collaborative object detection method and system based on collaborative graph fusion, which improves the detection performance of the central vehicle through the cooperation of local features by screening the blind areas of the central vehicle at a coarse granularity and fusing the local features of the collaborative graph at a fine granularity, so as to solve at least one of the technical problems existing in the above-mentioned background technology.
[0008] In order to achieve the above object, the present invention adopts the following technical solutions:
[0009] On the one hand, the present invention provides a collaborative object detection method based on collaborative graph fusion, including:
[0010] Obtain the point cloud data of the target to be detected, and generate a two-dimensional bird's-eye view and candidate region boxes;
[0011] Select the detection blind area of the central vehicle of the candidate region box based on a coarse-grained method, and screen the local features of the two-dimensional bird's-eye view of the neighboring vehicles according to the detection blind area;
[0012] Fuse the local features of the two-dimensional bird's-eye view of the neighboring vehicles based on a fine-grained method using collaborative graph fusion to obtain new collaborative features of the central vehicle;
[0013] Based on the new collaborative features of the central vehicle, perform classification and regression prediction on each candidate region, and after threshold screening, obtain the final detection result.
[0014] Preferably, obtaining the point cloud data of the target to be detected and generating a two-dimensional bird's-eye view and candidate region boxes includes:
[0015] In the autonomous driving scenario, for the point cloud data of each vehicle target, use a feature extractor to extract the three-dimensional point cloud data and convert it into two-dimensional bird's-eye view features as global features;
[0016] Input the two-dimensional bird's-eye view of each vehicle into a 3D region generation network to generate 3D candidate region boxes for the corresponding vehicle;
[0017] After obtaining the candidate region boxes of the vehicle, then pass through the 3D region of interest pooling layer to obtain the two-dimensional bird's-eye view features of each candidate region box as local features.
[0018] Preferably, each 3D candidate region box has a corresponding classification confidence, and the classification confidence represents the probability that the corresponding candidate box belongs to each category and the background class. When the probability that the candidate region box belongs to the foreground is less than a preset threshold, the candidate region box belongs to the detection blind area of the current vehicle.
[0019] Preferably, selecting the detection blind area of the central vehicle of the candidate region box based on a coarse-grained method and screening the local features of the two-dimensional bird's-eye view of the neighboring vehicles according to the detection blind area includes:
[0020] Select neighboring vehicles within a preset range around the central vehicle as collaborative targets; for each candidate region box of the central vehicle, determine whether it is a blind spot; obtain a series of candidate region boxes and their local features of the central vehicle and its collaborative vehicles.
[0021] Use the intersection over union (IoU) to select the regions of neighboring vehicles that can collaborate with the blind spots of the central vehicle; for each blind spot candidate box of the central vehicle, traverse the candidate region boxes of the neighboring vehicles. If the IoU between the blind spot candidate box and the neighboring candidate box is greater than the threshold, it is very likely that the neighboring candidate box and the blind spot candidate box of the central vehicle represent the same region, and their collaboration can enhance the central vehicle's recognition ability for this region.
[0022] Preferably, determining whether each candidate region box of the central vehicle is a blind spot includes: if the confidence distribution of the candidate region box is significantly different and the confidence in a certain category is greater than the preset confidence threshold, it indicates that the central vehicle can clearly detect that the target belongs to the background or a specific category, and this candidate box has significance; on the contrary, if the confidence of the candidate region box is less than the preset threshold and the vehicle cannot determine the category to which the target belongs, then the region where this candidate box is located is the blind spot of the central vehicle, which is the region that the central vehicle needs to collaborate with, and is added to the blind spot set of the central vehicle.
[0023] Preferably, using a fine-grained method to fuse the local features of the two-dimensional bird's-eye view of neighboring vehicles with a collaborative graph to obtain new collaborative features of the central vehicle, including:
[0024] Fuse local features using an attention map-based method; traverse each blind spot candidate box of the central vehicle, construct an attention map for the candidate box and its neighboring collaborative boxes, and update the local features of this blind spot candidate box; construct an attention map for the blind spot candidate box and its neighboring collaborative boxes, where the nodes of the map are the BEV local features of the blind spot candidate box and the collaborative neighboring candidate boxes, the direction is from each collaborative neighboring candidate box to the blind spot candidate box, and from the blind spot candidate box to itself; after obtaining the weights of each edge, use an aggregation function to update the local features of the blind spots of the central vehicle.
[0025] In a second aspect, the present invention provides a collaborative target detection system based on collaborative graph fusion, including:
[0026] An acquisition module, configured to acquire point cloud data of a target to be detected, generate a two-dimensional bird's-eye view and candidate region boxes.
[0027] A screening module, configured to select the detection blind spots of the candidate region boxes of the central vehicle based on a coarse-grained method, and screen the local features of the two-dimensional bird's-eye view of neighboring vehicles according to the detection blind spots.
[0028] A cooperation module, which is used to fuse the local features of the two-dimensional bird's-eye view of neighboring vehicles based on a fine-grained method using a cooperation graph to obtain new cooperation features of the central vehicle;
[0029] A detection module, which is used to perform classification and regression prediction on each candidate region based on the new cooperation features of the central vehicle, and obtain the final detection result after threshold screening.
[0030] In a third aspect, the present invention provides a computer device, including a memory and a processor, the processor and the memory communicate with each other, the memory stores program instructions executable by the processor, and the processor calls the program instructions to execute the above-mentioned collaborative target detection method based on cooperation graph fusion.
[0031] In a fourth aspect, the present invention provides an electronic device, including a memory and a processor, the processor and the memory communicate with each other, the memory stores program instructions executable by the processor, and the processor calls the program instructions to execute the above-mentioned collaborative target detection method based on cooperation graph fusion.
[0032] In a fifth aspect, the present invention provides a computer-readable storage medium, which stores a computer program, and when the computer program is executed by a processor, it implements the above-mentioned collaborative target detection method based on cooperation graph fusion.
[0033] Advantages of the present invention: Consider the cooperation of local features from both coarse-grained and fine-grained perspectives for the first time; through the transmission of local features, collaborative detection can relieve the pressure on computing resources, more accurately cooperate on the blind area of the central vehicle, and effectively improve the performance of collaborative detection; while effectively improving the detection accuracy, it reduces the overhead of communication resources.
[0034] Additional aspects and advantages of the present invention will be given in part in the following description, and these will become obvious from the following description, or can be understood through the practice of the present invention. Description of the Drawings
[0035] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required for the description of the embodiments will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0036] Figure 1 It is a flowchart of the collaborative target detection method based on coarse-to-fine cooperation graph fusion described in the embodiments of the present invention.
[0037] Figure 2This is the collaborative target detection framework diagram based on the coarse-to-fine collaborative graph fusion described in the embodiments of the present invention.
[0038] Figure 3 This is the flowchart of the coarse-grained collaborative work described in the embodiments of the present invention.
[0039] Figure 4 This is the flowchart of the fine-grained collaborative work described in the embodiments of the present invention. Detailed implementation manners
[0040] The following details the implementation manners of the present invention. The examples of the implementation manners are shown in the drawings, where the same or similar reference numerals indicate the same or similar elements or elements with the same or similar functions throughout. The implementation manners described below with reference to the drawings are exemplary and are only used to explain the present invention, and should not be construed as a limitation to the present invention.
[0041] Those skilled in the art of the present technology can understand that, unless otherwise defined, all terms (including technical terms and scientific terms) used here have the same meaning as the general understanding of those of ordinary skill in the art in the field to which the present invention belongs.
[0042] It should also be understood that terms such as those defined in a general dictionary should be understood to have a meaning consistent with the meaning in the context of the prior art, and will not be interpreted with an idealized or overly formal meaning unless defined as here.
[0043] Those skilled in the art of the present technology can understand that, unless specifically stated, the singular forms "a", "an", "the" and "said" used here may also include the plural forms. It should be further understood that the term "including" used in the description of the present invention means the presence of the described features, integers, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements and / or their groups.
[0044] In the description of this specification, the description with reference to terms such as "one embodiment", "some embodiments", "example", "specific example", or "some examples" means that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. Moreover, the specific features, structures, materials, or characteristics described may be combined in a suitable manner in any one or more embodiments or examples. In addition, without contradiction, those skilled in the art can combine and combine the different embodiments or examples described in this specification and the features of the different embodiments or examples.
[0045] To facilitate the understanding of the present invention, the following further explains the present invention with specific examples in conjunction with the drawings, and the specific examples do not constitute a limitation to the embodiments of the present invention.
[0046] Those skilled in the art should understand that the accompanying drawings are only schematic diagrams of the embodiments, and the components in the drawings are not necessarily essential for implementing the present invention.
[0047] Embodiment 1
[0048] The collaborative transmission of raw data at the data level brings excessive bandwidth pressure, and some target information has been lost in the detection results of the decision-level collaboration. In order to maintain the balance between accuracy and bandwidth, Embodiment 1 of the present invention selects graph-based feature-level collaborative object detection.
[0049] First, a collaborative object detection system based on collaborative graph fusion is provided, including:
[0050] An acquisition module, configured to acquire the point cloud data of the target to be detected, and generate a two-dimensional bird's-eye view and candidate region boxes;
[0051] A screening module, configured to select the detection blind area of the vehicle at the center of the candidate region box based on a coarse-grained method, and screen the local features of the two-dimensional bird's-eye view of the neighboring vehicles according to the detection blind area;
[0052] A collaboration module, configured to fuse the local features of the two-dimensional bird's-eye view of the neighboring vehicles using collaborative graph fusion based on a fine-grained method to obtain new collaborative features of the center vehicle;
[0053] A detection module, configured to perform classification and regression prediction on each candidate region based on the new collaborative features of the center vehicle, and obtain the final detection result after threshold screening.
[0054] Secondly, in this embodiment, using the above system, a collaborative object detection method based on collaborative graph fusion is implemented, including:
[0055] Acquire the point cloud data of the target to be detected, and generate a two-dimensional bird's-eye view and candidate region boxes;
[0056] Select the detection blind area of the vehicle at the center of the candidate region box based on a coarse-grained method, and screen the local features of the two-dimensional bird's-eye view of the neighboring vehicles according to the detection blind area;
[0057] Fuse the local features of the two-dimensional bird's-eye view of the neighboring vehicles using collaborative graph fusion based on a fine-grained method to obtain new collaborative features of the center vehicle;
[0058] Perform classification and regression prediction on each candidate region based on the new collaborative features of the center vehicle, and obtain the final detection result after threshold screening.
[0059] Acquiring the point cloud data of the target to be detected and generating a two-dimensional bird's-eye view and candidate region boxes includes:
[0060] In the autonomous driving scenario, for the point cloud data of each vehicle target, a feature extractor is used to extract the three-dimensional point cloud data, which is then converted into two-dimensional bird's-eye view features as global features.
[0061] The two-dimensional bird's-eye view of each vehicle is input into a 3D region generation network to generate 3D candidate region boxes for the corresponding vehicle.
[0062] After obtaining the candidate region boxes of the vehicle, the two-dimensional bird's-eye view features of each candidate region box are obtained through a 3D region of interest pooling layer as local features.
[0063] Among them, each 3D candidate region box has a corresponding classification confidence. The classification confidence represents the probability that the corresponding candidate box belongs to each category and the background class. When the probability that the candidate region box belongs to the foreground is less than a preset threshold, the candidate region box belongs to the detection blind area of the current vehicle.
[0064] Based on a coarse-grained method, the detection blind areas of the vehicles centered on the candidate region boxes are selected, and the two-dimensional bird's-eye view local features of neighboring vehicles are filtered according to the detection blind areas, including:
[0065] Select neighboring vehicles within a preset range around the central vehicle as cooperative targets; determine whether each candidate region box of the central vehicle is a blind area; obtain a series of candidate region boxes and their local features of the central vehicle and its cooperative vehicles.
[0066] Use the intersection over union (IoU) to select neighboring vehicle regions that can cooperate with the blind area of the central vehicle. For each blind area candidate box of the central vehicle, traverse the candidate region boxes of neighboring vehicles. If the IoU between the blind area candidate box and the neighboring candidate box is greater than the threshold, it is very likely that the neighboring candidate box and the blind area candidate box of the central vehicle represent the same region, and their cooperation can enhance the central vehicle's recognition ability of this region.
[0067] Among them, determining whether each candidate region box of the central vehicle is a blind area includes: if the confidence distribution of the candidate region box is significantly different and the confidence in a certain category is greater than the preset confidence threshold, it means that the central vehicle can clearly detect that the target belongs to the background or a specific category, and the candidate box is significant; on the contrary, if the confidence of the candidate region box is less than the preset threshold and the vehicle cannot determine the category to which the target belongs, the area where the candidate box is located is the blind area of the central vehicle, which is the area that the central vehicle needs to cooperate with and is added to the blind area set of the central vehicle.
[0068] Based on a fine-grained method, the two-dimensional bird's-eye view local features of neighboring vehicles are fused using a collaborative graph to obtain new collaborative features of the central vehicle, including:
[0069] Fuse local features using an attention map-based approach; traverse each blind spot candidate box of the central vehicle, construct an attention map for the candidate box and its neighboring collaborative boxes, and update the local features of the blind spot candidate box; construct an attention map for the blind spot candidate box and its neighboring collaborative boxes, where the nodes of the map are the BEV local features of the blind spot candidate box and the collaborative neighboring candidate boxes, the direction is from each collaborative neighboring candidate box to the blind spot candidate box, and from the blind spot candidate box to itself; after obtaining the weights of each edge, use an aggregation function to update the local features of the blind spot of the central vehicle.
[0070] In summary, in this Embodiment 1, for the first time, an attempt is made to consider the collaboration of local features from both coarse-grained and fine-grained perspectives. Coarse-grained collaboration is to judge the area that the central vehicle needs to collaborate through information such as detection confidence, select and transmit the local features of the collaborative vehicle, and fine-grained collaboration is to assign weights to the local features of the collaborative area through a graph fusion method and update the local features of the blind spot. By transmitting local features, collaborative detection can relieve the pressure on computing resources, collaborate more precisely on the blind spot of the central vehicle, and effectively improve the performance of collaborative detection. The two-stage detection model is adopted, which can effectively improve the accuracy of the detection model while reducing the communication resource overhead.
[0071] Embodiment 2
[0072] As Figures 1 to 4 shown, in this Embodiment 2, a collaborative object detection method based on coarse-to-fine collaborative graph fusion is proposed. This method improves the detection performance of the central vehicle through local feature collaboration by screening the blind spot of the central vehicle at a coarse grain and fusing the local feature collaborative graph at a fine grain.
[0073] In this Embodiment 2, based on the multi-agent collaborative object detection model, the collaborative object detection task is divided into four steps, as Figure 1 shown. The first step is to generate candidate boxes and their local features based on the 3D region generation network, the second step is to screen the blind spot of the central vehicle at a coarse grain and transmit collaborative features, the third step is to fuse the local feature collaborative graph at a fine grain, and the fourth step is to perform object detection based on the collaborative features to obtain the final detection result.
[0074] The flowchart of a collaborative object detection method based on coarse-to-fine collaborative graph fusion provided by an embodiment of the present invention is as Figure 1 shown, including the following processing steps:
[0075] S10, generate BEV features and candidate region boxes through CNN based on the point cloud data to be subjected to object detection:
[0076] In this second embodiment, the task faced is two-stage 3D point cloud object detection. In the autonomous driving scenario, assuming there are a total of C categories of objects and n agents (vehicles), for each agent A i the point cloud data X i (i = 1, 2, 3,..., n), using the feature extractor to convert the three-dimensional point cloud data into two-dimensional Bird's Eye View (BEV) features F i ∈R h×w×k , where h, w, and k respectively represent the height, width, and number of channels of the BEV feature. This feature is the BEV global feature of the vehicle.
[0077] Next, input the BEV feature F of each agent i into the 3D Region Proposal Network (RPN) to generate the 3D candidate region boxes of the corresponding agent where N i represents the number of candidate region boxes of agent A i . Each 3D candidate region box P ij generated by the RPN has a corresponding classification confidence S ij ∈R (C+1) . This confidence represents the probability that the candidate box belongs to each category and the background class. When the probability that the candidate region box belongs to the foreground is very small, it is difficult to determine the object contained in the candidate region box, which belongs to the blind area of the current agent.
[0078] After obtaining the candidate region boxes of the vehicle using the 3D RPN module, the BEV feature f of each candidate region box P ij is obtained through the 3D ROI pooling layer ij ∈R m×m×d , where m and d respectively refer to the feature size and dimension of the candidate region box. The BEV feature of the candidate region box is a local feature.
[0079] S20. Based on the coarse-grained method, select the detection blind area of the central vehicle, and filter the BEV local features of neighboring agents according to the detection blind area.
[0080] Feature-level collaborative object detection requires transmitting the features of agents. Traditional feature-level collaborative detection transmits and fuses global features, consuming a large amount of computing resources. The method for screening blind areas based on coarse-grained provided in this embodiment is as Figure 3 shown and includes the following processing procedures:
[0081] In collaborative object detection, not all agents can be used for collaboration. Select the neighboring vehicles within the range D around the central vehicle as collaborative targets.
[0082] For the central vehicle Ac , for each candidate region box P of (c = 1, 2,..., n) cj Determine whether it is a blind area: If the confidence distribution of the candidate region box is significantly different and the confidence in a certain category is greater than the threshold T cs , it indicates that the vehicle can clearly detect that the target belongs to the background or a specific category, and this candidate box has significance; on the contrary, if the confidence distribution of the candidate region box is relatively flat and the vehicle cannot determine the category to which the target belongs, then the area where this candidate box is located is the blind area of the central vehicle, which is the area that the central vehicle needs to cooperate with, and is added to the blind area set P of the central vehicle cN .
[0083] After the above steps, the central vehicle and its cooperative vehicles both obtain a series of candidate region boxes and their local features. To relieve the resource consumption pressure of traditional cooperative detection for propagating and fusing global features, in this embodiment, a creative approach is proposed to propagate and fuse local features, that is, only the features of the regions that the central vehicle needs to cooperate with are propagated
[0084] In this embodiment, the Intersection over Union (IoU) is used to select the neighboring vehicle regions that can cooperate with the blind areas of the central vehicle. For each blind area candidate box of the central vehicle Traverse the neighboring vehicle A i 's candidate region box P i , if the IOU between the blind area candidate box and the neighboring candidate box is greater than the threshold T iou , as shown in formula (1), then this neighboring candidate box and the blind area candidate box of the central vehicle are very likely to represent the same region, and their cooperation can enhance the central vehicle's recognition ability of this region
[0085]
[0086] For each blind area candidate box of the central vehicle Initialize a collaborative region local feature set S i , which is used to store the BEV local features of the blind area candidate box and the local BEV features of the collaborative region. Transmit the BEV features of the neighboring candidate boxes that satisfy formula (1) to the central vehicle. Since each vehicle has its own observation, the features transmitted by the cooperative vehicle need to be converted to the angle of the central vehicle and then put into the collaborative region local feature set S i . During cooperation, only the features in the collaborative region local feature set S i need to be fused
[0087] S30, based on a fine-grained method, fuse the local features of neighboring agents using a collaborative graph, and the central vehicle obtains new collaborative features
[0088] The present invention uses an attention map-based method to fuse local features. Each blind spot candidate box of the central vehicle is traversed An attention map is constructed for the candidate box and its neighboring collaborative boxes, and the local features of the blind spot candidate box are updated. The process of constructing the attention map and the process of updating the local features are as follows Figure 2 shown
[0089] Construct an attention map G for the blind spot candidate box and its neighboring collaborative boxes i , where the nodes of the graph are the BEV local features h of the blind spot candidate box and the collaborative neighboring candidate boxes j→i ∈R m×m×d , the direction is from each collaborative neighboring candidate box to the blind spot candidate box (j→i), and from the blind spot candidate box to itself (i→j). The edges of the graph take the form of a matrix, and the weight W of each edge j→i is calculated based on formula (2),
[0090] W j→i =Π(h j , h i )∈R m×m #(2)
[0091] where Π concatenates the features of adjacent nodes and uses a 1×1 convolutional layer to reduce the number of channels of the edge features from d to 1. Through this calculation, the matrix weight of the edge can be obtained. In addition, the weight of each edge is input into a Softmax layer for regularization, and the matrix weight of the edge can reflect the spatial weight of the BEV local features
[0092] After obtaining the weight of each edge, the present invention uses an aggregation function to update the local features of the blind spot of the central vehicle. The aggregation function is shown in formula (3),
[0093]
[0094] where ⊙ denotes point multiplication in the channel direction, M is the number of features in the local feature set S of the collaborative region of the blind spot candidate box i , and H i is the updated BEV feature of the blind spot candidate box
[0095] S40. The central vehicle performs object detection based on the collaborative features
[0096] After the fusion of the collaborative graph, the features of the blind spot candidate box of the central vehicle contain more sufficient information. The above is the first stage of the two-stage 3D object detection
[0097] In the second stage, the candidate boxes are further adjusted and recognized. The features of all candidate boxes of the central vehicle are input into the head of the detection model, and classification and regression predictions are made for each candidate box. After threshold screening, more accurate detection results can be obtained
[0098] It should be noted that in the scenario of collaborative target detection, each vehicle can be the central vehicle or the collaborative vehicle of other vehicles. Therefore, the entire collaborative detection process is parallel.
[0099] In summary, in this Embodiment 2, the proposed collaborative target detection method based on coarse-to-fine collaborative graph fusion, based on a two-stage 3D target detection framework, greatly saves the resource consumption of collaboration and achieves accurate target detection through the selection and propagation of local features in the collaborative area at a coarse granularity and the local feature fusion based on the collaborative graph at a fine granularity. Considering collaboration from two perspectives of coarse and fine granularities, propagating and fusing the local features of other intelligent vehicles, where the coarse-grained collaboration selects the area of neighboring vehicles that need collaboration based on the confidence of the blind area of the central vehicle, and the fine-grained collaboration is based on the collaborative graph fusion to learn the weights of the collaborative area and fuse them, thus solving the problem of resource consumption of global features.
[0100] Embodiment 3
[0101] Embodiment 3 of the present invention provides an electronic device, including a memory and a processor, where the processor and the memory communicate with each other, the memory stores program instructions executable by the processor, and the processor calls the program instructions to execute the collaborative target detection method based on collaborative graph fusion, and this method includes the following process steps:
[0102] Obtain the point cloud data of the target to be detected, and generate a two-dimensional bird's-eye view and candidate region boxes;
[0103] Select the detection blind area of the central vehicle of the candidate region boxes based on a coarse-grained method, and screen the local features of the two-dimensional bird's-eye view of neighboring vehicles according to the detection blind area;
[0104] Fuse the local features of the two-dimensional bird's-eye view of neighboring vehicles using collaborative graph fusion based on a fine-grained method to obtain new collaborative features of the central vehicle;
[0105] Based on the new collaborative features of the central vehicle, perform classification and regression predictions on each candidate region, and obtain the final detection result after threshold screening.
[0106] Embodiment 4
[0107] Embodiment 4 of the present invention provides a computer-readable storage medium, which stores a computer program, and when the computer program is executed by a processor, it implements the collaborative target detection method based on collaborative graph fusion, and this method includes the following process steps:
[0108] Obtain the point cloud data of the target to be detected, and generate a two-dimensional bird's-eye view and candidate region boxes;
[0109] Select the detection blind area of the vehicle at the center of the candidate region box based on a coarse-grained method, and screen the local features of the two-dimensional bird's-eye view of neighboring vehicles according to the detection blind area;
[0110] Based on a fine-grained method, fuse the local features of the two-dimensional bird's-eye view of neighboring vehicles using a collaborative graph to obtain new collaborative features of the central vehicle;
[0111] Based on the new collaborative features of the central vehicle, perform classification and regression prediction on each candidate region, and obtain the final detection result after threshold screening.
[0112] Embodiment 5
[0113] Embodiment 5 of the present invention provides a computer device, including a memory and a processor, the processor and the memory communicate with each other, the memory stores program instructions executable by the processor, and the processor calls the program instructions to execute a collaborative object detection method based on collaborative graph fusion. The method includes the following steps:
[0114] Obtain the point cloud data of the target to be detected, and generate a two-dimensional bird's-eye view and candidate region boxes;
[0115] Select the detection blind area of the vehicle at the center of the candidate region box based on a coarse-grained method, and screen the local features of the two-dimensional bird's-eye view of neighboring vehicles according to the detection blind area;
[0116] Based on a fine-grained method, fuse the local features of the two-dimensional bird's-eye view of neighboring vehicles using a collaborative graph to obtain new collaborative features of the central vehicle;
[0117] Based on the new collaborative features of the central vehicle, perform classification and regression prediction on each candidate region, and obtain the final detection result after threshold screening.
[0118] In summary, the collaborative object detection method and system based on collaborative graph fusion described in the embodiments of the present invention first attempt to consider the collaboration of local features from two perspectives: coarse-grained and fine-grained. Coarse-grained collaboration is to judge the blind area of the central vehicle through detection confidence, and screen and propagate the collaborative region features of neighboring vehicles. Fine-grained collaboration is to assign weights to the collaborative region through the method of fusing graphs and update the features of the blind area of the central vehicle. By transmitting local features, the pressure on computing resources can be alleviated, and the blind area of the central vehicle can be collaborated more precisely, thereby improving the collaborative detection performance. Aiming at the problem of low accuracy of the existing single-stage 3D object detection model, a two-stage 3D object detection model is adopted. The two-stage 3D detection model has higher accuracy than the single-stage 3D object detection. At the same time, the proposed local feature extraction and collaboration method reduces the computing time and resources of 3D object detection, thereby further exerting the accuracy advantage of the two-stage 3D object detection.
[0119] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a system, or a computer program product. Therefore, the present invention can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memory, CD-ROM, optical memory, etc.) that contain computer-usable program code.
[0120] The present invention is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of the present invention. It should be understood that each flow and / or block in the flowchart and / or block diagram, and the combination of flows and / or blocks in the flowchart and / or block diagram, can be realized by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate a device for realizing the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0121] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing devices to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured product including an instruction device, and the instruction device realizes the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0122] These computer program instructions can also be loaded onto a computer or other programmable data processing devices, and a series of operation steps are executed on the computer or other programmable devices to generate a computer-implemented process. Therefore, the instructions executed on the computer or other programmable devices provide steps for realizing the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0123] Although the specific implementation manners of the present invention are described above in conjunction with the accompanying drawings, it is not a limitation to the protection scope of the present invention. Those skilled in the art should understand that based on the technical solutions disclosed in the present invention, various modifications or deformations that can be made by those skilled in the art without creative efforts should be covered within the protection scope of the present invention.
Claims
1. A collaborative object detection method based on collaborative graph fusion, characterized in that Including: Obtain the point cloud data of the target to be detected, and generate a two-dimensional bird's-eye view and candidate region boxes; Select the detection blind area of the vehicle at the center of the candidate region box based on a coarse-grained method, and filter the local features of the two-dimensional bird's-eye view of neighboring vehicles according to the detection blind area; Fuse the local features of the two-dimensional bird's-eye view of neighboring vehicles using a collaborative graph based on a fine-grained method to obtain new collaborative features of the central vehicle; Based on the new collaborative features of the central vehicle, perform classification and regression prediction on each candidate region, and after threshold screening, obtain the final detection result; Among them, selecting the detection blind area of the vehicle at the center of the candidate region box based on a coarse-grained method and filtering the local features of the two-dimensional bird's-eye view of neighboring vehicles according to the detection blind area includes: Select neighboring vehicles within a preset range around the central vehicle as collaborative targets; judge whether each candidate region box of the central vehicle is a blind area; obtain a series of candidate region boxes and their local features of the central vehicle and its collaborative vehicles; Use the intersection over union (IoU) to select the regions of neighboring vehicles that can collaborate with the blind area of the central vehicle; for each blind area candidate box of the central vehicle, traverse the candidate region boxes of neighboring vehicles. If the IoU between the blind area candidate box and the neighboring candidate box is greater than the threshold, then the neighboring candidate box and the blind area candidate box of the central vehicle are very likely to represent the same region, and their collaboration can enhance the central vehicle's recognition ability of this region; Fusing the local features of the two-dimensional bird's-eye view of neighboring vehicles using a collaborative graph based on a fine-grained method to obtain new collaborative features of the central vehicle includes: Fuse local features using an attention map-based method; traverse each blind area candidate box of the central vehicle, construct an attention map for the candidate box and its neighboring collaborative boxes, and update the local features of this blind area candidate box; construct an attention map for the blind area candidate box and its neighboring collaborative boxes, where the nodes of the graph are the BEV local features of the blind area candidate box and the collaborative neighboring candidate boxes, the direction is from each collaborative neighboring candidate box to the blind area candidate box, and from the blind area candidate box to itself; after obtaining the weights of each edge, use an aggregation function to update the local features of the blind area of the central vehicle.
2. The collaborative target detection method based on collaborative graph fusion according to claim 1, wherein, Obtain the point cloud data of the target to be detected, and generate a two-dimensional bird's-eye view and candidate region boxes, including: In an autonomous driving scenario, for the point cloud data of each vehicle target, use a feature extractor to extract three-dimensional point cloud data and convert it into two-dimensional bird's-eye view features as global features; Input the two-dimensional bird's-eye view of each vehicle into a 3D region generation network to generate 3D candidate region boxes for the corresponding vehicle; After obtaining the candidate region boxes of the vehicle, pass through a 3D region of interest pooling layer to obtain the two-dimensional bird's-eye view features of each candidate region box as local features.
3. The collaborative target detection method based on collaborative graph fusion according to claim 2, wherein, Each 3D candidate region box has a corresponding classification confidence. The classification confidence represents the probability that the corresponding candidate box belongs to each category and the background category. When the probability that the candidate region box belongs to the foreground is less than the preset threshold, then the candidate region box belongs to the detection blind area of the current vehicle.
4. The collaborative target detection method based on collaborative graph fusion according to claim 1, characterized in that Determine whether each candidate region box of the central vehicle is a blind area, including: If the confidence distribution of the candidate region box is significantly different and the confidence in a certain category is greater than the preset confidence threshold, it indicates that the central vehicle can clearly detect that the target belongs to the background or a specific category, and the candidate box has significance; on the contrary, if the confidence of the candidate region box is less than the preset threshold and the vehicle cannot determine the category to which the target belongs, the area where the candidate box is located is the blind area of the central vehicle, which is the area that the central vehicle needs to cooperate with, and is added to the blind area set of the central vehicle.
5. A collaborative target detection system based on collaborative graph fusion, characterized in that, Including: An acquisition module, configured to acquire point cloud data of a target to be detected and generate a two-dimensional bird's-eye view and candidate region boxes; A screening module, configured to select the detection blind area of the candidate region box of the central vehicle based on a coarse-grained method and screen the local features of the two-dimensional bird's-eye view of neighboring vehicles according to the detection blind area; A cooperation module, configured to fuse the local features of the two-dimensional bird's-eye view of neighboring vehicles using a cooperation graph based on a fine-grained method to obtain new cooperation features of the central vehicle; A detection module, configured to perform classification and regression prediction on each candidate region based on the new cooperation features of the central vehicle, and obtain a final detection result after threshold screening; Among them, selecting the detection blind area of the candidate region box of the central vehicle based on a coarse-grained method and screening the local features of the two-dimensional bird's-eye view of neighboring vehicles according to the detection blind area includes: Select neighboring vehicles within a preset range around the central vehicle as cooperation targets; determine whether each candidate region box of the central vehicle is a blind area; obtain a series of candidate region boxes and their local features of the central vehicle and its cooperative vehicles; Use the intersection over union to select the neighboring vehicle areas that can cooperate with the blind area of the central vehicle; for each blind area candidate box of the central vehicle, traverse the candidate region boxes of neighboring vehicles. If the IOU of the blind area candidate box and the neighboring candidate box is greater than the threshold, it is very likely that the neighboring candidate box and the blind area candidate box of the central vehicle represent the same area, and their cooperation can enhance the central vehicle's recognition ability of this area; Fusing the local features of the two-dimensional bird's-eye view of neighboring vehicles using a cooperation graph based on a fine-grained method to obtain new cooperation features of the central vehicle includes: Fusing local features using an attention map-based method; traverse each blind area candidate box of the central vehicle, construct an attention map for the candidate box and its neighboring cooperative boxes, and update the local features of the blind area candidate box; construct an attention map for the blind area candidate box and its neighboring cooperative boxes, where the nodes of the graph are the BEV local features of the blind area candidate box and the cooperative neighboring candidate boxes, the direction is from each cooperative neighboring candidate box to the blind area candidate box, and from the blind area candidate box to itself; after obtaining the weights of each edge, use an aggregation function to update the local features of the blind area of the central vehicle.
6. A computer device, including a memory and a processor, the processor and the memory communicate with each other, the memory stores program instructions executable by the processor, and the processor calls the program instructions to execute the cooperative target detection method based on cooperation graph fusion according to any one of claims 1-4.
7. An electronic device, characterized in that, It includes a memory and a processor, the processor and the memory communicate with each other, the memory stores program instructions executable by the processor, and the processor calls the program instructions to execute the collaborative target detection method based on collaborative graph fusion according to any one of claims 1-4.
8. A computer-readable storage medium, characterized in that, It stores a computer program, and when the computer program is executed by a processor, it implements the collaborative target detection method based on collaborative graph fusion according to any one of claims 1-4.
Citation Information
Patent Citations
Vehicle blind area early warning method and device, MEC platform and storage medium
CN113284366A
Multi-vehicle cooperative environment perception method based on semantic-level information fusion
CN114091598A