Target detection method and electronic equipment
By combining a multi-view object detection method that combines 2D and 3D detection, and utilizing a backbone network and a joint decoder, more accurate object detection is achieved, solving the problems of incomplete and inaccurate detection in traditional methods and improving the accuracy and comprehensiveness of detection results.
Patent Information
- Application Number
- CN202511100280.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-07
- Publication Date
- 2025-09-09
- Estimated Expiration
- 2045-08-07
AI Technical Summary
Existing object detection technologies in the field of assisted driving suffer from incomplete and inaccurate detection. In particular, traditional 3D detection methods only detect 3D boxes, and the utilization efficiency of 2D detectors is limited, the 2D detection results are rough, and the 3D detector is highly dependent on the 2D detector.
A multi-view target detection method that combines 2D and 3D is adopted. Through the combination of backbone network, joint decoder and detection head, multiple image feature information and 3D query feature information are utilized, combined with 2D and 3D detection heads to achieve the combination of 2D and 3D target detection and obtain more comprehensive and accurate detection results.
It achieves more precise target detection, improves the accuracy and comprehensiveness of detection results, and can detect 2D and 3D targets at the same time, improving detection accuracy.
Smart Images

Figure CN120612554A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of target detection, and in particular to a target detection method and electronic equipment. Background Art
[0002] Object detection is a core technology in computer vision. It aims to identify the category of objects in images or videos and accurately locate their positions. The target object is the object to be detected, and its location is typically represented by a detection box. With the continuous development of deep learning, object detection technology has made significant progress and is widely used in fields such as autonomous driving, security monitoring, medical image analysis, and industrial quality inspection.
[0003] Taking the field of assisted driving as an example, the more accurate the target detection results obtained by target detection methods, the more precise the driving control based on the detection results will be. Therefore, how to more accurately detect targets and obtain more accurate detection results becomes particularly important. Summary of the Invention
[0004] The embodiments of the present application provide a target detection method and electronic device, which can perform target detection more accurately, thereby obtaining more accurate detection results and improving the accuracy of the detection results.
[0005] To solve the above technical problems, in the first aspect, an embodiment of the present application provides a target detection method, which is applied to a target detection model. The target detection model includes a backbone network, a joint decoder and a detection head. The detection head includes a two-dimensional detection head and a three-dimensional detection head. The method includes: the backbone network obtains multiple image feature information based on multiple images corresponding to different perspectives, and inputs the multiple image feature information into the joint decoder; the joint decoder obtains target two-dimensional query feature information and target three-dimensional query feature information corresponding to each target to be detected based on the three-dimensional feature information of multiple targets to be detected corresponding to the multiple images and the multiple image feature information, inputs the target two-dimensional query feature information into the two-dimensional detection head, and inputs the target three-dimensional query feature information into the three-dimensional detection head; the two-dimensional detection head obtains two-dimensional detection results corresponding to multiple targets to be detected based on the target two-dimensional query feature information; the three-dimensional detection head obtains three-dimensional detection results corresponding to multiple targets to be detected based on the target three-dimensional query feature information.
[0006] Using the above technical solution, a joint decoder based on a target detection model obtains target two-dimensional query feature information and target three-dimensional query feature information of each target to be detected based on multiple image feature information and three-dimensional query feature information of multiple targets to be detected. A two-dimensional detection head based on the joint decoder obtains two-dimensional detection results of multiple targets to be detected based on the target two-dimensional query feature information, and a three-dimensional detection head based on the joint decoder obtains three-dimensional detection results of multiple targets to be detected based on the target three-dimensional query feature information of each target to be detected. In this way, the target detection model performs two-dimensional target detection of the target to be detected based on the three-dimensional query feature information, and also performs three-dimensional target detection of the target to be detected based on the three-dimensional query feature information. The combination of two-dimensional target detection and three-dimensional target detection achieves more accurate target detection, and the detection results obtained include both two-dimensional detection results and three-dimensional detection results. Therefore, more comprehensive and accurate detection results can be obtained, thereby improving the accuracy of the detection results.
[0007] In a possible implementation of the first aspect above, the three-dimensional feature information includes three-dimensional query feature information and three-dimensional attribute feature information. The joint decoder obtains target two-dimensional query feature information and target three-dimensional query feature information corresponding to each target to be detected based on the three-dimensional feature information of multiple targets to be detected corresponding to multiple images and multiple image feature information, including: the joint decoder obtains multiple groups of two-dimensional query feature information corresponding to each target to be detected as target two-dimensional query feature information based on multiple image feature information, corresponding three-dimensional query feature information and three-dimensional attribute feature information, and obtains second three-dimensional query feature information as target three-dimensional query feature information based on the multiple groups of two-dimensional query feature information.
[0008] In a possible implementation of the first aspect above, the joint decoder obtains multiple groups of two-dimensional query feature information corresponding to each target to be detected as target two-dimensional query feature information based on multiple image feature information, corresponding three-dimensional query feature information and three-dimensional attribute feature information, and obtains second three-dimensional query feature information as target three-dimensional query feature information based on the multiple groups of two-dimensional query feature information, including: the joint decoder initializes the three-dimensional query feature information to obtain fourth three-dimensional query feature information, clusters the three-dimensional attribute feature information to obtain second three-dimensional attribute feature information; determines camera parameters of a camera used to shoot the multiple images, and obtains conversion matrices corresponding to the multiple targets to be detected based on the camera parameters and the second three-dimensional attribute information. matrix; obtain multiple two-dimensional query feature information corresponding to multiple targets to be detected according to the fourth three-dimensional query feature information and the transformation matrix, group the multiple two-dimensional query feature information to obtain multiple groups of two-dimensional query feature information of multiple targets to be detected; obtain updated multiple groups of two-dimensional query feature information corresponding to the multiple targets to be detected according to the multiple groups of two-dimensional query feature information and the multiple image feature information, as target two-dimensional query feature information; obtain third three-dimensional query feature information corresponding to each target to be detected according to the updated multiple groups of two-dimensional query feature information and the transformation matrix; obtain second three-dimensional query feature information corresponding to each target to be detected according to the multiple image feature information and the third three-dimensional query feature information, as target three-dimensional query feature information.
[0009] In a possible implementation of the first aspect above, the joint decoder includes multiple decoding layers, and the backbone network inputs multiple image feature information into the joint decoder, including: the backbone network inputs multiple image feature information into each decoding layer respectively; the three-dimensional feature information includes three-dimensional query feature information and three-dimensional attribute feature information, then the joint decoder obtains target two-dimensional query feature information and target three-dimensional query feature information corresponding to each target to be detected based on the three-dimensional query feature information of multiple targets to be detected corresponding to the input multiple images and multiple image feature information, including: each decoding layer obtains multiple groups of two-dimensional query feature information corresponding to each target to be detected based on the multiple image feature information, the corresponding first three-dimensional query feature information and the first three-dimensional attribute feature information, obtains second three-dimensional query feature information based on the multiple groups of two-dimensional query feature information, and The second three-dimensional query feature information is input into the next-level decoding layer as the first three-dimensional query feature information corresponding to the next-level decoding layer, so that the next-level decoding layer performs corresponding processing until the final decoding layer completes processing, thereby obtaining target two-dimensional query feature information and target three-dimensional query feature information corresponding to each target to be detected. The target three-dimensional query feature information includes the second three-dimensional query feature information corresponding to each target to be detected obtained by each decoding layer, and the target two-dimensional query feature information includes multiple sets of two-dimensional query feature information corresponding to each target to be detected obtained by each decoding layer, or the target three-dimensional query feature information includes the second three-dimensional query feature information corresponding to each target to be detected obtained by the final decoding layer, and the target two-dimensional query feature information includes multiple sets of two-dimensional query feature information corresponding to each target to be detected obtained by the final decoding layer.
[0010] The first three-dimensional query feature information corresponding to the first level decoding layer is the three-dimensional query feature information, and the first three-dimensional attribute feature information corresponding to each level decoding layer is the three-dimensional attribute feature information.
[0011] By adopting the above technical solution, the multi-level decoding layer updates the query feature information step by step, which can obtain more accurate and comprehensive query feature information, and thus obtain more accurate target detection results.
[0012] In a possible implementation of the first aspect above, the joint decoder includes multiple decoding layers, and the backbone network inputs multiple image feature information into the joint decoder included in the target detection model, including: the backbone network inputs multiple image feature information into the decoding layers at each level included in the joint decoder; the joint decoder obtains target two-dimensional query feature information and target three-dimensional query feature information corresponding to each target to be detected based on the three-dimensional feature information of multiple targets to be detected corresponding to the multiple images and the multiple image feature information, including: the decoding layers at each level obtain multiple groups of two-dimensional query feature information corresponding to each target to be detected based on the input multiple image feature information and three-dimensional feature information, as target two-dimensional query feature information, and obtain second three-dimensional query feature information based on the multiple groups of two-dimensional query feature information, as target three-dimensional query feature information.
[0013] In a possible implementation of the first aspect above, each level of decoding layer includes a two-dimensional decoding layer and a three-dimensional decoding layer, respectively, and the backbone network inputs multiple image feature information into each level of decoding layer, respectively, including: the backbone network inputs multiple image feature information into the two-dimensional decoding layer and the three-dimensional decoding layer included in the decoding layer at each level; the method also includes, each level of decoding layer obtaining multiple groups of two-dimensional query feature information and second three-dimensional query feature information corresponding to each target to be detected in the following manner: the two-dimensional decoding layer obtains multiple groups of two-dimensional query feature information corresponding to each target to be detected based on the multiple image feature information, the first three-dimensional query feature information and the first three-dimensional attribute feature information, and obtains third three-dimensional query feature information corresponding to each target to be detected based on the multiple groups of two-dimensional query feature information, and inputs the third three-dimensional query feature information into the corresponding three-dimensional decoding layer; the three-dimensional decoding layer obtains the second three-dimensional query feature information corresponding to each target to be detected based on the multiple image feature information and the third three-dimensional query feature information.
[0014] By adopting the above technical solution, the three-dimensional query feature information is gradually updated based on the two-dimensional decoding layer and the three-dimensional decoding layer to obtain multiple sets of two-dimensional query feature information and new three-dimensional query feature information. Target detection is performed based on the combination of the multiple sets of two-dimensional query feature information and the new three-dimensional query feature information, which can achieve more comprehensive and accurate target detection.
[0015] In a possible implementation of the first aspect above, the two-dimensional decoding layer includes a distribution module, a deformable attention module and an aggregation module. The method also includes that the two-dimensional decoding layer obtains multiple sets of two-dimensional query feature information and third three-dimensional query feature information in the following manner: the distribution module initializes the first three-dimensional query feature information to obtain fourth three-dimensional query feature information, clusters the first three-dimensional attribute feature information to obtain second three-dimensional attribute feature information; the distribution module determines the camera parameters of the camera used to shoot multiple images, obtains the transformation matrix corresponding to multiple targets to be detected according to the camera parameters and the second three-dimensional attribute information, and inputs the transformation matrix into the aggregation module; the distribution module performs initialization processing on the first three-dimensional query feature information to obtain fourth three-dimensional query feature information, and ... performs initialization processing on The fourth three-dimensional query feature information and the transformation matrix obtain multiple two-dimensional query feature information corresponding to multiple targets to be detected, and the multiple two-dimensional query feature information are grouped and processed to obtain multiple groups of two-dimensional query feature information of multiple targets to be detected, and the multiple groups of two-dimensional query feature information are input into the deformable attention module; the deformable attention module obtains multiple updated groups of two-dimensional query feature information corresponding to multiple targets to be detected based on the multiple groups of two-dimensional query feature information and multiple image feature information, and inputs the updated multiple groups of two-dimensional query feature information into the aggregation module; the aggregation module obtains the third three-dimensional query feature information corresponding to each target to be detected based on the updated multiple groups of two-dimensional query feature information and the transformation matrix.
[0016] By adopting the above technical solution, the query feature information is updated according to the first three-dimensional query feature information and the image feature information to obtain multiple sets of updated two-dimensional query feature information, and then two-dimensional target detection can be performed more accurately based on the multiple sets of two-dimensional query feature information.
[0017] In a possible implementation of the first aspect above, the three-dimensional decoding layer includes a deformable aggregation module, and the method further includes the three-dimensional decoding layer obtaining the second three-dimensional query feature information in the following manner: the deformable aggregation module obtains the second three-dimensional query feature information corresponding to each target to be detected based on multiple image feature information and the third three-dimensional query feature information.
[0018] Using this technical solution, the deformable aggregation module updates the input image features and the third 3D query feature information again, obtaining updated second 3D query feature information corresponding to each target to be detected. This process gradually updates the 3D query feature information, allowing for more accurate 3D object detection.
[0019] In a possible implementation of the first aspect, the method further includes the distribution module obtaining multiple sets of two-dimensional query feature information in the following manner:
[0020]
[0021] in, are multiple sets of two-dimensional query feature information, for The transformation matrix, where M is the number of two-dimensional query feature information, N is the number of first three-dimensional query feature information, It is the first three-dimensional query feature information.
[0022] In a possible implementation of the first aspect, the method further includes the aggregation module obtaining the third three-dimensional query feature information in the following manner:
[0023]
[0024] in, For the third 3D query feature information, is the transformation matrix, are multiple sets of updated two-dimensional query feature information.
[0025] In a possible implementation of the first aspect above, the two-dimensional detection result includes the category of each target to be detected and the two-dimensional detection frame information, and the three-dimensional detection result includes the category of each target to be detected and the three-dimensional detection frame information.
[0026] In a possible implementation of the first aspect above, the target detection model is a detection model obtained by updating the model based on the corresponding loss, the loss includes two-dimensional loss and three-dimensional loss, and the loss is determined according to historical two-dimensional detection results and historical three-dimensional detection results.
[0027] By adopting the above technical solution, the target detection model is updated based on the two-dimensional loss and the three-dimensional loss, so that the updated target detection model can perform target detection more accurately.
[0028] In the second aspect, the implementation of the present application also discloses a method for updating a target detection model, which determines historical two-dimensional detection results and historical three-dimensional detection results. The historical two-dimensional detection results and historical three-dimensional detection results are obtained according to the target detection method provided by any implementation method of the first aspect. According to the historical two-dimensional detection results and historical three-dimensional detection results, the loss corresponding to the target detection model is determined, and the loss includes two-dimensional loss and three-dimensional loss; the target detection model is updated based on the loss.
[0029] In a third aspect, the implementation of the present application further discloses an electronic device, including a target detection model, which implements the target detection method provided by any implementation of the first aspect above.
[0030] In a fourth aspect, the implementation of the present application further discloses a computer-readable storage medium, which stores a computer program. The computer program can be executed by an electronic device to implement the target detection method provided by any implementation of the first aspect above.
[0031] In a fifth aspect, the implementation of the present application further discloses a computer program product, including a computer program, which, when executed by an electronic device, implements the target detection method provided by any one of the implementations of the first aspect above. BRIEF DESCRIPTION OF THE DRAWINGS
[0032] In order to more clearly illustrate the technical solution of the present application, the following is a brief introduction to the drawings used in the description of the implementation methods.
[0033] Figure 1 A schematic diagram of the principle of a target detection method in the prior art;
[0034] Figure 2 A schematic diagram illustrating the principle of the target detection method provided by an embodiment of the present invention;
[0035] Figure 3 A schematic diagram of the structure of the target detection model provided by an embodiment of the present invention;
[0036] Figure 4 A schematic diagram of a flow chart of a target detection method provided by an embodiment of the present invention;
[0037] Figure 5 A schematic diagram illustrating the principle of grouping distribution modules provided in an embodiment of the present invention;
[0038] Figure 6 A schematic diagram of grouped two-dimensional query feature information provided by an embodiment of the present invention;
[0039] Figure 7 A schematic diagram illustrating the principle of two-dimensional target detection by a two-dimensional detection head provided in an embodiment of the present invention;
[0040] Figure 8 A schematic diagram illustrating the principle of three-dimensional target detection by a three-dimensional detection head provided in an embodiment of the present invention;
[0041] Figure 9 A schematic diagram of the principle of loss processing based on a loss function provided in an embodiment of the present invention;
[0042] Figure 10 A schematic diagram of a process for obtaining target two-dimensional query feature information and target three-dimensional query feature information by the target detection model provided by an embodiment of the present invention;
[0043] Figure 11 A schematic structural diagram of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0044] Taking the field of autonomous driving as an example, the more accurate the target detection results obtained by target detection methods, the more precise the driving control based on the detection results will be. Therefore, how to more accurately detect targets and obtain more accurate detection results has become particularly important.
[0045] like Figure 1 As shown in the figure, traditional target detection methods include Figure 1 The multi-view shown in part (a) Figure 3 3D object detection involves inputting images corresponding to different viewpoints into a 3D detector so that the 3D detector can obtain a 3D detection box (also known as a 3D box). However, this object detection method only detects the 3D box and its related attributes, and the detection content is single, which may lead to incomplete and inaccurate object detection. In addition, some existing object detection methods have proven that using the detection results of a 2D detector to initialize 3D query feature information (also known as a 3D query) can achieve better 3D performance, such as Figure 1 The multi-view based two-dimensional (2D) detector shown in part (b) Figure 3 For D target detection, images corresponding to different perspectives are input into the 2D detector and the 3D detector respectively, so that the 2D detector obtains a 2D detection frame, and the 2D detection frame is input into the 3D detector, so that the 3D detector obtains a 3D detection frame. In this way, the 2D detector is only used for initialization once, and the output 2D detection result is relatively rough. The utilization efficiency of the 2D detector is limited. The query initialized by the 3D detector depends on the result of the 2D detector, and has certain requirements on the performance of the 2D detector. The features of the 2D detection results are not explicitly used in the 3D detector.
[0046] Based on this, this application proposes a multi-view target detection method that combines 2D and 3D, such as Figure 2 As shown, images corresponding to different perspectives are input into the joint detector so that the joint detector outputs 2D detection frames and 2D detection frames, and the 2D detection frames and 3D detection frames are associated. In this way, the 2D and 3D features can be effectively combined to achieve simultaneous detection of 2D and 3D targets.
[0047] Specifically, the object detection model obtains multiple image feature information from multiple images corresponding to different viewpoints. Based on the multiple image information and the 3D query (as an example of 3D query feature information) for the multiple objects to be detected corresponding to the multiple images, 2D and 3D query feature information of the object is obtained, thereby obtaining 2D and 3D detection results. This combines 2D and 3D detection to achieve simultaneous 2D and 3D object detection, resulting in more comprehensive and accurate object detection results.
[0048] The target detection method provided by the implementation of this application is applied to the target detection model, such as Figure 3 As shown, the target detection model includes a backbone network, a joint decoder and a detection head, and the detection head includes a two-dimensional detection head and a three-dimensional detection head.
[0049] The target detection model can specifically be a neural network model, such as a feed forward neural network (FNN) or a recurrent neural network (RNN).
[0050] like Figure 4 As shown in Figure 2, the target detection model implements target detection including the following steps.
[0051] S100, the backbone network obtains multiple image feature information based on multiple input images corresponding to different perspectives.
[0052] S200: The backbone network inputs multiple image feature information into the joint decoder.
[0053] S300 , a joint decoder obtains target 2D query feature information and target 3D query feature information corresponding to each target to be detected based on 3D query feature information of multiple targets to be detected corresponding to multiple input images and multiple image feature information.
[0054] S400: Inputting target two-dimensional query feature information into a two-dimensional detection head.
[0055] S500: Inputting target three-dimensional query feature information into a three-dimensional detection head.
[0056] S600: The two-dimensional detection head obtains two-dimensional detection results corresponding to multiple targets to be detected based on the input two-dimensional query feature information of the target.
[0057] S700: The three-dimensional detection head obtains three-dimensional detection results corresponding to multiple targets to be detected based on the input three-dimensional query feature information of the target.
[0058] In the implementation method of the present application, a joint decoder based on a target detection model obtains target two-dimensional query feature information and target three-dimensional query feature information of each target to be detected based on multiple image feature information and three-dimensional query feature information of multiple targets to be detected; a two-dimensional detection head based on the joint decoder obtains two-dimensional detection results of multiple targets to be detected based on the target two-dimensional query feature information; and a three-dimensional detection head based on the joint decoder obtains three-dimensional detection results of multiple targets to be detected based on the target three-dimensional query feature information of each target to be detected. In this way, based on the target detection model, two-dimensional target detection of the target to be detected is performed based on the three-dimensional query feature information, and three-dimensional target detection of the target to be detected is performed based on the three-dimensional query feature information. The combination of two-dimensional target detection and three-dimensional target detection realizes more accurate target detection, and the detection results obtained include both two-dimensional detection results and three-dimensional detection results, so that more comprehensive and accurate detection results can be obtained.
[0059] Taking target detection during vehicle assisted driving as an example, first, step S100 is executed. The electronic device obtains multiple images corresponding to different perspectives of the vehicle's surroundings taken by the vehicle's surround-view camera and other devices, and inputs the multiple images into the backbone network of the target detection model. The backbone network extracts image features based on the input multiple images corresponding to different perspectives to obtain multiple image feature information.
[0060] Furthermore, step S200 is executed, and the backbone network inputs multiple image feature information into the joint decoder.
[0061] Among them, ResNet50 pre-trained on ImageNet can be used as the backbone network for image feature extraction.
[0062] Furthermore, step S300 is performed to obtain a 3D query (as an example of three-dimensional query feature information) for each target to be detected corresponding to the image of each perspective, and the image feature information and the 3D query are input into a joint decoder, so that the joint decoder obtains a target 2D query (as an example of target two-dimensional query feature information) and a target 3D query (as an example of target three-dimensional query feature information) for the multiple targets to be detected based on the multiple image feature information and the 3D query.
[0063] Among them, 3D query represents the three-dimensional feature information of each target to be detected.
[0064] Exemplarily, the three-dimensional feature information includes three-dimensional query feature information and three-dimensional attribute feature information. The joint decoder obtains target two-dimensional query feature information and target three-dimensional query feature information corresponding to each target to be detected based on the three-dimensional feature information of multiple targets to be detected corresponding to multiple images and multiple image feature information, including: the joint decoder obtains multiple groups of two-dimensional query feature information corresponding to each target to be detected as target two-dimensional query feature information based on multiple image feature information, corresponding three-dimensional query feature information and three-dimensional attribute feature information, and obtains second three-dimensional query feature information as target three-dimensional query feature information based on the multiple groups of two-dimensional query feature information.
[0065] Furthermore, the joint decoder obtains multiple groups of two-dimensional query feature information corresponding to each target to be detected according to multiple image feature information, corresponding three-dimensional query feature information and three-dimensional attribute feature information, as target two-dimensional query feature information, and obtains second three-dimensional query feature information according to the multiple groups of two-dimensional query feature information, as target three-dimensional query feature information, including: the joint decoder initializes the three-dimensional query feature information to obtain fourth three-dimensional query feature information, clusters the three-dimensional attribute feature information to obtain second three-dimensional attribute feature information; determines camera parameters of a camera used to shoot multiple images, and obtains transformation matrices corresponding to multiple targets to be detected according to the camera parameters and the second three-dimensional attribute information; obtains transformation matrices corresponding to the fourth three-dimensional query feature information according to the fourth three-dimensional query feature information; and obtains transformation matrices corresponding to the fourth three-dimensional query feature information according to the fourth three-dimensional query feature information. The three-dimensional query feature information and the transformation matrix are used to obtain multiple two-dimensional query feature information corresponding to multiple targets to be detected, and the multiple two-dimensional query feature information are grouped and processed to obtain multiple groups of two-dimensional query feature information of the multiple targets to be detected; based on the multiple groups of two-dimensional query feature information and the multiple image feature information, updated multiple groups of two-dimensional query feature information corresponding to the multiple targets to be detected are obtained as target two-dimensional query feature information; based on the updated multiple groups of two-dimensional query feature information and the transformation matrix, third three-dimensional query feature information corresponding to each target to be detected is obtained; based on the multiple image feature information and the third three-dimensional query feature information, second three-dimensional query feature information corresponding to each target to be detected is obtained as target three-dimensional query feature information.
[0066] In one implementation of this application, see Figure 3 ,The joint decoder includes multiple levels of decoding layers.
[0067] The backbone network inputs multiple image feature information into each level of decoding layer.
[0068] Furthermore, the joint decoder obtains target two-dimensional query feature information and target three-dimensional query feature information corresponding to each target to be detected based on the three-dimensional query feature information of multiple targets to be detected corresponding to the input multiple images and multiple image feature information, including: each level of decoding layer obtains multiple groups of two-dimensional query feature information corresponding to each target to be detected based on the input multiple image feature information and three-dimensional feature information, as target two-dimensional query feature information, and obtains second three-dimensional query feature information based on the multiple groups of two-dimensional query feature information, as target three-dimensional query feature information.
[0069] Exemplarily, the joint decoder includes three joint decoding layers. Each decoding layer obtains multiple groups of 2D queries corresponding to each target to be detected based on multiple image feature information and 3D queries, and obtains 3Dquery' based on the multiple groups of 2D queries (as an example of the second three-dimensional query feature information). The multiple groups of 2D queries obtained by each decoding layer are used as target 2D queries, and the 3D query' obtained by each decoding layer is used as the target 3D query.
[0070] Furthermore, the upper level decoding layer may use the output 3D query' as the input 3D query of the lower level decoding layer to perform hierarchical update.
[0071] Therefore, in the implementation method of the present application, the joint decoder obtains target two-dimensional query feature information and target three-dimensional query feature information corresponding to each target to be detected based on the three-dimensional feature information of multiple targets to be detected corresponding to the input multiple images and multiple image feature information, including: each level of decoding layer obtains multiple groups of two-dimensional query feature information corresponding to each target to be detected based on the input multiple image feature information, the first three-dimensional query feature information and the first three-dimensional attribute feature information, obtains the second three-dimensional query feature information based on the multiple groups of two-dimensional query feature information, and inputs the second three-dimensional query feature information as the first three-dimensional query feature information of the next level decoding layer to the next level decoding layer, so that the next level decoding layer performs corresponding processing until the processing of the last level decoding layer is completed, and the target two-dimensional query feature information and target three-dimensional query feature information corresponding to each target to be detected are obtained.
[0072] The target three-dimensional query feature information includes the second three-dimensional query feature information corresponding to each target to be detected obtained by each level of decoding layer, and the target two-dimensional query feature information includes multiple groups of two-dimensional query feature information corresponding to each target to be detected obtained by each level of decoding layer, or the target three-dimensional query feature information includes the second three-dimensional query feature information corresponding to each target to be detected obtained by the last level of decoding layer, and the target two-dimensional query feature information includes multiple groups of two-dimensional query feature information corresponding to each target to be detected obtained by the last level of decoding layer.
[0073] Exemplarily, taking the three-level decoding layer as an example, the first-level decoding layer obtains multiple groups of 2D queries based on multiple image feature information, 3D query (that is, the first three-dimensional query feature information corresponding to the first-level decoding layer), and 3D anchor (that is, the first three-dimensional attribute feature information), and obtains 3D query' (as an example of the second three-dimensional query feature information obtained by the first-level decoding layer) based on the multiple groups of 2D queries, and inputs the 3D query' into the second-level decoding layer. The second-level decoding layer obtains multiple groups of 2D queries based on multiple image feature information, 3D query' (that is, the first three-dimensional query feature information corresponding to the second-level decoding layer) and 3D anchor (that is, the first three-dimensional attribute feature information), and obtains 3D query' (as an example of the second three-dimensional query feature information obtained by the second-level decoding layer) based on the multiple groups of 2D queries. The 3D query' is input into the third-level decoding layer. The second decoding layer obtains multiple groups of 2D queries based on multiple image feature information, 3D query' (that is, the first three-dimensional query feature information corresponding to the third-level decoding layer) and 3D anchor (that is, the first three-dimensional attribute feature information), and obtains 3D based on the multiple groups of 2D queries. query' (as an example of the second three-dimensional query feature information corresponding to the third-level decoding layer).
[0074] The multiple sets of 2D queries obtained at each decoding layer are used as target 2D queries, and the 3D query' obtained at each decoding layer is used as the target 3D query. Alternatively, the multiple sets of 2D queries obtained at the last decoding layer are used as target 2D queries, and the 3D query' obtained at the last decoding layer is used as the target 3D query.
[0075] It should be noted that, for the first level decoding layer, the first three-dimensional query feature information (ie, 3D query') corresponding to the first level decoding layer is three-dimensional query feature information (3D query), and the first three-dimensional attribute feature information corresponding to each level decoding layer is three-dimensional attribute feature information.
[0076] Furthermore, in the present application, each level of decoding layer includes a two-dimensional decoding layer (i.e., a 2D decoding layer) and a three-dimensional decoding layer (i.e., a 3D decoding layer), and the backbone network inputs multiple image feature information into the two-dimensional decoding layer and the three-dimensional decoding layer included in each level of decoding layer.
[0077] Each decoding layer obtains the second three-dimensional query feature information and multiple sets of two-dimensional query feature information corresponding to each target to be detected in the following manner.
[0078] First, the two-dimensional decoding layer obtains multiple sets of two-dimensional query feature information corresponding to each target to be detected based on the input multiple image feature information, the first three-dimensional query feature information (or three-dimensional query feature information) and the first three-dimensional attribute feature information, and obtains third three-dimensional query feature information corresponding to each target to be detected based on the multiple sets of two-dimensional query feature information, and inputs the third three-dimensional query feature information into the corresponding three-dimensional decoding layer.
[0079] In the implementation of the present application, the input of the target detection model also includes a 3D anchor (as an example of the first three-dimensional attribute information), and the 3D anchor includes the reference position, size and related attribute information of the target to be detected.
[0080] For example, see Figure 3 ,The 2D decoding layer includes a distribution module, a deformable attention module, and an aggregation module.
[0081] The two-dimensional decoding layer obtains the third three-dimensional query feature information in the following manner.
[0082] First, the distribution module randomly initializes the input 3D query (for the distribution module of the first-level decoding layer, it is the three-dimensional query feature information; for the distribution modules of other levels of decoding layers, it is the first three-dimensional query feature information) to obtain an initialized 3D query with a dimension of [n, 256] (as an example of the fourth three-dimensional query feature information), and clusters the input 3D anchor (that is, the first three-dimensional attribute feature information, that is, the three-dimensional attribute feature information) to obtain a clustered 3D anchor with a dimension of [n, 11] (as an example of the second three-dimensional attribute information).
[0083] Among them, the clustering processing can be specifically Kmeans clustering, and the 11-dimensional attributes are [x, y, z, l, w, h, sin, cos, vx, vy, vz], where x, y, z are three-dimensional coordinates, l, w, h are the length, width and height of the three-dimensional detection box, sin and cos are sine and cosine values, and vx, vy, vz are the speeds of the target to be detected relative to the x, y and z directions.
[0084] Furthermore, the distribution module determines camera parameters of a camera used to capture the multiple images, obtains a transformation matrix corresponding to the multiple targets to be detected based on the camera parameters and the second three-dimensional attribute information, and inputs the transformation matrix into the aggregation module.
[0085] For example, Figure 5As shown in the figure, the distribution module projects the 3D anchor (also known as anchor3d) into the multi-view image according to the camera parameters to obtain the transformation matrix (Transform Matrix, also known as the transformation matrix). The transformation matrix dimension is a 0, 1 matrix of [n, m], which is used to implement the mapping relationship between the 3D query (also known as query3d) and the distribution to different camera perspectives. , , M is the number of two-dimensional query feature information, and N is the number of first three-dimensional query feature information.
[0086] Furthermore, the distribution module obtains multiple two-dimensional query feature information corresponding to multiple targets to be detected based on the fourth three-dimensional query feature information and the transformation matrix, groups the multiple two-dimensional query feature information to obtain multiple groups of two-dimensional query feature information of multiple targets to be detected, and inputs the multiple groups of two-dimensional query feature information into the deformable attention module.
[0087] Exemplarily, according to the transformation matrix, the distribution module can distribute the 3D query and 3D anchor to multiple views.
[0088] like Figure 5 and Figure 6 As shown in the figure, for example, if the 3D anchor corresponding to a 3D query for an object to be detected spans two viewpoints, it will be copied and used for 2D object detection from different cameras (reference_points2d). The 3D anchor is projected onto the multi-view image, and the center point of the bounding rectangle is selected as the reference point for the subsequent deformable attention module (DA).
[0089] That is, according to the number of 3D anchors corresponding to the 3D query of each target to be detected on each image, the N value of the corresponding transformation matrix can be determined, that is, the number of grouped 2D queries can be determined.
[0090] Furthermore, the distribution module obtains multiple 2D queries according to the 3D query and the transformation matrix.
[0091] The distribution module obtains multiple sets of two-dimensional query feature information in the following way:
[0092]
[0093] in, are multiple sets of two-dimensional query feature information, for The transformation matrix, where M is the number of two-dimensional query feature information, N is the number of first three-dimensional query feature information, It is the first three-dimensional query feature information.
[0094] Trans is a matrix consisting of only two values, {0, 1}. Trans[n, m] indicates that the nth 3D query is distributed to m 2D queries. If the nth 3D query is distributed as two 2D queries (with indexes m1 and m2, respectively), then Trans[n, m1] and Trans[n, m2] are equal to 1. Matrix multiplication naturally replicates the 3D query and distributes it to one or more locations.
[0095] Taking four targets to be detected as an example, the center point of the 3D anchor corresponding to the 3D query of target 1 to be detected and the 6 circumscribed corner points are projected to each camera perspective through the camera parameter K. When the point falls inside the image and the depth is greater than 0, it is considered that the 3D anchor needs to be distributed to this image perspective, and it can be distributed as two 2D queries (2D query1 and 2D query2) based on the transformation matrix. Similarly, the 3D query of target 2 to be detected is distributed as two 2D queries (2D query3 and 2D query4) based on the transformation matrix. The 3D query of target 3 to be detected is converted into a 2D query5 based on the transformation matrix, and the 3D query of target 4 to be detected is converted into a 2D query6 based on the transformation matrix. In this way, the 3D anchors of all targets to be detected are grouped in turn to obtain the relationship between n 3D anchors distributed to multiple views, and the Trans matrix is filled to obtain the Tran matrix, so as to obtain multiple 2D queries based on the Trans matrix.
[0096] For multiple 2D queries, based on the projection of the 3D anchor on the view (image), 2D query1 and 2D query3 in the same view can be grouped together, and 2D query2, 2D query4, 2D query5, and 2D query6 in another view can be grouped together, thus obtaining multiple groups of 2D queries.
[0097] Furthermore, the deformable attention module obtains updated multiple sets of two-dimensional query feature information corresponding to multiple targets to be detected based on the multiple sets of two-dimensional query feature information and the multiple image feature information, and inputs the updated multiple sets of two-dimensional query feature information into the aggregation module.
[0098] Exemplarily, the deformable attention module samples the image features at each view separately to update each set of 2D queries combined with the grouped multi-view image features to obtain multiple sets of updated 2D queries'.
[0099] The aggregation module obtains the third three-dimensional query feature information corresponding to each target to be detected based on the updated multiple sets of two-dimensional query feature information and the transformation matrix.
[0100] Exemplarily, the aggregation module reassembles the updated 2D query' into a 3D query for use in subsequent 3D decoding layers.
[0101] The aggregation module obtains the third-dimensional query feature information in the following way:
[0102]
[0103] in, For the third 3D query feature information, is the transformation matrix, are multiple sets of updated two-dimensional query feature information.
[0104] Furthermore, the three-dimensional decoding layer obtains second three-dimensional query feature information corresponding to each target to be detected based on the input multiple image feature information and the third three-dimensional query feature information.
[0105] Continue to see Figure 3 , the three-dimensional decoding layer includes a deformable aggregation module.
[0106] The 3D decoding layer obtains the second 3D query feature information in the following manner: the deformable aggregation module obtains the second 3D query feature information corresponding to each target to be detected based on the input multiple image feature information and the third 3D query feature information.
[0107] Exemplarily, the deformable aggregation module samples image feature information from multiple views, samples image feature information of multiple targets to be detected from the multiple views, and dynamically aggregates these image features through learnable weights to interact with the 3D query, so that useful 3D query information of the targets to be detected can be retrieved from the multiple views based on the 3D query, thereby obtaining an updated 3D query' (i.e., the second three-dimensional query feature information).
[0108] Furthermore, the updated multiple sets of 2D query' output by each level of two-dimensional decoding layer are used as target 2D query and input into the two-dimensional detection head to obtain two-dimensional detection results. The updated 3D query' output by each level of three-dimensional decoding layer is used as target 3D query and input into the three-dimensional detection head to obtain three-dimensional detection results. In this way, the updated multiple sets of 2D query' output by each level of decoding layer are input into the two-dimensional detection head for two-dimensional target detection, which can obtain more accurate two-dimensional detection results. The updated 3D query' output by each level of decoding layer is input into the three-dimensional detection head for three-dimensional target detection, which can obtain more accurate three-dimensional detection results, greatly improving the accuracy of target detection results.
[0109] Alternatively, the updated multiple sets of 2D queries' output by the last two-dimensional decoding layer are used as target 2D queries and input into the two-dimensional detection head to obtain two-dimensional detection results. The updated 3D query' output by the last three-dimensional decoding layer is used as the target 3D query and input into the three-dimensional detection head to obtain three-dimensional detection results. In this way, because the multi-level decoding layers are updated layer by layer, the 2D query' output by the last decoding layer is the latest 2D query after multiple updates, and the 3D query' output by the last decoding layer is the latest 3D query after multiple updates. Therefore, target detection is performed only based on the multiple sets of 2D queries' and 3D queries' output by the last decoding layer, reducing the detection data for the two-dimensional and three-dimensional detection heads while still obtaining accurate detection results.
[0110] It should be noted that the number of input 3D queries is the same as the number of final output 3D queries', so the 3D anchors corresponding to the 3D queries can be reused in the multi-level decoding layers.
[0111] Furthermore, step S400 is executed, and the two-dimensional decoding layer inputs the target 2D query into the two-dimensional detection head (ie, the 2D detection head).
[0112] Exemplarily, the 2D detection head is placed after the 2D decoding layer to receive the grouped target 2D queries.
[0113] Furthermore, step S500 is executed, and the 3D decoding layer inputs the target 3D query to the 3D detection head (ie, the 3D detection head).
[0114] Exemplarily, the 3D detection head is placed after the 3D decoding layer to receive the target 3D query.
[0115] Furthermore, step S600 is executed, where the 2D detection head obtains two-dimensional detection results corresponding to a plurality of targets to be detected based on the input two-dimensional query feature information of the target.
[0116] The two-dimensional detection results include the category of each target to be detected and the two-dimensional detection frame information.
[0117] For example, Figure 7 As shown in the figure, the 2D detection head receives grouped 2D queries (i.e., target 2D query feature information). The 2D queries corresponding to multiple views from different perspectives share the weights of the same 2D detection head, and obtain the category and 2D detection box information of each target to be detected (i.e., 2D detection results on multiple views).
[0118] Furthermore, step S700 is executed, where the three-dimensional detection head obtains three-dimensional detection results corresponding to a plurality of targets to be detected based on the input three-dimensional query feature information of the target.
[0119] The 3D detection results include the category of each target to be detected and the 3D detection frame information.
[0120] For example, Figure 8 As shown in the figure, the 3D detection head receives a 3D query (i.e., the target's three-dimensional query feature information) and outputs the category and 3D detection box of the target to be detected, as well as the relevant attribute information of the target to be detected (i.e., the three-dimensional detection result from the bird's eye view (BEV)).
[0121] Furthermore, in the implementation of the present application, the target detection model is a detection model obtained by updating the model based on the corresponding loss, and the loss includes two-dimensional loss and three-dimensional loss, and the loss is determined according to historical two-dimensional detection results and historical three-dimensional detection results.
[0122] For example, similar to DETR (a method for 2D object detection) and DETR3D (a method for 3D object detection), the grouped 2D and 3D detection boxes output by the object detection model are first matched against their ground truth values based on a loss function. A loss is then applied to the matching results, and the model gradient is updated through backpropagation. In 2D object detection, each group of grouped 2D detection boxes is equivalent to an independent image, generating a loss. The 2D detection losses for each group are then combined to update the gradient.
[0123] Among them, such as Figure 9 As shown, the loss function can perform 2D loss and 3D loss. The 2D loss includes 2D category classification loss, 2D detection box IOU (Intersection over Union) loss, and 2D detection box regression loss. The 3D loss includes 3D category classification loss, 3D detection box regression loss, and 3D attribute regression loss.
[0124] In the implementation of this application, S400 and S500 can be executed synchronously, and S600 and S700 can be executed synchronously.
[0125] In another implementation of the present application, Figure 10 As shown in the figure, the specific principles of target detection by the target detection model are as follows:
[0126] Multi-view images (i.e., images corresponding to different perspectives) are input into the backbone network, and image features are extracted by the backbone network to obtain multi-view image feature information. Then, the 3D queries corresponding to the multiple images are input into the joint decoder. Each joint decoding layer of the joint decoder includes a two-dimensional decoding layer and a three-dimensional decoding layer. The distribution module (Allocation) of the two-dimensional decoding layer obtains grouped 2D queries and a transformation matrix (Transform Matrix) according to the input 3D query (first three-dimensional query feature information), sends the transformation matrix to the aggregation module, and sends the grouped 2D queries to the corresponding deformable attention modules (Deformable Attention), wherein the deformable attention module includes at least one, and the number of deformable attention modules corresponds to the number of 2D query groups. The deformable attention module obtains multiple updated groups of 2D queries based on the input grouped 2D query and multiple image feature information, and sends the updated multiple groups of 2D queries to the aggregation module (Aggregation). The aggregation module obtains an updated 3D query based on the transformation matrix and the updated multiple groups of 2D queries, and sends the updated 3D query to the aggregation module. The query is sent to the deformable aggregation module of the 3D decoding layer. The deformable aggregation module updates and fuses the input multiple image feature information and the updated 3D query again to obtain the updated 3D query' (target three-dimensional query feature information).
[0127] In this way, the initialized 3D query will be gradually updated through the joint decoder to obtain multiple groups of 2D query' and 3D query'. The updated multiple groups of 2D query' and 3D query' will respectively output 2D detection results and 3D detection results through their respective detection heads. Among them, in each decoding layer, the 3D query will pass through the 2D decoding layer and the 3D decoding layer, and interact with the image features of multiple views in turn. In the 2D decoding layer, the 3D query will be distributed to each view for grouped single-view 2D target detection, and then the 2D queries' of these targets to be detected will be fused in the aggregation module to obtain the 3D query, and then sent to the subsequent 3D decoding layer for 3D target detection to obtain the 3D query'.
[0128] The object detection method provided by the implementation of this application is actually a multi-view object detection method that combines 2D and 3D. Based on the joint decoder, it gradually updates the 3D query in combination with image feature information from multiple views, obtaining multiple sets of updated 2D queries and updated 3D queries. The 2D detection head obtains 2D detection results based on the multiple sets of 2D queries, and the 3D detection head obtains 3D detection results based on the 3D queries. In this way, the combined 2D and 3D object detection method can better perform target detection, and obtain more accurate and comprehensive target detection.
[0129] The target detection method provided by the implementation of this application can be applied to electronic devices, in which a target detection model is deployed to perform the target detection method. Specifically, the electronic device can be a vehicle, a remote computer, or other device.
[0130] The target detection method provided by the implementation of this application can also be used for target detection in virtual scenes.
[0131] The present application also discloses a target detection model updating method, which determines historical two-dimensional detection results and historical three-dimensional detection results. The historical two-dimensional detection results and historical three-dimensional detection results are obtained according to the aforementioned target detection method. Based on the historical two-dimensional detection results and historical three-dimensional detection results, the loss corresponding to the target detection model is determined, and the loss includes two-dimensional loss and three-dimensional loss; the target detection model is updated based on the loss.
[0132] See Figure 11 , Figure 11 The figure shows a schematic diagram of the structure of the electronic device provided in the embodiment of the present application. Figure 11 As shown, the electronic device may include: a transceiver 121 , a processor 122 , and a memory 123 .
[0133] The processor 122 executes the computer-executable instructions stored in the memory, so that the processor 122 performs the technical solution of the target detection method in the above-mentioned embodiment. The processor 122 can be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it can also be a digital data processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components.
[0134] The memory 123 is connected to the processor 122 via a system bus and communicates with the processor 122. The memory 123 is used to store computer program instructions.
[0135] By way of example and not limitation, memory 123 may include a hard disk drive (HDD), a floppy disk drive, flash memory, an optical disk, a magneto-optical disk, a magnetic tape, or a universal serial bus (USB) drive, or a combination of two or more thereof. Where appropriate, memory 123 may include removable or non-removable (or fixed) media. Where appropriate, memory 123 may be internal or external to the integrated gateway device. In certain embodiments, memory 123 is a non-volatile solid-state memory. In certain embodiments, memory 123 includes read-only memory (ROM). Where appropriate, the ROM may be a mask-programmed ROM, a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), an electrically alterable read-only memory (EAROM), or flash memory, or a combination of two or more thereof. The transceiver 121 may be used to obtain tasks to be executed and configuration information of the tasks to be executed.
[0136] The system bus can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, among others. System buses can be divided into address buses, data buses, and control buses. For ease of illustration, the diagram uses only a single thick line, but this does not imply a single bus or type of bus. Transceivers enable communication between the database access device and other computers (such as clients, read-write libraries, and read-only libraries). Memory may include random access memory (RAM) and non-volatile memory.
[0137] An embodiment of the present application also provides a chip for executing instructions, which is used to execute the technical solution of the target detection method in the above embodiment.
[0138] An embodiment of the present application also provides a computer-readable storage medium, which stores computer instructions. When the computer instructions are executed on a processor of an electronic device, the processor of the electronic device executes the technical solution of the target detection method of the above embodiment.
[0139] In some possible implementations, various aspects of the method provided in the present application can also be implemented in the form of a program product, which includes program code. When the program product runs on a processor of an electronic device, the program code is used to enable the processor of the electronic device to execute the steps of the method according to various exemplary implementations of the present application described above in this specification. For example, the electronic device can execute the target detection method recorded in the embodiments of the present application.
[0140] The program product may employ any combination of one or more readable media. The readable medium may be a readable data medium or a readable storage medium. The readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or component, or any combination thereof. More specific examples of readable storage media (a non-exhaustive list) include: an electrical connection having one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CDROM), an optical storage device, a magnetic storage device, or any suitable combination thereof.
[0141] The implementation method of the present application also provides a computer program product, which includes a computer program stored in a computer-readable storage medium. At least one processor can read the computer program from the computer-readable storage medium. When the at least one processor executes the computer program, the technical solution of the target detection method in the above embodiment can be implemented.
[0142] It should be noted that, in addition to the implementation of the present application described in the above-mentioned specific embodiments, those skilled in the art can easily understand other advantages and effects of the present application from the contents disclosed in this specification. Although the description of the present application is introduced in conjunction with the preferred embodiment, this does not mean that the features of this invention are limited to this implementation. On the contrary, the purpose of introducing the invention in conjunction with the implementation is to cover other options or modifications that may be extended from the present application. In order to provide an in-depth understanding of the present application, the above description contains many specific details, and the present application can also be implemented without using these details. In addition, in order to avoid confusion or blurring the focus of the present application, some specific details will be omitted in the description. It should be noted that, in the absence of conflict, the embodiments in the present application and the features in the embodiments can be combined with each other.
[0143] It should be noted that in this specification, similar reference numerals and letters denote similar items in the following drawings, and therefore, once an item is defined in one drawing, it does not need to be further defined or explained in subsequent drawings.
[0144] It should be noted that the terms "first", "second", "third", "fourth", etc. are only used to distinguish and describe, and cannot be understood as indicating or implying relative importance.
[0145] It should be noted that in the accompanying drawings, some structural or method features may be shown in a specific arrangement and / or order. However, it should be understood that such a specific arrangement and / or order may not be required. Rather, in some embodiments, these features may be arranged in a manner and / or order different from that shown in the illustrative drawings. In addition, the inclusion of structural or method features in a particular figure does not imply that such features are required in all embodiments, and in some embodiments, these features may not be included or may be combined with other features.
[0146] Although the present application has been illustrated and described with reference to certain preferred implementations of the present application, those skilled in the art should understand that the above description is provided as a further detailed explanation of the present application in conjunction with specific implementations, and that the specific implementation of the present application should not be limited to these descriptions. Those skilled in the art may make various changes in form and detail, including simple deductions or substitutions, without departing from the spirit and scope of the present application.
Claims
1. A target detection method, characterized in that: Applied to a target detection model, the target detection model includes a backbone network, a joint decoder and a detection head, the detection head includes a two-dimensional detection head and a three-dimensional detection head, the method includes: The backbone network obtains a plurality of image feature information according to a plurality of images corresponding to different perspectives, and inputs the plurality of image feature information into the joint decoder; The joint decoder obtains target two-dimensional query feature information and target three-dimensional query feature information corresponding to each target to be detected based on the three-dimensional feature information of the multiple targets to be detected corresponding to the multiple images and the feature information of the multiple images, inputs the target two-dimensional query feature information to the two-dimensional detection head, and inputs the target three-dimensional query feature information to the three-dimensional detection head; The two-dimensional detection head obtains two-dimensional detection results corresponding to the multiple targets to be detected based on the target two-dimensional query feature information; The three-dimensional detection head obtains three-dimensional detection results corresponding to the multiple targets to be detected based on the target three-dimensional query feature information.
2. The target detection method according to claim 1, wherein: The joint decoder comprises multiple decoding layers, The backbone network inputs the plurality of image feature information into the joint decoder, including: The backbone network inputs the plurality of image feature information into the decoding layers at each level respectively; The three-dimensional feature information includes three-dimensional query feature information and three-dimensional attribute feature information. The joint decoder obtains target two-dimensional query feature information and target three-dimensional query feature information corresponding to each target to be detected based on the three-dimensional query feature information of the multiple targets to be detected corresponding to the multiple images and the feature information of the multiple images, including: Each decoding layer obtains multiple sets of two-dimensional query feature information corresponding to each target to be detected based on the multiple image feature information, the corresponding first three-dimensional query feature information, and the first three-dimensional attribute feature information. Second three-dimensional query feature information is obtained based on the multiple sets of two-dimensional query feature information, and the second three-dimensional query feature information is input into the next decoding layer as the first three-dimensional query feature information corresponding to the next decoding layer, so that the next decoding layer performs corresponding processing until the last decoding layer completes processing, thereby obtaining target two-dimensional query feature information and target three-dimensional query feature information corresponding to each target to be detected. The target three-dimensional query feature information includes the second three-dimensional query feature information corresponding to each target to be detected obtained by each decoding layer, and the target two-dimensional query feature information includes the multiple sets of two-dimensional query feature information corresponding to each target to be detected obtained by each decoding layer, or the target three-dimensional query feature information includes the second three-dimensional query feature information corresponding to each target to be detected obtained by the last decoding layer, and the target two-dimensional query feature information includes the multiple sets of two-dimensional query feature information corresponding to each target to be detected obtained by the last decoding layer.
3. The target detection method according to claim 2, wherein: The decoding layers at each level include a two-dimensional decoding layer and a three-dimensional decoding layer. The backbone network inputs the plurality of image feature information into the decoding layers at each level, including: The backbone network inputs the plurality of image feature information to the two-dimensional decoding layer and the three-dimensional decoding layer included in each level of the decoding layer respectively; The method further includes obtaining, at each level of the decoding layer, the multiple sets of two-dimensional query feature information and the second three-dimensional query feature information corresponding to each target to be detected by: The two-dimensional decoding layer obtains the multiple sets of two-dimensional query feature information corresponding to each of the to-be-detected objects based on the multiple sets of image feature information, the first three-dimensional query feature information, and the first three-dimensional attribute feature information, obtains third three-dimensional query feature information corresponding to each of the to-be-detected objects based on the multiple sets of two-dimensional query feature information, and inputs the third three-dimensional query feature information into the corresponding three-dimensional decoding layer; The three-dimensional decoding layer obtains the second three-dimensional query feature information corresponding to each of the to-be-detected objects based on the multiple image feature information and the third three-dimensional query feature information.
4. The target detection method according to claim 3, wherein: The two-dimensional decoding layer includes a distribution module, a deformable attention module, and an aggregation module. The method further includes the two-dimensional decoding layer obtaining the third three-dimensional query feature information by: The distribution module performs initialization processing on the first three-dimensional query feature information to obtain fourth three-dimensional query feature information, and performs clustering processing on the first three-dimensional attribute feature information to obtain second three-dimensional attribute feature information; The distribution module determines camera parameters of a camera used to capture the plurality of images, obtains a transformation matrix corresponding to the plurality of objects to be detected based on the camera parameters and the second three-dimensional attribute feature information, and inputs the transformation matrix into the aggregation module; The distribution module obtains a plurality of two-dimensional query feature information corresponding to the plurality of to-be-detected targets based on the fourth three-dimensional query feature information and the transformation matrix, groups the plurality of two-dimensional query feature information to obtain the plurality of groups of two-dimensional query feature information corresponding to the plurality of to-be-detected targets, and inputs the plurality of groups of two-dimensional query feature information into the deformable attention module; The deformable attention module obtains the updated multiple sets of two-dimensional query feature information corresponding to the multiple targets to be detected based on the multiple sets of two-dimensional query feature information and the multiple image feature information, and inputs the updated multiple sets of two-dimensional query feature information into the aggregation module; The aggregation module obtains the third three-dimensional query feature information corresponding to each of the to-be-detected objects based on the updated multiple sets of two-dimensional query feature information and the transformation matrix.
5. The target detection method according to claim 4, characterized in that: The three-dimensional decoding layer includes a deformable aggregation module. The method further includes the three-dimensional decoding layer obtaining the second three-dimensional query feature information by: The deformable aggregation module obtains the second three-dimensional query feature information corresponding to each of the to-be-detected objects according to the plurality of image feature information and the third three-dimensional query feature information.
6. The target detection method according to claim 5, characterized in that: The method further includes, the distribution module obtaining the multiple sets of two-dimensional query feature information by: in, are the multiple sets of two-dimensional query feature information, for The conversion matrix, wherein M is the number of the two-dimensional query feature information, N is the number of the first three-dimensional query feature information, The first three-dimensional query feature information.
7. The target detection method according to claim 6, characterized in that: The method further includes, the aggregation module obtaining the third three-dimensional query feature information in the following manner: in, is the third three-dimensional query feature information, is the transformation matrix, are the multiple sets of two-dimensional query feature information after update.
8. The target detection method according to claim 7, wherein: The two-dimensional detection result includes the category and two-dimensional detection frame information of each target to be detected, and the three-dimensional detection result includes the category and three-dimensional detection frame information of each target to be detected.
9. The target detection method according to any one of claims 1 to 8, characterized in that: The target detection model is a detection model obtained by updating the model based on the corresponding loss, wherein the loss includes a two-dimensional loss and a three-dimensional loss, and the loss is determined according to historical two-dimensional detection results and historical three-dimensional detection results.
10. An electronic device, characterized in that: The electronic device includes a target detection model, and the target detection model is used to implement the target detection method according to any one of claims 1 to 9.
Citation Information
Patent Citations
Three-dimensional target detection method and device and computer readable storage medium
CN114359892A
Image processing method and related equipment
CN116758301A
Target detection model training method, target detection method, device and equipment
CN117274575A
Target detection method and device and electronic equipment
CN117541816A
Target detection method, target detection device, equipment and computer storage medium
CN120339574A