Weakly perceived target detection method and related equipment
By introducing weak-perceptual object detection method in the autonomous driving system, using point cloud feature encoding subnetwork and transformer encoder/decoder for shape feature completion and attention fusion, the problems of sparse and incomplete shape of long-distance target point cloud data are solved, accurate detection and recognition of weak-perceptual objects are achieved, and the safety of autonomous driving is improved.
Patent Information
- Application Number
- CN202210650722.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-06-09
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2042-06-09
AI Technical Summary
In autonomous driving tasks, due to the complex interference of actual scenes and the unbalanced radar acquisition equipment, the point cloud data of long-distance targets is sparse and incomplete in shape, making it difficult for weak-perceived targets to be accurately positioned, affecting the safety of autonomous driving.
A weak-perceptual object detection method is proposed. By obtaining the training sample set and inputting it into the weak-perceptual object detection network, the point cloud feature encoding subnet and transformer encoder/decoder are used for iterative operations to complete the shape characteristics of the weak-perceptual object, and the global characteristics of the weak-perceptual object are obtained through attention fusion operations, and finally the weak-perceptual object detection box and category information are generated.
It effectively improves the detection accuracy of weak-perceived targets, can accurately locate and identify weak-perceived targets in complex scenarios, and improves the safety of the autonomous driving system.
Smart Images

Figure CN115222954B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of target recognition, and more specifically, to a weak-perception target detection method and related equipment. Background Art
[0002] In recent years, frequent traffic accidents and low traffic efficiency have become the most serious problems in urban traffic development, causing huge economic losses and adverse effects on society. With the rise of artificial intelligence and computer vision technology, the rapid development of autonomous driving technology has provided new solutions for ensuring traffic safety and improving traffic efficiency. In order for the autonomous driving system to accurately perceive complex environments, it is necessary to accurately detect and locate the targets in the current scene and determine their categories. Therefore, studying target detection methods in complex scenes is crucial to autonomous driving technology.
[0003] At present, deep neural networks based on human brain intelligence have made certain progress in the field of computer vision and are widely used in target detection tasks. However, the two-dimensional target detection method based on deep learning lacks depth information, making it difficult to accurately locate three-dimensional space targets and is sensitive to lighting and weather conditions. The three-dimensional target detection method based on deep learning using lidar point cloud data can obtain high-precision depth information in complex environments and achieve better detection performance. Therefore, it is widely used in the field of autonomous driving. However, in autonomous driving tasks, due to the complex interference of actual scenes and the imbalance of radar acquisition equipment, the farther the target is from the sensor, the fewer points are collected, that is, the point cloud data of distant targets is relatively sparse and the shape is incomplete. Such incomplete weakly perceived targets are difficult to be accurately located due to the lack of shape feature information, which affects the safety of autonomous driving. Therefore, studying the detection method of weakly perceived targets in complex scenes is the key to solving the problem of complex scene perception in autonomous driving tasks and improving driving safety. Summary of the invention
[0004] A series of simplified concepts are introduced in the Summary of the Invention, which will be further described in detail in the Detailed Description of the Invention. The Summary of the Invention does not mean to attempt to define the key features and essential technical features of the claimed technical solution, nor does it mean to attempt to determine the scope of protection of the claimed technical solution.
[0005] In order to improve the recognition accuracy of weakly perceived targets, in a first aspect, the present invention proposes a weakly perceived target detection method, the method comprising:
[0006] Obtain a training sample set, input it into a weak-sensing target detection network, and perform preliminary detection through the first point cloud feature encoding subnetwork in the weak-sensing target detection network to obtain initial candidate data, wherein the initial candidate data includes an initial candidate box, candidate target point cloud data, and candidate target point cloud features;
[0007] Based on the candidate target point cloud data, the transformer encoder and the transformer decoder in the weak-perception target detection network are used to perform iterative operations to obtain a shape feature sequence of the missing part of the weak-perception target, and the reconstruction operation in the weak-perception target detection network is used to obtain a complete shape feature sequence of the weak-perception target according to the shape feature sequence of the missing part of the weak-perception target;
[0008] Based on the complete shape feature sequence of the weakly perceived target and the point cloud features of the candidate target, an attention fusion operation in the weakly perceived target detection network is performed to obtain a global feature of the weakly perceived target;
[0009] Based on the global features of the weakly perceived target, confidence calculation and position regression operations are performed in the weakly perceived target detection network to obtain a confidence score and a residual parameter of the weakly perceived target, and a loss value is calculated based on the confidence score and the residual parameter to adjust the parameters of the weakly perceived target detection network to generate a weakly perceived target detection model;
[0010] The above-mentioned weak-perception target detection model is used to detect the sample set to be detected, generate weak-perception target detection boxes and weak-perception target category information, and complete weak-perception target detection.
[0011] Optionally, the iterative operation includes a first iterative operation and a second iterative operation, and the reconstruction operation includes a folding operation and a feature extraction operation;
[0012] The above-mentioned method uses the transformer encoder and the transformer decoder in the above-mentioned weakly perceived target detection network to perform iterative operations based on the above-mentioned candidate target point cloud data to obtain the shape feature sequence of the missing part of the weakly perceived target, and uses the reconstruction operation in the above-mentioned weakly perceived target detection network to obtain the complete shape feature sequence of the weakly perceived target according to the shape feature sequence of the missing part of the weakly perceived target, including:
[0013] Performing the sampling convolution operation and the first embedding operation in the weak-sensing target detection network according to the candidate target point cloud data to obtain a local structural feature sequence of the embedding position;
[0014] Based on the above-mentioned local structural feature sequence of the embedded position, using the above-mentioned transformer encoder to perform the above-mentioned first iterative operation and dimension transformation operation to obtain the missing part center point sequence and the missing part center point local structural feature sequence;
[0015] Performing the second embedding operation in the weak-perception target detection network on the missing part center point sequence and the missing part center point local structure feature sequence to obtain a local shape feature sequence;
[0016] According to the local shape feature sequence, the local structure feature sequence of the center point of the missing part and the center point sequence of the missing part, the transformer decoder is used to perform the second iterative operation and feature transformation operation to obtain the shape feature sequence of the missing part of the weakly perceived target;
[0017] For the shape feature sequence of the missing part of the weakly perceived target, the folding operation is performed to obtain the complete point cloud data of the weakly perceived target by combining the center point sequence of the missing part and the point cloud data of the candidate target;
[0018] For the above-mentioned complete point cloud data of the weakly perceived target, the second point cloud feature encoding subnetwork is used to perform the above-mentioned feature extraction operation to obtain the above-mentioned complete shape feature sequence of the weakly perceived target.
[0019] Optionally, performing the sampling convolution operation and the first embedding operation in the weak-sensing target detection network according to the candidate target point cloud data to obtain a local structural feature sequence of the embedding position includes:
[0020] The candidate target point cloud data is used to obtain the center point sequence through the iterative farthest sampling method;
[0021] Utilize the above center point sequence to extract relevant feature sequences based on graph convolutional network;
[0022] The first embedding operation is performed on the center point sequence and the related feature sequence to obtain the local structure feature sequence of the embedding position.
[0023] Optionally, the dimensionality transformation operation includes a maximum pooling operation and a first multi-layer perceptron;
[0024] The above-mentioned local structural feature sequence based on the above-mentioned embedded position is used to perform the above-mentioned first iterative operation and dimension transformation operation using the above-mentioned transformer encoder to obtain the missing part center point sequence and the missing part center point local structural feature sequence, including:
[0025] Using the local structural feature sequence of the embedding position as the input of the transformer encoder to perform the first iterative operation to obtain the encoder output feature sequence, wherein the first iterative operation is encoded based on the first self-attention weighted operation and the first feedforward network;
[0026] According to the encoder output feature sequence, the maximum pooling operation and the first multi-layer perceptron are used to transform the feature dimension to obtain the missing part center point sequence and the missing part center point local structure feature sequence.
[0027] Optionally, the second iterative operation includes a second self-attention weighted operation, a cross-attention weighted operation and a second feedforward network;
[0028] The above-mentioned transformer decoder is used to perform the above-mentioned second iterative operation and feature transformation operation according to the above-mentioned local shape feature sequence, the above-mentioned missing part center point local structure feature sequence and the above-mentioned missing part center point sequence to obtain the above-mentioned weakly perceived target missing part shape feature sequence, including: using the above-mentioned local shape feature sequence as the first layer input shape feature sequence D of the above-mentioned transformer decoder (1) ;
[0029] The input shape feature sequence D of the kth layer of the above transformer decoder (k) Perform the second self-attention weighted operation to obtain the query vector E corresponding to the kth layer of the transformer decoder (k+1) ;
[0030] Based on the input shape feature sequence D of the kth layer of the above transformer decoder (k) The local structural feature sequence of the missing part center point and the missing part center point sequence are calculated by the following formula to obtain the mixed feature R of the kth layer of the transformer decoder (k) :
[0031] R (k) =Conv 2 (Cat(D (k) ,S)+Conv 1 (Y)),k=1,...,L
[0032] Among them, Cat represents the first concatenation operation, Conv 1 Represents the first convolution operation, Conv 2 represents the second convolution operation, Y represents the central point sequence of the missing part, S represents the local structural feature sequence of the central point of the missing part, and L represents the number of layers of the transformer decoder;
[0033] The above mixed feature R (k) As the key vector and value vector corresponding to the kth layer of the above transformer decoder, combined with the query vector E corresponding to the kth layer of the above transformer decoder (k+1), the cross-attention weighted shape feature U of the kth layer of the above transformer decoder is obtained by performing the cross-attention weighted operation as follows: (k+1) :
[0034]
[0035] Among them, H is the number of attention heads, and denote the first projection matrix, the second projection matrix, and the third projection matrix of the h-th attention head of the k-th layer of the transformer decoder, respectively. is the output linear projection matrix of the cross-attention weighted operation of the kth layer of the above transformer decoder, d is the scaling factor, T represents the matrix transpose operation, and δ represents the normalization operation;
[0036] Based on the cross-attention weighted shape feature U of the kth layer of the above transformer decoder (k+1) The output shape feature sequence D of the kth layer of the transformer decoder is obtained by using the second feedforward network. (k+1) ;
[0037] Based on the output shape feature sequence D of the kth layer of the above transformer decoder (k+1) The output shape feature sequence D of the last layer of the above transformer decoder is obtained by iterating (Lk) times through the remaining (Lk) layers of the above transformer decoder. (L+1) , as the decoder output feature sequence;
[0038] The feature transformation operation is performed on the feature sequence output by the decoder through a second multi-layer perceptron to obtain the shape feature sequence of the missing part of the weakly perceived target.
[0039] Optionally, the folding operation includes a second splicing operation, a third multilayer perceptron, and a third splicing operation;
[0040] The above-mentioned folding operation is performed on the shape feature sequence of the missing part of the weakly perceived target, combining the center point sequence of the missing part and the point cloud data of the candidate target, to obtain the complete point cloud data of the weakly perceived target, including:
[0041] Performing the second splicing operation on the missing part center point sequence and the weakly perceived target missing part shape feature sequence, and using the third multi-layer perceptron mapping on the splicing result of the second splicing operation to obtain the weakly perceived target missing part point cloud data;
[0042] The missing part of the point cloud data of the weakly perceived target and the point cloud data of the candidate target are subjected to the third splicing operation to obtain the complete point cloud data of the weakly perceived target.
[0043] Optionally, the attention fusion operation in the weakly perceived target detection network is performed based on the complete shape feature sequence of the weakly perceived target and the point cloud features of the candidate target to obtain the global features of the weakly perceived target, including:
[0044] Randomly collect the candidate target point cloud features to obtain an original sampling feature sequence, and use the corresponding neural network operation based on the original sampling feature sequence to obtain an original feature sequence;
[0045] Performing a fourth concatenation operation on the weakly perceived target complete shape feature sequence and the original feature sequence to obtain a concatenated feature sequence;
[0046] Perform channel-by-channel pooling and point-by-point pooling operations on the above concatenated feature sequence to obtain a point-by-point attention feature sequence and a channel-by-channel attention feature sequence respectively;
[0047] The point-by-point attention feature sequence and the channel-by-channel attention feature sequence are linearly transformed based on the first linear layer and the second linear layer, and then multiplied to obtain an attention feature product, and the attention feature product is normalized to obtain an overall attention weight map;
[0048] Multiply the above overall attention weight map and the above original feature sequence to redistribute the weights to obtain the original weighted feature sequence;
[0049] Obtaining a weighted feature sequence based on the original weighted feature sequence based on a third feedforward network;
[0050] A downsampling operation is performed on the weighted feature sequence to obtain geometric features, and corresponding neural network operations are used to encode the geometric features to obtain the weakly perceived target global features.
[0051] In a second aspect, the present application further proposes a weakly perceived target detection device, comprising:
[0052] A first acquisition unit is used to acquire a training sample set, input it into a weak-sensing target detection network, and perform preliminary detection through a first point cloud feature encoding subnetwork in the weak-sensing target detection network to acquire initial candidate data, wherein the initial candidate data includes an initial candidate box, candidate target point cloud data, and candidate target point cloud features;
[0053] an iterative operation unit, configured to perform iterative operations based on the candidate target point cloud data using the transformer encoder and the transformer decoder in the weakly perceived target detection network to obtain a shape feature sequence of a missing portion of the weakly perceived target, and to obtain a complete shape feature sequence of the weakly perceived target using a reconstruction operation in the weakly perceived target detection network according to the shape feature sequence of the missing portion of the weakly perceived target;
[0054] A fusion unit, configured to perform an attention fusion operation in the weakly perceived target detection network based on the complete shape feature sequence of the weakly perceived target and the point cloud features of the candidate target to obtain a global feature of the weakly perceived target;
[0055] a second acquisition unit, configured to perform confidence calculation and position regression operations in the weakly perceived target detection network based on the weakly perceived target global features to obtain a confidence score and a residual parameter of the weakly perceived target, and to calculate a loss value based on the confidence score and the residual parameter to adjust parameters of the weakly perceived target detection network to generate a weakly perceived target detection model;
[0056] The generation unit is used to use the above-mentioned weak perception target detection model to detect the sample set to be detected, generate a weak perception target detection frame and weak perception target category information, and complete weak perception target detection.
[0057] In a third aspect, an electronic device comprises: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor is used to implement the steps of the weak-perception target detection method of any one of the first aspects described above when executing the computer program stored in the memory.
[0058] In a fourth aspect, the present invention further proposes a computer-readable storage medium having a computer program stored thereon, and when the computer program is executed by a processor, the weak-perception target detection method of any one of the above items in the first aspect is implemented.
[0059] In summary, the present application proposes a weakly perceived target detection method, which inputs a training sample set into a weakly perceived target detection network, uses a first point cloud feature encoding subnetwork to obtain initial candidate data in the scene point cloud, and performs iterative operations through an improved transformer encoder and a transformer decoder to complete and reconstruct the overall shape of the candidate target point cloud data in the initial candidate frame to enhance the shape features of the weakly perceived target, and based on an improved attention fusion operation, aggregates the complete shape features of the weakly perceived target in the completed initial candidate frame and the original feature sequence containing the position information of the weakly perceived target in the initial candidate frame before completion to obtain the global features of the weakly perceived target with rich spatial geometric information, calculates the confidence score and residual parameters of the weakly perceived target based on the global features of the weakly perceived target, calculates and updates the loss value with the real label data in the training sample set to adjust the parameters of the weakly perceived target detection network, learns to obtain a weakly perceived target detection model to detect the sample set to be detected, generates a weakly perceived target detection frame and weakly perceived target category information, and completes the weakly perceived target detection. The completion and reconstruction operation in this weakly perceived target detection method is implemented through a structure-aware transformer model, which extracts the structural features of the candidate target point cloud data through the transformer encoder to obtain the local structural feature sequence of the missing part of the center point, and constructs the key-value vector corresponding to the current layer of the transformer decoder by aggregating the local structural feature sequence of the missing part of the center point and the shape features output by the previous layer of the transformer decoder, so that the query vector can query more low-level structural details, thereby guiding the decoder to generate a more accurate complete shape feature sequence of the weakly perceived target. Different from the existing splicing feature aggregation module, the attention fusion operation in this detection method can fuse semantic features of different scales, and redistribute the weights of the points by fusing the complete shape features of the weakly perceived target obtained by completion and the original feature sequence, so as to enhance the weight of the target key points and suppress the interference of non-key points on the detection performance. This detection method effectively improves the detection accuracy of weakly perceived targets by combining the completion and reconstruction operation with the attention fusion operation.
[0060] The weak-perception target detection method of the present invention, and other advantages, objectives and features of the present invention will be reflected in part through the following description, and in part will also be understood by technicians in this field through research and practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0061] Various other advantages and benefits will become apparent to those of ordinary skill in the art by reading the detailed description of the preferred embodiments below. The accompanying drawings are only for the purpose of illustrating the preferred embodiments and are not to be considered as limiting the present specification. Also, the same reference symbols are used throughout the accompanying drawings to represent the same components. In the accompanying drawings:
[0062] Figure 1 A schematic diagram of a weakly perceived target detection method provided in an embodiment of the present application;
[0063] Figure 2 A schematic diagram of a process for generating initial candidate data of a three-dimensional target provided in an embodiment of the present application;
[0064] Figure 3 A schematic diagram of a weakly perceived target detection model structure provided in an embodiment of the present application;
[0065] Figure 4 A schematic diagram comparing the detection accuracy of a weakly perceived target detection method provided in an embodiment of the present application and a benchmark method;
[0066] Figure 5 A schematic diagram of the effect of using the method to detect weakly perceived targets provided in an embodiment of the present application;
[0067] Figure 6 A weakly perceived target detection device provided in an embodiment of the present application;
[0068] Figure 7 A schematic diagram of the structure of a weak-sensing target detection electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0069] The terms "first", "second", "third", "fourth", etc. (if any) in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments described herein can be implemented in an order other than that illustrated or described herein. In addition, the terms "including" and "having" and any of their variations are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units that are clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices. The technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all of the embodiments.
[0070] See also Figure 1 , is a flow chart of a weakly perceived target detection method in an embodiment of the present application, the method comprising:
[0071] S110, obtaining a training sample set, inputting it into a weak-sensing target detection network, and performing preliminary detection through the first point cloud feature encoding subnetwork in the weak-sensing target detection network to obtain initial candidate data, wherein the initial candidate data includes an initial candidate box, candidate target point cloud data, and candidate target point cloud features;
[0072] For example, the training sample set is N×C n Scene point cloud Y scene Input to the weak perception target detection network, using Figure 2 The first point cloud feature encoding subnetwork PointNet++ in the initial candidate data generation module shown extracts features to obtain discriminative point-by-point features F of size N×C. scene The foreground points of the target are segmented and the region proposal network (RPN) in the weak-sensing target detection network performs preliminary detection to obtain the scene point cloud Y scene The length, width, height, angle and center point coordinates of the 3D candidate box corresponding to the foreground point of the target are obtained by regression and the classification probability is obtained. The 3D candidate box is generated by regression, and the ratio of the overlapping area and the merged area between the 3D candidate box and the real target box in the training sample set is calculated. Z 3D candidate boxes with a ratio greater than 0.55 are selected, and the parameters of the Z 3D candidate boxes are used to calculate the overlap area and the merged area of the scene point cloud Y. scene Z candidate targets are found in the scene point cloud. Since the number of weakly perceived targets in the scene point cloud is too small, they lack sufficient feature representation and are difficult to be perceived. Therefore, M candidate targets with sparse point clouds are selected from the Z candidate targets as part of the weakly perceived targets. Since the number of M candidate targets is too small and the spatial geometric information is lacking, the position parameters of the corresponding M candidate frames are inaccurate. Therefore, the candidate data in the M candidate frames are only part of the actual weakly perceived targets. The candidate frames corresponding to the M sparsely perceived targets constitute a M×C p The initial candidate box set, where the size of each initial candidate box in the initial candidate box set is M 0 ×C p , for a certain initial candidate box P, the point cloud data in P is the candidate target point cloud data Y in the initial candidate box P known , according to P in the point-by-point feature F scene The candidate target point cloud feature F in the initial candidate box P can be obtained by cropping known ; Initial candidate box P, candidate target point cloud data Y in the initial candidate box P known And the candidate target point cloud features F in the initial candidate box P known is the initial candidate data, where Y known Size is N 1 ×C n , F known Size is N 1×C, M is the number of initial candidate boxes in the initial candidate box set, M 0 is the dimension of P, C p is the initial candidate box set and the dimension of the initial candidate box P; N is the scene point cloud Y scene , point-wise features F scene The number of midpoints, C is the point-by-point feature F scene , candidate target point cloud features F known The feature dimension, N 1 The candidate target point cloud data Y known , candidate target point cloud features F known The number of midpoints, C n Y scene , Y known Dimensions of the midpoint, Z, N, C n , C, M, M 0 , C p , N 1 are all positive integers. For example, the scene point cloud Y with a size of 16384×3 in the training sample set scene The input is sent to the weak perception target detection network. First, the first point cloud feature encoding subnetwork PointNet++ in the weak perception target detection network extracts features to obtain a point-by-point feature F with a size of 16384×128. scene , F scene After preliminary detection by RPN in the weak perception target detection network, the classification probability of 16384×1 and the regression parameter of 16384×7 corresponding to the 3D candidate box are obtained, and the 3D candidate box is generated based on the classification probability and regression parameter. Then, 32 3D candidate boxes with an IOU ratio greater than 0.55 are selected and mapped on the scene point cloud Y according to the parameters of the 32 3D candidate boxes. scene 32 candidate targets are found in the dataset, and 10 candidate targets with point cloud numbers between 200 and 2048 are selected from the 32 candidate targets. For the candidate targets with less than 2048 points, the number of points is fixed to 2048 by filling zero values. The 10 3D candidate boxes corresponding to the 10 candidate targets constitute an initial candidate box set with a size of 10×7. The size of each initial candidate box in the initial candidate box set is 1×7. For a certain initial candidate box P, the point cloud data of the candidate target with a size of 2048×3 in P is the candidate target point cloud data Y. known , according to the initial candidate box P in the point-by-point feature F scene Cropping in P can obtain the candidate target point cloud feature F with a size of 2048×128 known .
[0073] S120, based on the candidate target point cloud data, using the transformer encoder and the transformer decoder in the weakly perceived target detection network to perform iterative operations to obtain a shape feature sequence of the missing part of the weakly perceived target, and using the reconstruction operation in the weakly perceived target detection network to obtain a complete shape feature sequence of the weakly perceived target according to the shape feature sequence of the missing part of the weakly perceived target;
[0074] Exemplarily, according to the candidate target point cloud data Y obtained in step S110 known The transformer encoder in the weak-sensing target detection network and the transformer decoder in the weak-sensing target detection network are iterated to obtain a size of N 2 ×C 3 The shape feature sequence D of the weakly perceived target is missing part fine , and the shape feature sequence D of the weakly perceived target is missing fine Perform the reconstruction operation in the weak perception target detection network to obtain a size of N 5 ×I 1 The complete shape feature sequence F of the weakly perceived target com , where N 2 The shape feature sequence D is missing for the weakly perceived target. fine The number of midpoints, C 3 D fine The feature dimension, N 5 is the complete shape feature sequence F of the weakly perceived target com The number of midpoints, I 1 F com The feature dimension, N 5 , I 1 , N 2 , C 3 All are positive integers.
[0075] It should be noted that the transformer encoder and transformer decoder in this application are improvements on the standard transformer, and a structure-aware transformer model with strong global contextual structure feature extraction capability is proposed to realize the weakly perceived target completion and reconstruction operation, thereby completing the shape point set of the weakly perceived target and enhancing its structural shape information to improve detection accuracy.
[0076] S130, performing an attention fusion operation in the weakly perceived target detection network based on the complete shape feature sequence of the weakly perceived target and the point cloud features of the candidate target to obtain a global feature of the weakly perceived target;
[0077] Exemplarily, the candidate target point cloud feature F in the initial candidate box P obtained in step S110 known and the complete shape feature sequence F of the weakly perceived target obtained in step S120 com , perform the attention fusion operation in the weakly perceptual target detection network to obtain a B×I 3 The weakly perceived target global feature F glo , where B is F glo The corresponding points, I 3 F glo The characteristic dimensions, B, I 3 All are positive integers.
[0078] S140, performing confidence calculation and position regression operations in the weakly perceived target detection network based on the weakly perceived target global features to obtain a confidence score and a residual parameter of the weakly perceived target, and calculating a loss value based on the confidence score and the residual parameter to adjust parameters of the weakly perceived target detection network, thereby generating a weakly perceived target detection model;
[0079] Exemplarily, the weakly perceived target global feature F obtained by example in step S130 glo , respectively perform confidence calculation and position regression operations in the weak perception target detection network, and generate a size M corresponding to the weak perception target 0 ×M 0 The confidence score cls and size M 0 ×C p The residual parameter reg of the weakly perceived target is generated according to the confidence score cls, and the classification loss is calculated using the cross entropy loss function based on the category information and the true target category label in the training sample set. The residual target is calculated based on the initial candidate box and the true target box in the training sample set, and the smooth-L1 loss is calculated based on the residual parameter reg and the residual target to obtain the position regression loss. The parameters of the weakly perceived target detection network are constrained and adjusted based on the sum of the classification loss and the position regression loss in the training sample set, and the weakly perceived target detection model is learned and generated. 0 is the dimension of cls and reg, C p is the dimension of reg, M 0 , C p All are positive integers.
[0080] S150, using the above-mentioned weak-perception target detection model to detect the sample set to be detected, generating a weak-perception target detection frame and weak-perception target category information, and completing weak-perception target detection.
[0081] For example, a sample set to be detected is obtained, and a size of N×C in the sample set to be detected is nThe scene point cloud is input into the weakly perceived target detection model with the optimal parameters of the weakly perceived target detection network. First, step S110 is executed for preliminary detection to obtain the initial candidate data in the sample set to be detected, including the initial candidate box, candidate target point cloud data and candidate target point cloud features in the sample set to be detected. Then, step S120 is executed on the initial candidate data in the sample set to be detected to reconstruct the complete shape feature sequence of the weakly perceived target in the sample set to be detected. Then, based on step S130, an attention fusion operation is performed on the complete shape feature sequence of the weakly perceived target in the sample set to be detected and the candidate target point cloud features to obtain the global features of the weakly perceived target in the sample set to be detected. Finally, the confidence calculation and position regression operation in step S140 are performed to obtain the weakly perceived target category information and the size of each of the Q weakly perceived targets in the sample set to be detected. 0 ×C p The weakly perceived target detection box is constructed to complete the weakly perceived target detection, where N is the number of points in the scene point cloud in the sample set to be detected, and C n is the dimension of the point in the scene point cloud in the sample set to be detected, M 0 is the dimension of the target detection box, C p is the dimension of the target detection box. N, C n , Q, M 0 , C pare all positive integers. For example, a scene point cloud of size 16384×3 in the sample set to be detected is input into a weakly perceived target detection model with optimal parameters of a weakly perceived target detection network, and the global features of the weakly perceived targets in the sample set to be detected are obtained by sequentially executing the preliminary detection of step S110, step S120, and step S130, and then the confidence calculation in step S140 is performed based on the global features of the weakly perceived targets in the sample set to be detected to obtain a confidence score, and the category information of the 8 weakly perceived targets in the sample set to be detected is determined based on the confidence score, and the confidence score is calculated based on the global features of the weakly perceived targets in the sample set to be detected. The target detection frames of 8 weakly perceived targets of size 1×7 in the sample set to be detected are obtained by the position regression operation in step S140; taking a certain weakly perceived target in the sample set to be detected as an example, the confidence score of the weakly perceived target is 0.98 after the confidence calculation operation, and the category information of the weakly perceived target is judged to be "vehicle" based on the confidence score. The corresponding weakly perceived target detection frame with a length of 4.35 meters, a width of 1.76 meters, a height of 1.78 meters, a center point at (2.36 meters, 1.59 meters, 19.10 meters), and an angle of 38 degrees is generated by the position regression operation to complete the weakly perceived target detection. In summary, the present application proposes a weakly perceived target detection method, which inputs a training sample set into a weakly perceived target detection network, first uses a first point cloud feature encoding subnetwork to obtain initial candidate data in the scene point cloud, and then uses an improved transformer encoder and transformer decoder to perform iterative operations to complete and reconstruct the overall shape of the candidate target point cloud data in the initial candidate frame to enhance the shape features of the weakly perceived target, and based on the improved attention fusion operation, aggregates the complete shape features of the weakly perceived target in the completed initial candidate frame and the original feature sequence containing the position information of the weakly perceived target in the initial candidate frame before completion to obtain the global features of the weakly perceived target with rich spatial geometric information, calculates the confidence score and residual parameters of the weakly perceived target based on the global features of the weakly perceived target, and calculates and updates the loss value with the real label data in the training sample set to adjust the parameters of the weakly perceived target detection network, learns to obtain a weakly perceived target detection model to detect the sample set to be detected, generates a weakly perceived target detection frame and weakly perceived target category information, and completes the weakly perceived target detection.The completion and reconstruction operation in this weakly perceived target detection method is implemented through a structure-aware transformer model, which extracts the structural features of the candidate target point cloud data through the transformer encoder to obtain the local structural feature sequence of the missing part of the center point, and constructs the key-value vector corresponding to the current layer of the transformer decoder by aggregating the local structural feature sequence of the missing part of the center point and the shape features output by the previous layer of the transformer decoder, so that the query vector can query more low-level structural details, thereby guiding the decoder to generate a more accurate complete shape feature sequence of the weakly perceived target. Different from the existing splicing feature aggregation module, the attention fusion operation in this detection method can fuse semantic features of different scales, and redistribute the weights of the points by fusing the complete shape features of the weakly perceived target obtained by completion and the original feature sequence, so as to enhance the weight of the target key points and suppress the interference of non-key points on the detection performance. This detection method effectively improves the detection accuracy of weakly perceived targets by combining the completion and reconstruction operation with the attention fusion operation.
[0082] In some examples, the iterative operation includes a first iterative operation and a second iterative operation, and the reconstruction operation includes a folding operation and a feature extraction operation;
[0083] The above step S120 may specifically include: step S1201 to step S1206;
[0084] S1201, performing a sampling convolution operation and a first embedding operation in the weakly perceptual target detection network according to the candidate target point cloud data to obtain a local structural feature sequence of the embedding position;
[0085] Exemplarily, the candidate target point cloud data Y obtained in step S110 is known Perform the sampling convolution operation in the weak perception target detection network and the first embedding operation in the weak perception target detection network to obtain a size of N 2 ×C 2 The local structural feature sequence F of the embedded position (1) , where N 2 is the local structural feature sequence F of the embedding position (1) The number of midpoints, C 2 is the local structural feature sequence F of the embedding position (1) The feature dimension, N 2 , C 2 All are positive integers.
[0086] In some examples, the above step S1201 specifically includes: step S12011 to step S12013;
[0087] S12011, using the candidate target point cloud data to obtain a center point sequence by an iterative farthest sampling method;
[0088] S12012. Using the center point sequence, extract relevant feature sequences based on a graph convolutional network;
[0089] S12013. Perform the first embedding operation on the center point sequence and the related feature sequence to obtain the local structural feature sequence of the embedding position.
[0090] Exemplarily, the candidate target point cloud data Y obtained in step S110 is known The size N is obtained by iterative farthest sampling method 2 ×C n The center point sequence X, based on X and Y known Use graph convolutional network to obtain a size of N 2 ×C 1 The first embedding operation is performed on the center point sequence X and the related feature sequence G, that is, the feature dimension of the center point sequence X is equal to the feature dimension of the related feature sequence G through the corresponding neural network mapping transformation to obtain the projected center point feature, and then the projected center point feature and the related feature sequence G are spliced in the channel dimension to obtain the projection splicing result, and then the projection splicing result is passed through the corresponding neural network to obtain the local structure feature sequence F of the embedded position. (1) , where N 2 is the number of center point sequence X and related feature sequence G points, C n is the dimension of the midpoint of X, C 1 is the feature dimension of the related feature sequence G, N 2 , C n , C 1 All are positive integers.
[0091] For example, for the candidate target point cloud data Y of size 2048×3 in the initial candidate box P obtained in the example of step S110 known , using the iterative farthest point sampling method from Y known The center point sequence X with a size of 128×3 is obtained by sampling, and then the relevant feature sequence G with a size of 128×512 is extracted based on the graph convolutional network. The center point sequence X is embedded into the relevant feature sequence G to obtain the embedded position local structure feature sequence F with a size of 128×384. (1) .
[0092] S1202: Based on the local structural feature sequence of the embedded position, use the transformer encoder to perform the first iterative operation and dimension transformation operation to obtain a missing part center point sequence and a missing part center point local structural feature sequence;
[0093] Exemplarily, the local structural feature sequence F of the embedded position obtained in step S1201 is (1) The transformer encoder in the weakly perceptual target detection network performs the first iteration operation and dimension transformation operation to obtain a size of N 2 ×C n The missing part of the center point sequence Y and size N 2 ×C 3 The local structural feature sequence S of the missing part center point, where N 2 is the number of points in the missing part center point sequence Y and the missing part center point local structure feature sequence S, C n C is the dimension of the midpoint of the missing part of the center point sequence Y. 3 is the feature dimension of the local structural feature sequence S of the missing part of the center point, N 2 , C 3 , C n All are positive integers.
[0094] In some examples, the dimensionality transformation operation includes a maximum pooling operation and a first multi-layer perceptron;
[0095] The above step S1202 may specifically include: step S12021 and step S12022;
[0096] S12021, using the above-mentioned local structure feature sequence of the embedded position as the input of the above-mentioned transformer encoder to perform the above-mentioned first iterative operation to obtain the encoder output feature sequence, wherein the above-mentioned first iterative operation is encoded based on the first self-attention weighted operation and the first feedforward network;
[0097] S12022. According to the encoder output feature sequence, use the maximum pooling operation and the first multi-layer perceptron to transform the feature dimension to obtain the missing part center point sequence and the missing part center point local structure feature sequence.
[0098] Exemplarily, the local structural feature sequence F of the embedded position obtained in step S1201 is (1) As the input feature sequence F of the first layer of the transformer encoder in the weakly perceptual object detection network (1) After L layers of transformer encoders, the first iterative operation is iteratively updated and refined. The first iterative operation includes L calculations as shown in formula (1). The input feature sequence F of the kth layer of the transformer encoder is (k)The first self-attention weighted operation is used to calculate the first self-attention of the encoder, and the first self-attention weighted feature sequence of the kth layer of the transformer encoder is generated. The first self-attention weighted feature sequence is refined through the first feedforward network to generate the attention weighted structural feature sequence F of the kth layer of the transformer encoder. (k+1) , F (k+1) is the output feature sequence of the kth layer of the transformer encoder, F (k) The size is N 2 ×C 2 , F (k+1) The size is N 2 ×C 2 , N 2 F (k) 、F (k+1) The number of midpoints, C 2 F (k) 、F (k+1) The feature dimension, N 2 , C 2 All are positive integers.
[0099] F (k+1) =Self_Att 1 (F (k) )+FFN 1 (Self_At t 1(F (k) )),k=1,...,L (1)
[0100] Among them, FFN 1 Represents the first feedforward network, Self_Att 1 The first self-attention weighted operation for the encoder's first self-attention calculation is shown in formula (2):
[0101]
[0102] Among them, H 1 is the number of self-attention heads of the first self-attention weighted operation, both of size C 2 ×(C 2 / H 1 )of denotes the projection matrix of the h-th self-attention head of the k-th layer of the transformer encoder, C 2 , (C 2 / H 1 )for The dimension is C 2 ×C 2 of is the output linear projection matrix of the first self-attention weighted operation of the kth layer of the transformer encoder, C 2 yes The dimensions of C 2 , H 1 are all positive integers;
[0103] F (k+1) It is sent to the k+1th layer of the transformer encoder and continues to be iteratively updated as the input feature sequence of the k+1th layer of the transformer encoder until the Lth layer;
[0104] According to the input feature sequence F of the first layer of the transformer encoder (1) Using L-layer transformer encoder iterative calculation, we get a size of N 2 ×C 2 The encoder outputs the feature sequence F (L+1) , where N 2 Output feature sequence F for the encoder (L +1) The number of midpoints, C 2 F (L+1) The feature dimension, N 2 , C 2 All are positive integers.
[0105] F (L+1) The maximum pooling operation and the first multi-layer perceptron are used to transform the feature dimension to obtain the missing part center point sequence Y and the missing part center point local structure feature sequence S.
[0106] For example, the embedded position local structure feature sequence F with a size of 128×384 obtained in step S1201 is (1) Input to the three-layer transformer encoder, and perform the first calculation as shown in formula (1) in the first layer of the transformer encoder to obtain the output feature sequence F of the first layer of the transformer encoder (2) , that is, F (1) First, the first self-attention weighted operation is used to perform the first self-attention calculation of the encoder with 8 self-attention heads. In the 8th self-attention head of the first layer of the transformer encoder, the input feature sequence F with a size of 128×384 is (1)Multiply them with the projection matrices of sizes 384×48, 384×48, and 384×48 respectively to obtain the first eigenvector, second eigenvector, and third eigenvector of sizes 128×48, 128×48, and 128×48. The first self-attention feature of size 128×48 is obtained by multiplying the first eigenvector, the second eigenvector, and the third eigenvector. Then, the 8 first self-attention features calculated by the 8 self-attention heads are spliced to obtain the attention splicing result of size 128×384, and the attention splicing result is combined with the output linear projection matrix of size 384×384. Multiply them together to get the first self-attention weighted feature sequence of size 128×384, and then update the first feedforward network to get the output feature sequence F of the first layer of the transformer encoder of size 128×384 (2) , F (2) As the input feature sequence of the second layer of the transformer encoder, the second layer of the transformer encoder performs a second calculation as shown in formula (1), and outputs the output feature sequence F of the second layer of the transformer encoder with a size of 128×384 (3) , F (3) As the input feature sequence of the third layer of the transformer encoder, the third layer of the transformer encoder performs the third calculation as shown in formula (1) to obtain the encoder output feature sequence F with a size of 128×384 (4) , then use the maximum pooling operation and the first multi-layer perceptron transformation F (4) The feature dimension is , and the missing part center point sequence Y with a size of 128×3 and the missing part center point local structure feature sequence S with a size of 128×1024 are obtained.
[0107] S1203, performing the second embedding operation in the weakly perceptual object detection network on the missing part center point sequence and the missing part center point local structure feature sequence to obtain a local shape feature sequence;
[0108] Exemplarily, the second embedding operation in the weakly perceptual target detection network is performed on the missing part center point sequence Y and the missing part center point local structure feature sequence S obtained in the example of step S1202 to obtain a size of N 2 ×C 2 The local shape feature sequence D (1) , where N 2 D (1) The number of midpoints, C 2 D (1) The feature dimension, N 2 , C 2For example, the second embedding operation in the weakly perceptual target detection network is performed on the missing part center point sequence Y of size 128×3 and the missing part center point local structure feature sequence S of size 128×1024 obtained in step S1202 to obtain a local shape feature sequence D of size 128×384. (1) .
[0109] S1204, performing the second iterative operation and feature transformation operation using the transformer decoder according to the local shape feature sequence, the local structure feature sequence of the center point of the missing part, and the center point sequence of the missing part, to obtain the shape feature sequence of the missing part of the weakly perceived target;
[0110] Exemplarily, the local shape feature sequence D obtained in the example of step S1203 is (1) The transformer decoder in the weakly perceived target detection network is used to perform the second iteration operation and feature transformation operation to obtain the shape feature sequence D of the missing part of the weakly perceived target. fine .
[0111] In some examples, the second iterative operation includes a second self-attention weighted operation, a cross-attention weighted operation, and a second feedforward network;
[0112] The above step S1204 may specifically include: step S12041 and step S12047;
[0113] S12041, using the above local shape feature sequence as the first layer input shape feature sequence D of the above transformer decoder (1) ;
[0114] S12042, the input shape feature sequence D of the kth layer of the transformer decoder (k) Perform the second self-attention weighted operation to obtain the query vector E corresponding to the kth layer of the transformer decoder (k+1) ;
[0115] S12043, based on the input shape feature sequence D of the kth layer of the above transformer decoder (k) The local structural feature sequence of the missing part center point and the missing part center point sequence are calculated by the following formula to obtain the mixed feature R of the kth layer of the transformer decoder (k) :
[0116] R (k) =Conv 2 (Cat(D (k),S)+Conv 1 (Y)),k=1,...,L
[0117] Among them, Cat represents the first concatenation operation, Conv 1 Represents the first convolution operation, Conv 2 represents the second convolution operation, Y represents the central point sequence of the missing part, S represents the local structural feature sequence of the central point of the missing part, and L represents the number of layers of the transformer decoder;
[0118] S12044, the above mixed feature R (k) As the key vector and value vector corresponding to the kth layer of the above transformer decoder, combined with the query vector E corresponding to the kth layer of the above transformer decoder (k+1) , the cross-attention weighted shape feature U of the kth layer of the above transformer decoder is obtained by performing the cross-attention weighted operation as follows: (k+1) :
[0119]
[0120] Among them, H is the number of attention heads, and denote the first projection matrix, the second projection matrix, and the third projection matrix of the h-th attention head of the k-th layer of the transformer decoder, respectively. is the output linear projection matrix of the cross-attention weighted operation of the kth layer of the above transformer decoder, d is the scaling factor, T represents the matrix transpose operation, and δ represents the normalization operation.
[0121] S12045, based on the cross-attention weighted shape feature U of the kth layer of the above transformer decoder (k+1) The output shape feature sequence D of the kth layer of the transformer decoder is obtained by using the second feedforward network. (k+1) ;
[0122] S12046, based on the output shape feature sequence D of the kth layer of the above transformer decoder (k+1) The output shape feature sequence D of the last layer of the above transformer decoder is obtained by iterating (Lk) times through the remaining (Lk) layers of the above transformer decoder. (L+1) , as the decoder output feature sequence;
[0123] S12047. Perform the feature transformation operation through a second multi-layer perceptron according to the feature sequence output by the decoder to obtain the shape feature sequence of the missing part of the weakly perceived target.
[0124] Exemplarily, construct an L-layer transformer decoder to decode the local shape feature sequence D obtained in step S1203. (1) To refine, the process is as follows:
[0125] D (1) The shape feature sequence of the first layer input to the transformer decoder in the weak perception target detection network is combined with the local structural feature sequence S of the missing part center point and the missing part center point sequence Y. The L-layer transformer decoder performs the second iterative operation to continuously refine the shape features, wherein the second iterative operation includes L calculations, where L is a positive integer; the kth calculation process in the second iterative operation is as follows:
[0126] The size of the kth layer of the transformer decoder in the weakly perceptual object detection network is N 2 ×C 2 The input shape feature sequence D (k) The second self-attention weight calculation of the decoder is performed through the second self-attention weighted operation as shown in formula (3), and the size of the k-th layer of the transformer decoder is obtained as N 2 ×C 2 The second self-attention weighted shape feature E (k+1) , where N 2 For E (k+1) , D (k) The number of midpoints, C 2 For E (k+1) , D (k) The feature dimension, N 2 , k, C 2 All are positive integers:
[0127] E (k+1) =Self_Att(D (k) ),k=1,...,L (3)
[0128] in, H 2 is the number of the second self-attention heads, both of size C 2 ×(C 2 / H 2 )of denotes the projection matrix of the h-th second self-attention head of the k-th layer of the transformer decoder, C 2 , (C 2 / H 2) is the projection matrix The dimension is C 2 ×C 2 of is the output linear projection matrix of the second self-attention weighted operation of the k-th layer of the transformer decoder, C 2 for The dimension, C 2 , H 2 are all positive integers;
[0129] The input shape feature sequence D of the kth layer of the transformer decoder (k) spliced to the missing part center point local structure feature sequence S obtained in the example of step S1202 (i.e., the first splicing operation) to obtain a size of N 2 ×C 4 The additive characteristic Cat(D (k) , S), and embed the missing part center point sequence Y obtained in the example of step S1202 into the additive feature Cat (D (k) ,S), the size of the k-th layer of the transformer decoder is N 2 ×C 2 The mixed feature R (k) , where N 2 YesR (k) Cat(D (k) ,S) the number of midpoints, C 2 YesR (k) The characteristic dimension, C 4 Represents the additive feature Cat(D (k) ,S) feature dimension, N 2 , C 2 , C 4 All are positive integers.
[0130] R (k) =Conv 2 (Cat(D (k) ,S)+Conv 1 (Y)),k=1,...,L (4)
[0131] Among them, Cat represents the first concatenation operation, Conv 1 Represents the first convolution operation, Conv 2 Represents the second convolution operation, Conv 1 (Y) indicates that the dimension of the missing part center point sequence Y is transformed by the first convolution operation, so that the dimension of the missing part center point sequence Y is equal to the additive feature Cat(D (k),S) have the same feature dimensions, resulting in a size of N 2 ×C 4 The position feature Y′ is combined with the additive feature Cat(D (k) ,S) and then use the second convolution operation Conv 2 Transform the feature dimension to get R (k) , N 2 represents the number of points in Y′, C 4 represents the feature dimension of Y′, N 2 , C 4 All are positive integers.
[0132] The second self-attention weighted shape feature E of the k-th layer of the transformer decoder (k+1) As the query vector for the cross-attention weighted operation, and the mixed feature R of the k-th layer of the transformer decoder (k) As the key vector and value vector of the cross-attention weighted operation, the decoder cross-attention calculation is performed as shown in formula (5), and the size of the k-th layer of the transformer decoder is N 2 ×C 2 The cross-attention weighted shape feature U (k+1) , where N 2 , C 2 For U (k+1) The feature dimension, N 2 , C 2 All are positive integers:
[0133]
[0134] Among them, H is the number of attention heads, and the size is C 2 ×(C 2 / H) and They represent the first, second, and third projection matrices of the h-th attention head of the k-th layer of the transformer decoder, respectively, with a size of C 2 ×C 2 of is the output linear projection matrix of the cross-attention weighted operation of the kth layer of the transformer decoder, d is the scaling factor, T represents the matrix transpose operation, δ represents the normalization operation, and C 2 yes and The dimension (C 2 / H)Yes The dimension, C 2 , H, d are all positive integers;
[0135] The cross-attention weighted shape feature U of the k-th layer of the transformer decoder (k+1) The second feedforward network of the kth layer of the transformer decoder is refined, as shown in formula (6), and the size of the kth layer of the transformer decoder is N 2 ×C 2 The output shape feature sequence D (k+1) , where N 2 D (k+1) The number of midpoints, C 2 D (k+1) The feature dimension, N 2 , C 2 All are positive integers;
[0136] D (k+1) =U (k+1) +FFN(U (k+1) ) (6)
[0137] Among them, FFN represents the cross-attention weighted shape feature U of the k-th layer of the transformer decoder (k+1) A second feed-forward network for refinement;
[0138] D (k+1) As the input shape feature sequence of the k+1th layer of the transformer decoder, it continues to be iteratively updated until the Lth layer;
[0139] After the above L layers of transformer decoders, the first layer input shape feature sequence D of the transformer decoder is updated and refined (1) , get the size N 2 ×C 2 The decoder outputs the feature sequence D (L+1) , and then transform D through the second multi-layer perceptron (L+1) The feature dimension of the weakly perceived target is generated by the shape feature sequence D of the missing part. fine , where N 2 D (L+1) The number of midpoints, C 2 D (L+1) The feature dimension, N 2 , C 2 All are positive integers.
[0140] For example, the local shape feature sequence D of size 128×384 obtained in step S1203 is (1)The output shape feature sequence D of the first layer of the transformer decoder is obtained by inputting the three-layer transformer decoder in the weak perception target detection network, combining the local structural feature sequence S of the missing part center point and the missing part center point sequence Y obtained in the example of step S1202, and performing the first calculation in the second iterative operation in the first layer of the transformer decoder. (2) , that is, D (1) First, the second self-attention weighted shape feature E of the first layer of the transformer decoder with a size of 128×384 is obtained through the second self-attention weighted operation (2) , and then concatenate the input shape feature sequence D of the first layer of the transformer decoder (1) The additive feature is obtained by combining the local structural feature sequence S of the missing part center point, and then the dimension of the missing part center point sequence Y is transformed by the first convolution operation to obtain Y′, Y′ is added to the additive feature, and then the mixed feature R with a size of 128×384 is obtained by the second convolution operation. (1) , the second self-attention weighted shape feature E (2) With mixed features R (1) Perform a cross-attention weighted operation to obtain a cross-attention weighted shape feature U of size 128×384 (2) , U (2) After the second feed-forward network of the first layer of the transformer decoder is updated, the output shape feature sequence D of the first layer of the transformer decoder with a size of 128×384 is obtained. (2) , D (2) As the input shape feature sequence of the second layer of the transformer decoder, combined with the local structural feature sequence S of the missing part center point and the missing part center point sequence Y, the second layer of the transformer decoder performs the second calculation in the second iteration operation to obtain the output shape feature sequence D of the second layer of the transformer decoder with a size of 128×384 (3) , D (3) The input is sent to the third layer of the transformer decoder, and the local structural feature sequence S of the missing part center point and the missing part center point sequence Y are combined. The third layer of the transformer decoder performs the third calculation in the second iteration operation to obtain the output shape feature sequence D of the third layer (the last layer) of the transformer decoder with a size of 128×384. (4) (i.e. the decoder outputs a feature sequence), and then the decoder outputs a feature sequence D through the second multi-layer perceptron transformation (4)The feature dimension is 128×1024, generating a shape feature sequence D of the missing part of the weakly perceived target. fine .
[0141] S1205, for the shape feature sequence of the missing part of the weakly perceived target, combining the center point sequence of the missing part and the point cloud data of the candidate target, performing the folding operation to obtain the complete point cloud data of the weakly perceived target;
[0142] For example, for a weakly perceived target, part of the shape feature sequence D is missing. fine , combined with the missing part center point sequence Y and the above candidate target point cloud data Y known , perform a folding operation to obtain a size of N 4 ×C n The complete point cloud data Y of the weakly perceived target completed , where N 4 Y completed The number of midpoints, C n Y completed The dimension of the midpoint, N 4 , C n All are positive integers.
[0143] In some examples, the folding operation includes a second splicing operation, a third multilayer perceptron, and a third splicing operation; the step S1205 may specifically include: step S12051 to step S12052;
[0144] S12051, performing the second splicing operation on the missing part center point sequence and the weakly perceived target missing part shape feature sequence, and using the third multi-layer perceptron mapping on the splicing result of the second splicing operation to obtain the weakly perceived target missing part point cloud data;
[0145] S12052. Perform the third splicing operation on the missing part of the point cloud data of the weakly perceived target and the point cloud data of the candidate target to obtain the complete point cloud data of the weakly perceived target.
[0146] Exemplarily, the missing part center point sequence Y obtained in the example of step S1202 and the shape feature sequence D of the missing part of the weakly perceived target obtained in the example of step S1204 are combined. fine The splicing result obtained by the second splicing operation is input into the third multi-layer perceptron to obtain the missing part of the point cloud data Y of the weakly perceived target missing , size is N 3 ×C n , Y missing and the candidate target point cloud data Y obtained in the example of step S110 known After the third splicing operation, the complete point cloud data Y of the weakly perceived target is obtainedcompleted , where N 3 Y missing The number of midpoints, C n Y missing The dimension of the midpoint, N 3 , C n For example, the missing part center point sequence Y of size 128×3 obtained in the example of step S1202 is concatenated with the shape feature sequence D of the missing part of the weakly perceived target of size 128×1024 obtained in the example of step S1204. fine , and the splicing result with a size of 128×1027 is obtained. The splicing result is mapped by the third multi-layer perceptron to obtain the missing part of the weak perception target point cloud data Y with a size of 4096×3 missing , and then the missing part of the point cloud data Y of the weakly perceived target missing The candidate target point cloud data Y with a size of 2048×3 obtained in step S110 known After splicing, we finally get the complete point cloud data Y of the weakly perceived target with a size of 6144×3 completed .
[0147] S1206. For the above-mentioned complete point cloud data of the weakly perceived target, use the second point cloud feature encoding subnetwork to perform the above-mentioned feature extraction operation to obtain the above-mentioned complete shape feature sequence of the weakly perceived target.
[0148] For example, for the complete point cloud data Y of the weakly perceived target obtained by the example in step S1205 completed , the second point cloud feature encoding subnetwork PointNet++ is used to extract its shape features and obtain the complete shape feature sequence F of the weakly perceived target com For example, according to the example in step S1205, the complete point cloud data Y of the weakly perceived target with a size of 6144×3 is obtained. completed , the second point cloud feature encoding subnetwork PointNet++ is used to extract shape features to obtain a complete shape feature sequence F of a weakly perceived target with a size of 512×128 com .
[0149] The above step S130 may specifically include: step S1301 to step S1307;
[0150] S1301, randomly collecting the candidate target point cloud features to obtain an original sampling feature sequence, and using a corresponding neural network operation based on the original sampling feature sequence to obtain an original feature sequence;
[0151] S1302, performing a fourth splicing operation on the above-mentioned weakly perceived target complete shape feature sequence and the above-mentioned original feature sequence to obtain a spliced feature sequence;
[0152] S1303, performing a channel-by-channel pooling operation and a point-by-point pooling operation on the above-mentioned spliced feature sequence to obtain a point-by-point attention feature sequence and a channel-by-channel attention feature sequence respectively;
[0153] S1304, performing linear transformation on the point-by-point attention feature sequence and the channel-by-channel attention feature sequence based on the first linear layer and the second linear layer, respectively, and then multiplying them to obtain an attention feature product, and normalizing the attention feature product to obtain an overall attention weight map;
[0154] S1305, multiplying the overall attention weight map and the original feature sequence to redistribute the weights to obtain an original weighted feature sequence;
[0155] S1306, obtaining a weighted feature sequence based on the third feedforward network according to the original weighted feature sequence;
[0156] S1307. Use a downsampling operation on the weighted feature sequence to obtain geometric features, and use a corresponding neural network operation to encode the geometric features to obtain the weakly perceived target global features.
[0157] For example, the complete point cloud data Y of the weakly perceived target obtained in step S110 known Sampling N 5 points, and the candidate target point cloud feature F obtained in step S110 known Extract N 5 The features corresponding to the points are used as the original sampling feature sequence, and then the original sampling feature sequence is operated using the corresponding neural network to obtain a feature sequence of size N. 5 ×I 1 The original feature sequence F ori . Concatenate the original feature sequence F containing spatial position information ori And the complete shape feature sequence F of the weakly perceived target obtained in step S120 com , get the splicing feature sequence F c , size is N 5 ×I 2 , N 5 F ori 、F c The number of corresponding points, I 1 F ori The characteristic dimension, I 2 F c The feature dimension, N 5 , I 1 , I 2 All are positive integers;
[0158] The concatenation feature sequence F cPerform channel-by-channel pooling and point-by-point pooling operations respectively to obtain a size of N 5 ×B’s point-by-point attention feature sequence F p and size is B×I 2 The channel-by-channel attention feature sequence F g , F g and F p The first linear layer and the second linear layer are transformed in dimension and multiplied to obtain the attention feature product, and the attention feature product is standardized to obtain a size of N 5 ×I 1 The overall attention weight map F b , F b The calculation process of is shown in formula (7), where N 5 F p 、F b The dimension I 1 F b The dimension I 2 F g The feature dimension of B is F p 、F g The dimension, N 5 , B, I 1 , I 2 All are positive integers;
[0159] F b =sigmoid(linear 2 (F p )×linear 1 (F g )) (7)
[0160] Among them, sigmoid means using sigmoid function for standardization, Linear 1 and Linear 2 denote the first linear layer and the second linear layer respectively;
[0161] The overall attention weight map F b With the original feature sequence F ori Multiply by F ori The weights of each point are redistributed to obtain a size of N 5 ×I 1 The original weighted feature sequence F e , and use the third feed-forward network to learn F e The attention features in the 5 ×I 1 The weighted feature sequence F k , and then the weighted feature sequence F kUsing downsampling operations and corresponding neural network operations, we can obtain the weakly perceived target global feature F glo , where N 5 F e 、F k The number of midpoints, I 1 F e 、F k The feature dimension, N 5 , I 1 All are positive integers.
[0162] In some implementations, for the candidate target point cloud data Y of size 2048×3 obtained in step S110, known And the candidate target point cloud feature F with a size of 2048×128 known , first in the candidate target point cloud data Y known 512 points are randomly collected from the dataset and the corresponding candidate target point cloud features F are extracted. known The original sampling feature sequence is obtained, and the corresponding neural network operation is used on the original sampling feature sequence to obtain the original feature sequence F with a size of 512×128 ori At the same time, for the complete shape feature sequence F of the weakly perceived target with a size of 512×128 obtained by the example of step S120 com , splicing F com With the original feature sequence F ori Get the concatenated feature sequence F of size 512×256 c , then to F c Perform channel-by-channel pooling and point-by-point pooling operations respectively to obtain a point-by-point attention feature sequence F of size 512×1. p and a channel-wise attention feature sequence F of size 1×256 g , and F g and F p The overall attention weight map F with a size of 512×128 is calculated by formula (7): b , and then F b With F ori Multiply them together to get the original weighted feature sequence F of size 512×128 e , and F e Input to the third feedforward network to obtain a weighted feature sequence F of size 512×128 k , and then weighted feature sequence F k Using the downsampling operation and the corresponding neural network operation, the weakly perceived target global feature F with a size of 1×512 is obtained. glo .
[0163] In some embodiments, Figure 3As shown in FIG. 1 , the main framework of the weakly perceived target detection model in the weakly perceived target detection method proposed in this application is divided into four parts: an initial candidate data generation module, a weakly perceived target completion and reconstruction operation based on a structure-aware transformer, an attention fusion operation, and a weakly perceived target detection result generation module. The initial candidate data generation module is shown in FIG. Figure 2 As shown. In order to obtain the detection results using the weak-perception target detection method of the present application, the 3D standard data set KITTI is used to construct the training sample set and the sample set to be detected of the weak-perception target detection network. First, the 3712 samples provided by the training set in the KITTI data set are used as the training sample set to train the weak-perception target detection network and learn to generate the weak-perception target detection model. Secondly, the 3769 samples to be detected provided by the validation set in the KITTI data set are used to form the sample set to be detected. The sample set to be detected is detected and evaluated according to the weak-perception target detection model to generate the weak-perception target detection results. According to the sparseness and incompleteness of the targets in the scene and the degree of weak perception, the samples in the sample set to be detected are divided into three difficulty levels: "weak perception easy level", "weak perception medium level" and "weak perception difficult level". Among them, there are a large number of weak-perception targets with small size at a long distance, severe truncation and incomplete shape in the weak-perception difficult level sample scene.
[0164] The effectiveness of this method can be verified by comparing it with the benchmark method through experiments. Figure 4 The figure is a comparison diagram of the detection accuracy of this method and five benchmark methods on samples of three weak perception difficulty levels in the sample set to be detected. Figure 4 It can be seen that the method has achieved remarkable detection performance on difficult-level scene samples with a large number of weakly perceived targets, and the detection accuracy is higher than that of the other five benchmark methods, which proves that the method can accurately detect and locate weakly perceived targets; at the same time, for weakly perceived simple-level and weakly perceived medium-level sample scenes with relatively few weakly perceived targets, the method of the present invention still achieved the highest detection accuracy, which proves the effectiveness and practicality of the method.
[0165] Figure 5 The schematic diagram shows the detection effect of the weak-sensing target detection method proposed in the present invention on the KITTI to-be-detected sample set. Figure 5 For each set of detection effect diagrams in , the upper figure is the 2D image corresponding to the detection result, and the lower figure is the 3D point cloud representation of the detection result. Figure 5 The scene point cloud in the first group of figures is subjected to the weak-sensing target detection method proposed in the present invention, and the target detection boxes and "vehicle" category information corresponding to the 8 weak-sensing targets in the scene are detected. Figure 5 The scene point cloud data in the other three groups shown in the figure can accurately generate the target detection frame and category information corresponding to each weakly perceived target by executing the method of the present invention, and effectively complete the weakly perceived target detection. Figure 5 It can be seen that for weakly perceived targets with sparse point clouds at a long distance that are easily ignored, as well as some weakly perceived vehicle targets with incomplete shapes that only contain part of the point clouds, the method of the present invention can generate corresponding accurate target detection boxes and category information.
[0166] The above experimental results prove that the present invention has excellent detection performance for complex scenes with a large number of weakly perceived targets with incomplete shapes. This is because the weakly perceived target detection method proposed in the present invention uses an improved transformer encoder and decoder and an attention fusion mechanism to effectively query lower-level target structure information, and can accurately complete and generate complete geometric shape information of weakly perceived targets and enhance the overall spatial geometric features of weakly perceived targets, thereby effectively improving the detection accuracy of weakly perceived targets. Of course, the KITTI training sample set and the KITTI validation set used for evaluation are only examples. In practice, training and verification can also be performed through other databases or point cloud data prepared by the user himself.
[0167] In summary, the weak-perception target detection method provided by the present invention has a high theoretical value. For different types of weak-perception difficulty level scenes with a large number of weak-perception targets, this method can effectively detect weak-perception targets in positioning scenes and achieve excellent detection performance. The detection accuracy is significantly higher than other benchmark methods. Moreover, this method has been implemented through software and has great engineering application value.
[0168] See also Figure 6 The present invention also proposes a weakly perceived target detection device, comprising:
[0169] The first acquisition unit 21 is used to acquire a training sample set, input it into a weak-sensing target detection network, and perform preliminary detection through the first point cloud feature encoding subnetwork in the weak-sensing target detection network to acquire initial candidate data, wherein the initial candidate data includes an initial candidate box, candidate target point cloud data, and candidate target point cloud features;
[0170] An iterative operation unit 22 is used to perform iterative operations based on the candidate target point cloud data using the transformer encoder and the transformer decoder in the weakly perceived target detection network to obtain a shape feature sequence of a missing portion of the weakly perceived target, and to obtain a complete shape feature sequence of the weakly perceived target using a reconstruction operation in the weakly perceived target detection network according to the shape feature sequence of the missing portion of the weakly perceived target;
[0171] A fusion unit 23 is used to perform an attention fusion operation in the weakly perceived target detection network based on the complete shape feature sequence of the weakly perceived target and the point cloud features of the candidate target to obtain a global feature of the weakly perceived target;
[0172] A second acquisition unit 24 is used to perform confidence calculation and position regression operations in the weakly perceived target detection network based on the weakly perceived target global features to obtain a confidence score and a residual parameter of the weakly perceived target, and calculate a loss value based on the confidence score and the residual parameter to adjust the parameters of the weakly perceived target detection network to generate a weakly perceived target detection model;
[0173] The generation unit 25 is used to use the above-mentioned weakly perceived target detection model to detect the sample set to be detected, generate a weakly perceived target detection frame and weakly perceived target category information, and complete weakly perceived target detection.
[0174] like Figure 7 As shown, an embodiment of the present application also provides an electronic device 300, including a memory 310, a processor 320, and a computer program 311 stored in the memory 320 and executable on the processor, and when the processor 320 executes the computer program 311, the steps of any of the above-mentioned methods for weak-perception target detection are implemented.
[0175] Since the electronic device introduced in this embodiment is a device used to implement a weak-sensing target detection device in the embodiment of the present application, based on the method introduced in the embodiment of the present application, technical personnel in this field can understand the specific implementation method of the electronic device of this embodiment and its various variations. Therefore, how the electronic device implements the method in the embodiment of the present application is not introduced in detail here. As long as the equipment used by technical personnel in this field to implement the method in the embodiment of the present application is within the scope of protection of this application.
[0176] In the specific implementation process, when the computer program 311 is executed by the processor, it can achieve Figure 1 Any implementation manner in the corresponding embodiments.
[0177] It should be noted that in the above embodiments, the description of each embodiment has its own emphasis, and for parts that are not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0178] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems, or computer program products. Therefore, the present application may adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment in combination with software and hardware. Moreover, the present application may adopt the form of a computer program product implemented in one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that include computer-usable program code.
[0179] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded computer, or other programmable data processing device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0180] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.
[0181] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.
[0182] The present application also provides a computer program product, which includes computer software instructions. When the computer software instructions are executed on a processing device, the processing device is caused to execute the following Figure 1 The process of the weak perception target detection method in the corresponding embodiment.
[0183] The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on the computer, the process or function according to the embodiment of the present application is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from a website site, a computer, a server or a data center by wired (e.g., coaxial cable, optical fiber, digital subscriber line (digital subscriber line, DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) mode to another website site, computer, server or data center. The computer-readable storage medium can be any available medium that a computer can store or a data storage device such as a server or a data center that includes one or more available media integration. The available medium can be a magnetic medium, (e.g., a floppy disk, a hard disk, a tape), an optical medium (e.g., a DVD), or a semiconductor medium (e.g., a solid state drive (SSD)), etc.
[0184] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
[0185] In the several embodiments provided in the present application, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are only schematic. For example, the division of units is only a logical function division. There may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be an indirect coupling or communication connection through some interfaces, devices or units, which can be electrical, mechanical or other forms.
[0186] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0187] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of software functional units.
[0188] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the various embodiments of the present application. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (Read-Only Memory, ROM), random access memory (Random Access Memory, RAM), disk or optical disk and other media that can store program codes.
[0189] The above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present application.
Claims
1. A weakly perceived target detection method, characterized in that: include: Acquire a training sample set, input it into a weak-sensing target detection network, and perform preliminary detection through a first point cloud feature encoding subnetwork in the weak-sensing target detection network to acquire initial candidate data, wherein the initial candidate data includes an initial candidate box, candidate target point cloud data, and candidate target point cloud features; The iterative operation includes a first iterative operation and a second iterative operation, and the reconstruction operation includes a folding operation and a feature extraction operation. The first iterative operation is encoded based on a first self-attention weighted operation and a first feedforward network, The second iterative operation includes a second self-attention weighted operation, a cross-attention weighted operation, and a second feedforward network, The folding operation includes a second splicing operation, a third multilayer perceptron and a third splicing operation, Performing a sampling convolution operation and a first embedding operation in the weakly perceptual target detection network according to the candidate target point cloud data to obtain a local structural feature sequence of the embedding position; Based on the local structural feature sequence of the embedded position, using a transformer encoder to perform the first iterative operation and dimensional transformation operation to obtain a missing part center point sequence and a missing part center point local structural feature sequence, wherein the dimensional transformation operation includes a maximum pooling operation and a first multi-layer perceptron; Performing the second embedding operation in the weakly perceptual object detection network on the missing part center point sequence and the missing part center point local structure feature sequence to obtain a local shape feature sequence; According to the local shape feature sequence, the local structure feature sequence of the center point of the missing part and the center point sequence of the missing part, using a transformer decoder to perform the second iterative operation and feature transformation operation to obtain the shape feature sequence of the missing part of the weakly perceived target; For the shape feature sequence of the missing part of the weakly perceived target, combining the center point sequence of the missing part and the candidate target point cloud data, performing the folding operation to obtain the complete point cloud data of the weakly perceived target; For the complete point cloud data of the weakly perceived target, a second point cloud feature encoding subnetwork is used to perform the feature extraction operation to obtain a complete shape feature sequence of the weakly perceived target; Based on the complete shape feature sequence of the weakly perceived target and the point cloud features of the candidate target, an attention fusion operation is performed in the weakly perceived target detection network to obtain a global feature of the weakly perceived target; Based on the global features of the weakly perceived target, confidence calculation and position regression operations are performed in the weakly perceived target detection network to obtain a confidence score and a residual parameter of the weakly perceived target, and a loss value is calculated based on the confidence score and the residual parameter to adjust the parameters of the weakly perceived target detection network to generate a weakly perceived target detection model; The weakly perceived target detection model is used to detect the sample set to be detected, generate a weakly perceived target detection frame and weakly perceived target category information, and complete weakly perceived target detection.
2. The method according to claim 1, characterized in that The step of performing a sampling convolution operation and a first embedding operation in the weakly-perceptual target detection network according to the candidate target point cloud data to obtain a local structural feature sequence of an embedding position includes: Using the candidate target point cloud data, a center point sequence is obtained by an iterative farthest sampling method; Utilizing the center point sequence to extract relevant feature sequences based on a graph convolutional network; The first embedding operation is performed on the center point sequence and the related feature sequence to obtain the local structure feature sequence of the embedding position.
3. The method according to claim 1, characterized in that The step of performing the first iterative operation and dimension transformation operation based on the local structural feature sequence of the embedded position using the transformer encoder to obtain a missing part center point sequence and a missing part center point local structural feature sequence includes: Using the local structural feature sequence of the embedded position as the input of the transformer encoder to perform the first iterative operation to obtain an encoder output feature sequence; The maximum pooling operation and the first multi-layer perceptron are used to transform the feature dimension according to the encoder output feature sequence to obtain the missing part center point sequence and the missing part center point local structure feature sequence.
4. The method according to claim 1, characterized in that The step of performing the second iterative operation and feature transformation operation using the transformer decoder according to the local shape feature sequence, the local structure feature sequence of the missing part center point and the missing part center point sequence to obtain the shape feature sequence of the missing part of the weakly perceived target includes: The local shape feature sequence is used as the first layer input shape feature sequence of the transformer decoder ; The input shape feature sequence of the kth layer of the transformer decoder Perform the second self-attention weighted operation to obtain the query vector corresponding to the kth layer of the transformer decoder ; Based on the input shape feature sequence of the kth layer of the transformer decoder The local structural feature sequence of the missing part center point and the missing part center point sequence are calculated by the following formula to obtain the mixed feature of the kth layer of the transformer decoder : in, represents the first splicing operation, represents the first convolution operation, represents the second convolution operation, Y represents the center point sequence of the missing part, S represents the local structural feature sequence of the center point of the missing part, and L represents the number of layers of the transformer decoder; The mixed feature As the key vector and value vector corresponding to the kth layer of the transformer decoder, combined with the query vector corresponding to the kth layer of the transformer decoder , the cross-attention weighted operation is performed by the following formula to obtain the cross-attention weighted shape feature of the kth layer of the transformer decoder : in, is the number of attention heads, , and denote the first projection matrix, the second projection matrix and the third projection matrix of the hth attention head of the kth layer of the transformer decoder, respectively. is the output linear projection matrix of the criss-cross attention weighted operation of the k-th layer of the transformer decoder, is the scaling factor, represents the matrix transpose operation, Represents a normalization operation; Cross-attention weighted shape features based on the k-th layer of the transformer decoder Using the second feedforward network to obtain the output shape feature sequence of the kth layer of the transformer decoder ; Based on the output shape feature sequence of the kth layer of the transformer decoder Iterate (Lk) times through the remaining (Lk) layers of the transformer decoder to obtain the output shape feature sequence of the last layer of the transformer decoder , as the decoder output feature sequence; The feature transformation operation is performed through a second multilayer perceptron according to the decoder output feature sequence to obtain the shape feature sequence of the missing part of the weakly perceived target.
5. The method according to claim 1, characterized in that The step of performing the folding operation on the shape feature sequence of the missing part of the weakly perceived target and combining the center point sequence of the missing part with the candidate target point cloud data to obtain the complete point cloud data of the weakly perceived target includes: Performing the second splicing operation on the missing part center point sequence and the weakly perceived target missing part shape feature sequence, and using the third multi-layer perceptron mapping on the splicing result of the second splicing operation to obtain the weakly perceived target missing part point cloud data; The missing part of the point cloud data of the weakly perceived target and the candidate target point cloud data are subjected to the third splicing operation to obtain the complete point cloud data of the weakly perceived target.
6. The method according to claim 1, characterized in that The step of performing an attention fusion operation in the weakly perceived target detection network based on the complete shape feature sequence of the weakly perceived target and the point cloud features of the candidate target to obtain a global feature of the weakly perceived target includes: Randomly collect the candidate target point cloud features to obtain an original sampling feature sequence, and use a corresponding neural network operation based on the original sampling feature sequence to obtain an original feature sequence; Performing a fourth splicing operation on the complete shape feature sequence of the weakly perceived target and the original feature sequence to obtain a spliced feature sequence; Performing a channel-by-channel pooling operation and a point-by-point pooling operation on the concatenated feature sequence to obtain a point-by-point attention feature sequence and a channel-by-channel attention feature sequence respectively; The point-by-point attention feature sequence and the channel-by-channel attention feature sequence are linearly transformed based on the first linear layer and the second linear layer, and then multiplied to obtain an attention feature product, and the attention feature product is normalized to obtain an overall attention weight map; Multiplying the overall attention weight map and the original feature sequence to redistribute weights to obtain an original weighted feature sequence; Acquire a weighted feature sequence based on the original weighted feature sequence based on a third feedforward network; A downsampling operation is performed on the weighted feature sequence to obtain geometric features, and a corresponding neural network operation is used to encode the geometric features to obtain the weakly perceived target global features.
7. A weakly perceived target detection device, characterized in that: include: A first acquisition unit is used to acquire a training sample set, input it into a weak-sensing target detection network, and perform preliminary detection through a first point cloud feature encoding subnetwork in the weak-sensing target detection network to acquire initial candidate data, wherein the initial candidate data includes an initial candidate box, candidate target point cloud data, and candidate target point cloud features; Iteration unit, The iterative operation includes a first iterative operation and a second iterative operation, and the reconstruction operation includes a folding operation and a feature extraction operation. The first iterative operation is encoded based on a first self-attention weighted operation and a first feedforward network, The second iterative operation includes a second self-attention weighted operation, a cross-attention weighted operation, and a second feedforward network, The folding operation includes a second splicing operation, a third multilayer perceptron and a third splicing operation, Performing a sampling convolution operation and a first embedding operation in the weakly perceptual target detection network according to the candidate target point cloud data to obtain a local structural feature sequence of the embedding position; Based on the local structural feature sequence of the embedded position, using a transformer encoder to perform the first iterative operation and dimensional transformation operation to obtain a missing part center point sequence and a missing part center point local structural feature sequence, wherein the dimensional transformation operation includes a maximum pooling operation and a first multi-layer perceptron; Performing the second embedding operation in the weakly perceptual object detection network on the missing part center point sequence and the missing part center point local structure feature sequence to obtain a local shape feature sequence; According to the local shape feature sequence, the local structure feature sequence of the center point of the missing part and the center point sequence of the missing part, using a transformer decoder to perform the second iterative operation and feature transformation operation to obtain the shape feature sequence of the missing part of the weakly perceived target; For the shape feature sequence of the missing part of the weakly perceived target, combining the center point sequence of the missing part and the candidate target point cloud data, performing the folding operation to obtain the complete point cloud data of the weakly perceived target; For the complete point cloud data of the weakly perceived target, a second point cloud feature encoding subnetwork is used to perform the feature extraction operation to obtain a complete shape feature sequence of the weakly perceived target; A fusion unit, configured to perform an attention fusion operation in the weakly perceived target detection network based on the complete shape feature sequence of the weakly perceived target and the point cloud features of the candidate target to obtain a global feature of the weakly perceived target; a second acquisition unit, configured to perform confidence calculation and position regression operations in the weakly perceived target detection network based on the weakly perceived target global features to obtain a confidence score and a residual parameter of the weakly perceived target, and calculate a loss value based on the confidence score and the residual parameter to adjust parameters of the weakly perceived target detection network, and generate a weakly perceived target detection model; The generation unit is used to use the weakly perceived target detection model to detect the sample set to be detected, generate a weakly perceived target detection frame and weakly perceived target category information, and complete weakly perceived target detection.
8. An electronic device comprising: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor is used to implement the steps of the weak-perception target detection method as described in any one of claims 1 to 6 when executing the computer program stored in the memory.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the weak-perception target detection method according to any one of claims 1 to 6 is implemented.