Training method and device of 3D target perception model, 3D target perception method and device, equipment and computer program product
By using a training method for a 3D target perception model, feature extraction and comparative learning are performed using the original network branch and the contrastive network branch. This solves the problem of multimodal feature misalignment and improves the accuracy and robustness of target localization.
Patent Information
- Application Number
- CN202511091595.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-05
- Publication Date
- 2025-11-21
AI Technical Summary
Existing roadside multimodal feature fusion methods require high calibration accuracy between sensors and have failed to effectively address the problem of decreased perception algorithm accuracy caused by multimodal feature misalignment.
The training method of 3D target perception model is adopted. Through the original network branch and the contrastive network branch, feature extraction and instance-level contrastive learning of multimodal data are performed respectively. The model parameters are optimized by BEV space feature fusion and contrastive loss function to enhance the multimodal feature alignment effect.
It improves the target localization accuracy and robustness of the 3D target perception model, solves the multimodal feature misalignment problem, balances multimodal feature fusion, and enhances the robustness of the perception algorithm.
Smart Images

Figure CN120995270A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of 3D target perception technology, and in particular to a training method and apparatus for a 3D target perception model, a 3D target perception method and apparatus, equipment, and computer program product. Background Technology
[0002] With the development of IoT (Internet of Things), communication, and AI technologies, roadside perception systems are able to provide real-time, wide-field-of-view road condition perception and are playing an increasingly important role in intelligent transportation systems. To improve the performance of perception systems, roadside systems are typically equipped with multiple sensors to collect multimodal perception data.
[0003] Existing roadside multimodal feature fusion methods include pre-fusion, mid-fusion, and post-fusion paradigms. With the development of deep learning technology, mainstream research methods have begun to focus on the mid-fusion paradigm. However, current multimodal feature mid-fusion paradigms mostly employ feature dimension concatenation and position-by-position addition fusion methods. These methods require high calibration accuracy between sensors and rarely consider the problem of decreased perception algorithm accuracy caused by misalignment of multimodal features (such as misalignment between laser feature data and camera feature data). Summary of the Invention
[0004] This application provides a training method and apparatus for a 3D target perception model, a 3D target perception method and apparatus, a device, and a computer program product to improve the perception accuracy of the 3D target perception model.
[0005] The embodiments of this application adopt the following technical solutions:
[0006] In a first aspect, embodiments of this application provide a method for training a 3D target perception model, the method comprising:
[0007] Acquire multimodal sensor data and corresponding 3D target ground truth data;
[0008] The multimodal sensor data and the corresponding 3D target ground truth data are input into the original network branch of the 3D target perception model to obtain the original loss of the original network branch.
[0009] The multimodal sensor data and the corresponding 3D target ground truth data are input into the contrast network branch of the 3D target perception model to obtain the contrast loss of the contrast network branch. The contrast network branch performs contrast learning based on the detection box of the 3D target.
[0010] The parameters of the 3D target perception model are updated based on the original loss of the original network branch and the contrastive loss of the contrastive network branch to obtain the trained 3D target perception model.
[0011] Optionally, the original network branch includes a single-modal encoder and a BEV encoder / decoder, with each modality's sensor data corresponding to a single-modal encoder. The step of inputting the multimodal sensor data and the corresponding 3D target ground truth data into the original network branch of the 3D target perception model to obtain the original loss of the original network branch includes:
[0012] The sensor data for each mode are input into each of the single-mode encoders to obtain the coding features for each mode;
[0013] The encoded features of each modality are transformed into the BEV space to obtain the BEV space features of each modality.
[0014] The BEV spatial features of each mode are fused to obtain the fused BEV spatial features;
[0015] The fused BEV spatial features are input into the BEV codec to obtain the 3D target perception results of the original network branch;
[0016] The original loss of the original network branch is calculated based on the 3D target perception results of the original network branch and the corresponding 3D target ground truth data.
[0017] Optionally, the 3D target ground truth data includes the ground truth detection box data of the 3D target, and the step of inputting the multimodal sensor data and the corresponding 3D target ground truth data into the contrast network branch of the 3D target perception model to obtain the contrast loss of the contrast network branch includes:
[0018] Based on the BEV space features of each modality and the ground truth detection box data of the 3D target, determine the instance features of each modality corresponding to the ground truth detection box of each 3D target;
[0019] Based on the instance features of each modality corresponding to the ground truth detection box of each 3D target, the contrast loss of the contrast network branch is calculated using a preset loss function.
[0020] Optionally, determining the instance features of each modality corresponding to the ground truth detection box of each 3D target based on the BEV space features of each modality and the ground truth detection box data of the 3D target includes:
[0021] Based on the ground truth detection box data of the 3D target, calculate the spatial region of the ground truth detection box of each 3D target on the BEV spatial features of each modality;
[0022] Based on the spatial region of the ground truth detection box of each 3D target on the BEV spatial features of each modality, the instance features at the corresponding positions are cropped out.
[0023] Optionally, the step of calculating the contrast loss of the contrast network branch using a preset loss function based on the instance features of each modality corresponding to the ground truth detection box of each 3D target includes:
[0024] Positive and negative samples are constructed based on the instance features of each modality corresponding to the ground truth detection box of each 3D target.
[0025] Based on the positive and negative samples, calculate the pairwise similarity between the instance features of each modality corresponding to the ground truth detection box of each 3D target;
[0026] Based on the pairwise similarity between instance features of each modality corresponding to the ground truth detection box of each 3D target, the contrast loss of the contrast network branch is calculated using a preset loss function.
[0027] Optionally, updating the parameters of the 3D object perception model based on the original loss of the original network branch and the contrastive loss of the contrastive network branch to obtain the trained 3D object perception model includes:
[0028] The network parameters of the single-modal encoder and BEV codec in the original network branch are updated using the original loss of the original network branch.
[0029] The network parameters of the single-modal encoder in the original network branch are updated using the contrast loss of the contrast network branch.
[0030] Secondly, embodiments of this application also provide a 3D target perception method, the 3D target perception method comprising:
[0031] Acquire multimodal sensor data;
[0032] Based on multimodal sensor data, a 3D target perception model is used to perform target perception and obtain 3D target perception results.
[0033] The 3D target perception model is trained based on any of the aforementioned training methods for the 3D target perception model.
[0034] Thirdly, embodiments of this application also provide a training apparatus for a 3D target perception model, the training apparatus for the 3D target perception model comprising:
[0035] The first acquisition unit is used to acquire multimodal sensor data and corresponding 3D target ground truth data;
[0036] The first input unit is used to input the multimodal sensor data and the corresponding 3D target ground truth data into the original network branch of the 3D target perception model to obtain the original loss of the original network branch;
[0037] The second input unit is used to input the multimodal sensor data and the corresponding 3D target ground truth data into the contrast network branch of the 3D target perception model to obtain the contrast loss of the contrast network branch, and the contrast network branch performs contrast learning based on the detection box of the 3D target.
[0038] The update unit is used to update the parameters of the 3D target perception model based on the original loss of the original network branch and the contrastive loss of the contrastive network branch, so as to obtain the trained 3D target perception model.
[0039] Fourthly, embodiments of this application also provide a 3D target sensing device, the 3D target sensing device comprising:
[0040] The second acquisition unit is used to acquire multimodal sensor data;
[0041] The target perception unit is used to perform target perception based on multimodal sensor data and a 3D target perception model to obtain 3D target perception results.
[0042] The 3D target perception model is trained based on the aforementioned training device for the 3D target perception model.
[0043] Fifthly, embodiments of this application also provide an apparatus, comprising:
[0044] A processor; and a memory arranged to store computer-executable instructions, which, when executed, cause the processor to perform a training method for any of the aforementioned 3D target perception models, and to perform the aforementioned 3D target perception method.
[0045] Sixthly, embodiments of this application also provide a computer program product, including a computer program / instruction, wherein when the computer program / instruction is executed by a processor, it implements a training method for any of the aforementioned 3D target perception models, and implements the aforementioned 3D target perception method.
[0046] The above-mentioned technical solutions adopted in the embodiments of this application can achieve the following beneficial effects: The training method of the 3D target perception model in the embodiments of this application first acquires multimodal sensor data and corresponding 3D target ground truth data; then, the multimodal sensor data and corresponding 3D target ground truth data are input into the original network branch of the 3D target perception model to obtain the original loss of the original network branch; then, the multimodal sensor data and corresponding 3D target ground truth data are input into the contrast network branch of the 3D target perception model to obtain the contrast loss of the contrast network branch, and the contrast network branch performs contrast learning based on the detection boxes of the 3D target; finally, the parameters of the 3D target perception model are updated according to the original loss of the original network branch and the contrast loss of the contrast network branch to obtain the trained 3D target perception model. The training method of the 3D target perception model in this application introduces a contrastive learning network branch based on instance-level features into the original training framework of the 3D target perception model. This enhances the alignment effect between multimodal features, solves the misalignment problem that occurs in the fusion stage of multimodal features, and improves the target localization accuracy of the perception algorithm. At the same time, it can balance the fusion of multimodal features, avoid over-reliance on a certain type of sensor, and improve the robustness of the multimodal perception algorithm. Attached Figure Description
[0047] The accompanying drawings, which are included to provide a further understanding of this application and form part of this application, illustrate exemplary embodiments and are used to explain this application, but do not constitute an undue limitation of this application. In the drawings:
[0048] Figure 1 This is a flowchart illustrating a training method for a 3D target perception model according to an embodiment of this application.
[0049] Figure 2 This is a schematic diagram of the training process of a 3D target perception model in an embodiment of this application;
[0050] Figure 3 This is a flowchart illustrating a 3D target perception method in an embodiment of this application;
[0051] Figure 4 This is a schematic diagram of the structure of a training device for a 3D target perception model according to an embodiment of this application;
[0052] Figure 5 This is a schematic diagram of the structure of a 3D target sensing device according to an embodiment of this application;
[0053] Figure 6 This is a schematic diagram of the structure of a device according to an embodiment of this application. Detailed Implementation
[0054] To make the objectives, technical solutions, and advantages of this application clearer, the technical solutions of this application will be clearly and completely described below in conjunction with specific embodiments and corresponding drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of them. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0055] The technical solutions provided by the various embodiments of this application are described in detail below with reference to the accompanying drawings.
[0056] This application provides a method for training a 3D object perception model, such as... Figure 1 The diagram illustrates a flowchart of a training method for a 3D target perception model according to an embodiment of this application. The training method for the 3D target perception model includes at least the following steps S110 to S140:
[0057] Step S110: Obtain multimodal sensor data and corresponding 3D target ground truth data.
[0058] Combination Figure 2 This application provides a schematic diagram of the training process for a 3D target perception model according to an embodiment of the present application. In practical applications, data is collected using various types of sensors. For example, in autonomous driving scenarios, sensors such as LiDAR, cameras, and millimeter-wave radar are mainly used. LiDAR can acquire 3D point cloud data of the surrounding environment, cameras can capture color image information of the environment, and millimeter-wave radar can provide data such as the distance and speed of the target. The data collected by these different modal sensors are collected, and at the same time, accurate 3D target ground truth data corresponding to these sensor data, such as the position, size, and category of the target object, are obtained through manual annotation or with the help of high-precision positioning equipment, providing supervision signals for subsequent model training.
[0059] Step S120: Input the multimodal sensor data and the corresponding 3D target ground truth data into the original network branch of the 3D target perception model to obtain the original loss of the original network branch.
[0060] The multimodal sensor data obtained in the preceding steps, along with the corresponding ground truth 3D target data, are input into the pre-designed original network branch of the 3D target perception model. The 3D target perception model can be, for example, a task model for 3D target detection or 3D target segmentation. The original network branch performs feature extraction, fusion, and target perception operations on the input multimodal data. During model processing, a loss value is calculated based on the output of the network branch and the input ground truth 3D target data. This loss value reflects the degree of difference between the prediction result of the original network branch and the true target, and serves as the loss of the original network branch. For example, if there is a deviation between the model's predicted target location and the true location, this deviation can be quantified using a specific loss function (such as the mean squared error loss function) to obtain the original loss value.
[0061] Step S130: Input the multimodal sensor data and the corresponding 3D target ground truth data into the contrast network branch of the 3D target perception model to obtain the contrast loss of the contrast network branch. The contrast network branch performs contrast learning based on the detection box of the 3D target.
[0062] Similarly, the multimodal sensor data and corresponding ground truth 3D target data obtained in the preceding steps are input into the contrastive network branch of the 3D target perception model. The contrastive network branch primarily performs contrastive learning based on the instance features of the 3D target detection boxes. It extracts instance-level features related to the 3D target detection boxes from the multimodal data, and then, through a contrastive learning strategy, makes the instance features of the same target more similar across different modalities, and the instance features of different targets more distinct. During this process, the contrastive loss value is calculated based on the results of the contrastive learning and its relationship with the ground truth 3D target data.
[0063] Step S140: Update the parameters of the 3D target perception model according to the original loss of the original network branch and the contrastive loss of the contrastive network branch to obtain the trained 3D target perception model.
[0064] After obtaining the original loss of the original network branch and the comparative loss of the comparative network branch, these two loss values are combined according to a certain weight ratio (the weights can be adjusted according to actual needs and model performance) to obtain the total loss value. Then, an optimization algorithm (such as stochastic gradient descent and its variants) is used to update the parameters of the 3D target perception model based on the total loss value. Through continuous iterative training, updating the model parameters according to the new loss value in each iteration, the total loss of the model gradually decreases until a preset stopping condition is met (such as reaching a specified number of training epochs, loss convergence, etc.). The 3D target perception model corresponding to the model parameters obtained at this point is the trained model, which has better performance in multimodal feature fusion and target localization.
[0065] The training method of the 3D target perception model in this application introduces a contrastive learning network branch based on instance-level features into the original training framework of the 3D target perception model. This enhances the alignment effect between multimodal features, solves the misalignment problem that occurs in the fusion stage of multimodal features, and improves the target localization accuracy of the perception algorithm. At the same time, it can balance the fusion of multimodal features, avoid over-reliance on a certain type of sensor, and improve the robustness of the multimodal perception algorithm.
[0066] In some embodiments of this application, the original network branch includes a single-modal encoder and a BEV codec. Sensor data for each modality corresponds to one of the single-modal encoders. The step of inputting the multimodal sensor data and the corresponding 3D target ground truth data into the original network branch of the 3D target perception model to obtain the original loss of the original network branch includes: inputting sensor data for each modality into each of the single-modal encoders to obtain the encoding features of each modality; converting the encoding features of each modality to the BEV space to obtain the BEV space features of each modality; fusing the BEV space features of each modality to obtain fused BEV space features; inputting the fused BEV space features into the BEV codec to obtain the 3D target perception result of the original network branch; and calculating the original loss of the original network branch based on the 3D target perception result of the original network branch and the corresponding 3D target ground truth data.
[0067] Multimodal sensor data comes from different sensor modes, such as LiDAR, cameras, and millimeter-wave radar. For each mode of sensor data, there is a corresponding single-mode encoder. The sensor data for each mode is input into its respective single-mode encoder. The function of the single-mode encoder is to extract and encode features from the input data of that specific mode, transforming the raw sensor data into higher-level, more representative coded features.
[0068] After obtaining the encoded features of each modality, these features are transformed into BEV (Bird's Eye View) space according to calibration parameters. BEV space is a top-down view of a scene, which is very useful for 3D target perception tasks because it visually shows the position and distribution of targets in space. Each modality's encoded features have their own corresponding transformation method. These are transformed into BEV space to obtain the BEV space features for each modality. For example, for LiDAR encoded features, coordinate transformation and projection operations can be used to transform its point cloud data into BEV space, generating a two-dimensional BEV feature map, where each pixel represents object information within a certain spatial region. For camera encoded features, perspective transformation and other methods can be used to transform them into BEV space.
[0069] After obtaining the BEV spatial features of each modality, these features need to be fused. The purpose of feature fusion is to integrate information from different modalities, fully utilize the advantages of each modality, and improve the model's ability to perceive 3D targets. The fusion process can employ various methods, such as simple concatenation, which concatenates the BEV spatial features of different modalities along the channel dimension to form a fused feature containing multimodal information; or more complex fusion strategies, such as attention mechanisms, which dynamically adjust the weights of different modal features based on their importance before fusion. Through fusion, a fused BEV spatial feature is obtained, which integrates information from multiple modalities and can more comprehensively describe 3D targets in the scene.
[0070] The fused BEV spatial features are input into the BEV codec. The BEV codec further processes and parses the fused features to obtain 3D target perception results. The codec may contain a multi-layered network structure, which extracts information such as the target's category, location, and size from the BEV spatial features through operations such as convolution and deconvolution, and finally outputs 3D target perception results, such as the location coordinates of the detection box and the category label.
[0071] After obtaining the 3D target perception results of the original network branch, the original loss needs to be calculated based on these results and the corresponding ground truth 3D target data. The ground truth 3D target data is accurate target information obtained in advance through manual annotation or other high-precision methods. Loss calculation can employ specific loss functions. For example, for target detection tasks, classification loss (such as cross-entropy loss) can be used to measure the difference between the predicted target category and the true category, while regression loss (such as SmoothL1 loss) can be used to measure the deviation of the predicted target position and size from the true values. Combining these loss values yields the original loss of the original network branch, which reflects the gap between the model's predictive performance on the current input data and the actual situation.
[0072] Each modality of sensor data corresponds to a single-modal encoder, enabling independent feature extraction from different modalities. This comprehensively captures various information related to 3D targets in the scene, avoiding the problem of insufficient information from a single modality. The encoded features of each modality are transformed into the BEV space, providing a unified perspective representation for multimodal data. Fusion of the BEV space features from different modalities fully utilizes the complementary information between them. Different modalities possess advantages in different aspects; fusion combines these advantages to obtain a more comprehensive and accurate feature representation.
[0073] In some embodiments of this application, the 3D target ground truth data includes 3D target ground truth detection box data. The step of inputting the multimodal sensor data and the corresponding 3D target ground truth data into the contrast network branch of the 3D target perception model to obtain the contrast loss of the contrast network branch includes: determining the instance features of each modality corresponding to the ground truth detection box of each 3D target based on the BEV spatial features of each modality and the 3D target ground truth detection box data; and calculating the contrast loss of the contrast network branch using a preset loss function based on the instance features of each modality corresponding to the ground truth detection box of each 3D target.
[0074] After acquiring the BEV spatial features of each modality and the ground truth bounding box data of the 3D target, for each 3D target, based on the position and range of its ground truth bounding box in the BEV space, corresponding instance features are extracted from the BEV spatial features of each modality. For example, for the BEV spatial features of the LiDAR modality, features corresponding to the LiDAR data within the spatial region defined by the ground truth bounding box are extracted. These features reflect information such as the shape and distance of the 3D target from the LiDAR's perspective. Similarly, for the BEV spatial features of the camera modality, image features related to the target, such as the target's color and texture, are extracted according to the range of the ground truth bounding box. In this way, instance features of each modality corresponding to the ground truth bounding box of each 3D target can be obtained.
[0075] After obtaining the instance features of each modality for each 3D target, the contrastive loss of the contrastive network branch is calculated using a pre-defined loss function. The pre-defined loss function is designed to measure the similarity and difference between instance features of different modalities, making the instance features of the same target more similar in different modalities, while making the instance features of different targets more different.
[0076] By determining the instance features of each modality based on the ground truth detection box data of the 3D target and calculating the contrast loss using a preset loss function, the model can be forced to learn similar feature representations of the same target under different modalities. This helps to enhance the alignment effect between multimodal features, enabling data from different modalities to better correspond and fuse at the feature level, thus avoiding the target localization deviation problem caused by multimodal feature misalignment.
[0077] The contrastive learning process enables the model to gain a deeper understanding of the relationship between different modalities of data and 3D targets, thereby extracting more discriminative features. These features can more accurately describe the target's position, shape, and other information, thus improving the target localization accuracy of the 3D target perception model.
[0078] In some embodiments of this application, determining the instance features of each modality corresponding to the ground truth detection box of each 3D target based on the BEV spatial features of each modality and the ground truth detection box data of the 3D target includes: calculating the spatial region of the ground truth detection box of each 3D target on the BEV spatial features of each modality based on the ground truth detection box data of the 3D target; and cropping out the instance features at the corresponding positions based on the spatial region of the ground truth detection box of each 3D target on the BEV spatial features of each modality.
[0079] After acquiring the BEV spatial features of each modality and the ground truth detection box data of the 3D target, the spatial region {S_11,S_12,…,S_1N,…,S_M1,…,S_MN} of a single Box_i on the BEV feature of each modality is calculated based on the center point position and size of the ground truth detection box {Box_1,…,Box_M} of each 3D target. Here, N represents the number of sensor modalities, and M represents the number of detection boxes. Since the coordinate system and spatial resolution of sensor data from different modalities may differ during the acquisition and transformation to BEV space, it is necessary to map the coordinates of the ground truth detection boxes onto the BEV spatial features of each modality according to the transformation relationship and parameters of each modality.
[0080] For example, for the LiDAR modality, after the point cloud data is converted to BEV space, it is necessary to determine the pixel coordinate range of the ground truth bounding box on the LiDAR BEV feature map. Similarly, for the camera modality, the region on the camera BEV feature map corresponding to the ground truth bounding box needs to be calculated based on the image-to-BEV space conversion rules. In this way, the spatial region covered by the ground truth bounding box of each 3D target on the BEV space features of each modality can be accurately determined, providing an accurate range for subsequent feature clipping.
[0081] After determining the spatial region of the ground truth detection box for each 3D target on the BEV spatial features of each modality, the BEV spatial features are clipped according to the calculated spatial region S_ij to extract the instance features Inst_ij at the corresponding position, where i represents the index of the sensor modality and j represents the index of the detection box.
[0082] For example, for the BEV spatial features of a LiDAR radar, assuming the region corresponding to the ground truth detection box is a rectangle, the feature data within this region is cropped from the LiDAR BEV feature map according to the boundary of this rectangle. This data represents the instance features of the 3D target from the LiDAR's perspective. Similarly, for other modalities such as cameras, cropping is performed according to the corresponding calculated spatial regions to obtain instance features for each modality. Through the cropping operation, the relevant features of each 3D target in each modality can be extracted for subsequent comparative learning and other processing.
[0083] By calculating and cropping the spatial regions of the ground truth detection boxes in the BEV space features of each modality, the instance features of each 3D target in each modality can be accurately extracted. This allows subsequent comparative learning and target perception to be based on more accurate and relevant features, thereby improving the model's accuracy in recognizing and locating 3D targets.
[0084] In contrastive learning, it is necessary to compare instance features of the same target under different modalities. Accurately cropped instance features can provide more accurate and representative data for contrastive learning, enabling the model to better learn the correspondence and similarity between features of different modalities.
[0085] Because the embodiments of this application can accurately extract the instance features of each 3D target in various modalities, the model can access richer and more accurate target feature information during the learning process, enabling the model to better understand the target features under different scenarios and different sensor configurations, and thus adapt and generalize more quickly and accurately when faced with new and unseen data.
[0086] In some embodiments of this application, the step of calculating the contrast loss of the contrast network branch using a preset loss function based on the instance features of each modality corresponding to the ground truth detection box of each 3D target includes: constructing positive and negative samples based on the instance features of each modality corresponding to the ground truth detection box of each 3D target; calculating the pairwise similarity between the instance features of each modality corresponding to the ground truth detection box of each 3D target based on the positive and negative samples; and calculating the contrast loss of the contrast network branch using a preset loss function based on the pairwise similarity between the instance features of each modality corresponding to the ground truth detection box of each 3D target.
[0087] After obtaining the instance features of each modality corresponding to the ground truth bounding box of each 3D target, the positive and negative samples are constructed. Assume there exists an instance feature set {Inst_i1,…,Inst_iM} for modality i and an instance feature set {Inst_s1,…,Inst_sM} for modality s.
[0088] For constructing positive samples, instance features of the same 3D target in different modalities are paired, i.e., {(Inst_ij,Inst_sj)} constitute positive samples. This means that for each target, the features extracted in modality i and modality s are considered similar features describing the same target and should have high similarity.
[0089] For negative sample construction, instance features of different targets in different modalities are paired, i.e., {(Inst_ip,Inst_sk)|p!=k} constitutes the negative samples. These negative sample pairs represent features of different targets and should have low similarity. In this way, a set of positive and negative samples for contrastive learning is constructed, providing a foundation for subsequent calculation of contrastive loss.
[0090] To facilitate the calculation of similarity between instance features, it is necessary to map two instance features to a space with the same spatial and feature dimensions. This can be achieved through adaptive pooling and multilayer perceptrons (MLPs). Adaptive pooling can perform pooling operations on features according to a preset output size, adjusting the spatial dimension of the features. MLPs, on the other hand, can map input features to a specified dimensional space through a series of linear transformations and nonlinear activation functions.
[0091] After mapping features to the same dimension, the mapped features are flattened, converting multidimensional features into one-dimensional vectors. Then, the flattened features are normalized to obtain a normalized set of instance features {q_ij | i = 1, ..., N; j = 1, ..., M}. Normalization ensures that the feature vectors have a uniform scale, avoiding the impact of different feature value ranges on similarity calculations.
[0092] Based on normalized instance features, the similarity between two instance features can be calculated, for example, using L2 distance or cosine similarity. L2 distance measures the Euclidean distance between two vectors in space; the smaller the distance, the more similar the two vectors are. Cosine similarity measures the similarity by calculating the cosine of the angle between two vectors; the closer the cosine value is to 1, the more similar the two vectors are.
[0093] Based on the pairwise similarity between instance features of each modality corresponding to the ground truth detection box of each 3D target, the contrastive loss of the contrastive network branch is calculated using a preset loss function. The preset loss function is usually designed based on the similarity distribution of positive and negative samples, with the aim of maximizing the similarity between positive samples and minimizing the similarity between negative samples, thereby guiding the model to learn more discriminative feature representations.
[0094] Taking the cosine similarity calculation method as an example, the preset loss function in this application embodiment can be expressed as follows:
[0095]
[0096] Where Lis represents the contrast loss between the i-th modality and the s-th modality. L represents the total contrast loss, which is the sum of the Lis losses of all modalities. M represents the number of detection boxes. τ is a hyperparameter used to adjust the sensitivity of the contrast loss, controlling the smoothness of the feature distribution. The smaller τ is, the higher the weight of highly similar samples, enhancing the model's ability to distinguish between positive and negative samples; the larger τ is, the less difference between positive and negative samples.
[0097] qij represents the feature representation of the j-th detection box in the i-th modality. qsj represents the feature representation of the j-th detection box in the s-th modality. exp(q ij ·q sj The formula ( / τ) calculates the dot product of the feature vectors of the j-th detection box in the i-th and s-th modalities, divides by τ, and then takes the exponent. This measures the similarity between the feature vectors of the same object detection box in two different modalities based on cosine distance. This means that for any target detection box j in the i-th modality, the similarity between its feature vectors and the feature vectors of all target detection boxes k in the s-th modality is calculated and summed. This means that for any target detection box j in the s-th modality, the similarity between its feature vectors and the feature vectors of all target detection boxes k in the ith modality is calculated and summed.
[0098] The loss function described above is equivalent to calculating the similarity between the instance features of any two detection boxes in any two modalities. By comparing the similarity of positive sample pairs (features of the same target in different modalities) and negative sample pairs (features of different targets in different modalities), the model is encouraged to learn more discriminative feature representations. The similarity of positive sample pairs should be as high as possible, while the similarity of negative sample pairs should be as low as possible, thereby optimizing the model's perceptual performance on multimodal data.
[0099] By constructing positive and negative samples and calculating the comparative loss using a preset loss function, and updating the model parameters based on the loss values, the distribution of features in the feature space can be optimized. The model will gradually adjust the feature extraction method, causing positive samples to cluster in similar regions of the feature space and negative samples to be scattered in different regions of the feature space, thereby improving the model's generalization ability.
[0100] Furthermore, in practical applications, sensor data may be affected by noise, interference, and other factors, leading to inaccurate feature extraction. By constructing positive and negative samples through contrastive learning and calculating contrastive loss, the model can learn more robust feature representations and has a certain tolerance for noise and interference in sensor data.
[0101] In some embodiments of this application, updating the parameters of the 3D target perception model based on the original loss of the original network branch and the contrastive loss of the contrastive network branch to obtain a trained 3D target perception model includes: updating the network parameters of the single-modal encoder and BEV codec in the original network branch using the original loss of the original network branch; and updating the network parameters of the single-modal encoder in the original network branch using the contrastive loss of the contrastive network branch.
[0102] Continue to refer to Figure 2 The original network branch contains a single-modal encoder and a BEV codec. Using the original loss, the gradient is calculated via backpropagation, which then updates the network parameters of the single-modal encoder and BEV codec in the original network branch. The backpropagation algorithm adjusts the parameter values of the single-modal encoder and BEV codec along the direction of gradient descent of the loss function, based on the sensitivity of the original loss to the network parameters.
[0103] Simultaneously, the contrastive loss of the contrastive network branch is used to calculate the gradient via backpropagation, updating the network parameters of the unimodal encoder in the original network branch. The contrastive loss reflects the requirements for multimodal feature alignment and discrimination in the contrastive learning task. Through backpropagation, the gradient information of the contrastive loss is passed to the unimodal encoder, adjusting its parameters so that the unimodal encoder can extract feature representations that are more conducive to multimodal feature comparison and fusion.
[0104] Updating the network parameters of the unimodal encoder and BEV codec using the original loss directly optimizes the model's performance on basic 3D target perception tasks. The unimodal encoder extracts initial features from sensor data of different modalities, while the BEV codec further processes and parses these features to achieve tasks such as 3D target detection and localization. By updating parameters based on the original loss, the model can better adapt to the requirements of the original task, improving the accuracy and reliability of target perception.
[0105] Updating the network parameters of a unimodal encoder using contrastive loss helps improve the alignment between multimodal features. Contrastive learning requires that features of the same target in different modalities have high similarity, while features of different targets have low similarity. By adjusting the parameters of the unimodal encoder, the features of the same target in different modalities become closer, thereby achieving better alignment at the feature level and providing a more accurate foundation for subsequent multimodal feature fusion.
[0106] Combining the two parameter update methods described above comprehensively considers both the requirements of the original task and the requirements of multimodal feature alignment and discrimination, thereby improving the overall performance of the 3D target perception model. The model not only performs well in basic 3D target perception tasks but also better handles multimodal data, improving target localization accuracy and robustness, and adapting to more complex scenarios and task requirements.
[0107] This application also provides a 3D target perception method, such as... Figure 3 The diagram provided illustrates a flowchart of a 3D target perception method according to an embodiment of this application. The 3D target perception method includes at least the following steps S310 to S320:
[0108] Step S310: Acquire multimodal sensor data;
[0109] Step S320: Based on the multimodal sensor data, use the 3D target perception model to perform target perception and obtain the 3D target perception result;
[0110] The 3D target perception model is trained based on any of the aforementioned training methods for the 3D target perception model.
[0111] When applying the 3D target perception model trained in the aforementioned embodiments, multimodal sensor data is first collected using various types of sensors, such as laser point cloud data and camera image data. The acquired multimodal sensor data is then input into the trained 3D target perception model. This 3D target perception model is trained based on the aforementioned 3D target perception model training method. During training, the model has learned how to effectively process multimodal data, perform feature extraction and fusion, and possesses the ability to accurately perceive 3D targets. The model analyzes and processes the input multimodal sensor data, ultimately outputting 3D target perception results, such as the target object's position, size, category, and motion state.
[0112] Because the 3D object perception model incorporates a contrastive learning network branch based on instance-level features during training, it enhances the alignment between multimodal features and resolves the misalignment problem that occurs during the fusion stage. Therefore, when using this model for 3D object perception, it can more accurately identify and locate target objects, improving target localization accuracy and reducing false positives and false negatives.
[0113] This application embodiment also provides a training device 400 for a 3D target perception model, such as... Figure 4The diagram shows a structural schematic of a training device for a 3D target perception model according to an embodiment of this application. The training device 400 for the 3D target perception model includes: a first acquisition unit 410, a first input unit 420, a second input unit 430, and an update unit 440, wherein:
[0114] The first acquisition unit 410 is used to acquire multimodal sensor data and corresponding 3D target ground truth data;
[0115] The first input unit 420 is used to input the multimodal sensor data and the corresponding 3D target ground truth data into the original network branch of the 3D target perception model to obtain the original loss of the original network branch;
[0116] The second input unit 430 is used to input the multimodal sensor data and the corresponding 3D target ground truth data into the contrast network branch of the 3D target perception model to obtain the contrast loss of the contrast network branch, and the contrast network branch performs contrast learning based on the detection box of the 3D target.
[0117] The update unit 440 is used to update the parameters of the 3D target perception model according to the original loss of the original network branch and the contrast loss of the contrast network branch, so as to obtain the trained 3D target perception model.
[0118] In some embodiments of this application, the original network branch includes a single-modal encoder and a BEV codec, with sensor data for each modality corresponding to one of the single-modal encoders. The first input unit 420 is specifically used for: inputting sensor data for each modality into each of the single-modal encoders to obtain the encoding features of each modality; converting the encoding features of each modality into the BEV space to obtain the BEV space features of each modality; fusing the BEV space features of each modality to obtain the fused BEV space features; inputting the fused BEV space features into the BEV codec to obtain the 3D target perception result of the original network branch; and calculating the original loss of the original network branch based on the 3D target perception result of the original network branch and the corresponding 3D target ground truth data.
[0119] In some embodiments of this application, the 3D target ground truth data includes 3D target ground truth detection box data, and the second input unit 430 is specifically used to: determine the instance features of each modality corresponding to the ground truth detection box of each 3D target based on the BEV space features of each modality and the 3D target ground truth detection box data; and calculate the contrast loss of the contrast network branch using a preset loss function based on the instance features of each modality corresponding to the ground truth detection box of each 3D target.
[0120] In some embodiments of this application, the second input unit 430 is specifically used to: calculate the spatial region of the ground truth detection box of each 3D target on the BEV spatial features of each modality based on the ground truth detection box data of the 3D target; and crop out the instance features at the corresponding positions based on the spatial region of the ground truth detection box of each 3D target on the BEV spatial features of each modality.
[0121] In some embodiments of this application, the second input unit 430 is specifically used to: construct positive and negative samples based on the instance features of each modality corresponding to the ground truth detection box of each 3D target; calculate the pairwise similarity between the instance features of each modality corresponding to the ground truth detection box of each 3D target based on the positive and negative samples; and calculate the contrast loss of the contrast network branch using a preset loss function based on the pairwise similarity between the instance features of each modality corresponding to the ground truth detection box of each 3D target.
[0122] In some embodiments of this application, the update unit 440 is specifically used to: update the network parameters of the single-modal encoder and BEV codec in the original network branch using the original loss of the original network branch; and update the network parameters of the single-modal encoder in the original network branch using the contrast loss of the contrast network branch.
[0123] It is understood that the training device for the 3D target perception model described above can implement each step of the training method for the 3D target perception model provided in the foregoing embodiments. The relevant explanations of the training method for the 3D target perception model are applicable to the training device for the 3D target perception model, and will not be repeated here.
[0124] This application embodiment also provides a 3D target sensing device 500, such as Figure 5 The diagram shows a structural schematic of a 3D target sensing device according to an embodiment of this application. The 3D target sensing device 500 includes: a second acquisition unit 510 and a target sensing unit 520, wherein:
[0125] The second acquisition unit 510 is used to acquire multimodal sensor data;
[0126] The target perception unit 520 is used to perform target perception based on multimodal sensor data and a 3D target perception model to obtain 3D target perception results.
[0127] The 3D target perception model is trained based on the training device for the aforementioned 3D target perception model.
[0128] It is understood that the above-mentioned 3D target perception device can realize all the steps of the 3D target perception method provided in the foregoing embodiments. The relevant explanations of the 3D target perception method are applicable to the 3D target perception device, and will not be repeated here.
[0129] Figure 6 This is a schematic diagram of the structure of a device according to an embodiment of this application. For example... Figure 6 As shown, the device includes one or more processors (or processing units), and may also include one or more memories coupled to the processors, and may also include a communication module coupled to the processors.
[0130] A communication module can be used to communicate with other devices or apparatuses, such as sending or receiving data and / or signals. A communication module may have at least one communication module for communication. A communication module may include any interface necessary for communicating with other devices. Exemplarily, a communication module may be a transceiver, circuit, bus, module, or other type of communication module.
[0131] The processor may include, but is not limited to, one or more of the following: a general-purpose computer, a special-purpose computer, a microcontroller, a digital signal processor (DSP), or a controller-based multi-core controller architecture. The device may have multiple processors, such as application-specific integrated circuit (ASIC) chips, which are time-dependent on a clock synchronized with the main processor.
[0132] The memory may include one or more non-volatile memories and one or more volatile memories. Examples of non-volatile memories include, but are not limited to, at least one of the following: read-only memory (ROM), electrically programmable read-only memory (EPROM), flash memory, hard disk, compact disc (CD), digital video disc (DVD), or other magnetic and / or optical storage. Examples of volatile memories include, but are not limited to, at least one of the following: random access memory (RAM), or other volatile memories that do not persist during the duration of a power outage.
[0133] A computer program consists of computer-executable instructions that are executed by an associated processor. Programs can be stored in ROM. A processor can perform any appropriate action and processing by loading the program into RAM.
[0134] Possible implementations of this application can be achieved through a program, enabling the communication device to execute any of the processes discussed in the foregoing embodiments. Possible implementations of this application can also be achieved through hardware or a combination of software and hardware.
[0135] In some implementations, the program may be tangibly contained in a computer-readable storage medium, which may include in a device (such as in memory) or other storage device accessible by the device. The program may be loaded from the computer-readable storage medium into RAM for execution. The computer-readable storage medium may include any type of tangible non-volatile memory, such as ROM, EPROM, flash memory, hard disk, CD, DVD, etc.
[0136] This application also provides a computer-readable storage medium storing computer instructions or program code thereon, which, when executed by a processor, causes the processor to perform the methods and functions involved in any of the above embodiments. A computer-readable medium can be any tangible medium that contains or stores a program for or relating to an instruction execution system, apparatus, or device. A computer-readable medium can be a computer-readable signal medium or a computer-readable storage medium. A computer-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination thereof. The computer-readable storage medium can be any available medium accessible to a computer or a data storage device such as a server or data center that integrates one or more available media. More detailed examples of computer-readable storage media include electrical connections with one or more wires, magnetic media (e.g., disks, floppy disks, hard disks, magnetic tapes, magnetic storage devices), optical media (e.g., optical storage devices, DVDs), semiconductor media (e.g., solid-state drives), random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), or any suitable combination thereof.
[0137] In the above embodiments, implementation can be achieved, in whole or in part, through software, hardware, firmware, or any combination thereof. When implemented in software, it can be implemented, in whole or in part, as a computer program product. Embodiments of this application also provide at least one computer program product tangibly stored on a non-transitory computer-readable storage medium. This computer program product includes one or more computer-executable instructions, such as instructions included in a program module, which execute in a device on a target's real or virtual processor to perform the processes, methods, and functions involved in any of the above embodiments. When the computer program instructions are loaded and executed on a computer, all or part of the flow or function according to the embodiments of this application is generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transferred from one computer-readable storage medium to another. For example, the computer instructions can be transferred from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line) or wireless (e.g., infrared, wireless, microwave, etc.) means.
[0138] This application also proposes a computer program product, including a computer program or instructions that, when run on a computer, cause the computer to perform the processes, methods, and functions described in the above embodiments. Typically, program modules include routines, programs, libraries, objects, classes, components, data structures, etc., that perform specific tasks or implement specific abstract data types. In various embodiments, the functionality of program modules can be combined or divided as needed. The machine-executable instructions for the program modules can be executed locally or in a distributed device. In a distributed device, the program modules can reside in both local and remote storage media.
[0139] Generally, the various embodiments of this application can be implemented in hardware or dedicated circuitry, software, logic, or any combination thereof. Some aspects can be implemented in hardware, while others can be implemented in firmware or software, which can be executed by a controller, microprocessor, or other computing device. Although various aspects of the embodiments of this disclosure are shown and described as block diagrams, flowcharts, or represented using some other illustration, it should be understood that the blocks, apparatuses, systems, techniques, or methods described herein can be implemented as, as non-limiting examples, in hardware, software, firmware, dedicated circuitry or logic, general-purpose hardware or controllers or other computing devices, or some combination thereof.
[0140] It should be noted that although embodiments of this application have been described above with reference to the accompanying drawings, these embodiments are not independent of each other, and they can be combined to obtain other embodiments. The methods, situations, categories, and classifications of embodiments in this application are only for the convenience of description and should not constitute a special limitation. Various methods, categories, situations, and features in embodiments can be combined with each other if logically consistent. The various embodiments of this application can be arbitrarily combined to achieve different technical effects. The embodiments of this application will not list various combinations.
[0141] Furthermore, although the operation of the methods of this disclosure is described in a specific order in the accompanying drawings, this does not require or imply that these operations must be performed in that specific order, or that all of the operations shown must be performed to achieve the desired result. Rather, the steps depicted in the flowcharts may be performed in a different order. Additionally or alternatively, certain steps may be omitted, multiple steps may be combined into one step, and / or one step may be broken down into multiple steps. It should also be noted that the features and functions of two or more devices according to this disclosure may be embodied in one device. Conversely, the features and functions of one device described above may be further divided and embodied by multiple devices.
[0142] It should also be noted that the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitation, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
[0143] The above description is merely an embodiment of this application and is not intended to limit this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principle of this application should be included within the scope of the claims of this application.
Claims
1. A training method for a 3D target perception model, characterized in that, The training method for the 3D target perception model includes: Acquire multimodal sensor data and corresponding 3D target ground truth data; The multimodal sensor data and the corresponding 3D target ground truth data are input into the original network branch of the 3D target perception model to obtain the original loss of the original network branch. The multimodal sensor data and the corresponding 3D target ground truth data are input into the contrast network branch of the 3D target perception model to obtain the contrast loss of the contrast network branch. The contrast network branch performs contrast learning based on the detection box of the 3D target. The parameters of the 3D target perception model are updated based on the original loss of the original network branch and the contrastive loss of the contrastive network branch to obtain the trained 3D target perception model.
2. The training method for the 3D target perception model according to claim 1, characterized in that, The original network branch includes a single-modal encoder and a BEV encoder / decoder. Sensor data for each modality corresponds to one of the single-modal encoders. The process of inputting the multimodal sensor data and the corresponding 3D target ground truth data into the original network branch of the 3D target perception model to obtain the original loss of the original network branch includes: The sensor data for each mode are input into each of the single-mode encoders to obtain the coding features for each mode; The encoded features of each modality are transformed into the BEV space to obtain the BEV space features of each modality. The BEV spatial features of each mode are fused to obtain the fused BEV spatial features; The fused BEV spatial features are input into the BEV codec to obtain the 3D target perception results of the original network branch; The original loss of the original network branch is calculated based on the 3D target perception results of the original network branch and the corresponding 3D target ground truth data.
3. The training method for the 3D target perception model according to claim 2, characterized in that, The 3D target ground truth data includes the ground truth detection box data of the 3D target. The step of inputting the multimodal sensor data and the corresponding 3D target ground truth data into the contrast network branch of the 3D target perception model to obtain the contrast loss of the contrast network branch includes: Based on the BEV space features of each modality and the ground truth detection box data of the 3D target, determine the instance features of each modality corresponding to the ground truth detection box of each 3D target; Based on the instance features of each modality corresponding to the ground truth detection box of each 3D target, the contrast loss of the contrast network branch is calculated using a preset loss function.
4. The training method for the 3D target perception model according to claim 3, characterized in that, The step of determining the instance features of each modality corresponding to the ground truth detection box of each 3D target based on the BEV space features of each modality and the ground truth detection box data of the 3D target includes: Based on the ground truth detection box data of the 3D target, calculate the spatial region of the ground truth detection box of each 3D target on the BEV spatial features of each modality; Based on the spatial region of the ground truth detection box of each 3D target on the BEV spatial features of each modality, the instance features at the corresponding positions are cropped out.
5. The training method for the 3D target perception model according to claim 3, characterized in that, The step of calculating the contrast loss of the contrast network branch based on the instance features of each modality corresponding to the ground truth detection box of each 3D target using a preset loss function includes: Positive and negative samples are constructed based on the instance features of each modality corresponding to the ground truth detection box of each 3D target. Based on the positive and negative samples, calculate the pairwise similarity between the instance features of each modality corresponding to the ground truth detection box of each 3D target; Based on the pairwise similarity between instance features of each modality corresponding to the ground truth detection box of each 3D target, the contrast loss of the contrast network branch is calculated using a preset loss function.
6. The training method for the 3D target perception model according to claim 1, characterized in that, The step of updating the parameters of the 3D object perception model based on the original loss of the original network branch and the contrastive loss of the contrastive network branch to obtain the trained 3D object perception model includes: The network parameters of the single-modal encoder and BEV codec in the original network branch are updated using the original loss of the original network branch. The network parameters of the single-modal encoder in the original network branch are updated using the contrast loss of the contrast network branch.
7. A 3D target perception method, characterized in that, The 3D target perception method includes: Acquire multimodal sensor data; Based on multimodal sensor data, a 3D target perception model is used to perform target perception and obtain 3D target perception results. The 3D target perception model is trained based on the training method of the 3D target perception model according to any one of claims 1 to 6.
8. A training device for a 3D target perception model, characterized in that, The training device for the 3D target perception model includes: The first acquisition unit is used to acquire multimodal sensor data and corresponding 3D target ground truth data; The first input unit is used to input the multimodal sensor data and the corresponding 3D target ground truth data into the original network branch of the 3D target perception model to obtain the original loss of the original network branch. The second input unit is used to input the multimodal sensor data and the corresponding 3D target ground truth data into the contrast network branch of the 3D target perception model to obtain the contrast loss of the contrast network branch, and the contrast network branch performs contrast learning based on the detection box of the 3D target. The update unit is used to update the parameters of the 3D target perception model based on the original loss of the original network branch and the contrastive loss of the contrastive network branch, so as to obtain the trained 3D target perception model.
9. A 3D target sensing device, characterized in that, The 3D target sensing device includes: The second acquisition unit is used to acquire multimodal sensor data; The target perception unit is used to perform target perception based on multimodal sensor data and a 3D target perception model to obtain 3D target perception results. The 3D target perception model is trained based on the training device of the 3D target perception model of claim 8.
10. An apparatus comprising: processor; And a memory arranged to store computer-executable instructions, which, when executed, cause the processor to perform the training method of any one of claims 1 to 6 for the 3D target perception model, and to perform the 3D target perception method of claim 7.
11. A computer program product, comprising a computer program / instructions, characterized in that, When the computer program / instruction is executed by the processor, it implements the training method of any one of the 3D target perception models of claims 1 to 6, and the 3D target perception method of claim 7.