Model training and target recognition method
Patent Information
- Application Number
- CN202610940209.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-26
- Publication Date
- 2026-09-22
AI Technical Summary
[0007]本申请提供了一种模型训练及目标识别方法,以至少解决相关目标检测识别技术在自动驾驶功能安全实车测试中存在复杂路况抗干扰能力弱,单模态失效时检测性能骤降,边缘细节定位不准且无法适配自动驾驶功能安全实车测试的迭代优化机制的技术问题
[0018]在本申请中,采用获取训练数据,其中,训练数据包括车辆所在道路的可见光图像和深度图像;利用跨模态反向引导模型中的编码器分别对可见光图像和深度图像进行编码特征提取,得到第一编码特征以及第二编码特征;利用跨模态反向引导模型中的解码器分别对第一编码特征和第二编码特征进行引导降噪,得到第一降噪特征和第二降噪特征;对第一降噪特征和第二降噪特征进行跨模态注意力交互,得到第一增强编码特征和第二增强编码特征;对第一增强编码特征和第二增强编码特征分别进行全局池化处理,并对处理结果进行拼接,得到融合特征;对融合特征进行逐层解码,得到末级解码特征,并对末级解码特征和可见光图像进行边缘细化处理,得到边缘细化特征;将边缘细化特征上采样至可见光图像的尺寸,得到目标检测图像;将目标检测图像发送至自动驾驶决策系统,以指示自动驾驶决策系统根据目标检测图像进行功能安全实车测试,并接收自动驾驶决策系统发送的与功能安全测试相关的反馈数据;基于反馈数据,对跨模态反向引导模型进行参数迭代优化,直至满足预设停止条件,得到完成训练的跨模态反向引导模型的方式,达到了提升模型在复杂路况下的感知精度、鲁棒性及实时性,并建立功能安全闭环优化机制的目的,从而实现了单模态失效时检测性能稳定、边缘细节定位精准且模型具备实车场景泛化能力的技术效果,进而解决了相关目标检测识别技术在自动驾驶功能安全实车测试中存在复杂路况抗干扰能力弱,单模态失效时检测性能骤降,边缘细节定位不准且无法适配自动驾驶功能安全实车测试的迭代优化机制的技术问题。
Smart Images

Figure CN122799397A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of autonomous driving environmental perception and functional safety testing, and more specifically, to a model training and target recognition method. Background Technology
[0002] Significant object detection is a core technology for environmental perception in autonomous driving, and its detection accuracy and robustness directly determine the overall vehicle functional safety level. Currently, RGB-D multimodal detection, due to its fusion of visual and depth information, has become the mainstream solution for real-world functional safety testing. However, related technologies still have significant limitations in scenarios oriented towards functional safety testing in real-world vehicles:
[0003] Insufficient anti-interference and robustness: The encoding and decoding features are not fully utilized, and low-level features are easily affected by sensor noise and road clutter; moreover, the model is overly dependent on dual-modal features, and the fusion results are easily distorted in single-modal failure scenarios such as backlight, rain, and tunnels, resulting in a sharp drop in detection performance and a lack of sufficient functional safety redundancy.
[0004] Low edge positioning accuracy: The vehicle-mounted target edge detail extraction capability is weak, and the boundary recognition of pedestrian limbs, road markings, etc. is blurry, which can easily lead to decision-making errors and cannot meet the stringent requirements of functional safety for perception accuracy.
[0005] Lack of real-vehicle iterative optimization mechanism: Most of the relevant models are trained on public datasets and have not been specifically optimized for complex road conditions in real vehicles. Furthermore, there is a lack of iterative optimization mechanism based on feedback from real vehicle tests, resulting in a significant gap between laboratory accuracy and actual real-vehicle performance, making it difficult to support the test data requirements for functional safety certification.
[0006] There is currently no effective solution to the above problems. Summary of the Invention
[0007] This application provides a model training and target recognition method to at least solve the technical problems of related target detection and recognition technologies in real-vehicle testing of autonomous driving functional safety, such as weak anti-interference ability in complex road conditions, sharp drop in detection performance when single-mode failure occurs, inaccurate edge detail localization, and inability to adapt to the iterative optimization mechanism of real-vehicle testing of autonomous driving functional safety.
[0008] According to one aspect of this application, a model training method is provided, comprising: acquiring training data, wherein the training data includes a visible light image and a depth image of a road where a vehicle is located; extracting coded features from the visible light image and the depth image respectively using an encoder in a cross-modal back-guided model to obtain a first coded feature and a second coded feature; performing guided denoising on the first coded feature and the second coded feature respectively using a decoder in the cross-modal back-guided model to obtain a first denoised feature and a second denoised feature; performing cross-modal attention interaction on the first denoised feature and the second denoised feature to obtain a first enhanced coded feature and a second enhanced coded feature; and performing full-processing on the first enhanced coded feature and the second enhanced coded feature respectively. Local pooling is performed, and the results are concatenated to obtain fused features. The fused features are then decoded layer by layer to obtain final-level decoded features. Edge thinning is then performed on the final-level decoded features and the visible light image to obtain edge thinning features. The edge thinning features are upsampled to the size of the visible light image to obtain the target detection image. The target detection image is sent to the autonomous driving decision system to instruct the system to conduct functional safety real-vehicle testing based on the target detection image, and feedback data related to functional safety testing is received from the system. Based on the feedback data, the parameters of the cross-modal reverse guidance model are iteratively optimized until a preset stopping condition is met, resulting in a fully trained cross-modal reverse guidance model.
[0009] Optionally, the encoder in the cross-modal back-guided model is used to extract coded features from the visible light image and the depth image respectively, to obtain a first coded feature and a second coded feature. This includes: extracting coded features from four stages through the visible light coding branch to obtain a first coded feature containing features from multiple stages; and extracting coded features from four stages through the depth coding branch to obtain a second coded feature containing features from multiple stages. The visible light coding branch and the depth coding branch share weights, and both are constructed using neural networks based on a self-attention mechanism. In both the visible light coding branch and the depth coding branch, adjacent stages are downsampled through block merging operations, making the feature resolution of stage i+1 half that of stage i, and the number of feature channels in stage i+1 twice that of stage i, where i is a positive integer not less than 1. The output feature of stage i of the visible light coding branch is used as part of the input of stage i+1 of the depth coding branch, and the output feature of stage i of the depth coding branch is used as part of the input of stage i+1 of the visible light coding branch, so that the visible light coding branch and the depth coding branch can interact across modalities during the encoding process.
[0010] Optionally, the decoder in the cross-modal reverse guidance model includes cascaded multi-level decoding units, which sequentially output multi-level decoding features, and guided denoising and decoding are performed alternately at each level. Guided denoising is performed on the first and second encoded features using the decoder in the cross-modal reverse guidance model to obtain the first and second denoised features, including: for the current level of guided denoising, using the decoded features output by the previous level decoding unit in the decoder as the guiding features, where the guiding features are used to represent the semantic information of the image; upsampling the guiding features to obtain adjusted guiding features; performing activation processing on the adjusted guiding features to obtain an adaptive weight vector; multiplying the adaptive weight vector element-wise with the first encoded feature to obtain a first weighted feature, and adding the first weighted feature element-wise with the adjusted guiding feature to obtain the first denoised feature; multiplying the adaptive weight vector element-wise with the second encoded feature to obtain a second weighted feature, and adding the second weighted feature element-wise with the adjusted guiding feature to obtain the second denoised feature.
[0011] Optionally, cross-modal attention interaction is performed on the first denoising feature and the second denoising feature to obtain a first enhanced coding feature and a second enhanced coding feature, including: performing a linear transformation on the first denoising feature to obtain a first query vector, a first key vector, and a first value vector; performing a linear transformation on the second denoising feature to obtain a second query vector, a second key vector, and a second value vector; calculating a first similarity between the first query vector and the second key vector; normalizing the first similarity to obtain a first attention weight; using the first attention weight to perform a weighted summation on the second value vector to obtain a first cross-modal attention feature; calculating a second similarity between the second query vector and the first key vector; normalizing the second similarity to obtain a second attention weight; using the second attention weight to perform a weighted summation on the first value vector to obtain a second cross-modal attention feature; adding the first cross-modal attention feature and the first denoising feature element-wise to obtain a first enhanced coding feature; and adding the second cross-modal attention feature and the second denoising feature element-wise to obtain a second enhanced coding feature.
[0012] Optionally, global pooling is performed on the first enhanced coding feature and the second enhanced coding feature respectively, and the processing results are concatenated to obtain a fused feature, including: performing global average pooling on the first enhanced coding feature to obtain a first global feature; performing global average pooling on the second enhanced coding feature to obtain a second global feature; concatenating the first global feature and the second global feature to obtain a concatenated global feature; inputting the concatenated global feature into a fully connected layer, and performing activation processing on the output of the fully connected layer to obtain an adaptive channel vector; performing a weighted multiplication of the adaptive channel vector and the first enhanced coding feature, and then multiplying the weighted multiplication result with... The first enhanced coding features are added element-wise to obtain the first improved feature; the adaptive channel vector is multiplied with the second enhanced coding feature by weight, and the result of the weighted multiplication is added element-wise with the second enhanced coding feature to obtain the second improved feature; the first improved feature and the second improved feature are activated respectively to obtain the first fusion weight and the second fusion weight; the first improved feature is enhanced by weighting with the first fusion weight to obtain the first weighted enhanced feature; the second improved feature is enhanced by weighting with the second fusion weight to obtain the second weighted enhanced feature; the first weighted enhanced feature and the second weighted enhanced feature are concatenated to obtain the fused feature.
[0013] Optionally, the decoder includes cascaded multi-level decoding units, each built based on a self-attention mechanism; the fused features are decoded layer by layer to obtain the final-level decoded features, and the final-level decoded features and the visible light image are subjected to edge refinement processing to obtain edge refinement features, including: inputting the fused features into the decoder, performing decoding processing through the multi-level decoding units and restoring the spatial resolution of the features layer by layer to obtain multi-level decoded features; taking the last level of the multi-level decoded features as the final-level decoded features; performing feature transformation on the final-level decoded features and performing feature transformation on the visible light image so that the feature dimensions of the transformed final-level decoded features and the transformed visible light image are the same; stitching the transformed final-level decoded features and the transformed visible light image to obtain stitched features; performing multi-scale attention processing on the stitched features to obtain multi-scale edge interaction features; performing feature transformation and multi-scale attention processing on the multi-scale edge interaction features and the visible light image to obtain optimized edge interaction features; upsampling the optimized edge interaction features so that the size of the optimized edge interaction features is the same as the size of the visible light image to obtain edge refinement features.
[0014] Optionally, the spatial resolution of the features is restored step by step through multi-level decoding units, including: performing decoding processing step by step through multi-level decoding units, such that the spatial resolution of the decoded features output by each level of decoding unit is twice the spatial resolution of the decoded features output by the level above it; performing feature transformation on the final-level decoded features, including: sequentially performing convolution processing, batch normalization processing, and activation processing on the final-level decoded features; performing feature transformation on the visible light image, including: sequentially performing convolution processing, batch normalization processing, and activation processing on the visible light image; performing feature transformation and multi-scale attention processing on the multi-scale edge interaction features and the visible light image, including: sequentially performing convolution processing, batch normalization processing, and activation processing on the multi-scale edge interaction features; sequentially performing convolution processing, batch normalization processing, and activation processing on the visible light image; and concatenating the processed multi-scale edge interaction features and the processed visible light image along the channel dimension, and performing multi-scale attention processing on the concatenated result to obtain the optimized edge interaction features.
[0015] According to another aspect of this application, a target recognition method is also provided, comprising: acquiring an image to be recognized; and performing target recognition on the image to be recognized using a cross-modal back-guiding model, wherein the cross-modal back-guiding model is trained using the above-described model training method.
[0016] According to another aspect of this application, a model training apparatus is also provided, comprising: an acquisition module for acquiring training data, wherein the training data includes a visible light image and a depth image of the road where the vehicle is located; an extraction module for extracting coded features from the visible light image and the depth image respectively using an encoder in a cross-modal back-guided model to obtain a first coded feature and a second coded feature; a denoising module for guiding denoising of the first coded feature and the second coded feature respectively using a decoder in a cross-modal back-guided model to obtain a first denoised feature and a second denoised feature; performing cross-modal attention interaction on the first denoised feature and the second denoised feature to obtain a first enhanced coded feature and a second enhanced coded feature; and a processing module for processing the first enhanced coded feature and the second enhanced coded feature. The code features are globally pooled and the results are concatenated to obtain fused features. The fused features are then decoded layer by layer to obtain final-level decoded features. The final-level decoded features and the visible light image are then edge-refined to obtain edge-refined features. The edge-refined features are upsampled to the size of the visible light image to obtain the target detection image. The receiving module sends the target detection image to the autonomous driving decision system to instruct the autonomous driving decision system to conduct functional safety real-vehicle testing based on the target detection image, and receives feedback data related to functional safety testing sent by the autonomous driving decision system. The training module iteratively optimizes the parameters of the cross-modal reverse guidance model based on the feedback data until a preset stopping condition is met, resulting in a trained cross-modal reverse guidance model.
[0017] According to another aspect of this application, an electronic device is also provided, comprising: a memory and a processor, the processor being configured to run a program stored in the memory, wherein the program executes the above-described model training method during runtime.
[0018] In this application, the following steps are employed: Training data is acquired, including visible light and depth images of the road where the vehicle is located; the encoder in a cross-modal back-guided model extracts coded features from the visible light and depth images to obtain first and second coded features; the decoder in the cross-modal back-guided model performs guided denoising on the first and second coded features to obtain first and second denoised features; cross-modal attention interaction is performed on the first and second denoised features to obtain first and second enhanced coded features; global pooling is performed on the first and second enhanced coded features respectively, and the results are concatenated to obtain fused features; the fused features are decoded layer by layer to obtain final-level decoded features, and edge thinning processing is performed on the final-level decoded features and the visible light image to obtain edge thinning features; the edge thinning features are upsampled to the size of the visible light image to obtain the target image. The system detects targets and sends the detected images to the autonomous driving decision-making system, instructing the system to conduct functional safety real-vehicle testing based on the images. It also receives feedback data related to the functional safety test from the autonomous driving decision-making system. Based on this feedback data, iterative optimization of the parameters of the cross-modal reverse guidance model is performed until a preset stopping condition is met, resulting in a fully trained cross-modal reverse guidance model. This approach improves the model's perception accuracy, robustness, and real-time performance under complex road conditions and establishes a functional safety closed-loop optimization mechanism. This achieves stable detection performance under single-modal failure, accurate edge detail localization, and the model's ability to generalize to real-vehicle scenarios. Furthermore, it solves the technical problems of related target detection and recognition technologies in autonomous driving functional safety real-vehicle testing, such as weak anti-interference capability under complex road conditions, sharp drop in detection performance under single-modal failure, inaccurate edge detail localization, and inability to adapt to the iterative optimization mechanism of autonomous driving functional safety real-vehicle testing. Attached Figure Description
[0019] The accompanying drawings, which are included to provide a further understanding of this application and form part of this application, illustrate exemplary embodiments of this application and are used to explain this application, but do not constitute an undue limitation of this application. In the drawings:
[0020] Figure 1 This is a flowchart of a model training method according to an embodiment of this application;
[0021] Figure 2 This is a schematic diagram of a model training method according to an embodiment of this application;
[0022] Figure 3 This is a schematic diagram of a multi-level feature reverse guidance enhancement according to an embodiment of this application;
[0023] Figure 4This is a schematic diagram of a dual-stream interactive fusion according to an embodiment of this application;
[0024] Figure 5 This is a schematic diagram of edge refinement sensing according to an embodiment of this application;
[0025] Figure 6 This is a structural diagram of a model training device according to an embodiment of this application;
[0026] Figure 7 This is a hardware structure block diagram of a computer terminal for a model training method according to an embodiment of this application. Detailed Implementation
[0027] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present application, and not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative effort should fall within the scope of protection of the present application.
[0028] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0029] According to an embodiment of this application, a method embodiment for model training is provided. It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. Furthermore, although a logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order than that shown here.
[0030] Figure 1 This is a flowchart of a model training method according to an embodiment of this application, such as... Figure 1 As shown, the method includes the following steps:
[0031] Step S101: Obtain training data, which includes visible light images and depth images of the road where the vehicle is located.
[0032] In step S101, data of the road scene where the vehicle is located is collected using an onboard multimodal perception system. The training data includes at least visible light images (RGB images) and depth images simultaneously collected during vehicle movement. The visible light images provide rich texture and color information, while the depth images provide three-dimensional spatial geometry information of the scene. To ensure data quality, the collected raw data can be preprocessed, including but not limited to size normalization (e.g., uniformly adjusted to 224×224 or the standard size required for network input), denoising, and filtering of invalid data (such as blurred images caused by sensor occlusion or severe bumps), thereby obtaining high-quality training samples.
[0033] Step S102: The encoder in the cross-modal reverse guidance model is used to extract coded features from the visible light image and the depth image respectively, to obtain the first coded feature and the second coded feature.
[0034] In step S102, the preprocessed visible light image is input to the RGB encoding branch, and the depth image is input to the depth encoding branch. The encoder can be constructed using a weight-sharing Transformer backbone network (such as Swin-Transformer), which can effectively capture local details and global contextual information in the image. Through the forward propagation of the encoder, multi-level encoded features are extracted, denoted as the first encoded feature (corresponding to the visible light modality) and the second encoded feature (corresponding to the depth modality). The first and second encoded features retain the key semantic information of the original image and provide a basis for feature enhancement.
[0035] Step S103: The decoder in the cross-modal reverse guidance model is used to guide the denoising of the first coding feature and the second coding feature respectively to obtain the first denoised feature and the second denoised feature; cross-modal attention interaction is performed on the first denoised feature and the second denoised feature to obtain the first enhanced coding feature and the second enhanced coding feature.
[0036] Guided denoising refers to a processing mechanism that adaptively weights and suppresses noise in low-level features based on high-level semantic information. Specifically, in the cross-modal reverse guided model of this embodiment, the guided denoising process can use the high-level decoded features output by the decoder (which contain significant geometric structure information and target semantic information and are less affected by low-level noise) as a guiding signal to process the low-level encoded features output by the encoder (which contain rich texture details but are susceptible to sensor noise, lighting changes, and road clutter).
[0037] In step S103, the high-level decoded features output by the decoder (which have stronger semantic information and significant region localization capabilities) are used to reverse-engineer the first and second encoded features obtained in step S102. Specifically, the high-level decoded features can be reshaped to the same spatial resolution as the low-level encoded features, and an adaptive weight vector can be generated using an activation function (such as sigmoid) to weight the low-level encoded features, thereby obtaining the first and second denoising features. This process can effectively suppress sensor noise and road clutter interference in vehicular scenarios.
[0038] The first and second denoising features are converted into query, key, and value vectors, respectively. A cross-modal attention mechanism is used to achieve information complementarity between the RGB and deep modalities. For example, attention is calculated using the query vector from the RGB modality and the key / value vector from the deep modality, and vice versa. Through this interaction, the model can integrate the advantages of both modalities to obtain the first and second enhanced coding features, thereby improving the robustness of feature representation.
[0039] Step S104: Perform global pooling on the first enhanced coding feature and the second enhanced coding feature respectively, and concatenate the processing results to obtain the fused feature; decode the fused feature layer by layer to obtain the final-level decoded feature, and perform edge thinning processing on the final-level decoded feature and the visible light image to obtain the edge thinning feature; upsample the edge thinning feature to the size of the visible light image to obtain the target detection image.
[0040] In this context, the final-level decoded feature refers to the last layer of decoded features output by the decoder module in the cross-modal backpropagation model. Specifically, during the forward propagation of the model, the fused features are processed layer by layer by the decoder module. The decoder consists of multiple decoder blocks (such as Swin-Transformer-based decoder blocks) stacked together. Each decoder block is responsible for progressively restoring the spatial resolution of the features (e.g., through upsampling operations) and extracting higher-level semantic information. After the fused features have been processed by all decoder layers in sequence, the final feature output from the end of the decoder is the final-level decoded feature.
[0041] In step S104, global average pooling is performed on the first and second enhanced coding features respectively to extract global contextual information. Then, the global information from the two modalities is concatenated, and an adaptive channel vector is generated through a fully connected layer and activation function. This channel vector is used to weight the original enhanced features, and a fusion weight is generated using the Sigmoid function. Finally, the weighted bimodal features are concatenated to obtain the fused features. This fusion strategy reduces the model's dependence on a single modality and improves robustness in scenarios where a single modality fails.
[0042] The fused features are input into the decoder module. The decoder gradually recovers the spatial resolution of the features through layer-by-layer upsampling operations, generating multi-stage decoded features. Finally, the final-level decoded features are output, which contain rich semantic information.
[0043] To accurately reconstruct the target boundary, the final-level decoded features are fused with the original visible light image. Specifically, both are convolved to map to the same feature dimension and then concatenated. An efficient multi-scale attention (EMA) mechanism is then input to extract edge interaction information at different scales. After further optimization, the edge refinement features are upsampled to the size of the original visible light image, ultimately yielding a high-precision target detection image (salient target detection map).
[0044] Step S105: Send the target detection image to the autonomous driving decision system to instruct the autonomous driving decision system to conduct functional safety real vehicle testing based on the target detection image, and receive feedback data related to functional safety testing sent by the autonomous driving decision system.
[0045] In step S105, the target detection image generated in step S104 is sent to the autonomous driving decision-making system in real time. The autonomous driving decision-making system performs road condition target recognition, warning, and decision-making actions in the functional safety real-vehicle test based on this target detection image. During this process, the system collects feedback data related to functional safety, including but not limited to: detailed detection errors under complex road conditions, missed / false detection data for core targets (such as pedestrians and obstacles), edge recognition error data, and perception accuracy indicators under different scenarios. The feedback data accurately reflects the model's performance in complex real-vehicle environments and is an important basis for model iterative optimization.
[0046] Step S106: Based on the feedback data, perform iterative optimization of the parameters of the cross-modal back-guided model until the preset stopping condition is met, and obtain the cross-modal back-guided model that has been trained.
[0047] In step S106, based on the functional safety test feedback data received in step S105, the parameters of the cross-modal back-guided model are iteratively optimized. Specifically, an optimizer (such as the Adam optimizer) is used to calculate the loss function, and the network parameters are updated through backpropagation. During this process, an initial learning rate and decay strategy can be set (e.g., the learning rate decays by a factor of 5 every 40 training epochs) to balance convergence speed and final accuracy. Steps S101 to S105 are repeated (or the process is simulated during offline training) until the model performance meets the preset stopping conditions (e.g., the mean absolute error (MAE) on the validation set is lower than a threshold, or the F-measure reaches a predetermined standard), thereby obtaining a cross-modal back-guided model that has been trained and is adapted to the requirements of functional safety real-vehicle testing.
[0048] The cross-modal reverse guidance model is primarily used to identify and locate salient targets in autonomous driving scenarios. Specifically, the cross-modal reverse guidance model can identify dynamic and static targets including the following categories: Traffic participants: Pedestrians: including pedestrians walking normally, crossing the road, or those at the edge of the road; Non-motorized vehicles: including bicycles, electric vehicles, motorcycles, and their drivers; Motorized vehicles: including cars, buses, trucks, lorries, and other road vehicles. Traffic facilities and static obstacles: Road markings: including lane lines, zebra crossings, stop lines, and other road markings with significant geometric features; Static obstacles: including traffic cones, construction barriers, fallen objects, guardrails, and other objects that block passageways. Other salient areas: In specific scenarios, this may include visually impactful elements such as traffic signs and traffic lights.
[0049] The recognition results are presented as follows: The target detection image (or saliency map) output by the model exists in the form of a pixel-level probability map, where: High response regions: correspond to the identified salient targets (such as pedestrians and vehicles) and their outlines, with a high confidence value (close to 1); Low response regions: correspond to the background environment (such as roads, sky, and building backgrounds), with a low confidence value (close to 0).
[0050] The steps S101 to S106 above achieve the goal of improving the model's perception accuracy, robustness, and real-time performance under complex road conditions, and establishing a functional safety closed-loop optimization mechanism. This achieves the technical effect of stable detection performance, accurate edge detail localization, and model generalization ability in real vehicle scenarios when a single mode fails.
[0051] The following are Figure 1 The steps shown are illustrated and explained by way of example.
[0052] According to some optional embodiments of this application, the encoder in the cross-modal back-guided model is used to extract coded features from visible light images and depth images respectively, obtaining first coded features and second coded features. This can be achieved through the following steps: extracting coded features from four stages through the visible light coding branch to obtain first coded features containing features from multiple stages; extracting coded features from four stages through the depth coding branch to obtain second coded features containing features from multiple stages; wherein the visible light coding branch and the depth coding branch share weights, and both the visible light coding branch and the depth coding branch are constructed based on a self-attention mechanism neural network; In both the visible light coding branch and the depth coding branch, downsampling is performed between adjacent stages through block merging operations, so that the feature resolution of the (i+1)th stage is half that of the i-th stage, and the number of feature channels in the (i+1)th stage is twice that of the i-th stage, where i is a positive integer not less than 1. The output features of the i-th stage of the visible light coding branch are used as part of the input of the i+1th stage of the depth coding branch, and the output features of the i-th stage of the depth coding branch are used as part of the input of the i+1th stage of the visible light coding branch, so that the visible light coding branch and the depth coding branch can interact with each other across modalities during the coding process.
[0053] In this embodiment, the encoder includes a visible light coding branch and a depth coding branch. The two branches adopt a weight-sharing strategy, that is, they share the same network parameters to reduce the number of model parameters and adapt to the computing power limitations of automotive embedded hardware. Each coding branch is constructed by a neural network based on a self-attention mechanism (such as Swin-Transformer) to capture global contextual information and local detail features of the image.
[0054] During the feature extraction process, the visible light coding branch encodes the input visible light image and extracts the coding features of four stages in sequence, thereby obtaining the first coding feature containing multiple stage features; at the same time, the depth coding branch encodes the input depth image and similarly extracts the coding features of four stages in sequence, thereby obtaining the second coding feature containing multiple stage features.
[0055] The first and second coding features each contain four stages of features, denoted as first-stage features, second-stage features, third-stage features, and fourth-stage features, respectively. These four stages constitute a multi-scale feature representation from low-level details to high-level semantics, specifically defined as follows:
[0056] The first-stage features are the initial output features of the encoding process. If the input image resolution is H×W, the spatial resolution of the first-stage features is typically one-quarter of the input resolution, i.e., H / 4×W / 4 (or H / 2×W / 2 in some implementations, depending on the stride setting of the preceding convolutional layers). The number of channels for this stage is set to the base number C. This stage preserves the richest low-level texture and edge details of the image, but has a smaller receptive field and is more susceptible to noise.
[0057] The second-stage features are obtained by downsampling the first-stage features. Their spatial resolution is reduced to half that of the first-stage features, i.e., H / 8 × W / 8 (or H / 4 × W / 4). Simultaneously, their channel count is doubled compared to the first-stage features, i.e., 2C. This stage of features retains some spatial details while beginning to incorporate local contextual information.
[0058] The third-stage features are obtained by further downsampling the second-stage features. Their spatial resolution is reduced to half that of the second-stage features, i.e., H / 16×W / 16 (or H / 8×W / 8). Their channel count is doubled compared to the second-stage features, i.e., 4C. These features have a moderately sized receptive field, enabling the extraction of more discriminative semantic information.
[0059] The fourth-stage feature, obtained by further downsampling the third-stage feature, is the highest-level feature in the encoding process. Its spatial resolution is reduced to half that of the third-stage feature, i.e., H / 32×W / 32 (or H / 16×W / 16). Its channel count is doubled to 8C. This stage feature has the largest receptive field, contains the strongest high-level semantic information, and has the strongest representational ability for the overall structure of the target and scene, but is relatively vague in terms of subtle spatial location information.
[0060] In the above process, for any positive integer i, the spatial resolution of the features in the (i+1)th stage is half that of the features in the i-th stage, and the number of channels in the (i+1)th stage features is twice that of the features in the i-th stage. This hierarchical feature extraction method enables the model to capture features of salient targets at different scales, providing multi-level information support for subsequent reverse guidance and edge refinement.
[0061] Within each encoding branch, adjacent stages undergo downsampling through block merging. Specifically, for any positive integer i not less than 1, the feature resolution of the (i+1)th stage is set to half that of the i-th stage to compress spatial dimensions; simultaneously, the number of feature channels in the (i+1)th stage is doubled that of the i-th stage to increase the semantic expressiveness of the features. This hierarchical downsampling and channel expansion mechanism allows the model to gradually transition from low-level texture features to high-level semantic features.
[0062] To achieve early information complementarity between modalities, the visible light coding branch and the deep coding branch engage in cross-modal information exchange during the encoding process. Specifically, the output features of the i-th stage of the visible light coding branch are used as part of the input to the (i+1)-th stage of the deep coding branch, and vice versa. Through this interleaved feature transfer mechanism, the two coding branches can refer to each other's semantic information when extracting features at different stages, thereby achieving deep fusion of RGB modalities and deep modalities during the encoding stage.
[0063] The above steps, through a weight-sharing dual-stream self-attention encoder combined with cross-modal feature interaction between adjacent stages, effectively promote the complementary fusion of RGB and deep modalities while achieving multi-scale feature extraction, thereby improving the robustness of feature representation and reducing the number of model parameters.
[0064] According to some optional embodiments of this application, the decoder in the cross-modal reverse guidance model includes cascaded multi-level decoding units, which sequentially output multi-level decoding features, and guided denoising and decoding are performed alternately at each level. Guided denoising is performed on the first and second encoded features using the decoder in the cross-modal reverse guidance model to obtain the first and second denoised features, which can be achieved through the following steps: For guided denoising at the current level, the decoding features output by the previous level decoding unit in the decoder are used as guiding features, wherein the guiding features are used to characterize the semantic information of the image; the guiding features are upsampled to obtain adjusted guiding features; the adjusted guiding features are activated to obtain an adaptive weight vector; the adaptive weight vector is multiplied element-wise by the first encoded feature to obtain a first weighted feature, and the first weighted feature is added element-wise by the adjusted guiding feature to obtain the first denoised feature; the adaptive weight vector is multiplied element-wise by the second encoded feature to obtain a second weighted feature, and the second weighted feature is added element-wise by the adjusted guiding feature to obtain the second denoised feature.
[0065] In this embodiment, the decoder in the cross-modal reverse guidance model adopts a cascaded structure, consisting of multiple multi-level decoding units connected sequentially. Each level of decoding unit is responsible for outputting the corresponding multi-level decoded features, thereby achieving a gradual recovery from low-resolution high-level semantics to high-resolution detailed features. In this architecture, guided denoising operations and decoding operations are performed alternately at each level, that is, guided denoising steps are embedded before or during each level of decoding to ensure that high-level semantic information can be used to optimize low-level encoded features at each stage of feature decoding.
[0066] Specifically, the process of guided denoising using the decoder on the first coded features (i.e., features extracted from the visible light coding branch) and the second coded features (i.e., features extracted from the deep coding branch) is as follows: For the guided denoising operation of the current level, the decoded features output by the previous level decoding unit are first obtained from the decoder as guided features. Since the higher-level decoded features contain richer global semantic information and prior knowledge of the location of salient targets, they are used as guided signals to characterize the salient regions that the current coded features to be processed should focus on.
[0067] Then, to align high-level semantic information with low-level encoded features spatially, the guiding features are upsampled to match their spatial resolution with the first and second encoded features of the current level input, resulting in adjusted guiding features. These adjusted guiding features are then activated (e.g., using the sigmoid activation function) to generate an adaptive weight vector with values between 0 and 1. This adaptive weight vector reflects the importance of each spatial location to the salient target; regions with higher weights correspond to salient targets, while regions with lower weights correspond to background noise.
[0068] Based on this adaptive weight vector, the first and second encoded features are weighted separately. Specifically, the adaptive weight vector is multiplied element-wise with the first encoded feature to obtain the first weighted feature, thereby suppressing noise in non-salient regions and enhancing information in salient regions. Simultaneously, the adaptive weight vector is multiplied element-wise with the second encoded feature to obtain the second weighted feature. Finally, to inject guiding semantics into the encoded features, the first weighted feature is added element-wise with the adjusted guiding feature to obtain the first denoising feature; similarly, the second weighted feature is added element-wise with the adjusted guiding feature to obtain the second denoising feature. Through this mechanism, the high-level semantic information of the decoder effectively guides and purifies the low-level features of the encoder, achieving collaborative denoising and enhancement of cross-modal features.
[0069] The above steps utilize high-level semantic features of the decoder to generate adaptive weights to guide and weight-enhance low-level encoded features, achieving multi-level feature denoising and semantic injection, effectively improving the purity and discriminative power of features in significant target regions.
[0070] In some optional embodiments of this application, cross-modal attention interaction is performed on the first denoising feature and the second denoising feature to obtain the first enhanced coding feature and the second enhanced coding feature. This can be achieved through the following steps: performing a linear transformation on the first denoising feature to obtain a first query vector, a first key vector, and a first value vector; performing a linear transformation on the second denoising feature to obtain a second query vector, a second key vector, and a second value vector; calculating a first similarity between the first query vector and the second key vector; normalizing the first similarity to obtain a first attention weight; using the first attention weight to perform a weighted summation on the second value vector to obtain a first cross-modal attention feature; calculating a second similarity between the second query vector and the first key vector; normalizing the second similarity to obtain a second attention weight; using the second attention weight to perform a weighted summation on the first value vector to obtain a second cross-modal attention feature; adding the first cross-modal attention feature and the first denoising feature element-wise to obtain a first enhanced coding feature; and adding the second cross-modal attention feature and the second denoising feature element-wise to obtain a second enhanced coding feature.
[0071] In this embodiment, to further explore the complementary information between the visible light mode and the depth mode, after obtaining the first and second denoising features, a cross-modal attention interaction step is performed to generate the first and second enhanced coding features. This achieves information flow and fusion between modalities through an attention mechanism, enabling features from each modality to focus on key information in another modality, thereby improving the expressive power and robustness of the features.
[0072] Specifically, the first denoised feature (i.e., the visible light feature after guided denoising) undergoes a linear projection transformation, mapping it to a query, key, and value space, resulting in a first query vector, a first key vector, and a first value vector, respectively. Simultaneously, the second denoised feature (i.e., the depth feature after guided denoising) undergoes the same linear transformation, yielding a second query vector, a second key vector, and a second value vector. This transformation converts the features into a form suitable for calculating attention weights.
[0073] Then, cross-modal attention computation is performed. On one hand, the similarity between the first query vector and the second key vector is calculated (e.g., through a dot product operation) to measure the correlation between query information in visible light features and key information in depth features. This similarity is normalized (e.g., using a softmax function) to obtain the first attention weight. The second value vector is then weighted and summed using this first attention weight to extract the depth semantic information related to the visible light features, resulting in the first cross-modal attention feature. Essentially, this process allows the visible light features to pay attention to and acquire useful information from the depth features.
[0074] On the other hand, a symmetrical operation is performed. The similarity between the second query vector and the first key vector is calculated to obtain a second similarity, which is then normalized to obtain a second attention weight. The first value vector is weighted and summed using the second attention weight to extract visible light semantic information related to the depth features, resulting in a second cross-modal attention feature. This process allows the depth features to pay attention to and acquire useful information from the visible light features. Through this bidirectional cross-modal attention interaction, the features of the two modalities achieve depth information complementarity.
[0075] Finally, to preserve the information of the original features and fuse the interacting information, the first cross-modal attention feature and the first denoising feature are added element-wise to obtain the first enhanced coding feature; similarly, the second cross-modal attention feature and the second denoising feature are added element-wise to obtain the second enhanced coding feature. This residual connection method not only preserves the basic features of each modality but also injects complementary information from another modality, thus enabling the enhanced coding features to retain their respective advantages while possessing stronger cross-modal correlation and anti-interference capabilities.
[0076] The above steps achieve bidirectional information complementarity and fusion of visible light and depth features through a cross-modal attention mechanism. While preserving the basic information of each modality, the semantic relevance and discriminative power of the features are enhanced, thereby improving the model's robustness to perception of complex scenes.
[0077] As some optional embodiments of this application, the first enhanced coding feature and the second enhanced coding feature are subjected to global pooling processing respectively, and the processing results are concatenated to obtain fused features. This can be achieved through the following steps: performing global average pooling processing on the first enhanced coding feature to obtain a first global feature; performing global average pooling processing on the second enhanced coding feature to obtain a second global feature; concatenating the first global feature and the second global feature to obtain a concatenated global feature; inputting the concatenated global feature into a fully connected layer, and performing activation processing on the output of the fully connected layer to obtain an adaptive channel vector; and performing a weighted multiplication of the adaptive channel vector with the first enhanced coding feature. The weighted multiplication result is added element-wise to the first enhanced coding feature to obtain the first improved feature; the adaptive channel vector is multiplied with the second enhanced coding feature using weights, and the result is added element-wise to the second enhanced coding feature to obtain the second improved feature; the first and second improved features are activated respectively to obtain the first fusion weight and the second fusion weight; the first improved feature is enhanced using the first fusion weight to obtain the first weighted enhanced feature; the second improved feature is enhanced using the second fusion weight to obtain the second weighted enhanced feature; the first and second weighted enhanced features are concatenated to obtain the fused feature.
[0078] In this embodiment, to further reduce the model's over-reliance on single-modal features, improve the robustness of dual-modal fusion, and increase functional safety redundancy, a two-stream interactive fusion operation is performed on the first and second enhanced coding features. Adaptive channel weights are generated by extracting global context information, thereby dynamically adjusting the importance of each channel and achieving robust fusion with fault tolerance capabilities.
[0079] Specifically, the first enhanced encoded feature is first subjected to global average pooling to extract its global contextual information, resulting in the first global feature. Simultaneously, the second enhanced encoded feature is subjected to the same global average pooling process to obtain the second global feature. Global average pooling effectively compresses spatial dimensions while retaining the most representative semantic information. The first and second global features are then concatenated to integrate the bimodal global semantics, resulting in the concatenated global feature.
[0080] Then, the concatenated global features are input into a fully connected layer for processing to learn the global dependencies between modalities. The output of the fully connected layer is processed by an activation function (such as ReLU) to generate an adaptive channel vector. This channel vector contains channel-level attention information after fusing global information from both modalities, reflecting the importance distribution of each feature channel in the current scene.
[0081] The original enhanced features are reweighted using this adaptive channel vector. Specifically, the adaptive channel vector is multiplied with the first enhanced coding feature in a weighted manner to highlight important channels and suppress secondary channels, resulting in the first weighted feature. The first weighted feature is then added element-wise to the original first enhanced coding feature to obtain the first improved feature. Similarly, the adaptive channel vector is multiplied with the second enhanced coding feature in a weighted manner to obtain the second weighted feature, which is then added element-wise to the original second enhanced coding feature to obtain the second improved feature. This residual connection method introduces adaptive weight adjustment while preserving the basic information of the original features.
[0082] Finally, to generate the final fusion weights, the first and second improved features are activated (e.g., using the sigmoid function) to obtain the first and second fusion weights. These weights reflect the contribution of each modality feature to the fusion process.
[0083] The first improved feature is weighted and enhanced using a first fusion weight to obtain a first weighted enhanced feature; the second improved feature is weighted and enhanced using a second fusion weight to obtain a second weighted enhanced feature. The first weighted enhanced feature and the second weighted enhanced feature are then concatenated to obtain the final fused feature.
[0084] Through the above steps, the model can automatically adjust its dependence on RGB and depth modes according to the input scenario. When a certain mode fails or is noisy, the features of another mode can be more fully utilized through adaptive weights, thereby significantly improving the robustness and functional safety level of the model under complex road conditions.
[0085] In some optional embodiments of this application, the decoder includes cascaded multi-level decoding units, each built based on a self-attention mechanism; the fused features are decoded layer by layer to obtain the final-level decoded features, and the final-level decoded features and the visible light image are subjected to edge refinement processing to obtain edge refinement features. This can be achieved through the following steps: inputting the fused features into the decoder, performing decoding processing layer by layer through the multi-level decoding units and restoring the spatial resolution of the features layer by layer to obtain multi-level decoded features; taking the last level of the multi-level decoded features as the final-level decoded features; performing feature transformation on the final-level decoded features and performing feature transformation on the visible light image so that the feature dimensions of the transformed final-level decoded features and the transformed visible light image are the same; stitching the transformed final-level decoded features and the transformed visible light image together to obtain stitched features; performing multi-scale attention processing on the stitched features to obtain multi-scale edge interaction features; performing feature transformation and multi-scale attention processing on the multi-scale edge interaction features and the visible light image to obtain optimized edge interaction features; upsampling the optimized edge interaction features so that the size of the optimized edge interaction features is the same as the size of the visible light image to obtain edge refinement features.
[0086] Preferably, the spatial resolution of the features is restored step by step by multi-level decoding units through the following steps: the spatial resolution of the decoded features output by each level of decoding unit is twice that of the decoded features output by the level above it.
[0087] Furthermore, feature transformation of the final-level decoded features can be achieved through the following steps: sequentially performing convolution processing, batch normalization processing, and activation processing on the final-level decoded features.
[0088] Furthermore, feature transformation of visible light images can be achieved through the following steps: sequentially performing convolution processing, batch normalization processing, and activation processing on the visible light image.
[0089] Furthermore, feature transformation and multi-scale attention processing of multi-scale edge interaction features and visible light images can be achieved through the following steps: performing convolution, batch normalization, and activation processing on the multi-scale edge interaction features sequentially; performing convolution, batch normalization, and activation processing on the visible light image sequentially; concatenating the processed multi-scale edge interaction features and the processed visible light image along the channel dimension, and performing multi-scale attention processing on the concatenation result to obtain optimized edge interaction features.
[0090] In this embodiment, to restore the spatial resolution of features and accurately restore the edge details of onboard targets, a decoder based on a self-attention mechanism is used to decode the fused features step by step, and edge refinement processing is performed in conjunction with the detailed information of the visible light image. This compensates for the ambiguity of high-level semantic features in spatial positioning, ensures the clear restoration of salient target boundaries (such as pedestrian limbs, road markings, etc.), thereby improving perception and positioning accuracy and avoiding misjudgments in autonomous driving decisions caused by blurred edge recognition.
[0091] Specifically, the fused features are input into the decoder, which consists of cascaded multi-level decoding units, each containing a structure based on a self-attention mechanism. The fused features are sequentially decoded through each level of the decoding unit. During this process, the spatial resolution of the features is gradually restored, while the semantic information is further abstracted, generating multi-level decoded features. The feature output after processing by the last level of the decoding unit is defined as the final-level decoded feature, which has the highest level of semantic information but a relatively low spatial resolution.
[0092] To effectively fuse high-level semantic information with low-level detail information, feature transformations are first performed on the final-level decoded features and the original visible light image. The feature transformation can consist of convolutional layers, batch normalization layers, and activation functions, aiming to map the final-level decoded features and the visible light image to the same feature dimension space, facilitating subsequent feature interactions. The transformed final-level decoded features (containing global semantics) are then concatenated with the transformed visible light image (containing rich texture and edge details) to obtain the concatenated features.
[0093] Then, multi-scale attention processing (e.g., using an efficient multi-scale attention EMA mechanism) is performed on the stitched features to extract the edge interaction information of the vehicle target at different scales. This process captures subtle changes in the target edges at different resolutions, generating multi-scale edge interaction features. To further optimize the edge information, the multi-scale edge interaction features are combined with the original visible light image again through feature transformation and multi-scale attention processing to obtain optimized edge interaction features. This dual refinement mechanism ensures the accuracy and completeness of the edge information.
[0094] Finally, the optimized edge interaction features are upsampled (e.g., by bilinear interpolation) to restore their spatial dimensions to the same size as the original visible light image. The resulting edge refinement features contain high-precision information about salient target boundaries and can be directly used to generate the final salient target detection map.
[0095] By employing the above methods, the model not only preserves the accuracy of high-level semantics but also makes full use of the original details of visible light images, effectively reducing the security risks caused by positioning errors.
[0096] Optionally, the visible light image and depth image are images synchronously acquired in a vehicle-mounted scene using an in-vehicle image acquisition device; the decoder in the cross-modal reverse guidance model is used to guide denoising of the first and second encoded features respectively to obtain the first and second denoised features, specifically including the following steps:
[0097] The decoding features output by the previous decoding unit in the decoder are obtained. These decoding features are used to characterize the location information of salient targets in the vehicle scene.
[0098] Based on the location information represented by the decoding features, the feature responses located outside the significant target location in the first coding features are suppressed to obtain the first noise reduction feature, so that the non-significant area responses caused by rain, fog, backlight or low illumination in the first noise reduction feature are weaker than the responses of the significant target area.
[0099] Based on the location information represented by the decoding features, the feature responses located outside the salient target location in the second coding features are suppressed to obtain the second noise reduction features, so that the non-salient region responses caused by holes in the depth image or measurement noise in the second noise reduction features are weaker than the responses in the salient target region.
[0100] Cross-modal attention interaction is performed on the first and second noise reduction features to obtain the first and second enhanced coding features, specifically including the following steps:
[0101] Using the first denoising feature as the query source and the second denoising feature as the query source, cross-modal attention calculation is performed to obtain the first enhanced coding feature, which makes the geometric information of the salient target region in the depth image compensate for the lack of salient target features in the visible light image under rain, fog, backlight or low illumination scenes.
[0102] Using the second noise reduction feature as the query source and the first noise reduction feature as the query source, cross-modal attention calculation is performed to obtain the second enhanced coding feature, which enables the color and texture information of the salient target region in the visible light image to compensate for the loss of salient target features in the depth image caused by holes or measurement noise.
[0103] Furthermore, in the first noise reduction feature, the non-significant region response caused by rain, fog, backlight, or low illumination does not propagate to the second enhanced coding feature through cross-modal attention calculation, and in the second noise reduction feature, the non-significant region response caused by holes or measurement noise does not propagate to the first enhanced coding feature through cross-modal attention calculation.
[0104] In this embodiment, visible light images and depth images are simultaneously acquired in a vehicle-mounted scene using an onboard image acquisition device to construct an input data source that meets functional safety requirements. The decoder in the cross-modal reverse guidance model is used to guide noise reduction on the first coded feature (visible light coded feature) and the second coded feature (depth coded feature), aiming to solve the feature failure problem of a single modality under complex road conditions. The specific process is as follows:
[0105] First, the decoded features output by the previous decoding unit in the decoder are obtained. These features contain high-level semantic and positional prior information about salient targets in the vehicle scene. Using this positional information as a guiding signal, suppression processing is applied to regions outside the salient target locations in the first encoded features. This involves reducing the feature response intensity of non-salient regions, thus obtaining the first denoised feature. This process effectively reduces noise interference in visible light images caused by environmental factors such as rain, fog, backlight, or low illumination, resulting in a stronger response in salient target regions than in non-salient regions in the first denoised feature, thereby improving the purity of the visible light features.
[0106] Similarly, the decoded features are used to guide denoising of the second encoded features. By suppressing the responses of non-salient target regions in the second encoded features, the second denoised features are obtained. This step effectively reduces the abnormal responses of non-salient regions caused by holes in the depth image, measurement noise, or sensor errors, enabling the second denoised features to more accurately reflect the true distribution of depth information and providing high-quality basic features for subsequent feature interaction.
[0107] After obtaining the first and second denoising features, cross-modal attention interaction is performed to generate the first and second enhanced coding features, achieving complementary enhancement between modalities. Specifically, the first denoising feature is used as the query source, and the second denoising feature is used as the key and value for cross-modal attention calculation to obtain the first enhanced coding feature. This process enables significant geometric structure information in the depth image to compensate for the loss of color and texture features in the visible light image due to insufficient lighting or occlusion in rain, fog, backlight, or low-light scenes, thereby enhancing the perception ability of the visible light branch for salient targets.
[0108] On the other hand, using the second denoising feature as the query source and the first denoising feature as the key and value, cross-modal attention computation is performed to obtain the second enhanced coding feature. This process enables the rich color, texture, and edge detail information in the visible light image to compensate for the loss of geometric information in the depth image caused by holes, measurement noise, or long-distance attenuation, thereby enhancing the depth branch's ability to identify salient targets.
[0109] Furthermore, the key to this embodiment lies in establishing an isolation mechanism for feature propagation. Since insignificant region responses caused by rain, fog, backlight, or low illumination have been suppressed during the first denoising feature generation process, noise features from these insignificant regions will not effectively propagate to the second enhanced coding features through attention weights in subsequent cross-modal attention calculations. Similarly, since insignificant region responses caused by depth image holes or measurement noise have been suppressed during the second denoising feature generation process, these depth noises will not propagate to the first enhanced coding features. This bidirectional isolation mechanism ensures that cross-modal interactions only transmit effective complementary information, while blocking cross-modal noise contamination, thereby significantly improving the model's robustness and perception accuracy in complex scenarios of functional safety real-vehicle testing.
[0110] This application also provides a target recognition method, specifically including the following steps: acquiring an image to be recognized; and performing target recognition on the image to be recognized using a cross-modal back-guiding model, wherein the cross-modal back-guiding model is a method that... Figure 1 The model was trained using the method shown.
[0111] Understandably, the target recognition method aims to apply the aforementioned trained cross-modal backpropagation model to real-world vehicle environment perception scenarios to achieve highly reliable salient target detection. The execution flow of this method mainly includes two core stages: data acquisition and model inference.
[0112] First, the process of acquiring the image to be identified is performed. In practical applications, the vehicle-mounted perception system collects data on the current driving scene in real time through multimodal sensors (such as RGB-D cameras and LiDAR) deployed on the vehicle. Specifically, the system simultaneously acquires visible light images (RGB images) and depth images of the scene to be identified. This image data constitutes the input source for target recognition, reflecting the real-time state of the vehicle's surrounding environment, including road conditions, obstacle locations, pedestrian dynamics, and lighting conditions. The acquired image data undergoes standardization processing, such as size normalization, noise reduction, and format conversion, to ensure that the data meets the input requirements of the cross-modal reverse guidance model.
[0113] Then, a pre-trained cross-modal inverse guidance model is used to perform target recognition on the image to be recognized. This cross-modal inverse guidance model is trained according to the model training method described in the foregoing embodiments of this application, and it includes core modules such as dual-stream coding, multi-level feature inverse guidance enhancement, dual-stream interactive fusion, decoding, and edge refinement. During inference, the preprocessed visible light image and depth image are first input into the model's dual-stream encoder to extract multi-stage encoded features. The model uses the high-level semantic features output by the decoder to perform inverse guidance noise reduction on the low-level encoded features, and achieves complementarity and enhancement of dual-modal features through a cross-modal attention mechanism. After step-by-step decoding and edge refinement, the model finally outputs a high-precision salient target detection map.
[0114] The advantage of the aforementioned target recognition method lies in its strong robustness and generalization ability, as the model has undergone closed-loop iterative optimization using feedback data from functional safety vehicle testing during the training phase. Even in complex or single-modal failure extreme scenarios such as backlighting, rain, and tunnels, the model can still accurately identify key targets such as pedestrians, vehicles, and road markings through complementary and redundant design of cross-modal information, and precisely restore the edge details of the targets. This not only meets the stringent requirements of perception accuracy and real-time performance in autonomous driving functional safety vehicle testing, but also provides reliable environmental perception data for upper-level autonomous driving decision-making systems, thereby effectively reducing the risk of driving safety accidents caused by perception failure.
[0115] The aforementioned target recognition method is specifically designed for real-world testing scenarios of autonomous driving functional safety. It enhances the functional safety level of the environmental perception module from the perception layer. Verified through real-world road testing, under complex road conditions, the model's mean absolute error is as low as 0.019, with a maximum F-measure of 0.956. The false negative rate for core targets such as pedestrians and obstacles is less than 1%, and the false positive rate is less than 0.8%, fully meeting the stringent requirements for perception accuracy in real-world functional safety testing. Even in single-modal failure scenarios such as backlighting, rain, and tunnels, the model can still achieve stable detection, with core indicator errors controlled within 1.1%, significantly improving the functional safety redundancy of the perception module and matching the fault tolerance requirements of functional safety.
[0116] Furthermore, the model's edge recognition error rate is reduced to below 0.8%, and fine-grained edge features such as road markings and non-motorized vehicle outlines are accurately located, avoiding decision-making errors at the perception level and strengthening the first line of defense for functional safety. The network parameters of the aforementioned target recognition method are only 109.4M, with an inference frame rate of 23.7FPS. After quantization, it can be deployed on mainstream automotive embedded chips. The transmission and inference latency meet the time requirements of functional safety vehicle testing, with no real-time risks, and complies with the hardware deployment specifications for automotive functional safety. Combined with feedback from functional safety vehicle testing, a closed-loop iteration is achieved, with an average detection accuracy of 95.2% across multiple scenarios. Its generalization capability adapts to the diverse scenarios of functional safety vehicle testing, providing real and reliable test data for functional safety certification. Overall, it provides an accurate, reliable, real-time, and robust environmental perception solution for autonomous driving functional safety vehicle testing, improving functional safety levels from the source of the perception module, effectively reducing driving safety accidents caused by perception failure, and providing core technical support for the functional safety testing and certification of autonomous vehicles.
[0117] Figure 2 This is a schematic diagram of a model training method according to an embodiment of this application. The following is in conjunction with... Figure 2 This section describes the implementation steps of another model training method.
[0118] Step S1: Vehicle-mounted multimodal data acquisition and preprocessing, conforming to functional safety data acquisition specifications;
[0119] Step S2: Dual-stream encoding feature extraction, constructing a dual-stream encoder based on weight-sharing Swin-Transformer;
[0120] Step S3: Multi-level feature reverse guidance enhancement, which achieves feature denoising and enhancement through the core module;
[0121] Step S4: Dual-stream interaction fusion reduces dual-modal dependency and achieves robust fusion;
[0122] Step S5: Fuse feature decoding to restore feature space resolution;
[0123] Step S6: Edge refinement perception, complete the edge details of the on-board target;
[0124] Step S7 involves functional safety real-vehicle test feedback and model iterative optimization to form a closed-loop optimization mechanism.
[0125] To ensure data quality from the source and meet the stringent functional safety requirements for sensed data, step S1 specifically includes:
[0126] Step S11: Using an RGB-D camera vehicle perception system, road condition RGB images and depth images are simultaneously acquired during functional safety vehicle testing.
[0127] Step S12: Perform size normalization processing on the acquired RGB image and depth image, and uniformly adjust them to the standard size of 224×224;
[0128] Step S13: Use random flipping and rotation data augmentation on the training set data to improve the model's adaptability to diverse road conditions in functional safety real vehicle testing.
[0129] Step S14: Filter out invalid data caused by vehicle bumps, sensor obstruction, etc., and finally obtain standard feature data that meets the network input requirements.
[0130] To efficiently extract dual-modal basic features while also considering the lightweight requirements of automotive applications, step S2 specifically includes:
[0131] Step S21: Construct a dual-stream encoder with a weight-sharing Swin-Transformer as the backbone, which is divided into an RGB encoding branch and a depth encoding branch;
[0132] Step S22: Input the standard feature data obtained in step S1 into the two coding branches respectively;
[0133] Step S23: The two branches simultaneously extract the encoding features of the four stages, with the RGB encoding branch outputting the RGB encoding features. The deep coding branch outputs deep coding features. ;
[0134] Step S24: Utilize the window self-attention mechanism of Swin-Transformer to capture global contextual information of the vehicle scene, providing high-quality basic features for subsequent feature processing.
[0135] To reduce noise interference in vehicle-mounted scenarios and enhance feature representation capabilities to improve perception accuracy, step S3 is specifically as follows: Figure 3 As shown, this is achieved through a multi-level feature reverse guidance enhancement module. Step S3 specifically includes:
[0136] Step S31: Integrate the multi-level feature reverse guidance enhancement module into each layer of the encoder. This module decodes features at a higher level of the decoder. and encoder low-level coding features , For input;
[0137] Step S32, decode the high-level features Perform reshaping operation Adjust it to the same size as the low-level encoded features;
[0138] Step S33, activate the Sigmoid function. Generate adaptive weight vectors to encode low-level RGB features. and deep encoding features Weighted noise reduction is performed to quickly locate salient target regions. The calculation process satisfies the following:
[0139]
[0140] in To sum element by element, For element-wise multiplication, , This refers to the post-guided noise reduction characteristics;
[0141] Step S34, the post-guidance features , These are respectively converted into query (Q), key (K), and value (V) vectors, i.e. correspond , , , correspond , , ;
[0142] Step S35: Information interaction and complementarity of bimodal features are achieved through a cross-attention mechanism, and the calculation process satisfies:
[0143]
[0144] in , For cross-attention enhancement features;
[0145] Step S36: Element-wise sum the cross-attention enhancement features and the guided denoising features to obtain the final enhanced encoded features. The calculation process satisfies:
[0146]
[0147] in , To enhance the output features.
[0148] To improve the robustness of dual-modal fusion and reduce unimodal dependency to match functional safety fault tolerance requirements, step S4 is specifically as follows: Figure 4 As shown, this is achieved through a dual-stream interactive fusion module. Step S4 specifically includes:
[0149] Step S41, Enhanced features output in step S3 , Perform average pooling operations separately Extract global contextual information of bimodal features;
[0150] Step S42: Concatenate the extracted dual-modal global information into features. Integrate bimodal global semantics;
[0151] Step S43: Input the spliced global information into the fully connected layer. After ReLU activation function The process involves processing and generating adaptive channel vectors, with the following calculation steps:
[0152]
[0153] Step S44, generate the channel vector , Compared with the original enhanced features , By performing weighted multiplication and summing element by element, we can obtain more prominent improved features in the significant regions.
[0154] Step S45: Generate fusion weights for the improved features using the Sigmoid activation function, and then perform weighted enhancement on the improved features;
[0155] Step S46: Concatenate the weighted and enhanced dual-modal features to obtain the final fused features. .
[0156] To restore the feature space resolution and provide accurate semantic support for subsequent edge refinement, step S5 specifically includes:
[0157] Step S51: Construct a decoder based on the Swin-Transformer decoder block, with each layer of the decoder containing 2 Swin-Transformer decoder blocks;
[0158] Step S52, fused features obtained in step S4 The input is processed by the decoder, and the decoding process is performed layer by layer.
[0159] Step S53: During the decoding process, the spatial resolution of the features is gradually restored to generate multi-stage decoded features. ;
[0160] Step S54: Output the decoding features of the last layer of the decoder. This provides accurate semantic feature support for subsequent edge refinement.
[0161] To accurately complete the edge details of the onboard target and improve positioning accuracy to avoid misjudgment, step S6 is as follows: Figure 5 As shown, this is achieved through the edge refinement perception module. Step S6 specifically includes:
[0162] Step S61, process the original RGB image ( ) and decoder final features Perform Bconv operations (a combination of 3×3 convolution + BN + ReLU) on both to map them to the same feature dimension;
[0163] Step S62: The mapped features are concatenated to integrate semantic and detailed information;
[0164] Step S63: Input the stitched features into an efficient multi-scale attention (EMA) mechanism to extract the interaction information of the vehicle target edge at different scales. The calculation process satisfies:
[0165]
[0166] in For edge interaction features;
[0167] Step S64, edge interaction features The Bconv and EMA operations are performed again on the original RGB image to further optimize the edge information;
[0168] Step S65: Through upsampling, the optimized features are restored to the original image size acquired in the functional safety vehicle test, generating a high-precision road condition salient target detection map.
[0169] To establish a functional safety closed-loop optimization mechanism and improve the model's generalization ability to support real-vehicle certification, step S7 specifically includes:
[0170] Step S71: The salient target detection map generated in step S6 is output to the autonomous driving decision system in real time to complete the road condition target recognition, warning and decision-making in the functional safety vehicle test;
[0171] Step S72: Collect functional safety-related test data through the functional safety vehicle test feedback module, including detection errors of low-quality modes under complex road conditions, missed detection and false detection data of core targets, edge recognition error data, perception accuracy data under different road conditions, etc.
[0172] Step S73: Based on the collected feedback data, the Adam optimizer is used to iteratively optimize the parameters of the cross-modal backpropagation network;
[0173] Step S74: Set the initial learning rate for iterative optimization to 0.0001, and decrease the learning rate by a factor of 5 every 40 training rounds, for a total of 100 training rounds;
[0174] Step S75: After iterative optimization, output the optimized model adapted to the functional safety real vehicle test scenario, forming a closed-loop mechanism of detection-feedback-optimization.
[0175] This application also provides a cross-modal reverse guidance detection method for real-vehicle testing of functional safety on urban roads, which specifically includes the following steps.
[0176] Step S1: Test environment and equipment configuration, establishing a real vehicle test foundation that meets functional safety requirements;
[0177] Step S2: Multimodal data acquisition and standardized preprocessing to ensure the functional safety and compliance of the input data;
[0178] Step S3: Cross-modal reverse guidance for network deployment and operation, executing full-process detection logic;
[0179] Step S4: Record test data and iteratively optimize the model to verify the improvement of functional safety performance;
[0180] Step S5: Optimized performance verification to confirm compliance with functional safety real-vehicle testing standards.
[0181] Furthermore, step S1 specifically includes:
[0182] Step S11: The test vehicle is an L2 level autonomous driving test vehicle. The vehicle's onboard computing platform has passed ISO26262 functional safety standard certification and has the hardware foundation to adapt to functional safety real vehicle testing.
[0183] Step S12: The vehicle computing unit is equipped with an NVIDIA RTX A5000 embedded GPU. This GPU is optimized for vehicle scenarios, supports high concurrency and low latency model inference, and meets the functional safety requirements for computing stability.
[0184] In step S13, the RGB-D camera of the perception system is responsible for simultaneously acquiring RGB images and initial depth data, forming a redundant perception link;
[0185] Step S2 specifically includes:
[0186] Step S21: During the functional safety vehicle test, road condition data is collected synchronously through the vehicle perception system. The RGB image resolution is set to 1920×1080, the frame rate is 30FPS, and the effective measurement range of the depth image is 0.5m-100m, which meets the perception distance requirements of urban road scenarios.
[0187] Step S22: The LiDAR calibrates and completes the initial depth data collected by the RGB-D camera in real time. To address the problem of missing depth data in low-light scenarios such as rainy tunnels, the LiDAR point cloud interpolation algorithm fills in the depth gaps, ensuring the integrity of the depth data and reducing the risk of detection failure due to data distortion from the source.
[0188] In step S23, the data preprocessing module receives the acquired RGB image and the completed depth image, and first performs size normalization processing to uniformly adjust it to a standard size of 224×224 to adapt to the input requirements of the cross-modal backpropagation network.
[0189] Step S24: Data augmentation strategies are applied to the dataset during the training phase, including random horizontal flipping and random rotation, to simulate the posture changes of the vehicle during driving and improve the model's adaptability to different shooting angles.
[0190] Step S25: Set data filtering rules to automatically remove blurry images caused by severe vehicle bumps and invalid data caused by sensor lens obstruction, and retain only clear and effective multimodal data for model input to ensure the quality of input features;
[0191] Step S26: After the above processing, standard RGB feature and depth feature data that meet the input requirements of the cross-modal back-guided network are output, and the pixel values are normalized to the [0,1] interval.
[0192] Step S3 specifically includes:
[0193] Step S31: The standard feature data output in step S2 is transmitted to the cross-modal reverse guidance network via the vehicle Ethernet. The transmission process adopts a verification mechanism to ensure that the data transmission is free of loss and error.
[0194] In step S32, the feature data is first input into the dual-stream encoder. RGB features enter the RGB encoding branch, and deep features enter the deep encoding branch. Based on the weight-sharing Swing-Transformer backbone network, the encoded features of the four stages are extracted simultaneously. The number of feature channels is gradually increased to achieve feature extraction from the underlying texture to the high-level semantics.
[0195] Step S33: The multi-level feature backpropagation enhancement (MFGE) module corresponding to the input of the encoded features of each stage uses the decoded features output from the higher layers of the decoder for backpropagation: taking the encoded features of the third stage as an example, the higher-level decoded features of the fourth stage of the decoder are received. After being reshaped to the same size as the third-stage encoded features, an adaptive weight vector is generated through Sigmoid activation to apply the third-stage RGB encoded features. With deep encoding features Weighted noise reduction is performed, and then cross-attention mechanism is used to realize information interaction of dual-modal features, outputting enhanced encoded features. and ;
[0196] Step S34: The enhanced dual-modal encoded features are input into the dual-stream interactive fusion (DSIF) module. First, global context information is extracted through average pooling. Then, adaptive channel vectors are generated by feature concatenation, fully connected layers, and ReLU activation. After weighted enhancement of the original enhanced features, fusion weights are generated through Sigmoid. Finally, the fused features are concatenated to obtain the fused features. ;
[0197] Step S35: The fused features are input into the decoder in stages. Each decoder layer contains two Swin-Transformer decoder blocks. The feature space resolution is restored by upsampling layer by layer, generating multi-stage decoded features. to , where represents the high-level semantic features output from the last layer of the decoder;
[0198] Step S36, the decoder finally outputs features The image is input into the Edge Refinement Awareness (ERA) module along with the original RGB image. First, the two images are mapped to a 512-dimensional feature space through a combination of 3×3 convolution + BN + ReLU (Bconv). After feature concatenation, the image is input into the Efficient Multi-Scale Attention (EMA) mechanism to extract edge interaction information at different scales. After optimization by Bconv and EMA operations, the image is upsampled to 1920×1080 resolution through bilinear interpolation to generate the final salient object detection map.
[0199] Step S4 specifically includes:
[0200] Step S41: In the initial testing phase (before real vehicle feedback optimization), record the model detection performance data under different complex scenarios;
[0201] Step S42: Collect functional safety-related feedback data during the initial testing phase, including details of detection errors in each scenario, types and scenario characteristics of missed / false detection targets, and specific location information of edge recognition errors, and establish a feedback database.
[0202] Step S43: Based on the data in the feedback database, the Adam optimizer is used to iteratively optimize the parameters of the cross-modal back-guided network. The initial learning rate is set to 0.0001, and the learning rate decays by 5 times every 40 training rounds. The total number of training rounds is 100. During the optimization process, the focus is on strengthening the training of samples in complex scenarios such as rainy tunnels and backlit intersections.
[0203] Step S44: During the iterative optimization process, intermediate verification is performed every 20 rounds to evaluate the model's core metrics such as MAE, F-measure, E-measure, and S-measure on the validation set, and the optimization direction is dynamically adjusted to ensure that the model iterates in a direction that meets functional safety requirements.
[0204] Step S5 specifically includes:
[0205] Step S51: After completing 100 rounds of iterative optimization, the performance was retested in the same urban road functional safety real vehicle test scenario, and the core indicators after optimization were recorded: the detection error in the rainy tunnel scenario was reduced to 1.1%, the pedestrian missed detection rate at backlit intersections was reduced to 0.5%, and the electric vehicle contour edge recognition error rate was reduced to 0.8%. All indicators met the stringent requirements of functional safety real vehicle testing.
[0206] Step S52, comprehensively evaluate the overall performance of the model: the mean absolute error (MAE) reached 0.019, the maximum F-measure reached 0.956, the maximum E-measure reached 0.975, and the S-measure reached 0.948. All indicators are better than the initial testing stage.
[0207] Step S53, verify the real-time performance of the model: On the NVIDIA RTX A5000 embedded GPU, the model inference frame rate is stable at 23.7 FPS, which meets the real-time perception requirements of real vehicle testing on urban roads.
[0208] Step S54, final verification conclusion: After real vehicle testing and iterative optimization, the cross-modal reverse-guided RGB-D detection method of the present invention meets the functional safety real vehicle test requirements in terms of detection accuracy, robustness and real-time performance in complex urban road scenarios. The functional safety level of the perception module has been significantly improved, and it can provide reliable environmental perception basis for autonomous driving decision-making system.
[0209] Figure 6 This is a structural diagram of a model training device according to an embodiment of this application, such as... Figure 6 As shown, the device includes:
[0210] The acquisition module 61 is used to acquire training data, which includes visible light images and depth images of the road where the vehicle is located.
[0211] Extraction module 62 is used to extract coded features from visible light images and depth images respectively using the encoder in the cross-modal reverse guidance model to obtain first coded features and second coded features.
[0212] The noise reduction module 63 is used to guide the first coded feature and the second coded feature to perform noise reduction using the decoder in the cross-modal reverse guidance model, respectively, to obtain the first noise reduction feature and the second noise reduction feature; and to perform cross-modal attention interaction on the first noise reduction feature and the second noise reduction feature to obtain the first enhanced coded feature and the second enhanced coded feature.
[0213] The processing module 64 is used to perform global pooling processing on the first enhanced coding feature and the second enhanced coding feature respectively, and to concatenate the processing results to obtain the fused feature; to perform layer-by-layer decoding on the fused feature to obtain the final-level decoding feature, and to perform edge thinning processing on the final-level decoding feature and the visible light image to obtain the edge thinning feature; and to upsample the edge thinning feature to the size of the visible light image to obtain the target detection image.
[0214] The receiving module 65 is used to send the target detection image to the autonomous driving decision system to instruct the autonomous driving decision system to conduct functional safety real vehicle testing based on the target detection image, and to receive feedback data related to functional safety testing sent by the autonomous driving decision system.
[0215] Training module 66. Used to iteratively optimize the parameters of the cross-modal back-guided model based on feedback data until a preset stopping condition is met, thus obtaining a fully trained cross-modal back-guided model.
[0216] It should be noted that the above Figure 6 The modules in can be program modules (e.g., a set of program instructions that implements a specific function) or hardware modules. For the latter, they can be represented in the following forms, but are not limited to these: each of the above modules is represented by a processor, or the functions of each of the above modules are implemented by a processor.
[0217] It should be noted that, Figure 6 Preferred embodiments of the shown examples can be found in [reference needed]. Figure 1 The relevant descriptions of the embodiments shown will not be repeated here.
[0218] Figure 7 A hardware block diagram of a computer terminal for implementing a model training method is shown. Figure 7 As shown, the computer terminal 70 may include one or more processors 702 (shown as 702a, 702b, ..., 702n in the figure) 702 (processor 702 may include, but is not limited to, a microprocessor MCU or a programmable logic device FPGA, etc.), a memory 704 for storing data, and a transmission module 706 for communication functions. In addition, it may also include: a display, an input / output interface (I / O interface), a universal serial bus (USB) port (which may be included as one of the ports of a BUS bus), a network interface, a power supply, and / or a camera. Those skilled in the art will understand that... Figure 7 The structure shown is for illustrative purposes only and does not limit the structure of the aforementioned electronic device. For example, the computer terminal 70 may also include... Figure 7 The more or fewer components shown, or having the same Figure 7 The different configurations shown.
[0219] It should be noted that the aforementioned one or more processors 702 and / or other data processing circuits are generally referred to herein as "data processing circuits". These data processing circuits may be embodied, in whole or in part, in software, hardware, firmware, or any other combination thereof. Furthermore, the data processing circuits may be a single, independent processing module, or may be integrated, in whole or in part, into any other element within the computer terminal 70. As involved in the embodiments of this application, the data processing circuits serve as processor control (e.g., selection of a variable resistor termination path connected to an interface).
[0220] The memory 704 can be used to store software programs and modules of application software, such as the program instructions / data storage device corresponding to the model training method in the embodiments of this application. The processor 702 executes various functional applications and data processing by running the software programs and modules stored in the memory 704, thereby realizing the above-mentioned model training method. The memory 704 may include high-speed random access memory, and may also include non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory 704 may further include memory remotely located relative to the processor 702, and these remote memories can be connected to the computer terminal 70 via a network. Examples of the above-mentioned networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof.
[0221] The transmission module 706 is used to receive or send data via a network. Specific examples of the network described above may include a wireless network provided by the communication provider of the computer terminal 70. In one example, the transmission module 706 includes a network interface controller (NIC), which can connect to other network devices via a base station to communicate with the Internet. In another example, the transmission module 706 may be a radio frequency (RF) module, used for wireless communication with the Internet.
[0222] The display may be, for example, a touchscreen liquid crystal display (LCD) that allows the user to interact with the user interface of the computer terminal 70.
[0223] It should be noted here that, in some optional embodiments, the above... Figure 7 The computer terminal shown may include hardware components (including circuitry), software components (including computer code stored on a computer-readable medium), or a combination of both hardware and software components. It should be noted that... Figure 7This is only one instance of a specific particular instance, and is intended to illustrate the types of components that may exist in the aforementioned computer terminal.
[0224] It should be noted that, Figure 7 The computer terminal shown is used to execute Figure 1 The model training method shown above is also applicable to this electronic device, and will not be repeated here.
[0225] This application also provides a non-volatile storage medium, which includes a stored program, wherein the program, when running, controls the device where the storage medium is located to execute the above model training method.
[0226] A program on a non-volatile storage medium performs the following functions: acquires training data, including visible light and depth images of the road where the vehicle is located; extracts coded features from the visible light and depth images using the encoder in a cross-modal back-guided model, obtaining a first coded feature and a second coded feature; performs guided denoising on the first and second coded features using the decoder in the cross-modal back-guided model, obtaining a first denoised feature and a second denoised feature; performs cross-modal attention interaction on the first and second denoised features, obtaining a first enhanced coded feature and a second enhanced coded feature; and performs global pooling on the first and second enhanced coded features respectively. The system processes the data and stitches the results together to obtain fused features. It then decodes the fused features layer by layer to obtain final-level decoded features, and performs edge refinement processing on the final-level decoded features and the visible light image to obtain edge refinement features. The edge refinement features are upsampled to the size of the visible light image to obtain the target detection image. The target detection image is sent to the autonomous driving decision-making system to instruct it to conduct functional safety vehicle testing based on the target detection image, and the system receives feedback data related to functional safety testing from the autonomous driving decision-making system. Based on the feedback data, the parameters of the cross-modal reverse guidance model are iteratively optimized until a preset stopping condition is met, resulting in a fully trained cross-modal reverse guidance model.
[0227] This application also provides an electronic device, including: a memory and a processor, wherein the processor is used to run a program stored in the memory, wherein the program executes the above-described model training method when it runs.
[0228] The processor is used to run a program that performs the following functions: acquiring training data, wherein the training data includes visible light images and depth images of the road where the vehicle is located; extracting coded features from the visible light images and depth images respectively using the encoder in the cross-modal back-guided model to obtain first coded features and second coded features; performing guided denoising on the first coded features and second coded features respectively using the decoder in the cross-modal back-guided model to obtain first denoised features and second denoised features; performing cross-modal attention interaction on the first denoised features and second denoised features to obtain first enhanced coded features and second enhanced coded features; and performing global pooling on the first enhanced coded features and second enhanced coded features respectively. The system processes the data and stitches the results together to obtain fused features. It then decodes the fused features layer by layer to obtain final-level decoded features, and performs edge refinement processing on the final-level decoded features and the visible light image to obtain edge refinement features. The edge refinement features are upsampled to the size of the visible light image to obtain the target detection image. The target detection image is sent to the autonomous driving decision-making system to instruct it to conduct functional safety vehicle testing based on the target detection image, and the system receives feedback data related to functional safety testing from the autonomous driving decision-making system. Based on the feedback data, the parameters of the cross-modal reverse guidance model are iteratively optimized until a preset stopping condition is met, resulting in a fully trained cross-modal reverse guidance model.
[0229] The sequence numbers of the embodiments in this application are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments.
[0230] In the above embodiments of this application, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions of other embodiments.
[0231] In the above embodiments of this application, the information collected is information and data authorized by the user or fully authorized by all parties, and the collection, storage, use, processing, transmission, provision, disclosure and application of the relevant data all comply with relevant laws, regulations and standards, take necessary protective measures, do not violate public order and good morals, and provide corresponding operation entry points for users to choose to authorize or refuse.
[0232] In the several embodiments provided in this application, it should be understood that the disclosed technical content can be implemented in other ways. The device embodiments described above are merely illustrative; for example, the division of units can be a logical functional division, and in actual implementation, there may be other division methods. For instance, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the displayed or discussed mutual coupling, direct coupling, or communication connection may be through some interfaces; the indirect coupling or communication connection between units or modules may be electrical or other forms.
[0233] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0234] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0235] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to related technologies, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, read-only memory (ROM), random access memory (RAM), portable hard drives, magnetic disks, or optical disks.
[0236] The above description is only a preferred embodiment of this application. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the principle of this application, and these improvements and modifications should also be considered within the scope of protection of this application.
Claims
1. A model training method, characterized in that, include: Acquire training data, wherein the training data includes visible light images and depth images of the road where the vehicle is located; The encoder in the cross-modal back-guided model is used to extract coded features from the visible light image and the depth image respectively, to obtain the first coded feature and the second coded feature; The decoder in the cross-modal reverse guidance model is used to guide denoising on the first coded feature and the second coded feature respectively to obtain the first denoised feature and the second denoised feature; cross-modal attention interaction is performed on the first denoised feature and the second denoised feature to obtain the first enhanced coded feature and the second enhanced coded feature; The first enhanced coding feature and the second enhanced coding feature are respectively subjected to global pooling, and the processing results are concatenated to obtain a fused feature; the fused feature is decoded layer by layer to obtain a final-level decoded feature, and the final-level decoded feature and the visible light image are subjected to edge thinning processing to obtain an edge thinning feature; the edge thinning feature is upsampled to the size of the visible light image to obtain a target detection image; The target detection image is sent to the autonomous driving decision system to instruct the autonomous driving decision system to conduct functional safety real vehicle testing based on the target detection image, and to receive feedback data related to functional safety testing sent by the autonomous driving decision system. Based on the feedback data, the parameters of the cross-modal back-guiding model are iteratively optimized until a preset stopping condition is met, thus obtaining a cross-modal back-guiding model that has completed training.
2. The method according to claim 1, characterized in that, The encoder in the cross-modal back-guided model is used to extract coded features from the visible light image and the depth image respectively, resulting in first coded features and second coded features, including: The coding features of four stages are extracted by the visible light coding branch to obtain the first coding feature containing features of multiple stages; The coding features of the four stages are extracted by deep coding branches to obtain the second coding features containing features of multiple stages; The visible light coding branch and the deep coding branch share weights, and both the visible light coding branch and the deep coding branch are constructed based on a neural network with a self-attention mechanism. In the visible light coding branch and the depth coding branch, downsampling is performed between adjacent two stages through block merging operations, so that the feature resolution of the (i+1)th stage is half of the feature resolution of the i-th stage, and the number of feature channels in the (i+1)th stage is twice the number of feature channels in the i-th stage, where i is a positive integer not less than 1. The output features of the i-th stage of the visible light coding branch are used as part of the input of the (i+1)-th stage of the depth coding branch, and the output features of the i-th stage of the depth coding branch are used as part of the input of the (i+1)-th stage of the visible light coding branch, so that the visible light coding branch and the depth coding branch can perform cross-modal information interaction during the coding process.
3. The method according to claim 1, characterized in that, The decoder in the cross-modal reverse guidance model includes cascaded multi-level decoding units, which sequentially output multi-level decoding features, and the guidance noise reduction and the decoding are performed alternately at each level. The decoder in the cross-modal reverse guidance model is used to guide denoising on the first coded feature and the second coded feature respectively to obtain the first denoised feature and the second denoised feature, including: For the guided noise reduction of the current stage, the decoding features output by the previous stage decoding unit in the decoder are used as guided features, wherein the guided features are used to characterize the semantic information of the image; The guiding features are upsampled to obtain the adjusted guiding features; The adjusted guiding features are activated to obtain an adaptive weight vector; The adaptive weight vector is multiplied element-wise with the first encoded feature to obtain the first weighted feature, and the first weighted feature is added element-wise with the adjusted guiding feature to obtain the first noise reduction feature; The adaptive weight vector is multiplied element-wise by the second encoded feature to obtain the second weighted feature, and the second weighted feature is added element-wise by the adjusted guiding feature to obtain the second noise reduction feature.
4. The method according to claim 1, characterized in that, Cross-modal attention interaction is performed on the first denoising feature and the second denoising feature to obtain a first enhanced coding feature and a second enhanced coding feature, including: A linear transformation is performed on the first denoising feature to obtain a first query vector, a first key vector, and a first value vector; a linear transformation is performed on the second denoising feature to obtain a second query vector, a second key vector, and a second value vector. Calculate the first similarity between the first query vector and the second key vector; normalize the first similarity to obtain the first attention weight; The first cross-modal attention feature is obtained by weighting and summing the second value vector using the first attention weight; Calculate the second similarity between the second query vector and the first key vector; normalize the second similarity to obtain the second attention weight; The first value vector is weighted and summed using the second attention weights to obtain the second cross-modal attention feature; The first cross-modal attention feature is added element-wise to the first noise reduction feature to obtain the first enhanced coding feature; The second cross-modal attention feature is added element-wise to the second noise reduction feature to obtain the second enhanced coding feature.
5. The method according to claim 1, characterized in that, The first enhanced coding feature and the second enhanced coding feature are respectively subjected to global pooling, and the processing results are concatenated to obtain the fused feature, including: The first enhanced coding feature is subjected to global average pooling to obtain the first global feature; the second enhanced coding feature is subjected to global average pooling to obtain the second global feature. The first global feature and the second global feature are concatenated to obtain the concatenated global feature; The concatenated global features are input into a fully connected layer, and the output of the fully connected layer is activated to obtain an adaptive channel vector. The adaptive channel vector is multiplied by the first enhanced coding feature in a weighted manner, and the weighted multiplication result is added element by element to the first enhanced coding feature to obtain the first improved feature; The adaptive channel vector is multiplied by the second enhanced coding feature in a weighted manner, and the result of the weighted multiplication is added element by element to the second enhanced coding feature to obtain the second improved feature; The first improved feature and the second improved feature are activated respectively to obtain the first fusion weight and the second fusion weight; The first improved feature is weighted and enhanced using the first fusion weight to obtain a first weighted enhanced feature; the second improved feature is weighted and enhanced using the second fusion weight to obtain a second weighted enhanced feature. The first weighted enhancement feature and the second weighted enhancement feature are concatenated to obtain the fused feature.
6. The method according to claim 1, characterized in that, The decoder comprises cascaded multi-level decoding units, each of which is constructed based on a self-attention mechanism; The fused features are decoded layer by layer to obtain the final-level decoded features. The final-level decoded features and the visible light image are then subjected to edge thinning processing to obtain edge thinning features, including: The fused features are input into the decoder, and the multi-level decoding unit performs decoding processing step by step and restores the spatial resolution of the features step by step to obtain multi-level decoded features. The last level of the multi-level decoding features is taken as the final-level decoding feature; The final-level decoding features are transformed, and the visible light image is also transformed, so that the transformed final-level decoding features and the transformed visible light image have the same feature dimensions. The transformed final-level decoded features and the transformed visible light image are stitched together to obtain the stitched features; Multi-scale attention processing is applied to the spliced features to obtain multi-scale edge interaction features; The multi-scale edge interaction features and the visible light image are subjected to feature transformation and multi-scale attention processing to obtain optimized edge interaction features. The optimized edge interaction features are upsampled so that their size is the same as that of the visible light image, thus obtaining the edge thinning features.
7. The method according to claim 6, characterized in that, The multi-level decoding unit performs decoding processing step by step and restores the spatial resolution of features step by step, including: Decoding is performed step by step by the multi-level decoding units, so that the spatial resolution of the decoding feature output by each level of the decoding unit is twice the spatial resolution of the decoding feature output by the previous level of the decoding unit. The feature transformation of the final-level decoded features includes: sequentially performing convolution processing, batch normalization processing, and activation processing on the final-level decoded features; The feature transformation of the visible light image includes: sequentially performing convolution processing, batch normalization processing, and activation processing on the visible light image; Feature transformation and multi-scale attention processing are performed on the multi-scale edge interaction features and the visible light image, including: The multi-scale edge interaction features are sequentially subjected to convolution processing, batch normalization processing, and activation processing; The visible light image is sequentially subjected to convolution processing, batch normalization processing, and activation processing; The processed multi-scale edge interaction features are stitched together with the processed visible light image along the channel dimension, and the stitching result is subjected to multi-scale attention processing to obtain the optimized edge interaction features.
8. A target recognition method, characterized in that, include: Acquire the image to be recognized; The image to be identified is used to perform target recognition using a cross-modal back-guided model, wherein the cross-modal back-guided model is trained by the model training method described in any one of claims 1 to 7.
9. A model training device, characterized in that, include: An acquisition module is used to acquire training data, wherein the training data includes visible light images and depth images of the road where the vehicle is located; The extraction module is used to extract coded features from the visible light image and the depth image respectively using the encoder in the cross-modal back-guided model to obtain the first coded feature and the second coded feature; The noise reduction module is used to guide noise reduction on the first coded feature and the second coded feature respectively using the decoder in the cross-modal reverse guidance model to obtain a first noise reduction feature and a second noise reduction feature; and to perform cross-modal attention interaction on the first noise reduction feature and the second noise reduction feature to obtain a first enhanced coded feature and a second enhanced coded feature. The processing module is used to perform global pooling on the first enhanced coding feature and the second enhanced coding feature respectively, and to concatenate the processing results to obtain a fused feature; to perform layer-by-layer decoding on the fused feature to obtain a final-level decoded feature, and to perform edge thinning processing on the final-level decoded feature and the visible light image to obtain an edge thinning feature; and to upsample the edge thinning feature to the size of the visible light image to obtain a target detection image. The receiving module is used to send the target detection image to the autonomous driving decision system to instruct the autonomous driving decision system to conduct functional safety real vehicle testing based on the target detection image, and to receive feedback data related to functional safety testing sent by the autonomous driving decision system. The training module is used to iteratively optimize the parameters of the cross-modal back-guided model based on the feedback data until a preset stopping condition is met, thereby obtaining a cross-modal back-guided model that has completed training.
10. An electronic device, characterized in that, include: A memory and a processor, the processor being configured to run a program stored in the memory, wherein the program, when running, executes the model training method according to any one of claims 1 to 7.