Multi-modal target detection model training method and device, equipment and medium
By evaluating and correcting the feature reliability of the multimodal target detection model, the problem of unreliable features during the fusion of visible light and infrared images is solved, improving detection accuracy and robustness, and enhancing the ability to adapt to complex scenes.
Patent Information
- Application Number
- CN202511481782.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-16
- Publication Date
- 2025-12-16
Smart Images

Figure CN121147501A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of target detection technology, and specifically to a training method, apparatus, equipment, and medium for a multimodal target detection model. Background Technology
[0002] Object detection is a core task in computer vision, aiming to identify and locate targets in complex scenes. However, single-modal methods perform poorly in complex environments such as varying lighting conditions, inclement weather, or low visibility, failing to meet practical application requirements. To overcome this issue, multimodal information is introduced, combining the complementary advantages of different sensors to improve detection accuracy and robustness. Among these, visible light-infrared fusion, with its strong adaptability to various complex scenes, has become an important development direction in multimodal object detection. Visible light images provide high-resolution detail information under sufficient lighting, while infrared images maintain stable imaging in low-light or low-visibility environments by capturing stable thermal radiation signals. The fusion of these two technologies achieves complementary advantages, significantly enhancing detection accuracy and robustness, and has wide applications.
[0003] When fusing visible and infrared light, the same target may exhibit representational differences in the two modalities due to variations in imaging mechanisms and information distribution. This can lead to semantic conflicts, weakening the fusion effect and reducing the accuracy of localization and classification. To address this issue, existing research mainly falls into two categories. One type of method aligns modal features before fusion, utilizing relevant features from the other modality to correct the representation of the current modality, thereby enhancing cross-modal semantic consistency.
[0004] Specifically, existing technologies propose an end-to-end semantic alignment strategy that utilizes a mutual attention mechanism to achieve deep feature interaction between visible light and infrared images, and combines teacher network distillation to enhance semantic consistency, thereby effectively mitigating cross-modal semantic differences and generating high-precision fused images with consistent cross-modal semantics. Existing technologies also employ cross-modal cross-attention to enhance feature representation, and combine global affine and locally deformable convolution to finely align cross-modal features, thereby fully mitigating intermodal misalignment and semantic conflicts before fusion, providing more consistent feature representation for subsequent multi-scale fusion, and effectively improving the accuracy and robustness of multimodal target detection in complex scenes.
[0005] Another approach is to dynamically adjust the fusion weights of modes based on heuristic cues such as illumination intensity, highlighting modes with richer information under different conditions to mitigate the impact of semantic inconsistencies on the fusion effect. For example, one existing technology implements an adaptive fusion strategy through illumination intensity measurement: weighted bounding box fusion is used when the illumination is sufficient to comprehensively utilize the complementary information of visible light and infrared, while non-maximum suppression is switched under low illumination conditions to highlight the stable advantage of infrared modes, thereby dynamically adjusting the modal contribution, avoiding interference from low-quality features and mitigating the impact of semantic inconsistencies on the fusion effect; another existing technology uses a local illumination estimation module to predict the probabilities of high and low illumination, and adaptively adjusts the weights of infrared and visible light features during the fusion process accordingly: infrared information is enhanced in low illumination areas, and visible light details are highlighted in high illumination areas, thereby weakening the interference of low-quality features on the fusion process and effectively improving the detail clarity and robustness of the fused image.
[0006] However, existing methods generally overlook a crucial issue: the features extracted by the model are not always reliable, often exhibiting insufficient attention to the target region or incorrect attention to the background region. These unreliable features interfere with the effective fusion of cross-modal information, thereby weakening the model's overall detection performance. Therefore, establishing a scientific and effective modal reliability assessment mechanism and utilizing this information to guide the model in learning more stable, robust, and semantically consistent modal features remains a significant challenge that urgently needs to be addressed in this field. Summary of the Invention
[0007] This invention provides a training method, apparatus, device, and medium for a multimodal target detection model to address the problem in the prior art of neglecting the reliability of features when fusing visible light and infrared light.
[0008] In a first aspect, the present invention provides a training method for a multimodal target detection model, the multimodal target detection model comprising a feature extraction module, a feature fusion module, and a detection module, wherein the feature extraction module comprises a first modality feature extraction unit and a second modality feature extraction unit, and the method comprises: In each training iteration, forward inference and backpropagation are performed on the multimodal target detection model to obtain the gradients corresponding to all parameters in the multimodal target detection model; The reliability score of each parameter in the feature extraction module is obtained based on the validity of the parameters and the sensitivity of the gradient. The validity is used to characterize the contribution of the parameters to the model, and the sensitivity is used to characterize the degree of optimization of the parameters. The conflict set of parameters is determined based on the difference in gradient direction and reliability score in two modal feature extraction units. Based on the parameters in the conflict set, the gradient of the reliable mode is used to correct the gradient of the unreliable mode. The reliability score of the parameters in the reliable mode is greater than the reliability score of the parameters in the unreliable mode. The model parameters are updated based on the corrected gradient, and the next training iteration is performed until the iteration termination condition is met, resulting in the trained multimodal object detection model.
[0009] This invention dynamically evaluates the reliability of each modal feature by quantifying the effectiveness of parameters (their contribution to the model) and the sensitivity of gradients (the degree of parameter optimization). This allows for the accurate identification and correction of gradient conflicts between bimodal feature extraction units, ensuring that the model optimization process is always guided by more reliable modal information. This method significantly alleviates the semantic conflict problem in multimodal data, not only improving detection accuracy and reducing the risk of overfitting, but also enhancing the model's adaptability in complex scenarios (such as changes in lighting and occlusion).
[0010] In one alternative implementation, a reliability score for each parameter is obtained based on the validity of the parameter and the sensitivity of the gradient, including: determining the validity score of each parameter in the feature extraction module using counterfactual reasoning; determining the sensitivity score of the gradient based on the exponential moving average of the squared deviation between the current gradient and its historical mean; and determining the reliability score of each parameter based on the product of the validity score and the sensitivity score.
[0011] In this invention, a reliability score is calculated by fusing parameter validity and gradient sensitivity, providing refined optimization guidance for multimodal model training. Specifically, the validity score for each parameter is determined based on counterfactual reasoning, accurately measuring the contribution of each parameter; the sensitivity score is determined by the exponential moving average of the squared deviation between the current gradient and its historical mean, accurately reflecting the magnitude of gradient change over time.
[0012] In one optional implementation, counterfactual reasoning is used to determine the validity score of each parameter in the feature extraction module, including: determining the counterfactual loss based on the change in the loss value of the multimodal target detection model before and after parameter removal; performing a second-order Taylor expansion on the counterfactual loss to obtain first-order and second-order terms; approximating the diagonal of the Hessian matrix of the second-order terms using the Hutchinson method, and determining the validity score of each parameter by combining the calculation results with the first-order terms; and normalizing the validity scores.
[0013] Since directly intervening in each parameter individually incurs enormous computational costs, this invention employs a second-order Taylor expansion to approximate the counterfactual loss, thereby efficiently estimating the validity of the parameters. Simultaneously, in the calculation of the Hessian second-order term, this invention uses the Hutchinson method to reduce the computational complexity from the quadratic order to the linear order, significantly improving computational efficiency and scalability while ensuring estimation accuracy. Furthermore, the validity scores are normalized to ensure the comparability of values between different parameters and different modes.
[0014] In one optional implementation, determining a conflict set of parameters based on the difference in gradient direction and reliability score between two modality feature extraction units includes: determining whether the gradient direction is consistent in any output channel of any layer of the two modality feature extraction units based on the gradient cosine similarity of the corresponding parameters of the two modalities; determining the reliability difference in any output channel of any layer of the two modality feature extraction units based on the reliability score of the corresponding parameters of the two modalities; and constructing a conflict set based on the parameters of the corresponding output channel of the corresponding layer where the gradient direction is inconsistent and the reliability difference is greater than a threshold.
[0015] In this invention, by simultaneously considering gradient direction inconsistency and reliability differences, this conflict detection mechanism can establish a more robust and accurate discrimination criterion, effectively avoiding over-correction.
[0016] In one optional implementation, the gradient of the unreliable mode is corrected based on the parameters in the conflict set using the gradient of the reliable mode, including: determining the reliable mode and unreliable mode corresponding to each conflict position based on the parameter reliability score of each conflict position in the conflict set; at each conflict position, performing orthogonal decomposition on the gradient of the unreliable mode, removing the gradient projection of the reliable mode gradient direction, and obtaining the gradient projection in the orthogonal direction; and weightedly fusing the gradient projection in the orthogonal direction and the gradient of the reliable mode to obtain the gradient corrected for the unreliable mode.
[0017] This invention utilizes parameter reliability scores to dynamically distinguish between reliable and unreliable modes, ensuring that the correction process has a clear basis. Secondly, by orthogonally decomposing the gradients of unreliable modes, gradient components that conflict with reliable modes are effectively removed, fundamentally eliminating internal friction in the optimization direction. Finally, by weighted fusion to retain beneficial information from orthogonal directions and injecting guiding signals from reliable modes, the model significantly improves training stability and convergence efficiency while maintaining the complementary advantages of multimodal information. This effectively solves the gradient conflict problem in multimodal learning, enabling the model to exhibit stronger robustness and accuracy in complex scenarios.
[0018] In one alternative implementation, the weights for weighted fusion are determined based on the difference in reliability scores between reliable and unreliable modes at conflict locations.
[0019] In one alternative implementation, the conflict set is represented by the following formula:
[0020] In the formula, Indicates the first Layer For each output channel, the gradient cosine similarity of the parameters corresponding to the two modes is given. Indicates the first The mean of the negative cosine similarity of all output channels of the layer. Indicates the first Layer For each output channel, the reliability score of the parameters corresponding to the RGB mode is... Indicates the first Layer For each output channel, the reliability score of the corresponding IR mode parameter is... Indicates the first Average reliability difference across all output channels of the layer.
[0021] In this invention, when performing reliability difference analysis, the first... The average reliability difference of all output channels in the layer is used as the threshold for judging reliability differences, which effectively avoids misjudging minor and insignificant directional inconsistencies as conflicts, thereby triggering unnecessary corrections.
[0022] Secondly, the present invention provides a training apparatus for a multimodal target detection model, the multimodal target detection model including a feature extraction module, a feature fusion module, and a detection module, the feature extraction module including a first modality feature extraction unit and a second modality feature extraction unit, the apparatus comprising: The inference and propagation phase is used to perform forward inference and backward propagation on the multimodal target detection model during each training iteration to obtain the gradients corresponding to all parameters in the multimodal target detection model. The reliability estimation module is used to obtain the reliability score of each parameter in the feature extraction module based on the validity of the parameters and the sensitivity of the gradient. The validity is used to characterize the contribution of the parameters to the model, and the sensitivity is used to characterize the degree of optimization of the parameters. The conflict resolution module is used to determine the conflict set of parameters based on the difference in gradient direction and reliability score between two modal feature extraction units, and to correct the gradient of the unreliable mode based on the parameters in the conflict set using the gradient of the reliable mode, wherein the reliability score of the parameters in the reliable mode is greater than the reliability score of the parameters in the unreliable mode. The iteration module is used to update the model parameters based on the corrected gradient and the uncorrected gradient, and to perform the next training iteration until the iteration termination condition is met, so as to obtain the trained multimodal object detection model.
[0023] Thirdly, the present invention provides an electronic device, comprising: a memory and a processor, wherein the memory and the processor are communicatively connected to each other, the memory stores computer instructions, and the processor executes the computer instructions to perform the training method of the multimodal target detection model described in the first aspect or any corresponding embodiment thereof.
[0024] Fourthly, the present invention provides a computer-readable storage medium storing computer instructions for causing a computer to execute the training method of the multimodal target detection model described in the first aspect or any corresponding embodiment thereof.
[0025] Fifthly, the present invention provides a computer program product, including computer instructions for causing a computer to execute a training method for a multimodal target detection model according to the first aspect or any corresponding embodiment thereof. Attached Figure Description
[0026] To more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the drawings used in the description of the specific embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.
[0027] Figure 1 This is a schematic diagram of the first process of training a multimodal target detection model according to an embodiment of the present invention; Figure 2 This is a structural block diagram of a training device for a multimodal target detection model according to an embodiment of the present invention; Figure 3 This is a second structural block diagram of a training device for a multimodal target detection model according to an embodiment of the present invention; Figure 4 This is a schematic diagram of the gradient correction process according to an embodiment of the present invention; Figure 5 This is a schematic diagram of the hardware structure of an electronic device according to an embodiment of the present invention. Detailed Implementation
[0028] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0029] It is understood that before using the technical solutions disclosed in the various embodiments of the present invention, users should be informed of the types, scope of use, and usage scenarios of the personal information involved in the present invention and their authorization should be obtained in accordance with relevant laws and regulations through appropriate means.
[0030] The terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of that feature. In the description of this invention, "a plurality of" means two or more, unless otherwise explicitly specified.
[0031] According to an embodiment of the present invention, a training method for a multimodal target detection model is provided. It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. Furthermore, although a logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order than that shown here.
[0032] This embodiment provides a training method for a multimodal target detection model. Figure 1 This is a flowchart of a training method for a multimodal target detection model according to an embodiment of the present invention, such as... Figure 1 As shown, the process includes the following steps: Step S101: In each training iteration, forward inference and backpropagation are performed on the multimodal target detection model to obtain the gradients corresponding to all parameters in the multimodal target detection model.
[0033] Specifically, this multimodal target detection model is used to detect targets based on multimodal information. The target can be determined according to actual needs, such as a pedestrian or a vehicle. Target detection involves the localization and classification of targets, such as the localization, detection, and classification of pedestrians and vehicles. In this embodiment, the multimodal information includes visible light image (RGB) and infrared image (IR), which can be acquired by visible light sensors and infrared sensors, or through other methods. This embodiment does not specifically limit the method of acquiring multimodal information.
[0034] The multimodal target detection model is primarily used to extract and fuse features from different modalities, and then perform target detection based on the fused features. Therefore, this multimodal target detection model includes a feature extraction module, a feature fusion module, and a detection module. When the multimodal information includes visible light images and infrared images, the feature extraction module includes a first modality feature extraction unit and a second modality feature extraction unit, used to extract features from the two modalities respectively. For example, the first modality feature extraction unit for extracting visible light image features uses a visible light backbone network, which can be a network architecture such as a convolutional neural network, to capture texture, color, and detail information in the visible light image. The second modality feature extraction unit for extracting infrared image features uses an infrared backbone network, which has the same network structure as the visible light backbone network but with different parameter weights, to extract the thermal radiation features of the target, etc.
[0035] The feature fusion module is used to fuse features extracted by two modality feature extraction units. Fusion can be achieved through element-wise addition, concatenation, or weighted integration of features from different modalities using attention mechanisms. The detection module includes a shared detection head, which comprises a classification subnetwork or a localization subnetwork for classification or localization.
[0036] Based on the aforementioned multimodal object detection model, it needs to be trained before use for object detection to optimize its parameters and improve prediction accuracy. The parameters differ depending on the network used in the model. For example, if the first modality feature extraction unit uses a convolutional network, its parameters mainly include the weights of the convolutional kernels.
[0037] Specifically, model training is an iterative process, with each iteration typically involving forward inference, loss calculation, and backpropagation. Forward inference involves inputting information from different modalities into the model, and through layer-by-layer computation within the network, obtaining the predicted output. Loss calculation uses a loss function to calculate the difference between the predicted output and the true label (i.e., the loss value). Backpropagation optimizes the model parameters based on the loss value. Specifically, during parameter optimization, the partial derivative of each parameter (i.e., the gradient) is calculated according to the chain rule, and the backpropagation process is completed based on this gradient. In this embodiment, to address the issue of unreliable features extracted by trained models in related technologies, reliability estimation and conflict resolution processes are used to correct the parameter gradients, thereby promoting the learning of more reliable modal features.
[0038] Step S102: Based on the validity of the parameters and the sensitivity of the gradient, obtain the reliability score of each parameter in the feature extraction module. The validity is used to characterize the contribution of the parameter to the model, and the sensitivity is used to characterize the degree of optimization of the parameter.
[0039] Specifically, in the reliability estimation process, it is mainly used to generate a comprehensive reliability score. Since the features extracted by the model are not always reliable, and unreliable features may negatively impact detection performance, and given that features are generated from their corresponding model parameters, this embodiment performs reliability estimation based on parameter attribution. Therefore, when performing reliability estimation, this embodiment first evaluates the effectiveness of each parameter, that is, measures the contribution of each parameter to the model, for example, by measuring the impact of removing each parameter on the model's prediction results to obtain an effectiveness index. Furthermore, considering that different parameters may receive different degrees of optimization attention during training, a sensitivity score is introduced in the reliability estimation to reflect the magnitude of gradient changes over time, in order to measure the degree of parameter optimization during training. A higher sensitivity score means that the gradient change is more significant, indicating that the model more actively optimizes the parameter during training.
[0040] Step S103: Determine the conflict set of parameters based on the difference in gradient direction and reliability score in the two modal feature extraction units, and correct the gradient of the unreliable mode based on the parameters in the conflict set, wherein the reliability score of the parameters in the reliable mode is greater than the reliability score of the parameters in the unreliable mode.
[0041] Effective gradient correction relies on accurate conflict detection during training. Traditional methods typically treat any gradient cosine similarity that is negative as a conflict and adjust accordingly. However, this single-threshold-based criterion often triggers unnecessary corrections when the gradient direction deviates only slightly, leading to suboptimal optimization results. Therefore, this embodiment employs a dual-criteria conflict identification strategy. When determining a conflict, it considers both the inconsistency of gradient directions and the reliability gap between modes (i.e., the difference in reliability scores), and constructs a conflict set to complete conflict identification. For the gradients of parameters in the conflict set, the gradients of reliable modes are used to correct the gradients of unreliable modes.
[0042] Step S104 involves updating the model parameters based on the corrected gradient and performing the next training iteration until the iteration termination condition is met, resulting in the trained multimodal object detection model. Specifically, the gradient correction process primarily corrects the parameter gradients of the two modal feature extraction units in the feature extraction module. For other modules in the model, such as the feature fusion module and the detection module, the gradient correction process can be implemented using relevant technologies and will not be elaborated here. Furthermore, the process of parameter optimization based on the corrected gradient is not limited in this application. It should be noted that the above steps for gradient correction are performed in each iteration of the model training until the iteration termination condition is met, thus completing the model training. This iteration termination condition can be determined based on actual conditions, such as reaching a preset number of iterations.
[0043] This embodiment provides a training method for a multimodal target detection model, which includes the following steps: Step S201: During each training iteration, forward inference and backpropagation are performed on the multimodal object detection model to obtain the gradients corresponding to all parameters in the multimodal object detection model; for details, please refer to [link to relevant documentation]. Figure 1 Step S101 of the illustrated embodiment will not be described again here.
[0044] Step S202: Based on the validity of the parameters and the sensitivity of the gradient, obtain the reliability score of each parameter in the feature extraction module. The validity is used to characterize the contribution of the parameter to the model, and the sensitivity is used to characterize the degree of optimization of the parameter.
[0045] Specifically, step S202 includes: Step S2021: Counterfactual reasoning is used to determine the validity score of each parameter in the feature extraction module. Specifically, in order to measure the impact of removing each parameter on the model prediction results, this embodiment uses the principle of counterfactual reasoning. The parameter is set to zero through an intervention operation to obtain the counterfactual parameter set after removing the parameter, and the impact of this intervention on the model loss is determined. Thus, the validity score of the parameter is determined by the change in model loss after removing the parameter.
[0046] In an optional implementation, step S2021 includes: Step a1: Determine the counterfactual loss based on the change in loss value of the multimodal target detection model before and after parameter removal. Specifically, given the changes in loss value from the modal... Input samples and a complete set of model parameters The modality-specific parameters Representative mode In the backbone network Layer The parameter weights corresponding to each output channel. Through intervention operations... Setting it to zero yields the counterfactual parameter set after removing the parameter. And observe the effect of this intervention on the model loss. Based on this, the parameters The effectiveness, or counterfactual loss, is expressed by the following formula:
[0047] In the formula, This represents the loss function. When... When, it indicates that the parameter is removed. A negative value indicates that the parameter contributes positively to the prediction result, leading to increased losses. Conversely, a negative value indicates that the parameter may have a negative impact on the prediction.
[0048] Step a2 involves performing a second-order Taylor expansion on the counterfactual loss to obtain first-order and second-order terms. Specifically, deep learning models and other network models have millions or even hundreds of millions of parameters. Performing the above intervention operation on each parameter would be extremely costly. Therefore, this embodiment uses a second-order Taylor expansion to approximate the counterfactual loss, thereby efficiently estimating the validity of the parameters. This approximation process is expressed by the following formula:
[0049] Specifically, this approximation operation is mathematically equivalent to using the current parameter value. As the expansion point, for the loss function Perform a second-order Taylor expansion, and then estimate the value when the parameter is set to zero (i.e. The loss function value at time ). Wherein, the first-order term obtained in the above formula is expressed as: The second-order term is represented as In the second-order terms This indicates that the Hessian matrix corresponds to the parameter Its own second-order partial derivative.
[0050] Step a3 involves approximating the diagonal of the Hessian matrix for the second-order terms using the Hutchinson method, and then determining the validity score for each parameter by combining the calculation results with the first-order terms. Specifically, the first-order terms obtained from the second-order Taylor expansion can be calculated directly with almost no additional cost. The computational complexity of the second-order terms is on the order of squares; therefore, this embodiment uses the Hutchinson method to reduce the computational complexity from the order of squares to the order of linearities. This method can efficiently estimate the trace or specific quadratic forms of large-scale matrices without explicitly calculating and storing the entire matrix. This method allows for the parallel computation of the second-order terms for all parameters, and combining them with the first-order terms simultaneously determines the validity scores for all parameters. The computation process of the Hutchinson method can be implemented with reference to relevant technologies and will not be elaborated here. Thus, this embodiment significantly improves computational efficiency and scalability while ensuring estimation accuracy.
[0051] Step a4: Normalize the validity scores. Specifically, to ensure comparability between values of different parameters and different modes, the validity scores are normalized using z-score and then mapped to the (0,1) interval using the Sigmoid function, as expressed by the following formula:
[0052] In the formula, and These are the mean and standard deviation of the cross-modal validity score, respectively. It is the numerical stability constant. The Sigmoid function. Normalized validity score. As a scalar, the higher its value, the greater the contribution of the corresponding parameter to the model's prediction.
[0053] Step S2022: Determine the gradient sensitivity score based on the exponential moving average of the squared deviation between the current gradient and its historical mean; specifically, in this embodiment, sensitivity is defined as the exponential moving average of the squared deviation between the current gradient and its historical mean. Specifically, in the training... Step, parameters The gradient is denoted as Its historical average is denoted as Then the sensitivity score Defined as:
[0054] In the formula, This is a decay factor used to balance the influence of current and historical gradient information. The sensitivity score undergoes a similar normalization process to the validity score, ultimately yielding a normalized scalar result. The higher the score, the more optimization attention this parameter received during training.
[0055] Step S2023: Determine the reliability score for each parameter based on the product of the effectiveness score and the sensitivity score. Specifically, the reliability score for each modality-specific parameter is expressed by the following formula:
[0056] The score ranges from (0,1), with higher values indicating greater reliability of the corresponding modal-specific parameters, meaning the extracted features are more reliable. This reliability score will serve as an important guiding signal for subsequent conflict detection and gradient correction, providing a basis for cross-modal conflict optimization.
[0057] Step S203: Determine the conflict set of parameters based on the difference in gradient direction and reliability score in the two modal feature extraction units, and correct the gradient of the unreliable mode based on the parameters in the conflict set, wherein the reliability score of the parameters in the reliable mode is greater than the reliability score of the parameters in the unreliable mode.
[0058] Specifically, step S203 includes: Step S2031: In any output channel of any layer of the two modality feature extraction units, determine whether the gradient directions are consistent based on the gradient cosine similarity of the corresponding parameters of the two modalities; specifically, for the first... Layer For each output channel, first obtain the gradients of the corresponding parameters of the first mode (RGB) and the second mode (IR), which are expressed as follows: and Then calculate the cosine similarity between the two gradients. The cosine similarity metric measures the similarity at a specific location. The first step is to check if the optimization directions of the two modalities are consistent. If the cosine similarity is negative, it indicates that the gradient directions of the two modalities are inconsistent, and the optimization direction of one modality will cancel out the learning effect of the other modality. However, simply using a negative cosine similarity is not enough to reliably determine a conflict. Therefore, this embodiment introduces a dynamic threshold, which is the mean of the negative cosine similarities of all channels in this layer. When the calculated cosine similarity is less than the mean, the gradient directions are considered inconsistent. The mean can be calculated by filtering out cosine similarities less than 0 after calculating the cosine similarity of all channels in the layer, and then averaging these filtered cosine similarities.
[0059] Step S2032: In any output channel of any layer of the two modality feature extraction units, determine the reliability difference based on the reliability scores of the corresponding parameters of the two modalities; specifically, for the first... Layer Each output channel can calculate the difference in reliability scores between two modes based on the reliability scores obtained in the above steps, thus obtaining the reliability difference.
[0060] Step S2033: Construct a conflict set based on the parameters of the corresponding output channels of the corresponding layers where the gradient directions are inconsistent and the reliability differences exceed a threshold. Specifically, the conflict set can be represented by the following formula:
[0061] In the formula, Indicates the first Layer For each output channel, the gradient cosine similarity of the parameters corresponding to the two modes is given. Indicates the first The mean of the negative cosine similarity of all output channels of the layer. Indicates the first Layer For each output channel, the reliability score of the parameters corresponding to the RGB mode is... Indicates the first Layer For each output channel, the reliability score of the corresponding IR mode parameter is... Indicates the first The average reliability difference across all output channels of the layer. Specifically, by simultaneously considering gradient direction inconsistency and reliability differences, this conflict detection mechanism can establish a more robust and accurate discrimination criterion, effectively avoiding over-correction.
[0062] Step S2034: Determine the reliable and unreliable modes corresponding to each conflict location based on the parameter reliability score of each conflict location in the conflict set; specifically, as can be seen from the above steps, the corresponding channel (i.e., conflict location) of the corresponding layer with conflict can be determined by constructing the conflict set. For example, in the first... Layer If the gradient directions of the output channels are inconsistent and the reliability difference is greater than a threshold, then the... Layer Each output channel represents a conflict location. For each conflict location, the mode with the higher reliability score of the parameters is identified as the reliable mode. Then the other mode is an unreliable mode. Therefore, reliable modes It is expressed by the following formula:
[0063] Step S2035: At each conflict location, orthogonally decompose the gradient of the unreliable mode, removing the gradient projection along the reliable mode gradient direction to obtain the gradient projection along the orthogonal direction. Specifically, to mitigate the adverse effects of the unreliable mode gradient on the optimization process, its gradient is first orthogonally decomposed to remove its projection along the reliable mode gradient direction, retaining only the effective components along the orthogonal direction, thereby eliminating interference from the conflicting parts. Specifically, the obtained gradient projection along the orthogonal direction is expressed by the following formula:
[0064] In the formula, and These are the unreliable and reliable modes at the conflict locations. gradient vector, This represents the vector dot product operation. Let represent the squared 2-norm of a vector.
[0065] Step S2036 involves weighted fusion of the gradient projections along the orthogonal directions and the gradients of the reliable modes to obtain the gradients corrected for the unreliable modes. Specifically, after obtaining the orthogonal components... Subsequently, this embodiment employs a reliability-guided weighted fusion strategy, fusing the gradient with the gradient of the reliable modality according to their respective reliability weights, thereby obtaining the corrected gradient and further promoting optimization consistency in the cross-modal feature learning process. This weighted fusion process is expressed by the following formula:
[0066] In the formula, the fusion weight The weights are used to control the contribution of reliable mode gradients in the weighted fusion. A larger value indicates a stronger guiding effect provided by the reliable mode. The fusion weights are determined based on the difference in reliability scores between reliable and unreliable modes at conflict locations. Specifically, the fusion weights are expressed by the following formula:
[0067] In the formula, This represents the Sigmoid function.
[0068] Step S204: Update the model parameters based on the corrected gradient and proceed with the next training iteration until the iteration termination condition is met, resulting in the trained multimodal object detection model. For details, please refer to [link to details]. Figure 1 Step S104 of the illustrated embodiment will not be described again here.
[0069] The method provided in this embodiment utilizes modal reliability information to correct conflict gradients during the optimization process, thereby promoting stable and reliable modal feature learning, mitigating cross-modal semantic conflicts, and improving detection accuracy and robustness. It also possesses good versatility and can be applied to most visible-infrared target detection frameworks.
[0070] This embodiment also provides a training device for a multimodal target detection model, which is used to implement the above embodiments and preferred embodiments; details already described will not be repeated. As used below, the term "module" can refer to a combination of software and / or hardware that performs a predetermined function. Although the device described in the following embodiments is preferably implemented in software, hardware implementation, or a combination of software and hardware, is also possible and contemplated.
[0071] This embodiment provides a training device for a multimodal target detection model, such as... Figure 2 As shown, it includes: The inference and propagation phase 21 is used to perform forward inference and backward propagation on the multimodal target detection model during each training iteration to obtain the gradients corresponding to all parameters in the multimodal target detection model. The reliability estimation module 22 is used to obtain the reliability score of each parameter in the feature extraction module based on the validity of the parameters and the sensitivity of the gradient. The validity is used to characterize the contribution of the parameters to the model, and the sensitivity is used to characterize the degree of optimization of the parameters. The conflict resolution module 23 is used to determine the conflict set of parameters based on the difference between the gradient direction and reliability score in the two modal feature extraction units, and to correct the gradient of the unreliable mode based on the parameters in the conflict set using the gradient of the reliable mode, wherein the reliability score of the parameters in the reliable mode is greater than the reliability score of the parameters in the unreliable mode. Iteration module 24 is used to update the model parameters based on the corrected gradient and the uncorrected gradient, and to perform the next training iteration until the iteration termination condition is met, so as to obtain the trained multimodal object detection model.
[0072] In one alternative implementation, such as Figure 3As shown, the training device for the multimodal target detection model includes a detection framework based on reliable gradient correction. This framework comprises a multimodal target detection model that receives visible light and infrared images. The visible light image is input to a visible light backbone network for feature extraction, yielding visible light features. The infrared image is input to an infrared backbone network for feature extraction, yielding infrared features. The visible light and infrared features are then fused and detected to obtain the detection result. Specifically, training the multimodal target detection model involves multiple iterations, each including forward inference and backpropagation. Gradient correction is performed during backpropagation to optimize model parameter updates. In this embodiment, gradient correction is implemented through a reliability estimation module based on parameter attribution and a conflict resolution module guided by reliability.
[0073] The reliability estimation module based on parameter attribution first acquires the parameters (i.e., factual parameters) in the modal backbone network (i.e., visible light backbone network and infrared backbone network). For any parameter (i.e., the parameter to be intervened), counterfactual inference is performed to obtain the counterfactual parameter. The effectiveness is determined by the difference in loss between the factual and counterfactual parameters. Simultaneously, the sensitivity is determined by the exponential moving average of the gradient changes of the parameter over multiple training steps. Finally, the reliability of each parameter is obtained by combining the effectiveness and sensitivity, and reliable and unreliable parameters are distinguished based on reliability.
[0074] The reliability-guided conflict resolution module comprises a dual-criteria conflict identification submodule and a reliability-guided gradient correction submodule. The dual-criteria conflict identification submodule first acquires the original gradients (original gradients of the visible light mode and the infrared mode) and reliability scores (parameter reliability scores of the visible light mode and the infrared light mode) of the parameters in the modal backbone network for conflict identification. It identifies conflict indices (sets of conflict locations) with significant conflicts and non-conflict indices (sets of non-conflict locations) that require no intervention, and constructs a conflict set based on the conflict indices with significant conflicts. The reliability-guided gradient correction submodule corrects the gradients of unreliable modes using reliable modes within the conflict set. During the correction process, the unreliable gradient mode is decomposed into different components, and reliable modes are used to guide the correction of the orthogonal components, resulting in corrected gradients (such as the corrected gradients for visible light or infrared light).
[0075] Specifically, such as Figure 4 The diagram shows a reliability-guided gradient correction submodule. This process can be divided into two stages: First, the conflict set obtained through conflict identification is acquired, and then the gradient of the reliable mode at each conflict location in the conflict set is obtained. gradients of unreliable modes gradient for unreliable modes Perform orthogonal decomposition to remove components that conflict with reliable mode gradients, retaining only the components in the orthogonal directions. Secondly, based on the difference in modal reliability, this orthogonal component is compared with the gradient of the reliable mode. Reliability-weighted fusion yields two weighted components. and Together they constitute the corrected gradient. .
[0076] To test the detection accuracy of the multimodal object detection model trained in this embodiment, the model was tested on the public datasets VEDAI, LLVIP, and DroneVehicle. This was done in conjunction with the parameter attribution-based reliability estimation module and the reliability-guided conflict resolution module provided in this embodiment, as well as related object detection models (such as S...). 2 A comparison of the detection accuracy and other metrics between A-NET, YOLOv5, Faster R-CNN and directly using relevant object detection models is shown in Tables 1, 2, and 3 below: Table 1 Comparison results of the VEDAI dataset
[0077] Table 2 Comparison results of the LLVIP dataset
[0078] Table 3 Comparison results of the DroneVehicle dataset
[0079] As can be seen, on the public datasets VEDAI, LLVIP, and Drone-Vehicle, both the parameter attribution-based reliability estimation module and the reliability-guided conflict resolution module of this embodiment achieve stable performance improvements, effectively enhancing detection accuracy and robustness. Furthermore, this module can be seamlessly integrated into most existing visible-infrared target detection frameworks, exhibiting stable performance improvements across different frameworks, fully demonstrating its broad applicability and scalability. Simultaneously, it improves the reliability of modal features, enabling the model to focus more intently on the target region and reduce background interference, thereby effectively reducing false positives and false negatives, further enhancing the stability and accuracy of the detection results.
[0080] The reliability estimation module based on parameter attribution designed in this invention is based on the principle that "features are generated by their corresponding parameters". It evaluates modal reliability from the perspective of parameters by combining two complementary dimensions: 1. Effectiveness, which quantifies the impact of parameters on prediction performance through counterfactual reasoning; 2. Sensitivity, which reflects the degree to which parameters are optimized based on the dynamic changes of gradients during training.
[0081] The conflict resolution module based on reliability guidance designed in this invention uses modal reliability to identify conflicts and guide gradient correction, thereby achieving stable and consistent cross-modal optimization and improving detection performance.
[0082] The training apparatus for the multimodal target detection model provided in this embodiment of the invention can execute the training method for the multimodal target detection model provided in any embodiment of the invention, and has the corresponding functional modules and beneficial effects for executing the method. Further functional descriptions of the various modules and units described above are the same as in the corresponding embodiments described above, and will not be repeated here.
[0083] Figure 5 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present invention.
[0084] The following is a detailed reference. Figure 5 This diagram illustrates a structural schematic suitable for implementing an electronic device according to embodiments of the present invention. The electronic device may include a processor (e.g., a central processing unit, a graphics processing unit, etc.) 11, which can perform various appropriate actions and processes according to a program stored in read-only memory (ROM) 12 or a program loaded from memory 18 into random access memory (RAM) 13. The RAM 13 also stores various programs and data required for the operation of the electronic device. The processor 11, ROM 12, and RAM 13 are interconnected via a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.
[0085] Typically, the following devices can be connected to I / O interface 15: input devices 16 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 17 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; memory devices 18 including, for example, magnetic tapes, hard disks, etc.; and communication devices 19. Communication device 19 allows electronic devices to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 5 Electronic devices with various devices are shown, but it should be understood that it is not required to implement or have all of the devices shown, and more or fewer devices may be implemented or have instead.
[0086] In particular, according to embodiments of the present invention, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of the present invention include a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device 19, or installed from a memory 18, or installed from a ROM 12. When the computer program is executed by the processor 11, it performs the functions defined in the training method of the multimodal target detection model of the embodiments of the present invention.
[0087] Figure 5 The electronic device shown is merely an example and should not be construed as limiting the functionality and scope of use of the embodiments of the present invention.
[0088] This invention also provides a computer-readable storage medium. The methods described above according to embodiments of the invention can be implemented in hardware or firmware, or implemented as computer code that can be recorded on a storage medium, or implemented as computer code downloaded via a network and originally stored on a remote storage medium or a non-transitory machine-readable storage medium and then stored on a local storage medium. Thus, the methods described herein can be processed by software stored on a storage medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware. The storage medium can be a magnetic disk, optical disk, read-only memory, random access memory, flash memory, hard disk, or solid-state drive, etc.; further, the storage medium can also include combinations of the above types of memory. It is understood that computers, processors, microprocessor controllers, or programmable hardware include storage components capable of storing or receiving software or computer code. When the software or computer code is accessed and executed by the computer, processor, or hardware, the training method of the multimodal target detection model shown in the above embodiments is implemented.
[0089] A portion of this invention can be applied as a computer program product, such as computer program instructions, which, when executed by a computer, can invoke or provide the methods and / or technical solutions according to the invention through the operation of the computer. Those skilled in the art will understand that the forms in which computer program instructions exist in a computer-readable medium include, but are not limited to, source files, executable files, installation package files, etc. Correspondingly, the ways in which computer program instructions are executed by a computer include, but are not limited to: the computer directly executing the instructions, or the computer compiling the instructions and then executing the corresponding compiled program, or the computer reading and executing the instructions, or the computer reading and installing the instructions and then executing the corresponding installed program. Here, the computer-readable medium can be any available computer-readable storage medium or communication medium accessible to a computer.
[0090] Although embodiments of the invention have been described in conjunction with the accompanying drawings, those skilled in the art can make various modifications and variations without departing from the spirit and scope of the invention, and such modifications and variations all fall within the scope defined by the appended claims.
Claims
1. A training method for a multimodal target detection model, characterized in that, The multimodal target detection model includes a feature extraction module, a feature fusion module, and a detection module. The feature extraction module includes a first modality feature extraction unit and a second modality feature extraction unit. The method includes: In each training iteration, forward inference and backpropagation are performed on the multimodal target detection model to obtain the gradients corresponding to all parameters in the multimodal target detection model; The reliability score of each parameter in the feature extraction module is obtained based on the validity of the parameters and the sensitivity of the gradient. The validity is used to characterize the contribution of the parameters to the model, and the sensitivity is used to characterize the degree of optimization of the parameters. The conflict set of parameters is determined based on the difference in gradient direction and reliability score in two modal feature extraction units. Based on the parameters in the conflict set, the gradient of the reliable mode is used to correct the gradient of the unreliable mode. The reliability score of the parameters in the reliable mode is greater than the reliability score of the parameters in the unreliable mode. The model parameters are updated based on the corrected gradient, and the next training iteration is performed until the iteration termination condition is met, resulting in the trained multimodal object detection model.
2. The method according to claim 1, characterized in that, The reliability score for each parameter is derived based on parameter validity and gradient sensitivity, including: Counterfactual reasoning is used to determine the validity score of each parameter in the feature extraction module; The sensitivity score of the gradient is determined by the exponential moving average of the squared deviation between the current gradient and its historical mean. The reliability score for each parameter is determined by multiplying the validity score and the sensitivity score.
3. The method according to claim 2, characterized in that, Counterfactual reasoning is used to determine the validity score of each parameter in the feature extraction module, including: The counterfactual loss is determined based on the change in the loss value of the multimodal target detection model before and after parameter removal; Performing a second-order Taylor expansion on the counterfactual loss yields first-order and second-order terms; The Hutchinson method is used to approximate the diagonal of the Hessian matrix of the second-order terms, and the validity score of each parameter is determined by combining the calculation results with the first-order terms. The validity scores are then normalized.
4. The method according to claim 1, characterized in that, The conflict set of parameters is determined based on the differences in gradient direction and reliability score between two modal feature extraction units, including: In any output channel of any layer of the two modality feature extraction units, the gradient direction is determined to be consistent based on the gradient cosine similarity of the corresponding parameters of the two modalities. In any output channel of any layer of the two modality feature extraction units, the reliability difference is determined based on the reliability scores of the corresponding parameters of the two modalities; A conflict set is constructed based on the parameters of the corresponding output channels of the corresponding layers where the gradient directions are inconsistent and the reliability differences are greater than a threshold.
5. The method according to claim 1, characterized in that, Based on the parameters in the conflict set, the gradient of the unreliable mode is corrected using the gradient of the reliable mode, including: The reliable and unreliable modes corresponding to each conflict location are determined based on the parameter reliability score of each conflict location in the conflict set. At each conflict location, the gradient of the unreliable mode is orthogonally decomposed, and the gradient projection of the reliable mode gradient direction is removed to obtain the gradient projection in the orthogonal direction. The gradient projections in the orthogonal direction and the gradients of the reliable modes are weighted and fused to obtain the gradients after correction for the unreliable modes.
6. The method according to claim 5, characterized in that, The weights for weighted fusion are determined based on the difference in reliability scores between reliable and unreliable modes at conflict locations.
7. The method according to claim 4, characterized in that, The conflict set is represented by the following formula: In the formula, Indicates the first Layer For each output channel, the gradient cosine similarity of the parameters corresponding to the two modes is given. Indicates the first The mean of the negative cosine similarity of all output channels of the layer. Indicates the first Layer For each output channel, the reliability score of the parameters corresponding to the RGB mode is... Indicates the first Layer For each output channel, the reliability score of the corresponding IR mode parameter is... Indicates the first Average reliability difference across all output channels of the layer.
8. A training device for a multimodal target detection model, characterized in that, The multimodal target detection model includes a feature extraction module, a feature fusion module, and a detection module. The feature extraction module includes a first modality feature extraction unit and a second modality feature extraction unit. The device includes: The inference and propagation phase is used to perform forward inference and backward propagation on the multimodal target detection model during each training iteration to obtain the gradients corresponding to all parameters in the multimodal target detection model. The reliability estimation module is used to obtain the reliability score of each parameter in the feature extraction module based on the validity of the parameters and the sensitivity of the gradient. The validity is used to characterize the contribution of the parameters to the model, and the sensitivity is used to characterize the degree of optimization of the parameters. The conflict resolution module is used to determine the conflict set of parameters based on the difference in gradient direction and reliability score between two modal feature extraction units, and to correct the gradient of the unreliable mode based on the parameters in the conflict set using the gradient of the reliable mode, wherein the reliability score of the parameters in the reliable mode is greater than the reliability score of the parameters in the unreliable mode. The iteration module is used to update the model parameters based on the corrected gradient and the uncorrected gradient, and to perform the next training iteration until the iteration termination condition is met, so as to obtain the trained multimodal object detection model.
9. An electronic device, characterized in that, include: A memory and a processor are communicatively connected, the memory stores computer instructions, and the processor executes the computer instructions to perform the training method of the multimodal target detection model according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions for causing the computer to execute the training method of the multimodal target detection model according to any one of claims 1 to 7.
Citation Information
Cited By
Large model training full-process fault-tolerant analysis system and method
CN121860093A