Model training methods, devices, electronic equipment, and storage media

CN122572579APending Publication Date: 2026-08-14IFLYTEK CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-08
Publication Date
2026-08-14

AI Technical Summary

Technical Problem

[0004]本发明提供一种模型训练方法、装置、电子设备和存储介质,用以解决相关技术中多模态模型训练难以兼顾单模态特异信息和跨模态协同信息的缺陷

Benefits of technology

[0016]本发明提供的模型训练方法、装置、电子设备和存储介质,在每个模态下设置单模态教师模型和多模态教师模型两类教师知识来源,并基于两类教师模型提取的单模态特征和/或协同模态特征对多模态学生模型的特征提取分支进行蒸馏训练,使得多模态学生模型能够更加合理、充分地同时吸收单模态教师提供的模态特定细节信息以及多模态教师提供的跨模态协同判别信息,从而在多模态协同建模信息与模态特异信息保留之间取得最佳平衡,显著提升了多模态学生模型在复杂场景下的训练稳定性、鲁棒性以及整体性能。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122572579A_ABST
    Figure CN122572579A_ABST
Patent Text Reader

Abstract

This invention provides a model training method, apparatus, electronic device, and storage medium. The method includes: for each modality of a multimodal student model, inputting sample data of the modality into the corresponding unimodal teacher model and multimodal teacher model respectively, to obtain unimodal features extracted by the unimodal teacher model and comodal modal features extracted by the multimodal teacher model; and performing distillation training on the feature extraction branches of the modality in the multimodal student model based on the sample data, as well as the unimodal features and / or comodal modal features. The method, apparatus, electronic device, and storage medium provided by this invention enable the multimodal student model to more reasonably and fully absorb modality-specific detailed information provided by the unimodal teacher and cross-modal collaborative discriminative information provided by the multimodal teacher, achieving an optimal balance between multimodal collaborative modeling information and modality-specific information preservation, significantly improving the training stability, robustness, and overall performance of the model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of artificial intelligence technology, and in particular to a model training method, apparatus, electronic device, and storage medium. Background Technology

[0002] Currently, modeling methods for multimodal models are mainly divided into two categories: modality-independent modeling and cross-modality unified modeling. To combine the advantages of these two modeling methods, related technologies attempt to introduce knowledge distillation mechanisms to guide the training of multimodal models.

[0003] However, in existing multimodal knowledge distillation methods, features extracted by a single-source teacher model are usually used directly as supervision signals. Such supervision signals cannot comprehensively and specifically cover fine-grained features and global collaborative features, making it impossible for the trained multimodal model to simultaneously take into account cross-modal collaborative modeling and the preservation of single-modal specific information. This results in significant performance bottlenecks when dealing with complex samples. Summary of the Invention

[0004] This invention provides a model training method, apparatus, electronic device, and storage medium to address the shortcomings of related technologies in multimodal model training, which struggle to simultaneously incorporate single-modal specific information and cross-modal collaborative information.

[0005] This invention provides a model training method, comprising: For each modality of the multimodal student model, the sample data of the modality is input into the corresponding unimodal teacher model and multimodal teacher model to obtain the unimodal features extracted by the unimodal teacher model and the collaborative modal features extracted by the multimodal teacher model. Based on the sample data, as well as the single-modal features and / or the collaborative modal features, the feature extraction branches of the modal in the multimodal student model are subjected to distillation training.

[0006] According to a model training method provided by the present invention, the step of performing distillation training on the feature extraction branches of the modalities in the multimodal student model based on the sample data, the single-modal features, and / or the co-modal features includes: Based on the similarity between the unimodal features and the similarity between the comodal features, the differences in sample relationship cognition between the unimodal teacher model and the multimodal teacher model are determined. Based on the cognitive differences in sample relationships, the sample data, and the single-modal features and / or the collaborative modal features, the feature extraction branches of the modal in the multimodal student model are subjected to distillation training.

[0007] According to a model training method provided by the present invention, the step of performing distillation training on the feature extraction branches of the multimodal student model based on the cognitive differences in sample relationships, the sample data, and the single-modal features and / or the co-modal features includes: If the difference in the perceived relationship between the samples is greater than or equal to a first threshold, the single-modal feature or the collaborative modal feature is determined as a reference feature based on the reliability of the single-modal teacher model and the reliability of the multimodal teacher model; the reliability is determined based on the modal feature similarity of positive sample pairs and / or the modal feature similarity of negative sample pairs in the sample data. Based on the sample data and the reference features, the feature extraction branches of the modal in the multimodal student model are trained by distillation.

[0008] According to a model training method provided by the present invention, the step of performing distillation training on the feature extraction branches of the modal in the multimodal student model based on the sample data and the reference features includes: If the difference between the reliability of the unimodal teacher model and the reliability of the multimodal teacher model is greater than or equal to the second threshold, the feature extraction branch is distilled and trained with the goal of making the features of the sample data extracted by the feature extraction branch close to the reference features and far away from the counterexample features. The counterexample feature is a modal feature that is different from the reference feature.

[0009] According to a model training method provided by the present invention, the step of performing distillation training on the feature extraction branches of the multimodal student model based on the cognitive differences in sample relationships, the sample data, and the single-modal features and / or the co-modal features includes: If the difference in the perceived relationship between the samples is less than a first threshold, the single-modal features and the collaborative modal features are fused to obtain fused features; Based on the sample data and the fusion features, the feature extraction branches of the modal in the multimodal student model are trained by distillation.

[0010] According to a model training method provided by the present invention, the fusion of the single-modal features and the co-modal features to obtain fused features includes: The fusion weights are determined based on the reliability of the unimodal teacher model and the reliability of the multimodal teacher model; the reliability is determined based on the modal feature similarity of positive sample pairs and / or the modal feature similarity of negative sample pairs in the sample data. Based on the fusion weights, the single-modal features and the collaborative modal features are weighted and fused to obtain the fused features.

[0011] According to a model training method provided by the present invention, the multimodal student model, the unimodal teacher model, and the multimodal teacher model are all target re-identification models, and the modality is a spectral modality.

[0012] The present invention also provides a model training apparatus, comprising: The feature extraction module is used to input the sample data of each modality of the multimodal student model into the corresponding unimodal teacher model and multimodal teacher model respectively, to obtain the unimodal features extracted by the unimodal teacher model and the collaborative modal features extracted by the multimodal teacher model. The distillation training module is used to perform distillation training on the feature extraction branches of the modalities in the multimodal student model based on the sample data, as well as the single-modal features and / or the collaborative modal features.

[0013] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the model training method described above.

[0014] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the model training method as described above.

[0015] The present invention also provides a computer program product, including a computer program that, when executed by a processor, implements the model training method as described above.

[0016] The model training method, apparatus, electronic device, and storage medium provided by this invention set up two types of teacher knowledge sources in each modality: a unimodal teacher model and a multimodal teacher model. Based on the unimodal features and / or comodal features extracted by the two types of teacher models, the feature extraction branch of the multimodal student model is distilled and trained. This enables the multimodal student model to more reasonably and fully absorb the modality-specific detailed information provided by the unimodal teacher and the cross-modal collaborative discrimination information provided by the multimodal teacher. Thus, it achieves the best balance between multimodal collaborative modeling information and modality-specific information preservation, significantly improving the training stability, robustness, and overall performance of the multimodal student model in complex scenarios. Attached Figure Description

[0017] To more clearly illustrate the technical solutions in this invention or related technologies, the accompanying drawings used in the description of the embodiments or related technologies will be briefly introduced below. Obviously, the accompanying drawings described below are some embodiments of this invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0018] Figure 1 This is one of the flowcharts illustrating the model training method provided by the present invention.

[0019] Figure 2 This is the second flowchart of the model training method provided by the present invention.

[0020] Figure 3 This is a flowchart illustrating the training method of the multispectral target re-identification model provided by the present invention.

[0021] Figure 4 This is a schematic diagram of the model training device provided by the present invention.

[0022] Figure 5 This is a schematic diagram of the structure of the electronic device provided by the present invention. Detailed Implementation

[0023] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.

[0024] All actions involving the acquisition of signal information or data in this invention are carried out in compliance with the relevant data protection laws and policies of the country where the device is located, and with the authorization granted by the owner of the device.

[0025] With the development of artificial intelligence technology, multimodal models have been widely used in many fields such as computer vision and natural language processing. Multimodal data are highly complementary in information representation, and fusing multimodal features can effectively improve the robustness and generalization ability of artificial intelligence models.

[0026] Taking target re-identification technology in multimodal scenarios as an example, target re-identification technology aims to identify and associate the same target from images or videos acquired by multiple non-overlapping cameras. It is one of the key technologies in systems such as smart cities, public safety monitoring, intelligent transportation, and behavior analysis. In practical applications, targets often need to be matched at different times, in different spaces, and under different imaging conditions, which places high demands on the robustness and generalization ability of the system.

[0027] Traditional target re-identification systems are mostly based on visible light imaging equipment for modeling. However, in complex environments such as nighttime, low light, backlight, and inclement weather, the appearance information of targets in visible light images is easily degraded, thus affecting re-identification performance. To improve the stability and applicability of the system in complex environments, multimodal target re-identification technology introduces spectral imaging data from multiple spectral modalities, including visible light, near-infrared, and thermal infrared. This enables the acquisition of structural and thermal radiation information of targets under low light or no light conditions, offering significant advantages in all-weather monitoring and complex environmental perception scenarios. The complementarity of information representation among different spectral modalities provides a new technical approach to improving the robustness and accuracy of target re-identification models.

[0028] Currently, there are two main modeling approaches for multimodal models: modal-independent modeling and cross-modal unified modeling.

[0029] Modality-independent modeling involves constructing independent feature extraction networks for different modalities, thereby enhancing complementarity and modality specificity. This modeling approach has certain advantages in preserving modality-specific details, but the collaborative modeling capability between different modalities usually relies on explicitly designed interaction modules, resulting in limited overall multimodal collaboration. It is also prone to performance bottlenecks when dealing with difficult samples with significant cross-modal differences.

[0030] Cross-modal unified modeling employs a shared structure to extract features. For example, it can leverage the powerful cross-modal consistency capabilities of large models such as visual and language models to uniformly process data from different modalities, thereby enhancing cross-modal semantic alignment and collaborative understanding. While this modeling approach can effectively mine the collaborative semantics of multimodal data and demonstrates a strong advantage in overall matching performance, it tends to suppress modality-specific fine-grained discriminative information while emphasizing the commonalities of multimodalities, thus affecting the model's ability to characterize key details under certain modal conditions.

[0031] To combine the advantages of the two modeling approaches mentioned above, related technologies attempt to introduce knowledge distillation mechanisms to guide the training of multimodal models. However, existing training methods such as knowledge distillation typically use the features of one teacher model directly as the supervision signal, or simply take a weighted average of the features of multiple teacher models as the supervision signal. The supervision signal cannot take into account both fine-grained modality-specific features and global cross-modal collaborative features, making it difficult for existing model training methods to fully and reasonably utilize the knowledge provided by the teacher models.

[0032] To address the aforementioned problems, embodiments of the present invention provide a model training method that can be applied to the training of multimodal models, such as multimodal target re-identification models, as well as multimodal target detection models, multimodal generative models, and other models that require feature extraction from multimodal inputs.

[0033] Figure 1 This is one of the flowcharts illustrating the model training method provided by this invention, such as... Figure 1 As shown, the method includes: Step 110: For each modality of the multimodal student model, input the sample data of the modality into the corresponding unimodal teacher model and multimodal teacher model respectively to obtain the unimodal features extracted by the unimodal teacher model and the collaborative modal features extracted by the multimodal teacher model.

[0034] Specifically, modality can refer to the form of data representation, information source channel, or sensor type. For example, in a multimodal target re-identification scenario, modality can include visible light modality, near-infrared modality, thermal infrared modality, etc.; in a multimodal autonomous driving environmental perception scenario, modality can include image modality, point cloud modality, etc.; and in a multimodal medical image-assisted diagnosis scenario, modality can include CT (Computer tomography) modality, MRI (Magnetic Resonance Imaging) modality, etc.

[0035] The multimodal student model is the target model to be trained in the embodiments of the present invention. The multimodal student model has the ability to process multimodal data, and in the model training method provided in the embodiments of the present invention, the multimodal student model is regarded as the student model in knowledge distillation.

[0036] It should be noted that for each supported modality, the multimodal student model includes a feature extraction branch corresponding to that modality, which is used to extract features of the corresponding modality input data independently or collaboratively.

[0037] In some embodiments, the multimodal student model can be a modality-independent model. That is, for each modality in the multimodal student model, there is an independent feature extraction branch that corresponds one-to-one with that modality, and the feature extraction branch corresponding to that modality is only used to extract features from the input data of that modality.

[0038] In other embodiments, the multimodal student model can also be a model that unifies modeling across modalities. That is, in the multimodal student model, the same feature extraction branch is shared for each modality, and this feature extraction branch can collaboratively extract features from the input data of each modality.

[0039] In this embodiment of the invention, for each modality supported by the multimodal student model, two corresponding teacher networks are configured, namely a unimodal teacher model and a multimodal teacher model.

[0040] A unimodal teacher model refers to a teacher model trained using only training data corresponding to a single modality. It can also be understood as a unimodal teacher model capable only of processing data from a single modality. The role of a unimodal teacher model is to provide fine-grained, modality-specific knowledge within a given modality. Taking a multimodal target re-identification scenario as an example, for the visible light modality, the unimodal teacher model in the visible light modality is a target re-identification model trained solely on visible light data. Furthermore, the features extracted by this target re-identification model for the visible light modality can reflect the unique color, texture, edge, and other detailed information of visible light images.

[0041] A multimodal teacher model refers to a teacher model trained jointly using data from multiple different modalities. It can also be understood as a multimodal teacher model capable of processing multimodal data. For each modal input, the multimodal teacher model can output the features of that modality through its corresponding feature extraction branch. Because the training of a multimodal teacher model is influenced by the collaborative constraints of multimodal joint training and global discriminative information, its role is to provide cross-modal collaborative modal representations and multimodal collaborative knowledge. Taking a multimodal target re-identification scenario as an example, the multimodal teacher model can be a target re-identification model trained using multimodal data such as visible light data, near-infrared data, and thermal infrared data. For the visible light modality, the features extracted by this target re-identification model can reflect the modal representation of visible light images under multimodal collaboration of visible light, near-infrared, and thermal infrared.

[0042] For each modality supported by the multimodal student model, the corresponding sample data can be input into the corresponding unimodal teacher model and multimodal teacher model. Thus, the unimodal teacher model can extract features based on the input sample data, outputting the unimodal features of the sample data; the multimodal teacher model can also extract features based on the input sample data, outputting the comodal features of the sample data.

[0043] It is understandable that unimodal features and comodal features are features obtained by extracting features from the same sample data. The difference between the two is that unimodal features come from unimodal teacher models and can effectively preserve modality specificity, while comodal features come from multimodal teacher models and contain cross-modal co-operational semantics.

[0044] Taking a multimodal target re-identification scenario as an example, for the visible light modality, single-modal features and collaborative modal features of the visible light modality can be extracted using visible light images as sample data; for the near-infrared modality, single-modal features and collaborative modal features of the near-infrared modality can be extracted using near-infrared images as sample data; and for the thermal infrared modality, single-modal features and collaborative modal features of the thermal infrared modality can be extracted using thermal infrared images as sample data.

[0045] Step 120: Based on the sample data, and the single-modal features and / or the collaborative modal features, perform distillation training on the feature extraction branches of the modal in the multimodal student model.

[0046] Specifically, for each modality supported by the multimodal student model, after obtaining the unimodal and collaborative features of that modality, distillation training can be performed on the feature extraction branch of that modality in the multimodal student model based on these features. This distillation training can be implemented using a knowledge distillation mechanism, aiming to guide the feature extraction branch of that model in the multimodal student model to imitate and approximate the output feature representation of the teacher model.

[0047] Specifically, sample data can be input into the feature extraction branch corresponding to the modality in the multimodal student model to obtain student features for that modality extracted by the multimodal student model itself. Then, reference features are constructed using the unimodal and / or comodal features of that modality. The difference between the student features and the reference features is calculated to generate a distillation loss, and the parameters of the feature extraction branch for that modality in the multimodal student model are updated based on this distillation loss, thereby achieving distillation training for the feature extraction branch of that modality. Furthermore, to ensure the overall performance of the multimodal student model, in actual training, in addition to the distillation loss mentioned above, joint optimization can also be performed using cross-entropy loss or triplet loss based on the model output results. This embodiment of the invention does not specifically limit this approach.

[0048] Taking a multimodal target re-identification scenario as an example, for the visible light modality, distillation training can be performed on the feature extraction branch of the multimodal student model based on visible light images and the single-modal and collaborative modal features of the visible light modality. Similarly, for the near-infrared modality, distillation training can be performed on the feature extraction branch of the multimodal student model based on near-infrared images and the single-modal and collaborative modal features of the near-infrared modality. Furthermore, for the thermal infrared modality, distillation training can be performed on the feature extraction branch of the multimodal student model based on thermal infrared images and the single-modal and collaborative modal features of the thermal infrared modality. Thus, the feature extraction branch of each modality in the multimodal student model can fully learn multimodal collaborative modeling information and modality-specific information.

[0049] It should be noted that, for each modality supported by the multimodal student model, the reference features used during distillation training of the feature extraction branch for that modality can be unimodal features and / or comodal features. For example, unimodal and comodal features can be directly fused as reference features, thereby ensuring that the feature extraction branch can simultaneously absorb modality-specific details provided by the unimodal teacher model and cross-modal collaborative discriminative information provided by the multimodal teacher model during distillation training, thus achieving the best balance between multimodal collaborative modeling information and modality-specific information preservation. Alternatively, the reliability of feature extraction from sample data by the unimodal and multimodal teacher models can be assessed, and highly reliable modality features can be selected as reference features while low-reliability modality features are discarded. For example, unimodal features can be used as reference features while comodal features are discarded, or comodal features can be used as reference features while unimodal features are discarded, thereby ensuring that the feature extraction branch can selectively and reasonably absorb the knowledge provided by the teacher model during distillation training.

[0050] That is, in the embodiments of the present invention, the distillation training for the feature extraction branches of each modality can dynamically determine whether to select a more reliable single feature from single-modal features and collaborative modal features as the reference feature, or to fuse the two as the reference feature. Through this flexible distillation training mechanism, the multimodal student model can be guided to learn reasonably and fully, avoiding being misled by unreliable teacher knowledge or contradictory knowledge.

[0051] In the model training method provided in this embodiment of the invention, two types of teacher knowledge sources are set up in each modality: a unimodal teacher model and a multimodal teacher model. Based on the unimodal features and / or collaborative modal features extracted by the two types of teacher models, the feature extraction branch of the multimodal student model is distilled and trained. This enables the multimodal student model to more reasonably and fully absorb the modality-specific detailed information provided by the unimodal teacher and the cross-modal collaborative discrimination information provided by the multimodal teacher. This achieves the best balance between multimodal collaborative modeling information and modality-specific information preservation, significantly improving the training stability, robustness and overall performance of the multimodal student model in complex scenarios.

[0052] Based on the above embodiments, step 120, which involves performing distillation training on the feature extraction branches of the modalities in the multimodal student model based on the sample data, the single-modal features, and / or the collaborative modal features, includes: Based on the similarity between the unimodal features and the similarity between the comodal features, the differences in sample relationship cognition between the unimodal teacher model and the multimodal teacher model are determined. Based on the cognitive differences in sample relationships, the sample data, and the single-modal features and / or the collaborative modal features, the feature extraction branches of the modal in the multimodal student model are subjected to distillation training.

[0053] Specifically, in the training process of a multimodal student model, sample data is usually input in batches. For different sample data within the current batch, the unimodal features can be obtained through the unimodal teacher model, and the comodal features can be obtained through the multimodal teacher model.

[0054] Based on this, for the single-modal features of different sample data within the current batch, the similarity between the single-modal features of different sample data can be calculated; and for the collaborative modal features of different sample data within the current batch, the similarity between the collaborative modal features of different sample data can be calculated. It can be understood that the similarity between single-modal features reflects the distribution relationship of different sample data in the single-modal feature space; the similarity between collaborative modal features reflects the distribution relationship of different sample data in the collaborative modal feature space.

[0055] For example, the similarity between the unimodal features of every two samples within the current batch can be calculated, thereby constructing a similarity matrix describing the relationship structure of the unimodal teacher model with respect to the samples in the current batch. Similarly, the similarity between the comodal features of every two samples within the current batch can be calculated, thereby constructing a similarity matrix describing the relationship structure of the multimodal teacher model with respect to the samples in the current batch.

[0056] By calculating the difference between the two similarity matrices, the difference in perceived sample relationships can be quantified. Here, the difference in perceived sample relationships refers to the degree of disagreement between the two types of teacher models regarding the internal correlation structure between sample data in the same batch. Understandably, a smaller difference between the similarity matrices indicates that the sample relationships formed by the two types of teacher models in the current batch are similar, thus confirming that the knowledge provided by the two types of teacher models does not conflict significantly, and that the two types of teacher models have high synergy. Conversely, a larger difference between the similarity matrices indicates a significant disagreement between the two types of teacher models regarding the sample relationships in the current batch.

[0057] After quantifying the differences in sample relationship cognition, these differences can be used as the basis for dynamically adjusting the distillation training strategy. Unlike related technologies that blindly and fixedly use a single teacher model or a simple average multi-teacher model, this invention determines whether the knowledge provided by the two teacher models is synergistic based on the differences in sample relationship cognition between the two types of teacher models. This allows for a further determination of whether both single-modal features and synergistic modal features should be used to determine the reference features, or whether only single-modal features or synergistic modal features should be used.

[0058] For example, when differences in sample relationship cognition indicate a high degree of collaboration between the two teacher models, unimodal and collaborative modal features can be integrated as reference features. This guides the feature extraction branch of the multimodal student model to simultaneously absorb modality-specific details provided by the unimodal teacher and cross-modal collaborative discrimination information provided by the multimodal teacher. When differences in sample relationship cognition indicate a significant divergence between the two teacher models, a screening mechanism is triggered to select highly reliable features from the unimodal and collaborative modal features as reference features. This ensures that the reference features provided for the feature extraction branch of the multimodal student model are always stable and reasonable, thereby achieving adaptive knowledge distillation.

[0059] In the model training method provided in this embodiment of the invention, by calculating the similarity between unimodal features and the similarity between co-modal features, the difference in sample relationship cognition between unimodal teacher models and multimodal teacher models is evaluated in real time and explicitly from the perspective of sample relationship structure. This reflects the consistency and synergistic state of the knowledge provided by the two types of teacher models in the current training batch. Introducing the difference in sample relationship cognition into the distillation training process enables the distillation training to adaptively adjust according to the synergistic state between the two types of teacher models, effectively avoiding the risk of blind learning and negative transfer when the multimodal student model faces contradictory or conflicting teacher knowledge, and ensuring the stability and accuracy of model feature learning.

[0060] Based on any of the above embodiments, in step 120, the step of performing distillation training on the feature extraction branches of the modality in the multimodal student model based on the cognitive differences in sample relationships, the sample data, and the single-modal features and / or the collaborative modal features includes: If the difference in the perceived relationship between the samples is greater than or equal to a first threshold, the single-modal feature or the collaborative modal feature is determined as a reference feature based on the reliability of the single-modal teacher model and the reliability of the multimodal teacher model; the reliability is determined based on the modal feature similarity of positive sample pairs and / or the modal feature similarity of negative sample pairs in the sample data. Based on the sample data and the reference features, the feature extraction branches of the modal in the multimodal student model are trained by distillation.

[0061] Specifically, the first threshold is a pre-set boundary value used to determine whether the two types of teacher models can collaborate. After obtaining the differences in sample relationship cognition, the magnitude of the differences in sample relationship cognition can be compared with the first threshold.

[0062] Specifically, a difference in the perceived sample relationships greater than or equal to the first threshold indicates a significant difference in the sample relationships formed by the two teacher models within the current batch. This suggests a discrepancy in their understanding of the same batch of sample data, indicating a lack of collaboration between the two teacher models. In this case, directly fusing or simultaneously distilling the unimodal and collaborative modal features output by each of the two teacher models may introduce incorrect supervision signals into the multimodal student model, misleading its feature learning.

[0063] To resolve the conflict between the two types of teacher models, this embodiment of the invention calculates the reliability of each of the two types of teacher models, thereby achieving an explicit evaluation of the discriminative ability of the two types of teacher models.

[0064] Here, reliability is used to evaluate the reliability of the modal features output by the teacher model. Reliability can be determined based on the similarity of modal features between positive and negative sample pairs in the sample data.

[0065] In this context, a positive sample pair refers to a pair of sample data belonging to the same identity or category of target within the current batch; a negative sample pair refers to a pair of sample data belonging to different identities or categories of target within the current batch. For example, in a target re-identification scenario, different images of the same pedestrian can form a positive sample pair, while images of different pedestrians can form a negative sample pair.

[0066] Understandably, a reliable teacher model should enable the extracted modal features to exhibit the property of clustering similar features and repelling dissimilar features. Therefore, the reliability of a teacher model can be evaluated by calculating the similarity between the modal features output by the teacher model for all positive sample pairs, and / or the similarity between the modal features output by the teacher model for all negative sample pairs. Specifically, the higher the similarity of modal features for positive sample pairs and the lower the similarity of modal features for negative sample pairs, the higher the reliability of the teacher model; conversely, the lower the similarity of modal features for positive sample pairs and the higher the similarity of modal features for negative sample pairs, the lower the reliability of the teacher model.

[0067] In some embodiments, the difference between the modal feature similarity of positive sample pairs and the modal feature similarity of negative sample pairs can be calculated as the reliability of the teacher model, specifically expressed as the following formula: In the formula, For the first The teacher model is for the first Reliability of batch sample data. For the first The teacher model is for the first Batch sample data and sample data The similarity of the output modal features. Denotes the set of positive sample pairs. , and For sample data and sample data Each with its own tags; Represents the set of negative sample pairs. .

[0068] Understandably, the higher the reliability of the teacher model, the better it can make similar samples closer and different samples farther apart, meaning that the teacher model is more reliable in extracting modal features in the current batch.

[0069] After calculating the reliability of the unimodal and multimodal teacher models separately, the teacher model with higher reliability can be determined by comparing the reliability of the two models. The modal features output by the more reliable teacher model can then be used as reference features. For example, if the reliability of the unimodal teacher model is higher than that of the multimodal teacher model, the unimodal features can be used as reference features; conversely, if the reliability of the multimodal teacher model is higher than that of the unimodal teacher model, the comodal features can be used as reference features.

[0070] After this, the sample data can be input into the feature extraction branch of the multimodal student model to obtain the student features extracted by the multimodal student model. Subsequently, using the above-selected, more reliable reference features as the supervision target, a distillation loss function is constructed in combination with the student features, and backpropagation is performed based on this loss function, thereby realizing distillation training for the feature extraction branch.

[0071] In the model training method provided in this embodiment of the invention, when the difference in the cognitive relationship between samples is greater than or equal to a first threshold, a teacher model reliability evaluation mechanism based on the similarity of modal features of positive and negative sample pairs is introduced. Based on the reliability of the teacher model, a more reliable modal feature in the current batch is dynamically selected as the only reference feature for distillation training, thereby effectively avoiding the performance loss caused by the multimodal student model learning contradictory knowledge at the same time, and ensuring the recognition accuracy and generalization performance of the multimodal student model under complex or difficult samples.

[0072] Based on any of the above embodiments, step 120, which involves performing distillation training on the feature extraction branch of the modality in the multimodal student model based on the sample data and the reference features, includes: If the difference between the reliability of the unimodal teacher model and the reliability of the multimodal teacher model is greater than or equal to the second threshold, the feature extraction branch is distilled and trained with the goal of making the features of the sample data extracted by the feature extraction branch close to the reference features and far away from the counterexample features. The counterexample feature is a modal feature that is different from the reference feature.

[0073] Specifically, the second threshold is a pre-defined criterion used to measure the degree of difference in reliability between the two types of teacher models. When the difference in the perceived sample relationship between the two types of teacher models is greater than or equal to the first threshold, that is, when the two types of teacher models are in a non-cooperative state, the gap in reliability between the two types of teacher models can be further calculated, and the magnitude of this gap can be compared with the second threshold.

[0074] In this case, if the difference in reliability between the two types of teacher models is greater than or equal to the second threshold, it means that the modal features provided by the low-reliability teacher model have significant bias and unreliability. In this situation, the modal features output by the low-reliability teacher model can be used as negative example features. It can be understood that negative example features are the modal features that were not selected as reference features among the unimodal and comodal features. For example, if the reliability of the unimodal teacher model is much higher than that of the multimodal teacher model, then the unimodal feature is determined as the reference feature, and the comodal feature is determined as the negative example feature; conversely, if the reliability of the multimodal teacher model is much higher than that of the unimodal teacher model, then the comodal feature is determined as the reference feature, and the unimodal feature is determined as the negative example feature.

[0075] In response to the significant bias and unreliability of modal features provided by low-reliability teacher models, this invention introduces a low-reliability knowledge suppression mechanism. Thus, the distillation training of the feature extraction branch of the multimodal student model aims not only to integrate the extracted student features with the reference features, but also to keep the extracted student features away from counterexample features.

[0076] Specifically, the objectives of distillation training can include the following two parts: One part is used to achieve positive guidance, that is, with the goal of the features of the sample data extracted by the feature extraction branch being close to the reference features, the alignment loss, such as mean square error, is calculated to make the student features continuously approach the feature distribution of the high-reliability teacher model. Another part is used to implement reverse suppression, which aims to prevent student features from approaching the feature distribution of an unreliable teacher model by calculating knowledge suppression loss, with the goal of keeping the features of the sample data extracted by the feature extraction branch far away from the negative example features. It can be understood that the closer the features of the sample data extracted by the feature extraction branch are to the negative example features, the higher the knowledge suppression loss; conversely, the further the features of the sample data extracted by the feature extraction branch are from the negative example features, the lower the knowledge suppression loss.

[0077] In some embodiments, a safety distance can be set. When calculating the knowledge inhibition loss, if the distance between the student feature and the negative example feature is less than this safety distance, a knowledge inhibition loss is incurred as a penalty; if the distance between the student feature and the negative example feature is greater than or equal to this safety distance, the knowledge inhibition loss is defaulted to zero, and no penalty is applied. For example, the knowledge inhibition loss can be expressed as the following formula: In the formula, In order to target the Knowledge suppression loss of batch sample data. For safe distance. For the first Batch sample data Based on multimodal student model Extracted student characteristics For the first Batch sample data Teacher model based on low reliability Extracted counterexample features.

[0078] By combining these two objectives, a joint loss function can be used to backpropagate the feature extraction branch and update the network parameters, thereby guiding the multimodal student model to actively avoid low-reliability knowledge while absorbing highly reliable knowledge.

[0079] In the model training method provided in this embodiment of the invention, when the reliability difference between the two types of teacher models is greater than or equal to a second threshold, distillation training is introduced with the goal of approaching the reference features and staying away from the counterexample features. Thus, in the distillation training process of the feature extraction branch, not only is the effective absorption of high-reliability knowledge achieved, but also the explicit knowledge suppression mechanism effectively avoids the feature extraction branch from approaching the low-reliability modal feature distribution in the feature space. This avoids the serious misleading effect of low-quality supervision signals on the multimodal student model, and further improves the robustness of the multimodal student model in extreme or complex scenarios.

[0080] Based on any of the above embodiments, in step 120, the step of performing distillation training on the feature extraction branches of the modality in the multimodal student model based on the cognitive differences in sample relationships, the sample data, and the single-modal features and / or the collaborative modal features includes: If the difference in the perceived relationship between the samples is less than a first threshold, the single-modal features and the collaborative modal features are fused to obtain fused features; Based on the sample data and the fusion features, the feature extraction branches of the modal in the multimodal student model are trained by distillation.

[0081] Specifically, if the difference in the perceived sample relationships is less than the first threshold, it means that the two types of teacher models have relatively small differences in the sample relationships formed within the current batch, and the knowledge they provide does not contradict or conflict, so they can be used together in the subsequent distillation process. In other words, the two types of teacher models are in a synergistic state.

[0082] In this context, unimodal features output by unimodal teachers can be combined with collaborative modal features output by multimodal teachers to construct fused features that incorporate the strengths of both teacher models. For example, unimodal and collaborative modal features can be directly concatenated, averaged, or added together to obtain fused features. Alternatively, unimodal and collaborative modal features can be assigned corresponding weights to achieve weighted fusion, thus yielding fused features.

[0083] After obtaining the fused features, they can be used as reference features for distillation training of the corresponding feature extraction branch in the multimodal student model. The current batch of sample data can be input into the corresponding feature extraction branch in the multimodal student model to obtain the student features output by that branch. Subsequently, by aligning the student features with the aforementioned fused features, the feature extraction branch in the multimodal student model can directly learn the knowledge provided by the two highly consistent teacher models after collaborative evaluation.

[0084] In the model training method provided in this embodiment of the invention, when the difference in sample relationship cognition is less than a first threshold, it is determined that the unimodal teacher model and the multimodal teacher model are in a collaborative state. The fusion features of the two are then obtained as reference features for distillation training, thereby ensuring that the multimodal student model can learn both the modality-specific details of the unimodal features and the cross-modal global collaborative expression of the collaborative modal features. Under the premise that the knowledge of the two types of teacher models does not conflict, the maximum utilization of information is achieved, which greatly improves the multimodal student model's ability to balance multimodal collaborative modeling information and modality-specific information retention, as well as its overall feature expression ability.

[0085] Based on any of the above embodiments, step 120, which involves fusing the single-modal features and the collaborative modal features to obtain the fused features, includes: The fusion weights are determined based on the reliability of the unimodal teacher model and the reliability of the multimodal teacher model; the reliability is determined based on the modal feature similarity of positive sample pairs and / or the modal feature similarity of negative sample pairs in the sample data. Based on the fusion weights, the single-modal features and the collaborative modal features are weighted and fused to obtain the fused features.

[0086] Specifically, in order to further ensure the reliability of the fused features, the fusion weights required for weighted fusion of single-modal features and collaborative modal features can be determined based on the reliability of the two types of teacher models.

[0087] Here, the reliability of the unimodal teacher model can be determined based on the unimodal feature similarity of positive sample pairs and / or the unimodal feature similarity of negative sample pairs in the sample data; the reliability of the multimodal teacher model can be determined based on the collaborative modal feature similarity of positive sample pairs and / or the collaborative modal feature similarity of negative sample pairs in the sample data. The reliability calculation methods for the above two types of teacher models can refer to the methods provided in the above embodiments, and will not be repeated here.

[0088] The fusion weights reflect the proportion of each teacher model's output modal features in the final fused features. Understandably, among the two types of teacher models, the modal features output by the teacher model with higher reliability have higher fusion weights, while those output by the teacher model with lower reliability have lower fusion weights. For example, the fusion weights can be obtained by normalizing the reliability of the two types of teacher models using the Softmax function.

[0089] Based on this, single-modal features can be multiplied by their corresponding fusion weights, and collaborative modal features can be multiplied by their corresponding fusion weights. Then, the two are added element-wise to obtain the fused features. For example, the fused features can be expressed as the following formula: In the formula, In order to target the Batch sample data The fusion characteristics. For the first Batch sample data The single-modal characteristics, For the first Batch sample data The collaborative modal characteristics. For the first The fusion weights corresponding to the single-modal features of the batch. For the first The fusion weights corresponding to the collaborative modal features of the batch.

[0090] In the model training method provided in this embodiment of the invention, the reliability of the unimodal teacher model and the multimodal teacher model are combined to perform weighted fusion of unimodal features and collaborative modal features. This enables the multimodal student model to achieve the best balance between preserving multimodal collaborative modeling information and modality-specific information when it subsequently receives distillation training based on the fused features. This further improves the training stability, robustness and overall performance of the multimodal student model in complex scenarios.

[0091] Based on any of the above embodiments, the multimodal student model, the unimodal teacher model, and the multimodal teacher model are all target re-identification models, and the modality is a spectral modality.

[0092] Specifically, in a multimodal target re-identification scenario, to train the target re-identification model, the multimodal student model, the unimodal teacher model, and the multimodal teacher model can all be configured as models to perform the target re-identification task. The modality referred to here can be limited to spectral modalities, specifically including visible light modalities, near-infrared modalities, thermal infrared modalities, etc.

[0093] In this scenario, the unimodal teacher model for each spectral modality is responsible for extracting fine-grained unimodal features that focus on distinguishing different target identities under specific spectral modalities; the multimodal teacher model is responsible for extracting collaborative modal features that are cross-modal and have global identity discrimination capabilities; and the multimodal student model learns the above-mentioned unimodal and collaborative modal features through distillation training, and finally outputs a highly robust feature vector for target retrieval and matching.

[0094] During the model inference stage, it is only necessary to use the multimodal student model that has been distilled to extract features from the input target and complete target matching and retrieval based on the extracted feature representation to achieve target re-identification in complex environments.

[0095] Based on any of the above embodiments Figure 2 This is the second flowchart illustrating the model training method provided by this invention, as shown below. Figure 2 As shown, the method includes: First, input multimodal data to obtain sample data.

[0096] Specifically, it can obtain sample data for each modality in the current batch.

[0097] Secondly, construct unimodal teacher models and multimodal teacher models.

[0098] Specifically, for each modality of the multimodal student model to be trained, a unimodal teacher model and a multimodal teacher model corresponding to that modality can be constructed.

[0099] Next, single-modal features and collaborative modal features are extracted.

[0100] Specifically, for each modality of the multimodal student model to be trained, the sample data of that modality can be input into the corresponding unimodal teacher model and multimodal teacher model, respectively. After forward processing by the two types of teacher models, the unimodal features extracted by the unimodal teacher model and the collaborative modal features extracted by the multimodal teacher model can be obtained.

[0101] Subsequently, the differences and reliability of the sample relationship perceptions were assessed.

[0102] Specifically, for each modality of the multimodal student model to be trained, the similarity between each unimodal feature and the similarity between each comodal feature within the current batch are calculated. Then, based on the similarity between unimodal features and the similarity between comodal features, the differences in sample relationship cognition between the unimodal teacher model and the multimodal teacher model are determined.

[0103] Furthermore, using the sample labels within the current batch, the reliability of the unimodal teacher model is calculated based on the unimodal feature similarity of positive sample pairs and / or the unimodal feature similarity of negative sample pairs in the sample data; and the reliability of the multimodal teacher model is calculated based on the collaborative modal feature similarity of positive sample pairs and / or the collaborative modal feature similarity of negative sample pairs in the sample data.

[0104] Next, we will determine whether the two types of teacher models can work together.

[0105] Specifically, for each modality of the multimodal student model to be trained, the differences in sample relationship cognition can be compared with a preset first threshold to determine whether the two types of teacher models can collaborate.

[0106] For cases where the two types of teacher models can collaborate, unimodal features and collaborative modal features are integrated.

[0107] Specifically, when the difference in sample relationship perception is less than a first threshold, the two types of teacher models are determined to be in a collaborative state. At this point, fusion weights can be determined based on the reliability of the unimodal teacher model and the multimodal teacher model. Then, based on the fusion weights, the unimodal features and collaborative modal features are weighted and fused to obtain fused features. Under this branch, the fused features can be identified as reference features for guiding training.

[0108] To address the issue of non-cooperation between the two types of teacher models, high-reliability modal features are referenced while low-reliability modal features are suppressed.

[0109] Specifically, when the difference in the perceived relationship between samples is greater than or equal to a first threshold, the two types of teacher models are determined to be in a non-cooperative state. At this point, the reliability of the two types of teacher models is compared, and the modal features output by the teacher model with high reliability are determined as reference features, while the modal features output by the teacher model with low reliability are determined as counterexample features that need to be suppressed.

[0110] Next, multimodal student model distillation training is performed.

[0111] Specifically, for cases where the two types of teacher models can collaborate, the feature extraction branch of that modality in the multimodal student model can be distilled and trained based on sample data and fused features. For cases where the two types of teacher models cannot collaborate, if the reliability difference between the two types of teacher models is less than a second threshold, the feature extraction branch of that modality in the multimodal student model can be distilled and trained based on sample data and reference features; if the reliability difference between the two types of teacher models is greater than or equal to the second threshold, the feature extraction branch can be distilled and trained with the goal of ensuring that the features extracted by the feature extraction branch are close to the reference features and far away from the negative example features.

[0112] Finally, the inference output.

[0113] Specifically, the multimodal student model is trained through one or more batches of iterative distillation training. During the actual inference phase, inference output can be performed based on the trained multimodal student model.

[0114] In the model training method provided in this embodiment of the invention, two types of teacher models are introduced simultaneously as knowledge sources for each modality. The unimodal teacher model provides modality-specific knowledge to preserve modality-specific details, while the multimodal teacher model provides multimodal collaborative knowledge to preserve collaborative discrimination information after multimodal joint training. Therefore, the multimodal student model can fully and reasonably learn both unimodal detailed information and multimodal collaborative information.

[0115] Furthermore, to avoid inconsistent teacher knowledge interfering with the multimodal student model, this embodiment of the invention determines whether the two types of teacher models can collaborate, and when they are not collaborative, selects the modal features output by the more reliable teacher model as the reference features. This avoids the multimodal student model blindly learning inconsistent or unreliable teacher knowledge, reducing the risk of negative transfer.

[0116] Furthermore, in this embodiment of the invention, a reference teacher target is dynamically generated based on the synergy and reliability of the two types of teacher models in the current batch, thereby ensuring that the multimodal student model can learn stable supervision signals after screening and integration, rather than being driven by multiple potentially inconsistent modal features at the same time.

[0117] Finally, by conducting collaborative evaluation of the two types of teacher models, adaptive reference teacher target generation, feature extraction branch distillation for each modality, and suppressing low-reliability teacher knowledge, the method provided in this embodiment of the invention helps to improve the robustness of multimodal student models in multimodal tasks, enabling them to have better robustness and generalization performance in complex scenarios.

[0118] Figure 3This is a flowchart illustrating the training method of the multispectral target re-identification model provided by the present invention, as shown below. Figure 3 As shown, the method includes: First, multispectral sample images are input into the unimodal teacher model and the multimodal teacher model to obtain the unimodal features and comodal features of each spectral mode.

[0119] Here, multispectral modes include visible light modes, near-infrared modes, and thermal infrared modes. Single-mode teacher models can be set for each of these modes, resulting in visible light single-mode teacher models, near-infrared single-mode teacher models, and thermal infrared single-mode teacher models, respectively. This allows for the acquisition of visible light single-mode features, near-infrared single-mode features, and thermal infrared single-mode features, respectively. Furthermore, multi-mode teacher models can be set for each of these modes. These multi-mode teacher models can extract features for each of the visible light, near-infrared, and thermal infrared modes and output corresponding co-modal features, namely visible light co-modal features, near-infrared co-modal features, and thermal infrared co-modal features.

[0120] Secondly, synergy assessment is performed based on single-modal features and cooperative modal features.

[0121] Here, the synergy assessment is performed separately for each mode. That is, the synergy assessment of the visible light mode can be performed based on the visible light single-mode characteristics and the visible light synergy mode characteristics; the synergy assessment of the near-infrared mode can be performed based on the near-infrared single-mode characteristics and the near-infrared synergy mode characteristics; and the synergy assessment of the thermal infrared mode can be performed based on the thermal infrared single-mode characteristics and the thermal infrared synergy mode characteristics.

[0122] For each modality, the synergy assessment includes assessing differences in sample relationship cognition, evaluating the reliability of the teacher model, and determining the state of synergy.

[0123] Specifically, assessing the differences in sample relationship perception involves calculating the similarity matrix between single-modal features and the similarity matrix between collaborative modal features based on the sample images of each modality within the current batch. By comparing the differences between these two similarity matrices, the differences in sample relationship perception between the single-modal teacher model and the multimodal teacher model can be calculated. Thus, the differences in sample relationship perception for the visible light modality, the near-infrared modality, and the thermal infrared modality can be obtained.

[0124] To assess the reliability of the teacher model, specifically, the modal feature similarity of positive sample pairs and negative sample pairs in the sample images of each modality within the current batch can be used to calculate the reliability of the single-modal teacher model and the multi-modal teacher model in that modality. This allows us to obtain the reliability of the two types of teacher models in the visible light modality, the near-infrared modality, and the thermal infrared modality.

[0125] The collaborative state determination can specifically involve comparing the calculated differences in sample relationship cognition with a first threshold to determine whether the two types of teacher models are in a collaborative or non-collaborative state. Here, collaborative state determination for two types of teacher models in the visible light mode, the near-infrared mode, and the thermal infrared mode can be implemented.

[0126] Next, distillation training was performed on the multimodal student model.

[0127] For each modality, if the two teacher models are in a collaborative state, perform A. Feature Fusion. That is, when the difference in sample relationship cognition is less than a first threshold, determine the fusion weights based on the reliability of the unimodal teacher model and the reliability of the multimodal teacher model, and then perform a weighted fusion of the unimodal features and collaborative modal features based on the fusion weights to obtain the fused features. For each modality, if the two teacher models are in a non-cooperative state, proceed to step B: Select a reference feature. That is, when the difference in sample relationship perception is greater than or equal to a first threshold, a reliable teacher model selection mechanism is executed. Specifically, the reliability of the unimodal teacher model is directly compared with the reliability of the multimodal teacher model. If the reliability of the unimodal teacher model is higher, the unimodal feature is determined as the reference feature; if the reliability of the multimodal teacher model is higher, the cooperative modality feature is determined as the reference feature.

[0128] Next, perform C. Branch distillation based on fused features or reference features. That is, fused features or reference features can be used as the reference teacher target to construct an alignment loss function, and then perform distillation training on the feature extraction branch of that modality in the multimodal student model.

[0129] Finally, the trained multispectral target re-identification model is obtained.

[0130] Here, while the feature extraction branches are being trained by distillation, the student features output by each feature extraction branch of each modality can be fused together to obtain a feature representation for target matching. This feature representation is then used for regular re-identification loss calculation to achieve overall joint optimization for the multimodal student model.

[0131] When the model completes training and enters the inference phase, the multimodal student model can be directly regarded as a trained multispectral target re-identification model. The input data can be processed based on the feature extraction branches in the multispectral target re-identification model, and the features output by each branch can be fused to finally output a feature representation for target matching, thereby completing the re-identification retrieval task.

[0132] In the method provided in the embodiments of the present invention, by conducting collaborative evaluation of two types of teacher models, adaptive reference teacher target generation, feature extraction branch distillation for each modality, and suppressing low-reliability teacher knowledge, it is possible to improve the robust discrimination ability of multimodal student models in multispectral target re-identification tasks, and enable them to have better robustness and generalization performance in scenarios such as complex lighting, occlusion, thermal noise, and large modal differences.

[0133] The model training apparatus provided by the present invention is described below. The model training apparatus described below and the model training method described above can be referred to in correspondence.

[0134] Figure 4 This is a schematic diagram of the model training device provided by the present invention, as shown below. Figure 4 As shown, the device includes: The feature extraction module 410 is used to input the sample data of each modality of the multimodal student model into the corresponding unimodal teacher model and multimodal teacher model respectively, to obtain the unimodal features extracted by the unimodal teacher model and the collaborative modal features extracted by the multimodal teacher model. The distillation training module 420 is used to perform distillation training on the feature extraction branches of the modal in the multimodal student model based on the sample data, the single-modal features and / or the collaborative modal features.

[0135] In the model training device provided in this embodiment of the invention, two types of teacher knowledge sources are set up in each modality: a unimodal teacher model and a multimodal teacher model. Based on the unimodal features and / or collaborative modal features extracted by the two types of teacher models, the feature extraction branch of the multimodal student model is distilled and trained. This enables the multimodal student model to more reasonably and fully absorb the modality-specific detailed information provided by the unimodal teacher and the cross-modal collaborative discrimination information provided by the multimodal teacher. This achieves the best balance between multimodal collaborative modeling information and modality-specific information preservation, significantly improving the training stability, robustness and overall performance of the multimodal student model in complex scenarios.

[0136] Based on any of the above embodiments, the distillation training module is specifically used for: Based on the similarity between the unimodal features and the similarity between the comodal features, the differences in sample relationship cognition between the unimodal teacher model and the multimodal teacher model are determined. Based on the cognitive differences in sample relationships, the sample data, and the single-modal features and / or the collaborative modal features, the feature extraction branches of the modal in the multimodal student model are subjected to distillation training.

[0137] Based on any of the above embodiments, the distillation training module is specifically used for: If the difference in the perceived relationship between the samples is greater than or equal to a first threshold, the single-modal feature or the collaborative modal feature is determined as a reference feature based on the reliability of the single-modal teacher model and the reliability of the multimodal teacher model; the reliability is determined based on the modal feature similarity of positive sample pairs and / or the modal feature similarity of negative sample pairs in the sample data. Based on the sample data and the reference features, the feature extraction branches of the modal in the multimodal student model are trained by distillation.

[0138] Based on any of the above embodiments, the distillation training module is specifically used for: If the difference between the reliability of the unimodal teacher model and the reliability of the multimodal teacher model is greater than or equal to the second threshold, the feature extraction branch is distilled and trained with the goal of making the features of the sample data extracted by the feature extraction branch close to the reference features and far away from the counterexample features. The counterexample feature is a modal feature that is different from the reference feature.

[0139] Based on any of the above embodiments, the distillation training module is specifically used for: If the difference in the perceived relationship between the samples is less than a first threshold, the single-modal features and the collaborative modal features are fused to obtain fused features; Based on the sample data and the fusion features, the feature extraction branches of the modal in the multimodal student model are trained by distillation.

[0140] Based on any of the above embodiments, the distillation training module is specifically used for: The fusion weights are determined based on the reliability of the unimodal teacher model and the reliability of the multimodal teacher model; the reliability is determined based on the modal feature similarity of positive sample pairs and / or the modal feature similarity of negative sample pairs in the sample data. Based on the fusion weights, the single-modal features and the collaborative modal features are weighted and fused to obtain the fused features.

[0141] Based on any of the above embodiments, the multimodal student model, the unimodal teacher model, and the multimodal teacher model are all target re-identification models, and the modality is a spectral modality.

[0142] Figure 5 An example is a schematic diagram of the physical structure of an electronic device, such as... Figure 5 As shown, the electronic device may include: a processor 510, a communications interface 520, a memory 530, and a communication bus 540, wherein the processor 510, the communications interface 520, and the memory 530 communicate with each other via the communication bus 540. The processor 510 can call logical instructions in the memory 530 to execute a model training method, which includes: For each modality of the multimodal student model, the sample data of the modality is input into the corresponding unimodal teacher model and multimodal teacher model to obtain the unimodal features extracted by the unimodal teacher model and the collaborative modal features extracted by the multimodal teacher model. Based on the sample data, as well as the single-modal features and / or the collaborative modal features, the feature extraction branches of the modal in the multimodal student model are subjected to distillation training.

[0143] Furthermore, the logical instructions in the aforementioned memory 530 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to related technologies, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0144] On the other hand, the present invention also provides a computer program product, the computer program product comprising a computer program that can be stored on a non-transitory computer-readable storage medium, wherein when the computer program is executed by a processor, the computer is able to execute the model training method provided by the above methods, the method comprising: For each modality of the multimodal student model, the sample data of the modality is input into the corresponding unimodal teacher model and multimodal teacher model to obtain the unimodal features extracted by the unimodal teacher model and the collaborative modal features extracted by the multimodal teacher model. Based on the sample data, as well as the single-modal features and / or the collaborative modal features, the feature extraction branches of the modal in the multimodal student model are subjected to distillation training.

[0145] In another aspect, the present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, is implemented to perform the model training methods provided by the methods described above, the method comprising: For each modality of the multimodal student model, the sample data of the modality is input into the corresponding unimodal teacher model and multimodal teacher model to obtain the unimodal features extracted by the unimodal teacher model and the collaborative modal features extracted by the multimodal teacher model. Based on the sample data, as well as the single-modal features and / or the collaborative modal features, the feature extraction branches of the modal in the multimodal student model are subjected to distillation training.

[0146] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.

[0147] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the parts that contribute to the related technology, can be embodied in the form of software products. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.

[0148] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A model training method, characterized in that, include: For each modality of the multimodal student model, the sample data of the modality is input into the corresponding unimodal teacher model and multimodal teacher model to obtain the unimodal features extracted by the unimodal teacher model and the collaborative modal features extracted by the multimodal teacher model. Based on the sample data, as well as the single-modal features and / or the collaborative modal features, the feature extraction branches of the modal in the multimodal student model are subjected to distillation training.

2. The model training method according to claim 1, characterized in that, The step of performing distillation training on the feature extraction branches of the modalities in the multimodal student model based on the sample data, the single-modal features, and / or the collaborative modal features includes: Based on the similarity between the unimodal features and the similarity between the comodal features, the differences in sample relationship cognition between the unimodal teacher model and the multimodal teacher model are determined. Based on the cognitive differences in sample relationships, the sample data, and the single-modal features and / or the collaborative modal features, the feature extraction branches of the modal in the multimodal student model are subjected to distillation training.

3. The model training method according to claim 2, characterized in that, The step of performing distillation training on the feature extraction branches of the multimodal student model based on the cognitive differences in sample relationships, the sample data, and the single-modal features and / or the collaborative modal features includes: If the difference in the perceived relationship between the samples is greater than or equal to a first threshold, the single-modal feature or the collaborative modal feature is determined as a reference feature based on the reliability of the single-modal teacher model and the reliability of the multimodal teacher model; the reliability is determined based on the modal feature similarity of positive sample pairs and / or the modal feature similarity of negative sample pairs in the sample data. Based on the sample data and the reference features, the feature extraction branches of the modal in the multimodal student model are trained by distillation.

4. The model training method according to claim 3, characterized in that, The step of performing distillation training on the feature extraction branch of the multimodal student model based on the sample data and the reference features includes: If the difference between the reliability of the unimodal teacher model and the reliability of the multimodal teacher model is greater than or equal to the second threshold, the feature extraction branch is distilled and trained with the goal of making the features of the sample data extracted by the feature extraction branch close to the reference features and far away from the counterexample features. The counterexample feature is a modal feature that is different from the reference feature.

5. The model training method according to claim 2, characterized in that, The step of performing distillation training on the feature extraction branches of the multimodal student model based on the cognitive differences in sample relationships, the sample data, and the single-modal features and / or the collaborative modal features includes: If the difference in the perceived relationship between the samples is less than a first threshold, the single-modal features and the collaborative modal features are fused to obtain fused features; Based on the sample data and the fusion features, the feature extraction branches of the modal in the multimodal student model are trained by distillation.

6. The model training method according to claim 5, characterized in that, The fusion of the single-modal features and the collaborative modal features to obtain fused features includes: The fusion weights are determined based on the reliability of the unimodal teacher model and the reliability of the multimodal teacher model; the reliability is determined based on the modal feature similarity of positive sample pairs and / or the modal feature similarity of negative sample pairs in the sample data. Based on the fusion weights, the single-modal features and the collaborative modal features are weighted and fused to obtain the fused features.

7. The model training method according to any one of claims 1 to 6, characterized in that, The multimodal student model, the unimodal teacher model, and the multimodal teacher model are all target re-identification models, and the modality is a spectral modality.

8. A model training device, characterized in that, include: The feature extraction module is used to input the sample data of each modality of the multimodal student model into the corresponding unimodal teacher model and multimodal teacher model respectively, to obtain the unimodal features extracted by the unimodal teacher model and the collaborative modal features extracted by the multimodal teacher model. The distillation training module is used to perform distillation training on the feature extraction branches of the modalities in the multimodal student model based on the sample data, as well as the single-modal features and / or the collaborative modal features.

9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that, When the processor executes the computer program, it implements the model training method as described in any one of claims 1 to 7.

10. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the model training method as described in any one of claims 1 to 7.

11. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the model training method as described in any one of claims 1 to 7.