Cross-modal knowledge model training method and system, and cross-modal knowledge model identification method and system

By introducing two-way knowledge migration and dynamic distillation strategies into the knowledge distillation network model, the modal gap problem in cross-modal knowledge distillation is solved, the model performance and knowledge diversity are improved, and the degree of personalization of the model is enhanced.

CN120338039APending Publication Date: 2025-07-18THE HONG KONG UNIV OF SCI & TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510337370.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2024-03-25
Filing Date
2025-03-21
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

The existing knowledge distillation technology has caused the modal gap problem caused by modal data imbalance during the cross-modal knowledge distillation process, resulting in a degradation of model performance and insufficient knowledge diversity.

Method used

By introducing a two-way knowledge transfer mechanism into the knowledge distillation network model, the two-way interaction between the teacher model and the student model is used to migrate the first soft label and the second soft label, combining dynamic distillation strategy and dual-agent two-way distillation strategy, the modal gap problem is alleviated and the model performance is improved.

Benefits of technology

It effectively suppresses the modal gap, improves the model performance and knowledge diversity of the cross-modal knowledge model, and enhances the degree of personalization of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120338039A_ABST
    Figure CN120338039A_ABST
Patent Text Reader

Abstract

The invention discloses a cross-modal knowledge model training method, identification method and system, and the method comprises the steps: inputting the medium data of a first modal into a teacher model for knowledge extraction, obtaining a first soft label, inputting the medium data of a second modal into a student model for knowledge extraction, and obtaining a second soft label; performing bidirectional knowledge migration processing according to the first soft label and the second soft label to obtain a teacher migration loss value and a student migration loss value; according to the teacher migration loss value and the student migration loss value, performing parameter updating on the knowledge distillation network model to obtain a trained knowledge distillation network model; and extracting the cross-modal knowledge model from the trained knowledge distillation network model. According to the method, the modal gap problem caused by modal data imbalance in the knowledge distillation process can be inhibited, and the model performance of the cross-modal knowledge model can be improved. The invention relates to the technical field of deep learning.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross-reference to related patent applications

[0002] This invention claims priority to U.S. Provisional Patent Application No. 63 / 569,220, filed on March 25, 2024, the entire content of which is incorporated herein by reference. Technical Field

[0003] This application relates to the field of deep learning technology, and particularly relates to a training method, an identification method, and a system for a cross-modal knowledge model. Background Art

[0004] With the increasing popularity of multi-modal sensors, traditional knowledge distillation methods have been extended, which can achieve knowledge transfer of cross-modal data and play an important role in improving the performance of downstream tasks.

[0005] Currently, existing knowledge distillation (KD) techniques usually use a modality with high accuracy or high annotation quality as the teacher model to transfer knowledge to a modality with low accuracy or no annotation. However, due to the imbalance in the amount of information between different modality data, the modality gap that appears in the cross-modal knowledge distillation (CMKD) process will make existing knowledge distillation techniques ineffective in the cross-modal knowledge distillation process.

[0006] Therefore, the problems existing in the prior art still need to be solved and optimized urgently. Summary of the Invention

[0007] An object of the present invention is to solve at least to some extent one of the technical problems existing in the related art.

[0008] To this end, an object of an embodiment of the present invention is to provide a training method, an identification method, and a system for a cross-modal knowledge model. Among them, the training method can suppress the modality gap problem caused by the imbalance of modality data in the knowledge distillation process, and thus is beneficial to improving the model performance of the cross-modal knowledge model.

[0009] To achieve the above technical object, the technical solutions adopted in the embodiments of the present application include:

[0010] In a first aspect, an embodiment of the present application provides a training method for a cross-modal knowledge model, which is applied to a knowledge distillation network model. The knowledge distillation network model includes a teacher model and a student model. The method includes:

[0011] Obtain media data of a first modality and media data of a second modality, where the types of the first modality and the second modality are different;

[0012] Input the media data of the first modality into the teacher model for knowledge extraction to obtain a first soft label, and input the media data of the second modality into the student model for knowledge extraction to obtain a second soft label;

[0013] Perform bidirectional knowledge transfer processing based on the first soft label and the second soft label to obtain a teacher transfer loss value of the teacher model and a student transfer loss value of the student model;

[0014] Update the parameters of the knowledge distillation network model according to the teacher transfer loss value and the student transfer loss value to obtain a trained knowledge distillation network model;

[0015] Extract the cross-modal knowledge model from the trained knowledge distillation network model.

[0016] In a second aspect, an embodiment of the present application provides a method for identifying a cross-modal knowledge model, including:

[0017] Obtain target media data to be identified;

[0018] Input the target media data into the above cross-modal knowledge model for identification to obtain an identification result of the target media data.

[0019] In a third aspect, an embodiment of the present application provides a training system for a cross-modal knowledge model, which is applied to a knowledge distillation network model. The knowledge distillation network model includes a teacher model and a student model. The system includes:

[0020] A first processing unit, configured to obtain media data of a first modality and media data of a second modality, where the types of the first modality and the second modality are different;

[0021] A second processing unit, configured to input the media data of the first modality into the teacher model for knowledge extraction to obtain a first soft label, and input the media data of the second modality into the student model for knowledge extraction to obtain a second soft label;

[0022] A third processing unit, configured to perform bidirectional knowledge transfer processing based on the first soft label and the second soft label to obtain a teacher transfer loss value of the teacher model and a student transfer loss value of the student model;

[0023] A fourth processing unit, configured to update the parameters of the knowledge distillation network model according to the teacher transfer loss value and the student transfer loss value to obtain a trained knowledge distillation network model;

[0024] A fifth processing unit, configured to extract the cross-modal knowledge model from the trained knowledge distillation network model.

[0025] In a fourth aspect, an embodiment of the present application further provides an electronic device, including:

[0026] At least one processor;

[0027] At least one memory for storing at least one program;

[0028] When the at least one program is executed by the at least one processor, the at least one processor implements the above training method or recognition method.

[0029] In a fifth aspect, an embodiment of the present application further provides a computer-readable storage medium, in which a program executable by a processor is stored, and the program executable by the processor is used to implement the above training method or recognition method when executed by the processor.

[0030] The advantages and beneficial effects of the present application will be partially given in the following description, partially become obvious from the following description, or be understood through the practice of the present application:

[0031] A training method, recognition method and system for a cross-modal knowledge model disclosed in an embodiment of the present application are applied to a knowledge distillation network model, and the knowledge distillation network model includes a teacher model and a student model. The method obtains media data of a first modality and media data of a second modality, where the types of the first modality and the second modality are different; inputs the media data of the first modality into the teacher model for knowledge extraction to obtain a first soft label, and inputs the media data of the second modality into the student model for knowledge extraction to obtain a second soft label; performs bidirectional knowledge transfer processing according to the first soft label and the second soft label to obtain a teacher transfer loss value of the teacher model and a student transfer loss value of the student model; updates the parameters of the knowledge distillation network model according to the teacher transfer loss value and the student transfer loss value to obtain a trained knowledge distillation network model; extracts the cross-modal knowledge model from the trained knowledge distillation network model. By performing bidirectional knowledge transfer on the first soft label and the second soft label, specifically, while transferring the first soft label of the teacher model to the student model, the second soft label of the student model is also transferred to the teacher model, and the media data of different modalities are complemented through the bidirectional interaction between the teacher model and the student model, thereby alleviating and suppressing the modality gap problem caused by unbalanced modality data during the knowledge distillation process, and further being beneficial to improving the model performance of the cross-modal knowledge model. BRIEF DESCRIPTION OF THE DRAWINGS

[0032] Figure 1 is a schematic flowchart of a training method for a cross-modal knowledge model provided by an embodiment of the present application;

[0033] Figure 2 It is a schematic diagram of the framework of the first knowledge distillation network model provided by the embodiments of the present application;

[0034] Figure 3 It is a detailed flowchart of the first two-way knowledge transfer provided by the embodiments of the present application;

[0035] Figure 4 It is a schematic diagram of the framework of the second knowledge distillation network model provided by the embodiments of the present application;

[0036] Figure 5 It is a detailed flowchart of the second two-way knowledge transfer provided by the embodiments of the present application;

[0037] Figure 6 It is a schematic diagram of the framework of the third knowledge distillation network model provided by the embodiments of the present application;

[0038] Figure 7 It is a schematic diagram of the process for verifying label relevance provided by the embodiments of the present application;

[0039] Figure 8 It is a schematic diagram of the process of an identification method for a cross-modal knowledge model provided by the embodiments of the present application;

[0040] Figure 9 It is a schematic diagram of the framework of a training system for a cross-modal knowledge model provided by the embodiments of the present application;

[0041] Figure 10 It is a schematic diagram of the hardware structure of an electronic device provided by the embodiments of the present application. Detailed implementation manners

[0042] In order to make the objectives, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that when referring to "embodiments" in this document, it means that specific features, structures or characteristics described in connection with the embodiments can be included in at least one embodiment of the present application. The phrase does not necessarily refer to the same embodiment at all positions in the specification, nor is it an independent or alternative embodiment mutually exclusive with other embodiments. Those skilled in the art will explicitly and implicitly understand that the embodiments described herein can be combined with other embodiments. The implementation manners described in the following exemplary embodiments do not represent all implementation manners consistent with the embodiments of the present application. They are only examples of devices and methods consistent with some aspects of the embodiments of the present application as detailed in the appended claims.

[0043] It can be understood that the terms "first", "second", etc. used in this application may be used herein to describe various concepts, but unless otherwise specified, these concepts are not limited by these terms. These terms are only used to distinguish one concept from another. For example, without departing from the scope of the embodiments of this application, the first loss function may also be referred to as the second loss function, and similarly, the second loss function may also be referred to as the first loss function. Depending on the context, words such as "if" and "when" used herein may be interpreted as "when...", "while...", or "in response to determining".

[0044] The terms "at least one", "a plurality of", "each", "any one", etc. used in this application, at least one includes one, two or more than two, a plurality of includes two or more than two, each refers to each one of the corresponding plurality, and any one refers to any one of the plurality.

[0045] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which this application belongs. The terms used herein are only for the purpose of describing the embodiments of this application and are not intended to limit this application.

[0046] In the following description, for the purpose of illustration rather than limitation, specific details such as specific system architectures and technologies are set forth in order to provide a thorough understanding of the embodiments of this application. However, those skilled in the art should understand that this application can also be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid unnecessary details from obscuring the description of this application.

[0047] Currently, existing knowledge distillation (KD) techniques typically use a high-precision or high-annotation-quality modality as the teacher model to transfer knowledge to a low-precision or unannotated modality. However, due to the imbalance in information content between different modality data, the modality gap that appears in the cross-modal knowledge distillation (CMKD) process will cause existing knowledge distillation techniques to fail in the cross-modal knowledge distillation process.

[0048] In addition, there are still problems of inconsistent soft labels in the process of cross-modal knowledge distillation (CMKD). Specifically, there are label inconsistencies or misalignments between different modal data. For example, in video data, the labels among audio, video, and subtitles may be inconsistent, making it difficult for the model to accurately align these modalities during the training process, thus resulting in a modality gap in the cross-modal knowledge distillation (CMKD) process. Moreover, existing technologies usually unidirectionally transfer modal knowledge with high precision or high annotation quality from the teacher model to the student model. The knowledge generated in the cross-modal knowledge distillation process is relatively fixed, with poor knowledge diversity and low personalization degree of knowledge.

[0049] In view of this, the present application proposes a training method for a cross-modal knowledge model, which is applied to a knowledge distillation network model. The knowledge distillation network model includes a teacher model and a student model. This method performs bidirectional knowledge transfer on the first soft label and the second soft label. Specifically, while transferring the first soft label of the teacher model to the student model, it also transfers the second soft label of the student model to the teacher model. Through the bidirectional interaction between the teacher model and the student model, the media data of different modalities are made complementary, thereby alleviating and suppressing the modality gap problem caused by unbalanced modal data in the knowledge distillation process, and further facilitating the improvement of the model performance of the cross-modal knowledge model.

[0050] In addition, this method verifies the correlation between the first soft label and the second soft label through an online filtering soft distillation (OFSD) strategy. It can filter out non-extractable samples (i.e., filter soft label misalignment samples), and at the same time, it can enable the teacher model or the student model to inherit knowledge from non-target classes, avoiding the phenomenon of inconsistent soft labels and effectively suppressing the modality gap phenomenon in the cross-modal knowledge distillation (CMKD) process.

[0051] Moreover, this method introduces a bidirectional distillation strategy with dual proxies. Specifically, corresponding proxy soft labels or inference soft labels are output through different classification heads in the target model. By gradually transferring cross-modal knowledge, it can improve the generated knowledge diversity and personalization degree, thus facilitating the improvement of the model performance of the subsequent model.

[0052] The method provided by the embodiments of the present application can be applied to a terminal, a server, or software running on a terminal or a server. In some embodiments, the terminal can be a smart phone, a tablet computer, a laptop computer, a desktop computer, a smart speaker, a smart watch, a vehicle-mounted terminal, etc., but is not limited thereto; the server side can be configured as an independent physical server, or as a server cluster or distributed system composed of multiple physical servers, or as a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms. The server can also be a node server in a blockchain network; the software can be an application implementing the method, etc., but is not limited to the above forms.

[0053] The present application can be used in many general or special computer system environments or configurations. For example: personal computers, server computers, handheld or portable devices, tablet-type devices, multi-processor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics devices, network PCs, minicomputers, mainframe computers, distributed computing environments including any of the above systems or devices, and so on. The present application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. The present application can also be practiced in a distributed computing environment where tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules can be located in local and remote computer storage media including storage devices.

[0054] Refer to Figure 1 and Figure 2 , Figure 1 FIG. is an optional flowchart of a method for training a cross-modal knowledge model provided by the embodiments of the present application. Figure 1 The method in FIG. is applied to a knowledge distillation network model, and the knowledge distillation network model includes a teacher model and a student model. The method may include, but is not limited to, steps S110 to S150.

[0055] S110. Obtain media data of a first modality and media data of a second modality, where the types of the first modality and the second modality are different;

[0056] In the embodiments of the present application, the knowledge distillation network model can be a neural network model constructed based on the Knowledge Distillation technology. The knowledge distillation network model includes a teacher model and a student model. The first modality can be any one of the image modality, text modality, audio modality, video modality, RGB-depth modality, etc., and the second modality can be any modality other than the first modality.

[0057] It can be understood that for the media data of the first modality, if the first modality is the image modality, the corresponding media data can be image data; or, if the first modality is the text modality, the corresponding media data can be text data; or, if the first modality is the audio modality, the corresponding media data can be audio data; or, if the first modality is the video modality, the corresponding media data can be video data; or, if the first modality is the RGB-depth modality, the corresponding media data can be RGB-depth data. As for the media data of the second modality, it is similar to the media data of the aforementioned first modality and can be simply deduced by analogy.

[0058] S120. Input the media data of the first modality into the teacher model for knowledge extraction to obtain a first soft label, and input the media data of the second modality into the student model for knowledge extraction to obtain a second soft label;

[0059] In the embodiments of the present application, for a certain training process, the media data of the first modality can be input into the teacher model, and the teacher model can predict the media data of the first modality to obtain a soft label corresponding to the media data of the first modality, and the obtained soft label is determined as the first soft label; the content of the student model is similar to that of the aforementioned teacher model and can be simply deduced by analogy.

[0060] S130. Perform two-way knowledge transfer processing according to the first soft label and the second soft label to obtain a teacher transfer loss value of the teacher model and a student transfer loss value of the student model;

[0061] In the embodiments of the present application, for the teacher model, the second soft label output by the student model can be used as knowledge and transferred into the teacher model to obtain the teacher transfer loss value of the teacher model in the current training process; the content of the student model is similar to that of the aforementioned teacher model and can be simply deduced by analogy.

[0062] S140. Update the parameters of the knowledge distillation network model according to the teacher transfer loss value and the student transfer loss value to obtain a trained knowledge distillation network model;

[0063] In an embodiment of the present application, in one feasible implementation manner, step S140 may be to perform backpropagation update on the teacher model in the knowledge distillation network model based on the teacher transfer loss value, and perform backpropagation update on the student model in the knowledge distillation network model based on the student transfer loss value. After obtaining the trained teacher model and the trained student model respectively, a trained knowledge distillation network model is constructed based on the trained teacher model and the trained student model.

[0064] It can be understood that, in another feasible implementation manner, step S140 may be to perform joint training update on the initial knowledge distillation network model based on the teacher transfer loss value and the student transfer loss value by using the backpropagation algorithm, so as to obtain a trained knowledge distillation network model.

[0065] S150. Extract the cross-modal knowledge model from the trained knowledge distillation network model.

[0066] In an embodiment of the present application, the cross-modal knowledge model for inference can be determined from the trained knowledge distillation network model.

[0067] Refer to Figure 3 and continue to refer to Figure 2 In the first feasible implementation manner, step S130, according to the first soft label and the second soft label, perform two-way knowledge transfer processing to obtain the teacher transfer loss value of the teacher model and the student transfer loss value of the student model, including:

[0068] A1. Obtain the first true label of the media data of the first modality and the second true label of the media data of the second modality;

[0069] A2. Calculate the hard label loss of the first soft label according to the first true label to obtain a first loss value, and calculate the hard label loss of the second soft label according to the second true label to obtain a second loss value;

[0070] A3. Calculate the soft label loss of the first soft label according to the second soft label to obtain a third loss value;

[0071] A4. Obtain the teacher transfer loss value according to the first loss value and the third loss value;

[0072] A5. Obtain the student transfer loss value according to the second loss value and the third loss value.

[0073] In the embodiments of the present application, the first true label is used to indicate the true attribute category or attribute information of the media data of the first modality, and the second true label is used to indicate the true attribute category or attribute information of the media data of the second modality. For the first true label, the hard label loss calculation can be the hard label loss between the first true label and the first soft label, and the specific calculation method can be based on the cross-entropy loss function to obtain the first loss value; the second loss value is similar to the content of the foregoing first loss value and can be simply deduced by analogy.

[0074] For the first feasible implementation manner, the soft label loss calculation in step A3 can be to calculate the soft label loss value between the first soft label and the second soft label based on the KD loss function. The specific KD loss function can be constructed based on any one of the KL divergence method, the DKD (Decoupled Knowledge Distillation) method, etc., so as to obtain the third loss value. Specifically, for the teacher model, its specific implementation manner can be to transfer the second soft label in the student model to the teacher model, and then the teacher model uses the second soft label as its learning target in the knowledge distillation stage, calculates the KD loss value of the teacher model through the KD loss function, and records it as the third loss value; the third loss value of the student model is similar to the third loss value of the foregoing teacher model and can be simply deduced by analogy.

[0075] It can be understood that after obtaining the first loss value of the teacher model, the second loss value of the student model, and the soft label loss (i.e., the third loss value) between the teacher model and the student model, the teacher transfer loss value of the teacher model in the current training process can be obtained based on the first loss value and the third loss value through calculation methods such as direct summation or weighted summation, and the student transfer loss value of the student model in the current training process is the same by analogy.

[0076] Refer to Figure 4 and Figure 5 , in the second feasible implementation manner, step S130, performing two-way knowledge transfer processing according to the first soft label and the second soft label to obtain the teacher transfer loss value of the teacher model and the student transfer loss value of the student model, includes:

[0077] B1. Obtain a dynamic distillation strategy, the first true label of the media data of the first modality, and the second true label of the media data of the second modality;

[0078] B2. According to the dynamic distillation strategy, perform label correlation verification on the first soft label and the second soft label to obtain a first label verification result;

[0079] B3. If the first label verification result is that the label is distorted, then based on the first true label and the first soft label, obtain the teacher transfer loss value, and based on the second true label and the second soft label, obtain the student transfer loss value;

[0080] Or, B4. If the first label verification result is that the label is not distorted, then based on the first true label, the first soft label, and the second soft label, obtain the teacher transfer loss value, and based on the second true label, the first soft label, and the second soft label, obtain the student transfer loss value.

[0081] In the embodiments of the present application, on the basis of the first feasible implementation manner described above, a dynamic distillation (OFSD) strategy can be introduced in the two-way knowledge transfer process to sample and select soft label misalignment, so as to implement filtering processing on the samples with soft label misalignment. Specifically, the dynamic distillation strategy can be expressed as:

[0082]

[0083] where η is the strategy representation symbol of the dynamic distillation strategy; ω is the threshold of the dynamic distillation strategy; KRC(f T , f S ) is the Kendall rank correlation coefficient between the first soft label f T and the second soft label f S .

[0084] The Kendall rank correlation coefficient can be expressed as:

[0085]

[0086] where KRC is the Kendall rank correlation coefficient; C is the number of soft label pairs with consistent relative order of elements between the first soft label and the second soft label; is the first soft label with serial number i; the first soft label with serial number j; is the second soft label with serial number i; is the second soft label with serial number j.

[0087] It can be understood that step B2 can be to determine the first label verification result between the first soft label and the second soft label based on the dynamic distillation strategy. Specifically, if η = 0, then a first label verification result indicating label distortion can be generated. At this time, the knowledge transfer path between the teacher model and the student model is disconnected, and no knowledge transfer occurs between the teacher model and the student model. The teacher transfer loss value in step B3 can be calculated based on the cross-entropy loss function of the first true label and the first soft label, and the student transfer loss value is the same.

[0088] Alternatively, if η = 1, a first label verification result with undistorted labels can be generated. At this time, the teacher transfer loss value in step B4 is similar to the content of the aforementioned steps A2 to A4. Specifically, it is to calculate the hard label loss between the first true label and the first soft label, and calculate the soft label loss between the first soft label and the second soft label, and generate a teacher transfer loss value based on the calculated hard label loss value and soft label loss value; the student transfer loss value in step B4 is similar to the teacher transfer loss value in the aforementioned step B4 and can be simply deduced by analogy, so it will not be elaborated herein.

[0089] It should be noted that if the first label verification result is that the label is undistorted, for the specific implementation method of the soft label loss in step B4, it can be that the first soft label is migrated to the student model through the conduction path of the dynamic distillation strategy, and the second soft label is migrated to the teacher model through the conduction path of the dynamic distillation strategy, which will not be elaborated herein.

[0090] Referring to Figure 6 , in the third feasible implementation manner, the inputting the media data of the first modality into the teacher model for knowledge extraction to obtain a first soft label, and the inputting the media data of the second modality into the student model for knowledge extraction to obtain a second soft label includes:

[0091] C1. Inputting the target modality data into the target model for knowledge extraction to obtain a target soft label;

[0092] Wherein, the target model is the teacher model or the student model, and the classification head that outputs the proxy soft label in the target model is different from the classification head that outputs the inference soft label. The target modality data is the media data of the first modality or the media data of the second modality; when the target modality data is the media data of the first modality, the target model is the teacher model, and the target soft label is the first soft label; when the target modality data is the media data of the second modality, the target model is the student model, and the target soft label is the second soft label.

[0093] In the embodiments of the present application, on the basis of the aforementioned second feasible implementation manner, a proxy classification head can be introduced to gradually transfer cross-modal knowledge. Specifically, taking the target model as the teacher model in the embodiments of the present application as an example, the target modality data input into the teacher model is the media data of the first modality, and step C1 can be inputting the media data of the first modality into the teacher model. This media data of the first modality passes through the teacher backbone layer (i.e., Figure 6 T in) and two parallel classification heads to obtain a first soft label, which includes the proxy soft label output by the proxy classification head (i.e., Figure 6 PT in) and the inference classification head (i.e.,Figure 6 The inference soft labels output by CH) connected to the teacher backbone layer T in the [specific context].

[0094] It can be understood that when introducing a proxy classification head for the teacher model, the initial proxy classification head introduced can have the same parameters as the original inference classification head in the teacher model. Additionally, for the case where the target model is a student model, it can be simply analogously derived in a similar manner to the case where the target model is a teacher model.

[0095] It should be noted that the teacher backbone layer T in the embodiments of this application can include a global average pooling layer (GAP) and a feature adaptation layer; and the student backbone layer S is similar to the aforementioned teacher backbone layer T.

[0096] Exemplarily, for the inference soft labels output by the backbone layer and the inference classification head, it can be expressed as:

[0097]

[0098] where f m is the inference soft label; is the classification function of the inference classification head; GAP(·) is the global average pooling function; B m is the overall symbolic representation of the teacher backbone layer T or the student backbone layer S; F m is the media data of the first modality input to the teacher model, or the media data of the second modality input to the student model; T is used to indicate the teacher backbone layer; S is used to indicate the student backbone layer.

[0099] And the proxy soft labels output by the backbone layer and the proxy classification head can be expressed as:

[0100]

[0101] where, is the proxy soft label; is the classification function of the proxy classification head; A is the adaptation function of the feature adaptation layer, and this feature adaptation layer includes a Conv - BN - ReLU block.

[0102] Step S130, according to the first soft label and the second soft label, perform two - way knowledge transfer processing to obtain the teacher transfer loss value of the teacher model and the student transfer loss value of the student model, including:

[0103] D1. Obtain a dynamic distillation strategy, the first true label of the media data of the first modality, and the second true label of the media data of the second modality;

[0104] D2. Verify the label correlation between the first soft label and the second soft label according to the dynamic rectification strategy to obtain a second label verification result. The first soft label includes a first proxy soft label and a first inference soft label, and the second soft label includes a second proxy soft label and a second inference soft label;

[0105] D3. If the second label verification result indicates label distortion, obtain the teacher migration loss value according to the first true label, the first proxy soft label, and the first inference soft label, and obtain the student migration loss value according to the second true label, the second proxy soft label, and the second inference soft label;

[0106] Alternatively, D4. If the second label verification result indicates that the label is not distorted, obtain a target migration loss value according to the target true label, the first proxy soft label, the first inference soft label, the second proxy soft label, and the second inference soft label. The target true label is the first true label or the second true label, and the target migration loss value is the teacher migration loss value or the student migration loss value corresponding to the target true label.

[0107] In the embodiments of the present application, the first implementation manner of step D2 may be similar to the foregoing step B2 and can be simply analogized. For the second implementation manner of step D2, it may be to verify the label correlation between the first proxy soft label and the second proxy soft label respectively, and verify the label correlation between the first inference soft label and the second inference soft label, and determine the final second label verification result based on the two obtained verification results.

[0108] In addition, for the third implementation manner of step D2, it may be to verify the label correlation between two soft labels of another model respectively based on each soft label, and determine the final second label verification result based on the obtained multiple verification results. For example, for the second inference soft label of the student model, it may verify the label correlation with the first inference soft label and the first proxy soft label of the teacher model respectively.

[0109] In some embodiments, the target migration loss value includes a proxy migration loss value and an inference migration loss value. Obtaining the target migration loss value according to the target true label, the target proxy soft label, and the target inference soft label includes:

[0110] E1. Calculate the soft label loss of the target inference soft label according to the target proxy soft label to obtain the proxy migration loss value;

[0111] E2. Calculate the hard label loss of the target inference soft label according to the target true label to obtain a fourth loss value;

[0112] E3. Obtain the inference transfer loss value according to the fourth loss value and the proxy transfer loss value.

[0113] In the embodiments of the present application, if the second label verification result is label distortion, it indicates that the knowledge transfer path between the teacher model and the student model is disconnected, and no knowledge transfer occurs between the teacher model and the student model. At this time, the target proxy soft label and the target inference soft label are respectively the proxy soft label and the inference soft label corresponding to the target true label. Specifically, if the target true label is the first true label, its corresponding target proxy soft label is the first proxy soft label, and the corresponding target inference soft label is the first inference soft label; or, if the target true label is the second true label, its corresponding target proxy soft label is the second proxy soft label, and the corresponding target inference soft label is the second inference soft label.

[0114] It can be understood that for Figure 6 the teacher model in, its teacher transfer loss value includes a proxy transfer loss value and an inference transfer loss value, where the proxy transfer loss value is the loss value corresponding to the teacher model proxy classification head, and this proxy transfer loss value can be calculated based on the soft labels output by two different classification heads of the teacher model. The specific calculation method can be similar to the content of the foregoing step A3 and can be simply analogized.

[0115] For the inference transfer loss value, it is the loss value corresponding to the teacher model inference classification head. This inference transfer loss value includes the hard label loss (i.e., the fourth loss value) between the first true label and the first inference soft label and the soft label loss (i.e., the proxy transfer loss value) between the first inference soft label and the first proxy soft label. Specifically, step E2 is similar to the content of the foregoing step A2 and can be simply analogized. After obtaining the fourth loss value and the proxy transfer loss value, the inference transfer loss value can be obtained through calculation methods such as direct summation or weighted summation.

[0116] It should be noted that for Figure 6 the student model in, its student transfer loss value is similar to the content of the foregoing teacher transfer loss value.

[0117] In some embodiments, the target transfer loss value includes a proxy transfer loss value and an inference transfer loss value. The obtaining of the target transfer loss value according to the target true label, the first proxy soft label, the first inference soft label, the second proxy soft label, and the second inference soft label includes:

[0118] F1. Calculate the hard label loss for the target inference soft label according to the target true label to obtain a fifth loss value, where the target inference soft label is the first inference soft label or the second inference soft label corresponding to the target true label;

[0119] F2. Calculate the soft label loss of each of the first proxy soft label, the first inference soft label, the second proxy soft label, and the second inference soft label to obtain a soft label loss value set;

[0120] F3. Obtaining the proxy migration loss value according to the soft label loss value set;

[0121] F4. Obtain the inference migration loss value according to the fifth loss value and the soft label loss value set.

[0122] In an embodiment of the present application, if the second label verification result is that the label is not distorted, it means that the knowledge transfer path between the teacher model and the student model is connected, and knowledge transfer is required between the teacher model and the student model, and the target true label of step F1 is similar to the target true label of step E2.

[0123] It is understandable that if the second label verification result is that the label is not distorted, it can be further subdivided into a label verification result indicating that at least one soft label pair is not distorted, or a label verification result indicating that all soft label pairs are not distorted.

[0124] Exemplarily, for the second label verification result obtained in step D2 in the aforementioned third embodiment, if the label verification result is specifically a label verification result in which at least one soft label pair is not distorted. For example, if the second label verification result indicates that the soft label pair of the first proxy soft label and the second inference soft label is distorted, and the other soft label pairs are not distorted, then step F2 may be to calculate the soft label loss value of the soft label pair of the first proxy soft label and the second proxy soft label respectively, recorded as the sixth loss value; calculate the soft label loss value of the soft label pair of the first inference soft label and the second inference soft label, recorded as the seventh loss value; calculate the soft label loss value of the soft label pair of the first inference soft label and the second proxy soft label, recorded as the eighth loss value; calculate the soft label loss value of the soft label pair of the first proxy soft label and the first inference soft label, recorded as the ninth loss value; and calculate the soft label loss value of the soft label pair of the second proxy soft label and the second inference soft label, recorded as the tenth loss value; then, construct a soft label loss value set based on the sixth loss value, the seventh loss value, the eighth loss value, the ninth loss value and the tenth loss value.

[0125] Alternatively, still for the second label verification result obtained in the foregoing third implementation manner in step D2, if the label verification result is specifically that all soft label pairs are not distorted, it may be based on each soft label to calculate the soft label loss for the two soft labels of another model respectively, thereby constructing a set of soft label loss values. In addition to the foregoing sixth loss value, seventh loss value, eighth loss value, ninth loss value, and tenth loss value, this set of soft label loss values further includes an eleventh loss value, and this eleventh loss value is the soft label loss value of the soft label pair of the first proxy soft label and the second inference soft label.

[0126] For the target model, the proxy migration loss value of its proxy classification head can be calculated by direct summation, weighted summation, or other methods based on a number of soft label loss values obtained in step F3. For example, in the case where the second label verification result indicates that the soft label pair of the first proxy soft label and the second inference soft label is distorted and the remaining soft label pairs are not distorted, for the proxy classification head of the teacher model, its proxy migration loss value can be determined based on the sixth loss value and the ninth loss value in the set of soft label loss values; for the proxy classification head of the student model, its proxy migration loss value can be determined based on the sixth loss value, the eighth loss value, and the tenth loss value in the set of soft label loss values.

[0127] For the target model, the inference migration loss value of its inference classification head can be calculated by direct summation, weighted summation, or other methods based on the fifth loss value obtained in step F1 and a number of soft label loss values obtained in step F3. For example, in the case where the second label verification result indicates that the soft label pair of the first proxy soft label and the second inference soft label is distorted and the remaining soft label pairs are not distorted, for the inference classification head of the teacher model, its inference migration loss value can be determined based on the fifth loss value and the seventh loss value, the eighth loss value, and the ninth loss value in the set of soft label loss values; for the inference classification head of the student model, its inference migration loss value can be determined based on the fifth loss value and the seventh loss value and the tenth loss value in the set of soft label loss values.

[0128] It should be noted that the examples in this application are only for illustration, and do not limit the method for obtaining the second label verification result, nor the method for obtaining the target migration loss value under the constraint of the second label verification result. There can be many other specific examples. For example, the second label verification result can also be obtained based on the first implementation manner or the second implementation manner of step D2, and the corresponding example of step F2 can be obtained by making corresponding changes based on the foregoing examples. This application will not elaborate further here.

[0129] In addition, for the case where the soft label pair of the first proxy soft label and the second inference soft label is distorted while the other soft label pairs are not distorted, the specific implementation manners of steps F2 to F4 may be that in the knowledge transfer path between the teacher model and the student model, the knowledge transfer path of the soft label pair of the first proxy soft label and the second inference soft label is disconnected, while the knowledge transfer paths between the first proxy soft label and the second proxy soft label, between the first inference soft label and the second inference soft label, and between the first inference soft label and the second proxy soft label are connected. This application will not elaborate further here.

[0130] Referring Figure 7 , in some embodiments, according to the dynamic distillation strategy, the label correlation between the first soft label and the second soft label is verified to obtain a target label verification result, including:

[0131] G1. Obtain the strategy threshold of the dynamic distillation strategy;

[0132] G2. Calculate the correlation metric for the second soft label based on the first soft label to obtain the soft label correlation degree;

[0133] G3. Compare the soft label correlation degree with the strategy threshold to obtain the target label verification result;

[0134] Wherein, the target label verification result is the first label verification result or the second label verification result.

[0135] In the embodiments of the present application, the target label verification result may be the first label verification result obtained in step B2 or the second label verification result obtained in step D2. The strategy threshold may be the threshold ω of the aforementioned dynamic distillation strategy.

[0136] It can be understood that the soft label correlation degree between the first soft label and the second soft label can be determined based on the aforementioned Kendall rank correlation coefficient; and step G3 may be to compare the size relationship between the strategy threshold and several soft label correlation degrees, and determine the first label verification result or the second label verification result based on the obtained several comparison information.

[0137] Specifically, for a certain soft label correlation degree, if the soft label correlation degree is greater than the strategy threshold, it indicates that the knowledge transfer path between the first soft label and the second soft label needs to be connected, and the strategy value of the corresponding dynamic distillation strategy is 1. At this time, a target label verification result indicating that the label is not distorted can be generated; or, if the soft label correlation degree is less than or equal to the strategy threshold, it indicates that the knowledge transfer path between the first soft label and the second soft label needs to be disconnected, and the strategy value of the corresponding dynamic distillation strategy is 0. At this time, a target label verification result indicating that the label is distorted can be generated.

[0138] It should be noted that in practical applications, forFigure 2 or Figure 4 For the knowledge distillation network model, the cross-modal knowledge model obtained by extraction can be a student model, which can specifically be composed of a student backbone layer (i.e., Figure 2 or Figure 4 the S in Figure 2 or Figure 4 and a student inference classification head (i.e., Figure 6 For the knowledge distillation network model, in order to reduce the inference overhead required for the cross-modal knowledge model during actual inference, the corresponding cross-modal knowledge model can be a partial network structure of the student model, and this cross-modal knowledge model can specifically be composed of Figure 6 the student backbone layer S of Figure 6 and the inference classification head CH connected to the student backbone layer S in

[0139] Figure 8 FIG. Figure 8 The method in

[0140] S210. Obtain target media data to be recognized;

[0141] S220. Input the target media data into the above cross-modal knowledge model for recognition to obtain the recognition result of the target media data.

[0142] In the embodiments of the present application, in practical applications, the target media data can be any one of image media data, text media data, audio media data, video media data, RGB-depth media data, etc.; after inputting the target media data into the cross-modal knowledge model, the cross-modal knowledge model can recognize the input target media data to obtain the recognition result.

[0143] Specifically, any one of image classification, image reconstruction, etc. can be performed on the image media data; or, any one of text generation, text keyword extraction, text semantic recognition, sentiment analysis, etc. can be performed on the text media data; or, any one of audio-text conversion, audio sentiment analysis, target audio extraction (such as extracting the audio of a target person from multi-person mixed audio), etc. can be performed on the audio media data; or, any one of audio recognition, image recognition, subtitle text recognition, etc. can be performed on the video media data.

[0144] Please refer to Figure 9 , the embodiments of the present application further provide a training system for a cross-modal knowledge model, which is applied to a knowledge distillation network model. The knowledge distillation network model includes a teacher model and a student model. The system includes:

[0145] A first processing unit 810, configured to obtain media data of a first modality and media data of a second modality, where the types of the first modality and the second modality are different;

[0146] A second processing unit 820, configured to input the media data of the first modality into the teacher model for knowledge extraction to obtain a first soft label, and input the media data of the second modality into the student model for knowledge extraction to obtain a second soft label;

[0147] A third processing unit 830, configured to perform bidirectional knowledge transfer processing according to the first soft label and the second soft label to obtain a teacher transfer loss value of the teacher model and a student transfer loss value of the student model;

[0148] A fourth processing unit 840, configured to update parameters of the knowledge distillation network model according to the teacher transfer loss value and the student transfer loss value to obtain a trained knowledge distillation network model;

[0149] A fifth processing unit 850, configured to extract the cross-modal knowledge model from the trained knowledge distillation network model.

[0150] It can be understood that the content in the above method embodiments is applicable to the system embodiments. The functions specifically implemented by the system embodiments are the same as those in the above method embodiments, and the beneficial effects achieved are also the same as those in the above method embodiments.

[0151] An embodiment of the present application further provides an electronic device. The electronic device includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the above method is implemented. The electronic device may be any intelligent terminal including a tablet computer, an in-vehicle computer, etc.

[0152] It can be understood that the content in the above method embodiments is applicable to the device embodiments. The functions specifically implemented by the device embodiments are the same as those in the above method embodiments, and the beneficial effects achieved are also the same as those in the above method embodiments.

[0153] Please refer to Figure 10 , Figure 10 , which schematically shows the hardware structure of an electronic device according to an embodiment. The electronic device includes:

[0154] The processor 901 can be implemented in the form of a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, etc., and is used to execute relevant programs to implement the technical solutions provided in the embodiments of the present application;

[0155] The memory 902 can be implemented in the form of a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM), etc. The memory 902 can store an operating system and other application programs. When implementing the technical solutions provided in the embodiments of this specification through software or firmware, the relevant program codes are stored in the memory 902 and are called by the processor 901 to execute the methods in the embodiments of the present application;

[0156] The input / output interface 903 is used to implement information input and output;

[0157] The communication interface 904 is used to implement communication interaction between this device and other devices, and can implement communication through wired means (such as USB, network cable, etc.) or wireless means (such as mobile network, WIFI, Bluetooth, etc.);

[0158] The bus 905 transmits information between the various components of the device (such as the processor 901, the memory 902, the input / output interface 903, and the communication interface 904);

[0159] Among them, the processor 901, the memory 902, the input / output interface 903, and the communication interface 904 achieve communication connections with each other inside the device through the bus 905.

[0160] The embodiments of the present application also provide a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, the above-mentioned method is implemented.

[0161] It can be understood that the content in the above method embodiments is applicable to the embodiments of this storage medium. The functions specifically implemented by the embodiments of this storage medium are the same as those in the above method embodiments, and the beneficial effects achieved are also the same as those in the above method embodiments.

[0162] The memory, as a non-transitory computer-readable storage medium, can be used to store non-transitory software programs and non-transitory computer-executable programs. In addition, the memory may include high-speed random access memory, and may also include non-transitory memory, such as at least one magnetic disk storage device, a flash memory device, or other non-transitory solid-state storage devices. In some embodiments, the memory may optionally include memories remotely located relative to the processor, and these remote memories can be connected to the processor through a network. Examples of the above networks include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0163] The embodiments described in the embodiments of the present application are for more clearly illustrating the technical solutions of the embodiments of the present application, and do not constitute a limitation on the technical solutions provided by the embodiments of the present application. Those skilled in the art will know that with the evolution of technology and the emergence of new application scenarios, the technical solutions provided by the embodiments of the present application are equally applicable to similar technical problems.

[0164] Those skilled in the art can understand that the technical solutions shown in the figures do not constitute a limitation on the embodiments of the present application, and may include more or fewer steps than shown in the figures, or combine certain steps, or different steps.

[0165] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0166] Those of ordinary skill in the art can understand that all or some of the steps in the methods disclosed above, and the functional modules / units in the systems and devices, can be implemented as software, firmware, hardware, and appropriate combinations thereof.

[0167] The terms "first", "second", "third", "fourth", etc. (if any) in the specification of the present application and the above-mentioned drawings are used to distinguish similar objects, and do not have to be used to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present application described here can be implemented in an order other than those illustrated or described here. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units does not have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products, or devices.

[0168] It should be understood that in this application, "at least one (item)" means one or more, and "a plurality" means two or more. "And / or" is used to describe the association relationship of associated objects and indicates that there can be three relationships. For example, "A and / or B" can mean: only A exists, only B exists, and both A and B exist at the same time. Among them, A and B can be singular or plural. The character " / " generally represents an "or" relationship between the associated objects before and after. "At least one (item) of the following" or its similar expression refers to any combination of these items, including any combination of single item (item) or plural items (items). For example, at least one (item) of a, b, or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, and c can be single or multiple.

[0169] In several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the above division of units is only a logical function division. In actual implementation, there can be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection to each other can be through some interfaces. The indirect coupling or communication connection of devices or units can be in electrical, mechanical or other forms.

[0170] The units described above as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place, or they can be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0171] In addition, in each embodiment of this application, the functional units can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above integrated unit can be implemented in the form of hardware or in the form of a software functional unit.

[0172] When the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes multiple instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods of various embodiments of this application. The aforementioned storage medium includes: various media that can store programs such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs.

[0173] The preferred embodiments of the embodiments of this application have been described above with reference to the accompanying drawings, and thus do not limit the scope of rights of the embodiments of this application. Any modifications, equivalent replacements, and improvements made by those skilled in the art without departing from the scope and essence of the embodiments of this application shall be within the scope of rights of the embodiments of this application.

Claims

1. A training method for a cross-modal knowledge model, characterized in that Applied to a knowledge distillation network model, the knowledge distillation network model includes a teacher model and a student model, and the method includes: Obtain media data of a first modality and media data of a second modality, where the types of the first modality and the second modality are different; Input the media data of the first modality into the teacher model for knowledge extraction to obtain a first soft label, and input the media data of the second modality into the student model for knowledge extraction to obtain a second soft label; Perform bidirectional knowledge transfer processing based on the first soft label and the second soft label to obtain a teacher transfer loss value of the teacher model and a student transfer loss value of the student model; Update the parameters of the knowledge distillation network model according to the teacher transfer loss value and the student transfer loss value to obtain a trained knowledge distillation network model; Extract the cross-modal knowledge model from the trained knowledge distillation network model.

2. The method according to claim 1, wherein The performing bidirectional knowledge transfer processing based on the first soft label and the second soft label to obtain the teacher transfer loss value of the teacher model and the student transfer loss value of the student model includes: Obtain a first ground truth label of the media data of the first modality and a second ground truth label of the media data of the second modality; Calculate a hard label loss for the first soft label according to the first ground truth label to obtain a first loss value, and calculate a hard label loss for the second soft label according to the second ground truth label to obtain a second loss value; Calculate a soft label loss for the first soft label according to the second soft label to obtain a third loss value; Obtain the teacher transfer loss value according to the first loss value and the third loss value; Obtain the student transfer loss value according to the second loss value and the third loss value.

3. The method according to claim 1, characterized in that The performing bidirectional knowledge transfer processing based on the first soft label and the second soft label to obtain the teacher transfer loss value of the teacher model and the student transfer loss value of the student model includes: Obtain a dynamic distillation strategy, a first ground truth label of the media data of the first modality, and a second ground truth label of the media data of the second modality; Verify the label correlation of the first soft label and the second soft label according to the dynamic distillation strategy to obtain a first label verification result; If the first label verification result is label distortion, obtain the teacher transfer loss value according to the first ground truth label and the first soft label, and obtain the student transfer loss value according to the second ground truth label and the second soft label; or, if the first label verification result is label non-distortion, obtain the teacher transfer loss value according to the first ground truth label, the first soft label, and the second soft label, and obtain the student transfer loss value according to the second ground truth label, the first soft label, and the second soft label.

4. The method according to claim 1, wherein Inputting the media data of the first modality into the teacher model for knowledge extraction to obtain a first soft label, and inputting the media data of the second modality into the student model for knowledge extraction to obtain a second soft label, includes: Inputting the target modality data into the target model for knowledge extraction to obtain a target soft label, where the target soft label includes a proxy soft label and an inference soft label; Among them, the classification head for outputting the proxy soft label in the target model is different from the classification head for outputting the inference soft label; the target modality data is the media data of the first modality or the media data of the second modality; when the target modality data is the media data of the first modality, the target model is the teacher model, and the target soft label is the first soft label; when the target modality data is the media data of the second modality, the target model is the student model, and the target soft label is the second soft label.

5. The method according to claim 4, wherein Performing two-way knowledge transfer processing based on the first soft label and the second soft label to obtain a teacher transfer loss value of the teacher model and a student transfer loss value of the student model, includes: Obtaining a dynamic distillation strategy, a first ground truth label of the media data of the first modality, and a second ground truth label of the media data of the second modality; According to the dynamic distillation strategy, performing label correlation verification on the first soft label and the second soft label to obtain a second label verification result, where the first soft label includes a first proxy soft label and a first inference soft label, and the second soft label includes a second proxy soft label and a second inference soft label; If the second label verification result is label distortion, then obtaining the teacher transfer loss value according to the first ground truth label, the first proxy soft label, and the first inference soft label, and obtaining the student transfer loss value according to the second ground truth label, the second proxy soft label, and the second inference soft label; or, if the second label verification result is label non-distortion, then obtaining a target transfer loss value according to a target ground truth label, the first proxy soft label, the first inference soft label, the second proxy soft label, and the second inference soft label, where the target ground truth label is the first ground truth label or the second ground truth label, and the target transfer loss value is the teacher transfer loss value or the student transfer loss value corresponding to the target ground truth label.

6. The method according to claim 5, wherein The target transfer loss value includes a proxy transfer loss value and an inference transfer loss value. Obtaining the target transfer loss value according to the target ground truth label, a target proxy soft label, and a target inference soft label, includes: Performing soft label loss calculation on the target inference soft label according to the target proxy soft label to obtain the proxy transfer loss value; Performing hard label loss calculation on the target inference soft label according to the target ground truth label to obtain a fourth loss value; Obtaining the inference transfer loss value according to the fourth loss value and the proxy transfer loss value.

7. The method according to claim 5, characterized in that The target migration loss value includes a proxy migration loss value and an inference migration loss value. Obtaining the target migration loss value according to the target true label, the first proxy soft label, the first inference soft label, the second proxy soft label, and the second inference soft label includes: Performing hard label loss calculation on the target inference soft label according to the target true label to obtain a fifth loss value, where the target inference soft label is the first inference soft label or the second inference soft label corresponding to the target true label; Performing pairwise soft label loss calculation on the first proxy soft label, the first inference soft label, the second proxy soft label, and the second inference soft label to obtain a set of soft label loss values; Obtaining the proxy migration loss value according to the set of soft label loss values; Obtaining the inference migration loss value according to the fifth loss value and the set of soft label loss values.

8. The method according to claim 3 or 5, characterized in that, Performing label correlation verification on the first soft label and the second soft label according to the dynamic distillation strategy to obtain a target label verification result, including: Obtaining the strategy threshold of the dynamic distillation strategy; Performing correlation metric calculation on the second soft label according to the first soft label to obtain a soft label correlation degree; Comparing the soft label correlation degree with the strategy threshold to obtain the target label verification result; wherein the target label verification result is a first label verification result or a second label verification result.

9. A recognition method for a cross-modal knowledge model, characterized in that, Including: Obtaining target media data to be recognized; Inputting the target media data into the cross-modal knowledge model according to any one of claims 1 to 8 for recognition to obtain an identification result of the target media data.

10. A training system for a cross-modal knowledge model, characterized in that, Applied to a knowledge distillation network model, the knowledge distillation network model includes a teacher model and a student model, and the system includes: A first processing unit, configured to obtain media data of a first modality and media data of a second modality, where the types of the first modality and the second modality are different; A second processing unit, configured to input the media data of the first modality into the teacher model for knowledge extraction to obtain a first soft label, and input the media data of the second modality into the student model for knowledge extraction to obtain a second soft label; A third processing unit, configured to perform bidirectional knowledge migration processing according to the first soft label and the second soft label to obtain a teacher migration loss value of the teacher model and a student migration loss value of the student model; A fourth processing unit, configured to update parameters of the knowledge distillation network model according to the teacher migration loss value and the student migration loss value to obtain a trained knowledge distillation network model; A fifth processing unit, configured to extract the cross-modal knowledge model from the trained knowledge distillation network model.