Process collaborative learning method based on multi-modal characteristic deviation Gaussian

Through the process collaborative learning method based on multimodal feature deviation Gaussianization, the problem of lack of deep interaction in offline knowledge distillation is solved, and the deep interaction between student networks and model compression efficiency is improved.

CN120012867AActive Publication Date: 2025-05-16HUAIYIN TEACHERS COLLEGE
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510102266.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-22
Publication Date
2025-05-16
Estimated Expiration
2045-01-22

AI Technical Summary

Technical Problem

The existing offline knowledge distillation methods lack deep interaction in the learning process, which affects the model compression performance.

Method used

Using a process collaborative learning method based on multimodal feature deviation Gaussianization, a process collaborative learning function is established to improve the deep interaction between student networks by constructing multi-stream student network groups and multimodal feature converters, and calculating and Gaussian multimodal feature deviations.

Benefits of technology

It improves the deep interaction between student networks, enhances feature representation performance, improves the ability to fight against abnormal samples, and improves model compression efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120012867A_ABST
    Figure CN120012867A_ABST
Patent Text Reader

Abstract

The invention discloses a process collaborative learning method based on multi-modal feature deviation Gaussian. According to the method, a multi-stream student network group and image features, text features and sound features of each multi-stream student network in the group are constructed. Then, the multi-modal feature converter converts each stream feature to obtain multi-modal features which can be adapted among the multi-stream student networks; next, calculating the deviation between the output characteristic of each sub-network of the student network and the output characteristic of each sub-network of other student networks, and carrying out Gaussian processing on the deviation by using segmented power-level conversion; and finally, establishing a process collaborative learning function for the feature signals using the multi-modal feature deviation Gaussian technology, and completing optimization of the multi-stream student network.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a process collaborative learning method based on Gaussianization of multimodal feature deviations. Background Art

[0002] With the widespread application of deep learning in various fields, model compression technology is essential to achieve real-time inference and reduce energy consumption on resource-constrained devices. Through model compression, large models can be converted into small models, making them easier to deploy on mobile devices, embedded systems or IoT devices, thereby promoting the practical application of deep learning technology in various scenarios. Knowledge distillation is a model compression method that transfers the knowledge of a large and complex model (teacher model) to a small and simple model (student model) in the hope of achieving similar performance to the large model, but with lower computational complexity and storage requirements.

[0003] The knowledge distillation method can be divided into online distillation and offline distillation. Offline distillation is to first train the teacher model, then fix its parameters and use it to guide the training of the student model. In online distillation, the teacher and student models are updated simultaneously, which can be regarded as a collaborative learning process with multiple deep neural networks, where any network can be used as a student model and the rest of the networks as teacher models.

[0004] Online distillation methods have attracted extensive attention in the field of deep learning research due to their good compression performance. However, the current offline knowledge distillation methods still lack deep interaction in the learning process, which affects the model compression performance. In order to solve the above problems, the present invention discloses a process collaborative learning algorithm based on Gaussianization of multimodal feature deviations to improve the deep interactive learning between student networks in the online distillation algorithm. Summary of the invention

[0005] In view of the lack of deep interaction in the learning process in current offline knowledge distillation methods, the present invention provides a process collaborative learning method based on Gaussianization of multimodal feature deviations to improve the deep interactive learning between student networks in online distillation algorithms.

[0006] The technical solutions adopted in the present invention are:

[0007] A process collaborative learning method based on Gaussianization of multimodal feature deviations includes the following steps:

[0008] S1: Construct a multi-stream student network group consisting of several multi-stream student networks. Each multi-stream student network includes an image sub-network, a text sub-network and a sound sub-network. Input the image sample, text sample and sound sample into the corresponding sub-network of each multi-stream student network respectively, and obtain the feature output of the image, text and sound of each multi-stream student network;

[0009] S2: Receive the feature output of each multi-stream student network through the constructed multi-modal feature converter, perform corresponding feature conversion, and output the converted multi-modal features;

[0010] S3: For each student network in the multi-stream student network group, calculate the deviation between the output features of each sub-network in each multi-stream student network and the output features of the corresponding sub-network in other multi-stream student networks, and Gaussianize the deviation using piecewise power level transformation;

[0011] S4: For the feature signal that is Gaussianized using multimodal feature deviations, a process collaborative learning function is established.

[0012] Furthermore, S1 specifically includes:

[0013] S11: Initialize a total of S multi-stream student networks. The s-th multi-stream student network is represented as in, is the image subnetwork with parameter θ, is the text sub-network with parameter μ, The parameters are The sound sub-network; S multi-stream student networks form a multi-stream student network group

[0014] S12: Randomly select a batch of samples N from the image sample set, text sample set and sound sample set respectively, where any sample is represented as x. i ,y i , z i , respectively input into the sub-network of the s-th multi-stream student network, and obtain the following features:

[0015]

[0016] Furthermore, the multimodal feature converter is a fully connected layer with multi-signal input and multi-output signals, the multi-signal input respectively receives the feature output of each sub-network of the multi-stream student network, and the multi-signal output respectively outputs the converted multimodal features.

[0017] Furthermore, S2 specifically includes:

[0018] S21: The multimodal feature converter is represented as Tran = {O w1 ,O w2 ,O w3 ,B w ,P w1 ,P w2 ,P w3}, where O w1 ,O w2 ,O w3Denote the characteristic input converter with parameters w1, w2, w3 respectively, B w is a multi-modal feature hybrid transformer with parameter w, P w1 ,P w2 ,P w3 The characteristic output terminal converters are represented by parameters w1, w2, and w3 respectively;

[0019] S22: The features first enter the feature input converter to prepare for entering the multi-modal hybrid converter. The features of each stream signal are converted as follows at the input:

[0020]

[0021]

[0022] Then, the signal features of each stream enter the multimodal feature hybrid converter for the following conversion:

[0023]

[0024] Finally, the mixed multimodal features F i The characteristic output converter performs the following decoupling operations:

[0025]

[0026] Furthermore, S3 specifically includes:

[0027] S31: Randomly select a multi-stream student network from the multi-stream student network group, represented as Any sample x in the image sample set, text sample set and sound sample set i ,y i , z i , are input into the sub-networks of the multi-stream student network M respectively, and the features are obtained After the multimodal feature converter of S2 is operated, the converted feature representation is obtained as follows:

[0028] S32: Calculate the deviation between the converted features of the multi-stream student network M and the s-th multi-stream student network, and the calculation formula is as follows:

[0029]

[0030] S33: Use piecewise power level transformation to Gaussianize the multimodal feature deviation, where the expression of Gaussianization of image feature deviation is as follows:

[0031]

[0032] In the formula, γ is an adjustable parameter;

[0033] The Gaussian expression of text feature deviation is as follows:

[0034]

[0035] The Gaussian expression of the sound feature deviation is as follows:

[0036]

[0037] Furthermore, S4 specifically includes:

[0038] S41: For the multi-stream student network M, for its image sub-network, establish the following process collaborative learning function:

[0039]

[0040] For its text sub-network, the following process collaborative learning function is established:

[0041]

[0042] For its sound sub-network, the following process collaborative learning function is established:

[0043]

[0044] S42: For the multi-stream student network M, the overall objective function of the collaborative learning process is established:

[0045] L=L I +L T +L V .

[0046] The present invention has the following beneficial effects:

[0047] (1) The present invention designs a collaborative learning algorithm that can achieve collaborative learning between multimodal features between student networks, thereby improving the deep interaction between student networks and enhancing the feature representation performance of student networks;

[0048] (2) The present invention uses the Gaussianization technology of the intermediate hidden layer features to correct the multimodal feature distribution, thereby improving the ability of the student network to resist abnormal samples;

[0049] (3) The online test of the present invention only uses a trained student network and does not introduce additional parameter capacity, which can improve the model compression efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0050] Figure 1 It is a flow chart of the present invention. DETAILED DESCRIPTION

[0051] The present invention will be further described below in conjunction with the accompanying drawings.

[0052] The present invention provides a process collaborative learning method based on Gaussianization of multimodal feature deviations, comprising the following steps:

[0053] S11: Initialize a total of S multi-stream student networks. The s-th multi-stream student network is represented as in, is the image subnetwork with parameter θ, is the text sub-network with parameter μ, The parameters are The sound sub-network; S multi-stream student networks form a multi-stream student network group

[0054] S12: Randomly select a batch of samples N from the image sample set, text sample set and sound sample set, where any sample is represented by x. i ,y i , z i , respectively input into the sub-network of the s-th multi-stream student network, and obtain the following features:

[0055]

[0056] S13: Construct a multimodal feature converter. The multimodal feature converter is a fully connected layer with multi-signal input and multi-output signals. The multi-signal input receives the feature output of each sub-network of the multi-stream student network respectively, and the multi-signal output outputs the converted multimodal features respectively.

[0057] S14: The multimodal feature converter is represented as Tran = {O w1 ,O w2 ,O w3 ,B w ,P w1 ,P w2 ,P w3}, where O w1 ,O w2 ,O w3 Denote the characteristic input converter with parameters w1, w2, w3 respectively, B w is a multi-modal feature hybrid transformer with parameter w, P w1 ,P w2 ,P w3 The characteristic output terminal converters are represented by parameters w1, w2, and w3 respectively;

[0058] S15: The features first enter the feature input converter to prepare for entering the multi-modal hybrid converter. The features of each stream signal are converted as follows at the input:

[0059]

[0060] Then, the signal features of each stream enter the multimodal feature hybrid converter for the following conversion:

[0061]

[0062] Finally, the mixed multimodal features F i The characteristic output converter performs the following decoupling operations:

[0063]

[0064] S16: Randomly select a multi-stream student network from the multi-stream student network group, represented as Any sample x in the image sample set, text sample set and sound sample set i ,y i , z i , are input into the sub-networks of the multi-stream student network M respectively, and the features are obtained After the multimodal feature converter of S2 is operated, the converted feature representation is obtained as follows:

[0065] S17: Calculate the deviation between the multi-stream student network M and the converted features of the s-th multi-stream student network. The calculation formula is as follows:

[0066]

[0067] S18: Use piecewise power level transformation to Gaussianize the multimodal feature deviation, where the expression of Gaussianization of image feature deviation is as follows:

[0068]

[0069] In the formula, γ is an adjustable parameter;

[0070] The Gaussian expression of text feature deviation is as follows:

[0071]

[0072] The Gaussian expression of the sound feature deviation is as follows:

[0073]

[0074] S19: For the multi-stream student network M, for its image sub-network, establish the following process collaborative learning function:

[0075]

[0076] For its text sub-network, the following process collaborative learning function is established:

[0077]

[0078] For its sound sub-network, the following process collaborative learning function is established:

[0079]

[0080] S20: For the multi-stream student network M, the overall objective function of collaborative learning is established:

[0081] L=L I +L T +L V .

[0082] The implementation method of the present invention is demonstrated by taking an image classification task as an example. Figure 1 This is a flowchart of the collaborative learning algorithm based on Gaussianization of multimodal feature deviations in the present invention. Figure 1 The specific steps are as follows:

[0083] (1): Initialize a total of S multi-stream student networks, and the s-th multi-stream student network is represented as in is the image subnetwork with parameter θ, is the text sub-network with parameter μ, The parameters are The sound sub-network of S multi-stream student networks constitutes a multi-stream student network group. A batch of samples are randomly selected from the image sample set, text sample set and sound sample set, the number is N. Any of the samples is represented as x i ,y i , z i , respectively input into the sub-network of the s-th multi-stream student network, and obtain the feature as well as

[0084] (2): The feature first enters the input converter, is transformed into Then it is sent to the multimodal feature hybrid converter to get F si , and finally the output converter decouples it into

[0085] (3): Randomly select a student network M from the multi-stream student network group, which extracts image features, text features, and sound features After the multimodal feature converter operation in the second step, Convert to Calculate the deviation between the multi-stream student network M and the transformed features of the s-th multi-stream student network The piecewise power level transformation is Gaussianized to

[0086] (4): For the multi-stream student network M, establish a process collaborative learning function L for its image sub-network I , establish a process collaborative learning function L for its text sub-network T , establish a process collaborative learning function L for its sound sub-network V Based on L I , L T and L V The overall objective function of process collaborative learning is established, and the parameters in the multi-stream student network M are optimized using the fastest Newton descent method.

[0087] The above description is only a preferred embodiment of the present invention. It should be pointed out that a person skilled in the art can make several improvements without departing from the principle of the present invention, and these improvements should also be regarded as within the protection scope of the present invention.

Claims

1. A process collaborative learning method based on Gaussianization of multimodal feature deviations, characterized by: The steps include: S1: Construct a multi-stream student network group consisting of several multi-stream student networks. Each multi-stream student network includes an image sub-network, a text sub-network and a sound sub-network. Input the image sample, text sample and sound sample into the corresponding sub-network of each multi-stream student network respectively, and obtain the feature output of the image, text and sound of each multi-stream student network; S2: Receive the feature output of each multi-stream student network through the constructed multi-modal feature converter, perform corresponding feature conversion, and output the converted multi-modal features; S3: For each student network in the multi-stream student network group, calculate the deviation between the output features of each sub-network in each multi-stream student network and the output features of the corresponding sub-network in other multi-stream student networks, and Gaussianize the deviation using piecewise power level transformation; S4: For the feature signal that is Gaussianized using multimodal feature deviations, a process collaborative learning function is established.

2. The process collaborative learning method based on Gaussianization of multimodal feature deviations as claimed in claim 1, characterized in that: S1 specifically includes: S11: Initialize a total of S multi-stream student networks. The s-th multi-stream student network is represented as in, is the image subnetwork with parameter θ, is the text sub-network with parameter μ, The parameters are The sound sub-network; S multi-stream student networks form a multi-stream student network group S12: Randomly select a batch of samples N from the image sample set, text sample set and sound sample set, where any sample is represented by x. i ,y i , z i , respectively input into the sub-network of the s-th multi-stream student network, and obtain the following features:

3. The process collaborative learning method based on Gaussianization of multimodal feature deviations as claimed in claim 1, characterized in that: The multimodal feature converter is a fully connected layer with multi-signal input and multi-output signals. The multi-signal input respectively receives the feature output of each sub-network of the multi-stream student network, and the multi-signal output respectively outputs the converted multimodal features.

4. The process collaborative learning method based on Gaussianization of multimodal feature deviations as claimed in claim 3, characterized in that: S2 specifically includes: S21: The multimodal feature converter is represented as Tran = {O w1 ,O w2 ,O w3 ,B w ,P w1 ,P w2 ,P w3 }, where O w1 ,O w2 ,O w3 Denote the characteristic input converter with parameters w1, w2, w3 respectively, B w is a multi-modal feature hybrid transformer with parameter w, P w1 ,P w2 ,P w3 The characteristic output terminal converters are represented by parameters w1, w2, and w3 respectively; S22: The features first enter the feature input converter to prepare for entering the multi-modal hybrid converter. The features of each stream signal are converted as follows at the input: Then, the signal features of each stream enter the multimodal feature hybrid converter for the following conversion: Finally, the mixed multimodal features F i The characteristic output converter performs the following decoupling operations:

5. The process collaborative learning method based on Gaussianization of multimodal feature deviations as claimed in claim 4, characterized in that: S3 specifically includes: S31: Randomly select a multi-stream student network from the multi-stream student network group, represented as Any sample x in the image sample set, text sample set and sound sample set i ,y i , z i , are input into the sub-networks of the multi-stream student network M respectively, and the features are obtained After the multimodal feature converter of S2 is operated, the converted feature representation is obtained as follows: S32: Calculate the deviation between the converted features of the multi-stream student network M and the s-th multi-stream student network, and the calculation formula is as follows: S33: Use piecewise power level transformation to Gaussianize the multimodal feature deviation, where the expression of Gaussianization of image feature deviation is as follows: In the formula, γ is an adjustable parameter; The Gaussian expression of text feature deviation is as follows: The Gaussian expression of the sound feature deviation is as follows:

6. The process collaborative learning method based on Gaussianization of multimodal feature deviations as claimed in claim 5, characterized in that: S4 specifically includes: S41: For the multi-stream student network M, for its image sub-network, establish the following process collaborative learning function: For its text sub-network, the following process collaborative learning function is established: For its sound sub-network, the following process collaborative learning function is established: S42: For the multi-stream student network M, the overall objective function of the collaborative learning process is established: L=L I +L T +L V .

Citation Information

Patent Citations

  • Video recommendation method and related device

    CN110609955A

  • Multi-modal pre-training method, device, equipment and medium

    CN114118417A

  • Semi-supervised medical image segmentation method and system based on mutual learning

    CN114418954A

  • Multimodal medical image conversion method and system based on knowledge distillation and adversarial attack

    CN116596910A

  • Multi-modal image super-resolution reconstruction method based on structured knowledge distillation

    CN117911246A