Process collaborative learning method based on multi-modal characteristic deviation Gaussian
Through the process collaborative learning method based on multimodal feature deviation Gaussianization, the problem of lack of deep interaction in offline knowledge distillation is solved, and the deep interaction between student networks and model compression efficiency is improved.
Patent Information
- Application Number
- CN202510102266.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-22
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2045-01-22
AI Technical Summary
The existing offline knowledge distillation methods lack deep interaction in the learning process, which affects the model compression performance.
Using a process collaborative learning method based on multimodal feature deviation Gaussianization, a process collaborative learning function is established to improve the deep interaction between student networks by constructing multi-stream student network groups and multimodal feature converters, and calculating and Gaussian multimodal feature deviations.
It improves the deep interaction between student networks, enhances feature representation performance, improves the ability to fight against abnormal samples, and improves model compression efficiency.
Smart Images

Figure CN120012867A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a process collaborative learning method based on Gaussianization of multimodal feature deviations. Background Art
[0002] With the widespread application of deep learning in various fields, model compression technology is essential to achieve real-time inference and reduce energy consumption on resource-constrained devices. Through model compression, large models can be converted into small models, making them easier to deploy on mobile devices, embedded systems or IoT devices, thereby promoting the practical application of deep learning technology in various scenarios. Knowledge distillation is a model compression method that transfers the knowledge of a large and complex model (teacher model) to a small and simple model (student model) in the hope of achieving similar performance to the large model, but with lower computational complexity and storage requirements.
[0003] The knowledge distillation method can be divided into online distillation and offline distillation. Offline distillation is to first train the teacher model, then fix its parameters and use it to guide the training of the student model. In online distillation, the teacher and student models are updated simultaneously, which can be regarded as a collaborative learning process with multiple deep neural networks, where any network can be used as a student model and the rest of the networks as teacher models.
[0004] Online distillation methods have attracted extensive attention in the field of deep learning research due to their good compression performance. However, the current offline knowledge distillation methods still lack deep interaction in the learning process, which affects the model compression performance. In order to solve the above problems, the present invention discloses a process collaborative learning algorithm based on Gaussianization of multimodal feature deviations to improve the deep interactive learning between student networks in the online distillation algorithm. Summary of the invention
[0005] In view of the lack of deep interaction in the learning process in current offline knowledge distillation methods, the present invention provides a process collaborative learning method based on Gaussianization of multimodal feature deviations to improve the deep interactive learning between student networks in online distillation algorithms.
[0006] The technical solutions adopted in the present invention are:
[0007] A process collaborative learning method based on Gaussianization of multimodal feature deviations includes the following steps:
[0008] S1: Construct a multi-stream student network group consisting of several multi-stream student networks. Each multi-stream student network includes an image sub-network, a text sub-network and a sound sub-network. Input the image sample, text sample and sound sample into the corresponding sub-network of each multi-stream student network respectively, and obtain the feature output of the image, text and sound of each multi-stream student network;
[0009] S2: Receive the feature output of each multi-stream student network through the constructed multi-modal feature converter, perform corresponding feature conversion, and output the converted multi-modal features;
[0010] S3: For each student network in the multi-stream student network group, calculate the deviation between the output features of each sub-network in each multi-stream student network and the output features of the corresponding sub-network in other multi-stream student networks, and Gaussianize the deviation using piecewise power level transformation;
[0011] S4: For the feature signal that is Gaussianized using multimodal feature deviations, a process collaborative learning function is established.
[0012] Furthermore, S1 specifically includes:
[0013] S11: Initialize a total of S multi-stream student networks. The s-th multi-stream student network is represented as in, is the image subnetwork with parameter θ, is the text sub-network with parameter μ, The parameters are The sound sub-network; S multi-stream student networks form a multi-stream student network group
[0014] S12: Randomly select a batch of samples N from the image sample set, text sample set and sound sample set respectively, where any sample is represented as x. i ,y i , z i , respectively input into the sub-network of the s-th multi-stream student network, and obtain the following features:
[0015]
[0016] Furthermore, the multimodal feature converter is a fully connected layer with multi-signal input and multi-output signals, the multi-signal input respectively receives the feature output of each sub-network of the multi-stream student network, and the multi-signal output respectively outputs the converted multimodal features.
[0017] Furthermore, S2 specifically includes:
[0018] S21: The multimodal feature converter is represented as Tran = {O w1 ,O w2 ,O w3 ,B w ,P w1 ,P w2 ,P w3}, where O w1 ,O w2 ,O w3Denote the characteristic input converter with parameters w1, w2, w3 respectively, B w is a multi-modal feature hybrid transformer with parameter w, P w1 ,P w2 ,P w3 The characteristic output terminal converters are represented by parameters w1, w2, and w3 respectively;
[0019] S22: The features first enter the feature input converter to prepare for entering the multi-modal hybrid converter. The features of each stream signal are converted as follows at the input:
[0020]
[0021]
[0022] Then, the signal features of each stream enter the multimodal feature hybrid converter for the following conversion:
[0023]
[0024] Finally, the mixed multimodal features F i The characteristic output converter performs the following decoupling operations:
[0025]
[0026] Furthermore, S3 specifically includes:
[0027] S31: Randomly select a multi-stream student network from the multi-stream student network group, represented as Any sample x in the image sample set, text sample set and sound sample set i ,y i , z i , are input into the sub-networks of the multi-stream student network M respectively, and the features are obtained After the multimodal feature converter of S2 is operated, the converted feature representation is obtained as follows:
[0028] S32: Calculate the deviation between the converted features of the multi-stream student network M and the s-th multi-stream student network, and the calculation formula is as follows:
[0029]
[0030] S33: Use piecewise power level transformation to Gaussianize the multimodal feature deviation, where the expression of Gaussianization of image feature deviation is as follows:
[0031]
[0032] In the formula, γ is an adjustable parameter;
[0033] The Gaussian expression of text feature deviation is as follows:
[0034]
[0035] The Gaussian expression of the sound feature deviation is as follows:
[0036]
[0037] Furthermore, S4 specifically includes:
[0038] S41: For the multi-stream student network M, for its image sub-network, establish the following process collaborative learning function:
[0039]
[0040] For its text sub-network, the following process collaborative learning function is established:
[0041]
[0042] For its sound sub-network, the following process collaborative learning function is established:
[0043]
[0044] S42: For the multi-stream student network M, the overall objective function of the collaborative learning process is established:
[0045] L=L I +L T +L V .
[0046] The present invention has the following beneficial effects:
[0047] (1) The present invention designs a collaborative learning algorithm that can achieve collaborative learning between multimodal features between student networks, thereby improving the deep interaction between student networks and enhancing the feature representation performance of student networks;
[0048] (2) The present invention uses the Gaussianization technology of the intermediate hidden layer features to correct the multimodal feature distribution, thereby improving the ability of the student network to resist abnormal samples;
[0049] (3) The online test of the present invention only uses a trained student network and does not introduce additional parameter capacity, which can improve the model compression efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0050] Figure 1 It is a flow chart of the present invention. DETAILED DESCRIPTION
[0051] The present invention will be further described below in conjunction with the accompanying drawings.
[0052] The present invention provides a process collaborative learning method based on Gaussianization of multimodal feature deviations, comprising the following steps:
[0053] S11: Initialize a total of S multi-stream student networks. The s-th multi-stream student network is represented as in, is the image subnetwork with parameter θ, is the text sub-network with parameter μ, The parameters are The sound sub-network; S multi-stream student networks form a multi-stream student network group
[0054] S12: Randomly select a batch of samples N from the image sample set, text sample set and sound sample set, where any sample is represented by x. i ,y i , z i , respectively input into the sub-network of the s-th multi-stream student network, and obtain the following features:
[0055]
[0056] S13: Construct a multimodal feature converter. The multimodal feature converter is a fully connected layer with multi-signal input and multi-output signals. The multi-signal input receives the feature output of each sub-network of the multi-stream student network respectively, and the multi-signal output outputs the converted multimodal features respectively.
[0057] S14: The multimodal feature converter is represented as Tran = {O w1 ,O w2 ,O w3 ,B w ,P w1 ,P w2 ,P w3}, where O w1 ,O w2 ,O w3 Denote the characteristic input converter with parameters w1, w2, w3 respectively, B w is a multi-modal feature hybrid transformer with parameter w, P w1 ,P w2 ,P w3 The characteristic output terminal converters are represented by parameters w1, w2, and w3 respectively;
[0058] S15: The features first enter the feature input converter to prepare for entering the multi-modal hybrid converter. The features of each stream signal are converted as follows at the input:
[0059]
[0060] Then, the signal features of each stream enter the multimodal feature hybrid converter for the following conversion:
[0061]
[0062] Finally, the mixed multimodal features F i The characteristic output converter performs the following decoupling operations:
[0063]
[0064] S16: Randomly select a multi-stream student network from the multi-stream student network group, represented as Any sample x in the image sample set, text sample set and sound sample set i ,y i , z i , are input into the sub-networks of the multi-stream student network M respectively, and the features are obtained After the multimodal feature converter of S2 is operated, the converted feature representation is obtained as follows:
[0065] S17: Calculate the deviation between the multi-stream student network M and the converted features of the s-th multi-stream student network. The calculation formula is as follows:
[0066]
[0067] S18: Use piecewise power level transformation to Gaussianize the multimodal feature deviation, where the expression of Gaussianization of image feature deviation is as follows:
[0068]
[0069] In the formula, γ is an adjustable parameter;
[0070] The Gaussian expression of text feature deviation is as follows:
[0071]
[0072] The Gaussian expression of the sound feature deviation is as follows:
[0073]
[0074] S19: For the multi-stream student network M, for its image sub-network, establish the following process collaborative learning function:
[0075]
[0076] For its text sub-network, the following process collaborative learning function is established:
[0077]
[0078] For its sound sub-network, the following process collaborative learning function is established:
[0079]
[0080] S20: For the multi-stream student network M, the overall objective function of collaborative learning is established:
[0081] L=L I +L T +L V .
[0082] The implementation method of the present invention is demonstrated by taking an image classification task as an example. Figure 1 This is a flowchart of the collaborative learning algorithm based on Gaussianization of multimodal feature deviations in the present invention. Figure 1 The specific steps are as follows:
[0083] (1): Initialize a total of S multi-stream student networks, and the s-th multi-stream student network is represented as in is the image subnetwork with parameter θ, is the text sub-network with parameter μ, The parameters are The sound sub-network of S multi-stream student networks constitutes a multi-stream student network group. A batch of samples are randomly selected from the image sample set, text sample set and sound sample set, the number is N. Any of the samples is represented as x i ,y i , z i , respectively input into the sub-network of the s-th multi-stream student network, and obtain the feature as well as
[0084] (2): The feature first enters the input converter, is transformed into Then it is sent to the multimodal feature hybrid converter to get F si , and finally the output converter decouples it into
[0085] (3): Randomly select a student network M from the multi-stream student network group, which extracts image features, text features, and sound features After the multimodal feature converter operation in the second step, Convert to Calculate the deviation between the multi-stream student network M and the transformed features of the s-th multi-stream student network The piecewise power level transformation is Gaussianized to
[0086] (4): For the multi-stream student network M, establish a process collaborative learning function L for its image sub-network I , establish a process collaborative learning function L for its text sub-network T , establish a process collaborative learning function L for its sound sub-network V Based on L I , L T and L V The overall objective function of process collaborative learning is established, and the parameters in the multi-stream student network M are optimized using the fastest Newton descent method.
[0087] The above description is only a preferred embodiment of the present invention. It should be pointed out that a person skilled in the art can make several improvements without departing from the principle of the present invention, and these improvements should also be regarded as within the protection scope of the present invention.
Claims
1. A process collaborative learning method based on Gaussianization of multimodal feature deviations, characterized by: The steps include: S1: Construct a multi-stream student network group consisting of several multi-stream student networks. Each multi-stream student network includes an image sub-network, a text sub-network and a sound sub-network. Input the image sample, text sample and sound sample into the corresponding sub-network of each multi-stream student network respectively, and obtain the feature output of the image, text and sound of each multi-stream student network; S2: Receive the feature output of each multi-stream student network through the constructed multi-modal feature converter, perform corresponding feature conversion, and output the converted multi-modal features; S3: For each student network in the multi-stream student network group, calculate the deviation between the output features of each sub-network in each multi-stream student network and the output features of the corresponding sub-network in other multi-stream student networks, and Gaussianize the deviation using piecewise power level transformation; S4: For the feature signal that is Gaussianized using multimodal feature deviations, a process collaborative learning function is established.
2. The process collaborative learning method based on Gaussianization of multimodal feature deviations as claimed in claim 1, characterized in that: S1 specifically includes: S11: Initialize a total of S multi-stream student networks. The s-th multi-stream student network is represented as in, is the image subnetwork with parameter θ, is the text sub-network with parameter μ, The parameters are The sound sub-network; S multi-stream student networks form a multi-stream student network group S12: Randomly select a batch of samples N from the image sample set, text sample set and sound sample set, where any sample is represented by x. i ,y i , z i , respectively input into the sub-network of the s-th multi-stream student network, and obtain the following features:
3. The process collaborative learning method based on Gaussianization of multimodal feature deviations as claimed in claim 1, characterized in that: The multimodal feature converter is a fully connected layer with multi-signal input and multi-output signals. The multi-signal input respectively receives the feature output of each sub-network of the multi-stream student network, and the multi-signal output respectively outputs the converted multimodal features.
4. The process collaborative learning method based on Gaussianization of multimodal feature deviations as claimed in claim 3, characterized in that: S2 specifically includes: S21: The multimodal feature converter is represented as Tran = {O w1 ,O w2 ,O w3 ,B w ,P w1 ,P w2 ,P w3 }, where O w1 ,O w2 ,O w3 Denote the characteristic input converter with parameters w1, w2, w3 respectively, B w is a multi-modal feature hybrid transformer with parameter w, P w1 ,P w2 ,P w3 The characteristic output terminal converters are represented by parameters w1, w2, and w3 respectively; S22: The features first enter the feature input converter to prepare for entering the multi-modal hybrid converter. The features of each stream signal are converted as follows at the input: Then, the signal features of each stream enter the multimodal feature hybrid converter for the following conversion: Finally, the mixed multimodal features F i The characteristic output converter performs the following decoupling operations:
5. The process collaborative learning method based on Gaussianization of multimodal feature deviations as claimed in claim 4, characterized in that: S3 specifically includes: S31: Randomly select a multi-stream student network from the multi-stream student network group, represented as Any sample x in the image sample set, text sample set and sound sample set i ,y i , z i , are input into the sub-networks of the multi-stream student network M respectively, and the features are obtained After the multimodal feature converter of S2 is operated, the converted feature representation is obtained as follows: S32: Calculate the deviation between the converted features of the multi-stream student network M and the s-th multi-stream student network, and the calculation formula is as follows: S33: Use piecewise power level transformation to Gaussianize the multimodal feature deviation, where the expression of Gaussianization of image feature deviation is as follows: In the formula, γ is an adjustable parameter; The Gaussian expression of text feature deviation is as follows: The Gaussian expression of the sound feature deviation is as follows:
6. The process collaborative learning method based on Gaussianization of multimodal feature deviations as claimed in claim 5, characterized in that: S4 specifically includes: S41: For the multi-stream student network M, for its image sub-network, establish the following process collaborative learning function: For its text sub-network, the following process collaborative learning function is established: For its sound sub-network, the following process collaborative learning function is established: S42: For the multi-stream student network M, the overall objective function of the collaborative learning process is established: L=L I +L T +L V .
Citation Information
Patent Citations
Video recommendation method and related device
CN110609955A
Multi-modal pre-training method, device, equipment and medium
CN114118417A
Semi-supervised medical image segmentation method and system based on mutual learning
CN114418954A
Multimodal medical image conversion method and system based on knowledge distillation and adversarial attack
CN116596910A
Multi-modal image super-resolution reconstruction method based on structured knowledge distillation
CN117911246A