Neural Network Model Training Method, Device, Corresponding Equipment, and Interaction System
By constructing the auxiliary network and feature distribution alignment loss function, the problem of large differences in feature distribution in heterogeneous model transfer learning is solved, efficient knowledge transfer between heterogeneous models is achieved, and the classification performance of the target model is improved.
Patent Information
- Application Number
- CN202110184424.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-02-10
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2041-02-10
AI Technical Summary
In the transfer learning of heterogeneous model, especially in the source data is not available and the model structure is different, it is difficult to achieve effective knowledge transfer, resulting in large differences in feature distribution and difficulty in directly aligning, which affects the performance of the target model.
By constructing an auxiliary network, the output features of the source model are mapped to the auxiliary subspace of the same dimension as the output of the target model, and the parameters of the auxiliary network and the target model are adjusted through the feature distribution alignment loss function, thereby realizing transfer learning from the source model to the target model.
It effectively reduces the difference in feature distribution, improves the accuracy of classification prediction of the target model, and realizes efficient knowledge migration among heterogeneous models.
Smart Images

Figure CN114913362B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of neural networks, and in particular, to a method, apparatus, corresponding device, and interaction system for training a neural network model. Background Art
[0002] Deep neural networks (DNNs) have achieved great success in various tasks driven by large-scale labeled data (i.e., supervised training). However, the process of manually labeling data is both expensive and time-consuming. At the same time, a large amount of labeled data and well-evaluated datasets (such as Imagenet) may have general semantic features. The existing feature extractors can be further utilized to provide guidance for related tasks. For this purpose, the concept of transfer learning is proposed. Transfer learning belongs to a research field of machine learning, and its purpose is to transfer knowledge from a known source task to a new target task.
[0003] Although various solutions have been proposed for transfer learning in the prior art, these solutions have relatively strict constraints on the structures of the source model and the target model, as well as the model training data, and are not applicable to more common transfer scenarios.
[0004] Therefore, there is a need for an improved solution for training a target model using an existing model. Summary of the Invention
[0005] One technical problem to be solved by the present disclosure is to provide an improved solution for training a target model using an existing model. This solution realizes the transfer of model capabilities from the source model to the target model (especially between heterogeneous models) by establishing an auxiliary network (corresponding to an auxiliary feature subspace).
[0006] According to a first aspect of the present disclosure, there is provided a method for training a neural network model, including: constructing an auxiliary network, where the auxiliary network takes the output features of a source model as input and outputs auxiliary network output features having the same dimension as the output features of the target model; and aligning the distribution of the auxiliary network output features with the output features of the target model to achieve transfer learning from the source model to the target model.
[0007] According to a second aspect of the present disclosure, there is provided a method for predicting with a neural network model, including: obtaining information to be processed; and feeding the information to be processed into the target model obtained by the method according to the first aspect to obtain prediction information.
[0008] According to a third aspect of the present disclosure, there is provided a neural network model training device, including: an auxiliary network construction unit configured to construct an auxiliary network, the auxiliary network obtaining output features of a source model as input and outputting auxiliary network output features having the same dimension as the output of a target model; and a feature distribution alignment unit configured to align the distribution of the auxiliary network output features with the output features of the target model, so as to implement transfer learning from the source model to the target model.
[0009] According to a fourth aspect of the present disclosure, there is provided an intelligent device, including: an acquisition module configured to acquire information to be processed; and a networking module configured to upload the acquired information to be processed, the information to be processed being sent to the target model obtained as described in the first aspect to obtain a prediction result of the target model.
[0010] According to a fifth aspect of the present disclosure, there is provided a computing device, including: a processor; and a memory having executable code stored thereon, when the executable code is executed by the processor, causing the processor to execute the method as described in the above first and / or second aspect.
[0011] According to a sixth aspect of the present disclosure, there is provided a non-transitory machine-readable storage medium having executable code stored thereon, when the executable code is executed by a processor of an electronic device, causing the processor to execute the method as described in the above first and / or second aspect.
[0012] The neural network knowledge transfer solution of the present invention is particularly applicable to conversion between heterogeneous models lacking source data, and dynamically adapts source features to the target distribution by constructing an auxiliary network. Further, multiple subspaces and multi-distance measurements can be performed to formulate a framework for MAS, and MAS can achieve effective transfer learning by alleviating the overall transfer difficulty. When partial source data is available, the transfer effect can also be improved by constructing a second auxiliary network. BRIEF DESCRIPTION OF THE DRAWINGS
[0013] By describing the exemplary embodiments of the present disclosure in more detail in conjunction with the accompanying drawings, the above and other objects, features, and advantages of the present disclosure will become more apparent, wherein, in the exemplary embodiments of the present disclosure, the same reference numerals generally represent the same components.
[0014] Figure 1 Shows the differences between existing transfer learning and heterogeneous model transfer learning (HMTL).
[0015] Figure 2 Shows the basic operation principle of the present invention for knowledge transfer in HMTL.
[0016] Figure 3 Shows a neural network model training method according to an embodiment of the present invention.
[0017] Figure 4A -B shows an example of knowledge transfer between heterogeneous networks using an auxiliary network according to the present invention.
[0018] Figure 5A -B shows examples of MAS and M 2 AS based on multiple auxiliary subspaces.
[0019] Figure 6 Shows an example of using a dual auxiliary network when part of the source data is available.
[0020] Figure 7 Shows an example of visualizing the output feature distribution according to the present invention.
[0021] Figure 8 Shows a schematic diagram of the composition of a neural network model training device according to an embodiment of the present invention.
[0022] Figure 9 Shows a schematic diagram of the structure of a computing device that can be used to implement the above neural network model training and prediction methods according to an embodiment of the present invention. Detailed implementation manners
[0023] The preferred embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although the preferred embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be limited by the embodiments set forth herein. On the contrary, these embodiments are provided so that the present disclosure will be more thorough and complete, and can fully convey the scope of the present disclosure to those skilled in the art.
[0024] Deep neural networks (DNNs) have achieved great success in various tasks driven by large-scale labeled data (i.e., supervised training). However, the process of manually labeling data is both expensive and time-consuming. At the same time, a large amount of labeled data and well-evaluated data sets (such as Imagenet) may have general semantic features. The existing feature extractors can be further utilized to provide guidance for related tasks. For this purpose, the concept of transfer learning is proposed. Transfer learning belongs to a research field of machine learning, and its purpose is to transfer knowledge from known source tasks to new target tasks.
[0025] In the real world, transfer learning may face more challenging and practical problems. On the one hand, considering deployment and computational consumption issues, we often need to transfer knowledge from a large model to a smaller model. The model requirements indicate that the source model and the target model are likely to be heterogeneous. Under heterogeneous models, transfer learning is widely accepted as heterogeneous transfer learning (HTL). On the other hand, due to data privacy issues, the source data may not be available. In this case, we can only utilize the model learned from the source data. In the situation where the source data is not available and there is a heterogeneous network structure, both of the above two problems will occur simultaneously. Here, this problem can be referred to as heterogeneous model transfer learning (HMTL).
[0026] Figure 1 shows the differences between existing transfer learning and heterogeneous model transfer learning (HMTL). As shown in the figure, in a relatively simple transfer learning scenario (such as Figure 1 shown in the upper left), both the source data and the target data are available, and the target model G T can have the same or similar structure as the source model G S . At this time, transfer learning can be achieved through (data) domain adaptation. For example, the source model can be fine-tuned using the target data. Further, as Figure 1 shown by the dotted "source data" box in the upper right, when the source data is not available due to reasons such as sensitive data, then the existing knowledge can be used to accelerate the training of the target model G S by initializing the parameters (Init.) using the parameters of the source model G T .
[0027] Considering that large source models require a large amount of memory and are time-consuming during inference, in practice, knowledge needs to be transferred to small models. When the network structures between tasks are different, transfer learning becomes heterogeneous transfer learning (HTL). For this reason, in response to the two problems encountered, namely, the unavailability of source data and heterogeneous models, it is necessary to turn to heterogeneous model transfer learning (HMTL).
[0028] In Figure 1 heterogeneous model transfer learning (HMTL) shown in the lower left, due to the unavailability of the source data and the fact that the target model G T can have a different structure from the source model G S , for example, the target model G T is a model for a dedicated task (such as a dedicated image classification task) that is much smaller in scale than the source model G S . Due to the large distribution deviation between the source features and the target features, it is difficult to achieve direct alignment under such a large distribution deviation, especially for the output of heterogeneous models where the source data is not available. Therefore, how to achieve knowledge transfer from the source model G S to the target model GT Migration to improve the feature distribution is a huge problem faced by existing technologies.
[0029] To this end, the present invention proposes a method for knowledge transfer using an auxiliary subspace (AS, Auxiliary subspace). Specifically, the present invention constructs an auxiliary network, and the output feature distribution obtained therefrom is in the corresponding auxiliary subspace. By mapping the final output features of the source model to the same auxiliary subspace as that of the target model (i.e., the dimensions of the source model output features and the auxiliary subspace output features are the same) and adjusting the parameters, the alignment of the feature space dimensions is completed.
[0030] Figure 2 Shows the basic operation principle of the present invention for knowledge transfer in HMTL. As Figure 2 shown, the present invention constructs an auxiliary network following the source model for mapping source features to the target distribution. To maintain the existing knowledge in the source model, the network parameters of the source model will be frozen. The auxiliary network is trained to perform the mapping from the source distribution to the target distribution. As the training process progresses, the target model and the auxiliary network can be iteratively updated to facilitate the learning process of the neural network. The auxiliary network can be regarded as the dynamic part of the source model and enriches its transferable knowledge. As Figure 2 shown, the trained auxiliary network can achieve a better feature distribution for HMTL, and the target model G that obtains knowledge from the source model via the auxiliary network T can further optimize the feature distribution. For example, Figure 2 as shown in the lower right, a feature distribution with a smaller intra-class distance and a larger inter-class distance, thereby improving the classification prediction accuracy of the target model.
[0031] The auxiliary subspace scheme of the present invention can adapt to various situations and improve its performance. By constructing multiple auxiliary subspaces, the alignment between source features and target features can be enhanced to one-to-many and many-to-many settings, denoted as MAS and M 2 AS respectively. The present invention is particularly suitable for applications in visual classification tasks. The simple structure of AS can achieve state-of-the-art performance. MAS and M 2 AS can achieve further improvements compared to AS. If some unlabeled source data is available, the training method of the present invention can also construct a second (multiple) auxiliary network for the source data to further improve the prediction performance, such as the classification performance of the target model.
[0032] It should be understood that although the present invention is particularly suitable for implementing the heterogeneous model transfer learning scenario where the models are heterogeneous and the source data is unavailable, the principle of the present invention is also applicable to various scenarios of non-heterogeneous models and / or available source data.
[0033] Figure 3 A neural network model training method according to an embodiment of the present invention is shown. This method is used to implement transfer learning from the source model to the target model. In step S310, an auxiliary network is constructed. The auxiliary network can obtain the output features of the source model as input and output auxiliary network output features with the same dimension as the output of the target model. Generally, a classification model may include a feature extractor part for feature extraction and a classifier part for classification. The feature extractor can be composed of, for example, a multi-layer convolutional neural network (CNN), pooling, and activation function layers. The classifier usually consists of fully connected layers, which play a role in mapping the learned distributed feature representations to the sample label space. Here, what is obtained from the source model G as shown by the auxiliary network Figure 2 can be the distributed feature representations extracted by the feature extractor of the source model. In other words, at this time, G S specifically refers to the feature extraction part of the source model. S can specifically refer to the feature extraction part of the source model.
[0034] The auxiliary network can then obtain the feature output of the feature extraction part of the source model and output auxiliary network output features with the same dimension as the output of the target model. For example, the source model can be a large model trained with a large amount of data, while the target model is a heterogeneous small model for a specific task. At this time, the feature extraction part of the source model outputs, for example, a 1x2048-dimensional feature output, while the feature extraction part of the target model outputs, for example, a 1x512-dimensional output. For this reason, after the auxiliary network obtains the feature output of the feature extraction part of the source model, it can output a 1x512-dimensional output to align with the feature output of the target model as follows. Specifically, the auxiliary network can be implemented by a fully connected layer (FC). In another embodiment, the auxiliary network can also be implemented using a squeeze-and-excitation block structure (SEblock).
[0035] In step S320, the auxiliary network output features can be aligned with the output feature distribution of the target model to achieve transfer learning from the source model to the target model. Since the auxiliary network output features and the output features of the target model have the same dimension, for example, both are 1x512-dimensional, the difference (objective function) between the two output features can be directly calculated, and feature alignment can be achieved through the backpropagation algorithm.
[0036] Specifically, an alignment loss function of the auxiliary network output features and the output features of the target model can be calculated, and based on the alignment loss function, the parameters of the auxiliary network can be adjusted.
[0037] Subsequently, the above-mentioned auxiliary network can be used for training the target model. To this end, the method may further include: calculating a matching loss function between the output features of the auxiliary network and the output features of the target model; and adjusting the parameters of the target model based on the matching loss function. Similarly, the parameters of the target model adjusted herein may be the parameters of the feature extractor of the target model. The parameters of the feature extractor and the classifier of the target model can be further adjusted by a classification loss function.
[0038] To this end, the method may further include: calculating a classification loss function between the classification result obtained after classifying the output features of the (feature extractor of) the target model by the target classifier and the target data label; and adjusting the parameters of the target model and the target classifier according to the classification loss function. Here, for the convenience of expression, the feature extractor part of the target model can be denoted as G T , and the classifier part can be denoted as F T .
[0039] Here, the parameters of the auxiliary network G A can be first adjusted based on the alignment loss, and then the parameters of the target network (including the feature extraction part and the classifier part) can be adjusted based on the matching loss and the classification loss.
[0040] Specifically, the parameters of the source network can be frozen (frozen throughout the entire training process of the present invention), and the parameters of the auxiliary network G A can be adjusted based on the alignment loss function. After adjusting the parameters of the auxiliary network based on the alignment loss function, the parameters of the auxiliary network are fixed, and the parameters of the target model and the target classifier are adjusted based on the matching loss function and the classification loss function. The adjustment of the parameters of the auxiliary network and the parameters of the target model and the target classifier can be iterated to enable the auxiliary network to learn the knowledge of the source network round by round and transfer the knowledge to the target network round by round.
[0041] To improve the training efficiency, before performing the above parameter adjustment, the target data and the target data label can be used to pre-train the target model (feature extractor) G T and the target classifier F T to obtain the basic parameters of G T and F T .
[0042] Here, since in many scenarios the source data is unavailable (e.g., due to the sensitivity of the source data), it is necessary to use the target data for training the auxiliary network and knowledge transfer. At this time, the training method of the present invention may further include: respectively inputting the target data into the source model and the target model to obtain the output features of the source model and the output features of the target model obtained based on the target data, wherein the output features of the source model are used as the input of the auxiliary network. In other words, as described above, the parameter adjustment based on each loss function is achieved based on the prediction for the target data.
[0043] As described above, the present invention can achieve knowledge transfer between heterogeneous network models through AS (auxiliary subspace). The AS scheme realizes knowledge transfer by aligning the output features of the auxiliary network and the target network. Further, the present invention may further include a MAS (one-to-many AS) scheme, and M 2 AS (many-to-many AS) scheme to achieve more sufficient knowledge transfer learning through alignment in more dimensions.
[0044] To this end, when implemented as the MAS scheme, the model training method of the present invention may include: respectively aligning the output features of the auxiliary network with the output feature distributions of multiple layers of the target model to achieve transfer learning from the source model to the target model. Further, respectively aligning the output features of the auxiliary network with the output feature distributions of multiple layers of the target model includes: transforming the output features of the auxiliary network to have the same dimension distribution as the output features of each layer of the target model; calculating the sum of the mean square errors between the transformed output features of the auxiliary network and the output features of each layer of the target model as the second alignment loss function; and adjusting the parameters of the auxiliary network based on the second alignment loss function.
[0045] Further, when implemented as M 2 AS scheme, the method may further include: respectively aligning the output features of the auxiliary network and the output features of multiple layers of the source model with the output feature distributions of multiple layers of the target model to achieve transfer learning from the source model to the target model.
[0046] When at least part of the source data is available, for example, when at least part of the source data without labels can be obtained, the second auxiliary network can also be trained by using the features obtained from the source data. Different from the auxiliary network trained in the target data domain which is used to map the disordered source feature distribution into an ordered distribution, the second auxiliary network trained in the source data domain can be used to map the ordered source feature distribution into a disordered distribution. To this end, the training method of the present invention may further include: constructing a second auxiliary network; using part of the source data to input into the source model and the target model respectively, so as to obtain the source data output features of the source model and the source data output features of the target model obtained from the part of the source data respectively, wherein the source data output features of the source model are used as the input of the second auxiliary network; calculating a second alignment loss function between the output features of the second auxiliary network and the source data output features of the target model; and adjusting the parameters of the auxiliary network based on the second alignment loss function.
[0047] Further, a matching loss function between the output features of the second auxiliary network and the source data output features of the target model can be calculated; and the parameters of the target model can be adjusted based on the matching loss function, thereby realizing the knowledge transfer from the second auxiliary network to the target model.
[0048] The following will be specifically described in combination with Figure 4A -B, Figure 5A -B and Figure 6 the details of MAS, M 2 AS and the scheme using the source data domain.
[0049] Figure 4A -B shows an example of knowledge transfer between heterogeneous networks using an auxiliary network according to the present invention. Specifically, Figure 4A a schematic diagram of knowledge transfer between networks is given, Figure 4B and specific training steps are given.
[0050] As mentioned above, the model training method of the present invention is particularly applicable to heterogeneous model transfer learning (HMTL) where the source data is unavailable and the source model and the target model are heterogeneous. In HMTL, we provide a pre-trained source model G S , and a target model G T to be optimized and a target classifier F T . There are N target samples where y i ∈{1,..., C T}. The category C TIs completely different from the category of the source task. For example, the source model is a 1000-class image classification model for identifying categories such as cats, dogs, birds, buildings, and cars, while the target model is a 200-class image classification model for identifying different birds. Although the source data is not available, the present invention still attempts to use the output of G S to guide the training process of G T .
[0051] The goal of HMTL is to obtain good prediction results of G S with the help of the source model G T . Here, G S (x i ) can be regarded as source features (obtained by training with target data), G T (x i ) can be regarded as target features (also obtained by training with target data), and G S (x i ) and G T (x i ) can be regarded as paired features for further alignment.
[0052] Under different network structures and distribution biases, it is difficult to obtain good results by directly aligning the outputs of G S and G T . As Figure 4A illustrated, G S has a more complex network structure compared to G T . For example, G S has more hidden layers and a more complex network structure to complete more diverse classification tasks. Therefore, the present invention constructs a new auxiliary subspace for improving source features for more effective distribution alignment. Since there is a situation where source data cannot be utilized, the present invention designs an auxiliary network G A to transfer source features to the auxiliary subspace. The input of G A is G S (x i ), and the output G A (G T (x i )) has the same dimension as G T (x i ). For example, G S (x i ) is 1x2048-dimensional, and G A (G T (x i )) is the same as G T (x i ), which is 1x512-dimensional. Therefore, the auxiliary network G AIt can be implemented as a fully connected layer (FC) or a squeeze-and-excitation block (SE block) structure.
[0053] Subsequently, G A (G T (x i )) can be aligned with G T (x i ) to reduce the distribution difference between the auxiliary subspace and the target features. For example, an alignment loss can be used to update G A , a matching loss can be used to update G T , and a classification loss, which is usually a cross-entropy loss function, can be used to update G T and F T . Thus, the knowledge of the source model is transferred to the target model via the auxiliary subspace.
[0054] To prevent the distribution difference from expanding, the present invention designs a set of training steps to synchronize the auxiliary network with the target model to improve the ability of the target model. The training steps are as Figure 4B shown, including three steps for iteratively updating G A and G T .
[0055] In step 1, the target model G T and the target classifier F T can be simply trained using the target data and its labels through a classification loss. This step sets the target data representation that the auxiliary network can utilize. The classification loss can be calculated based on the labeled target data using a cross-entropy loss function. In one embodiment, the cross-entropy loss function is as follows:
[0056]
[0057] where L CE represents the cross-entropy loss.
[0058] In step 2, G A can be trained by aligning the features of the auxiliary space and the target space with the help of an alignment loss. In this step, G A obtains knowledge from the target domain. Subsequently, the alignment function between G A (G T (x i )) and G T (x i ) can be obtained, for example, the mean squared error (MSE) for feature alignment:
[0059]
[0060] In step 3, G A and GS , retrain G using classification loss and matching loss T and F T In this step, the knowledge is transferred to the target model, and equation (1) can be used as the classification loss function and equation (2) can be used as the matching loss function. Repeat steps 2 and 3 until the training is completed.
[0061] These three steps form the basic training process of the present invention, are also the components of MAS described below, and can gradually transfer knowledge to the target task. Among these three steps, step 1 can obtain G T and F T Step 2 can obtain G with good performance through the alignment process. A In practice, step 2 can be repeated multiple times to obtain a good performance G A Finally, step 3 in G A and G S Modify G with the help of the output T The transferable knowledge can be obtained from G S Pulled to G A During the training process, G A The output of the source model can be transferred to an auxiliary subspace, which has smaller differences than the target features. Through the auxiliary subspace, the distance between features can also be reduced. Preferably, MSE is always used to calculate the matching loss, but other objective functions such as COS (cosine distance) loss function, MMD (maximum mean difference) loss function, or cross entropy loss function can be used to calculate the alignment loss.
[0062] Since the connection between the source model and the target model should have different weights for multiple layers of output, the weights can be learned during the automatic training process of the network. In the present invention, the auxiliary subspace can be further enhanced to multiple auxiliary subspaces through multiple connections. Multiple connections can be implemented on the layer features of the network.
[0063] Figure 5A -B shows the MAS and M based on multiple auxiliary subspaces 2 Example of AS. Figure 5A Three different alignment methods are shown. Auxiliary subspace (AS) achieves one-to-one alignment, MAS achieves one-to-many alignment, and M 2 AS implements many-to-many alignment. Specifically, in AS, the source model features can be mapped to the auxiliary subspace for direct migration. In MAS, obtain G S The last layer outputs G as input A , whose outputs need to be matched with the outputs of multiple layers of the target model respectively. 2In AS, not only the auxiliary subspace, but also the outputs of multiple layers of the source model need to be respectively matched with the outputs of multiple layers of the target model.
[0064] The present invention can implement multiple auxiliary subspaces through the one-to-many and many-to-many matching as described above. For the convenience of description, the total number of layers involved in the source model and the target model can be set as P and Q respectively, and p and q are the layer numbers in the source model and the target model respectively. Based on Figure 5A -B's description, can be used to represent the output of the source model from the beginning to the p-th layer, and is used to represent the output of the target model from the beginning to the q-th layer. Further, can be used to represent the output feature set of multiple layers of the source model, and is used to represent the output feature set of multiple layers of the target model.
[0065] In MAS, one-to-many matching between the output of and each layer output of can be performed. At this time, Equation (2) as the alignment function can be replaced by:
[0066]
[0067] where represents a network for realizing the same dimension between the p-th layer source feature and the q-th layer target feature. includes a linear interpolation layer and a convolutional layer with a convolution kernel of 1x1. 's output can be expressed as multiple auxiliary subspaces with different q values, which is also the reason why this method is called MAS. In MAS, P is a constant, and G A only obtains the output of the last layer of G S
[0068] Figure 5B shows the alignment details of MAS. As Figure 5B shown, the auxiliary network can obtain the output of the last layer of G S , for example, a 1x2048-dimensional feature, and output , for example, a 1x512-dimensional feature. This 1x512-dimensional feature can only be directly matched with the output feature of the last layer of G T which is also 1x512-dimensional. It has different dimensions from other layers of G T , as shown in the first three layer outputs of the row where Figure 5B is located in . At this time, the linear interpolation and convolution included in are required to transform the 1x512-dimensional feature layer in the auxiliary subspace to The previous layers have the same dimension respectively and perform respective matching, such as the matching shown in Equation (3).
[0069] Furthermore, considering that different layers of the source model can also output various semantic features for transfer, for this purpose, it can be further achieved by aligning the outputs of all selected layers in and to implement many-to-many matching (i.e., M 2 AS). In one embodiment, the alignment process can be calculated as follows:
[0070]
[0071]
[0072] Among them, 's output can represent a many-to-many auxiliary subspace (M 2 AS). The only difference between AS, MAS, and M 2 AS lies in the different alignment methods shown in Figure 5A . Since the connection weights between the source model and the target model can be learned during the automatic training of the network, only Equation (3) or Equation (5) needs to be used instead of Equation (2) for the AS matching loss function (for example, adjusting G A in step 2 and adjusting G T in step 3), then MAS or M 2 AS can be achieved. When the latent features are sensitive, it is preferred to use AS or MAS for transfer, otherwise M 2 AS can be used to achieve better transfer performance.
[0073] As mentioned above, the typical measurement value of MSE such as Equation (2) can be used as the alignment loss function for training the auxiliary network. In other embodiments, the present invention can also use loss functions such as COS (cosine distance) loss function, MMD (maximum mean discrepancy) loss function, or pseudo-label loss function. Among them, the cosine loss can be calculated as follows:
[0074]
[0075] Among them, cos(·) is the cosine similarity metric. MMD can be interpreted in the "reproducing kernel Hilbert space (RKHS)" to reduce the distribution difference through different kernels. To simplify the calculation, the identity mapping can be used as the kernel, and at this time, the calculation method of the alignment loss is as follows:
[0076]
[0077] Among them, ||·||1 represents the L1-norm. In addition, G S, G A and F T to obtain the prediction of the source model for the target data. Thus, the label of the target data can be represented as the pseudo-label predicted by the network. Using these pseudo-labels, reliable prediction can be achieved by calculating the cross-entropy loss as follows:
[0078]
[0079] In L CE , the annotation of the target data is required. Experiments have verified that the results are stable under different distance measurements, which indicates the general effectiveness of the auxiliary subspace.
[0080] As mentioned above, when at least part of the source data is available, for example, when at least part of the unlabeled source data can be obtained, the second auxiliary network can also be trained by the features obtained from the source data. Figure 6 shows an example of using a dual auxiliary network when part of the source data is available. When part of the unlabeled source data can be obtained in the model scenario, another auxiliary network can be generated to utilize this source data and provide more transferable knowledge to improve the transfer performance. In the present invention, the auxiliary network is used to map the source feature space distribution to the auxiliary subspace. When part of the source data is available, two auxiliary networks are constructed and and as Figure 6 shown, different domain data are used for training.
[0081] Based on the target data, is used to map the disordered source feature distribution into an ordered distribution. Based on the source data, is used to map the ordered source feature distribution into a disordered distribution. These two opposite mapping functions and can both provide information about the relationship between G S and G T . For this purpose, Figure 4B in the training steps shown in step 2 can train the two auxiliary networks independently, while steps 1 and 3 can remain unchanged. With the additional knowledge generated by the additional auxiliary subspace, it can be transferred to the target network G Figure 5B with the help of the 1x1 convolutional layer and the matching loss shown in T .
[0082] Although shown in different figures, it should be understood that when the source data is available, the MAS and M 2 AS schemes can also further obtain the knowledge of the source model through the second auxiliary subspace.
[0083] As can be seen from the above, through the introduction of the auxiliary network, the present invention can improve the classification and prediction ability of the network through multiple iterative adjustments of the parameters of the auxiliary network and the target network. Figure 7 An example of visualizing the output feature distribution according to the present invention is shown. For the convenience of explanation, 5 categories are randomly selected from the test set for visualization (labeled as 1, 2, 3, 4, 5 in each figure respectively). It should be understood that in actual use, each model can predict more or fewer classifications.
[0084] Figure 7 The upper left shows the visualization of the output feature distribution of the source model for the target data. Since the source model itself is not a model for classifying the target model classification, as shown in the figure, categories 1-5 are basically aliased together. After the first round of iteration, the auxiliary network has learned part of the feature extraction ability. Figure 7 The upper right shows the visualization of the output feature distribution of the auxiliary network. It can be seen that the intra-class distances of the features of categories 1-5 start to approach, and the inter-class distances are pulled apart.
[0085] After multiple rounds of iteration, Figure 7 The visualization of the output feature distribution presented in the lower left shows that the auxiliary network not only has better feature distinguishability than the source model, but also the feature distribution is closer to the feature distribution of the target domain model.
[0086] Figure 7 The lower right shows the visualization of the feature distribution of the finally output target model. It can be seen that the intra-class distances of the features of categories 1-5 are close, and the inter-class distances are far apart.
[0087] After obtaining the trained target model based on transfer learning, the target model can be used for prediction. To this end, the present invention can also be implemented as a neural network model prediction method, including: obtaining information to be processed; and sending the information to be processed into the target model obtained by the method as described above to obtain prediction information.
[0088] In a preferred embodiment, the obtained model can be a classification model. For this purpose, the information to be processed can be information to be classified, and the target model is a model for classifying input information. The source model and the target model of the present invention can be particularly implemented as a picture classification model. Among them, the source model is a large model that can classify more picture types, and the target model is a model with a relatively smaller scale compared to the source model, and can usually perform sub-classification on a certain (or certain) specific object. In one embodiment, the source model can be a model trained using a large picture library as source data and having, for example, 1000 prediction classifications, and can classify the objects included in the input picture. For example, the picture includes people, buildings, cats, dogs, and birds, etc. The target model can be a sub-prediction model (for example, 200 classifications) for performing a specific type (for example, birds). At this time, the target model can be trained through a small labeled bird picture library. With the help of the auxiliary space of the present invention, at least part of the feature extraction ability already learned in the source model can be transferred to the target model to help identify specific type pictures.
[0089] In addition to the picture classification scenario, the auxiliary subspace construction scheme of the present invention can also be applied to other heterogeneous model transfer scenarios. For example, classification model training and knowledge transfer can be performed on text, speech, and even video (essentially also pictures) data sets, and can be applied to many scenarios such as text annotation, text translation, speech recognition, and video tagging.
[0090] Figure 8 The composition schematic diagram of a neural network model training device according to an embodiment of the present invention is shown. As shown in the figure, the device 800 can include an auxiliary network construction unit 810, a feature distribution alignment unit 820, and an optional target model pre-training unit 830.
[0091] The auxiliary network construction unit 810 can be used to construct an auxiliary network. The auxiliary network takes the output features of the source model as input and outputs auxiliary network output features with the same dimension as the output of the target model. In some embodiments, the auxiliary network can be a fully connected layer or an SEblock. The feature distribution alignment unit 820 can be used to align the distribution of the auxiliary network output features with the output features of the target model to achieve transfer learning from the source model to the target model.
[0092] To improve the training speed, the target data can be pre-trained. For this purpose, the device 800 can also include a target model pre-training unit 830, which is used to pre-train the target model and the target classifier using the target data and the target data labels.
[0093] Feature distribution alignment may include parameter tuning for the auxiliary network and parameter tuning for the target network. To this end, the feature distribution alignment unit 820 may include: an auxiliary network adjustment subunit. This subunit is used to: calculate the alignment loss function between the output features of the auxiliary network and the output features of the target model; and based on the alignment loss function, adjust the parameters of the auxiliary network.
[0094] Further, the feature distribution alignment unit 820 may further include: a target model adjustment subunit. This subunit is used to: calculate the matching loss function between the output features of the auxiliary network and the output features of the target model, and the classification loss function of the target model; and based on the matching loss function and the classification loss function, adjust the parameters of the target model. The above-mentioned auxiliary network adjustment subunit and target model adjustment subunit may operate alternately to iteratively adjust the parameters of the auxiliary network and the parameters of the target model.
[0095] Further, in the case where part of the source data is available, the auxiliary network construction unit 810 may be used to: construct a second auxiliary network, where part of the source data is respectively input into the source model and the target model to obtain the source data output features of the source model and the source data output features of the target model obtained based on part of the source data, and the source data output features of the source model are used as the input of the second auxiliary network. Accordingly, the feature distribution alignment unit 820 may be used to: align the output features of the second auxiliary network with the source data output feature distribution of the target model to achieve transfer learning from the source model to the target model.
[0096] To implement MAS or M 2 AS, the auxiliary network construction unit 810 may be used to: align the output features of the auxiliary network with the output features of multiple layers of the target model respectively to achieve transfer learning from the source model to the target model; and / or align the output features of the auxiliary network and the output features of multiple layers of the source model with the output features of multiple layers of the target model respectively to achieve transfer learning from the source model to the target model.
[0097] The present invention may also be implemented as an intelligent device. The device includes: an acquisition module for acquiring information to be processed; and a networking module for uploading the acquired information to be processed, and the above-mentioned information to be processed is sent to the target model obtained as described above to obtain the prediction result of the target model.
[0098] The above-mentioned intelligent device can also be a device included in an interaction system. For this purpose, the present invention can also be implemented as an interaction system, including: a plurality of intelligent devices as described above; and a processing device on which a target model trained according to the present invention is arranged. In a small system, the processing device can be implemented as the central node of the system. In a large system, the processing device can be a server accessible by a plurality of intelligent devices via a wireless connection.
[0099] Figure 9 FIG. shows a schematic structural diagram of a computing device that can be used to implement the above-mentioned neural network model training and prediction method according to an embodiment of the present invention.
[0100] See Figure 9 , the computing device 900 includes a memory 910 and a processor 920.
[0101] The processor 920 can be a multi-core processor or can include multiple processors. In some embodiments, the processor 920 can include a general-purpose main processor and one or more special coprocessors, such as a graphics processing unit (GPU), a digital signal processor (DSP), and so on. In some embodiments, the processor 920 can be implemented using customized circuits, such as an application specific integrated circuit (ASIC) or a field programmable gate array (FPGA).
[0102] The memory 910 may include various types of storage units, such as system memory, read-only memory (ROM), and permanent storage devices. Among them, the ROM may store static data or instructions required by the processor 920 or other modules of the computer. The permanent storage device may be a readable and writable storage device. The permanent storage device may be a non-volatile storage device that does not lose the stored instructions and data even when the computer is powered off. In some embodiments, the permanent storage device employs a mass storage device (such as a magnetic or optical disk, flash memory) as the permanent storage device. In some other embodiments, the permanent storage device may be a removable storage device (such as a floppy disk, optical drive). The system memory may be a readable and writable storage device or a volatile readable and writable storage device, such as dynamic random access memory. The system memory may store some or all of the instructions and data required by the processor during operation. In addition, the memory 910 may include any combination of computer-readable storage media, including various types of semiconductor storage chips (DRAM, SRAM, SDRAM, flash memory, programmable read-only memory), and magnetic disks and / or optical disks may also be employed. In some embodiments, the memory 910 may include a removable storage device that is readable and / or writable, such as a compact disc (CD), read-only digital versatile disc (such as DVD-ROM, dual-layer DVD-ROM), read-only Blu-ray disc, ultra density optical disc, flash memory card (such as SD card, min SD card, Micro-SD card, etc.), magnetic floppy disk, and so on. The computer-readable storage media do not include carrier waves and instantaneous electronic signals transmitted wirelessly or by wire.
[0103] Executable code is stored on the memory 910, and when the executable code is processed by the processor 920, it can cause the processor 920 to execute the neural network training and prediction methods described above.
[0104] The neural network knowledge transfer solution according to the present invention has been described in detail above with reference to the accompanying drawings. The present invention is particularly applicable to the conversion between heterogeneous models lacking source data, and dynamically adapts the source features to the target distribution by constructing an auxiliary network. Further, multiple subspaces and multi-distance measurements can be performed to formulate the framework of MAS, and MAS can achieve effective transfer learning by alleviating the overall transfer difficulty. When some source data is available, the transfer effect can also be improved by constructing a second auxiliary network.
[0105] In addition, the method according to the present invention can also be implemented as a computer program or a computer program product, and the computer program or the computer program product includes computer program code instructions for performing the above steps defined in the above method of the present invention.
[0106] Alternatively, the present invention may also be implemented as a non-transitory machine-readable storage medium (or computer-readable storage medium, or machine-readable storage medium) having stored thereon executable code (or computer program, or computer instruction code), which when executed by a processor of an electronic device (or computing device, server, etc.) causes the processor to perform the various steps of the above-described method according to the present invention.
[0107] Those skilled in the art will also understand that the various exemplary logical blocks, modules, circuits, and algorithm steps described in connection with the disclosure herein can be implemented as electronic hardware, computer software, or a combination of both.
[0108] The flowcharts and block diagrams in the figures illustrate the architecture, functionality, and operation of possible implementations of systems and methods according to various embodiments of the present invention. In this regard, each block in the flowchart or block diagram may represent a module, segment, or portion of code that contains one or more executable instructions for implementing the specified logical function. It should also be noted that, in some alternative implementations, the functions noted in the blocks may occur in a different order than noted in the figures. For example, two consecutive blocks may in fact be executed substantially in parallel, or they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block of the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented by a dedicated hardware-based system that performs the specified function or operation, or by a combination of dedicated hardware and computer instructions.
[0109] The embodiments of the present invention have been described above. The above description is exemplary and not exhaustive, and is not limited to the disclosed embodiments. Many modifications and variations will be apparent to those of ordinary skill in the art without departing from the scope and spirit of the described embodiments. The choice of terms used herein is intended to best explain the principles of the embodiments, the practical application, or the improvement of the technology in the market, or to enable other ordinary skilled persons in the art to understand the embodiments disclosed herein.
Claims
1. A neural network model training method for heterogeneous model transfer learning, comprising: Constructing an auxiliary network, where the auxiliary network takes the output features of the source model as input and outputs auxiliary network output features with the same dimension as the output of the target model; Aligning the distribution of the auxiliary network output features with the output features of the target model to achieve transfer learning from the source model to the target model, where aligning the distribution of the auxiliary network output features with the output features of the target model includes: Freezing the parameters of the source model, calculating the alignment loss function between the auxiliary network output features and the output features of the target model, and Adjusting the parameters of the auxiliary network based on the alignment loss function; and Freezing the parameters of the source model and the auxiliary network parameters, calculating the matching loss function between the auxiliary network output features and the output features of the target model, and adjusting the parameters of the target model based on the matching loss function, where the source model and the target model are image classification models.
2. The method according to claim 1, further comprising: Calculating the classification loss function between the classification result obtained after classifying the output features of the target model by the target classifier and the target data label; and Adjusting the parameters of the target model and the target classifier according to the classification loss function.
3. The method according to claim 2, wherein, After adjusting the parameters of the auxiliary network based on the alignment loss function, fix the parameters of the auxiliary network and perform parameter adjustment of the target model and the target classifier based on the matching loss function and the classification loss function.
4. The method according to claim 1, comprising: Using the target data to input the source model and the target model respectively to obtain the output features of the source model and the output features of the target model obtained based on the target data respectively, where the output features of the source model are used as the input of the auxiliary network.
5. The method according to claim 1, further comprising: Constructing a second auxiliary network; Using part of the source data to input the source model and the target model respectively to obtain the source data output features of the source model and the source data output features of the target model obtained based on the part of the source data respectively, where the source data output features of the source model are used as the input of the second auxiliary network; Calculating the second alignment loss function between the output features of the second auxiliary network and the source data output features of the target model; and Adjusting the parameters of the auxiliary network based on the second alignment loss function.
6. The method according to claim 5, further comprising: Calculating the matching loss function between the output features of the second auxiliary network and the source data output features of the target model; and Adjusting the parameters of the target model based on the matching loss function.
7. The method according to claim 1, further comprising: Aligning the distribution of the auxiliary network output features with the output features of multiple layers of the target model respectively to achieve transfer learning from the source model to the target model.
8. The method according to claim 7, further comprising: Align the output features of the auxiliary network and the output features of multiple layers of the source model with the output feature distributions of multiple layers of the target model respectively, so as to achieve transfer learning from the source model to the target model.
9. A neural network model prediction method, comprising: Obtain information to be processed; And Send the information to be processed into the target model obtained by the method according to any one of claims 1-8 to obtain prediction information.
10. The method according to claim 9, wherein, The information to be processed is information to be classified, and the target model is a model for classifying input information.
11. A neural network model training device for heterogeneous model transfer learning, comprising: An auxiliary network construction unit for constructing an auxiliary network, the auxiliary network obtaining the output features of the source model as input and outputting auxiliary network output features with the same dimension as the output of the target model; And A feature distribution alignment unit for aligning the output feature distribution of the auxiliary network with the output feature distribution of the target model to achieve transfer learning from the source model to the target model Wherein, the feature distribution alignment unit includes: An auxiliary network adjustment subunit for: Freezing the parameters of the source model, calculating an alignment loss function between the output features of the auxiliary network and the output features of the target model; and Adjusting the parameters of the auxiliary network based on the alignment loss function, A target model adjustment subunit for: Freezing the parameters of the source model and the parameters of the auxiliary network, calculating a matching loss function between the output features of the auxiliary network and the output features of the target model, and a classification loss function of the target model; and Adjusting the parameters of the target model based on the matching loss function and the classification loss function, Wherein, the auxiliary network adjustment subunit and the target model adjustment subunit operate alternately to iteratively adjust the parameters of the auxiliary network and the parameters of the target model, Wherein, the source model and the target model are picture classification models.
12. The device according to claim 11, wherein, The auxiliary network construction unit is used for: Constructing a second auxiliary network, wherein part of the source data is respectively input into the source model and the target model to respectively obtain the source data output features of the source model and the source data output features of the target model obtained based on the part of the source data, the source data output features of the source model being used as the input of the second auxiliary network, and The feature distribution alignment unit is used for: Aligning the output feature distribution of the second auxiliary network with the source data output feature distribution of the target model to achieve transfer learning from the source model to the target model.
13. The device according to claim 11, wherein, The auxiliary network construction unit is used for: Aligning the output features of the auxiliary network with the output feature distributions of multiple layers of the target model respectively to achieve transfer learning from the source model to the target model; and / or Aligning the output features of the auxiliary network and the output features of multiple layers of the source model with the output feature distributions of multiple layers of the target model respectively to achieve transfer learning from the source model to the target model.
14. An intelligent device, comprising: An acquisition module for acquiring information to be processed; A networking module for uploading the information to be processed obtained, and the information to be processed is fed into the target model obtained by the method according to any one of claims 1-8 to obtain the prediction result of the target model.
15. The intelligent device according to claim 14, wherein, The information to be processed includes the information of the picture to be classified.
16. A computing device, comprising: A processor; And A memory storing executable code thereon, which when executed by the processor, causes the processor to execute the method according to any one of claims 1-10.
17. A non-transitory machine-readable storage medium storing executable code thereon, which when executed by a processor of an electronic device, causes the processor to execute the method according to any one of claims 1-10.
Citation Information
Patent Citations
Task-towed feature distillation deep neural network learning training method and system, and readable storage medium
CN112132268A