Model migration optimization method, system, equipment and medium

By covariance matrix decomposition and population optimization of the transfer learning model, the filter fine-tuning layer is determined, which solves the problem of difficult to determine the filter fine-tuning, and improves the optimization efficiency and classification accuracy of the model.

CN120296487APending Publication Date: 2025-07-11GUIZHOU MINZU UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510262970.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-06
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

The existing transfer learning methods are difficult to directly determine which filters in the model need to be fine-tuned, making it difficult to guarantee classification accuracy.

Method used

By decomposing the covariance matrix of the optimization model, a decision subset is obtained and population optimization is performed on the decision subset, a set of filters that need to be fine-tuned is determined, and precise optimization and adjustment is performed.

Benefits of technology

The workload of model optimization is reduced, the optimization efficiency and classification accuracy of the model are improved, the knowledge of the source task set is fully utilized, and the classification performance of the target model in the training task set is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120296487A_ABST
    Figure CN120296487A_ABST
Patent Text Reader

Abstract

The invention discloses a model migration optimization method, system and device and a medium. The method comprises the steps of obtaining a training task set; performing covariance matrix decomposition on the to-be-optimized model by utilizing the training task set to obtain a decision subset; wherein the to-be-optimized model is a model obtained by training a source task set, and the decision subset represents a set composed of to-be-adjusted filters in a network layer of the to-be-optimized model; performing population optimization on the decision subset to obtain a target decision subset; and performing optimization adjustment on the to-be-optimized model based on the target decision subset to obtain a target model. The problem that in existing transfer learning, which filters in the model need to be finely adjusted are difficult to directly determine, so that the model classification accuracy obtained through transfer learning is difficult to guarantee is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Background Art

[0002] Transfer learning aims to utilize the knowledge obtained from one task (source domain) to improve the learning performance of another task (target domain), and the transfer of knowledge is as shown Figure 1 below. It is usually used to solve the problem of insufficient training data in the target task and has achieved remarkable success in various fields. The model-based transfer method completes the transmission by constructing a model with shared parameters from the source domain (source task) and the target domain (target task), as shown Figure 2 below. The parameters learned in the source task with a large amount of training data are transferred as "knowledge" to the target task and partial parameters are fine-tuned to assist the learning of the target task.

[0003] As the scale and structural complexity of the model continue to increase, when the pre-trained model is transferred from the source task to the target task, some filters need to be fine-tuned. However, it is difficult to directly determine which filters need to be fine-tuned due to the large scale of the number of filters in the model, making it difficult to guarantee the classification accuracy of the model obtained by transfer learning. Summary of the Invention

[0004] In order to overcome the problem in existing transfer learning that it is difficult to directly determine which filters in the model need to be fine-tuned, making it difficult to guarantee the classification accuracy of the model obtained by transfer learning, this application provides a model transfer optimization method, system, device and medium.

[0005] In a first aspect, to solve the above technical problem, this application provides a model transfer optimization method, including:

[0006] Obtain a training task set;

[0007] Perform covariance matrix decomposition on the model to be optimized using the training task set to obtain a decision subset; wherein, the model to be optimized is a model trained using a source task set, and the decision subset represents a set composed of the filters to be adjusted in the network layer of the model to be optimized;

[0008] Perform population optimization on the decision subset to obtain a target decision subset;

[0009] Optimize and adjust the model to be optimized based on the target decision subset to obtain a target model.

[0010] In a second aspect, this application also provides a model transfer optimization system, including:

[0011] An acquisition module, configured to obtain a training task set;

[0012] A decomposition module is used to perform covariance matrix decomposition on the model to be optimized using a training task set, obtaining decision subsets. Here, the model to be optimized is a model trained using a source task set, and the decision subsets represent the set of filters to be adjusted in the network layer of the model to be optimized.

[0013] A population optimization module is used to perform population optimization on the decision subsets to obtain target decision subsets.

[0014] An optimization and adjustment module is used to optimize and adjust the model to be optimized based on the target decision subsets to obtain a target model.

[0015] In a third aspect, the present application also provides a computing device, including a memory, a processor, and a program stored on the memory and running on the processor. When the processor executes the program, it implements the steps of a model migration optimization method as described above.

[0016] In a fourth aspect, the present application also provides a computer-readable storage medium. Instructions are stored in the computer-readable storage medium. When the instructions run on a terminal device, the terminal device is caused to execute the steps of a model migration optimization method.

[0017] The beneficial effects of the present application are as follows: First, covariance matrix decomposition is performed on the model to be optimized using a training task set, obtaining multiple decision subsets composed of filters that need to be adjusted. Here, the model to be optimized is a model trained using a source task set, and the decision subsets represent the set of filters to be adjusted in the network layer of the model to be optimized. Population optimization is performed on the decision subsets to finely adjust the filters that need to be adjusted, obtaining optimal target decision subsets. Then, targeted and precise optimization and adjustment are performed on the model to be optimized based on the target decision subsets to obtain a target model. In this way, it is not necessary to adjust all the filters in the model. Only partial filters in the model need to be precisely fine-tuned through transfer learning, which can not only reduce the optimization workload of the model and improve the optimization efficiency of the model, but also improve the classification accuracy of the optimized target model. At the same time, it can enable the knowledge of the source task set corresponding to the model to be optimized to be fully utilized in the model training of the training task set, thereby improving the classification performance of the obtained target model on the training task set and increasing the classification accuracy rate. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] Figure 1 is a schematic diagram of transfer learning;

[0019] Figure 2 is a schematic diagram of the fine-tuning process of the model in transfer learning;

[0020] Figure 3 is a schematic flowchart of a model migration optimization method shown in an exemplary embodiment of the present application;

[0021] Figure 4 In an exemplary embodiment of the present application, it is a flowchart for determining the filter fine-tuning layer of the model to be optimized based on the covariance matrix;

[0022] Figure 5 In an exemplary embodiment of the present application, it is a flowchart for clustering the eigenvalues and eigenvectors;

[0023] Figure 6 It is a schematic structural diagram of a model migration optimization system shown in an exemplary embodiment of the present application. Detailed implementation manners

[0024] The following embodiments are further explanations and supplements to the present application and do not constitute any limitation to the present application.

[0025] Regarding the problem of difficult selection of pre-trained model parameters in transfer learning, existing research can be classified into the following categories:

[0026] Traditional methods: Traditional methods rely on expert experience to select fine-tuning parameters. Either all parameters of the pre-trained model are fine-tuned, or the previous layers in the network are frozen and the last few layers of the network are fine-tuned. However, fine-tuning all parameters will lead to overfitting problems in the target task, and for the method of only fine-tuning the last few layers, it is difficult to determine the number of fine-tuned and frozen layers, resulting in poor performance of model fine-tuning.

[0027] Methods based on policy networks: The fine-tuning method optimized based on a policy network uses an additional neural network to be trained to determine which parameters in the pre-trained model need to be fine-tuned. However, this method requires training a large number of parameters. Due to the discreteness and non-differentiability of the fine-tuning parameter selection problem, approximately applying gradient descent brings challenges, ultimately affecting the accuracy of the fine-tuning scheme.

[0028] Methods based on evolutionary optimization: The methods based on evolutionary optimization regard each fine-tuning scheme as an individual in the population, and through iterative operations such as crossover, mutation, and selection, search for an optimal fine-tuning scheme. This method models the problem of whether to fine-tune a single layer or filter as an optimization model of a binary mask, and designs an optimization algorithm to solve the optimal binary mask through continuous iteration to determine whether to fine-tune the layer or filter.

[0029] It can be seen that although existing methods effectively improve the fine-tuning performance, they still essentially search for a suitable fine-tuning scheme by trial and error, and it is difficult to represent the relationship between the pre-trained model parameters and the target task. When the scale (number of layers and number of parameters) of the pre-trained model continues to increase, the difficulty of the search process also increases, and its fine-tuning accuracy is difficult to guarantee.

[0030] To solve the above problems, embodiments of the present application provide a model migration optimization method, system, device, and medium, which will be described in detail below.

[0031] A model migration optimization method provided by an embodiment of the present application can be specifically executed by a server. It should be noted that the server can be an independent server or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, Content Delivery Network (CDN), and big data and artificial intelligence platforms. There is no limitation here.

[0032] Please refer to Figure 3 , Figure 3 which shows a model migration optimization method according to an exemplary embodiment of the present application. As Figure 3 shown, the present application provides a model migration optimization method, including:

[0033] Step S31, obtaining a training task set;

[0034] Step S32, performing covariance matrix decomposition on the model to be optimized using the training task set to obtain a decision subset; wherein, the model to be optimized is a model trained using a source task set, and the decision subset represents a set composed of filters to be adjusted in the network layer of the model to be optimized;

[0035] Step S33, performing population optimization on the decision subset to obtain a target decision subset;

[0036] Step S34, performing optimization adjustment on the model to be optimized based on the target decision subset to obtain a target model.

[0037] In this embodiment provided by the present application, a model migration optimization method is disclosed. First, the covariance matrix of the model to be optimized is decomposed by a training task set to obtain multiple decision subsets composed of filters that need to be adjusted. Here, the model to be optimized is a model trained using a source task set, and the decision subset represents a set composed of the filters to be adjusted in the network layer of the model to be optimized. Then, population optimization is performed on the decision subsets to finely tune the filters that need to be adjusted, so as to obtain the optimal target decision subset. Next, targeted and precise optimization adjustment is performed on the model to be optimized based on the target decision subset to obtain the target model. In this way, instead of adjusting all the filters in the model, only partial filters in the model are precisely fine-tuned through transfer learning, which can not only reduce the optimization workload of the model and improve the optimization efficiency of the model, but also improve the classification accuracy of the optimized target model. At the same time, the knowledge of the source task set corresponding to the model to be optimized can be fully utilized in the model training of the training task set, thereby improving the classification performance of the obtained target model on the training task set and increasing the classification accuracy rate.

[0038] In another exemplary embodiment provided by the present application, when performing optimization adjustment on the model to be optimized based on the target decision subset, the filter parameters corresponding to the target decision subset are replaced correspondingly in the model to be optimized to form the optimized target model, reducing the training workload of the model and thus improving the optimization efficiency of the model.

[0039] Optionally, decomposing the covariance matrix of the model to be optimized by using the training task set to obtain the decision subset includes:

[0040] Constructing the covariance matrix between the layer filters of each network layer in the model to be optimized and the training task set;

[0041] Determining the filter fine-tuning layer of the model to be optimized based on the covariance matrix;

[0042] Performing eigenvalue decomposition on the covariance matrix corresponding to the filter fine-tuning layer to obtain eigenvalues and eigenvectors;

[0043] Performing clustering processing on the eigenvalues and eigenvectors to obtain the decision subset of the filter fine-tuning layer.

[0044] In this embodiment provided by the present application, first, a covariance matrix between the layer filters of each network layer in the model to be optimized and the training task set is constructed to measure the correlation between the layer filters and the training task set, so as to determine the filter fine-tuning layer in the model to be optimized that needs to be fine-tuned based on the correlation corresponding to the covariance matrix. Then, eigenvalue decomposition is performed on the covariance matrix corresponding to the filter fine-tuning layer to obtain eigenvalues and eigenvectors, and clustering processing is performed on the eigenvalues and eigenvectors to obtain the decision subset of the filter fine-tuning layer, so as to find the filters to be adjusted in the filter fine-tuning layer that needs to be fine-tuned, which is convenient for subsequent precise optimization and adjustment of the filters to be adjusted in the decision subset, thereby improving the accuracy of model optimization and further improving the classification accuracy of the target model obtained by optimization. Among them, the clustering processing technology can be K-means clustering, and the eigenvalues and eigenvectors are the characteristics of the filters of the filter fine-tuning layer.

[0045] In another exemplary embodiment provided by the present application, since it is difficult to directly determine which filters need to be fine-tuned due to the large scale of the number of filters in the model to be optimized, a covariance matrix decomposition strategy is designed to decompose the large-scale filter adjustment problem into multiple sub-problems by analyzing the correlation between the filters of the model and the training task set, and this sub-problem is the optimization and adjustment problem of the filter fine-tuning layer.

[0046] In order to select the filter layer in the model to be optimized that has a greater impact on the training task set, first define the filter output of the model to be optimized and the characteristics of the training task set. Let the filter output of each network layer of the model to be optimized be The characteristics of the training task set are D = [d1, d2,..., d m , where i represents the i-th network layer, n i represents the number of filters in the i-th network layer, and m represents the number of characteristics of the training task set.

[0047] Then, a covariance matrix COV(F i , D) between all the filters in the i-th network layer and the characteristics of the training task set is constructed, and its form is:

[0048]

[0049] Among them, cov(f ij , d k ) represents the covariance between the j-th filter in the i-th network layer of the model to be optimized and the k-th characteristic of the training task set.

[0050] Optionally, determining the filter fine-tuning layer of the model to be optimized based on the covariance matrix includes:

[0051] Calculate the first covariance value of the covariance matrix, and determine the feature correlation between the corresponding layer filter and the training task set based on the first covariance value;

[0052] When the feature correlation is greater than the first threshold, regard the corresponding layer filter as the filter fine-tuning layer.

[0053] In this embodiment provided by the present application, since there is a significant correlation between the model filter parameters and the training data, and this correlation can be quantified by calculating the covariance between the two, therefore, the first covariance value of the covariance matrix corresponding to each layer filter calculated can reflect the correlation between the corresponding layer filter and the training task set, that is, the higher the first covariance value, the more relevant the layer filter is. Thus, based on the first covariance value, the feature correlation between the corresponding layer filter and the training task set can be directly determined, and when the feature correlation is greater than the first threshold, the corresponding layer filter is regarded as the filter fine-tuning layer to find out the filter fine-tuning layer that meets the feature correlation requirements with the training task set from the model to be trained, which is convenient for quickly and accurately positioning and adjusting the filters to be adjusted in the filter fine-tuning layer later, reducing the useless adjustment of the layer filters to be learned, being able to accurately adjust the filters to be adjusted in the model to be optimized later, and improving the model optimization efficiency based on transfer learning.

[0054] Please refer to Figure 4 , Figure 4 which is a flowchart for determining the filter fine-tuning layer of the model to be optimized based on the covariance matrix in an exemplary embodiment of the present application. Figure 4 In , the pre-trained model obtained by training with the source task set is the model to be optimized, which includes multiple network layers. For each network layer, by constructing the covariance matrix between the layer filter of the network layer and the target task (training task set), and calculating the first covariance value of the covariance matrix, the feature correlation between the layer filter and the target task (training task set) is determined. If the feature correlation meets the requirements, then the layer filter is used as the filter fine-tuning layer highly relevant to the target task (training task set) for fine-tuning.

[0055] In this embodiment provided by the present application, according to the covariance matrix, the correlation strength between the shallow part of the model to be optimized and the training task set can be obtained. It can be known from the papers of three authors, Yosinski, Girshick, and Zeiler and Ferhus, that the shallow layer of the convolutional layer mainly learns low-level features such as edges and textures, and the low-level features learned by the shallow part are often general. The deep part mainly learns high-level abstract features, and these features are more targeted at specific tasks. Therefore, the present application selects to freeze the shallow part of the model to be optimized and optimize the deep network so that the model to be optimized can better adapt to the training task set.

[0056] Optionally, clustering processing is performed on the eigenvalue and eigenvector to obtain a decision subset of the filter fine-tuning layer, including:

[0057] Sort the eigenvalues of the filter fine-tuning layer in descending order to obtain a sorting result with values from large to small;

[0058] Obtain a preset number of eigenvalues from the sorting result in the order from front to back as target eigenvalues, and determine the eigenvectors corresponding to the target eigenvalues as target vectors;

[0059] Obtain the task correlation between every two target vectors;

[0060] Take the task correlation greater than the second threshold as the target correlation;

[0061] Cluster the target vectors corresponding to the target correlations with the same direction to form a subspace;

[0062] Obtain the vector correlation between the target vectors in the subspace, and take the subspace with the vector correlation greater than the third threshold as the decision subset of the filter fine-tuning layer.

[0063] In this embodiment provided by the present application, since the filter parameters with weak features have little influence on the classification of the input training task set, only a preset number of eigenvalues from the front to the back among all the eigenvalues sorted in the filter fine-tuning layer are used as target eigenvalues, and the eigenvectors corresponding to the target eigenvalues are determined as target vectors to select the target vectors with stronger features in the filter fine-tuning layer. Then, obtain the task correlation between every two target vectors, cluster the target vectors corresponding to the target correlations with the same direction to form a subspace, and obtain the vector correlation between the target vectors in the subspace. Take the subspace with the vector correlation greater than the third threshold as the decision subset of the filter fine-tuning layer. In this way, the target vectors in the decision subset have the characteristics of strong features, consistent directions, and both task correlation and vector correlation meeting the correlation requirements of the training task set, realizing the targeted and accurate selection of the decision subset, thereby improving the classification accuracy of the target model obtained by subsequent filter fine-tuning based on the decision subset. At the same time, the targeted and accurate selection of the decision subset reduces the number of target vectors in the decision subset, thereby reducing the number of filter parameters that need to be adjusted specifically, and further improving the efficiency of subsequent model optimization.

[0064] In another exemplary embodiment provided by this application, based on the calculation of the covariance matrix of the filter fine-tuning layer and the training task set, the features input by the training task set are standardized. The covariance matrix can be decomposed into eigenvalues and eigenvectors through eigenvalue decomposition, thereby analyzing the variability of the filter fine-tuning layer in different directions. The eigenvalues obtained from the decomposition represent the variance of the data in the direction of the corresponding eigenvectors, and each eigenvector corresponds to an eigenvalue, indicating the projection of the data in a certain direction. The calculation formula of the covariance matrix of the filter fine-tuning layer and the training task set and the formula of eigenvalue decomposition are as follows:

[0065]

[0066] ∑ l =V l Λ l (V l ) T

[0067] Among them, ∑ l represents the covariance matrix of the filter fine-tuning layer and the training task set, f l represents the output of the layer filter of the k-th feature of the training task set in the l-th network layer, d k represents the task label of the k-th feature of the training task set, represents the mean of the layer filter of the l-th network layer, represents the mean of each task label of the training task set. V l represents the matrix containing the eigenvectors of the covariance matrix ∑ l , that is, the main change direction of the training task set. Λ l represents a diagonal matrix containing the eigenvalues corresponding to each eigenvector, that is, the variability of the layer filter of the l-th network layer in different directions.

[0068] The eigenvalue decomposition of the covariance matrix can help understand the variance distribution of the data and find the main change direction of the data (i.e., the direction pointed by the eigenvector). The larger the eigenvalue, the greater the variance of the data in this direction, the more significant the change, and the stronger the correlation with the training task set. Select the eigenvectors corresponding to the top K largest eigenvalues according to the eigenvalue size as the new principal components. The principal components selected in this application can already explain at least 90% of the total variance.

[0069] After eigenvalue decomposition, the most important directions in the covariance matrix are obtained. However, the outputs of the filter fine-tuning layer in the model may have different variabilities and task correlations in multiple different directions. The feature vectors are divided into multiple subspaces through K-means clustering, and the feature vectors with strong correlations are grouped together, making the feature vectors within the same subspace as similar as possible and the similarity between different subspaces relatively low. The formula for clustering is as follows:

[0070]

[0071] where J represents the subspaces obtained by clustering, v j represents the j-th target vector of the filter fine-tuning layer, μ i represents the center of the i-th subspace, and k represents the number of subspaces. K-means clustering continuously adjusts the subspace centers through an iterative method until convergence (i.e., the positions of the target vectors within the subspace no longer change). Finally, the average covariance between all target vectors within each subspace is calculated to obtain the vector correlation of each subspace, and the third threshold is set according to the vector correlation. The subspaces with vector correlations greater than the third threshold can be used as the decision subsets of the filter fine-tuning layer for subsequent adaptive micro-search optimization.

[0072] Optionally, population optimization is performed on the decision subsets to obtain the target decision subsets, including:

[0073] Calculating the first fitness of the decision subsets based on a preset fitness function;

[0074] Taking the decision subsets as the parent set and performing adaptive crossover and adaptive mutation on the parent set based on the first fitness to obtain the offspring set;

[0075] Calculating the second fitness of the offspring set based on the fitness function and taking the offspring set corresponding to the maximum second fitness as the target decision subset.

[0076] In this embodiment provided by the present application, adaptive crossover and adaptive mutation are performed on the decision subsets used as the parent set based on the calculated first fitness of the decision subsets to obtain the offspring set, and the second fitness of the offspring set is calculated. The offspring set corresponding to the maximum second fitness is taken as the target decision subset to realize the adaptive micro-search of the filter to be adjusted, which can effectively construct an effective decision subspace containing the optimal solution, that is, the target decision subset, so as to allocate more computing resources to the effective design search algorithm, improve the search efficiency of the target decision subset, and further improve the optimization efficiency in the model optimization process. Among them, the calculation process of the second fitness is the same as that of the first fitness, and there is an optimal target decision subset in the filter fine-tuning layer.

[0077] Adaptive Micro Search Strategy: According to the covariance matrix decomposition strategy, the decision set of the deep part (filter fine-tuning layer) is decomposed into multiple decision subsets. Search the subspaces of strongly correlated decision subsets to solve the filter to be adjusted for fine-tuning. Inspired by the idea of population evolution, based on this decomposition strategy, in this algorithm, the existing crossover and mutation are improved, and an adaptive operator is designed that can dynamically adjust the search process according to the fitness of each offspring, improving the search space and diversity of the population, thereby improving the accuracy of the target decision subset obtained by population optimization.

[0078] Please refer to Figure 5 , Figure 5 which is a flowchart for clustering eigenvalues and eigenvectors in an exemplary embodiment of this application. Figure 5 In it, the eigenvalues and eigenvectors of the filters in the filter fine-tuning layer are clustered to obtain subspaces (decision subspaces), the vector correlation between the target vectors in the subspaces is obtained, and the subspaces (decision subspaces) with vector correlation greater than the third threshold are used as the decision subsets (effective decision subspaces) of the filter fine-tuning layer, and the filters outside the decision subsets (effective decision subspaces) are frozen, and population optimization is performed on the decision subsets (effective decision subspaces) subsequently to adjust the filters in the decision subsets (effective decision subspaces) to determine the target decision subset containing the fine-tuning filter (filter to be adjusted). Based on the target decision subset, the model to be optimized is optimized and adjusted, and finally a target model that can accurately classify the target task (training task set) is obtained.

[0079] Optionally, calculating the first fitness of the decision subset based on a preset fitness function includes:

[0080] Adjust the model to be optimized based on the filter parameters corresponding to the decision subset to obtain the adjusted model to be optimized and the predicted classification corresponding to the training task set output by the adjusted model to be optimized;

[0081] Based on the predicted classification and the true classification corresponding to the preset training task set, obtain the classification accuracy;

[0082] The calculation formula of the classification accuracy is as follows:

[0083]

[0084] where f(χ i ) represents the classification accuracy of the i-th decision subset χ i of the filter fine-tuning layer, D t represents the training task set, represents the model obtained by fine-tuning the model to be optimized by selecting the decision subset χ i ;

[0085] Construct the optimized covariance matrix between each filter in the adjusted model to be optimized and the training task set, and calculate the second covariance value of the optimized covariance matrix;

[0086] Based on the classification accuracy and the second covariance value, use a preset fitness function to calculate the first fitness of the decision subset;

[0087] The calculation formula of the first fitness is as follows:

[0088]

[0089] where f fitness (χ i ) represents the i-th decision subset χ of the filter fine-tuning layer i , represents that after selecting the decision subset χ i to fine-tune the model to be optimized to obtain the model , the classification accuracy of the training task set D t on the model , and λ is used to balance the weight of the penalty term between the classification accuracy and the covariance matrix. Cov(F i , Target) represents the first covariance value between the filter fine-tuning layer F i of the model to be optimized and the training task set Target.

[0090] In this embodiment provided by the present application, first, based on the filter parameters corresponding to the decision subset, the model to be optimized is adjusted, the filter parameters are correspondingly replaced in the model to be optimized to obtain the adjusted model to be optimized, and the predicted classification corresponding to the training task set output by the adjusted model to be optimized, and based on the predicted classification and the true classification corresponding to the preset training task set, the classification accuracy is obtained. Then, construct the optimized covariance matrix between each filter in the adjusted model to be optimized and the training task set, and calculate the second covariance value of the optimized covariance matrix, and based on the classification accuracy and the second covariance value, use a preset fitness function to calculate the first fitness of the decision subset, which is convenient for subsequent adaptive crossover and adaptive mutation of the decision subset as the parent set to obtain the corresponding target decision subset, so as to improve the search efficiency of the target decision subset and the optimization efficiency in the model optimization process.

[0091] Optionally, performing adaptive crossover and adaptive mutation on the parent set based on the first fitness to obtain a child set includes:

[0092] Obtain the children of each parent element in the parent set, as well as the child crossover rate and child mutation rate of the children;

[0093] Update the offspring crossover rate and the offspring mutation rate based on the first fitness, a preset initial crossover rate, and an initial mutation rate to obtain an updated offspring crossover rate and an updated offspring mutation rate, and perform crossover and mutation on the parent set based on the updated offspring crossover rate and the updated offspring mutation rate to obtain a new parent set;

[0094] Calculate the new first fitness of the new parent set using the fitness function;

[0095] Iterate on the new parent set, the updated offspring crossover rate, and the updated offspring mutation rate based on the new first fitness until the offspring crossover rate and the offspring mutation rate converge to obtain the corresponding offspring set.

[0096] In the traditional population optimization GA (Genetic Algorithm) algorithm, the mutation rate and the crossover rate are usually statically set and basically remain unchanged throughout the evolution process, which may lead to the inability to adjust according to the current state of the search during the evolution process, possibly resulting in low search efficiency and getting trapped in local optimal solutions, and not being able to make full use of the current information for optimization. In the embodiment provided in this application, the operation strategy is intelligently adjusted according to the fitness information in the current search stage, that is, during the population evolution process, the offspring crossover rate and the offspring mutation rate are dynamically updated through the first fitness of the parent set, and crossover and mutation are performed on the parent set based on the updated offspring crossover rate and the updated offspring mutation rate to obtain a new parent set until the new first fitness corresponding to the new parent set (offspring set) reaches the maximum. At this time, the offspring crossover rate and the offspring mutation rate converge, and the population evolution ends. The new parent set obtained during the population evolution process is the offspring set. In this way, by dynamically updating the offspring crossover rate and the offspring mutation rate, the diversity of the obtained offspring set is increased and the fitness is improved, thereby being able to improve the search efficiency and search accuracy of the offspring set.

[0097] In another exemplary embodiment provided in this application, for the adaptive offspring crossover rate: The crossover rate determines the exchange frequency of parent elements in the parent set. The crossover operation generates offspring elements by exchanging the gene information of parent individuals. When the first fitness of the parent set is relatively high, the offspring crossover rate can be increased to combine more excellent genes. When the first fitness of the parent set is relatively low, the offspring crossover rate can be decreased to increase the mutation operation, realizing the adaptive dynamic adjustment of the offspring crossover rate.

[0098] The dynamic adjustment formula of the offspring crossover rate is as follows:

[0099]

[0100] Among them, crossover_rate(g) represents the offspring crossover rate of the current generation g, initial_crossover_rate represents the offspring crossover rate corresponding to the new parent set of the (g - 1)-th generation, fitness g represents the new first fitness value of the optimal offspring set in the current generation g. The optimal offspring set is the new parent set corresponding to the largest new first fitness value in the current generation g. avg_fitness represents the average fitness of all offspring sets in the current generation g. G represents the maximum number of generations for the preset population optimization, which can be set to any value between 20 and 100. When the new first fitness is relatively high, the offspring crossover rate increases, indicating frequent exchange of genetic information of the parents; while when the new first fitness is relatively low, the offspring crossover rate decreases, reducing information exchange and increasing the mutation operation.

[0101] Regarding the adaptive offspring mutation rate: The mutation rate determines the degree of random change that will occur in the new parent set (offspring set) in each generation. When the new first fitness is relatively low, the offspring mutation rate is increased to increase the breadth of the search space; while when the new first fitness is relatively high, the offspring mutation rate is decreased to converge near the current excellent solution (the largest first fitness).

[0102] The dynamic adjustment formula for the offspring mutation rate is as follows:

[0103]

[0104] Among them, mutation_rate(g) represents the offspring mutation rate of the current generation g, and initial_mutation_rate represents the offspring mutation rate corresponding to the new parent set of the (g - 1)-th generation. When the new first fitness is relatively large, the offspring mutation rate will gradually decrease; while when the new first fitness is relatively small, the offspring mutation rate will increase.

[0105] In another exemplary embodiment provided by the present application, it is assumed that X_k = (x_k,1, x_k,2, x_k,n^k) represents a decision subset of the k-th sub-problem (the optimization adjustment problem of the filter fine-tuning layer), and n^k represents the dimension of X_k. The candidate solutions (offspring sets) of the sub-problems can be defined as follows:

[0106]

[0107] Among them, χ i represents the i-th offspring set of the current filter fine-tuning layer, Np represents the size of the population (the current filter fine-tuning layer), x i,j represents the j-th dimension of χ i and n krepresents the number of filters in the k-th sub-problem (the optimization adjustment problem of the filter fine-tuning layer). In addition, x i,j ∈ {1, 0} indicates whether the j-th filter in the i-th sub-problem (the optimization adjustment problem of the filter fine-tuning layer) is fine-tuned or frozen.

[0108] For each offspring set (new parent set) in each generation, calculate its new first fitness. If the new first fitness of this offspring set is higher than the previous optimal solution, update the optimal solution. During the update process, ensure that the constraints of the covariance matrix are satisfied, that is, the selected filters should have a high correlation with the training task set. For the subset of filters selected for each offspring set (new parent set), perform fine-tuning for a small number of epochs to calculate the classification accuracy And calculate the updated new first fitness of each offspring set (new parent set). If a better solution is found, update the global optimum. Determine whether the termination condition is met. Finally, output the offspring set corresponding to the maximum value of the new first fitness value as the optimal offspring set, and the filter selection scheme of the optimal offspring set is the filter selection scheme of the target decision subset. Record the combination of the optimal classification accuracy and the filter selection scheme of the target decision subset to optimize and adjust the model to be optimized to obtain the target model.

[0109] In an exemplary embodiment provided by the present application, experiments are conducted based on a model transfer optimization method. The training task sets are selected as Stanford dog, MIT indoors, Caltech256-30, and Caltech256-60 shown in Table 1, and the model to be optimized is selected as ResNet50. Based on the training task sets Stanford dog, MIT indoors, Caltech256-30, and Caltech256-60 respectively, the model to be optimized ResNet50 is optimized and adjusted using the traditional fine-tuning method and the model optimization method of the present application respectively to obtain corresponding target models, and the classification accuracy of each target model is calculated for comparison. The comparison of the classification accuracy is shown in Table 2. In Table 2, the traditional fine-tuning methods include Train-From-Scrach (training from scratch), Standard Fine-Tuning (standard fine-tuning), L2-SP (L2 Spectral Penalty), AdaFilter (adaptive filter), ALS (Alternating Least Squares), Auto-RGN (Automatic Region Generation Network), and the model optimization method of the present application is OFT (Optimal Filter Tuning).

[0110] Table 1

[0111]

[0112] Table 2

[0113]

[0114] Please refer to Figure 6 , Figure 6 which shows a model transfer optimization system according to an exemplary embodiment of the present application. As Figure 6 shown, the present application provides a model transfer optimization system 600, comprising:

[0115] An acquisition module 601, configured to acquire a training task set;

[0116] A decomposition module 602, configured to perform covariance matrix decomposition on the model to be optimized using the training task set to obtain a decision subset; wherein, the model to be optimized is a model trained using a source task set, and the decision subset represents a set composed of the filters to be adjusted in the network layer of the model to be optimized;

[0117] The population optimization module 603 is used to optimize the population of the decision subsets to obtain the target decision subset;

[0118] The optimization and adjustment module 604 is used to optimize and adjust the model to be optimized based on the target decision subset to obtain the target model.

[0119] In this embodiment provided by the present application, for a model migration optimization system, first, the decomposition module 602 performs covariance matrix decomposition on the model to be optimized through the training task set obtained by the acquisition module 601, and obtains multiple decision subsets composed of filters that need to be adjusted. Among them, the model to be optimized is a model trained using the source task set, and the decision subset represents a set composed of the filters to be adjusted in the network layer of the model to be optimized. Then, the population optimization module 603 is used to optimize the population of the decision subsets so that the filters that need to be adjusted are finely tuned to obtain the optimal target decision subset. Then, the optimization and adjustment module 604 is used to perform targeted and precise optimization and adjustment on the model to be optimized based on the target decision subset to obtain the target model. In this way, it is not necessary to adjust all the filters in the model, and only need to perform precise fine-tuning on some filters in the model through transfer learning, which can not only reduce the optimization workload of the model, improve the optimization efficiency of the model, but also improve the classification accuracy of the optimized target model. At the same time, it can make the knowledge of the source task set corresponding to the model to be optimized be fully utilized in the model training of the training task set, thereby improving the classification performance of the obtained target model in the training task set and improving the classification accuracy.

[0120] Optionally, the decomposition module is specifically used for:

[0121] Construct the covariance matrix between the layer filters of each network layer in the model to be optimized and the training task set;

[0122] Determine the filter fine-tuning layer of the model to be optimized based on the covariance matrix;

[0123] Perform eigenvalue decomposition on the covariance matrix corresponding to the filter fine-tuning layer to obtain eigenvalues and eigenvectors;

[0124] Perform clustering processing on the eigenvalues and eigenvectors to obtain the decision subsets of the filter fine-tuning layer.

[0125] Optionally, the decomposition module is specifically used for:

[0126] Calculate the first covariance value of the covariance matrix, and determine the feature correlation between the corresponding layer filter and the training task set based on the first covariance value;

[0127] When the feature correlation is greater than the first threshold, use the corresponding layer filter as the filter fine-tuning layer.

[0128] Optionally, the decomposition module is specifically configured to:

[0129] Sort the eigenvalues of the filter fine-tuning layer in terms of magnitude to obtain a sorting result with the values arranged from largest to smallest;

[0130] Obtain a preset number of eigenvalues from the sorting result in the order from front to back as target eigenvalues, and determine the eigenvectors corresponding to the target eigenvalues as target vectors;

[0131] Obtain the task correlation between every two target vectors;

[0132] Take the task correlation greater than the second threshold as the target correlation;

[0133] Cluster the target vectors corresponding to the target correlations with the same direction to form a subspace;

[0134] Obtain the vector correlation between the target vectors in the subspace, and take the subspace with the vector correlation greater than the third threshold as the decision subset of the filter fine-tuning layer.

[0135] Optionally, the population optimization module is specifically configured to:

[0136] Calculate the first fitness of the decision subset based on a preset fitness function;

[0137] Take the decision subset as the parent set, and perform adaptive crossover and adaptive mutation on the parent set based on the first fitness to obtain an offspring set;

[0138] Calculate the second fitness of the offspring set based on the fitness function, and take the offspring set corresponding to the largest second fitness as the target decision subset.

[0139] Optionally, the population optimization module is specifically configured to:

[0140] Adjust the model to be optimized based on the filter parameters corresponding to the decision subset to obtain an adjusted model to be optimized and the predicted classification corresponding to the training task set output by the adjusted model to be optimized;

[0141] Obtain the classification accuracy based on the predicted classification and the true classification corresponding to the preset training task set;

[0142] Construct an optimization covariance matrix between each filter in the adjusted model to be optimized and the training task set, and calculate the second covariance value of the optimization covariance matrix;

[0143] Calculate the first fitness of the decision subset based on the classification accuracy and the second covariance value using a preset fitness function.

[0144] Optionally, the population optimization module is specifically configured to:

[0145] Obtain the offspring of each parental element in the parental set, as well as the offspring crossover rate and offspring mutation rate of the offspring;

[0146] Update the offspring crossover rate and offspring mutation rate based on the first fitness, a preset initial crossover rate, and an initial mutation rate to obtain an updated offspring crossover rate and an updated offspring mutation rate, and perform crossover and mutation on the parental set based on the updated offspring crossover rate and the updated offspring mutation rate to obtain a new parental set;

[0147] Calculate the new first fitness of the new parental set using a fitness function;

[0148] Iterate on the new parental set, the updated offspring crossover rate, and the updated offspring mutation rate based on the new first fitness until the offspring crossover rate and the offspring mutation rate converge to obtain the corresponding offspring set.

[0149] It should be noted that a model migration optimization system provided in the above embodiment and a model migration optimization method provided in the above embodiment belong to the same concept. The specific ways in which each module and unit perform operations have been described in detail in the method embodiment and will not be repeated here. In actual application, a model migration optimization system provided in the above embodiment can, as needed, allocate the above functions to different functional modules, that is, divide the internal structure of the system into different functional modules to complete all or part of the functions described above, and this will not be limited here either.

[0150] A computing device according to an embodiment of the present application includes a memory, a processor, and a program stored in the memory and running on the processor. When the processor executes the program, it implements some or all of the steps of the above model migration optimization method.

[0151] Among them, the computing device can select a computer. Correspondingly, its program is computer software, and the parameters and steps in the above computing device of the present application can refer to the parameters and steps in the embodiment of the above model migration optimization method and will not be elaborated here.

[0152] A computer-readable storage medium in an embodiment of the present application stores instructions that, when running, execute the steps of the above model migration optimization method.

[0153] Among them, the computer-readable storage medium can be a transient computer-readable storage medium or a non-transient computer-readable storage medium.

[0154] The technical solution of the embodiments of the present disclosure can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes one or more instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the method of the embodiments of the present disclosure. The aforementioned computer-readable storage medium may be a non-transitory computer-readable storage medium, including: various media such as USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs that can store program codes, or it may also be a transitory computer-readable storage medium.

[0155] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present application. Among them, each block in the flowchart or block diagram may represent a module, a program segment, or a part of the code, and the above-mentioned module, program segment, or part of the code includes one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than that marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram or flowchart, and the combination of blocks in the block diagram or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.

[0156] Those skilled in the art of the relevant technology know that the present application can be implemented as a system, a method, or a computer program product. Therefore, the present disclosure can be specifically implemented in the following forms: it can be completely hardware, can also be completely software (including firmware, resident software, microcode, etc.), or can also be in the form of a combination of hardware and software, which is generally referred to as a "module" or "system" in this article. In addition, in some embodiments, the present application can also be implemented in the form of a computer program product in one or more computer-readable media, and the computer-readable media contains computer-readable program codes. The computer-readable storage medium can be, for example, but not limited to - electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices, or components, or any combination of the above.

[0157] In the description of this specification, the descriptions with reference to the terms "one embodiment", "some embodiments", "examples", "specific examples", or "some examples", etc. mean that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of this application. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described may be combined in any one or more embodiments or examples in a suitable manner. In addition, without contradiction, those skilled in the art can combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples.

[0158] Although the embodiments of this application have been shown and described above, it can be understood that the above embodiments are exemplary and should not be construed as limiting this application. Those of ordinary skill in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of this application.

Claims

1. A model migration optimization method, characterized in that, Including: Obtain a training task set; Perform covariance matrix decomposition on the model to be optimized using the training task set to obtain a decision subset; wherein, the model to be optimized is a model trained using a source task set, and the decision subset represents a set composed of filters to be adjusted in the network layer of the model to be optimized; Perform population optimization on the decision subset to obtain a target decision subset; Based on the target decision subset, optimize and adjust the model to be optimized to obtain a target model.

2. The method according to claim 1, characterized in that, The performing covariance matrix decomposition on the model to be optimized using the training task set to obtain a decision subset includes: Construct a covariance matrix between the layer filters of each network layer in the model to be optimized and the training task set; Determine the filter fine-tuning layer of the model to be optimized based on the covariance matrix; Perform eigenvalue decomposition on the covariance matrix corresponding to the filter fine-tuning layer to obtain eigenvalues and eigenvectors; Perform clustering processing on the eigenvalues and the eigenvectors to obtain the decision subset of the filter fine-tuning layer.

3. The method according to claim 2, characterized in that The determining the filter fine-tuning layer of the model to be optimized based on the covariance matrix includes: Calculate a first covariance value of the covariance matrix, and determine the feature correlation between the corresponding layer filter and the training task set based on the first covariance value; When the feature correlation is greater than a first threshold, use the corresponding layer filter as the filter fine-tuning layer.

4. The method according to claim 2, wherein The performing clustering processing on the eigenvalues and the eigenvectors to obtain the decision subset of the filter fine-tuning layer includes: Sort the eigenvalues of the filter fine-tuning layer in descending order to obtain a sorting result with values from large to small; Obtain a preset number of eigenvalues from the sorting result in the order from front to back as target eigenvalues, and determine the eigenvectors corresponding to the target eigenvalues as target vectors; Obtain the task correlation between every two of the target vectors; Use the task correlation greater than a second threshold as the target correlation; Cluster the target vectors corresponding to the target correlations with the same direction to form a subspace; Obtain the vector correlation between the target vectors in each subspace, and use the subspace with the vector correlation greater than a third threshold as the decision subset of the filter fine-tuning layer.

5. The method according to claim 1, characterized in that The performing population optimization on the decision subset to obtain a target decision subset includes: Calculate a first fitness of the decision subset based on a preset fitness function; Use the decision subset as a parent set, and perform adaptive crossover and adaptive mutation on the parent set based on the first fitness to obtain a child set; Calculate a second fitness of the child set based on the fitness function, and use the child set corresponding to the maximum second fitness as the target decision subset.

6. The method according to claim 5, characterized in that, The calculating a first fitness of the decision subset based on a preset fitness function includes: Adjust the model to be optimized based on the filter parameters corresponding to the decision subset to obtain the adjusted model to be optimized and the predicted classification corresponding to the training task set output by the adjusted model to be optimized; Based on the predicted classification and the true classification corresponding to the preset training task set, obtain the classification accuracy; Construct an optimized covariance matrix between each filter in the adjusted model to be optimized and the training task set, and calculate the second covariance value of the optimized covariance matrix; Based on the classification accuracy and the second covariance value, use a preset fitness function to calculate the first fitness of the decision subset.

7. The method according to claim 5, wherein The adaptive crossover and adaptive mutation of the parent set based on the first fitness to obtain a child set includes: Obtain the children of each parent element in the parent set, as well as the child crossover rate and child mutation rate of the children; Update the child crossover rate and the child mutation rate based on the first fitness, a preset initial crossover rate, and an initial mutation rate to obtain an updated child crossover rate and an updated child mutation rate, and perform crossover and mutation on the parent set based on the updated child crossover rate and the updated child mutation rate to obtain a new parent set; Use the fitness function to calculate the new first fitness of the new parent set; Iterate on the new parent set, the updated child crossover rate, and the updated child mutation rate based on the new first fitness until the child crossover rate and the child mutation rate converge, and obtain the corresponding child set.

8. A model migration optimization system, characterized in that, Including: An acquisition module, configured to acquire a training task set; A decomposition module, configured to perform covariance matrix decomposition on the model to be optimized using the training task set to obtain a decision subset; wherein, the model to be optimized is a model trained using a source task set, and the decision subset represents a set composed of filters to be adjusted in the network layer of the model to be optimized; A population optimization module, configured to perform population optimization on the decision subset to obtain a target decision subset; An optimization and adjustment module, configured to optimize and adjust the model to be optimized based on the target decision subset to obtain a target model.

9. A computing device, comprising a memory, a processor, and a program stored on the memory and running on the processor, characterized in that, When the processor executes the program, it implements the steps of a model migration optimization method according to any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, Instructions are stored in the computer-readable storage medium, and when the instructions are run on a terminal device, the terminal device is caused to execute the steps of a model migration optimization method according to any one of claims 1 to 7.