A deep memory transfer learning method based on topological entropy decomposition
Through the deep memory transfer learning method based on topological entropy decomposition, low topological entropy nodes are identified and pruned, model structure and weights are dynamically adjusted, and deep memory modules are embedded, the problem of negative transfer and insufficient robustness of the model in traditional transfer learning methods is solved, and a more efficient, flexible and adaptive transfer learning effect is achieved.
Patent Information
- Application Number
- CN202411751256.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-02
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2044-12-02
AI Technical Summary
Traditional transfer learning methods face negative transfer problems in practical applications, insufficient model robustness and generalization capabilities, and difficulty in dynamically adjusting the transfer ratio of source task knowledge, resulting in the model lacking flexibility and adaptability in the face of dynamically changing data or tasks.
A deep memory transfer learning method based on topological entropy decomposition is proposed. Low topological entropy nodes are identified and pruned through topological entropy analysis, model structure and weight initialization method are dynamically adjusted, and deep memory modules are embedded to store and utilize historical experience.
It significantly improves the robustness, efficiency and adaptability of transfer learning, reduces the number of parameters and calculation overhead of the model, and improves the learning effect of the target task and the flexibility of the model.
Smart Images

Figure CN119514647B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of deep memory transfer learning, and in particular to a deep memory transfer learning method based on topological entropy decomposition. Background Art
[0002] With the rapid development of artificial intelligence and deep learning technology, transfer learning, as an important research direction in the field of artificial intelligence, has gradually been widely used in image recognition, natural language processing, and medical diagnosis. Transfer learning can significantly improve the learning efficiency and effect of the target task by transferring the knowledge of the source task to the target task. However, traditional transfer learning methods still face many challenges in practical applications.
[0003] At present, most transfer learning methods rely on deep neural networks for feature extraction and transfer, usually by pre-training the model on a large-scale source task dataset and then directly applying the model's weights and structure to the target task. Although this method has achieved good results in some scenarios, due to the differences in data distribution and feature space between the source task and the target task, the negative transfer problem of transfer learning often occurs, that is, the knowledge of the source task cannot be effectively adapted to the target task, and even has a negative impact on the learning effect of the target task. In addition, due to the complexity and black-box characteristics of the deep learning model, traditional methods find it difficult to understand and optimize the learning behavior of the model during the transfer process, resulting in insufficient robustness and generalization ability of the model.
[0004] In response to the above-mentioned existing problems, some studies in recent years have attempted to introduce network pruning technology and feature selection methods to optimize the transfer learning process and reduce the interference of redundant features on the model. However, the methods in recent years lack systematic analysis of the complex internal behaviors of the model and the correlation between nodes, and are unable to fully identify and retain useful features. At the same time, the existing technology is difficult to dynamically adjust the migration ratio of the source task knowledge when processing the target task, resulting in the model's lack of flexibility and adaptability when facing dynamically changing data or tasks.
[0005] In addition, there is little research on the use of historical experience and features of source tasks in the existing technology, and the advantages of memory modules in the transfer learning process have not been fully utilized. Most methods only focus on the immediate learning process of the target task, ignoring the long-term accumulation and effective use of source and historical task features, which further limits the efficiency and effectiveness of the transfer learning model. Summary of the invention
[0006] One object of the present invention is to propose a deep memory transfer learning method based on topological entropy decomposition, which improves the robustness, efficiency and adaptability of transfer learning.
[0007] A deep memory transfer learning method based on topological entropy decomposition according to an embodiment of the present invention includes the following steps:
[0008] S1. Obtain the source task dataset and the target task dataset, and perform data cleaning, feature extraction, and standardization on the source task dataset and the target task dataset respectively;
[0009] S2. Use a deep neural network to train the preprocessed source task dataset to generate a source task training model;
[0010] S3. Perform topological entropy analysis on the source task training model, calculate the topological entropy value of each network node in the source task training model, and quantify the contribution of each node to the overall performance of the source task training model;
[0011] S4. Based on the topological entropy value, low topological entropy nodes and high topological entropy nodes are screened out according to a preset pruning threshold, and the low topological entropy nodes are pruned from the source task training model to obtain the pruned optimized source task model;
[0012] S5. Perform migration adaptation on the optimized source task model, calculate the similarity index between each layer structure of the optimized source task model and the target task in combination with the feature distribution of the target task data set, and adjust the weight initialization method and structural parameters of the optimized source task model;
[0013] S6. Use the target task dataset to perform preliminary training on the optimized source task model, generate a preliminary migration model, and record the performance indicators of the preliminary migration model on the target task;
[0014] S7. Perform topological entropy analysis on the preliminary migration model again, calculate the topological entropy values of the newly added network nodes during the migration process, optimize the structure of the preliminary migration model in combination with the performance indicators, further remove redundant low topological entropy nodes during the migration process, and generate the target task network model;
[0015] S8. embed a deep memory module in the target task network model, and use the deep memory module to store the high topological entropy node features in the source task model and the historical experience in the migration process;
[0016] S9. Perform final training on the target task network model including the deep memory module, dynamically monitor the changes in the model's topological entropy during the training process, and adjust the usage strategy and pruning rules of the deep memory module;
[0017] S10. Output the trained target task network model and record the pruning optimization process to form a reusable model optimization template.
[0018] Optionally, the S1 specifically includes:
[0019] S11. Obtain source task dataset And the target task dataset in, and Represent the feature vectors of the source task dataset and the target task dataset respectively, and Represents the label values of the source task dataset and the target task dataset, respectively, s and N t Represents the number of data samples for the source task and the target task respectively;
[0020] S12. Source task dataset D s And the target task dataset D t Perform data cleaning to remove abnormal samples in the data set;
[0021] S13. Perform feature extraction on the cleaned source task dataset and target task dataset, extract effective features through the predefined feature mapping function φ(x), and extract the source task feature set and the target task feature set
[0022] S14. Extract the source task feature set F s and the target task feature set F t Perform standardization to obtain the standardized source task feature set F′ s and the target task feature set F′ t .
[0023] Optionally, the S2 specifically includes:
[0024] S21. Use the standardized source task feature set and label set as input data to build a deep neural network model;
[0025] S22. Initialize the weight parameters and bias items of the deep neural network model, wherein the deep neural network includes an input layer, several hidden layers and an output layer, and the weight matrix and bias items of each layer are respectively Among them, L represents the number of layers of the neural network, W k is the weight matrix of the kth layer, b k is the bias term of the kth layer;
[0026] S23. Use the forward propagation algorithm to train the deep neural network and transform the standardized source task features As input, the output is calculated through linear transformation and activation function processing of each layer.
[0027] a k =σ k (W k ·ak-1 +b k +∈ k ·Δ k +λ·Ω k ),k=1,2,…,L;
[0028] in, σ k (·) is the activation function of the kth layer, ∈ k is the topological perturbation factor related to the kth layer, which is used to introduce topological entropy perturbation in training, Δ k represents the interaction matrix between nodes, λ is the regularization coefficient, Ω k is the topological constraint term, a k is the activation output of the kth layer, represents the predicted output of the model;
[0029] S24. Use topological weighted mean square error loss function Calculate the loss value of the source task training model:
[0030]
[0031] in, is the true label, is the predicted output, H k is the topological entropy value of the kth layer, β k is the topological weighting coefficient, which is used to adjust the contribution of each layer of topological entropy to the overall loss in the loss function;
[0032] S25. Use the back-propagation algorithm to calculate the gradient of the loss value relative to the weights and bias terms of each layer and
[0033]
[0034] Among them, γ is the topological regularization factor, which is used to balance the influence of topological entropy on the gradient, η k is the dynamic adjustment factor, Δ k represents the contribution of the influence of the interaction matrix between nodes to the gradient;
[0035] S26. Update weights and bias terms using gradient descent:
[0036]
[0037] Among them, η is the learning rate, ν k is the topological constraint influence factor of the kth layer, Ω k It is a topology constraint item, which optimizes the stability and balance of the network structure by constraining the topology structure of each layer of the network;
[0038] S27. Repeat steps S23 to S26 until the preset convergence conditions are met, and output the source task training model.
[0039] Optionally, the S3 specifically includes:
[0040] S31. Construct the topological structure diagram of the source task training model:
[0041] G = (V, E);
[0042] Among them, V represents the node set in the source task training model, E represents the connection relationship between nodes, and for each node, its connection weight is defined as w ij , represents the node v i and v j The weight strength between them;
[0043] S32. Calculate the local information entropy of each node of the source task training model Quantify the complexity of a node in the local topology:
[0044]
[0045] in, Represents node v i The neighborhood set of ;
[0046] S33. Calculate the global topological entropy of the source task training model as a whole Describe the overall complexity of the model:
[0047]
[0048] in, Represents node v i degree;
[0049] S34. Calculate the topological contribution C of each node by combining local information entropy and global topological entropy i :
[0050]
[0051] Among them, α∈[0,1] is the adjustment coefficient, which is used to balance the weights of local and global topological characteristics;
[0052] S35. Normalize the topological contributions of all nodes to obtain the normalized node topological entropy value
[0053]
[0054] in, Represents node v iThe relative contribution of the model to the whole model.
[0055] Optionally, the S4 specifically includes:
[0056] S41. Set the lower limit τ for pruning nodes with low topological entropy low and the upper limit τ of retaining nodes with high topological entropy high , τ low <τ high , according to the topological entropy Divide the model nodes into low topological entropy node sets V low and high topological entropy node set V high ;
[0057] S42. For low topological entropy node set V low The nodes and their connections in are pruned, and the topological structure diagram of the source task training model is updated to G ′ =(V ′ ,E ′ ):
[0058] V ′ =V\V low ;
[0059] E ′ =E\{e ij ∣v i ∈V low or v j ∈V low};
[0060] S43. Update the pruned node weight matrix W′ k and the bias term b′ k :
[0061]
[0062] Among them, w ij and b ij Respectively represent the contribution of the weight value and bias term associated with the pruned node;
[0063] S44. Recalculate the node connection strength for the updated topological structure graph G' and adjust the high topological entropy node set V high The network importance of redefines the connection weight w′ of high topological entropy nodes ij ;
[0064] S45. Recalculate the loss function of the pruned model and adjust the network weights and bias terms through optimization. The loss function is defined as:
[0065]
[0066] Among them, H′k is the global topological entropy value of the pruned model, β k and γ 1 is the weight coefficient;
[0067] S46. Output the pruned optimized source task model, including the updated network weight matrix, bias term, and optimized topology structure diagram.
[0068] Optionally, the S5 specifically includes:
[0069] S51. Obtain the feature distribution information of the target task data set and calculate the feature set of the target task data set The mean vector μ t and the covariance matrix Σ t ;
[0070] S52. The optimized source task model (W′ k ,b′ k ) to perform topological hierarchical analysis, extract the feature distribution of each layer of nodes in the optimization source task model, and calculate the feature mean vector μ′ of each layer k and the covariance matrix Σ′ k :
[0071]
[0072] in, To optimize the node activation value of the kth layer of the source task model, N s is the number of source task samples;
[0073] S53. Calculate the target task feature distribution (μ t ,Σ t ) and optimize the feature distribution of each layer of the source task model (μ′ k ,Σ′ k ) k , defining the similarity index as the weighted sum of the maximum mean difference and covariance deviation between feature distributions:
[0074]
[0075] Among them, α 1 and β 1 is the weighting coefficient, ∥·∥ represents the Frobenius norm of the matrix;
[0076] S54. Based on the calculated similarity index S k Adjust the weight initialization method of the optimized source task model and define the adjusted weight initialization matrix
[0077] S55. Initialize the matrix based on the adjusted weights and the bias term Redefine the structural parameters of the optimized source task model:
[0078]
[0079] in, η k1 is the dynamic adjustment factor of the bias term.
[0080] Optionally, the S6 specifically includes:
[0081] S61. Using the standardized feature set and corresponding label set of the target task dataset as input data, the target task features are input into the optimized source task model after transfer adaptation, and processed by linear transformation and activation function of each layer in turn to finally obtain the prediction output;
[0082] S62. Calculate the initial loss function of the target task training phase;
[0083] S63. Calculate the gradient of the loss function with respect to the model weights and bias terms using a back-propagation algorithm;
[0084] S64. Update the model weights and bias terms using the gradient descent method;
[0085] S65. Repeat steps S61 to S64, perform multiple rounds of iterative training on the target task data set until the preset training convergence conditions are met, generate a preliminary migration model, and record the performance indicators of the preliminary migration model on the target task, including the prediction error of the target task and the accuracy of the model on the target task.
[0086] Optionally, the S7 specifically includes:
[0087] S71. Perform topological entropy analysis on the network topology structure of the preliminary migration model, and calculate the local topological entropy and global topological entropy of the newly added network nodes during the migration process. The local topological entropy is used to describe the connection complexity between the newly added nodes and their neighboring nodes, and the global topological entropy is used to measure the contribution of the newly added nodes to the overall complexity of the entire network.
[0088] S72. The comprehensive importance weight of the newly added node is defined in combination with the performance index of the target task. The local topological entropy, global topological entropy, prediction error of the target task and the influence of the model accuracy on the importance of the node are comprehensively considered. The comprehensive importance weight is calculated as:
[0089]
[0090] Among them, α 2 , β 2 , γ 2 and δ 1is the importance weight adjustment coefficient, is the local topological entropy, is the global topological entropy;
[0091] S73. Set a pruning threshold of the comprehensive importance weight, filter out a low-weight node set, which is a newly added node with a weight value lower than the pruning threshold, and mark the low-weight node as a redundant node;
[0092] S74. Prune the low-weight nodes and their connections marked as redundant, update the topological structure of the preliminary migration model, and generate the target task network model.
[0093] Optionally, the S8 specifically includes:
[0094] S81. Embed a deep memory module in the target task network model, the deep memory module includes a storage unit, an update unit and a read unit, which are used to store high topological entropy node features in the source task model, dynamically update historical migration experience, and provide historical information required in the target task training process;
[0095] S82. Extract the features of high topological entropy nodes from the source task model and store them in the storage unit of the deep memory module, and record the historical high-weight node features in the migration process in time series to form a continuous feature history track;
[0096] S83. Define the update mechanism of the deep memory module, dynamically update the storage unit content according to the real-time training status and performance of the target task network model, remove redundant features and add new high-importance node features;
[0097] S84. Define the reading mechanism of the deep memory module, select historical features with high similarity to the target task features from the storage unit according to the requirements of the target task training, dynamically load relevant experience through the reading unit, and optimize the learning efficiency of the task network model.
[0098] The beneficial effects of the present invention are:
[0099] (1) The present invention introduces topological entropy analysis into the source task training model and the target task network model. By calculating the local and global topological entropies of the nodes in the network, the high topological entropy nodes that contribute more to the network performance are identified and retained, while the low topological entropy nodes and their redundant connections are pruned. The adaptive pruning method based on topological entropy is more accurate than the traditional node importance evaluation, can capture the nonlinear relationship and complex characteristics between network nodes, significantly reduce the redundant calculation of the network, and improve the efficiency and stability of the model.
[0100] (2) The similarity index calculation method proposed in the present invention dynamically adjusts and optimizes the weight initialization method and structural parameters of the source task model by analyzing the mean difference and covariance deviation between the feature distribution of the target task and the feature distribution of each layer of the source task. Compared with the traditional method of directly migrating weights, the method of the present invention can adjust the migration strategy in real time according to the feature similarity between the source task and the target task, thereby reducing the occurrence of negative migration. The dynamic adjustment mechanism makes the model more flexible when facing the dynamic changes in the target task data distribution.
[0101] (3) The present invention embeds a deep memory module in the target task network model to store the high topological entropy node features in the source task and the historical high-weight node features generated during the migration process. Through a dynamic update and reading mechanism, the historical knowledge related to the current task can be flexibly called during the target task training process. This can not only help the model make full use of historical experience when facing complex tasks, but also dynamically adapt to the needs of the target task by eliminating redundant memory, thereby improving the learning efficiency and effect of the model, compared with the model that does not contain a memory module. BRIEF DESCRIPTION OF THE DRAWINGS
[0102] The accompanying drawings are used to provide a further understanding of the present invention and constitute a part of the specification. Together with the embodiments of the present invention, they are used to explain the present invention and do not constitute a limitation of the present invention. In the accompanying drawings:
[0103] Figure 1 This is a flow chart of a deep memory transfer learning method based on topological entropy decomposition proposed by the present invention. DETAILED DESCRIPTION
[0104] The present invention will now be described in further detail with reference to the accompanying drawings. These drawings are simplified schematic diagrams, which only illustrate the basic structure of the present invention in a schematic manner, and therefore only show the components related to the present invention.
[0105] refer to Figure 1 , a deep memory transfer learning method based on topological entropy decomposition, comprising the following steps:
[0106] S1. Obtain the source task dataset and the target task dataset, and perform data cleaning, feature extraction, and standardization on the source task dataset and the target task dataset respectively;
[0107] S2. Use a deep neural network to train the preprocessed source task dataset to generate a source task training model;
[0108] S3. Perform topological entropy analysis on the source task training model, calculate the topological entropy value of each network node in the source task training model, and quantify the contribution of each node to the overall performance of the source task training model;
[0109] S4. Based on the topological entropy value, low topological entropy nodes and high topological entropy nodes are screened out according to a preset pruning threshold, and the low topological entropy nodes are pruned from the source task training model to obtain the pruned optimized source task model;
[0110] S5. Perform migration adaptation on the optimized source task model, calculate the similarity index between each layer structure of the optimized source task model and the target task in combination with the feature distribution of the target task data set, and adjust the weight initialization method and structural parameters of the optimized source task model;
[0111] S6. Use the target task dataset to perform preliminary training on the optimized source task model, generate a preliminary migration model, and record the performance indicators of the preliminary migration model on the target task;
[0112] S7. Perform topological entropy analysis on the preliminary migration model again, calculate the topological entropy values of the newly added network nodes during the migration process, optimize the structure of the preliminary migration model in combination with the performance indicators, further remove redundant low topological entropy nodes during the migration process, and generate the target task network model;
[0113] S8. embed a deep memory module in the target task network model, and use the deep memory module to store the high topological entropy node features in the source task model and the historical experience in the migration process;
[0114] S9. Perform final training on the target task network model including the deep memory module, dynamically monitor the changes in the model's topological entropy during the training process, and adjust the usage strategy and pruning rules of the deep memory module;
[0115] S10. Output the trained target task network model and record the pruning optimization process to form a reusable model optimization template.
[0116] In this implementation, S1 specifically includes:
[0117] S11. Obtain source task dataset And the target task dataset in, and Represent the feature vectors of the source task dataset and the target task dataset respectively, and Represents the label values of the source task dataset and the target task dataset, respectively, s and N t Represents the number of data samples for the source task and the target task respectively;
[0118] S12. Source task dataset D s And the target task dataset D t Perform data cleaning to remove abnormal samples in the data set;
[0119] S13. Perform feature extraction on the cleaned source task dataset and target task dataset, extract effective features through the predefined feature mapping function φ(x), and extract the source task feature set and the target task feature set
[0120] S14. Extract the source task feature set F s and the target task feature set F t Perform standardization to obtain the standardized source task feature set F′ s and the target task feature set F′ t .
[0121] In this implementation, S2 specifically includes:
[0122] S21. Use the standardized source task feature set and label set as input data to build a deep neural network model;
[0123] S22. Initialize the weight parameters and bias items of the deep neural network model. The deep neural network includes an input layer, several hidden layers and an output layer. The weight matrix and bias items of each layer are Among them, L represents the number of layers of the neural network, W k is the weight matrix of the kth layer, b k is the bias term of the kth layer;
[0124] S23. Use the forward propagation algorithm to train the deep neural network and transform the standardized source task features As input, the output is calculated through linear transformation and activation function processing of each layer.
[0125] a k =σ k (W k ·a k-1 +b k +∈ k Δ k +λ·Ω k ),k=1,2,…,L;
[0126] in, σ k (·) is the activation function of the kth layer, ∈ k is the topological perturbation factor related to the kth layer, which is used to introduce topological entropy perturbation in training, Δ k represents the interaction matrix between nodes, λ is the regularization coefficient, Ω k is the topological constraint, a k is the activation output of the kth layer, represents the predicted output of the model;
[0127] S24. Use topological weighted mean square error loss function Calculate the loss value of the source task training model:
[0128]
[0129] in, is the true label, is the predicted output, H k is the topological entropy value of the kth layer, β k is the topological weighting coefficient, which is used to adjust the contribution of each layer of topological entropy to the overall loss in the loss function;
[0130] S25. Use the back-propagation algorithm to calculate the gradient of the loss value relative to the weights and bias terms of each layer and
[0131]
[0132] Among them, γ is the topological regularization factor, which is used to balance the influence of topological entropy on the gradient, η k is the dynamic adjustment factor, Δ k represents the contribution of the influence of the interaction matrix between nodes to the gradient;
[0133] S26. Update weights and bias terms using gradient descent:
[0134]
[0135] Among them, η is the learning rate, ν k is the topological constraint influence factor of the kth layer, Ω k It is a topology constraint item, which optimizes the stability and balance of the network structure by constraining the topology structure of each layer of the network;
[0136] S27. Repeat steps S23 to S26 until the preset convergence conditions are met, and output the source task training model.
[0137] In this implementation, S3 specifically includes:
[0138] S31. Construct the topological structure diagram of the source task training model:
[0139] G = (V, E);
[0140] Among them, V represents the node set in the source task training model, E represents the connection relationship between nodes, and for each node, its connection weight is defined as w ij , represents the node v i and v j The weight strength between them;
[0141] S32. Calculate the local information entropy of each node of the source task training model Quantify the complexity of a node in the local topology:
[0142]
[0143] in, Represents node v i The neighborhood set of ;
[0144] S33. Calculate the global topological entropy of the source task training model as a whole Describe the overall complexity of the model:
[0145]
[0146] in, Represents node v i degree;
[0147] S34. Calculate the topological contribution C of each node by combining local information entropy and global topological entropy i :
[0148]
[0149] Among them, α∈[0,1] is the adjustment coefficient, which is used to balance the weights of local and global topological characteristics;
[0150] S35. Normalize the topological contributions of all nodes to obtain the normalized node topological entropy value
[0151]
[0152] in, Represents node v i The relative contribution of the model to the whole model.
[0153] In this implementation, S4 specifically includes:
[0154] S41. Set the lower limit τ for pruning nodes with low topological entropy low and the upper limit τ of retaining nodes with high topological entropy high , τ low <τ high , according to the topological entropy value T i Divide the model nodes into low topological entropy node sets V low and high topological entropy node set V high ;
[0155] S42. For low topological entropy node set V lowThe nodes and their connections in are pruned, and the topological structure diagram of the source task training model is updated to G ′ =(V ′ ,E ′ ):
[0156] V ′ =V\V low ;
[0157] E′=E\{e ij ∣v i ∈V low or v j ∈V low};
[0158] S43. Update the pruned node weight matrix W′ k and the bias term b′ k :
[0159]
[0160] Among them, w ij and b ij Respectively represent the contribution of the weight value and bias term associated with the pruned node;
[0161] S44. Recalculate the node connection strength for the updated topological structure graph G' and adjust the high topological entropy node set V high The network importance of redefines the connection weight w′ of high topological entropy nodes ij ;
[0162] S45. Recalculate the loss function of the pruned model and adjust the network weights and bias terms through optimization. The loss function is defined as:
[0163]
[0164] Among them, H′ k is the global topological entropy of the pruned model, β k and γ 1 is the weight coefficient;
[0165] S46. Output the pruned optimized source task model, including the updated network weight matrix, bias term, and optimized topology structure diagram.
[0166] In this implementation, S5 specifically includes:
[0167] S51. Obtain the feature distribution information of the target task data set and calculate the feature set of the target task data set The mean vector μ t and the covariance matrix Σ t ;
[0168] S52. The optimized source task model (W′ k ,b′ k ) to perform topological hierarchical analysis, extract the feature distribution of each layer of nodes in the optimization source task model, and calculate the feature mean vector μ′ of each layer k and the covariance matrix Σ′ k :
[0169]
[0170] in, To optimize the node activation value of the kth layer of the source task model, N s is the number of source task samples;
[0171] S53. Calculate the target task feature distribution (μ t ,Σ t ) and optimize the feature distribution of each layer of the source task model (μ′ k ,Σ′ k ) k , defining the similarity index as the weighted sum of the maximum mean difference and covariance deviation between feature distributions:
[0172]
[0173] Among them, α 1 and β 1 is the weighting coefficient, ∥·∥ represents the Frobenius norm of the matrix;
[0174] S54. Based on the calculated similarity index S k Adjust the weight initialization method of the optimized source task model and define the adjusted weight initialization matrix
[0175] S55. Initialize the matrix based on the adjusted weights and the bias term Redefine the structural parameters of the optimized source task model:
[0176]
[0177] in, η k1 is the dynamic adjustment factor of the bias term.
[0178] In this implementation, S6 specifically includes:
[0179] S61. Using the standardized feature set and corresponding label set of the target task dataset as input data, the target task features are input into the optimized source task model after transfer adaptation, and processed by linear transformation and activation function of each layer in turn to finally obtain the prediction output;
[0180] S62. Calculate the initial loss function of the target task training phase;
[0181] S63. Calculate the gradient of the loss function with respect to the model weights and bias terms using a back-propagation algorithm;
[0182] S64. Update the model weights and bias terms using the gradient descent method;
[0183] S65. Repeat steps S61 to S64, perform multiple rounds of iterative training on the target task data set until the preset training convergence conditions are met, generate a preliminary migration model, and record the performance indicators of the preliminary migration model on the target task, including the prediction error of the target task and the accuracy of the model on the target task.
[0184] In this implementation, S7 specifically includes:
[0185] S71. Perform topological entropy analysis on the network topology structure of the preliminary migration model, and calculate the local topological entropy and global topological entropy of the newly added network nodes during the migration process. The local topological entropy is used to describe the connection complexity between the newly added nodes and their neighboring nodes, and the global topological entropy is used to measure the contribution of the newly added nodes to the overall complexity of the entire network.
[0186] S72. The comprehensive importance weight of the newly added node is defined in combination with the performance index of the target task. The local topological entropy, global topological entropy, prediction error of the target task and the influence of the model accuracy on the importance of the node are comprehensively considered. The comprehensive importance weight is calculated as:
[0187]
[0188] Among them, α 2 , β 2 , γ 2 and δ 1 is the importance weight adjustment coefficient, is the local topological entropy, is the global topological entropy;
[0189] S73. Set a pruning threshold of the comprehensive importance weight, filter out a low-weight node set, which is a newly added node with a weight value lower than the pruning threshold, and mark the low-weight node as a redundant node;
[0190] S74. Prune the low-weight nodes and their connections marked as redundant, update the topological structure of the preliminary migration model, and generate the target task network model.
[0191] In this implementation, S8 specifically includes:
[0192] S81. Embed a deep memory module in the target task network model, the deep memory module includes a storage unit, an update unit and a read unit, which are used to store high topological entropy node features in the source task model, dynamically update historical migration experience, and provide historical information required in the target task training process;
[0193] S82. Extract the features of high topological entropy nodes from the source task model and store them in the storage unit of the deep memory module, and record the historical high-weight node features in the migration process in time series to form a continuous feature history track;
[0194] S83. Define the update mechanism of the deep memory module, dynamically update the storage unit content according to the real-time training status and performance of the target task network model, remove redundant features and add new high-importance node features;
[0195] S84. Define the reading mechanism of the deep memory module, select historical features with high similarity to the target task features from the storage unit according to the requirements of the target task training, dynamically load relevant experience through the reading unit, and optimize the learning efficiency of the task network model.
[0196] Embodiment 1:
[0197] In order to verify the practical feasibility and performance improvement of the present invention in the field of transfer learning, two self-constructed datasets, ChestX-ray14 (source task dataset) and VinDr-CXR (target task dataset), were selected for experimental verification. The two datasets respectively contain chest X-ray images, which are used to train and test the performance of different models in the detection and classification tasks of lung diseases.
[0198] ChestX-ray14 is a large-scale dataset containing 112,120 chest X-ray images, covering 14 different chest diseases. VinDr-CXR is a small-scale dataset containing 15,000 high-quality chest X-ray images, mainly used to detect lung nodules and abnormal areas. In the experiment, the two datasets are divided into training and test sets in a ratio of 8:2, respectively, so as to compare the performance in the source task and the target task.
[0199] Feature extraction and model setting: In the experiment, ResNet-50 was used as the backbone network to extract features from the ChestX-ray14 dataset and pre-trained to obtain a preliminary source task model. The VinDr-CXR dataset uses the same feature extraction network as the input of the target task dataset, and the source task model is migrated to the target task through transfer learning. The method based on topological entropy decomposition of the present invention is introduced to perform topological entropy analysis on the source task model, prune inefficient nodes, and dynamically adjust the migration weights and structural parameters of the target task.
[0200] Training configuration: Optimizer: Adaptive moment estimation (Adam) optimizer, initial learning rate 0.001. Batch size: 64. Number of training rounds: 150. Loss function: weighted cross entropy loss function, combined with the topological entropy regularization term in the present invention.
[0201] First, ResNet-50 was pre-trained on the ChestX-ray14 dataset. The node contribution of the network was analyzed by topological entropy decomposition to identify low topological entropy nodes and the model was pruned and optimized. In the pruned model, the number of parameters was reduced by 28%, the computational efficiency was improved by 35%, and the disease classification accuracy on the ChestX-ray14 dataset was maintained at 91.3%.
[0202] Next, the optimized source task model was migrated to the VinDr-CXR dataset. By analyzing the similarity of the feature distribution of the target task dataset, the weight initialization method of the source task model was adjusted. A deep memory module was added to the training process of VinDr-CXR to store the high topological entropy node features of the source task and dynamically call the historical experience in the migration process. The initial training model achieved an accuracy of 89.2% on the target task.
[0203] The preliminary migration model was analyzed again by topological entropy, and the final target task network model was generated by pruning the inefficient nodes newly added during the migration process. In the final model, the number of parameters was further reduced by 10%, and the accuracy was improved to 94.5%.
[0204] In the experiment, the method of the present invention was compared with the traditional transfer learning method (direct transfer of pre-trained model), and the results are shown in Table 1 below:
[0205] Table 1 Comparison between the method of the present invention and the traditional transfer learning method
[0206]
[0207] The average precision (mAP), recall rate (Recall) and F1 score were also used in the experiment to further evaluate the model performance. The results are shown in Table 2:
[0208] Table 2 Model performance evaluation indicators
[0209]
[0210] Experimental results show that the method of the present invention can effectively solve the negative transfer problem in the process of transfer learning, improve the transfer efficiency and the learning effect of the target task. Compared with the traditional transfer learning method, the present invention not only achieves a higher accuracy rate (94.5%) on the target task, but also significantly reduces the number of model parameters (35%) and training time (more than 2 hours). In addition, by introducing the deep memory module, the model can make full use of historical experience and show higher robustness and adaptability in complex task scenarios. The results of Example 1 fully verify the feasibility and effectiveness of the present invention in practical applications.
[0211] The present invention introduces topological entropy analysis into the source task training model and the target task network model. By calculating the local and global topological entropies of the nodes in the network, the high topological entropy nodes that contribute more to the network performance are identified and retained, while the low topological entropy nodes and their redundant connections are pruned. The adaptive pruning method based on topological entropy is more accurate than the traditional node importance evaluation, can capture the nonlinear relationship and complex characteristics between network nodes, significantly reduce the redundant calculation of the network, and improve the efficiency and stability of the model. Experiments show that the method of the present invention can maintain or even improve the performance of the transfer learning model after pruning, and reduce the computational overhead by 20%-30%.
[0212] The similarity index calculation method proposed in the present invention dynamically adjusts and optimizes the weight initialization method and structural parameters of the source task model by analyzing the mean difference and covariance deviation between the feature distribution of the target task and the feature distribution of each layer of the source task. Compared with the traditional method of directly migrating weights, the method of the present invention can adjust the migration strategy in real time according to the feature similarity between the source task and the target task, reducing the occurrence of negative migration. The dynamic adjustment mechanism makes the model more flexible when facing the dynamic changes of the target task data distribution. Experimental results show that the introduction of the similarity index shortens the training time of the target task model by 15%-20% and improves the migration performance by 12%.
[0213] The present invention embeds a deep memory module in the target task network model to store high topological entropy node features in the source task and historical high-weight node features generated during the migration process. Through a dynamic update and reading mechanism, historical knowledge related to the current task can be flexibly called during the target task training process. This can not only help the model make full use of historical experience when facing complex tasks, but also dynamically adapt to the needs of the target task by eliminating redundant memory, thereby improving the learning efficiency and effect of the model. Compared with the model that does not include a memory module, the method of the present invention improves the prediction accuracy in complex task scenarios by 8%, shortens the training time by 25%, and performs particularly well in dynamically changing scenarios.
[0214] The above description is only a preferred specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any technician familiar with the technical field can make equivalent replacements or changes according to the technical scheme and inventive concept of the present invention within the technical scope disclosed by the present invention, which should be covered by the protection scope of the present invention.
Claims
1. A deep memory transfer learning method based on topological entropy decomposition, characterized in that: The steps include: S1. Obtain the source task dataset and the target task dataset, and perform data cleaning, feature extraction, and standardization on the source task dataset and the target task dataset respectively; S2. Use a deep neural network to train the preprocessed source task dataset to generate a source task training model; S3. Perform topological entropy analysis on the source task training model, calculate the topological entropy value of each network node in the source task training model, and quantify the contribution of each node to the overall performance of the source task training model; S4. Based on the topological entropy value, low topological entropy nodes and high topological entropy nodes are screened out according to a preset pruning threshold, and the low topological entropy nodes are pruned from the source task training model to obtain the pruned optimized source task model; S5. Perform migration adaptation on the optimized source task model, calculate the similarity index between each layer structure of the optimized source task model and the target task in combination with the feature distribution of the target task data set, and adjust the weight initialization method and structural parameters of the optimized source task model; S6. Use the target task dataset to perform preliminary training on the optimized source task model, generate a preliminary migration model, and record the performance indicators of the preliminary migration model on the target task; S7. Perform topological entropy analysis on the preliminary migration model again, calculate the topological entropy values of the newly added network nodes during the migration process, optimize the structure of the preliminary migration model in combination with the performance indicators, further remove redundant low topological entropy nodes during the migration process, and generate the target task network model; S8. embed a deep memory module in the target task network model, and use the deep memory module to store the high topological entropy node features in the source task model and the historical experience in the migration process; S9. Perform final training on the target task network model including the deep memory module, dynamically monitor the changes in the model's topological entropy during the training process, and adjust the usage strategy and pruning rules of the deep memory module; S10. Output the trained target task network model and record the pruning optimization process to form a reusable model optimization template.
2. According to the deep memory transfer learning method based on topological entropy decomposition in claim 1, it is characterized in that: The S1 specifically includes: S11. Obtain source task dataset And the target task dataset in, and Represent the feature vectors of the source task dataset and the target task dataset respectively, and Represents the label values of the source task dataset and the target task dataset, respectively, s and N t Represents the number of data samples for the source task and the target task respectively; S12. Source task dataset D s And the target task dataset D t Perform data cleaning to remove abnormal samples in the data set; S13. Perform feature extraction on the cleaned source task dataset and target task dataset, extract effective features through the predefined feature mapping function φ(x), and extract the source task feature set and the target task feature set S14. Extract the source task feature set F s and the target task feature set F t Perform standardization to obtain the standardized source task feature set F′ s and the target task feature set F′ t .
3. According to the deep memory transfer learning method based on topological entropy decomposition in claim 1, it is characterized in that: The S2 specifically includes: S21. Use the standardized source task feature set and label set as input data to build a deep neural network model; S22. Initialize the weight parameters and bias items of the deep neural network model, wherein the deep neural network includes an input layer, several hidden layers and an output layer, and the weight matrix and bias items of each layer are respectively Among them, L represents the number of layers of the neural network, W k is the weight matrix of the kth layer, b k is the bias term of the kth layer; S23. Use the forward propagation algorithm to train the deep neural network and transform the standardized source task features As input, the output is calculated through linear transformation and activation function processing of each layer. a k =s k (W k ·a k-1 +b k +∈ k ·D k +l·O k ),k=1,2,…,L; in, σ k (·) is the activation function of the kth layer, ∈ k is the topological perturbation factor related to the kth layer, which is used to introduce topological entropy perturbation in training, Δ k represents the interaction matrix between nodes, λ is the regularization coefficient, Ω k is the topological constraint, a k is the activation output of the kth layer, represents the predicted output of the model; S24. Use topological weighted mean square error loss function Calculate the loss value of the source task training model: in, is the true label, is the predicted output, H k is the topological entropy value of the kth layer, β k is the topological weighting coefficient, which is used to adjust the contribution of each layer of topological entropy to the overall loss in the loss function; S25. Use the back-propagation algorithm to calculate the gradient of the loss value relative to the weights and bias terms of each layer and Among them, γ is the topological regularization factor, which is used to balance the influence of topological entropy on the gradient, η k is the dynamic adjustment factor, Δ k represents the contribution of the influence of the interaction matrix between nodes to the gradient; S26. Update weights and bias terms using gradient descent: Among them, η is the learning rate, ν k is the topological constraint influence factor of the kth layer, Ω k It is a topology constraint item, which optimizes the stability and balance of the network structure by constraining the topology structure of each layer of the network; S27. Repeat steps S23 to S26 until the preset convergence conditions are met, and output the source task training model.
4. According to the deep memory transfer learning method based on topological entropy decomposition according to claim 1, it is characterized in that: The S3 specifically includes: S31. Construct the topological structure diagram of the source task training model: G = (V, E); Among them, V represents the node set in the source task training model, E represents the connection relationship between nodes, and for each node, its connection weight is defined as w ij , represents the node v i and v j The weight strength between them; S32. Calculate the local information entropy of each node of the source task training model Quantify the complexity of a node in the local topology: in, Represents node v i The neighborhood set of ; S33. Calculate the global topological entropy of the source task training model as a whole Describe the overall complexity of the model: in, Represents node v i degree; S34. Calculate the topological contribution C of each node by combining local information entropy and global topological entropy i : Among them, α∈[0,1] is the adjustment coefficient, which is used to balance the weights of local and global topological characteristics; S35. Normalize the topological contributions of all nodes to obtain the normalized node topological entropy value in, Represents node v i The relative contribution of the model to the whole model.
5. According to the deep memory transfer learning method based on topological entropy decomposition in claim 1, it is characterized in that: The S4 specifically includes: S41. Set the lower limit τ for pruning nodes with low topological entropy low and the upper limit τ of retaining nodes with high topological entropy high , τ low <τ high , according to the topological entropy Divide the model nodes into low topological entropy node sets V low and high topological entropy node set V high ; S42. For low topological entropy node set V low The nodes and their connections in are pruned, and the topological structure diagram of the source task training model is updated to G′=(V′,E′): V′=V\V low ; E′=E\{e ij ∣v i ∈V low some v j ∈V low }; S43. Update the pruned node weight matrix W′ k and the bias term b′ k : Among them, w ij and b ij Respectively represent the contribution of the weight value and bias term associated with the pruned node; S44. Recalculate the node connection strength for the updated topological structure graph G' and adjust the high topological entropy node set V high The network importance of redefines the connection weight w′ of high topological entropy nodes ij ; S45. Recalculate the loss function of the pruned model and adjust the network weights and bias terms through optimization. The loss function is defined as: Among them, H′ k is the global topological entropy value of the pruned model, β k and γ1 are weight coefficients; S46. Output the pruned optimized source task model, including the updated network weight matrix, bias term, and optimized topology structure diagram.
6. The deep memory transfer learning method based on topological entropy decomposition according to claim 1 is characterized in that: The S5 specifically includes: S51. Obtain the feature distribution information of the target task data set and calculate the feature set of the target task data set The mean vector μ t and the covariance matrix Σ t ; S52. The optimized source task model (W′ k ,b′ k ) to perform topological hierarchical analysis, extract the feature distribution of each layer of nodes in the optimization source task model, and calculate the feature mean vector μ′ of each layer k and the covariance matrix Σ′ k : in, To optimize the node activation value of the kth layer of the source task model, N s is the number of source task samples; S53. Calculate the target task feature distribution (μ t ,Σ t ) and optimize the feature distribution of each layer of the source task model (μ′ k ,Σ′ k ) k , defining the similarity index as the weighted sum of the maximum mean difference and covariance deviation between feature distributions: Among them, α1 and β1 are weighting coefficients, and ||·|| represents the Frobenius norm of the matrix; S54. Based on the calculated similarity index S k Adjust the weight initialization method of the optimized source task model and define the adjusted weight initialization matrix S55. Initialize the matrix based on the adjusted weights and the bias term Redefine the structural parameters of the optimized source task model: in, η k1 is the dynamic adjustment factor of the bias term.
7. The deep memory transfer learning method based on topological entropy decomposition according to claim 1 is characterized in that: The S6 specifically includes: S61. Using the standardized feature set and corresponding label set of the target task dataset as input data, the target task features are input into the optimized source task model after transfer adaptation, and processed by linear transformation and activation function of each layer in turn to finally obtain the prediction output; S62. Calculate the initial loss function of the target task training phase; S63. Calculate the gradient of the loss function with respect to the model weights and bias terms using a back-propagation algorithm; S64. Update the model weights and bias terms using the gradient descent method; S65. Repeat steps S61 to S64, perform multiple rounds of iterative training on the target task data set until the preset training convergence conditions are met, generate a preliminary migration model, and record the performance indicators of the preliminary migration model on the target task, including the prediction error of the target task and the accuracy of the model on the target task.
8. The deep memory transfer learning method based on topological entropy decomposition according to claim 1 is characterized in that: The S7 specifically includes: S71. Perform topological entropy analysis on the network topology structure of the preliminary migration model, and calculate the local topological entropy and global topological entropy of the newly added network nodes during the migration process. The local topological entropy is used to describe the connection complexity between the newly added nodes and their neighboring nodes, and the global topological entropy is used to measure the contribution of the newly added nodes to the overall complexity of the entire network. S72. The comprehensive importance weight of the newly added node is defined in combination with the performance index of the target task. The local topological entropy, global topological entropy, prediction error of the target task and the influence of the model accuracy on the importance of the node are comprehensively considered. The comprehensive importance weight is calculated as: Among them, α2, β2, γ2 and δ1 are importance weight adjustment coefficients. is the local topological entropy, is the global topological entropy; S73. Set a pruning threshold of the comprehensive importance weight, filter out a low-weight node set, which is a newly added node with a weight value lower than the pruning threshold, and mark the low-weight node as a redundant node; S74. Prune the low-weight nodes and their connections marked as redundant, update the topological structure of the preliminary migration model, and generate the target task network model.
9. The deep memory transfer learning method based on topological entropy decomposition according to claim 1 is characterized in that: The S8 specifically includes: S81. Embed a deep memory module in the target task network model, the deep memory module includes a storage unit, an update unit and a read unit, which are used to store high topological entropy node features in the source task model, dynamically update historical migration experience, and provide historical information required in the target task training process; S82. Extract the features of high topological entropy nodes from the source task model and store them in the storage unit of the deep memory module, and record the historical high-weight node features in the migration process in time series to form a continuous feature history track; S83. Define the update mechanism of the deep memory module, dynamically update the storage unit content according to the real-time training status and performance of the target task network model, remove redundant features and add new high-importance node features; S84. Define the reading mechanism of the deep memory module, select historical features with high similarity to the target task features from the storage unit according to the requirements of the target task training, dynamically load relevant experience through the reading unit, and optimize the learning efficiency of the task network model.
Citation Information
Patent Citations
Deep migration learning system and method based on entropy minimization
CN110580496A
Novel power system energy supply prediction method and system based on transfer learning
CN118822145A