Training method and device for regression stacked memory network
Through sparse connections and binary Gaussian distribution prediction of stacked memory networks, the long training time and overfitting problems of traditional regression algorithms on lightweight devices are solved, and rapid training and dynamic model adjustment are achieved, adapting to hardware resource changes, and improving prediction accuracy.
Patent Information
- Application Number
- CN202510325605.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-19
- Publication Date
- 2025-07-22
AI Technical Summary
Traditional regression algorithms have long training time, large computing resource usage, difficulty in model scaling and overfitting, especially on lightweight devices, which are difficult to effectively deploy and adjust.
A stacked memory network is used to disseminate information through a graph network composed of sparsely connected nodes and fully connected nodes, and a joint probability distribution of binary Gaussian distribution approximate features and the target is predicted, and the number of model layers is dynamically adjusted to adapt to hardware resource changes.
It realizes rapid training and deployment on lightweight devices, avoids overfitting, can dynamically adjust the model size to adapt to hardware resource changes, and improves training efficiency and prediction accuracy.
Smart Images

Figure CN120354375A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of regression analysis, and particularly relates to a training method and device for a stacked memory network for regression. Background Art
[0002] Regression problems are a core issue widely existing in the field of computer science, which require the model to learn from continuous input data and predict the corresponding continuous response variables. As the cornerstone of predictive analysis, the application of regression analysis spans multiple disciplinary fields, such as stock market prediction in economics, pain prediction in medicine, precipitation prediction in environmental science, material life prediction in engineering, and many other practical engineering practices, which play an extremely important role in revealing the relationships between variables and trend prediction.
[0003] Traditional regression problems use methods such as linear regression, ridge regression, Lasso regression, and SVR regression. Although these methods are simple and reliable, the fitting accuracy is difficult to meet the actual engineering requirements. With the development of artificial intelligence, deep neural network technology has also been used in regression tasks. However, deep neural networks often have a large model scale, and during the training process, they need to perform gradient descent for learning, which requires a long training time and consumes a large amount of hardware resources. Therefore, it is difficult to train and deploy in lightweight edge devices. At the same time, the above methods also have the following problems: 1. It is difficult to meet the requirements of model pruning and expansion. For example, when the trained model needs to be migrated to a lighter or more resourceful platform, it usually requires retraining the entire network, with a high cost. 2. It is prone to overfitting, especially when the test set is much larger than the training set. If the hyperparameters are not adjusted, overfitting problems are likely to occur, resulting in the accuracy of the test set being far lower than that of the training set. Summary of the Invention
[0004] The main object of the present invention is to overcome the problems of excessively high time cost, overfitting, and difficult model scaling commonly existing in existing regression algorithms, and provide a training method and device for a stacked memory network for regression.
[0005] The object of the present invention can be achieved by adopting the following technical solutions: In a first aspect, the present invention provides a training method for a stacked memory network for regression, and the training method includes the following steps. S1. Preprocess the training input X and the model fitting target y, initialize the number of layers k of the stacked memory network to 0, set the predicted value Y of the stacked memory network to 0, set the maximum target number of layers of the stacked memory network to K, and set the target mean square error to ; S2. Initialize the memory network. The memory network is a graph composed of sparsely connected nodes and fully connected nodes. Initialize the connection methods and initial connection weights of each node in the memory network, construct the adjacency matrix of the graph, and at the same time add this memory network to the last layer of the stacked memory network, increment the number of layers k of the stacked memory network by 1, and the current layer is the k-th layer; S3. The input of the memory network is the training input X, and the prediction target is the difference between the model fitting target y and the predicted value Y of the current stacked memory network, = y - Y. The training input X performs multiple rounds of feature propagation in the graph, and after the propagation, features are left on the sparsely connected nodes and fully connected nodes in the graph; S4. Each node feature is independently learned and the joint probability distribution F of; S5. The sparsely connected nodes and fully connected nodes perform regression prediction on the training input X according to the joint probability distribution F, and add the regression prediction result to the predicted value Y of the stacked memory network. Determine whether the current layer number k reaches the set maximum target layer number K or the mean square error of the predicted value Y of the stacked memory network is less than . If so, the training ends; otherwise, jump to step S2 to add the next layer of the memory network; S6. According to the requirements of the actual mean square error and the model size that change dynamically, expand or delete the number of layers of the stacked memory network. For expansion, only the newly added memory network layer needs to be trained, while for deletion, the redundant memory network layer is directly deleted.
[0006] Furthermore, in step S1, preprocess the training input X and the model fitting target y. The purpose of preprocessing is to unify the feature scales, avoid the model being biased towards certain features due to differences in the feature value ranges, and improve the stability of the model. Initialize the number of layers k of the stacked memory network to 0, the predicted value Y of the stacked memory network to 0, set the maximum target layer number to K, set the target mean square error to , and denote the value of the i-th feature in the training input X as . The operations of the preprocessing are as follows:
[0007]
[0008] where, represents taking the average of , takes the variance of , represents taking the minimum value of the model fitting target y, represents taking the maximum value of the model fitting target y.
[0009] Further, the mean square error is calculated as follows: .
[0010] Further, in step S2, the adjacency matrix of the k-th layer of the stacked memory network is set , and the memory network is a graph composed of sparsely connected nodes and fully connected nodes. The fully connected nodes include input nodes and hidden nodes, and the sparsely connected nodes only contain enhanced nodes. The hidden nodes are fully connected to the input nodes, and the enhanced nodes are randomly and sparsely connected to the input nodes and hidden nodes. Such a design has two purposes. One is to simulate the random connections between neurons in the human brain. Some studies have shown that neurons are not fully connected, but selectively establish connections with some neurons. This sparse connection method can process information more efficiently. The other is to make the values stored in the adjacency matrix as sparse as possible, reduce the parameter training amount of the model, so as to achieve the purpose of accelerating the training speed and reducing the memory occupancy. The weights between nodes are randomly sampled from [-1, 1], and the connection weights between nodes are fixed by random initialization and will not be changed subsequently. Therefore, like other neural network methods, it is trained for multiple rounds. The weights between non-connected nodes are set to 0. Among them, the number of input nodes is equal to the number of features d of the training input X, and the numbers of hidden nodes and enhanced nodes are the parameters u and v set during training respectively. There are a total of n = d + u + v nodes.
[0011] Further, in step S3, the training input X starts from the input nodes, and the information is propagated in the graph through the adjacency matrix for a total of p rounds. Each node contains an input pool I, a memory pool H, and an output pool M. Here, the multi-round information propagation method of the traditional graph neural network is used, and its advantage is that it can effectively extract high-dimensional features. Through multi-round information propagation, each node can continuously obtain information from its neighbor nodes, aggregate and update this information, so as to capture more complex feature representations. At the same time, the signal is continuously propagated through the input pool and the memory pool for multiple rounds, and this process further enhances the memory ability of the model. The input pool is responsible for inputting the initial features into the network, while the memory pool updates the memory state of the nodes after each round of propagation to ensure that important information is retained and transmitted. This design enables the model to effectively utilize the multi-round propagation historical information, and finally, relying on the gating mechanism, selectively outputs some information to the output pool to extract effective high-dimensional features. Initially, only the input pool of the input nodes contains the information of the training input X, and the initial value of the memory pool is the same as that of the input pool. The memory pool is initialized to 0, that is , where represents the initial values of the input pool, memory pool, and output pool of the memory network of the k-th layer. The information propagation process in the t-th round is modeled as:
[0012]
[0013]
[0014] Among them, respectively represent the information in the input pool, memory pool, and output pool during the t-th round of propagation of the memory network in the k-th layer, respectively represent the information in the input pool, memory pool, and output pool during the t-th round of propagation.
[0015] Furthermore, in the step S4, the joint probability distribution F of each node feature and the target is learned separately. In order to avoid the multi-round training method of neural networks, we use the feature in the sense of statistical significance of the joint probability distribution, which can model the relationship between the feature and the target at one time. Only one round of learning is required, and it is less likely to overfit, and the training speed is faster. After p rounds of information propagation, the information in the output pool of each node is used as the extracted feature, denoted as The feature on the i-th node in the k-th layer of the memory network is denoted as Assume that the feature on the i-th node and the target value of the current k-th layer of the memory network are represented by a bivariate Gaussian distribution Then:
[0016]
[0017]
[0018] Among them, represents concatenated with to form a vector, is the covariance matrix of the bivariate Gaussian distribution and is the mean vector of the bivariate Gaussian distribution represents the variance of represents the covariance between and represents the variance of the mean of
[0019] Further, in step S5, the node performs regression prediction on the training input X according to the joint probability distribution F. To reduce the interference of noise nodes, improve the stability and accuracy of prediction, and reduce the computational complexity, only the confidence of the node with the maximum confidence is used for prediction. Using the mean of the conditional distribution of the feature on the i-th node as the prediction result of the node, and the variance as the confidence of the node. Therefore, the process of regression prediction by the k-th layer memory network is as follows: First, calculate for the conditional distribution. The conditional distribution of the binary Gaussian can be expressed as a unary Gaussian distribution:
[0020]
[0021]
[0022] where represents the predicted value of the i-th node in the k-th layer memory network, represents the confidence of the i-th node in the k-th layer memory network, The smaller the value of, the greater the confidence. In the output of each stacked layer, only the output value of the node with the maximum confidence is used. Let represent the output of the k-th layer memory network, which is expressed as:
[0023] where represents taking the node index with the smallest Therefore, the output of the stacked network at the k-th layer is used as the predicted value Y of the stacked memory network, which is specifically as follows: , so that the residual corresponding to
[0024] is the expected predicted value of the next stack :
[0025] Judge whether the number of layers k reaches the set maximum target number of layers K or whether the mean square error of Y is less than the given . If so, the training ends; otherwise, jump to step S2 to add the next layer of the memory network.
[0026] Further, in the step S6, according to the requirements of the actual mean square error that changes dynamically and the requirements of the model size, the number of layers of the stacked memory network is expanded or reduced. Since the allocable resources of the actual hardware are changing in real time, for some hardware that needs to process multiple tasks simultaneously, such as edge devices like mobile phones, the available hardware resources for a given task need to change dynamically. If the model can change the model size according to the dynamic hardware resources to adapt to the actual mean square error requirements, this is meaningful for actual production and is also the advantage of this method compared with other neural network-based methods, which usually cannot perform dynamic adjustment of the model size. For expansion, only the newly added memory network layer needs to be trained, and the new maximum target number of layers of the stacked memory network is set to K new , the new target mean square error is set to , and steps S2 to S5 are repeated, and new memory network layers are added until the maximum target number of layers or the new target mean square error is reached. For deletion, the redundant memory network layers are directly deleted, and the actual size of each layer of the memory network is the same as , and only the first L layers need to be retained:
[0027] where is the rounding symbol, and [x] represents the largest integer not exceeding x.
[0028] In the second aspect, the present invention also discloses a training device for a stacked memory network for regression, which is used to run the above-mentioned training method for a stacked memory network for regression. The training device includes: An operation processing module, which is used to preprocess the training input X and the model fitting target y, initialize the number of layers k of the stacked memory network to 0, set the predicted value Y of the stacked memory network to 0, set the maximum target number of layers of the stacked memory network to K, and set the target mean square error to ; An initialization memory network module, which is used to initialize the memory network. The memory network is a graph composed of sparsely connected nodes and fully connected nodes. The connection method and initial connection weight of each node in the memory network are initialized, and the adjacency matrix of the graph is constructed. At the same time, the memory network is added to the last layer of the stacked memory network, the number of layers k of the stacked memory network is incremented by 1, and the current layer is the kth layer; A feature propagation module, where the input of the memory network is the training input X, and the prediction target is the difference between the model fitting target y and the predicted value Y of the current stacked memory network, = y - Y. The training input X performs multiple rounds of feature propagation in the graph, and after the propagation, features are left on the sparsely connected nodes and fully connected nodes in the graph; A joint probability distribution learning module, which is used to separately learn the features of each node and The joint probability distribution F; A regression prediction module, which is used to perform regression prediction on the training input X by sparsely connecting nodes and fully connecting according to the joint probability distribution F, and add the regression prediction result to the predicted value Y of the stacked memory network, and determine whether the current layer number k reaches the set maximum target layer number K or the mean square error of the predicted value Y of the stacked memory network Whether it is less than , if so, the training ends, otherwise it jumps to the initialization memory network module to continue adding the next layer of memory network; A stacked memory network expansion and deletion module, which is used to expand or delete the number of layers of the stacked memory network according to the actual mean square error requirement and model size requirement that change dynamically. For expansion, only the newly added memory network layer needs to be trained, while for deletion, the redundant memory network layer is directly deleted.
[0029] The present invention has the following advantages and effects compared with the prior art: (1) The present invention uses a bivariate Gaussian distribution to fit and approximate the distribution of each node and the target value, making the learning process very simple and efficient, avoiding a large amount of computing resources required by gradient descent, and having a fast training speed, which is suitable for deployment on lightweight edge devices.
[0030] (2) The present invention uses the conditional distribution of the bivariate Gaussian distribution for prediction, uses the mean value as the predicted value of each layer, and the variance as the confidence level, and can effectively avoid the overfitting problem without changing the hyperparameters.
[0031] (3) The present invention uses a multi-layer stacking method, and the previous layers do not depend on the subsequent layers. Therefore, the number of layers of the model can be dynamically increased and deleted without retraining the entire network, which is beneficial to adapting to dynamically changing hardware resources. Description of the Drawings
[0032] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0033] Figure 1 It is a flowchart of the stacked memory network algorithm for regression provided by an embodiment of the present invention; Figure 2 It is a histogram of each input feature and label of the California housing price dataset provided by an embodiment of the present invention; Figure 3 It is an architecture diagram of the stacked memory network algorithm for regression provided by an embodiment of the present invention. Detailed Embodiments
[0034] To enable those skilled in the art to better understand the solution of this application, the following will clearly and completely describe the technical solution in the embodiments of this application in conjunction with the accompanying drawings in the embodiments of this application. Obviously, the described embodiments are only a part of the embodiments of this application, rather than all the embodiments. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without creative efforts fall within the scope of protection of this application.
[0035] In this application, the mention of "embodiment" means that the specific features, structures or characteristics described in conjunction with the embodiment may be included in at least one embodiment of this application. The phrase appears in various positions in the specification does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment mutually exclusive with other embodiments. Those skilled in the art explicitly and implicitly understand that the embodiments described in this application can be combined with other embodiments.
[0036] Embodiment 1 This embodiment discloses a training method for a stacked memory network for regression, which is implemented through the following Figure 1 shown implementation steps. The steps of this training method include: S1. Preprocess the training input X and the model fitting target y, initialize the number of layers k of the stacked memory network to 0, the prediction value Y of the stacked memory network to 0, set the maximum target number of layers of the stacked memory network to K, and set the target mean square error to ; The specific implementation of step S1 is: The selected dataset is the California housing price dataset, Figure 2 is the histogram of each input feature and label of the California housing price dataset, Figure 3 is the architecture diagram of the method of the present invention. Preprocess the training input X and the model fitting target y, initialize the number of layers k of the stacked memory network to 0, the prediction value Y of the stacked memory network to 0, set the maximum target number of layers K to 10, and set the target mean square error to is , and the value of the i-th feature in the training input X is denoted as . The operations of the preprocessing are as follows:
[0037]
[0038] Among them, represents taking the average of , takes the variance of , represents taking the minimum value of the model fitting target y, Indicates maximizing the model fitting objective y.
[0039] Mean squared error is calculated as follows: .
[0040] S2. Initialize the memory network. The memory network is a graph composed of sparsely connected nodes and fully connected nodes. Initialize the connection method and initial connection weights of each node in the memory network, construct the adjacency matrix of the graph, and at the same time add this memory network to the last layer of the stacked memory network. The number of layers k of the stacked memory network is incremented by 1, and the current layer is the k-th layer.
[0041] The specific implementation of step S2 is: Set the adjacency matrix of the k-th layer of the stacked memory network , and the memory network is a graph composed of sparsely connected nodes and fully connected nodes. Among them, the fully connected nodes include input nodes and hidden nodes, and the sparsely connected nodes only contain enhancement nodes. The hidden nodes are fully connected to the input nodes, and the enhancement nodes are randomly and sparsely connected to the input nodes and hidden nodes. The number of sparse connections is set to 10; the weights between nodes are randomly sampled from [-1, 1], and the weights between non-connected nodes are set to 0. Since the input features of the California housing price dataset are 8, the number of input nodes is equal to the number of features 8 of the training input X. The number of hidden nodes u is set to 10, and the number of enhancement nodes v is set to 300. There are a total of n = d + u + v = 318 nodes.
[0042] S3. The input of the memory network is the training input X, and the prediction target is the difference between the model fitting objective y and the predicted value Y of the current stacked memory network, = y - Y. The training input X performs multi-round feature propagation in the graph. After the propagation, features are left on the sparsely connected nodes and fully connected nodes in the graph; The specific implementation of step S3 is: The training input X starts from the input nodes and propagates information in the graph through the adjacency matrix for a total of p rounds. p is set to 2. Each node contains an input pool I, a memory pool H, and an output pool M; Initially, only the input pool of the input nodes contains the information of the training input X, and the initial value of the memory pool is the same as the input pool. The memory pool is initialized to 0, that is , where represents the initial values of the input pool, memory pool, and output pool of the k-th layer of the memory network. The information propagation process in the t-th round is modeled as:
[0043]
[0044]
[0045] Among them, respectively represent the information in the input pool, memory pool, and output pool during the t-th round of propagation of the k-th layer of the memory network, respectively represent the information in the input pool, memory pool, and output pool during the
[0046] S4. Each node feature is separately learned and the joint probability distribution F with The implementation of step S4 is: Each node feature is separately learned and the joint probability distribution F with the target. After two rounds of information propagation, the information in the output pool on each node is used as the extracted feature, denoted as The feature on the i-th node in the k-th layer of the memory network is denoted as Assume that the feature on the i-th node and the target value of the current k-th layer of the memory network are represented by a bivariate Gaussian distribution Then:
[0047]
[0048]
[0049] In , represents concatenated with to form a vector, is the covariance matrix of the bivariate Gaussian distribution , is the mean vector of the bivariate Gaussian distribution , represents the variance of represents the covariance between and represents the variance of represents the mean of represents the mean of
[0050] S5. The sparse connection nodes and the fully connected nodes perform regression prediction on the training input X according to the joint probability distribution F, and add the regression prediction result to the predicted value Y of the stacked memory network. Determine whether the current layer number k reaches the set maximum target layer number K or the mean square error of the predicted value Y of the stacked memory network is less than , if so, the training ends; otherwise, jump to step S2 to add the next layer of memory network; The implementation of step S5 is as follows: The node performs regression prediction on the training input X according to the joint probability distribution F, and only uses the confidence of the node with the highest confidence for prediction. Using the feature on the i-th node the mean of the conditional distribution as the prediction result of the node, and the variance as the confidence of the node. Therefore, the process of regression prediction by the k-th layer memory network is as follows: First, calculate for the conditional distribution. The conditional distribution of the bivariate Gaussian can be expressed as a univariate Gaussian distribution:
[0051]
[0052]
[0053] where represents the predicted value of the i-th node in the k-th layer memory network, represents the confidence of the i-th node in the k-th layer memory network. The smaller the value of , the greater the confidence. In the output of each stacked layer, only the output value of the node with the highest confidence is used. Let
[0054] where represents taking the node index with the smallest . , therefore, when reaching the k-th layer, the output of the stacked network is used as the predicted value Y of the stacked memory network, which is specifically as follows:
[0055] The corresponding residual is the expected predicted value of the next stack :
[0056] Judge whether the number of layers k reaches the set maximum target number of layers 10 or the mean square error of Y is less than . If so, the training ends; otherwise, jump to step S2 to add the next layer of memory network.
[0057] S6. Expand or delete the number of layers of the stacked memory network according to the requirements of the actual mean square error and the model size that change dynamically. For expansion, only the newly added memory network layer needs to be trained, while for deletion, the redundant memory network layers are directly deleted.
[0058] The implementation manner of step S6 is: expand or delete the number of layers of the stacked memory network according to the requirements of the actual mean square error and the model size that change dynamically. For expansion, only the newly added memory network layer needs to be trained. Set the new maximum target number of layers of the stacked memory network to K new to be 20, and set the new target mean square error to remain unchanged as , repeat steps S2 to S5, and add new memory network layers until the maximum target number of layers or the new target mean square error is reached.
[0059] Figure 1 Steps S1 to S6 are shown, which are the training process of the stacked memory network for regression.
[0060] After the training is completed, it also includes a testing step: input the test samples into the stacked network, and starting from the first layer, execute steps S3 to S5 to obtain the final predicted value of the network.
[0061] To verify the effectiveness of the method disclosed in the present invention, on the commonly used regression dataset of the California housing price dataset, this dataset records 20,640 housing price information, including 8 input features such as the median income, the median house age, the average number of rooms per household, the average number of bedrooms per household, the number of people in the block, the average number of family members, longitude, and latitude, and predicts the median value of the houses in the block. In this dataset, the method of the present invention is compared with other classic lightweight regression algorithms. And in order to highlight the advantages of the performance and anti-overfitting of the present invention, the ratios of the training set to the test set are divided into 4:1 (Table 1) and 5:95 (Table 2) respectively for testing.
[0062] Table 1. Mean square errors of different methods on the California housing price dataset with a 4:1 division of the training set and the test set
[0063] Table 2. Mean square errors of different methods on the California housing price dataset with a 5:95 division of the training set and the test set
[0064] It can be seen from Table 1 that the present invention achieves the lowest mean square error compared with other methods and has higher performance.
[0065] As can be seen from Table 2, in the case of an extremely imbalanced training set and test set division of 5:95, that is, when the number of samples used for training is much smaller than the test samples, the overfitting situation will be more serious. However, the present invention still has the best performance under such a small amount of training samples, and the mean squared error is even 0.103 of the second-best performance Lasso regression method. The rest of the lightweight regression methods have serious overfitting under a very small amount of training samples, demonstrating the excellent anti-overfitting ability of this method.
[0066] Table 3 shows the performance change of the present invention after the number of layers of the stacked memory network changes from 10 to 20 in step S6. It is not necessary to retrain all layers, only the newly added layers need to be trained. After the expansion of the number of layers, the performance of the present invention is further improved. Using this method can dynamically adapt to the changes in the allocated hardware in actual changes, which has practical significance and is also the characteristic of the present invention different from other lightweight regression models.
[0067] Table 3. Influence of the change in the number of layers of the stacked memory network on the mean squared error of the California housing price dataset
[0068] Example 2 Referring to step S1 in Example 1, change the set maximum target number of layers K to 30, and the dataset becomes the diabetes dataset. This dataset records the data information of 442 diabetes patients, including 10 input features such as age, gender, body mass index, mean blood pressure, total cholesterol, low-density lipoprotein, high-density lipoprotein, the ratio of total cholesterol to high-density lipoprotein, the logarithm of triglycerides, and blood glucose level, and predicts the value of glycated hemoglobin one year later.
[0069] Steps S2 to S5 in this example refer to steps S2 to S5 in Example 1, and the operation steps are the same.
[0070] Referring to step S6 in Example 1, set the new maximum target number of layers of the stacked memory network as K new to be 40.
[0071] Table 4. Mean squared error of different methods on the diabetes dataset
[0072] Table 5. Influence of the change in the number of layers of the stacked memory network on the mean squared error of the diabetes dataset
[0073] As can be seen from Table 4, the present invention achieves the lowest mean square error compared with other methods and has higher performance. When the samples used for training are much smaller than the test samples, the overfitting situation will be more serious. However, the present invention still has the best performance in the case of such a small number of training samples. The performance of the mean square error is better than that of other regression methods, demonstrating the excellent overfitting prevention ability of this method.
[0074] Table 5 shows the performance change of the present invention after the number of layers of the stacked memory network changes from 30 to 40 after step S6. It is not necessary to repeat the training for all layers. Only the newly added layers need to be trained. After the expansion of the number of layers, the performance of the present invention is further improved. Using this method can dynamically adapt to the changes in the allocated hardware in actual changes, which has practical significance and is also a feature of the present invention different from other lightweight regression models.
[0075] It should be noted that for the foregoing method embodiments, for the sake of simple description, they are all expressed as a series of action combinations. However, those skilled in the art should know that the present invention is not limited by the described action sequence, because according to the present invention, certain steps can be performed in other sequences or simultaneously.
[0076] The technical features of the above embodiments can be combined arbitrarily. For the sake of concise description, not all possible combinations of the technical features in the above embodiments are described. However, as long as the combinations of these technical features do not conflict, they should all be considered to be within the scope described in this specification.
[0077] The above embodiments are preferred embodiments of the present invention, but the embodiments of the present invention are not limited by the above embodiments. Any other changes, modifications, substitutions, combinations, and simplifications made without departing from the spirit and principle of the present invention shall be equivalent replacement methods and are all included in the protection scope of the present invention.
Claims
1. A training method for a stacked memory network for regression, characterized in that, The training method includes the following steps: S1. Preprocess the training input X and the model fitting target y. Initialize the number of layers k of the stacked memory network to 0, the predicted value Y of the stacked memory network to 0, set the maximum target number of layers of the stacked memory network to K, and set the target mean squared error to ; S2. Initialize the memory network. The memory network is a graph composed of sparsely connected nodes and fully connected nodes. Initialize the connection mode and initial connection weights of each node in the memory network, construct the adjacency matrix of the graph, and at the same time add this memory network to the last layer of the stacked memory network, increment the number of layers k of the stacked memory network by 1, and the current layer is the kth layer; S3. The input of the memory network is the training input X, and the prediction target is the difference between the model fitting target y and the predicted value Y of the current stacked memory network, i.e., = y - Y. The training input X undergoes multiple rounds of feature propagation in the graph. After the propagation, features are left on the sparse connection nodes and fully connected nodes in the graph; S4. Learn the joint probability distribution F of each node feature separately and ; S5. The sparse connection nodes and the fully connected nodes perform regression prediction on the training input X according to the joint probability distribution F, and add the regression prediction result to the predicted value Y of the stacked memory network. It is judged whether the current layer number k reaches the set maximum target layer number K or whether the mean square error of the predicted value Y of the stacked memory network is less than . If so, the training ends; otherwise, it jumps to step S2 to add the next layer of the memory network. S6. According to the dynamically changing actual mean square error requirement and the requirement of the model size, expand or delete the number of layers of the stacked memory network. For expansion, only the newly added memory network layer needs to be trained, while for deletion, the redundant memory network layer is directly deleted.
2. The training method of the stacked memory network for regression according to claim 1, characterized in that In the step S1, the training input X and the model fitting target y are preprocessed. The number of layers k of the stacked memory network is initialized to 0, the predicted value Y of the stacked memory network is 0, the maximum number of target layers is set to K, and the target mean square error is set to , the value of the i-th feature in the training input X is denoted as , and the preprocessing operations are as follows: Among them, represents taking the average value of taking the variance of represents taking the minimum value of the model fitting target y, represents taking the maximum value of the model fitting target y. 3. The training method of the stacked memory network for regression according to claim 1, characterized in that The mean squared error is calculated as follows: .
4. The training method of the stacked memory network for regression according to claim 1, characterized in that In the step S2, an adjacency matrix of the k-th layer of the stacked memory network is set , where the memory network is a graph composed of sparsely connected nodes and fully connected nodes. The fully connected nodes include input nodes and hidden nodes, and the sparsely connected nodes only contain enhancement nodes. The hidden nodes are fully connected to the input nodes, and the enhancement nodes are randomly and sparsely connected to the input nodes and the hidden nodes; the weights between the nodes are randomly sampled from [-1, 1], and the weights between non-connected nodes are set to 0. Among them, the number of input nodes is equal to the number of features d of the training input X, and the numbers of hidden nodes and enhancement nodes are parameters u and v set during training respectively. There are a total of n = d + u + v nodes.
5. The training method of the stacked memory network for regression according to claim 1, wherein In the step S3, the training input X starts from the input node and propagates information in the graph through the adjacency matrix for a total of p rounds. Each node contains an input pool I, a memory pool H, and an output pool M. Initially, only the input pool of the input node contains the information of the training input X, and the initial value of the memory pool is the same as that of the input pool (the memory pool is initialized to 0), that is , where represents the initial values of the input pool, memory pool, and output pool of the memory network in the k-th layer. The information propagation process in the t-th round is modeled as: Among them, respectively represent the information in the input pool, memory pool, and output pool during the t-th round of propagation of the k-th layer of the memory network, respectively represent the information in the input pool, memory pool, and output pool during the t-th round of propagation.
6. The training method of the stacked memory network for regression according to claim 5, wherein In the step S4, the joint probability distribution F of each node feature and the target is learned separately, and the information in the output pool on each node is output after p rounds of information propagation. As the extracted feature, it is denoted as , and the feature on the i-th node in the memory network of the k-th layer is denoted as . Assume that the feature on the i-th node and the target value of the current memory network of the k-th layer of the joint probability distribution is expressed as a bivariate Gaussian distribution, then: Among them, denotes and the vector formed by splicing is the covariance matrix of the bivariate Gaussian distribution . is the mean vector of the bivariate Gaussian distribution . denotes the variance of denotes and the covariance of denotes the variance of denotes the mean of denotes the mean of 7. The training method of the stacked memory network for regression according to claim 6, wherein In the step S5, the node performs regression prediction on the training input X according to the joint probability distribution F, and only uses the confidence of the node with the highest confidence for prediction. Using the features on the i-th node the mean of the conditional distribution is used as the prediction result of the node, and the variance is used as the confidence of the node. Therefore, the process of the k-th layer memory network for regression prediction is as follows: First, calculate For The conditional distribution. The conditional distribution of a bivariate Gaussian can be expressed as a univariate Gaussian distribution: Among them represents the predicted value of the $i$-th node in the $k$-th layer of the memory network, represents the confidence of the $i$-th node in the $k$-th layer of the memory network, The smaller the value of, the greater the confidence. In the output of each stacked layer, only the output value of the node with the highest confidence is used, and is represented by represents the output of the $k$-th layer of the memory network, expressed as: Among them represents taking the smallest node index , therefore, the predicted value Y of the stacked network at the k-th layer is specifically as follows: The corresponding residual is the expected predicted value for the next stack : Determine whether the number of layers k has reached the set maximum target number of layers K or the mean square error of Y is less than the given , if so, the training ends, otherwise jump to step S2 to add the next layer of the memory network.
8. The training method of the stacked memory network for regression according to claim 1, characterized in that In step S6, according to the requirements of the actual mean square error and the model size that change dynamically, the number of layers of the stacked memory network is expanded or reduced. For expansion, only the newly added memory network layer needs to be trained. Set the new maximum target number of layers of the stacked memory network to K new , set the new target mean square error to , repeat steps S2 to S5, add new memory network layers until the maximum target number of layers or the new target mean square error is reached; for reduction, directly delete the redundant memory network layers, set the expected model size C, and the actual size of each memory network layer is the same as , only the first L layers need to be retained: Among them is the rounding symbol, and [x] represents the largest integer not exceeding x.
9. A training device for a stacked memory network for regression, which is used to execute the training method for the stacked memory network for regression according to any one of claims 1 to 8 above, characterized in that, The training device includes: An operation processing module is used to preprocess the training input X and the model fitting target y, initialize the number of layers k of the stacked memory network to 0, set the predicted value Y of the stacked memory network to 0, set the maximum target number of layers of the stacked memory network to K, and set the target mean squared error to ; An initial memory network module, which is used to initialize the memory network. The memory network is a graph composed of sparsely connected nodes and fully connected nodes. Initialize the connection mode and initial connection weights of each node in the memory network, construct the adjacency matrix of the graph, and at the same time add this memory network to the last layer of the stacked memory network, increment the number of layers k of the stacked memory network by 1, and the current layer is the kth layer; The feature propagation module, where the input to the memory network is the training input X, and the prediction target is the difference between the model fitting target y and the predicted value Y of the current stacked memory network, i.e., = y - Y. The training input X undergoes multiple rounds of feature propagation in the graph. After the propagation, features are left on the sparse connection nodes and fully connected nodes in the graph; The joint probability distribution school module is used to separately learn the joint probability distribution F of each node feature and ; A regression prediction module, which is used to perform regression prediction on the training input X by the sparsely connected nodes and the fully connected nodes according to the joint probability distribution F, and add the regression prediction result to the predicted value Y of the stacked memory network, and judge whether the current layer number k reaches the set maximum target layer number K or the mean square error of the predicted value Y of the stacked memory network is less than , if so, the training ends, otherwise, jump to the initialization memory network module to continue adding the next layer of the memory network; A stacked memory network expansion and deletion module, which is used to expand or delete the number of layers of the stacked memory network according to the dynamically changing actual mean square error requirement and the requirement of the model size. For expansion, only the newly added memory network layer needs to be trained, while for deletion, the redundant memory network layer is directly deleted.