Efficient memory consolidation model fusing overlapped components and based on index mechanism
Through the combination of hippocampal simulation network and vector quantization variational autoencoder, the memory consolidation model is optimized, which solves the problems of insufficient quality of generated samples and low computational efficiency, and realizes efficient and stable memory management, which is suitable for memory-intensive tasks in the fields of brain-like computing and artificial intelligence.
Patent Information
- Application Number
- CN202510336669.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-21
- Publication Date
- 2025-07-08
AI Technical Summary
In the prior art, insufficient quality of generated samples, lack of multimodal evaluation and low computational efficiency lead to insufficient incremental learning performance, especially on edge devices, which is difficult to meet real-time requirements.
The hippocampus simulation network (MCHN) is used as the teacher network and the vector quantized variational autoencoder (VQVAE) is used as the student network. Through noise retrieval sample generation and multi-index joint optimization, the biological memory playback process is simulated to achieve efficient memory consolidation.
It improves the quality of generated samples, enhances the stability and fidelity of memory, and improves computing efficiency, which is especially suitable for the real-time requirements of lightweight storage and edge devices.
Smart Images

Figure CN120278196A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the fields of incremental learning, image generation, lightweight model deployment, etc., and relates to an efficient memory consolidation model based on a teacher-student learning framework, and particularly relates to a memory enhancement method based on noise retrieval sample generation, variational autoencoder reconstruction, and multi-index joint optimization. Background Art
[0002] Memory consolidation models are a research direction at the intersection of machine learning, neuroscience, and computer vision. Traditional continuous learning methods (such as regularization, dynamic architectures) have a long development history and have made certain progress in alleviating the catastrophic forgetting problem, but still face bottlenecks such as low generation efficiency and large storage overhead. Since traditional models rely on static sample libraries or knowledge replay mechanisms with fixed rules, their adaptability will be restricted by complex scenarios such as task dynamic changes and data distribution shifts. Therefore, how to simulate the biological memory replay mechanism and achieve efficient and robust memory consolidation is the key challenge to improving the performance of incremental learning.
[0003] Models based on the teacher-student learning framework generate samples through the teacher network and iteratively optimize through the student network, which can reduce the dependence on raw data. Such methods improve the generalization ability of the model under lightweight storage conditions by simulating the memory replay process in neuroscience (such as the hippocampus-neocortex cooperation mechanism). However, the existing technologies still have the following problems:
[0004] 1. Insufficient quality of generated samples: The pseudo-samples generated by the teacher network have a large difference from the real data distribution, resulting in a high reconstruction error (MSE) of the student network (usually >0.05), which affects the performance of subsequent tasks;
[0005] 2. Lack of multi-modal evaluation: Existing methods rely on a single index (such as classification accuracy) and lack the joint optimization of generation quality (PSNR, SSIM) and memory stability (storage capacity N);
[0006] 3. Low computational efficiency: The structure of the decoder of the traditional variational autoencoder (VAE) is complex (such as the stacking of fully connected layers), and the number of deconvolution layers is redundant, making it difficult to meet the real-time requirements of edge devices (inference speed <30FPS). Summary of the Invention
[0007] The present invention discloses an efficient memory consolidation model based on a teacher-student learning framework, and particularly relates to a memory enhancement method based on noise retrieval sample generation, variational autoencoder reconstruction, and multi-index joint optimization, mainly solving the problems of low sample generation quality, poor memory stability, and insufficient computational efficiency in existing continuous learning methods.
[0008] The specific steps are as follows:
[0009] An efficient memory consolidation model that integrates overlapping components and an index-based mechanism. The learning framework of this efficient memory consolidation model includes:
[0010] (1) Teacher network (MCHN, Modern Continuous Hopfield network): The hippocampal simulation network MCHN is used as the teacher network for storage and retrieval. During the retrieval process of the teacher network, by inputting Gaussian noise with a standard deviation σ = 0.1, it simulates the memory replay process of the hippocampus to generate high-fidelity replay samples. The energy function and retrieval rules of the teacher network are as follows:
[0011]
[0012] ξ new = Xp, Xp ≡ softmax(βX T ξ) (2)
[0013] where E is the energy function, β is the inverse temperature parameter in the energy function, N = 1000 is the maximum storage capacity, x i is the i-th storage pattern, ξ represents the network state, M is the maximum norm, ξ new is the newly updated network state, X is the matrix composed of all storage patterns x1,...,x N , and p is the probability vector calculated through the softmax function;
[0014] (2) Student network: The Vector Quantized Variational Autoencoder (VQVAE) is used as the student network. The specific structure of the student network is as follows:
[0015] Encoder maps the input data x from the original space to the latent space to obtain the feature vector z e (x) extracted by the encoder;
[0016] z e (x) = f(φ1, x) (3)
[0017] The codebook is initially defined as ε ∈ R K×D , where K represents the latent space size, D is the dimension of each vector in the codebook, and the codebook vector e j with serial number j is a subset of ε; The output z e (x) of the encoder is compared with the j-th codebook vector e jCalculate using the nearest neighbor algorithm formula (4). After calculating all the codebook vectors, generate J, which is an index matrix of w×h, and obtain the closest d-dimensional vector e in the codebook corresponding to each index. J * , Extract the w×h codebook vectors corresponding to the index matrix as the output z of the codebook. q (x);
[0018] z q (x) = e J * , J = argmin J ||z e (x) - e j ||2 (4)
[0019] Decoder Is responsible for reconstructing or generating data according to z q (x).;
[0020]
[0021] Among them, Is the image finally generated by the student network decoder. The encoder and decoder are respectively implemented by neural networks with parameters φ1 and φ2;
[0022] The goal of the student network is to update the parameters of the encoder, decoder, and codebook by optimizing the loss function formula (6), so that the codebook vectors e i , e i Is a subset of the codebook ε, approximates the encoder output z, and minimizes the difference between the input data x of the student network encoder and the data Generated by the student network decoder. The total loss function L is:
[0023]
[0024] Among them, z e (x) is the vector output by the encoder, Is a hyperparameter that controls the speed of update of the encoder parameters. sg is the stop gradient operator;
[0025] The specific steps are as follows:
[0026] (1) Initial encoding of memory: Simulate the initial learning of the biological hippocampus and form a memory mechanism. Use the storage function of the teacher network to perform the initial encoding and storage of events, that is, images;
[0027] (2) Memory replay: Input Gaussian noise into the teacher network, use formulas (1) and (2) to retrieve events, and pass the retrieval results to the student network;
[0028] (3) Vector Quantized Variational Autoencoder (VQVAE): The encoder in the student network extracts features from the events of memory replay according to formula (3) to obtain the feature vector z e (x), and then uses the VQVAE vector quantization technology to obtain z e (x) as the d-dimensional vector e closest to it in the corresponding codebook J * , and take out the w×h codebook vectors corresponding to the index matrix as the output z q (x) of the codebook. Here, through the update of the codebook parameters under the guidance of the loss function, the codebook vectors compress the learned semantic features. Features with similar semantics (such as the same semantics, such as shape and size) share the same codebook vector. Through the highly compressed characteristics of the codebook vectors, the encoding of events is realized, and the index matrix is stored in the teacher network to realize the restoration of the hippocampal index theory;
[0029] (4) Memory reconstruction: The codebook vectors are optimized through training so that features with similar semantics are close in the embedding space, and the consistency with the function of concept cells in the codebook vectors is verified through t-SNE visualization, and the decoder in the student network is used to reconstruct the memory events;
[0030] (5) Memory consolidation and replay: In the sleep or calm state, repeat step (2) to activate the teacher network through noise input to retrieve new content, and repeat steps (3) and (4), use the student network to reconstruct the memory events, convert short-term memory into long-term memory, and optimize memory storage using indexing technology;
[0031] (6) Evaluation of the memory model: Support memory recall tasks, and use the mean square error, structural similarity index, and peak signal-to-noise ratio to evaluate the reconstruction effect;
[0032]
[0033] where x i is the value of the i-th pixel of the original image, and x i ′ is the value of the i-th pixel of the reconstructed image, and N is the total number of pixels of the image;
[0034]
[0035] where MAX I is the maximum possible value of the pixel value;
[0036]
[0037] where μ x , μ x′ are the pixel means of the original image and the reconstructed image; σ x , σx′ is the pixel standard deviation of the original image and the reconstructed image; σ xx′ is the covariance between the original image and the reconstructed image; C1 = (k1L) 2 、C2 = (k2L) 2 are stability constants, k1 = 0.01, k2 = 0.03; L is the pixel dynamic range.
[0038] In addition, the memory consolidation model also supports memory imagination tasks and uses the Inception score (IS) to evaluate the reconstruction effect:
[0039]
[0040] where represents the expectation of the generated image x under the generated distribution p g , D KL is the KL divergence symbol, representing the difference between two distributions, p(y|x) is the class distribution predicted by the Inception model given the image x, and p(y) is the marginal distribution, representing the mean of the class prediction distributions of all generated images.
[0041] In addition, the memory consolidation model also supports using a support vector classifier (SVC) to predict image labels based on latent vectors in semantic memory experiments. As the training progresses, the reconstruction error of the model (measuring the difference between the reconstructed image of the model and the original image, using the mean squared error MSE, the smaller the error, the higher the reconstruction quality) gradually decreases, and the decoding accuracy (measuring the correct rate of the classifier predicting labels based on latent vectors, the higher the accuracy, the stronger the semantic separability of the latent representation) significantly improves, indicating that the class characteristics of the latent representation are gradually enhanced and the semantic features are gradually highlighted. The experimental results prove that the performance of the present invention in the semantic structure classification task is superior to the prior art.
[0042] Advantages of the present invention: The present invention proposes an efficient memory consolidation model based on overlapping components and an indexing mechanism. By combining vector quantization technology and a hippocampal simulation network (MCHN), the memory encoding and storage efficiency are optimized. While ensuring high-precision memory retrieval, the method uses overlapping component design to achieve shared storage of semantically similar events, significantly improving the memory capacity and retrieval speed. By noise activation and generative reconstruction, the memory consolidation mechanism during the biological sleep period is simulated, enhancing the stability and fidelity of memory. Finally, efficient and stable long-term memory management is achieved, which is particularly suitable for memory-intensive tasks in the fields of brain-like computing and artificial intelligence and has broad application potential. BRIEF DESCRIPTION OF THE DRAWINGS
[0043] Figure 1 is the model architecture diagram of the method of the present invention, showing the connections of the encoder, MCHN, and generation network.
[0044] Figure 2 This is the schematic diagram of various task processes for the method of the present invention.
[0045] Figure 3 This is the result of semantic classification for the method of the present invention. Among them, A) is the comparison of decoding accuracy between the baseline (orange) and the model of the present invention (yellow), and B) is the training process of the support vector classifier. Detailed implementation manners
[0046] The following further describes the detailed implementation manners of the present invention in combination with the accompanying drawings and technical solutions.
[0047] I. Dataset preprocessing
[0048] The following dataset is used for experiments: Shape3D. Shape3D is a dataset generated based on 6 independent latent factors (including floor color, wall color, object color, scale, shape, and orientation). All possible combinations appear only once, and a total of N = 480,000 images are generated. The Shape3D dataset provides rich scene diversity through a programmatic generation method and is suitable for simulating memory encoding, storage, and retrieval tasks.
[0049] II. Network structure construction
[0050] The network structure of the present invention is based on the teacher-student learning framework and consists of two parts: the teacher network (MCHN) and the student network (VQ-VAE), which are responsible for memory storage and retrieval and memory reconstruction tasks respectively.
[0051] (1) Teacher network (MCHN, Modern Continuous Hopfield network): The modern continuous Hopfield network (MCHN) stores N patterns represented by the matrix X = (x1,...,x N ) where the maximum norm is M = max i ‖x i ‖, and uses the energy function formula (1) to evaluate the similarity between the stored pattern and the current state pattern, thereby quantifying the network state and achieving efficient pattern retrieval and memory management. During the retrieval process, the teacher network inputs Gaussian noise with a standard deviation σ = 0.1 to simulate the memory replay process of the hippocampus and generate high-fidelity replay samples
[0052] (2) Student network: Use the vector quantization variational autoencoder (VQVAE) (Vector Quantized Variational Autoencoder) as the student network. The specific structure of the student network is as follows: the encoder Map the input data x from the original space X to the latent space according to formula (3). Obtain the feature vector z extracted by the encoder e The codebook is initially defined as ε ∈ R K×D , where K represents the latent space size, D is the dimension of each vector in the codebook, and the codebook vector e with serial number j j is a subset of ε; compare the output z e (x) of the encoder with the j-th codebook vector e j Calculate using the nearest neighbor algorithm formula (4). After calculating all the codebook vectors, generate J, which is the index matrix of w×h, and obtain the d-dimensional vector e in the codebook that is closest to each index J * , and extract the w×h codebook vectors corresponding to the index matrix as the output z q (x) of the codebook. The decoder is responsible for reconstructing or generating data based on z q (x).
[0053] III. Training Strategy
[0054] The training strategy of MCHN (Modern Continuous Hopfield Network) mainly includes the following steps: First, define the storage mode of the network through the energy function formula (1). Second, during the training process, simulate the memory retrieval process by inputting Gaussian noise, and update the network state using the retrieval rule formula (2) to gradually optimize the stability and retrieval efficiency of the storage mode. Finally, through the teacher-student learning framework, use the high-quality replay samples generated by MCHN to guide the student network (such as VQ-VAE) to ensure the efficient consolidation and reconstruction of memory.
[0055] The goal of the student network is to update the parameters of the encoder, decoder, and codebook by optimizing the loss function formula (6), so that the codebook vectors e i , e i are subsets of the codebook ε, approximate the encoder output z, and minimize the difference between the input data x of the student network encoder and the data generated by the student network decoder, that is, minimize the total loss function L formula (6).
[0056] III. Model Evaluation
[0057] In the real world, the recollection of memories is often not complete but rather subject to various interferences, including noise. By deliberately introducing noise into the input data and testing whether the model can accurately reconstruct the original scene under such interference, the present invention can test and optimize the model's ability to handle imperfect data in the real world. To quantitatively evaluate robustness, the present invention calculates the corresponding metrics between the reconstructed image and the image without added noise. The results show superior reconstruction quality. The evaluation process uses metrics such as mean squared error (MSE), structural similarity index (SSIM), and peak signal-to-noise ratio (PSNR) to evaluate the reconstruction effect, where MSE measures the pixel-level difference between the reconstructed image and the original image, SSIM evaluates the similarity of the image structure, and PSNR reflects the signal-to-noise ratio of the image; meanwhile, the quality of the generated images in the imagination experiment is evaluated through the Inception score (IS), which measures the diversity and authenticity of the generated distribution;
[0058] In the imagination experiment, the present invention applies PixelCNN (Pixel Convolutional Neural Network) to the index matrix to provide precise support for calculating the activation sequence for imagination modeling. PixelCNN is a generative model commonly used in image generation tasks, and its working principle is to gradually generate the pixels of an image, starting from the upper left corner and generating the value of each pixel row by row and column by column. The present invention models the generation process of PixelCNN as a process of imagination because the creative and random combination of imagination is satisfied. The present invention uses the IS metric to measure the quality and diversity of the images generated in the imagination experiment.
[0059] In the semantic memory experiment, a support vector classifier (SVC) is used to predict image labels based on latent vectors. As the training progresses, the reconstruction error gradually decreases and the decoding accuracy significantly improves, indicating that the semantic characteristics of the latent representation are gradually enhanced. The experimental results prove that the performance of the present invention in the semantic structure classification task is superior to the prior art.
[0060] Example 1
[0061] As Figure 3 shown, the present invention proposes an efficient memory consolidation model based on a teacher-student learning framework, aiming to significantly improve the storage and recall efficiency of memories through an innovative learning mechanism. The model uses Shape3D as the core dataset. In the experimental design, the present invention not only tests the accuracy of memory recall but also deeply explores the reproducibility of episodic memory and the creative performance in imagination experiments. The experimental results show that the present model is significantly superior to the prior art in all metrics, especially in terms of memory extraction and reconstruction capabilities in complex scenarios, demonstrating its unique advantages.
[0062] Table 1 Memory recall (i.e., partial image reconstruction), episodic memory (i.e., retrieval image (hippocampal replay) reconstruction), and imagination experiment.
[0063]
Claims
1. An efficient memory consolidation model that integrates overlapping components and an indexing mechanism, characterized in that The learning framework of the efficient memory consolidation model includes: (1) Teacher network: The hippocampal simulation network MCHN is used as the teacher network for storage and retrieval. During the retrieval process, the teacher network generates high-fidelity replay samples by inputting Gaussian noise with a standard deviation σ = 0.1 to simulate the memory replay process of the hippocampus. The energy function and retrieval rules of the teacher network are as follows: ξ new = Xp, Xp ≡ softmax(βX T ξ) (2) Among them, E is the energy function, β is the inverse temperature parameter in the energy function, N = 1000 is the maximum storage capacity, and x i is the i-th storage pattern, ξ represents the network state, M is the maximum norm, and ξ new is the newly updated network state, X is the matrix composed of all storage patterns x1,..., x N N, and p is the probability vector calculated through the softmax function; (2) Student network: The vector quantization variational autoencoder (VQVAE) is used as the student network. The specific structure of the student network is as follows: Encoder Maps the input data x from the original space To the latent space To obtain the feature vector z extracted by the encoder e (x); z e (x) = f(φ1, x) (3) The codebook is initially defined as ε ∈ R K×D , where K represents the size of the latent space, D is the dimension of each vector in the codebook, and the codebook vector e with index j j is a subset of ε; the output z e (x) of the encoder and the j-th codebook vector e j are calculated using the nearest neighbor algorithm formula (4). After calculating all the codebook vectors, J, which is an index matrix of w×h, is generated, and the d-dimensional vector e closest to each index in the codebook is obtained J * . The w×h codebook vectors corresponding to the index matrix are taken as the output z q (x) of the codebook; Decoder (g: Responsible for reconstructing or generating data according to z q (x). Among them, is the image finally generated by the student network decoder, and the encoder and decoder are respectively implemented by neural networks with parameters φ1 and φ2; The goal of the student network is to update the parameters of the encoder, decoder, and codebook by optimizing the loss function formula (6), so that the codebook vectors e i , e i which are subsets of the codebook ε, approximate the encoder output z, and minimize the difference between the input data x of the student network encoder and the data generated by the student network decoder . The total loss function L is as follows: where z e (x) is the vector output by the encoder, is the hyperparameter that controls the speed of the encoder parameter update change, and sg is the stop gradient operator; The specific steps are as follows: (1) Initial memory encoding: Simulate the initial learning of the biological hippocampus and form a memory mechanism. Use the storage function of the teacher network to perform the initial encoding and storage of events, i.e., images. (2) Memory replay: Input Gaussian noise into the teacher network, retrieve events using formulas (1) and (2), and pass the retrieval results to the student network. (3) Vector Quantized Variational Autoencoder (VQVAE): The encoder in the student network extracts features from the events of memory replay according to formula (3) to obtain the feature vector z e (x), and then uses the VQVAE vector quantization technology to obtain z e (x) as the d-dimensional vector e closest to it in the corresponding codebook J * , and take out the w×h codebook vectors corresponding to the index matrix as the output z q (x) of the codebook. Here, through the update of the codebook parameters under the guidance of the loss function, the learned semantic features are compressed in the codebook vectors, and semantically similar features share the same codebook vector. Through the highly compressed characteristics of the codebook vectors, the encoding of events is realized, and the index matrix is stored in the teacher network to realize the restoration of the hippocampal index theory; (4) Memory reconstruction: The codebook vectors are optimized through training so that features with similar semantics are close in the embedding space. Visualize the consistency with the function of concept cells in the codebook vectors through t-SNE, and use the decoder in the student network to reconstruct the memory events. (5) Memory consolidation and replay: In the sleep or calm state, repeat step (2) to activate the teacher network to retrieve new content through noise input, and repeat steps (3) and (4). Use the student network to reconstruct the memory events, convert short-term memory into long-term memory, and optimize memory storage using indexing techniques. (6) Evaluation of the memory model: Support memory recall tasks, and evaluate the reconstruction effect using the mean squared error, structural similarity index, and peak signal-to-noise ratio. where x i is the value of the i-th pixel of the original image, and x i ′ is the value of the i-th pixel of the reconstructed image; N is the total number of pixels in the image. where MAX I is the maximum possible value of the pixel value; Among them, μ x , μ x′ are the pixel means of the original image and the reconstructed image; σ x , σ x′ are the pixel standard deviations of the original image and the reconstructed image; σ xx′ is the covariance between the original image and the reconstructed image; C1 = (k1L) 2 , C2 = (k2L) 2 are stability constants, k1 = 0.01, k2 = 0.03; L is the pixel dynamic range.