Hebb learning and sleep enhancement strategy based on brain-like intelligent agent
By combining Herb learning rules and sleep enhancement mechanisms in brain-like agents, dynamically adjusting the weight of the association matrix and performing sparse optimization, the shortcomings of traditional time series learning algorithms in long-term dependence and noise interference are solved, and more efficient learning and memory effects are achieved.
Patent Information
- Application Number
- CN202510073130.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-16
- Publication Date
- 2025-05-13
AI Technical Summary
Traditional time series learning algorithms have problems such as gradient disappearance when dealing with long-term dependencies, and lack the ability to integrate memory in biological systems, and cannot effectively deal with noise interference.
A Herb learning and sleep enhancement strategy based on brain-like agents is adopted. The time-step sequence and action sequence are obtained through the time series input module, the association matrix is initialized, the association matrix weight is dynamically adjusted based on the Herb learning rules, and the spontaneous activity of neurons is simulated during the sleep stage. High confidence correlation is strengthened through the positive and negative sample adjustment mechanism, low confidence correlation is weakened, and matrix weights are sparsely optimized.
By simulating the sleep process of the biological brain, the learning efficiency and memory effect of complex time series data are improved, the learning and memory ability of artificial neural networks are enhanced, and the performance of agents in complex tasks is significantly improved.
Smart Images

Figure CN119990175A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the fields of artificial intelligence and neural computing, and in particular to a Hebbian learning and sleep enhancement strategy based on a brain-like intelligent agent. Background Art
[0002] With the rapid development of artificial intelligence technology, more and more intelligent systems need to process complex time series data, such as speech recognition, action prediction, and brain signal analysis. Such tasks require not only that the model can learn complex dependencies between time steps, but also that it has the ability to quickly adapt and resist interference in dynamic environments. However, traditional sequence learning algorithms such as recurrent neural networks (RNNs) and long short-term memory networks (LSTMs) have problems such as gradient vanishing when dealing with long-term dependencies, and lack the ability to integrate memory in biological systems, and cannot effectively deal with noise interference.
[0003] Hebbian Learning, as a biologically based learning mechanism, adjusts synaptic weights through the coordinated activation of neurons, and has significant advantages in simulating short-term learning behaviors. At the same time, studies have shown that the brain can integrate and optimize memory through spontaneous neuronal activity and synaptic plasticity during sleep. However, how to combine Hebbian learning rules with sleep enhancement mechanisms to build a brain-like intelligent entity with autonomous learning and memory consolidation capabilities remains an unresolved technical problem.
[0004] Therefore, designing an optimization method based on Hebbian learning and sleep enhancement strategy to solve the shortcomings of traditional sequence learning methods in long-term dependency capture and anti-interference by simulating the learning and memory mechanism of the biological nervous system is an urgent problem to be solved in the current technical field. Summary of the invention
[0005] In view of the above-mentioned defects of the prior art, the technical problem to be solved by the present invention is that the traditional time series learning algorithm has problems such as gradient vanishing when dealing with long-term dependencies, and lacks the ability of memory integration in biological systems and cannot effectively deal with noise interference.
[0006] To achieve the above object, the present invention provides a Hebbian learning and sleep enhancement strategy based on a brain-like intelligent agent, comprising the following steps:
[0007] (1) Obtain the time step sequence and action sequence through the time series input module;
[0008] (2) Initialize the association matrix and generate the initial association between time steps and actions based on regular weight assignment and random noise;
[0009] (3) Dynamically adjust the weights of the association matrix based on the Hebbian learning rule;
[0010] (4) Simulate the spontaneous activity of neurons during sleep, strengthen high-confidence associations and weaken low-confidence associations through the positive and negative sample adjustment mechanism, and perform sparse optimization on the matrix weights;
[0011] (5) Based on the optimized association matrix, the optimal action sequence corresponding to the inferred time step is estimated using the maximum a posteriori probability.
[0012] The present invention enhances sequence learning through learning and sleep enhancement, simulates the spontaneous activity of neurons and the memory replay process, optimizes memory through the positive and negative sample weight adjustment mechanism, accurately grasps the action sequence, and filters it out.
[0013] In a preferred embodiment of the present invention, step (2) also includes normalizing the matrix.
[0014] In another preferred embodiment of the present invention, the initialization of the correlation matrix in step (2) comprises the following steps:
[0015] (2.1) Divide the time step into at least two regular intervals, each of which is assigned a strong initial weight to a specific action group;
[0016] Among them, the initial weights assigned to each interval and a specific action group have obvious distinguishable patterns. According to the weight-normalized probability distribution, 1 is the strongest and 0 is the weakest (no correlation probability). The difference between strong and weak should be greater than 0.5 as a judgment with strong discrimination.
[0017] (2.2) Based on the regular distribution, random noise weights are introduced for the remaining actions;
[0018] (2.3) Each row of the matrix is normalized so that its weight sum is 1.
[0019] In another preferred embodiment of the present invention, the dynamic adjustment of the association matrix weights based on the Hebbian learning rule in step (3) includes strengthening the significant association between the current time step and the action, and performing recursive weight updates in combination with the previous and next time steps.
[0020] In another preferred embodiment of the present invention, the weight adjustment of the Hebbian learning rule in step (3) includes:
[0021] (3.1) For the current time step l i and action a i Increase its weight:
[0022] P[l i ,a i ]←P[l i ,a i ]+η
[0023] (3.2)3.2 If there is a previous time step l i-1 and action a i-1 , then recursively update its weight:
[0024]
[0025] (3.3) The rows corresponding to the updated time steps are normalized to ensure the numerical stability of the matrix.
[0026] In another preferred embodiment of the present invention, in step (3), the association weights of time steps and actions are adjusted by the Hebbian learning rule in the learning phase, and the specific rules include:
[0027] Co-activation enhancement rule: time l i and action a i When activated, its associated weight increases;
[0028] History association recursive rule: If time step l i-1 and action a i-1 is the predecessor of the current time step, its weight is recursively enhanced;
[0029] The adjusted correlation matrix is normalized to maintain the numerical stability of the weights.
[0030] The present invention strengthens the synaptic association through Hebbian learning and has a biomimetic basis through the memory association of timing-dependent synaptic plasticity.
[0031] Through multi-stage enhancement, the final memory effect is evaluated by MAP estimation of decision effect, which achieves good results, with an action sequence accuracy of 100% (under random noise) and a sparsity of more than 96%.
[0032] In another preferred embodiment of the present invention, in step (4), the association matrix is optimized by the timing-dependent synaptic plasticity mechanism during the sleep enhancement stage, and the specific steps include the following:
[0033] (4.1) During the sleep enhancement phase, the spontaneous activity of neurons is simulated by randomly selecting time steps and actions;
[0034] (4.2) Positive sample enhancement: If the association probability between a time step and an action is higher than a set threshold, its association weight is enhanced;
[0035] (4.3) Negative sample weakening: If the association probability between a time step and an action is lower than a set threshold, its association weight is weakened;
[0036] (4.4) Sparse processing: weights below the set threshold are clipped to zero;
[0037] (4.5) Weight normalization: Renormalize the matrix to maintain numerical balance and sparsity.
[0038] In another preferred embodiment of the present invention, when inferring the optimal action sequence corresponding to the time step in step (5), a maximum a posteriori probability estimation method is used, which specifically includes:
[0039] (5.1) For each time step, the optimal action is selected according to the following formula
[0040] (5.2) Output the complete inferred action sequence
[0041] In another preferred embodiment of the present invention, the strategy further comprises step (6): displaying the changing process of the weight distribution of the correlation matrix in the learning and sleep enhancement stages through a dynamic visualization module.
[0042] In another preferred embodiment of the present invention, in step (6), the dynamic visualization module is used to display the weight changes of the correlation matrix at different stages, including
[0043] (6.1) Matrix distribution after strengthening significant associations in the learning phase;
[0044] (6.2) Matrix distribution after sparse optimization in the sleep enhancement phase.
[0045] Technical Effects
[0046] The present invention relates to a Hebbian learning sleep expert strategy based on brain-like intelligent agents, which aims to improve the learning efficiency and memory effect of complex time series data and enhance the learning and memory ability of artificial neural networks by simulating the memory consolidation mechanism of biological brains during sleep. This strategy combines the Hebbian learning rule, the synaptic plasticity of neural networks and the replay process of neural activity during sleep to achieve adaptive learning and knowledge integration of brain-like intelligent agents. It is widely applicable to the fields of artificial intelligence, brain-computer interface and intelligent control.
[0047] The brain-like intelligent agent of the present invention includes two main stages: a learning stage and a sleep consolidation stage. In the learning stage, by inputting training data, the synaptic weights of the neural network are adjusted by applying the Hebbian learning rule to form a preliminary memory representation. The association matrix is updated through the coordinated activation between time steps and actions. The association of the current time step action is strengthened by the dynamic learning rate, and the intrinsic timing characteristics of the action sequence are captured by combining the historical memory information of the previous and next time steps. In this stage, a stable association model of time steps and actions is formed through multiple rounds of iterative training. In the sleep consolidation stage, the intelligent agent simulates the deep sleep and rapid eye movement sleep of the biological brain, activates neurons randomly or based on specific strategies, replays the neural activity pattern of the learning stage, and further strengthens the synaptic connection between related neurons. The association matrix is enhanced and sparsely optimized by the timing-dependent synaptic plasticity rule. In this stage, positive sample reinforcement is the main method, and negative sample weakening is the auxiliary method. Through spontaneous activation and random replay mechanisms, high-confidence associations are consolidated and low-confidence associations are weakened, further optimizing the memory capacity and anti-interference ability of the model.
[0048] By repeating the sleep consolidation process many times, the present invention can effectively enhance the memory stability and anti-interference ability of brain-like intelligent agents for learned information, and significantly improve the performance of intelligent agents in complex tasks. This strategy is not only applicable to traditional artificial neural network models, but can also be extended to various deep learning frameworks, with broad application prospects and potential value.
[0049] advantage:
[0050] (1) Improve memory consolidation efficiency: By simulating the memory consolidation mechanism during biological sleep, the memory stability of the intelligent agent for learned information can be significantly improved.
[0051] (2) Enhance anti-interference ability: Through multiple replays and synaptic weight adjustments, the anti-interference ability of memory is enhanced, making the intelligent agent more stable in complex environments.
[0052] (3) Adaptability to multiple models: This strategy can be applied to various artificial neural networks and deep learning frameworks, and has high versatility and portability.
[0053] The present invention is applicable to various intelligent systems that require efficient learning and memory, including but not limited to autonomous driving, intelligent robots, speech recognition, image processing and other artificial intelligence application fields. It has significant technical advantages and application value, especially in scenarios where complex tasks and long-term memory requirements are high.
[0054] Through the present invention, brain-like intelligent agents can be closer to the memory mechanism of biological brains, significantly improving the overall performance and application effects of the intelligent agents.
[0055] The concept, specific structure and technical effects of the present invention will be further described below in conjunction with the accompanying drawings to fully understand the purpose, characteristics and effects of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0056] Figure 1 It is a flow chart of a preferred embodiment of the present invention;
[0057] Figure 2 It is a flow chart of a learning algorithm for learning phase actions and time step correlation matrices based on Hebbian learning in a preferred embodiment of the present invention;
[0058] Figure 3 It is a schematic diagram of an association matrix of a series of ordered actions to be learned initialized in a preferred embodiment of the present invention;
[0059] Figure 4 is a schematic diagram of an action association matrix after a learning phase of a preferred embodiment of the present invention;
[0060] Figure 5 It is a schematic diagram of the final weight association matrix after sleep stage enhancement inference of a preferred embodiment of the present invention;
[0061] Among them, the darker the color, the higher the correlation. DETAILED DESCRIPTION
[0062] The following describes several preferred embodiments of the present invention with reference to the drawings in the specification, so that the technical content is clearer and easier to understand. The present invention can be embodied in many different forms of embodiments, and the protection scope of the present invention is not limited to the embodiments mentioned in the text.
[0063] In the drawings, components with the same structure are indicated by the same numerical reference numerals, and components with similar structures or functions are indicated by similar numerical reference numerals. The size and thickness of each component shown in the drawings are arbitrarily shown, and the present invention does not limit the size and thickness of each component. In order to make the illustration clearer, the thickness of the components is appropriately exaggerated in some places in the drawings.
[0064] like Figure 1 As shown, the present invention provides a Hebbian learning and sleep enhancement strategy based on a brain-like intelligent agent, comprising the following steps:
[0065] (1) Obtain the time step sequence and action sequence through the time series input module;
[0066] (2) Initialize the association matrix and generate the initial association between time steps and actions based on regular weight assignment and random noise;
[0067] (3) Dynamically adjust the weights of the association matrix based on the Hebbian learning rule;
[0068] (4) Simulate the spontaneous activity of neurons during sleep, strengthen high-confidence associations and weaken low-confidence associations through the positive and negative sample adjustment mechanism, and perform sparse optimization on the matrix weights;
[0069] (5) Based on the optimized association matrix, the optimal action sequence corresponding to the inferred time step is estimated using the maximum a posteriori probability.
[0070] In a preferred embodiment of the present invention, step (2) also includes normalizing the matrix.
[0071] In another preferred embodiment of the present invention, the initialization of the correlation matrix in step (2) comprises the following steps:
[0072] (2.1) Divide the time step into at least two regular intervals, each of which is assigned a strong initial weight to a specific action group;
[0073] (2.2) Based on the regular distribution, random noise weights are introduced for the remaining actions;
[0074] (2.3) Each row of the matrix is normalized so that its weight sum is 1.
[0075] In another preferred embodiment of the present invention, the dynamic adjustment of the association matrix weights based on the Hebbian learning rule in step (3) includes strengthening the significant association between the current time step and the action, and performing recursive weight updates in combination with the previous and next time steps.
[0076] In another preferred embodiment of the present invention, the weight adjustment of the Hebbian learning rule in step (3) includes:
[0077] (3.1) For the current time step l i and action a i Increase its weight:
[0078] P[l i ,a i ]←P[l i ,a i ]+η
[0079] (3.2)3.2 If there is a previous time step l i-1 and action a i-1 , then recursively update its weight:
[0080]
[0081] (3.3) The rows corresponding to the updated time steps are normalized to ensure the numerical stability of the matrix.
[0082] In another preferred embodiment of the present invention, in step (3), the association weights of time steps and actions are adjusted by the Hebbian learning rule in the learning phase, and the specific rules include:
[0083] Co-activation enhancement rule: time l i and action a i When activated, its associated weight increases;
[0084] History association recursive rule: If time step l i-1 and action a i-1 is the predecessor of the current time step, its weight is recursively enhanced;
[0085] The adjusted correlation matrix is normalized to maintain the numerical stability of the weights. The learning phase action and time step correlation matrix learning algorithm flow chart based on Hebbian learning Figure 2 shown.
[0086] In another preferred embodiment of the present invention, in step (4), the association matrix is optimized by the timing-dependent synaptic plasticity mechanism during the sleep enhancement stage, and the specific steps include the following:
[0087] (4.1) During the sleep enhancement phase, the spontaneous activity of neurons is simulated by randomly selecting time steps and actions;
[0088] (4.2) Positive sample enhancement: If the association probability between a time step and an action is higher than a set threshold, its association weight is enhanced;
[0089] (4.3) Negative sample weakening: If the association probability between a time step and an action is lower than a set threshold, its association weight is weakened;
[0090] (4.4) Sparse processing: weights below the set threshold are clipped to zero;
[0091] (4.5) Weight normalization: Renormalize the matrix to maintain numerical balance and sparsity.
[0092] In another preferred embodiment of the present invention, when inferring the optimal action sequence corresponding to the time step in step (5), a maximum a posteriori probability estimation method is used, which specifically includes:
[0093] (5.1) For each time step, the optimal action is selected according to the following formula
[0094] (5.2) Output the complete inferred action sequence
[0095] In another preferred embodiment of the present invention, the strategy further comprises step (6): displaying the changing process of the weight distribution of the correlation matrix in the learning and sleep enhancement stages through a dynamic visualization module.
[0096] In another preferred embodiment of the present invention, in step (6), the dynamic visualization module is used to display the weight changes of the correlation matrix at different stages, including
[0097] (6.1) Matrix distribution after strengthening significant associations in the learning phase;
[0098] (6.2) Matrix distribution after sparse optimization in the sleep enhancement phase.
[0099] Example 1: Brain-like time series learning system
[0100] 1. System Overview
[0101] The brain-like time series learning system of the present invention includes the following key stages: association matrix initialization, time series input, learning stage, sleep enhancement stage, inference and evaluation, and dynamic visualization stage. These stages work together through built-in parameter adjustment and data transmission mechanisms to form a complete time series learning and memory system.
[0102] 2. Initialization of functional partitions
[0103] Step 2.1 Initialize the association matrix. The association matrix is used to represent the association weights between time steps and actions. It has a dimension of n and is weighted according to regularity. Assume that the first k1 time steps are associated with actions {A, B, C, D, E}; the middle k2 time steps are strongly associated with actions {F, G, H, I, J}; and the last k3 time steps are strongly associated with actions {K, L, M, N, O}. Based on the above, random noise weights are introduced to ensure the diversity and exploration ability of the association matrix.
[0104] Step 2.2 Load the time series input information. Time series L = {l1,l2,...,l n} and the corresponding action sequence A={a1,a2,...,a n} is used as the input of the system for training and testing in the learning and inference phases.
[0105] 3. Learning phase
[0106] Step 3.1: Weight update based on Hebbian learning rule. Hebbian learning is a learning rule based on biological neuroscience, and its core idea is that "synchronous activation of neurons will enhance the strength of connections between them". In the present invention, Hebbian learning is used to dynamically adjust the weights of the association matrix between time steps and actions, and its rules specifically include:
[0107] Co-activation enhancement rule: For each time step l of the time series L and the action sequence A i and action a i, first enhance the current time step l i and action a i P[l i ,a i ]←P[l i ,a i ]+η
[0108] History association recursive rule: If there is a previous time step l i-1 and action a i-1 , then recursively update its associated weights:
[0109] Weight normalization: After each adjustment, i and l i-1 The weights are normalized to maintain numerical stability and biological plausibility.
[0110] The introduction of the Hebbian learning rule enables the system to effectively capture the temporal dependencies between time steps and simulate the memory and learning characteristics of the biological nervous system.
[0111] Step 3.2: Multiple rounds of iterative training. Through multiple rounds of iterations, the system strengthens the significant correlation between time steps and actions, and gradually forms a regular weight distribution.
[0112] 4. Sleep Stages
[0113] Step 4.1 Spontaneous activity and replay, initially the system randomly selects time step l i and action a i , simulating the spontaneous activity of neurons. This is a kind of timing-dependent synaptic plasticity (STDP), which is a learning mechanism based on time series, aiming to adjust the synaptic weight according to the time difference of activation between neurons. The present invention introduces the STDP mechanism in the sleep stage to optimize the sparsity and significance of the association matrix. Its rules include:
[0114] Positive sample enhancement: If time step l i and action a i If the probability of the association weight of is high (exceeding the set threshold), its association weight is enhanced: P[l i ,a i ]←P[l i ,a i ]+η.
[0115] Negative sample weakening: If the time step l i and action a i The probability of the associated weight of is low, then its associated weight is weakened: P[l i ,a i ]←P[l i ,a i ]-η.
[0116] Step 4.2 Sparsification and normalization. This is also an important step in implementing the STDP mechanism. Sparsify the association matrix by removing weights below the set threshold: P[l i ,a j = 0 if P[l i ,a j < threshold to optimize the sparsity of the matrix. Then renormalize each row of the matrix to ensure the numerical balance of the weights. The mechanism of the present invention can strengthen significant associations and reduce low-correlation weights during the sleep phase, thereby optimizing the memory retention ability.
[0117] Step 4.3 Through multiple rounds of iteration, strengthen high-confidence associations and weaken low-confidence associations to further consolidate memories.
[0118] 5. Inference and evaluation
[0119] Step 5.1 Inference based on maximum a posteriori probability. For each time step l i , select the action corresponding to the maximum a posteriori probability according to the association matrix P: Output the inferred action sequence
[0120] Step 5.2 Confidence analysis of the inference results. Calculate the confidence for the inferred action at each time step and output it.
[0121] Step 5.3 Inference accuracy evaluation: Compare the inferred sequence A* and the original sequence A, and calculate the inference accuracy:
[0122]
[0123] 6. Dynamic visualization analysis
[0124] Step 6.1 Dynamic display of the association matrix: Visualize the association matrix after the learning phase and the sleep phase respectively to show the changes in the weight distribution, as shown in Appendix Figure 3 、Appendix Figure 4 ,Appendix Figure 5 shown. As can be seen from Appendix Figure 3 , we initialized some regular action sequences, but they were not obvious in the noise background; Figure 4 is after the learning phase, and the learned action weights are clearly adjusted based on Hebbian learning; Figure 5 is after the simulated sleep phase, and the weights of the learned action sequence are enhanced.
[0125] Step 6.2 Comparison of regularity and sparsification effect: Demonstrate significant regularity after learning and weight sparsification effect after sleep enhancement. Initialize the initial test information with random noise, the sparsity is close to 0, and the sparsity after one round of Hebbian learning is about 1% (less than 0.001 is considered to be 0 sparsity), and after 100 sleep iterations, the sparsity can reach 90.56%. Figure 5 The image after 100 sleep-activation iterations is significantly better than Figure 4 Those who have only studied with Herb are much cleaner.
[0126] The preferred specific embodiments of the present invention are described in detail above. It should be understood that ordinary technicians in the field can make many modifications and changes based on the concept of the present invention without creative work. Therefore, all technical solutions that can be obtained by technicians in the technical field based on the concept of the present invention through logical analysis, reasoning or limited experiments on the basis of the prior art should be within the scope of protection determined by the claims.
Claims
1. A Hebbian learning and sleep enhancement strategy based on brain-like intelligent agents, characterized in that: The steps include: (1) Obtain the time step sequence and action sequence through the time series input module; (2) Initialize the association matrix and generate the initial association between time steps and actions based on regular weight assignment and random noise; (3) Dynamically adjust the weights of the association matrix based on the Hebbian learning rule; (4) Simulate the spontaneous activity of neurons during sleep, strengthen high-confidence associations and weaken low-confidence associations through the positive and negative sample adjustment mechanism, and perform sparse optimization on the matrix weights; (5) Based on the optimized association matrix, the optimal action sequence corresponding to the inferred time step is estimated using the maximum a posteriori probability.
2. The Hebbian learning and sleep enhancement strategy based on brain-like agents as claimed in claim 1, characterized in that: Step (2) also includes normalizing the matrix.
3. The Hebbian learning and sleep enhancement strategy based on brain-like agents as claimed in claim 1, characterized in that: The initialization of the association matrix in step (2) comprises the following steps: (2.1) Divide the time step into at least two regular intervals, each of which is assigned a strong initial weight to a specific action group; (2.2) Based on the regular distribution, random noise weights are introduced for the remaining actions; (2.3) Each row of the matrix is normalized so that its weight sum is 1.
4. The Hebbian learning and sleep enhancement strategy based on brain-like agents as claimed in claim 1, characterized in that: The step (3) dynamically adjusts the association matrix weights based on the Hebbian learning rule, including strengthening the significant association between the current time step and the action, and recursively updating the weights in combination with the previous and next time steps.
5. The Hebbian learning and sleep enhancement strategy based on brain-like agents as claimed in claim 1, characterized in that: The weight adjustment of the Hebbian learning rule in step (3) includes: (3.1) For the current time step l i and action a i Increase its weight: P[l i ,to i ]←P[l i ,to i ]+η (3.2)3.2 If there is a previous time step l i-1 and action a i-1 , then recursively update its weight: (3.3) The rows corresponding to the updated time steps are normalized to ensure the numerical stability of the matrix.
6. The Hebbian learning and sleep enhancement strategy based on brain-like agents as claimed in claim 1, characterized in that: In step (3), the association weights between time steps and actions are adjusted by the Hebbian learning rule during the learning phase. The specific rules include: Co-activation enhancement rule: time l i and action a i When activated, its associated weight increases; History association recursive rule: If time step l i-1 and action a i-1 is the predecessor of the current time step, its weight is recursively enhanced; The adjusted correlation matrix is normalized to maintain the numerical stability of the weights.
7. The Hebbian learning and sleep enhancement strategy based on brain-like agents as claimed in claim 1, characterized in that: In step (4), the association matrix is optimized through the timing-dependent synaptic plasticity mechanism during the sleep enhancement phase. The specific steps include the following: (4.1) During the sleep enhancement phase, the spontaneous activity of neurons is simulated by randomly selecting time steps and actions; (4.2) Positive sample enhancement: If the association probability between a time step and an action is higher than a set threshold, its association weight is enhanced; (4.3) Negative sample weakening: If the association probability between a time step and an action is lower than a set threshold, its association weight is weakened; (4.4) Sparse processing: weights below the set threshold are clipped to zero; (4.5) Weight normalization: Renormalize the matrix to maintain numerical balance and sparsity.
8. The Hebbian learning and sleep enhancement strategy based on brain-like agents as claimed in claim 1, characterized in that: When inferring the optimal action sequence corresponding to the time step in step (5), the maximum a posteriori probability estimation method is used, which specifically includes: (5.1) For each time step, the optimal action is selected according to the following formula (5.2) Output the complete inferred action sequence 9. The Hebbian learning and sleep enhancement strategy based on brain-like agents as claimed in claim 1, characterized in that: The strategy also includes step (6): displaying the changing process of the weight distribution of the association matrix during the learning and sleep enhancement stages through a dynamic visualization module.
10. The Hebbian learning and sleep enhancement strategy based on brain-like intelligent agent as claimed in claim 1, characterized in that: In step (6), the dynamic visualization module is used to display the weight changes of the correlation matrix at different stages, including (6.1) Matrix distribution after strengthening significant associations in the learning phase; (6.2) Matrix distribution after sparse optimization in the sleep enhancement phase.