A Brainwave Induction Method and Storage Medium Based on Deep Reinforcement Learning
Through the brain wave induction method based on deep reinforcement learning, the stimulation signal is automatically selected and adjusted, and the existing methods rely on expert knowledge and high cost are solved, and efficient and automated brain wave state induction is achieved.
Patent Information
- Application Number
- CN202211149229.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-09-21
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2042-09-21
AI Technical Summary
The existing brain wave induction methods mainly rely on expert knowledge, are costly, and require a lot of manpower to select and control stimulation signals.
The brain wave induction method based on deep reinforcement learning is adopted, and the stimulation signals (such as pictures) and their sequence are automatically selected and adjusted through the pre-trained reinforcement learning model module to achieve brain wave state induction.
This method can automatically select the most appropriate stimulation signal and its sequence based on brain wave data, reducing dependence on expert knowledge, reducing costs, and improving the efficiency of brain wave induction.
Smart Images

Figure HDA0003856214390000011 
Figure HDA0003856214390000021
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of brain wave data processing, and particularly relates to a brain wave induction method and a storage medium based on deep reinforcement learning. Background Art
[0002] Brain waves (English: brainwave) refer to the electrical oscillations generated when nerve cells in the human brain are active. Since this kind of oscillation appears on scientific instruments and looks like a wave, it is called a brain wave. To put it in one sentence, perhaps it can be said that it is the bioenergy generated by brain cells, or the rhythm of brain cell activities. Every second, no matter what a person is doing, even when sleeping, the brain will constantly generate "brain waves" like "electrical pulses". Brain waves can be divided into five categories according to frequency: β waves (consciousness 14 - 30HZ), α waves (bridging consciousness 8 - 14HZ), θ waves (subconsciousness 4 - 8Hz), δ waves (unconsciousness below 4Hz), and γ waves (focusing on something above 30HZ), etc. The combination of these consciousnesses forms a person's internal and external behaviors, emotions, and learning performances. β brain waves are a kind of conscious brain waves that operate at a frequency of 13 - 25 cycles per second. When people are in a state of being awake, concentrated, alert, or when thinking, analyzing, speaking, and actively acting, the brain will emit this kind of brain wave.
[0003] Brainwave entrainment, also called brain loading, is to change the brain wave from one mode to another through external stimuli, thereby intervening in people's emotions. Common brainwave entrainments include binaural beats, monaural beats, isochronic tones, spectral induction, etc. The first three are also called acoustic inductions and are the most commonly used. With the gradual maturity of brain wave acquisition devices, the research on brainwave entrainment is more extensive.
[0004] Currently, most of the methods for implementing brainwave entrainment are to study the relationship between the change of the stimulation signal and the change of the brain wave state through expert laboratory research, so as to determine the induction method.
[0005] Reinforcement learning is an algorithm for studying learning control strategies that has been studied for a long time. Its basic idea is to maximize the cumulative reward according to a certain policy combination under the given external environment, internal state, and policy set. Q-Learing is a relatively common form of reinforcement learning. This technology does not require a model and can learn the optimal policy value function, where the policy value function represents the long-term reward of the policy.
[0006] Currently, brainwave entrainment methods mainly rely on expert knowledge to select corresponding stimulation signals and stimulation times, and also need to control the stimulation duration, with a relatively high cost. Summary of the Invention
[0007] The present invention proposes a brainwave induction method based on deep reinforcement learning, which can solve at least one of the above-mentioned technical problems.
[0008] To achieve the above object, the present invention adopts the following technical solutions:
[0009] A brainwave induction method based on deep reinforcement learning comprises the following steps:
[0010] S1. Collect user brain wave data;
[0011] S2, input the user's brain wave data into the pre-trained reinforcement learning model module, and output the sequence number of the picture to be played;
[0012] S3, the picture playing module selects the corresponding picture to play according to the picture sequence number;
[0013] S4, after the picture is played for the set time, repeat steps S1 to S3.
[0014] Furthermore, the training process of the reinforcement learning model module is as follows:
[0015] (a) Preliminarily collect pictures that can induce brain waves to reach a fixed state based on specified knowledge, number them 1 to N, and form a picture set;
[0016] (b) Find volunteers and let them wear brain wave acquisition equipment in a specific environment. First, brain waves W1 of L duration are acquired, and then pictures are played. For a volunteer, the first picture played is randomly selected from the picture collection and recorded as picture n1. The volunteer watches the played picture n1, and the brain wave acquisition equipment continues to collect data. After watching for L time, brain wave data W2 is collected. At this time, the neural network defined later starts to process brain wave data W1 and brain wave data W2, and finally the neural network outputs a picture number n2. The program retrieves the picture collection according to the number and switches the played picture to picture n2. After watching for L seconds, the brain wave data collected during this period is recorded as W3, and the neural network starts to process brain wave data W2 and brain wave data W3. Finally, the neural network outputs a new picture number n3, and so on. Each time, the neural network processes the previous brain wave data and the brain wave data when watching the current picture, and then gives a new picture number.
[0017] Furthermore, the structure of the reinforcement learning model module includes:
[0018] BERT-like encoding model: This model imitates the training method of the BERT model in natural language processing. It trains a BERT-like pre-training model based on brain wave data and is built into the network structure here to encode brain wave data.
[0019] Encoder Block: The Encoder Block refers to a structure implemented based on the self-attention mechanism, which is equivalent to a module stacking N self-attention structures and feed-forward networks. N can be adjusted during the actual training process;
[0020] Fully connected: The fully connected is the implementation of the fully connected in traditional neural networks, in order to map the features output by the Encoder Block to the picture number.
[0021] Furthermore, the internal execution process of the structure of the reinforcement learning model module is as follows:
[0022] Assume that taking the α wave accounting for the main state in the induced brain wave as the training target of the final model, the data input during a certain network execution is: brain wave data W1, picture data n1, brain wave data W2;
[0023] (a), If it is the first training, initialize two network objects denoted as Net1 and Net2 using the above network structure; where Net1 is used to process the brain wave data W1, and its output result will be used to construct the subsequent loss function and calculate the picture number, and the output result of Net2 is used to assist in constructing the loss function used later; If it is not the first training, judge whether it has been trained for C rounds. If so, directly use Net1 to replace Net2, otherwise do not replace;
[0024] (b), Use Net1 to process the brain wave data W1 to obtain the output output1, and obtain the next picture number n2 according to output1;
[0025] (c), Use Net2 to process the brain wave data W1 to obtain the output output2;
[0026] (d), Analyze the brain wave data W2, calculate the proportion of the α wave in it, denoted as p1;
[0027] (e), Construct the loss function: loss = -(p1 + γ * output2 - output1);
[0028] (f), Perform backpropagation on loss to update the parameters of Net1;
[0029] (g), Loop and execute steps a to f.
[0030] On the other hand, a computer-readable storage medium of the present invention stores a computer program. When the computer program is executed by a processor, the processor is caused to execute the steps of the above method.
[0031] As can be seen from the above technical solutions, the brain wave induction method based on deep reinforcement learning of the present invention can automatically select the most appropriate stimulation signal and its order according to the set of stimulation signals (here, pictures are used), avoiding over-reliance on expert knowledge to a certain extent and reducing costs. BRIEF DESCRIPTION OF THE DRAWINGS
[0032] Figure 1 is a schematic diagram of the method of the present invention;
[0033] Figure 2 is a schematic structural diagram of an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0034] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Apparently, the described embodiments are some, but not all, of the embodiments of the present invention.
[0035] As Figure 1 shown, the brain wave induction method based on deep reinforcement learning described in this embodiment is based on a reinforcement learning model module. As Figure 1 shown, its main function is to use a pre-trained deep reinforcement learning model to obtain the serial number of the next picture to be played, and then the picture playback module selects the corresponding picture for playback according to the picture serial number.
[0036] The working process is as follows:
[0037] S1. Collect the brain wave data of the user (generally collect for 10 seconds);
[0038] S2. Input the brain wave data of the user into the reinforcement learning model module and output the serial number of the picture to be played;
[0039] S3. The picture playback module selects the corresponding picture for playback according to the picture serial number;
[0040] S4. After the picture is played for a certain duration (10 seconds), repeat steps S1 to S3.
[0041] The training process of the reinforcement learning model used is as follows:
[0042] (1). Training form
[0043] (a). Collect pictures that can induce brain waves to reach a fixed state (such as α waves, δ waves, etc.) according to relevant knowledge in advance, numbered from 1 to N, to form a picture set.
[0044] (b) Find volunteers and let them wear brain wave acquisition equipment in a specific environment. First, collect brain waves W1 for L (assuming 20 seconds), and then start playing pictures. For a volunteer, the first picture played is randomly selected from the picture collection and recorded as picture n1. The volunteer watches the played picture n1, and the brain wave acquisition device continues to collect. After watching for L time, brain wave data W2 is collected. At this time, the neural network defined later begins to process brain wave data W1 and brain wave data W2, and finally the neural network outputs a picture number n2. The program retrieves the picture collection according to the number and switches the played picture to picture n2. After watching for L seconds, the brain wave data collected during this period is recorded as W3, and the neural network begins to process brain wave data W2 and brain wave data W3. Finally, the neural network outputs a new picture number n3, and so on. Each time the neural network processes the previous brain wave data and the brain wave data when watching the current picture, and then gives a new picture number. Generally, a volunteer can participate for P (minutes or hours, according to the volunteer's wishes) according to the actual situation. L and P can be adjusted according to actual conditions.
[0045] (2) Model structure Figure 2 As shown;
[0046] The reinforcement learning model structure here is mainly to learn the strategy of selecting pictures. The network output result is generally the picture number in the picture collection. Its structure and implementation method are as follows:
[0047] (a) BERT-like encoding model: This model imitates the training method of the BERT (Bidirectional Encoder Representation from Transformers) model in natural language processing. It trains a BERT-like pre-trained model based on brain wave data and is built into the network structure here to encode brain wave data. It can also be replaced with other encoding methods.
[0048] (b) Encoder Block: Encoder Block refers to a structure implemented based on the self-attention mechanism, which can be seen as a module that stacks N self-attention structures and feedforward networks (refer to the implementation of the Encoder structure in Transformer in Pytorch, https: / / pytorch.org / docs / stable / generated / torch.nn.TransformerEncoderLayer.html#torch.nn.TransformerEncoderLayer). The default setting N=6 is generally used, which can be adjusted during the actual training process.
[0049] (c), Fully connected: The fully connected here is the implementation of the fully connected in the traditional neural network, mainly to map the features output by the Encoder Block to the picture number.
[0050] (3) Internal execution process during network training
[0051] Suppose we take the state where the alpha wave dominates in the induced brain waves as the training target of the final model. The data input during a certain network execution is: brain wave data W1, picture data n1, and brain wave data W2.
[0052] (a), If it is the first training, initialize two network objects using the above network structure and denote them as Net1 and Net2; where Net1 is used to process the brain wave data W1, and its output result will be used to construct the subsequent loss function and calculate the picture number. The output result of Net2 is used to assist in constructing the loss function used later. If it is not the first training, determine whether it has been trained for C rounds. If so, directly replace Net2 with Net1, otherwise do not replace.
[0053] (b), Use Net1 to process the brain wave data W1 to obtain the output output1, and obtain the next picture number n2 according to output1.
[0054] (c), Use Net2 to process the brain wave data W1 to obtain the output output2.
[0055] (d), Analyze the brain wave data W2 and calculate the proportion of the alpha wave in it, denoted as p1.
[0056] (e), Construct the loss function: loss = -(p1 + γ * output2 - output1)
[0057] (f), Perform backpropagation on the loss and update the parameters of Net1
[0058] (g), Loop through a to f.
[0059] Train the model for a certain period according to the above process, and the model can then be used.
[0060] The BERT-like encoding model of the embodiment of the present invention can be replaced with other encoding methods.
[0061] In summary, the present invention provides a method for inducing brain wave states based on deep reinforcement learning, which can automatically select the most appropriate stimulus signal and its order according to the set of stimulus signals (here pictures are used), avoiding over-reliance on expert knowledge to a certain extent and reducing costs.
[0062] In another aspect, the present invention also discloses a computer-readable storage medium storing a computer program, which when executed by a processor causes the processor to execute the steps of any of the above methods.
[0063] In yet another aspect, the present invention also discloses a computer device including a memory and a processor, the memory storing a computer program, which when executed by the processor causes the processor to execute the steps of any of the above methods.
[0064] In yet another embodiment provided by the present application, there is also provided a computer program product containing instructions, which when running on a computer causes the computer to execute the steps of any of the methods in the above embodiments.
[0065] It can be understood that the system provided by the embodiments of the present invention corresponds to the method provided by the embodiments of the present invention. For the explanations, examples and beneficial effects of the relevant content, reference can be made to the corresponding parts in the above methods.
[0066] Those of ordinary skill in the art can understand that all or part of the processes of implementing the methods in the above embodiments can be completed by instructing relevant hardware through a computer program. The program can be stored in a non-volatile computer-readable storage medium. When the program is executed, it may include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database or other medium used in the various embodiments provided in the present application may include non-volatile and / or volatile memories. Non-volatile memories may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memories may include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and Rambus dynamic RAM (RDRAM), etc.
[0067] The technical features of the above embodiments can be combined arbitrarily. For the sake of concise description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope recorded in this specification.
[0068] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the various embodiments of the present invention.
Claims
1. A brain wave induction method based on deep reinforcement learning, characterized in that, The following steps are included: S1. Collect user brain wave data; S2, input the user's brain wave data into the pre-trained reinforcement learning model module, and output the sequence number of the picture to be played; The structure of the reinforcement learning model module includes: BERT-like encoding model: This model imitates the training method of the BERT model in natural language processing. It trains a BERT-like pre-training model based on brain wave data and is built into the network structure here to encode brain wave data. Encoder Block: Encoder Block refers to a structure based on the self-attention mechanism, which is equivalent to stacking N modules of self-attention structures and feedforward networks. N can be adjusted during the actual training process. Full connection: Full connection is the implementation of full connection in traditional neural networks, in order to map the features output by the Encoder Block to the image number; S3, the picture playing module selects the corresponding picture to play according to the picture sequence number; S4, after the picture is played for the set time, repeat steps S1 to S3.
2. The brainwave induction method based on deep reinforcement learning according to claim 1, characterized in that: The reinforcement learning model module training process is as follows: (a) Collect pictures that can induce brain waves to reach a fixed state based on specified knowledge in advance, number them 1 to N, and form a picture set; (b) Find volunteers and let them wear brain wave acquisition equipment in a specific environment. First, brain waves W1 of L duration are acquired, and then pictures are played. For a volunteer, the first picture played is randomly selected from the picture collection and recorded as picture n1. The volunteer watches the played picture n1, and the brain wave acquisition equipment continues to collect data. After watching for L time, brain wave data W2 is collected. At this time, the neural network defined later starts to process brain wave data W1 and brain wave data W2, and finally the neural network outputs a picture number n2. The program retrieves the picture collection according to the number and switches the played picture to picture n2. After watching for L seconds, the brain wave data collected during this period is recorded as W3, and the neural network starts to process brain wave data W2 and brain wave data W3. Finally, the neural network outputs a new picture number n3, and so on. Each time, the neural network processes the previous brain wave data and the brain wave data when watching the current picture, and then gives a new picture number.
3. The brainwave induction method based on deep reinforcement learning according to claim 1, characterized in that: The internal structure execution process of the reinforcement learning model module is as follows: Assume that the training target of the final model is to induce the dominant state of alpha waves in brain waves. The input data during a network execution is: brain wave data W1, image data n1, brain wave data W2; (a) If it is the first training, use the above network structure to initialize two network objects, Net1 and Net2; Net1 is used to process the brain wave data W1, and its output result will be used to construct the subsequent loss function and calculate the image number, and the output result of Net2 is used to assist in constructing the subsequent loss function; if it is not the first training, determine whether it has been trained for C rounds. If so, directly use Net1 to replace Net2, otherwise do not replace it; (b), Process the brain wave data W1 using Net1 to obtain the output output1, and obtain the next picture number n2 according to output1; (c), Process the brain wave data W1 using Net2 to obtain the output output2; (d), Analyze the brain wave data W2, calculate the proportion of α waves therein, denoted as p1; (e), Construct a loss function: loss = -(p1 + γ * output2 - output1); (f), Perform backpropagation on loss to update the parameters of Net1; (g), Loop through steps a to f.
4. A computer-readable storage medium storing a computer program, which when executed by a processor causes the processor to execute the steps of the method according to any one of claims 1 to 3.
Citation Information
Patent Citations
System and method for potentiating effective brainwave by controling volume of sound
CN104254358A