Intelligent hypnosis method and system based on brain-computer interface and storage medium

Through brain-computer interface and deep reinforcement learning model, the music playback strategy is adjusted in real time, which solves the problems of personalization and inefficiency of existing hypnosis methods and realizes a personalized and efficient intelligent hypnosis experience.

CN120733205APending Publication Date: 2025-10-03SHENZHEN KUKAI BRAIN MACHINE INTELLIGENT TECHNOLOGY CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510868421.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-26
Publication Date
2025-10-03

AI Technical Summary

Technical Problem

Existing hypnosis methods lack personalization, are inefficient, and lack flexibility. They cannot be dynamically adjusted according to the user's real-time EEG state, resulting in low hypnosis efficiency and high service costs.

Method used

By adopting brain-computer interface technology, collecting the user's EEG signals in real time, using deep reinforcement learning model to dynamically adjust the music playback strategy, combining ∈-greedy strategy and Q-Learning algorithm to optimize the model, personalized, closed-loop hypnosis intervention is achieved.

Benefits of technology

It achieves a highly personalized hypnosis experience, significantly improves hypnosis efficiency, and can guide users from a wakeful state to a sleep state in a short period of time. The system also has self-optimization capabilities, and the hypnosis effect continues to improve with increasing usage.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120733205A_ABST
    Figure CN120733205A_ABST
Patent Text Reader

Abstract

The invention relates to an intelligent hypnosis method and system based on a brain-computer interface and a storage medium, and belongs to the technical field of human-computer interaction and artificial intelligence. The method comprises the following steps: acquiring an EEG signal of a user in real time through EEG acquisition equipment, and extracting a brain wave segment power value as state input through preprocessing; a deep reinforcement learning model (such as DRQN) selects music type actions (such as classical music and white noise) based on an epsilon-greedy strategy; calculating a reward value (maximizing delta wave increment and inhibiting beta wave) according to the electroencephalogram state change after playing, and optimizing model parameters by adopting Q-Learning; and dynamically adjusting the strategy through iterative interaction until the user reaches a preset sleep target. The system comprises an electroencephalogram acquisition module, a preprocessing module, a reinforcement learning module and a music control module, and realizes closed-loop regulation and control. The method has the advantages of high personalization, high hypnosis efficiency (induction to sleep in 1-7 minutes), flexible adaptation to different users and self-evolution optimization capability, and effectively improves the sleep induction effect.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of human-computer interaction technology and artificial intelligence technology, and in particular to an intelligent hypnosis method, system and storage medium based on a brain-computer interface. Background Art

[0002] Sleep is a fundamental physiological process essential for maintaining physical and mental health. However, with the accelerating pace of modern life and increasing work pressure, sleep problems such as insomnia and difficulty falling asleep are becoming increasingly common, seriously impacting people's quality of life and health. Hypnosis, as an effective aid in guiding individuals from wakefulness to sleep, has significant application value in the field of sleep health.

[0003] At present, the mainstream hypnosis methods mainly include artificial hypnosis, drug hypnosis, and fixed-pattern music or sound hypnosis.

[0004] Artificial hypnosis: Professional hypnotists use verbal guidance and psychological suggestion to help users relax and fall asleep. This method offers advantages such as high personalization and good results. However, its disadvantages are also significant: it relies heavily on the hypnotist's professional skills and clinical experience, and the scarcity of qualified hypnotists leads to high service costs, making it difficult to promote and popularize on a large scale.

[0005] Drug-induced hypnosis: Inducing sleep through the use of chemical drugs such as sleeping pills. This method is quick to work, but it often carries side effects, such as next-day sleepiness, dizziness, and decreased cognitive function. Long-term use can also lead to drug dependence, tolerance, and withdrawal symptoms, posing a potential threat to human health.

[0006] Fixed-pattern music or sound hypnosis: Creating a relaxing sleep environment by playing preset classical music, white noise, nature sounds, and other audio. This is currently the most common and easiest way to assist sleep. However, this method has fundamental flaws:

[0007] Lack of personalization: It uses a one-size-fits-all approach that fails to account for individual physiological differences. Different people have distinct preferences and physiological responses to sounds. The same piece of music may be a lullaby for one person, but ineffective or even irritating for another. Furthermore, even for the same person, their EEG state changes dynamically over time and under different emotional states, and a fixed music playlist cannot adapt to these changes.

[0008] Inefficiency: Due to the lack of a dynamic adjustment mechanism, this method cannot effectively intervene based on the user's real-time EEG state. It is essentially an open-loop system that only cares about "input" (playing sound) and not "output" (the user's actual response). As a result, users often take a long time to fall asleep, and even after repeated attempts, they still cannot successfully fall asleep, making hypnosis extremely inefficient.

[0009] Lack of flexibility: This method typically plays fixed audio clips in a preset order and duration, making it difficult to flexibly switch between different stimulation sources as needed during the hypnosis process. For example, one type of music may be needed to help with relaxation in the early stages of hypnosis, while another type of sound may be needed to deepen sleep as the hypnosis process nears. This fixed format fails to fully utilize the advantages of various hypnotic materials and is difficult to adapt to the dynamic needs of different hypnosis stages.

[0010] Therefore, developing an intelligent hypnosis method that can overcome the defects of existing technologies and achieve personalized, efficient and flexible results has important theoretical significance and broad market application prospects. Summary of the Invention

[0011] In view of the above situation, it is necessary to provide an intelligent hypnosis method based on brain-computer interface that solves at least one of the above problems, comprising the following steps:

[0012] a. Signal acquisition: Use EEG acquisition equipment to obtain the user's EEG signals (EEG) in real time during sleep;

[0013] b. State definition: Preprocess the collected EEG signals to extract the power values ​​of different EEG wave bands θ, α, β, and γ. This set of power values ​​is used as the state input of the deep reinforcement learning model.

[0014] c. Action selection: The deep reinforcement learning model uses the ∈-greedy strategy to output an action based on the current input state, which corresponds to selecting one of multiple preset music types to play;

[0015] d. Reward calculation: Based on the changes in the newly collected EEG signal state after the music is played, a reward value is calculated. This reward value is calculated to maximize the increase in theta wave power and suppress the power of alpha waves;

[0016] e. Model training: The state, action, reward, and next state are stored as experience data in the experience replay pool. When the data volume reaches a predetermined value, the Q-Learning algorithm is used to train and optimize the network parameters of the deep reinforcement learning model.

[0017] f. Iterative intervention: Repeat steps c to e, dynamically adjusting the music playback strategy through real-time interaction between the model and the user, until the user's EEG state reaches the preset sleep target, thereby guiding the user from the awake state to the sleep period.

[0018] Preferably, in the signal acquisition step, the EEG acquisition device used is a polysomnogram (PSG) device or a portable EEG acquisition device.

[0019] Preferably, the preprocessing in the state definition step includes: filtering and denoising the collected original EEG signal, segmenting it according to a fixed time window, calculating the power value of each band by fast Fourier transform (FFT) or wavelet transform, and finally performing normalization processing.

[0020] Preferably, the music types in the action selection step include at least one or more of classical music, natural white noise, electronic soothing music and mixed music.

[0021] Preferably, the specific function of the reward calculation is:

[0022] Among them, r t is the instant reward at the current moment, is the increment of theta wave power at the current moment, is the decrement of α wave power at the current moment (taken as the absolute value), and λ is the weight coefficient used to balance the two.

[0023] Preferably, the deep reinforcement learning model used in the model training is a deep Q network (DQN) or its variant, combined with a proximal policy optimization algorithm.

[0024] Preferably, the method further comprises:

[0025] Course learning strategy: In the early stages of training, set shorter sleep induction tasks to allow the model to learn basic hypnosis strategies; then gradually increase the time limit of the tasks to reduce the training difficulty and enable the model to better adapt to hypnosis tasks of different lengths.

[0026] Preferably, the method further comprises an attention mechanism: an attention layer is added to the deep reinforcement learning model to weight the input EEG power vector to enhance the model's sensitivity to changes in theta and alpha wave power, which are strongly correlated with hypnosis.

[0027] In the present invention, another solution is disclosed, an intelligent hypnosis system based on brain-computer interface, comprising:

[0028] EEG signal acquisition module, used to collect the user's EEG signals in real time;

[0029] A signal preprocessing module, connected to the EEG signal acquisition module, is used to process the acquired EEG signals and extract the power value of each band;

[0030] A deep reinforcement learning module, which contains a pre-trained deep reinforcement learning model that receives pre-processed power values ​​as state input and outputs a music genre selection action;

[0031] A music playback control module plays music clips of corresponding types according to the actions output by the deep reinforcement learning module;

[0032] Among them, the deep reinforcement learning module also calculates the reward value based on the changes in the EEG state fed back by the signal acquisition and preprocessing module after playing music, and continuously optimizes its own model.

[0033] In the present invention, another solution is disclosed, a computer-readable storage medium having a computer program stored thereon, wherein when the program is executed by a processor, the method according to any one of claims 1 to 8 is implemented.

[0034] Compared with the prior art, the present invention has the following significant beneficial effects:

[0035] Highly Personalized: This invention completely abandons the "one-size-fits-all" fixed hypnosis model. It uses a brain-computer interface to monitor each user's unique and dynamically changing EEG state in real time and bases its decisions on this. A deep reinforcement learning model learns each individual's unique physiological response patterns to different musical stimuli, thereby "tailoring" the optimal hypnosis strategy for the user, greatly enhancing the degree of personalization of hypnosis.

[0036] Significantly improved hypnosis efficiency: The present invention adopts a closed-loop, goal-oriented control mechanism. The core goal of the deep reinforcement learning algorithm is to maximize the cumulative reward, and the reward function is directly linked to the hypnosis efficiency (i.e., the speed at which the theta wave increases and the alpha wave decreases). This drives the model to actively and quickly find and execute the most effective hypnosis strategy, avoiding invalid or inefficient intervention. According to experiments, when the proportion of theta waves exceeds 50%, marking the completion of the hypnosis task, the present invention can efficiently guide the individual from the wakeful stage to the N1 stage in just 1-7 minutes. Compared with traditional methods that often take tens of minutes or even longer, the efficiency is greatly improved.

[0037] Robust Flexibility and Adaptability: The proposed action space encompasses a wide variety of music or sound sources, allowing the model to flexibly switch between different stimuli at any point during the hypnosis process, based on real-time EEG state changes. This dynamic adaptability leverages the strengths of various hypnotic materials to address the needs of different hypnosis stages (e.g., early relaxation, mid-stage induction, and late-stage deepening), thereby enhancing the flexibility and robustness of the entire hypnosis process.

[0038] Intelligence and Self-Evolution: This system is capable of learning and optimization. Through continuous interaction with users, the model accumulates experience, continuously iterating and optimizing its hypnosis strategies. This means the system "gets to know you better the more you use it," and its hypnosis effectiveness for a specific user will improve with repeated use. BRIEF DESCRIPTION OF THE DRAWINGS

[0039] Figure 1 Schematic diagram of the overall architecture of an intelligent hypnosis system provided by an embodiment of the present invention.

[0040] Figure 2 This is a detailed flow chart of an intelligent hypnosis method provided by an embodiment of the present invention.

[0041] Figure 3 2 is a schematic diagram of the structure of a deep recurrent Q network (DRQN) model in an embodiment of the present invention. DETAILED DESCRIPTION

[0042] In order to make the objectives, technical solutions, and advantages of the present invention more clearly understood, the following, in conjunction with the accompanying drawings and embodiments, further describes in detail the brain-computer interface-based intelligent hypnosis method, system, and storage medium of the present invention. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.

[0043] In the description of the present invention, unless otherwise specified, "plurality" means two or more; the terms "center", "longitudinal", "lateral", "upper", "lower", "left", "right", "inner", "outer", "front end", "rear end", "head", "tail", "vertical", "horizontal", "top", "bottom", "inner", "outer", etc., indicating the orientation or positional relationship, are based on the orientation or positional relationship shown in the accompanying drawings, and are only for the convenience of describing the present invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, be constructed and operate in a specific orientation, and therefore should not be understood as limiting the present invention. In addition, the terms "first", "second", "third", etc. are used for descriptive purposes only and should not be understood as indicating or implying relative importance.

[0044] In the description of this utility, it should be noted that, unless otherwise specified or limited, the terms "mounted," "connected," and "connected" should be understood in a broad sense. For example, they can refer to fixed connections, detachable connections, or integral connections; mechanical connections, electrical connections; direct connections, indirect connections through an intermediate medium, and internal connections between two components. Those skilled in the art will understand the specific meanings of the above terms in this utility based on specific circumstances. DETAILED DESCRIPTION

[0045] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.

[0046] Example

[0047] This embodiment describes in detail a deep reinforcement learning intelligent hypnosis method based on brain-computer interface and the implementation details of its system.

[0048] 1. Overall system architecture (refer to Figure 1 )

[0049] like Figure 1 As shown, the intelligent hypnosis system 100 of the present invention mainly includes: a user 101, an EEG signal acquisition module 102, a central processing unit 103 and a music playback control module 104.

[0050] User 101: An individual who needs hypnotic intervention.

[0051] EEG signal acquisition module 102: This is typically a non-invasive EEG acquisition device. This can be a professional, medical-grade polysomnography (PSG) device that acquires accurate signals with a high signal-to-noise ratio. Alternatively, it can be a consumer-grade portable EEG acquisition device, such as a headband from brands like Muse or MindWave, for convenient home use. In this embodiment, the Muse headband can be used. It has multiple dry electrodes (e.g., AF7, AF8, TP9, and TP10) located in the forehead area, effectively acquiring EEG signals from the prefrontal cortex, particularly signals from leads Fp1 and Fp2, which are closely associated with emotion and relaxation.

[0052] Central Processing Unit 103: This can be a personal computer, smartphone, or dedicated embedded device. It is the "brain" of the system, running the core algorithm of the present invention. It primarily consists of two submodules: a signal preprocessing module 103a and a deep reinforcement learning module 103b.

[0053] The music playing control module 104 is usually an audio playing device such as a speaker or earphones. It receives instructions from the central processing unit 103 and plays a specific type of music.

[0054] The entire system's workflow forms a closed loop: User 101 wears acquisition module 102, and EEG signals are sent in real time to the central processing unit 103. The signal preprocessing module 103a within the unit processes the signals and extracts features, forming a "state." Based on this "state," the deep reinforcement learning module 103b makes "action" decisions, instructing the music playback control module 104 to play the appropriate music. The music influences user 101, altering their brain state. This change is captured by acquisition module 102, forming a new "state," and the cycle repeats.

[0055] 2. Detailed process of intelligent hypnosis method (refer to Figure 2 )

[0056] Figure 2 The detailed steps of the method of the present invention are shown.

[0057] Step S201: Initialization

[0058] Before hypnosis begins, perform system initialization. This includes:

[0059] Randomly initialize the parameters θ of the deep reinforcement learning (DRQN) network.

[0060] Create a target network (Target Network) whose structure is exactly the same as the DRQN ​​network and set its parameters θ ― Copy to be the same as θ.

[0061] Initialize an empty experience replay pool (Replay Buffer), whose capacity can be set to, for example, 20,000 experiences.

[0062] Set hyperparameters such as learning rate, discount factor γ (such as 0.99), initial exploration rate ∈ of the ∈-greedy strategy start (such as 0.9) and the final exploration rate ∈ end (such as 0.1) etc.

[0063] Step S202: Signal acquisition and preprocessing (state definition)

[0064] At each time step (for example, every 30 seconds is a time step), do the following:

[0065] Collection: The EEG signal collection module 102 collects the user's raw EEG data in the past 30 seconds.

[0066] Preprocessing: The signal preprocessing module 103a processes the raw data.

[0067] Filtering: Apply a bandpass filter (e.g., 1-50 Hz) to remove DC drift and high-frequency noise, and a notch filter (e.g., 50 Hz or 60 Hz) to remove power frequency interference.

[0068] Segmentation: The continuous EEG signal is divided into 30-second windows.

[0069] Feature extraction: For the signal in each window, fast Fourier transform (FFT) or wavelet transform is used to calculate the absolute power values ​​of four key EEG wave bands: delta wave (1-4 Hz), theta wave (4-8 Hz, associated with drowsiness and light sleep), alpha wave (8-13 Hz, associated with a quiet, relaxed wakeful state), beta wave (13-30 Hz, associated with alertness and excitement), and gamma wave (>30 Hz).

[0070] Normalization: To eliminate individual differences and dimensionality, the extracted power values ​​for each band are z-score normalized. Normalization uses the mean μ and standard deviation σ calculated from a baseline period collected before hypnosis (e.g., the user sits quietly with their eyes closed for 2 minutes). The normalization formula is x' = (x - μ) / σ.

[0071] State construction: Combine the normalized power values ​​of each band (such as θ, α, β, γ) into a state vector s t If there are multiple leads, the feature vectors of all leads are concatenated. For example, if 4 leads are used and 4 band features are extracted from each lead, the state vector s t The dimension is 4x4=16. This vector s t Completely describes the user's brain physiological state at time t.

[0072] Step S203: Action selection

[0073] The deep reinforcement learning module 103b receives the state vector s t , and select an action a according to the ∈―greedy strategy t .

[0074] Action Space: A set of predefined actions, each corresponding to a music genre. For example, the action space could be defined as A = {Action 1: Classical Music, Action 2: Natural White Noise, Action 3: Electronic Soothing Music, Action 4: Mixed Genres (randomly switching between multiple genres)}.

[0075] ∈―greedy strategy: Generate a random number between 0 and 1. If the number is less than the current exploration rate ∈, then randomly select an action (exploration) in the action space A; otherwise, the state vector s tInput into the DRQN ​​network, the network will output a Q value for each possible action, and select the action with the largest Q value (utilization). The exploration rate ∈ will change from ∈ start (0.9) linearly decreases to ∈ end $(0.1), achieving a smooth transition from exploration-oriented to exploit-oriented.

[0076] Step S204: Execute actions and interact with the environment

[0077] The music playing control module 104 selects the action a t , play the corresponding type of music clip, the duration is one time step (30 seconds).

[0078] Step S205: Reward calculation and state transfer

[0079] After the 30-second music playback ends, the system executes step S202 again to collect and process new EEG signals to obtain the state s at the next moment. t+1 Then calculate the immediate reward r t .

[0080] Reward Function: The reward function is the key to guiding model learning. This embodiment adopts the following design:

[0081] Instant Rewards: in, is the increment of theta wave power from time t to time t+1, is the absolute value of the decrease in alpha wave power from time t to time t+1. λ is a weighting factor (e.g., 0.5) that is adjusted experimentally to balance the importance of promoting theta waves and suppressing alpha waves. This function explicitly incentivizes the model to seek musical strategies that simultaneously enhance sleepiness (theta waves) and reduce wakefulness (alpha waves).

[0082] Termination Reward: When an episode ends, a significant termination reward is given.

[0083] Successful termination: If the power ratio of theta waves exceeds a threshold (such as 50%), hypnosis is considered successful and a large positive reward is given, such as R success =+100.

[0084] Failure termination: If the hypnosis time exceeds the preset maximum time (such as 7 minutes, corresponding to 14 time steps), but still does not meet the success standard, the hypnosis is considered a failure and a large negative reward is given, such as R failure =-50.

[0085] Step S206: Storing Experience

[0086] The four-tuple generated by this interaction <st ,a t ,r t ,s t+1 >Store in the experience replay pool.

[0087] Step S207: Model training

[0088] Determine whether the amount of data in the experience replay pool is greater than a preset minimum training amount (such as 1000). If so, randomly extract a small batch (such as 64) of experience data from the pool. j ,a j ,r j ,s j+1 >. Then, this batch of data is used to calculate the loss function and update the parameters θ of the DRQN ​​network.

[0089] Loss function: The mean square error loss function (MSE Loss) is used, whose goal is to make the Q value Q(s) predicted by the network j ,a j ;θ) approaches the target Q value y j .

[0090] Target Q value: The idea of ​​Double DQN is used here, that is, the action selection is determined by the current network \theta, while the value evaluation of the action is determined by the more stable target network θ ― This can alleviate the problem of overestimation of Q value.

[0091] Parameter update: The gradient of the loss function with respect to the network parameters θ is calculated through the back-propagation algorithm, and an optimizer (such as Adam) is used to update θ.

[0092] Target network update: Every certain number of training steps (such as every 100 steps), the parameters θ of the current network are completely copied to the target network θ ― , in order to maintain the relative stability of the target network.

[0093] Step S208: Loop and termination judgment

[0094] Check whether the current episode has ended (i.e., whether the termination condition for success or failure has been met). If not, set t = t + 1 and return to step S203 to continue the next round of interaction. If it has ended, reset the environment and start a new episode. This entire process is repeated continuously, and the model continues to learn and evolve through interaction with users.

[0095] 3.DRQN model structure and optimization strategy (refer to Figure 3 )

[0096] ​In order to process EEG signals with temporal and spatial characteristics, this embodiment adopts the deep recurrent Q network (DRQN) as the core model, whose structure is as follows Figure 3 shown.

[0097] Input Layer 301: Receives the pre-processed state vector s t As mentioned above, if there are 8 channels and each channel extracts 5 band features (δ, θ, α, β, γ), the input dimension is 8x5=40.

[0098] 1D Convolutional Layer (1D-CNN Layer) 302: Following the input layer, a one-dimensional convolutional layer is set. This layer performs convolution along the channel dimension, aiming to extract spatial correlation features between different EEG leads. For example, 16 filters (filters=16) can be set, each with a kernel size of 5 (kernel_size=5). This layer helps the model understand the coordinated patterns of activity in different brain regions.

[0099] Long Short-Term Memory (LSTM) Layer 303: The output of the convolutional layer is fed into an LSTM layer. LSTM is a special type of recurrent neural network (RNN) that is very good at capturing and learning long-term dependencies in time series data. This is crucial for analyzing the dynamic evolution of EEG signals. For example, an LSTM layer with 32 hidden units (hidden_units=32) can be set.

[0100] Output Layer 304: The output of the LSTM layer is finally connected to a fully connected layer, the output layer. The number of neurons in this layer is equal to the size of the action space (4 in this example). Each neuron outputs the Q value corresponding to an action.

[0101] Optimization strategy:

[0102] Curriculum Learning: In order to accelerate the convergence of the model and improve its generalization ability, the present invention introduces a curriculum learning strategy. In the initial stage of training, a simple task is set for the model, that is, a shorter hypnosis time limit (such as 1-3 minutes). This allows the model to quickly learn some basic and effective hypnosis strategies. After the model performs well on this simple task, the difficulty of the task is gradually increased, and the time limit is extended to the final goal (such as 7 minutes). This easy-to-difficult training paradigm, like humans learning courses, can effectively reduce the difficulty of training and avoid the model from being difficult to learn due to continuous negative feedback in the early stages.

[0103] Attention Mechanism: To allow the model to focus more on EEG features strongly associated with hypnosis, an attention layer can be added to the DRQN ​​model (for example, between the CNN and LSTM). This mechanism learns a weight distribution to weight the input feature vector. In hypnosis tasks, changes in theta and alpha waves are the most critical indicators. The attention mechanism allows the model to automatically learn to assign higher weights to the power features of these two bands, thereby enhancing the model's sensitivity to key state changes and enabling more accurate decision-making.

[0104] In summary, this invention integrates multiple advanced technologies, including brain-computer interfaces, deep reinforcement learning, curriculum learning, and attention mechanisms, to construct an intelligent, adaptive, closed-loop hypnosis system. This system overcomes the drawbacks of existing technologies, providing users with an unprecedented personalized, efficient, and flexible hypnosis experience, and has enormous potential for application in the field of sleep health.

[0105] The above description is only a preferred embodiment of the present invention and is not intended to limit the scope of protection of the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present invention shall be included in the scope of protection of the present invention.

Claims

1. An intelligent hypnosis method based on brain-computer interface, characterized in that: The following steps are involved: a. Signal acquisition: Use EEG acquisition equipment to obtain the user's EEG signals (EEG) in real time during sleep; b. State definition: Preprocess the collected EEG signals to extract the power values ​​of different EEG wave bands θ, α, β, and γ. This set of power values ​​is used as the state input of the deep reinforcement learning model. c. Action selection: The deep reinforcement learning model uses the ∈-greedy strategy to output an action based on the current input state, which corresponds to selecting one of multiple preset music types to play; d. Reward calculation: Based on the changes in the newly collected EEG signal state after the music is played, a reward value is calculated. This reward value is calculated to maximize the increase in theta wave power and suppress the power of alpha waves; e. Model training: The state, action, reward, and next state are stored as experience data in the experience replay pool. When the data volume reaches a predetermined value, the Q-Learning algorithm is used to train and optimize the network parameters of the deep reinforcement learning model. f. Iterative intervention: Repeat steps c to e, dynamically adjusting the music playback strategy through real-time interaction between the model and the user, until the user's EEG state reaches the preset sleep target, thereby guiding the user from the awake state to the sleep period.

2. The method according to claim 1, characterized in that In the signal acquisition step, the EEG acquisition device used is a polysomnogram (PSG) device or a portable EEG acquisition device.

3. The method according to claim 1, characterized in that The preprocessing in the state definition step includes: filtering and denoising the collected original EEG signal, segmenting it according to a fixed time window, calculating the power value of each band through fast Fourier transform (FFT) or wavelet transform, and finally performing normalization processing.

4. The method according to claim 1, wherein The music types in the action selection step include at least one or more of classical music, natural white noise, electronic soothing music, and mixed music.

5. The method according to claim 1, wherein The specific function of the reward calculation: Among them, r t is the instant reward at the current moment, is the increment of theta wave power at the current moment, is the decrement of α wave power at the current moment (taken as the absolute value), and λ is the weight coefficient used to balance the two.

6. The method according to claim 1, wherein The deep reinforcement learning model used in the model training is the Deep Q Network (DQN) or its variant, combined with a proximal policy optimization algorithm.

7. The method according to claim 1, characterized in that The method further includes: Course learning strategy: In the early stages of training, set shorter sleep induction tasks to allow the model to learn basic hypnosis strategies; then gradually increase the time limit of the tasks to reduce the training difficulty and enable the model to better adapt to hypnosis tasks of different lengths.

8. The method according to claim 1, characterized in that The method also includes an attention mechanism: an attention layer is added to the deep reinforcement learning model to weight the input EEG power vector to enhance the model's sensitivity to changes in theta and alpha wave power, which are strongly associated with hypnosis.

9. An intelligent hypnosis system based on brain-computer interface, characterized in that: include: EEG signal acquisition module, used to collect the user's EEG signals in real time; A signal preprocessing module, connected to the EEG signal acquisition module, is used to process the acquired EEG signals and extract the power value of each band; A deep reinforcement learning module, which contains a pre-trained deep reinforcement learning model that receives pre-processed power values ​​as state input and outputs a music genre selection action; A music playback control module plays music clips of corresponding types according to the actions output by the deep reinforcement learning module; Among them, the deep reinforcement learning module also calculates the reward value based on the changes in the EEG state fed back by the signal acquisition and preprocessing module after playing music, and continuously optimizes its own model.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the method according to any one of claims 1 to 8 is implemented.

Citation Information

Cited By

  • Self-adaptive human factor illumination control method and system based on electroencephalogram feedback

    CN120916297A

  • Personalized sleep optimization method and system based on electromagnetic-acoustic wave cooperative control of multi-dimensional sleep data

    CN121667637A