Brain-computer interface closed-loop neuromodulation method, storage medium, system and program product
By combining transcranial magnetic stimulation (TMS) and electroencephalography (EEG) signal acquisition, and using a deep reinforcement learning network to update stimulation parameters in real time, the problem of insufficient precision of stimulation parameters in personalized TMS is solved, and precise control of brain state is achieved.
Patent Information
- Application Number
- CN202510344256.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-21
- Publication Date
- 2025-12-26
- Estimated Expiration
- 2045-03-21
AI Technical Summary
In existing personalized transcranial magnetic stimulation (TMS) techniques, the precision of stimulation parameters is insufficient, making it impossible to accurately respond to changes in brain state, thus limiting the therapeutic effect.
A closed-loop neuromodulation method using brain-computer interface was adopted, which combines transcranial magnetic stimulation equipment with electroencephalogram (EEG) signal acquisition, and uses a deep reinforcement learning network to process EEG signal sequences and update stimulation parameters in real time to adapt to changes in brain state.
It achieves precise temporal targeting of brain states, improves the accuracy and efficiency of neuromodulation, and enables personalized treatment tailored to individual characteristics.
Smart Images

Figure CN119896816B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present disclosure relates to the technical field of neuromodulation in general, and more particularly, to a brain-computer interface closed-loop neuromodulation method, a storage medium, a system and a program product. BACKGROUND
[0002] Neuromodulation technology is a technology that precisely regulates brain neural activity through physical or chemical means (such as electrical stimulation, magnetic stimulation, drugs, etc.), which is used to treat nervous system diseases, improve neural function or study brain mechanisms, and has broad application prospects in the fields of medical neuroscience and computational neuroscience. After decades of traditional physical therapy practice, neuroscientists and doctors have gradually realized that studying individual differences in environmental, biological and psychosocial factors is crucial for disease treatment. This concept is summarized as “precision medicine”. When exploring new treatment methods, taking into account these individual differences helps to achieve personalized regulation tailored to the specific characteristics of the patient.
[0003] However, it is particularly important to apply the concept of “precision medicine” to emerging treatment technologies such as transcranial magnetic stimulation (TMS). As a safe and non-invasive neuromodulation method, TMS has achieved remarkable results in clinical treatment. By combining TMS with electroencephalogram (EEG) technology, the activity of electrical neuronal groups in the brain can be monitored with the assistance of scalp electrodes, allowing precise TMS stimulation to be implemented in specific brain states targeting the target area of persistent neuronal oscillations. This process is called brain state-dependent stimulation. This method of precisely regulating brain activity lays the foundation for achieving “personalized TMS” treatment.
[0004] Currently, “personalized TMS” is still in its infancy. On the one hand, an open-loop control strategy is usually used, where the brain state of the subject is first detected, then the TMS stimulation parameters are determined based on the detected brain state, and subsequent stimulation is performed according to the stimulation parameters. Since the ongoing neuronal activity in the stimulated brain region is not understood, the precise temporal targeting of the brain state is limited. On the other hand, when determining the TMS stimulation parameters based on the detected brain state, a fixed state-parameter rule is used, i.e., a certain brain state corresponds to a certain set of stimulation parameters. However, the accurate detection of brain states itself is difficult, and the accuracy of the fixed rule, which is derived from human analysis and summary, is limited, resulting in insufficient precision of the determined stimulation parameters.
[0005] Therefore, how to improve the stimulation precision of “personalized TMS” is a problem that needs to be solved in applying the concept of “precision medicine” to neuromodulation technology. SUMMARY
[0006] The present disclosure provides a brain-computer interface closed-loop neuromodulation method, a storage medium, a system and a program product, which are used to solve at least one of the above problems.
[0007] According to an aspect of the present disclosure, a brain-computer interface closed-loop neuromodulation method is provided, comprising: performing the following steps in a first cycle until a termination condition is met: controlling a transcranial magnetic stimulation device to operate for a first duration according to a set stimulation parameter to stimulate a subject; collecting an electroencephalogram signal sequence of the subject at a target brain region; processing the electroencephalogram signal sequence to obtain a preset index, wherein the preset index is used to reflect the activity of the target brain region; processing the preset index using a deep reinforcement learning network to obtain a stimulation parameter for a next cycle.
[0008] Optionally, the collecting the electroencephalogram signal sequence of the subject at the target brain region comprises: collecting electroencephalogram signal sequences of multiple channels of the subject, wherein the multiple signal collection points corresponding to the multiple channels include multiple points of the target brain region, the target brain region includes a default mode network, and the multiple points of the target brain region include at least one of the following: a mid-parietal point, a mid-occipital point, a central parietal point, a left parietal point, and a right parietal point.
[0009] Optionally, the processing the electroencephalogram signal sequence to obtain a preset index comprises: performing compression preprocessing on the electroencephalogram signal sequence to obtain a compressed electroencephalogram signal sequence; and processing the compressed electroencephalogram signal sequence to obtain a preset index.
[0010] Optionally, the preset index includes at least one of the following: an average power spectral density, a frequency band phase synchronization index, and a phase locking value.
[0011] Optionally, the compression preprocessing on the electroencephalogram signal sequence to obtain a compressed electroencephalogram signal sequence comprises: sampling each electroencephalogram signal sequence of each channel according to a preset sampling frequency to obtain multiple signal subsequences; retaining signals less than a preset frequency and compressing signals greater than or equal to the preset frequency for each signal subsequence of each channel to obtain each compressed signal subsequence of each channel; determining a correlation measure between the compressed signal subsequence and a compressed signal subsequence of the same time sequence in an adjacent channel for each compressed signal subsequence of each channel; and removing compressed signal subsequences with a correlation measure less than a correlation threshold from all compressed signal subsequences to obtain the compressed electroencephalogram signal sequence.
[0012] Optionally, before the steps of executing the following steps according to the first cycle until the end condition is met, the method further comprises: executing the following steps according to a second cycle until an end length is reached, wherein the length of the second cycle is less than the length of the first cycle, and the end length and the length of the first cycle have the same order of magnitude: controlling the transcranial magnetic stimulation device to operate for a second length according to the set stimulation parameters to stimulate the subject; collecting a sequence of electroencephalogram signals of the subject in the target brain area within a third length, wherein the third length is greater than the second length; processing the sequence of electroencephalogram signals to obtain the preset index; and processing the preset index using the deep reinforcement learning network to obtain the stimulation parameters of the next cycle, wherein during the execution of the above steps according to the second cycle, the deep reinforcement learning network is initialized and trained using an experience replay mechanism, and the experience data used for training includes experience data collected during the execution of the above steps according to the second cycle.
[0013] Optionally, the experience data collected from the subject in the previous transcranial magnetic stimulation is recorded as historical experience data, wherein the experience data used for training further includes the historical experience data; or wherein the deep reinforcement learning network based on which the initialization training is performed is a deep reinforcement learning network trained using the historical experience data.
[0014] Optionally, the brain-computer interface closed-loop neuromodulation method further comprises: comparing the preset index with a preset value range to obtain modulation evaluation information, wherein the preset value range is obtained by statistically analyzing the values of the preset index of a target population, and the modulation evaluation information is used to represent the evaluation of the neuromodulation effect.
[0015] Optionally, during the initialization training of the deep reinforcement learning network, the instantaneous reward of the deep reinforcement learning is determined according to the preset index and a preset value range, wherein the preset value range is obtained by statistically analyzing the values of the preset index of a target population.
[0016] According to another aspect of the present disclosure, a brain-computer interface closed-loop neuromodulation device is provided, comprising: a first cycle unit configured to cyclically operate the following units according to a first cycle until an end condition is met; a first control unit configured to control a transcranial magnetic stimulation device to operate for a first length according to set stimulation parameters to stimulate a subject; a first acquisition unit configured to acquire a sequence of electroencephalogram signals of the subject in a target brain area; a first processing unit configured to process the sequence of electroencephalogram signals to obtain a preset index, wherein the preset index is used to reflect the activity of the target brain area; and a first modulation unit configured to process the preset index using a deep reinforcement learning network to obtain the stimulation parameters of the next cycle.
[0017] Optionally, the first acquisition unit is further configured to acquire a plurality of electroencephalogram signal sequences of the subject, wherein the plurality of signal acquisition points corresponding to the plurality of channels include a plurality of point positions of the target brain region, the target brain region includes a default mode network, and the plurality of point positions of the target brain region include at least one of the following: a mid-parietal point, a mid-occipital point, a central parietal point, a left parietal point, and a right parietal point.
[0018] Optionally, the first processing unit is further configured to: perform compression preprocessing on the electroencephalogram signal sequence to obtain a compressed electroencephalogram signal sequence; and perform processing on the compressed electroencephalogram signal sequence to obtain the preset index.
[0019] Optionally, the preset index includes at least one of the following: an average power spectral density, a frequency band phase synchronization index, and a phase locking value.
[0020] Optionally, the first processing unit is further configured to: for the electroencephalogram signal sequence of each channel of the plurality of channels, sample a plurality of signal subsequences at a preset sampling frequency; for each signal subsequence of each channel, retain signals less than a preset frequency and compress signals greater than or equal to the preset frequency to obtain each compressed signal subsequence of each channel; for each compressed signal subsequence of each channel, determine a correlation measure between the compressed signal subsequence and a compressed signal subsequence of the same time sequence in an adjacent channel; and from all compressed signal subsequences, remove compressed signal subsequences with a correlation measure less than a correlation threshold to obtain the compressed electroencephalogram signal sequence.
[0021] Optionally, the method further comprises a second loop unit configured to, before running the first loop unit, cyclically run the following units according to a second period until an ending length is reached, wherein the length of the second period is less than the length of the first period, and the ending length has the same order of magnitude as the length of the first period; a second control unit configured to control the transcranial magnetic stimulation device to run for a second length of time according to a set stimulation parameter to stimulate the subject; a second acquisition unit configured to acquire an electroencephalogram signal sequence of the subject at the target brain region within a third length of time, wherein the third length of time is greater than the second length of time; a second processing unit configured to process the electroencephalogram signal sequence to obtain the preset index; and a second regulation unit configured to use the deep reinforcement learning network to process the preset index to obtain a stimulation parameter of a next period, wherein in the process of cyclically executing the above steps according to the second period, the deep reinforcement learning network is initialized and trained using an experience replay mechanism, and the experience data used for training includes experience data collected in the process of cyclically executing the above steps according to the second period.
[0022] Optionally, experience data collected previously when the subject receives transcranial magnetic stimulation is recorded as historical experience data, wherein the experience data used for the training further comprises the historical experience data; or wherein the deep reinforcement learning network based on which the initialization training is performed is a deep reinforcement learning network trained using the historical experience data.
[0023] Optionally, the brain-computer interface closed-loop neuromodulation apparatus further comprises an evaluation unit configured to compare the preset index with a preset value range to obtain modulation evaluation information, wherein the preset value range is obtained by statistics of values of the preset index of a target population, and the modulation evaluation information is used to represent an evaluation of a neuromodulation effect.
[0024] Optionally, in the process of performing the initialization training on the deep reinforcement learning network, an instant reward of deep reinforcement learning is determined according to the preset index and a preset value range, wherein the preset value range is obtained by statistics of values of the preset index of a target population.
[0025] According to another aspect of the present disclosure, there is provided a computer-readable storage medium storing instructions, wherein the instructions, when executed by at least one computing device, cause the at least one computing device to perform the brain-computer interface closed-loop neuromodulation method as described above.
[0026] According to another aspect of the present disclosure, there is provided a system comprising at least one computing device and at least one storage device storing instructions, wherein the instructions, when executed by the at least one computing device, cause the at least one computing device to perform the brain-computer interface closed-loop neuromodulation method as described above.
[0027] Optionally, the at least one computing device comprises: a first computing device configured to control operation of the transcranial magnetic stimulation device and the electroencephalogram signal acquisition device; a second computing device configured to process the electroencephalogram signal sequence to obtain the preset index; and a third computing device configured to process the preset index using the deep reinforcement learning network to obtain stimulation parameters for a next cycle.
[0028] According to another aspect of the present disclosure, there is provided a computer program product comprising instructions, wherein the instructions, when executed by at least one computing device, cause the at least one computing device to perform the brain-computer interface closed-loop neuromodulation method as described above.
[0029] The brain-computer interface closed-loop neuromodulation method, storage medium, system and program product according to the exemplary embodiments of the present disclosure can respond to the brain state changes of the subject in time to adjust the stimulation parameters by first controlling the transcranial magnetic stimulation device to operate for a first duration according to the set stimulation parameters, then collecting and analyzing the electroencephalogram signal sequence of the subject, and updating the stimulation parameters according to the analysis results (i.e., the preset indicators), and repeating the above process in a loop, so as to realize closed-loop neuromodulation and help improve the accuracy of brain state time targeting. When updating the stimulation parameters, the analysis and processing of the collected electroencephalogram signal sequence can extract the preset indicators with higher information concentration to reflect the brain state and improve the efficiency of determining the new stimulation parameters. Meanwhile, the use of the deep reinforcement learning network to process the preset indicators to determine the new stimulation parameters can not only avoid giving explicit brain states compared with the fixed state-parameter rules in the related art, but also can mine implicit data correlations, and compared with other machine learning networks, the experience accumulation and learning can be realized by means of the experience replay mechanism (Experience Replay) to mine more implicit data correlations suitable for the current subject, so as to more efficiently determine the stimulation parameters suitable for the current subject and improve the accuracy of the determined stimulation parameters. In general, the present disclosure adjusts the brain state according to the accurate individual characteristics in a self-adaptive closed-loop manner, which helps to provide more accurate and effective neuromodulation schemes (i.e., stimulation parameters).
[0030] Additional aspects and / or advantages of the general inventive concept will be set forth in part in the description that follows, and in part will be obvious from the description, or can be learned by practice of the general inventive concept. BRIEF DESCRIPTION OF DRAWINGS
[0031] These and / or other aspects and advantages of the present disclosure will become apparent and more readily appreciated from the following description, taken in conjunction with the accompanying drawings, in which:
[0032] Figure 1 FIG. 1 is a flowchart illustrating a brain-computer interface closed-loop neuromodulation method according to exemplary embodiments of the present disclosure.
[0033] Figure 2 FIG. 2 is a schematic diagram illustrating a modulation process of a brain-computer interface closed-loop neuromodulation method according to exemplary embodiments of the present disclosure.
[0034] Figure 3 FIG. 3 is a block diagram illustrating a brain-computer interface closed-loop neuromodulation device according to exemplary embodiments of the present disclosure.
[0035] Figure 4 FIG. 4 is a schematic diagram illustrating a brain-computer interface closed-loop neuromodulation system according to exemplary embodiments of the present disclosure. DETAILED DESCRIPTION
[0036] The following description with reference to the accompanying drawings is provided to assist in a comprehensive understanding of embodiments of the application as defined by the claims and their equivalents. Various specific details are included to assist in understanding but are not intended to limit the application. Therefore, one of ordinary skill in the art will recognize that various changes and modifications of the embodiments described herein can be made without departing from the scope and spirit of the application. In addition, descriptions of well-known functions and constructions are omitted for clarity and conciseness.
[0037] It is to be noted that "at least one of a plurality of items" appearing in the present disclosure are intended to cover all of the following instances: "any one of the plurality of items", "combination of any two or more of the plurality of items", and "all of the plurality of items". For example, "including at least one of A and B" means any one of the following: (1) including A alone, (2) including B alone, and (3) including both A and B. As another example, "performing at least one of steps 1 and 2" means any one of the following: (1) performing step 1 alone, (2) performing step 2 alone, and (3) performing both steps 1 and 2.
[0038] Reference will now be made to Figures 1 to 4 The brain-computer interface closed-loop neuromodulation method, storage medium, system, and program product according to the exemplary embodiments of the present disclosure are described in detail below.
[0039] Figure 1 is a flowchart illustrating a brain-computer interface closed-loop neuromodulation method according to an exemplary embodiment of the present disclosure. The brain-computer interface closed-loop neuromodulation method according to the exemplary embodiments of the present disclosure can be implemented in a computing device having sufficient computing power.
[0040] Reference will now be made to Figure 1 Steps S101 to S104 are executed in a first cycle until an end condition is met. The end condition will be described together when step S104 is introduced, and will not be expanded here.
[0041] In step S101, the transcranial magnetic stimulation device is controlled to operate for a first duration according to a set stimulation parameter to stimulate the subject.
[0042] As an example, the stimulation parameter includes a stimulation intensity and / or a stimulation frequency. As for the stimulation intensity, a routine motor evoked potential (MEP) can be performed on the subject to determine the subject's stimulation response and determine an initial stimulation intensity motor threshold (MT), and in subsequent execution, the stimulation intensity is maintained within a safe range (for example, including but not limited to 80% to 120%) of the initial stimulation intensity motor threshold (MT).
[0043] As an example, the first time length is in a range of 3 min to 10 min, such as 4 min, 5 min, 7 min, etc., which is long enough to ensure reliable effect of neuromodulation, and short enough to facilitate timely updating of stimulation parameters, thus helping to achieve precise personalized neuromodulation.
[0044] In step S102, the EEG signal sequence of the subject in the target brain region is collected.
[0045] Steps S101 and S102 can implement a TMS-EEG combined control process, for example, a computing device can be used to control the transcranial magnetic stimulation device and the EEG signal collection device in real time to achieve stimulation and signal collection. As an example, the computing device can use a rotation control mechanism to turn off the transcranial magnetic stimulation device when the EEG signal collection device is enabled, and enable the EEG signal collection device during the interval period when the transcranial magnetic stimulation device performs stimulation, so as to achieve real-time regulation. It should be understood that "real-time" does not mean strict synchronization, but means continuous operation according to the set frequency, so as to allow interval within the design range and reasonable delay due to technical limitations. The same applies to "real-time" in the following description, which will not be repeated. Moreover, "real-time regulation" here emphasizes the closed-loop regulation of transcranial magnetic stimulation, which can collect EEG signals and analyze and update stimulation parameters in real time during stimulation.
[0046] As an example, the collection frequency of the EEG signal collection device can be fixed at 500 Hz to 2000 Hz, and a wired data interface can be used to provide data to the computing device for performing step S103 to analyze the EEG signal sequence, and the data interface can use the TCP / IP Socket communication protocol.
[0047] Optionally, step S102 includes collecting EEG signal sequences of multiple channels of the subject, wherein the multiple signal collection points corresponding to the multiple channels include multiple points in the target brain region, the target brain region includes the default mode network, and the multiple points in the target brain region include at least one of the following: midparietal point (Pz), midoccipital point (POz), central parietal point (CPz), left parietal point (P3), and right parietal point (P4). The multiple points in the target brain region as signal collection points can also be used as stimulation points of the transcranial magnetic stimulation device to achieve multi-point transcranial magnetic stimulation, so as to convert the transcranial magnetic stimulation into multi-node brain stimulation at the level of the default mode network of the brain.
[0048] The default mode network (DMN) in the brain is a key social function regulation system, which plays an important role in the pathogenesis of various neuropsychiatric diseases. Studies have shown that the DMN, as a highly integrated platform, is responsible for processing complex information related to social cognition and emotional processes. The DMN mainly includes the following brain regions: the posterior cingulate cortex (PCC), the anterior cingulate cortex (ACC), the precuneus cortex (Precuneus), the medial prefrontal cortex (mPFC), the dorsolateral prefrontal cortex (DLPFC), the temporo-parietal junction (TPJ), etc. These brain regions are connected by complex neural networks and jointly maintain the default state function of the brain.
[0049] As an example, the electroencephalogram signal acquisition device includes an electroencephalogram cap, which covers the monitoring range of the default mode network of the PCC and the Precuneus, and based on the standard 10-20 system electrode placement method stipulated by the International Society for Electroencephalography, the PCC includes two points Pz and POz, and the Precuneus includes three points CPz, P3 and P4. Of course, other points of the target brain region can also be used, and the present disclosure does not limit this.
[0050] In step S103, the electroencephalogram signal sequence is processed to obtain a preset index.
[0051] The preset index is used to reflect the activity of the target brain region and can realize indirect and quantitative representation of the brain state.
[0052] Optionally, the preset index includes at least one of the following: average power spectral density (PSD), frequency band phase synchronization index (PSI), and phase locking value (PLV).
[0053] Specifically, the average power spectral density is the average value of the power spectral density of the electroencephalogram signal at a specific frequency k. The processed signal sequence can be divided into several segments according to a time window, and the power spectral density of the signal at a specific frequency k in each time window is calculated, and then the average value of all windows is calculated. The formula can be expressed as:
[0054] .
[0055] In the formula, P m (k) is the power spectral density of the mth time window, M is the number of time windows, and the length of each time window is, for example, no more than 2 s. The power spectral density within each time window is calculated using a standard Discrete Fourier Transform (DFT) or Fast Fourier Transform (FFT) calculation method. As an example, for the case where the number of channels of the electroencephalogram signal sequence is multiple, since the average power spectral density can be calculated for a single channel, the average power spectral density of each channel can be obtained, and statistical values can be further calculated for the average power spectral densities of the channels, which is not limited by the present disclosure.
[0056] The band phase synchronization index and the phase locking value are both used to measure the consistency of the frequency phase of different measurement regions. The phase angle can be obtained using a Hilbert Transform, and then the band phase synchronization index and the phase locking value are calculated according to their respective formulas. It should be understood that since both of these two indicators are for calculating the phase consistency of two signals, when the number of channels of the electroencephalogram signal sequence is two, the calculation can be directly performed, and when the number of channels is greater than two, the signals of multiple channels can be paired first, then the indicators of the paired channel signals are calculated, and finally the indicators of multiple channel pairs are summarized, such as calculating the average value, mode, median, etc. statistical values, as the final indicators.
[0057] Optionally, step S103 comprises: performing compression preprocessing on the electroencephalogram signal sequence to obtain a compressed electroencephalogram signal sequence; and processing the compressed electroencephalogram signal sequence to obtain the preset indicator. By first performing real-time compression preprocessing on the acquired electroencephalogram signal sequence, the data amount can be reduced, the calculation speed when processing the preset indicator can be improved, the real-time processing capability can be improved, which helps to quickly realize real-time processing after the signal is acquired to determine the brain state (represented by the preset indicator) required to trigger the transcranial magnetic stimulation, and fully guarantee the response rate of the closed-loop neuromodulation. It should be understood that at this time, the calculation time of the preset indicator is extremely short and can be almost ignored.
[0058] Regarding the compression preprocessing, optionally, the operation of performing the compression preprocessing on the electroencephalogram signal sequence in step S103 to obtain the compressed electroencephalogram signal sequence includes: sampling the electroencephalogram signal sequence of each channel in the plurality of channels according to a preset sampling frequency to obtain a plurality of signal subsequences; for each signal subsequence of each channel, retaining signals less than a preset frequency, and compressing signals greater than or equal to the preset frequency to obtain each compressed signal subsequence of each channel. By sampling the signal sequence of each channel, the signal sequence can be split to obtain a plurality of signal subsequences, providing a stable and reliable data basis for signal compression. Specifically, for each signal subsequence, by retaining the required low-frequency signals (i.e., signals less than the preset frequency) and compressing the high-frequency signals, detailed key information can be selectively retained, reducing the data volume of other information, thereby improving the information concentration of the compressed data. As an example, the preset frequency is 10 Hz, and the signals less than 10 Hz (i.e., alpha waves, beta waves, delta waves, and theta waves) are retained. Of course, the preset frequency can also be set to other frequencies as needed, and the present disclosure does not limit this.
[0059] As an example, in order to obtain sufficient data to provide a reliable basis for the decision of neuromodulation, the collected electroencephalogram signal sequence can be required to have a sufficient time length, for example, including but not limited to reaching 1s to 2s. In actual execution, the signal can be stopped when the time length represented by the collected electroencephalogram signal sequence is long enough, and step S103 can be performed on the collected signal, or the compression preprocessing can be performed while collecting the signal until the time length represented by the final signal sequence is long enough, and the present disclosure does not limit this.
[0060] Further optionally, the operation of performing the compression preprocessing on the electroencephalogram signal sequence in step S103 to obtain the compressed electroencephalogram signal sequence further includes: for each compressed signal subsequence of each channel, determining a correlation measure between the compressed signal subsequence and a compressed signal subsequence of the same time sequence in an adjacent channel; from all compressed signal subsequences, removing compressed signal subsequences with a correlation measure less than a correlation threshold to obtain the compressed electroencephalogram signal sequence. By further calculating the correlation measure between the compressed signal subsequences of the adjacent channels of the same time sequence, and removing the compressed signal subsequences with insufficient correlation, the invalid data can be further reduced, the information concentration can be improved, and the real-time processing capability can be further improved.
[0061] As an example, the correlation measure includes, but is not limited to, covariance (Cov).
[0062] As an example, since this embodiment will delete the compressed signal subsequences with insufficient correlation, the final obtained compressed electroencephalogram signal sequence may be shorter than the original electroencephalogram signal sequence. Therefore, for each channel, the electroencephalogram signal sequence can be collected, sampled, compressed and removed according to the preset sampling frequency. That is, for each preset sampling frequency, when the number of collected electroencephalogram signals (for example, f / 10 electroencephalogram signals, where f is the preset sampling frequency) reaches the corresponding number, these signals are determined as a signal subsequence, compressed to obtain the compressed signal subsequences of multiple channels in the current time sequence, the correlation measure of adjacent channels is calculated, and then the compressed signal subsequences with insufficient correlation are removed. At the same time, the length of all obtained compressed signal subsequences is continuously counted. If the length is long enough, for example, but not limited to, reaching a length threshold of 1 s or 2 s, it is considered that the data volume is sufficient as a basis for decision of neuromodulation, and the collection of electroencephalogram signals is stopped. Of course, a certain length of electroencephalogram signal sequence can be collected and preprocessed first. If the length of the compressed signal sequence is not enough, the electroencephalogram signal is continuously collected. Other collection methods can also be used as long as the length of the final obtained compressed electroencephalogram signal sequence is sufficient. The specific collection and preprocessing methods are not limited in the present disclosure.
[0063] As an example, the method of real-time collection, compression and processing of electroencephalogram signals is to obtain the original signal data of electroencephalogram in real time, including the time stamp (Ts, accurate to 1 ms) of each channel and the original electrical signal (unit uV) of each time stamp, and convert the original signal data into a triple T(n) (each channel can obtain multiple triples), and push it into the compressed processing queue (Compressed Process Queue, CPQ), that is, the compressed electroencephalogram signal sequence. Each triple contains: 1, compressed value (Cpv); 2, variance value (Var); 3, time stamp (Ts). The specific steps include the following.
[0064] 1) The received f / 10 (where f is the sampling frequency) sampling electrical signals (i.e. a signal subsequence) of each channel are compressed using the Median value and algorithm, that is, the data from 10 Hz to 1 kHz is compressed, the median value (i.e. the middle value after sorting by size) of these data is taken as the compressed value, so as to reduce the data dimension while retaining the signals below the beta wave (alpha wave, beta wave, delta wave, theta wave) less than 10 Hz, and the variance value of the f / 10 sampling electrical signals is added to reflect the data change degree, and the middle point of the time stamp of the f / 10 sampling electrical signals is added to ensure the time correlation of the signals, and finally the triple (Tuple) T(n) is obtained, where n is the compression sequence number.
[0065] 2) Calculate the covariance Cov between the current compressed f / 10 signals and the time-aligned f / 10 signals in the adjacent channels as a correlation measure, if the correlation measure is less than a correlation threshold (for example, including but not limited to 0.1), discard the T(n), otherwise put the T(n) into the CPQ queue.
[0066] 3) When the data preprocessed by compression in the CPQ queue reaches a minimum processing time threshold (for example, 1s to 2s), start processing the data, calculate the power spectral density of the alpha band, beta band, etc. of each channel in the window period, and the phase synchronization index or phase locking value of each channel in the alpha band, beta band, etc. After processing, the average preset indicators of each channel are obtained, the processed data in the CPQ queue is cleared, and the calculated preset indicator results are given to the computing device for executing step S104 in real time.
[0067] In step S104, the preset indicators are processed using a deep reinforcement learning network to obtain the stimulation parameters of the next period.
[0068] This step uses a deep reinforcement learning network to make decisions for neuromodulation. Deep reinforcement learning is a technique that combines deep learning and reinforcement learning, used to learn the optimal policy in a given environment.
[0069] As an example, the deep reinforcement learning network is a DQN (Deep Q-Network) network. The core of DQN is to use a deep neural network to approximate the Q function, i.e. the state-action value function, which represents the expected return of taking a particular action in a particular state. DQN contains the following connotations.
[0070] Q-Learning basis: DQN is based on the Q-Learning algorithm, which is a model-free reinforcement learning algorithm designed to learn an optimal policy that maximizes cumulative rewards.
[0071] Neural network approximation: DQN uses a deep neural network to approximate the Q function, which receives the state (in this disclosure, the state is specifically the preset indicator) as input and outputs the Q value of each possible action (in this disclosure, the action is specifically the stimulation parameter).
[0072] Experience replay: To reduce the correlation between data and avoid overfitting, DQN uses an experience replay mechanism to store the experience of the agent interacting with the environment (including the state s t , the action a t , the immediate reward r, the state s t+1) in a replay memory and then sampled randomly for training. In an epoch of learning, several successive differentiable Q functions can be specified and updated iteratively according to a set of network weight steps.
[0073] Target Q Network: To improve the stability of training, DQN uses two neural networks: the main Q network and the target Q network. The main Q network is used to predict Q values, while the target Q network is used to generate target Q values (the output value of the target Q network needs to be processed to obtain the target Q value). The weights of the target Q network are periodically updated to the weights of the current main Q network to follow the main Q network.
[0074] The calculation of the target Q value can follow the formula:
[0075] .
[0076] In the formula, Q new (s t , a t ) is the target Q value, α is the learning rate, Q(s t , a t ) is the Q value calculated by the target Q network for taking action a t in state s t , r is the immediate reward, γ is the discount factor, and max a Q(s t+1 , a) is the maximum value among all possible actions a in the next state s t+1 .
[0077] Learning strategy: The learning strategy of DQN includes selection policy, reward policy and stopping policy. The selection policy determines how the agent selects an action in a given state, for example, selecting the action with the maximum predicted Q value. The reward policy is the determination of the reward under the given condition, that is, the immediate reward. The stopping policy stops decision-making when the corresponding condition is met, that is, the end condition of the first cycle of steps S101-S104 is recorded in the present disclosure. As an example, the end condition includes: 1. After the preset time length, the change rate of the preset indicator of the last preset round is less than the preset change rate. The preset time length can guarantee the basic neuroregulation time length, so as to meet the effect standard. The preset change rate indicates that the state of the subject tends to be stable, indicating that it may be difficult to further change the state of the subject by continuing transcranial magnetic stimulation, and thus the regulation can be ended. For example, the condition can be that after 5 minutes, the change rate of the PSD value of the last two rounds is less than 0.1%, and the change rate of the PSI / PLV value of the last two rounds is less than 0.1%. 2. From the beginning of transcranial magnetic stimulation on the subject, the total time length experienced reaches the preset regulation time length (for example, including but not limited to 20 minutes, 30 minutes), that is, whether the state of the subject can continue to be changed or not, the regulation is ended in the case that a long enough time is experienced throughout the regulation, which can reduce the physical and mental consumption of the subject caused by long-term regulation, and can reduce the physical load of the subject. When any of the two conditions is met, the cycle ends. The combination of the two end conditions can balance the effect of neuroregulation and the physical feeling of the subject, and realize more reasonable neuroregulation.
[0078] The training process of DQN includes the following steps.
[0079] 1) Initialize the weights of the main Q network and the target Q network.
[0080] 2) Interact with the environment, collect experiences and store them in the replay memory.
[0081] 3) Randomly sample a batch of experiences from the replay memory.
[0082] 4) Use the main Q network to predict the Q values of these experiences.
[0083] 5) Use the target Q network to calculate the target Q values according to the above formula.
[0084] 6) Calculate the loss according to the predicted Q values and the corresponding target Q values, and update the weights of the main Q network using gradient descent method.
[0085] 7) Update the weights of the target Q network periodically to approach the weights of the main Q network.
[0086] It should be noted that, as for the stimulation parameters of transcranial magnetic stimulation, high stimulation frequency (for example, including but not limited to greater than or equal to 1 Hz) and high stimulation intensity (for example, including but not limited to 100% to 120% of the initial stimulation intensity motor threshold MT) can achieve inhibition of the activity level of the target brain region, and low stimulation frequency (for example, including but not limited to less than 1 Hz) and low stimulation intensity (for example, including but not limited to 80% to 100% of the initial stimulation intensity motor threshold MT) can achieve stimulation of the activity level of the target brain region.
[0087] In addition, as for the relationship between the preset indicators and the activity state of the default mode network (DMN), taking the three preset indicators of the average power spectral density, the frequency band phase synchronization index, and the phase locking value listed in the foregoing as examples, the increase of any one of the three preset indicators can indicate that the DMN is active, at this time, the subject may be in a self-reflection or rest state, and otherwise, it indicates that the DMN is inhibited by an external task or high-frequency high-intensity transcranial magnetic stimulation, and this correlation can be gradually refined through the continuous learning of the deep reinforcement learning network.
[0088] As an example, throughout steps S101 to S104, the duration of the first period includes a first duration, a duration consumed by collecting the electroencephalogram signal sequence, and a duration consumed by processing to obtain the preset indicators and the stimulation parameters of the next period. The second item here is determined by the signal sequence duration required for processing the preset indicators, for example, 1s or 2s as an example in the foregoing, and since there is a possibility of removing part of the data, the actual collected signal duration will be greater than or equal to the duration. The processing process of the third item is completed almost instantaneously by the computing device. Therefore, it can be approximately considered that the duration of the first period is equal to the sum of the first duration and the duration of the collected electroencephalogram signal sequence, and since the duration of the second item may fluctuate slightly, the duration of the first period is not exactly fixed. The value range of the first duration is, for example, 3 min to 10 min, and specifically, for example, including but not limited to 4 min, 5 min, and 7 min.
[0089] The steps S101 to S104 above represent the formal neuromodulation process. In some embodiments, before the steps S101 to S104 are executed in a first cycle, the brain-computer interface closed-loop neuromodulation method according to the exemplary embodiments of the present disclosure further comprises: executing the following steps in a second cycle until a termination duration is reached, wherein the duration of the second cycle is less than the duration of the first cycle, and the termination duration has the same order of magnitude as the duration of the first cycle: controlling the transcranial magnetic stimulation device to operate for a second duration according to the set stimulation parameters to stimulate the subject; collecting a sequence of electroencephalogram signals of the subject in the target brain region within a third duration, wherein the third duration is greater than the second duration; processing the sequence of electroencephalogram signals to obtain a preset index; processing the preset index using the deep reinforcement learning network to obtain the stimulation parameters for the next cycle, wherein in the process of executing the above steps in the second cycle, the deep reinforcement learning network is initialized and trained using an experience replay mechanism, and the experience data used for training includes the experience data collected in the process of executing the above steps in the second cycle. The first cycle is the duration of a modulation cycle in the formal neuromodulation. By using a period of time (i.e., a period of time with a length equal to the termination duration) with the same order of magnitude as the first cycle for high-frequency cyclic modulation (i.e., cycling with a shorter second cycle) before the formal neuromodulation, a large amount of experience data can be collected in a short period of time, and based on these experience data, the deep reinforcement learning network is initialized and trained for the current subject, which helps to further improve the stimulation accuracy of the individualized transcranial magnetic stimulation. It should be understood that the stimulation parameters finally obtained in this stage are used as the stimulation parameters when the step S101 is first executed. It should also be understood that for the embodiment described above in which the initial stimulation intensity motor threshold of the subject is determined first, the initialization and training described here should be performed after the determination is completed. It should also be understood that for the second condition in the end condition of the above-described execution of the steps S101 to S104, i.e., the total duration experienced since the transcranial magnetic stimulation of the subject is performed reaches the preset modulation duration, this total duration should include the termination duration described here, i.e., the initialization and training stage should be included in the timing range. For example, if the preset modulation duration is 30 min and the termination duration is 5 min, the duration of the execution of the steps S101 to S104 should be no more than 25 min.
[0090] As an example, in the initialization training stage, the length of the second cycle includes the second length, the third length, and the length consumed by processing the preset index and the stimulation parameter of the next cycle. Since the processing of the last item is completed almost instantaneously by the computing device, the length of the second cycle can be approximately considered as the sum of the second length and the third length. The second length is in the range of, for example, 3s to 10s, and specifically includes, for example, but not limited to, 4s, 5s, 7s. The third length is in the range of, for example, 5s to 15s, and specifically includes, for example, but not limited to, 8s, 10s, 13s. The second length is less than the third length, so as to collect more electroencephalogram signals for processing by the deep reinforcement learning network, and improve the training efficiency. The end length is in the range of, for example, 3min to 10min, and specifically includes, for example, but not limited to, 5min, 8min. It should be understood that these operation steps are similar to the operations of steps S101 to S104, and some steps are different in parameters, and thus the specific operation mode can be referred to the foregoing description of steps S101 to S104, which will not be repeated here.
[0091] Regarding the initialization training, the experience data collected from the transcranial magnetic stimulation previously received by the subject is optionally denoted as historical experience data. The experience data used in the initialization training further includes the historical experience data; or the deep reinforcement learning network used in the initialization training is a deep reinforcement learning network trained using the historical experience data. It should be understood that the brain-computer interface closed-loop neuromodulation method according to the example embodiments of the present disclosure is used to perform a complete neuromodulation for a subject, but the same subject may need to receive multiple neuromodulations, and thus the historical experience data accumulated in the previous neuromodulation can be used as a reference for the later neuromodulation of the same subject. By adding the historical experience data to the experience data used in the initialization training, or directly using the deep reinforcement learning network previously trained and used by the same subject and continuing to train, the accuracy of the deep reinforcement learning network can be improved with the help of the historical experience.
[0092] Optionally, the brain-computer interface closed-loop neuromodulation method according to the example embodiments of the present disclosure further includes: comparing the preset index with a preset value range to obtain modulation evaluation information, wherein the preset value range is obtained by statistically analyzing the values of the preset index of a target population, and the modulation evaluation information is used to represent the evaluation of the neuromodulation effect. By selecting the target population to statistically analyze the values of the preset index, a cross-subject database can be constructed, a general health standard can be obtained, and a reference for evaluating the brain state of the subject and the neuromodulation effect can be provided.
[0093] As an example, the target population is a population that does not need to be neuro-modulated by transcranial magnetic stimulation, in other words, if transcranial magnetic stimulation is regarded as a treatment method for treating certain conditions, then the target population is a population without the condition, which can be a treatment target and provide a reference for evaluating the treatment effect. In addition, the target population can further include people of different ages, such as children, adolescents, adults, and the elderly, and the gender ratio in the target population can be reasonably controlled according to statistical needs, and the present disclosure does not limit this.
[0094] As an example, regarding the timing of evaluation, evaluation can be concentrated after the entire neuro-modulation is completed, evaluation can be performed at the end of each first period, and evaluation can be performed at other reasonable times, and the present disclosure does not limit this. When evaluating, the time can be taken as the abscissa, the preset index can be taken as the ordinate, the change curve of the evaluated preset index can be drawn, and the reference line can be drawn in the change curve according to the preset value range, so as to intuitively show the change of the preset index relative to the preset value range.
[0095] Optionally, in the process of initializing and training the deep reinforcement learning network, the instant reward of the deep reinforcement learning is determined according to the preset index and the preset value range, wherein the preset value range is obtained by statistics of the preset index of the target population. By determining the instant reward according to the preset index and the preset value range, specifically, the instant reward is made larger when the preset index is closer to the preset value range, and the instant reward reaches the maximum when the preset index enters the preset value range, which can guide the deep reinforcement learning network to learn how to adjust the stimulation parameter so that the preset index is as close to or even enters the preset value range, thereby providing a clear basis for the decision of the deep reinforcement learning network. In addition, such instant reward can also determine reasonable stimulation parameters, so that the preset index is lowered when it is too high, and the preset index is raised when it is too low, which can realize bidirectional regulation of the state of the target brain area, for example, when the preset index represents that the target brain area is too active, the appropriate inhibition of the target brain area is realized by enhancing the stimulation parameter, and when the preset index represents that the activity of the target brain area is insufficient, the appropriate activation of the target brain area is realized by weakening the stimulation parameter, which is helpful to realize reasonable neuro-modulation effect.
[0096] Figure 2 is a schematic diagram of a regulation process of a brain-computer interface closed-loop neuro-modulation method according to an example embodiment of the present disclosure.
[0097] Referring to Figure 2 The whole process includes an initialization phase and a formal regulation phase. In the initialization phase, the loop of the end length is used for high-frequency stimulation (each loop period includes transcranial magnetic stimulation of the second length, electroencephalogram signal acquisition of the third length, signal analysis of a short duration which can be ignored, Figure 2The initialization training of the deep reinforcement learning network for the subject is completed, and an initial stimulation scheme is determined. In the following formal regulation phase, transcranial magnetic stimulation of a first duration, a short electroencephalogram signal acquisition, and a short-to-negligible duration signal analysis Figure 2 The analysis part is also not shown in the middle), until the end condition is met.
[0098] Figure 3 is a block diagram illustrating a brain-computer interface closed-loop neuromodulation device according to an exemplary embodiment of the present disclosure.
[0099] Referring to Figure 3 , the brain-computer interface closed-loop neuromodulation device 300 comprises a first loop unit 301, a first control unit 302, a first acquisition unit 303, a first processing unit 304, and a first regulation unit 305.
[0100] The first loop unit 301 can cyclically operate the first control unit 302, the first acquisition unit 303, the first processing unit 304, and the first regulation unit 305 according to a first period until the end condition is met.
[0101] The first control unit 302 can control the transcranial magnetic stimulation device to operate for a first duration according to the set stimulation parameters to stimulate the subject.
[0102] The first acquisition unit 303 can acquire a sequence of electroencephalogram signals of the subject at the target brain region.
[0103] Optionally, the first acquisition unit 303 can also acquire a sequence of electroencephalogram signals of the subject at multiple channels, wherein the multiple channels correspond to multiple signal acquisition points including multiple points of the target brain region, and the target brain region includes a default mode network, and the multiple points of the target brain region include at least one of the following: a mid-parietal point, a mid-occipital point, a central-parietal point, a left parietal point, and a right parietal point.
[0104] The first processing unit 304 can process the sequence of electroencephalogram signals to obtain a preset index, wherein the preset index is used to reflect the activity of the target brain region.
[0105] Optionally, the first processing unit 304 can further: perform compression preprocessing on the sequence of electroencephalogram signals to obtain a compressed sequence of electroencephalogram signals; and process the compressed sequence of electroencephalogram signals to obtain the preset index.
[0106] Optionally, the first processing unit 304 can further: sample the electroencephalogram signal sequence of each channel in the plurality of channels to obtain a plurality of signal subsequences according to a preset sampling frequency; for each signal subsequence of each channel, retain signals less than a preset frequency and compress signals greater than or equal to the preset frequency to obtain each compressed signal subsequence of each channel; for each compressed signal subsequence of each channel, determine a correlation measure between the compressed signal subsequence and a compressed signal subsequence of the same time sequence in an adjacent channel; and remove compressed signal subsequences with a correlation measure less than a correlation threshold from all compressed signal subsequences to obtain a compressed electroencephalogram signal sequence.
[0107] The first regulation unit 305 can use a deep reinforcement learning network to process the preset index to obtain the stimulation parameter of the next period.
[0108] Optionally, the preset index includes at least one of the following: average power spectral density, frequency band phase synchronization index, and phase locking value.
[0109] Optionally, the brain-computer interface closed-loop neuromodulation device 300 further includes a second loop unit, a second control unit, a second acquisition unit, a second processing unit, and a second regulation unit (none of which is shown in the figure). The second loop unit can cyclically operate the following units according to a second period before operating the first loop unit until a termination duration is reached, wherein the duration of the second period is less than the duration of the first period, and the termination duration has the same order of magnitude as the duration of the first period; the second control unit can control the transcranial magnetic stimulation device to operate for a second duration according to the set stimulation parameter to stimulate the subject; the second acquisition unit can acquire an electroencephalogram signal sequence of the subject in the target brain region within a third duration, wherein the third duration is greater than the second duration; the second processing unit can process the electroencephalogram signal sequence to obtain a preset index; and the second regulation unit can use a deep reinforcement learning network to process the preset index to obtain the stimulation parameter of the next period, wherein during the cyclic execution of the above steps according to the second period, the deep reinforcement learning network is initialized and trained using an experience replay mechanism, and the experience data used for training includes experience data collected during the cyclic execution of the above steps according to the second period.
[0110] It should be understood that there are similar parts in the operations performed by the second loop unit and the first loop unit 301, and such similar operation parts can be referred to the first loop unit 301; similarly, the second control unit, the second acquisition unit, the second processing unit, and the second regulation unit have similar operations with the first control unit 302, the first acquisition unit 303, the first processing unit 304, and the first regulation unit 305, respectively, and can be referred to, and will not be described one by one. In addition, the second processing unit and the first processing unit 304 can be the same unit or different units, and the present disclosure does not limit this, and the second regulation unit and the first regulation unit 305 can also be the same unit or different units, and the present disclosure does not limit this.
[0111] Optionally, the experience data collected when the subject is subjected to transcranial magnetic stimulation in the past is recorded as historical experience data, wherein the experience data used for training further comprises the historical experience data; or wherein the deep reinforcement learning network based on which the initialization training is performed is a deep reinforcement learning network trained using the historical experience data.
[0112] Optionally, the brain-computer interface closed-loop neuromodulation apparatus 300 further comprises an evaluation unit (not shown in the figure), which can compare the preset index with a preset value range to obtain regulation evaluation information, wherein the preset value range is obtained by statistically analyzing the values of the preset index of the target population, and the regulation evaluation information is used to represent the evaluation of the neuromodulation effect.
[0113] Optionally, in the process of initializing and training the deep reinforcement learning network, the instant reward of the deep reinforcement learning is determined according to the preset index and the preset value range, wherein the preset value range is obtained by statistically analyzing the values of the preset index of the target population.
[0114] The above has been described with reference to Figures 1 to 3 The brain-computer interface closed-loop neuromodulation method and apparatus according to the exemplary embodiments of the present disclosure are described.
[0115] Figure 3 Each unit in the brain-computer interface closed-loop neuromodulation apparatus shown can be configured as software, hardware, firmware, or any combination of the above, which performs a specific function. For example, each unit can correspond to a dedicated integrated circuit, a pure software code, or a combination of software and hardware modules. In addition, one or more functions implemented by each unit can also be uniformly performed by components in a physical entity device (e.g., a processor, a client, or a server, etc.).
[0116] In addition, with reference to Figure 1The described brain-computer interface closed-loop neuromodulation method can be implemented by a program (or instructions) recorded on a computer readable storage medium. For example, according to an exemplary embodiment of the present disclosure, a computer readable storage medium storing instructions can be provided, wherein when the instructions are run by at least one computing device, the at least one computing device is caused to perform the brain-computer interface closed-loop neuromodulation method according to the present disclosure.
[0117] The computer program in the above computer readable storage medium can be run in an environment deployed in a computer device such as a client, a host, a proxy device, a server, etc. It should be noted that the computer program can also be used to perform additional steps or more specific processing when performing the above steps, the content of which has been described with reference to the above description of the related method, and thus will not be repeated here. Figure 1 During the description of the related method, the above steps and further processing have been mentioned, and thus will not be repeated here.
[0118] It should be noted that each unit in the brain-computer interface closed-loop neuromodulation device according to an exemplary embodiment of the present disclosure can be completely realized by running a computer program to realize the corresponding function, i.e., each unit corresponds to each step in the functional architecture of the computer program, so that the entire system is called by a special software package (e.g., lib library) to realize the corresponding function.
[0119] On the other hand, Figure 3 Each unit shown can also be realized by hardware, software, firmware, middleware, microcode or any combination thereof. When realized by software, firmware, middleware or microcode, the program code or code segment for performing the corresponding operation can be stored in a computer readable medium such as a storage medium, so that the processor can perform the corresponding operation by reading and running the corresponding program code or code segment.
[0120] For example, the exemplary embodiments of the present disclosure can also be implemented as a computing device including a storage component and a processor, the storage component storing a set of computer executable instructions, when the set of computer executable instructions is executed by the processor, performing the brain-computer interface closed-loop neuromodulation method according to the exemplary embodiments of the present disclosure.
[0121] Specifically, the computing device can be deployed in a server or a client, or on a node device in a distributed network environment. In addition, the computing device can be a PC computer, a tablet device, a personal digital assistant, a smart phone, a web application or other devices capable of executing the above instruction set.
[0122] Here, the computing device need not be a single computing device, but can be a collection of devices or circuits that individually or jointly execute the instructions (or sets of instructions) as described above. The computing device can also be part of a larger system that includes one or more integrated circuits or other processing devices that collectively implement the functionality described herein.
[0123] In the computing device, the processor can include a central processing unit (CPU), a graphics processing unit (GPU), a programmable logic device, a dedicated processor system, a microcontroller, or a microprocessor. By way of example and not limitation, the processor can also include an analog processor, a digital processor, a microprocessor, a multi-core processor, a processor array, a network processor, etc.
[0124] Some of the operations described in the brain-computer interface closed-loop neuromodulation method according to the exemplary embodiments of the present disclosure can be implemented in software, some can be implemented in hardware, and some can be implemented in a combination of software and hardware.
[0125] The processor can execute instructions or code stored in one of the storage components, which can also store data. The instructions and data can also be sent and received over a network via the network interface device, which can employ any known transmission protocol.
[0126] The storage components can be integrated with the processor, such as RAM or flash memory disposed within an integrated circuit microprocessor, etc. Further, the storage components can include separate devices, such as external disk drives, memory arrays, or other storage devices that can be used by any database system. The storage components and the processor can be operatively coupled, or can communicate with each other, for example, through I / O ports, network connections, etc., so that the processor can read files stored in the storage components.
[0127] Further, the computing device can also include a video display (such as a liquid crystal display) and a user interface interface (such as a keyboard, a mouse, a touch input device, etc.). All components of the computing device can be connected via a bus and / or a network.
[0128] The brain-computer interface closed-loop neuromodulation method according to the exemplary embodiments of the present disclosure can be described as various interconnected or coupled functional blocks or functional diagrams. However, these functional blocks or functional diagrams can be equally implemented as a single logic device or operate in non-exact boundaries.
[0129] Therefore, with reference to Figure 1 The brain-computer interface closed-loop neuromodulation method described can be implemented by a system including at least one computing device and at least one storage device storing instructions.
[0130] According to an example embodiment of the present disclosure, the at least one computing device is a computing device for performing the brain-computer interface closed-loop neuromodulation method according to an example embodiment of the present disclosure, and the storage device stores a set of computer executable instructions, which, when executed by the at least one computing device, performs the brain-computer interface closed-loop neuromodulation method according to an example embodiment of the present disclosure. Figure 1 The brain-computer interface closed-loop neuromodulation method is described.
[0131] Figure 4 is a schematic diagram showing a brain-computer interface closed-loop neuromodulation system according to an example embodiment of the present disclosure.
[0132] Optionally, referring to Figure 4 The at least one computing device in the system includes: a first computing device 41 for controlling the operation of the transcranial magnetic stimulation device 42 and the electroencephalogram signal acquisition device 43, which can be used to perform step S101 in Figure 1 , specifically to turn on, turn off, and control the parameters of the transcranial magnetic stimulation device 42 and the electroencephalogram signal acquisition device 43; a second computing device 44 for processing the electroencephalogram signal sequence collected by the electroencephalogram signal acquisition device 43 to obtain a preset index, which can be used to perform steps S102 and S103 in Figure 1 , specifically to receive and synchronously analyze the electroencephalogram signals collected by the electroencephalogram signal acquisition device 43 in real time; and a third computing device 45 for processing the preset index obtained by the second computing device 44 using a deep reinforcement learning network to obtain the stimulation parameters of the next period, which can be used to perform step S104 in Figure 1 and determine whether the end condition is met, specifically to determine the initialized stimulation parameters and continue to generate the iterative stimulation parameters, and also to determine whether to continue the regulation or stop the regulation, and transmit the related instructions to the first computing device 41 to achieve closed-loop neuromodulation.
[0133] Optionally, referring to Figure 4 The system further includes the transcranial magnetic stimulation device 42 and the electroencephalogram signal acquisition device 43 connected between the first computing device 41 and the second computing device 44 to provide a complete connected system. As an example, the transcranial magnetic stimulation device 42 and the electroencephalogram signal acquisition device 43 can be worn by the subject and fixed in a position that cannot be moved. As an example, the transcranial magnetic stimulation device 42 uses a fixed coil, the stimulation frequency range is 0.01 Hz to 50 Hz, the stimulation intensity is 1.5 Tesla to 6 Tesla, and both parameters are continuously adjustable; the number of channels of the electroencephalogram signal acquisition device 43 is greater than 8, and covers five key channels of the PCC brain area (Pz, POz) and the Precuneus brain area (CPz, P3, P4).
[0134] According to exemplary embodiments of the present disclosure, a computer program product including instructions, which, when executed by at least one computing device, cause the at least one computing device to perform a brain-computer interface closed-loop neuromodulation method according to the present disclosure, can also be provided.
[0135] The above describes various exemplary embodiments of the present disclosure, and it should be understood that the above description is merely exemplary and is not exhaustive, and the present disclosure is not limited to the disclosed exemplary embodiments. Many modifications and changes are obvious to those skilled in the art without departing from the scope and spirit of the present disclosure. Therefore, the scope of protection of the present disclosure should be subject to the scope of claims.
Claims
1. A brain-computer interface closed-loop neuromodulation system, comprising: The brain-computer interface closed-loop neuromodulation system comprises at least one computing device and at least one storage device storing instructions, which, when executed by the at least one computing device, cause the at least one computing device to perform the following brain-computer interface closed-loop neuromodulation method: In the formal regulation stage, the formal closed-loop neuromodulation is performed, wherein the formal closed-loop neuromodulation comprises the following steps executed in a first cycle until the end condition is met: Controlling the transcranial magnetic stimulation device to operate for a first duration according to the set stimulation parameters to stimulate the subject; Collecting an electroencephalogram signal sequence of the subject at the target brain region; Processing the electroencephalogram signal sequence to obtain a preset index, wherein the preset index is used to reflect the activity of the target brain region; Using a deep reinforcement learning network to process the preset index to obtain the stimulation parameters of the next cycle; Wherein, the collecting the electroencephalogram signal sequence of the subject at the target brain region comprises: Collecting the electroencephalogram signal sequence of multiple channels of the subject; Wherein, the processing the electroencephalogram signal sequence to obtain a preset index comprises: In the process of collecting the electroencephalogram signal sequence, repeatedly sampling the latest collected electroencephalogram signal sequence of each channel in the multiple channels according to a preset sampling frequency to obtain a signal sub-sequence of each channel at the current time sequence; Compressing the signal sub-sequence of each channel at the current time sequence into a compressed signal sub-sequence; For each channel at the current time sequence, determining a correlation measure between the compressed signal sub-sequence and the compressed signal sub-sequence of the adjacent channel at the current time sequence; Removing the compressed signal sub-sequence with a correlation measure less than a correlation threshold, and adding the remaining compressed signal sub-sequences to a compressed electroencephalogram signal sequence; Stopping collecting the electroencephalogram signal sequence when the time length represented by the compressed electroencephalogram signal sequence reaches a time length threshold; Processing the compressed electroencephalogram signal sequence to obtain a preset index.
2. The brain-computer interface closed-loop neuromodulation system of claim 1, wherein: The multiple signal collection points corresponding to the multiple channels comprise multiple points of the target brain region, the target brain region comprises a default mode network, and the multiple points of the target brain region comprise at least one of the following: a mid-parietal point, a mid-occipital point, a central parietal point, a left parietal point, and a right parietal point.
3. The brain-computer interface closed-loop neuromodulation system of claim 2, wherein: The preset index comprises at least one of the following: an average power spectral density, a frequency band phase synchronization index, and a phase locking value.
4. The brain-machine interface closed-loop neuromodulation system of claim 3, wherein, The compression of the signal sub-sequence of each channel at the current time sequence into a compressed signal sub-sequence comprises: For each channel at the current time sequence, retaining signals less than a preset frequency and compressing signals greater than or equal to the preset frequency to obtain a compressed signal sub-sequence of each channel at the current time sequence.
5. The brain-machine interface closed-loop neuromodulation system of any one of claims 1 to 4, wherein, Before the formal regulation stage, the brain-computer interface closed-loop neuromodulation method further comprises: In the initialization training phase, an initialization closed-loop neuromodulation is performed, wherein the initialization closed-loop neuromodulation comprises performing the following steps in a second cycle repeatedly until an end duration is reached, wherein the duration of the second cycle is less than the duration of the first cycle, and the end duration has the same order of magnitude as the duration of the first cycle: controlling the transcranial magnetic stimulation device to operate for a second duration according to the set stimulation parameters to stimulate the subject; collecting a sequence of electroencephalogram signals of the subject in the target brain region within a third duration, wherein the third duration is greater than the second duration; processing the sequence of electroencephalogram signals to obtain the preset index; processing the preset index using the deep reinforcement learning network to obtain stimulation parameters for a next cycle, wherein, in the process of performing the initialization closed-loop neuromodulation, an experience replay mechanism is used to initialize training of the deep reinforcement learning network, and the experience data used for training includes experience data collected in the process of performing the initialization closed-loop neuromodulation, wherein the deep reinforcement learning network obtained at the end of the initialization closed-loop neuromodulation is used to perform the formal closed-loop neuromodulation, and the stimulation parameters obtained at the end of the initialization closed-loop neuromodulation are used as the stimulation parameters for the initial cycle of the formal closed-loop neuromodulation.
6. The brain-computer interface closed-loop neuromodulation system of claim 5, wherein: the experience data collected when the subject was previously subjected to transcranial magnetic stimulation is denoted as historical experience data, wherein the experience data used for training further includes the historical experience data; or wherein the deep reinforcement learning network based on which the initialization training is performed is a deep reinforcement learning network trained using the historical experience data.
7. The brain-computer interface closed-loop neuromodulation system of claim 5, wherein: the brain-computer interface closed-loop neuromodulation method further comprises: comparing the preset index with a preset value range to obtain modulation evaluation information, wherein the preset value range is obtained by statistically analyzing the values of the preset index of a target population, and the modulation evaluation information is used to represent an evaluation of the effect of neuromodulation; and / or in the process of performing the initialization training on the deep reinforcement learning network, an instant reward of deep reinforcement learning is determined according to the preset index and the preset value range, wherein the preset value range is obtained by statistically analyzing the values of the preset index of a target population.
Citation Information
Patent Citations
Transcranial magnetic stimulation system with electroencephalogram signal feedback regulation function
CN221998650U