A method for optimizing a child magnetoencephalogram individualized stimulation paradigm based on reinforcement learning

By dynamically adjusting the magnetoencephalography (MEG) stimulation paradigm using a reinforcement learning-based approach, and by constructing a real-time state feature vector using a light-pumped magnetometer array to optimize the MEG stimulation paradigm, the problem of poor MEG data quality in children was solved, and the accuracy and efficiency of epileptic focus localization were improved.

CN121400841BActive Publication Date: 2026-04-10HANGZHOU ZERO MAGNETIC MEDICAL EQUIPMENT CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
HANGZHOU ZERO MAGNETIC MEDICAL EQUIPMENT CO LTD
Filing Date
2025-12-25
Publication Date
2026-04-10

AI Technical Summary

Technical Problem

Existing magnetoencephalography (MEG) stimulation paradigms are ill-suited to individual differences among children, resulting in poor data quality and low experimental efficiency, and failing to achieve personalized assessments.

Method used

A reinforcement learning-based approach was adopted to collect multi-channel magnetoencephalogram (MEG) signals, behavioral data, and physiological data through an optically pumped magnetometer array. Real-time state feature vectors were constructed, and the parameters of the MEG stimulation paradigm were dynamically adjusted. The results were then iteratively optimized using a reinforcement learning model.

Benefits of technology

It improved the quality of multi-channel magnetic brain signals, enhanced the accuracy and efficiency of epileptic focus localization, reduced artifacts, and improved children's cooperation and data quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121400841B_ABST
    Figure CN121400841B_ABST
Patent Text Reader

Abstract

The application provides a kind of child magnetoencephalogram personalized stimulation paradigm optimization method based on reinforcement learning, comprising: step S1, the multi-channel magnetoencephalogram signal of the child to be measured is continuously collected by using optically pumped magnetometer array, and behavior data and physiological data are synchronously collected;Step S2, the magnetoencephalogram signal feature, behavior feature and physiological feature are extracted, and real-time state feature vector is fused and constructed;Step S3, the real-time state feature vector is input into the reinforcement learning model to obtain the prediction Q value of stimulation adjustment action, and the parameters of magnetoencephalogram stimulation paradigm of the child to be measured are dynamically adjusted based on the prediction Q value;Step S4, based on the magnetoencephalogram stimulation paradigm after parameter adjustment, the new multi-channel magnetoencephalogram signal, behavior data and physiological data of the child to be measured are collected, and the comprehensive reward value is calculated to update the reinforcement learning model.The beneficial effect is that the application can intelligently and dynamically adjust the magnetoencephalogram stimulation paradigm, and improve the multi-channel magnetoencephalogram signal quality of the child to be measured.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of biomedical engineering, and in particular, to a method for optimizing a child magnetoencephalogram (MEG) individualized stimulation paradigm based on reinforcement learning. BACKGROUND

[0002] Magnetoencephalography (MEG) is a non-invasive neurophysiological technique that reflects brain function by measuring the weak magnetic fields generated by neuronal activity in the brain. Compared with electroencephalography (EEG), MEG is not affected by skull and scalp conductivity, has higher spatial resolution, and maintains millisecond-level temporal resolution, making it unique in brain function research and clinical diagnosis, especially in the localization of epileptic lesions.

[0003] Epilepsy is a common neurological disease in children, with a high incidence. Some children with epilepsy are refractory to medication and need to undergo surgery to remove the epileptogenic focus to control the disease. Accurate localization of the epileptogenic focus is crucial for the success of epilepsy surgery. Child MEG (MEG) can effectively capture the origin of epileptic discharges due to its high temporal and spatial resolution, and is widely used in preoperative assessment of epilepsy.

[0004] In recent years, with the development of optical pumping magnetometers (OPM), not only has the sensitivity and portability made great breakthroughs, but also the cost of manufacture and maintenance is lower, which makes up for the shortcomings of traditional MEG devices based on superconducting quantum interference devices, such as high operating cost and long distance between probe and measured object. The wearable feature of OPM-MEG also allows children to move their heads to some extent during MEG scanning, improving their compliance and comfort.

[0005] Although OPM-MEG brings new opportunities for the diagnosis of childhood epilepsy, how to fully utilize the advantages of OPM-MEG, solve the data quality problems caused by individual differences such as short attention span, easy distraction, and poor cooperation during the data acquisition process of child MEG, and dynamically optimize the stimulation paradigm to maximize effective signals and minimize artifacts, are still challenges. The existing MEG stimulation paradigms are mostly fixed and preset, which cannot adapt to the significant individual differences in cognitive ability, attention span, emotional state, and brain development stage of children, resulting in poor data quality, low experimental efficiency, and inability to achieve truly personalized assessment. Therefore, there is an urgent need for a method that can intelligently and dynamically adjust the MEG stimulation paradigm to improve the quality of child MEG signals and thus enhance the accuracy and efficiency of epileptic lesion localization. SUMMARY

[0006] The technical problem to be solved by the present application is how to intelligently and dynamically adjust the magnetoencephalogram stimulation paradigm to improve the multi-channel magnetoencephalogram signal quality of the child to be tested.

[0007] The present application provides a child magnetoencephalogram personalized stimulation paradigm optimization method based on reinforcement learning.

[0008] Step S1, continuously acquiring multi-channel magnetoencephalogram signals of the child to be tested by using the optical pumping magnetometer array during the execution of the emotional picture cognitive task by the child to be tested, and synchronously acquiring behavioral data and physiological data of the child to be tested;

[0009] Step S2, extracting magnetoencephalogram signal features, behavioral features and physiological features from the multi-channel magnetoencephalogram signals, the behavioral data and the physiological data respectively, and fusing to obtain a real-time state feature vector;

[0010] Step S3, inputting the real-time state feature vector into a pre-constructed reinforcement learning model to obtain a predicted Q value of a stimulation adjustment action, and dynamically adjusting parameters of the magnetoencephalogram stimulation paradigm of the child to be tested based on the predicted Q value of the stimulation adjustment action;

[0011] Step S4, acquiring new multi-channel magnetoencephalogram signals, behavioral data and physiological data of the child to be tested based on the magnetoencephalogram stimulation paradigm after parameter adjustment, and calculating a comprehensive reward value to update the reinforcement learning model, and then returning to step S2.

[0012] Compared with the prior art, the child magnetoencephalogram personalized stimulation paradigm optimization method based on reinforcement learning has the following advantages:

[0013] In the present application, the multi-channel magnetoencephalogram signals, behavior data and physiological data are collected through step S1, the feature extraction is performed and the real-time state feature vector is constructed through step S2, the parameter adjustment of the magnetoencephalogram stimulation paradigm is performed through step S3, the data re-collection and reinforcement learning model updating are performed through step S4, the real-time state feature vector composed of the magnetoencephalogram signal features, behavior features and physiological features is predicted and analyzed through the reinforcement learning model in the whole process, the real-time and individualized magnetoencephalogram stimulation paradigm adjustment is performed to reduce the artifacts caused by the non-compliance or distraction of the children to be tested, and then the new multi-channel magnetoencephalogram signals of higher quality are obtained, finally, the reinforcement learning model is iteratively updated combined with the new behavior data and physiological data to make it have higher prediction Q value accuracy, and the magnetoencephalogram stimulation paradigm is adjusted through continuous cycle optimization to improve the quality of the multi-channel magnetoencephalogram signals.

[0014] In a possible implementation, the optically pumped magnetometer array includes a plurality of extremely weak magnetic measurement sensors based on the spin-exchange relaxation effect arranged in an array, and in step S1, each of the extremely weak magnetic measurement sensors continuously collects the multi-channel magnetoencephalogram signals of the children to be tested.

[0015] Compared with the prior art, the above technical solution can utilize the characteristic of the spin-exchange relaxation effect with extremely high magnetic field sensitivity, is suitable for the weak magnetic field detection of the multi-channel magnetoencephalogram signals of the children to be tested, and can improve the collection accuracy of the multi-channel magnetoencephalogram signals.

[0016] In a possible implementation, in step S2, the multi-channel magnetoencephalogram signals are subjected to 0.1-100 Hz band-pass filtering, independent component analysis artifact removal and head motion compensation to obtain the signal-to-noise ratio of the magnetoencephalogram signals and the number of artifacts as the magnetoencephalogram signal features, and the task accuracy of the emotional picture cognitive task, the root mean square value of the three-dimensional displacement of the head, the emotional positivity score and the heart rate variability coefficient are extracted from the behavior data and the physiological data as the behavior features and the physiological features.

[0017] Compared with the prior art, the above technical solution can effectively remove noise and artifacts through the independent component analysis artifact removal and motion compensation technology, and significantly improve the purity of the magnetoencephalogram signal features; at the same time, the signal quality (signal-to-noise ratio of the magnetoencephalogram signals and the number of artifacts), the behavior performance (task accuracy) and the physiological state (emotional positivity score and heart rate variability coefficient) are fused as multi-dimensional state representation to comprehensively reflect the real-time state of the children to be tested.

[0018] In a possible implementation, in the step S3, a deep Q network is used as the reinforcement learning model, the deep Q network comprises an input layer, a plurality of hidden layers and an output layer connected in sequence, the real-time state feature vector is received through the input layer, and the predicted Q value of each stimulation adjustment action in a discrete action space is output through the output layer.

[0019] Compared with the prior art, the above technical solution can introduce a deep Q network as a reinforcement learning model, is suitable for complex multi-modal feature input, can effectively learn a state-action mapping relationship, select an optimal stimulation adjustment action through Q value prediction, and realize dynamic adaptive magnetoencephalogram stimulation paradigm optimization.

[0020] In a possible implementation, a rectified linear unit is used as an activation function of each hidden layer of the deep Q network, an Adam optimizer is used for network training, a learning rate is set to 0.001, a discount factor γ = 0.95, an experience replay mechanism is used, an experience replay pool size is 10000, a batch size is 64, and a target network update frequency is 100 steps.

[0021] Compared with the prior art, the above technical solution can reasonably set a learning rate, a discount factor, a batch size and other hyperparameters, and improve training efficiency.

[0022] In a possible implementation, parameters of the magnetoencephalogram stimulation paradigm in the step S3 include stimulation intensity, stimulation duration, stimulation presentation interval, task difficulty and rest interval, and each stimulation adjustment action in the discrete action space includes increasing stimulation intensity, decreasing stimulation intensity, shortening stimulation duration, lengthening stimulation duration, shortening stimulation presentation interval, lengthening stimulation presentation interval, increasing rest interval, shortening rest interval, pausing an experiment and ending an experiment.

[0023] Compared with the prior art, the above technical solution clearly defines parameters of the magnetoencephalogram stimulation paradigm and specific contents of the stimulation adjustment action, can flexibly adjust stimulation properties, and is suitable for different experimental scenarios.

[0024] In a possible implementation, the multi-channel magnetoencephalogram signal includes a magnetoencephalogram signal signal-to-noise ratio and a number of artifacts, the behavior data is a task accuracy rate, the physiological data is an emotion positivity score, and in the step S4, the comprehensive reward value is obtained through the following calculation formula:

[0025] ;

[0026] wherein,

[0027] represents the comprehensive reward value.

[0028] represents a first weight coefficient;

[0029] represents a reward value positively correlated with the signal-to-noise ratio of the brain magnetic signal;

[0030] represents a second weight coefficient;

[0031] represents a penalty value positively correlated with the number of artifacts;

[0032] represents a third weight coefficient;

[0033] represents a reward value positively correlated with the task accuracy;

[0034] represents a fourth weight coefficient;

[0035] represents a reward value positively correlated with the emotional positivity score.

[0036] Compared with the prior art, the above technical solution can effectively guide the reinforcement learning model to learn in the direction of improving data quality and child experience, balance multiple optimization goals such as signal quality, task performance, and child cooperation degree, and by adjusting four weight coefficients, different experimental goals can be optimized and focused on.

[0037] In a possible implementation, the step of updating the reinforcement learning model in the step S4 includes:

[0038] Step A1, calculating the comprehensive reward value of the current time step, integrating the complete interaction information of the current time step into an experience tuple and storing it in an experience replay pool, wherein, represents an action of performing stimulus adjustment the real-time state feature vector before, represents the stimulus adjustment action selected and performed by the reinforcement learning model, represents an action of performing stimulus adjustment the new comprehensive reward value after, represents an action of performing stimulus adjustment the new real-time state feature vector of the child to be tested after,

[0039] Step A2, when the amount of data in the experience replay pool reaches a preset threshold, a small batch of experience tuples in the experience replay pool is randomly drawn for the reinforcement learning model to learn, a loss function is defined by calculating the difference between the target Q value and the predicted Q value, then the Adam optimizer is used to calculate the gradient of the loss function with respect to the parameters in the reinforcement learning model through the backpropagation algorithm, and the weights of the reinforcement learning model are updated along the gradient descent direction, so that the predicted Q value approximates the target Q value.

[0040] Compared with the prior art, the above technical solution describes the updating mechanism of the reinforcement learning model, breaks the data correlation by using the experience replay mechanism, improves the sample utilization rate and training stability, forms a closed-loop optimization of "state-action-reward-update", and realizes continuous self-improvement. BRIEF DESCRIPTION OF DRAWINGS

[0041] Figure 1 The step flowchart of the present application;

[0042] Figure 2 The optimization flowchart of the magnetoencephalogram stimulation paradigm of the present application;

[0043] Figure 3 The interaction schematic diagram of the reinforcement learning model and the environment of the present application. DETAILED DESCRIPTION

[0044] First of all, those skilled in the art should understand that these embodiments are only used to explain the technical principles of the embodiments of the present application, and are not intended to limit the protection scope of the embodiments of the present application. Those skilled in the art can adjust them as needed in order to adapt to specific application occasions.

[0045] The present application will be further described in detail below in combination with the drawings and specific embodiments.

[0046] Reference Figure 1 The embodiments of the present application disclose a child magnetoencephalogram personalized stimulation paradigm optimization method based on reinforcement learning, which pre-wears an optically pumped magnetometer array on a child to be tested, and makes the child to be tested observe positive, neutral and negative three kinds of emotional pictures alternately presented on a screen, and press the corresponding keys after each emotional picture is displayed to classify and perform an emotional picture cognitive task. The child magnetoencephalogram personalized stimulation paradigm optimization method comprises the following steps:

[0047] Step S1, continuously acquiring multi-channel magnetoencephalogram signals of the child to be tested by using the optically pumped magnetometer array during the execution of the emotional picture cognitive task by the child to be tested, and synchronously acquiring behavior data and physiological data of the child to be tested;

[0048] Step S2, brain magnetic signal features, behavior features and physiological features are extracted from the multi-channel brain magnetic signal, behavior data and physiological data respectively, and a real-time state feature vector is constructed by fusion;

[0049] Step S3, the real-time state feature vector is input into the pre-constructed reinforcement learning model to obtain a predicted Q value of the stimulation adjustment action, and the parameters of the magnetoencephalogram stimulation paradigm of the child to be tested are dynamically adjusted based on the predicted Q value of the stimulation adjustment action;

[0050] Step S4, new multi-channel brain magnetic signals, behavior data and physiological data of the child to be tested are collected based on the parameter-adjusted magnetoencephalogram stimulation paradigm, and a comprehensive reward value is calculated to update the reinforcement learning model, and then the step S2 is returned.

[0051] This embodiment details how to apply the method of the present application to the magnetoencephalogram (OPM-MEG) data collection of the child to be tested when performing an emotional picture cognitive task, and dynamically optimize the individualized stimulation paradigm;

[0052] In this embodiment, when the child to be tested participates in the emotional picture cognitive task, the task requires the child to watch three types of emotional pictures, i.e. positive, neutral and negative, presented alternately on the screen, and to classify them by pressing the corresponding keys after each emotional picture is displayed, see Figure 2 and Figure 3 , and the specific implementation process is as follows:

[0053] Experimental setup and data collection

[0054] Multi-channel brain magnetic signal collection: an optical pumping magnetometer (OPM) array is used, which is composed of multiple extremely weak magnetic measurement sensors based on the non-spin exchange relaxation effect, for continuously collecting multi-channel brain magnetic signals of the child during the task;

[0055] Synchronous acquisition of behavior and physiological data: the task response data (key accuracy and reaction time), head motion data, facial expression data and physiological data of the child to be tested are synchronously collected;

[0056] The head motion data is collected by a three-axis acceleration sensor with a sampling frequency of 200 Hz and a measurement range of ±10° rotation angle and ±5 cm displacement; the facial expression data is collected by a camera, and the MTCNN+FER-2013 model is used to calculate the emotional state;

[0057] Construction of real-time state feature vector

[0058] The original multi-channel brain magnetic signal is applied with a 4th order Butterworth band-pass filter of 0.1-100 Hz to remove low-frequency drift and high-frequency noise, and a FastICA algorithm is used for independent component analysis to automatically identify and remove physiological artifacts such as eye movement, blinking, heartbeat, etc. At the same time, a signal space projection (SSP) algorithm is used in combination with head motion sensor data for motion artifact correction.

[0059] The brain magnetic signal features and behavior / physiological features are extracted, specifically, the brain magnetic signal features include extracting the signal-to-noise ratio (SNR) of the event-related magnetic field brain magnetic signal induced by the picture stimulation and the number of artifacts in the brain magnetic signal, and the behavior / physiological features calculate the real-time accuracy and reaction time of the emotional picture cognitive task; the facial expression recognition algorithm is used to analyze the emotional state of the child in real time, and quantified as an emotional positivity score to evaluate the impact of the stimulus on the child's emotion; the root mean square value of the three-dimensional displacement of the head is calculated to reflect the intensity of the head movement, and a smaller RMS value indicates that the child has a high degree of cooperation and a stable head; the heart rate variability (HRV) is calculated to reflect the balance state of the autonomic nervous system, which is related to the stress and emotional state of the child; the above features are standardized and fused into a real-time state feature vector State as the input of the reinforcement learning model;

[0060] Action selection and execution based on DQN

[0061] In this embodiment, a deep Q network (DQN) is used as the reinforcement learning model, and the specific network structure and training parameter settings are as follows:

[0062] The deep Q network includes 1 input layer, 3 hidden layers and 1 output layer, the dimension of the input layer matches the dimension of the state feature vector, the hidden layer is composed of three fully connected layers, the number of neurons is 512, 256 and 128 respectively, and ReLU activation function is used to increase the non-linear expression ability of the deep Q network, the number of neurons in the output layer is the same as the number of stimulus adjustment actions in the discrete action space, and the predicted Q value corresponding to each stimulus adjustment action is output, the Adam optimizer is used to train the deep Q network, the learning rate is set to 0.001, β1=0.9, β2=0.999, in order to break the data correlation and stabilize the training process, the deep Q network uses the experience replay mechanism, the size of the experience replay pool is set to 10000, 64 samples (batch size) are randomly selected from the experience replay pool for training and updating each time, and the update is performed every 100 steps;

[0063] The input of the reinforcement learning model is the real-time state feature vector constructed in step S2, and the discrete action space is designed as a series of discrete adjustment operations on the magnetoencephalogram stimulation paradigm, mainly including:

[0064] (1) Lengthen / shorten the display time of the next emotional picture;

[0065] (2) Lengthen / shorten the interval time of emotional picture presentation;

[0066] (3) Add a short break;

[0067] (4) Adjust the emotional type of the next presented emotional picture (for example, when the child shows negative emotions, prefer to present positive or neutral emotional pictures);

[0068] The reinforcement learning model outputs the predicted Q value of each possible stimulus adjustment action according to the current real-time state feature vector, and selects the stimulus adjustment action with the highest predicted Q value to adjust the emotional picture stimulus parameters in the next round;

[0069] Reward calculation and model update

[0070] After the reinforcement learning model performs a stimulus adjustment action, a comprehensive reward value is calculated according to the newly collected multi-channel magnetoencephalogram, behavior data and physiological data, which is used to evaluate the pros and cons of the executed stimulus adjustment action. The function formula of the comprehensive reward value is as follows:

[0071] ;

[0072] Wherein, represents the comprehensive reward value; represents the first weight coefficient; represents the reward value positively correlated with the signal-to-noise ratio of magnetoencephalogram; represents the second weight coefficient; represents the penalty value positively correlated with the number of artifacts; represents the third weight coefficient; represents the reward value positively correlated with the task accuracy rate; represents the fourth weight coefficient; represents the reward value positively correlated with the emotional positivity score;

[0073] Specifically, SNR_reward = (SNR_current-10) / 10, with a baseline SNR of 10 dB; Artifact_penalty = Artifact_rate / 0.2, with a baseline artifact rate of 20%; Task_Completion_Rate = 0.7 x (N_correct / N_total) + 0.3 x (1 - T_response / 2), where N_correct is the number of times the corresponding correct key is pressed, N_total is the total number of times the key is pressed, and T_response is the response time in seconds; and Compliance_reward represents the degree of cooperation score of the child. After calculating the comprehensive reward value Reward of the current time step, the complete interaction information of the current time step is integrated into an experience tuple and stored in the experience replay pool, wherein, represents the execution of the stimulus adjustment action the previous real-time state feature vector, represents the stimulus adjustment action selected and executed by the reinforcement learning model, represents the execution of the stimulus adjustment action the new comprehensive reward value, represents the execution of the stimulus adjustment action the new real-time state feature vector of the child to be tested after,

[0074] When the amount of data in the experience replay pool reaches a certain size, the reinforcement learning model randomly selects a small batch (mini-batch) of experience tuples from the experience replay pool for learning. The calculation formula of the target Q value is:

[0075] ;

[0076] the predicted Q value is the real-time state feature vector the Q value after the execution of the stimulus adjustment action , that is, The loss function is defined by calculating the difference between the target Q value and the predicted Q value:

[0077] ;

[0078] Subsequently, the reinforcement learning model uses the Adam optimizer to calculate the gradient of the loss function with respect to the main network parameters in the reinforcement learning model through the backpropagation algorithm, and updates the weights of the main network in the direction of gradient descent, so that the predicted Q value gradually approaches the target Q value.

[0079] To further increase the stability of training, the parameters of the main network are not updated at every step, but every fixed number of steps, the main network parameters are completely copied to the target network, this delay update mechanism can effectively prevent the oscillation and divergence in the training process.

[0080] This embodiment illustrates how to apply the method proposed in the application to individualize the optimization of the flash stimulation paradigm when performing magnetoencephalogram examination to induce epileptiform discharges in children with epilepsy. This scenario applies the application to clinical diagnosis, especially to epilepsy lesion localization.

[0081] The first two steps are the same in the specific implementation process, in the third step, the input of the reinforcement learning model is the real-time state feature vector constructed in step S2, the discrete action space is designed as a series of discrete adjustment operations on the magnetoencephalogram stimulation paradigm, aiming to dynamically adapt to the real-time state of children with epilepsy, improve the efficiency of inducing epileptiform discharges and the quality of data, mainly including:

[0082] (1) Increase / decrease the brightness level of flash stimulation;

[0083] (2) Increase / decrease the frequency of flash stimulation;

[0084] (3) Extend / shorten the duration of a single flash stimulation;

[0085] (4) Insert a non-stimulus rest interval;

[0086] The reinforcement learning model outputs the predicted Q value of each possible stimulation adjustment action according to the current real-time state feature vector, and selects the stimulation adjustment action with the highest predicted Q value to adjust the picture stimulation parameters in the next round;

[0087] After the reinforcement learning model executes a stimulation adjustment action, a comprehensive reward value is calculated according to the newly acquired multi-channel magnetoencephalogram signal, behavior data and physiological data, which is used to evaluate the pros and cons of the executed stimulation adjustment action, and the function formula of the comprehensive reward value is as follows:

[0088] ;

[0089] Since the flash stimulation paradigm has no behavioral feedback, the function formula of the comprehensive reward value does not contain the task completion related component, wherein, SNR_reward = (SNR_current - 10) / 10, the baseline SNR is set to 10dB; Artifact_penalty = Artifact_rate / 0.2, the baseline artifact rate is set to 20%; Compliance_reward represents the cooperation score of the child to be tested; after calculating the comprehensive reward value Reward of the current time step, the complete interaction information of the current time step is integrated into an experience tuple and stored in the experience replay pool, when the amount of data in the experience replay pool reaches a certain size, a small batch (mini-batch) of experience tuples is randomly extracted from the pool for learning, and the subsequent steps are the same as in embodiment one.

[0090] In the description of the present application, the description of the terms "one embodiment", "some embodiments", "in this embodiment", "specific examples" or "some examples" means that the specific features, mechanisms, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present application. In the present specification, the illustrative description of the above terms does not necessarily refer to the same embodiment or example. Moreover, the specific features, mechanisms, materials or characteristics described can be combined in any appropriate manner in any one or more embodiments or examples. In addition, different embodiments or examples described in the specification and the features of different embodiments or examples can be combined and combined by those skilled in the art without contradiction.

[0091] The above is only a specific embodiment of the present application, but the protection scope of the present application is not limited thereto, any person skilled in the art can easily think of changes or replacements within the technical range disclosed by the present application, which should be covered within the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

Claims

1. A method for optimizing personalized stimulation paradigms based on reinforcement learning in children's magnetoencephalography (MEG), characterized in that, The children to be tested are first fitted with an optically pumped magnetometer array, and then observe alternating images of positive, neutral, and negative emotions on a screen. After each emotion image is displayed, the children press the corresponding button to classify it, thus performing an emotion image recognition task. The personalized stimulation paradigm optimization method for children's magnetoencephalography includes the following steps: Step S1: During the process of the child performing the emotion picture recognition task, the optical pump magnetometer array continuously collects the multi-channel brain magnetic signals of the child and simultaneously collects the child's behavioral and physiological data. Step S2: Extract the magnetic brain signal features, behavioral features, and physiological features from the multi-channel magnetic brain signal, the behavioral data, and the physiological data, respectively, and fuse them to construct a real-time state feature vector; Step S3: Input the real-time state feature vector into the pre-built reinforcement learning model to obtain the predicted Q value of the stimulus adjustment action, and dynamically adjust the parameters of the magnetoencephalogram stimulation paradigm of the child to be tested based on the predicted Q value of the stimulus adjustment action. Step S4: Based on the adjusted magnetoencephalogram (MEG) stimulation paradigm, new multi-channel MEG signals, behavioral data, and physiological data of the child to be tested are collected, and a comprehensive reward value is calculated to update the reinforcement learning model. Then, the process returns to step S2. Among these, the characteristics of the magnetic field signal include the signal-to-noise ratio (SNR) of the magnetic field signal induced by the picture stimulus and the number of artifacts in the magnetic field signal; behavioral / physiological characteristics are used to calculate the real-time accuracy and reaction time of the emotional picture recognition task; facial expression recognition algorithm is used to analyze the child's emotional state in real time and quantify it into an emotional positivity score to assess the impact of the stimulus on the child's emotions; the root mean square value of the three-dimensional head displacement is calculated to reflect the intensity of head movement, and a smaller RMS value indicates that the child has high cooperation and head stability; the heart rate variability coefficient (HRV) is calculated to reflect the balance of the autonomic nervous system and is related to the child's stress and emotional state.

2. The method for optimizing personalized stimulation paradigms of children's magnetoencephalography based on reinforcement learning according to claim 1, characterized in that, The optically pumped magnetometer array includes multiple extremely weak magnetic measurement sensors arranged in an array based on the spin-free exchange relaxation effect. In step S1, each of the extremely weak magnetic measurement sensors is controlled to continuously acquire the multi-channel magnetoencephalogram (MEG) signals of the child to be tested.

3. The method for optimizing personalized stimulation paradigms based on reinforcement learning in children's magnetoencephalography (MEG) according to claim 1, characterized in that, In step S2, the multi-channel magnetoencephalography (MEG) signal is subjected to 0.1-100Hz bandpass filtering, independent component analysis for artifact removal, and head motion compensation to obtain the MEG signal-to-noise ratio and the number of artifacts as MEG signal features. The task accuracy, root mean square value of head three-dimensional displacement, emotional positivity score, and heart rate variability coefficient of the emotion picture cognition task are extracted from the behavioral data and the physiological data as the behavioral features and the physiological features.

4. The method for optimizing personalized stimulation paradigms of children's magnetoencephalography based on reinforcement learning according to claim 1, characterized in that, In step S3, a deep Q-network is used as the reinforcement learning model. The deep Q-network includes an input layer, multiple hidden layers and an output layer connected in sequence. The input layer receives the real-time state feature vector, and the output layer outputs the predicted Q value of each stimulus adjustment action in the discrete action space.

5. The method for optimizing personalized stimulation paradigms of children's magnetoencephalography based on reinforcement learning according to claim 4, characterized in that, The hidden layers of the deep Q-network use modified linear units as activation functions. The network is trained using the Adam optimizer with a learning rate of 0.001, a discount factor γ of 0.95, and an experience replay mechanism. The experience replay pool size is 10,000, the batch size is 64, and the target network update frequency is 100 steps.

6. The method for optimizing personalized stimulation paradigms of children's magnetoencephalography based on reinforcement learning according to claim 4, characterized in that, The parameters of the magnetoencephalogram stimulation paradigm in step S3 include stimulation intensity, stimulation duration, stimulation presentation interval, task difficulty, and rest interval. The stimulation adjustment actions in the discrete action space include increasing stimulation intensity, decreasing stimulation intensity, shortening stimulation duration, extending stimulation duration, shortening stimulation presentation interval, extending stimulation presentation interval, increasing rest interval, shortening rest interval, pausing experiment, and ending experiment.

7. The method for optimizing personalized stimulation paradigms of children's magnetoencephalography based on reinforcement learning according to claim 1, characterized in that, The multi-channel magnetoencephalogram (MEG) signal includes the MEG signal-to-noise ratio and the number of artifacts; the behavioral data is the task accuracy; the physiological data is the emotional positivity score; and in step S4, the comprehensive reward value is obtained using the following formula: , in, This represents the overall reward value; Indicates the first weighting coefficient; This represents a reward value that is positively correlated with the signal-to-noise ratio of the brain magnetic signal; This represents the second weighting coefficient; This represents a penalty value that is positively correlated with the number of artifacts. This represents the third weighting coefficient; This represents a reward value that is positively correlated with the accuracy of the task. This represents the fourth weighting coefficient; The reward value represents the positive correlation with the emotional positivity score.

8. The method for optimizing personalized stimulation paradigms of children's magnetoencephalography based on reinforcement learning according to claim 1, characterized in that, The step of updating the reinforcement learning model in step S4 includes: Step A1: Calculate the comprehensive reward value for the current time step, and integrate the complete interaction information of the current time step into an experience tuple. And store it in the experience replay pool, where, Indicates the execution of stimulus adjustment actions The previously mentioned real-time state feature vector, This refers to the stimulus adjustment action selected and executed by the reinforcement learning model. Indicates the execution of stimulus adjustment actions The new comprehensive reward value, Indicates the execution of stimulus adjustment actions Then, the new real-time state feature vector of the child to be tested; Step A2: When the amount of data in the experience replay pool reaches a preset threshold, a small batch of experience tuples is randomly extracted from the experience replay pool for the reinforcement learning model to learn. The loss function is defined by calculating the difference between the target Q value and the predicted Q value. Then, the Adam optimizer is used to calculate the gradient of the loss function with respect to the intrinsic parameters of the reinforcement learning model through the backpropagation algorithm, and the weights of the reinforcement learning model are updated along the gradient descent direction so that the predicted Q value approaches the target Q value.

Citation Information

Patent Citations

  • Brain wave induction method based on deep reinforcement learning and storage medium

    CN115462806A

  • Three-dimensional brain operation planning method

    CN117770953A