Training method, system and equipment for improving execution capability

By collecting and fusing prefrontal EEG, task response, and eye movement data from children with ADHD, training tasks can be adjusted in real time, solving the problems of inaccurate assessment and poor effectiveness of existing training methods, and achieving more efficient executive function training.

CN120932819APending Publication Date: 2025-11-11SHANGHAI SHUZHIYAO INTELLIGENT TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510982580.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-16
Publication Date
2025-11-11

AI Technical Summary

Technical Problem

Existing methods for training children with ADHD in their executive function have problems such as inaccurate assessment results and poor training effects. This is mainly because the brain region information extracted from EEG signals contains a lot of irrelevant information, resulting in large errors in assessment results and a lack of targeted training.

Method used

The system collects prefrontal EEG data, task execution response data, and eye movement data from users when they are performing tasks. By determining power feature maps, response feature maps, and fixation point heatmaps, it performs synchronous sliding interception and fusion, inputs the data into the performance assessment model, and adjusts the task stimulus parameters and difficulty level in real time to achieve dual feedback training within and between tasks.

Benefits of technology

It improves the targeting of training and the accuracy of indicators of brain execution ability, resulting in better training effects. It also dynamically adjusts training tasks to enhance concentration and execution response capabilities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120932819A_ABST
    Figure CN120932819A_ABST
Patent Text Reader

Abstract

The invention provides a training method, system and device for improving the execution ability. The method comprises the steps that forehead lobe electroencephalogram data, task execution response data and eye movement data are collected; according to the prefrontal lobe electroencephalogram data, the task execution response data and the eye movement data, a power feature map, a response feature map and a fixation point thermodynamic map are determined respectively; processing the power feature map, the response feature map and the fixation point thermodynamic map to obtain a corresponding same-time map segment sequence, further processing the same-time map segment sequence to obtain a plurality of single-mode images, and fusing the single-mode images to obtain a multi-mode image; inputting the multi-modal image into an execution capability evaluation model to obtain an attention index, an execution response index and an eye movement fixation index; and according to the attention index, the eye movement fixation index and a preset rule, adjusting a visual stimulation parameter and / or an auditory stimulation parameter of the current task, judging an execution response index, and adjusting a task type and / or a difficulty level of the next task so as to realize real-time double feedback in the tasks and between the tasks.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of rehabilitation training technology, and more particularly to a training technique for improving executive function. Background Technology

[0002] ADHD (Attention Deficit Hyperactivity Disorder) is a common childhood mental disorder. Globally, the prevalence of ADHD is approximately 5%-10%. Affected children (or those with ADHD) suffer from inattention, hyperactivity, and / or impulsive behavior, impacting their learning, social interactions, and daily lives. Extensive research in neuroscience and cognitive psychology has confirmed that the neural maturation of the prefrontal cortex affects the development of the brain's executive function systems, influencing multiple subsystems such as attention control, impulse inhibition, cognitive flexibility, short-term working memory, and attention maintenance. This results in affected children exhibiting significant developmental delays or functional deficits compared to their peers.

[0003] Typically, training relies on face-to-face interaction between professional medical staff and sick children, or guiding them to use hand-controlled games to train their continuous attention, stimulating corresponding areas of the brain and providing personalized attention to improve their executive function and sustained focus. However, existing training methods have many drawbacks. For example, the availability of hand-controlled game equipment is limited, and the game mechanics are not adjusted according to the individual needs of each child. Children often show little interest in the games themselves. The limited number of medical staff makes it difficult to continuously monitor a large number of children over extended periods. Furthermore, the psychological sensitivity of sick children may hinder their recovery and reduce the effectiveness of improving their executive function.

[0004] In existing technologies, some methods involve collecting whole-brain electroencephalogram (EEG) signals from affected children, extracting a large amount of brain region information, analyzing and assessing the children's brain activity and cognitive state, and then having medical staff conduct interactive training based on the assessment results. However, because the large amount of brain region information obtained after collecting general EEG signals contains much irrelevant information that is not strongly correlated with the brain's executive function system, the assessment results have significant errors, leading to a lack of targeted interactive training by medical staff based on the assessment results.

[0005] Therefore, existing methods for improving the executive function of children with ADHD suffer from technical problems such as inaccurate assessment results and poor training effectiveness due to a lack of interaction between assessment results and training. Summary of the Invention

[0006] The purpose of this application is to provide a training method, system, and device for improving executive function, so as to at least partially solve the technical problems of inaccurate assessment results and poor training effects of existing methods for training the brain executive function of children with ADHD.

[0007] According to one aspect of this application, a training method for improving execution capabilities is provided, wherein the method includes:

[0008] When performing executive function training, the prefrontal cortex EEG data, task execution response data and eye movement data of the user are collected when processing each task. The content of the task is determined according to the task parameters, which include at least: task type, difficulty level, visual stimulus parameters and auditory stimulus parameters.

[0009] Based on the prefrontal EEG data, a power feature map is determined; based on the task execution response data, a response feature map is determined; and based on the eye movement data and screen parameters, a fixation heatmap is determined.

[0010] The power feature map, the response feature map, and the gaze point heatmap are simultaneously slidably truncated to obtain corresponding time-series image segments. All time-series image segments are then processed to obtain multiple single-modal images. Finally, the multiple single-modal images are fused to obtain a multimodal image.

[0011] The multimodal images are input into the execution capability assessment model to obtain attention concentration index, execution response index, and eye movement fixation concentration index corresponding to the multimodal images;

[0012] Based on the attention concentration index, the eye-tracking fixation concentration index, and preset rules, the visual and / or auditory stimulus parameters of the current task are adjusted in real time, and it is determined whether the execution response index meets the execution response index threshold. If it does not meet the threshold, the task type and / or difficulty level of the next task are adjusted to achieve real-time dual feedback execution ability training within and between tasks.

[0013] Preferably, the acquisition of prefrontal cortex EEG data, task execution response data, and eye movement data during user processing of each task includes:

[0014] When the user processes each task

[0015] The user's prefrontal EEG signals were collected by the acquisition points located in the Fp1 and Fp2 brain regions of the prefrontal cortex, and the prefrontal EEG signals were preprocessed to obtain prefrontal EEG data.

[0016] Record the response time of each task successfully processed by the user as task execution response data;

[0017] Data acquired through a depth camera is used as eye-tracking data.

[0018] Preferably, the step of determining the response feature map based on the task execution response data includes any one of the following:

[0019] Based on the response time and corresponding difficulty level of each successfully processed task, a response feature map is determined;

[0020] Based on the response time of each successfully processed task, the response time variability is determined, and based on the response time variability and the corresponding task difficulty level, a response feature map is determined.

[0021] Preferably, the step of simultaneously sliding and truncating the power feature map, the response feature map, and the gaze point heatmap to obtain the corresponding time-series segmented map includes:

[0022] Using a sliding window with a preset window size and step size, starting from the same time point, the power feature map, the response feature map, and the fixation point heatmap are synchronously slidably truncated to obtain the same time map segment sequence corresponding to the power feature map, the same time map segment sequence corresponding to the response feature map, and the same time map segment sequence corresponding to the fixation point heatmap.

[0023] Preferably, the step of processing all time-segmented image sequences to obtain multiple single-modal images includes:

[0024] The Markov transfer field algorithm is used to encode the time-segmented sequences corresponding to the power feature map, the response feature map, and the gaze point heatmap, respectively. The encoding matrices are then fused to obtain the R-modal image.

[0025] The recursive graph algorithm is used to encode the time-segmented sequences of the power feature map, the response feature map, and the gaze point heatmap, respectively. The resulting encoding matrices are then fused to obtain the G-modal image.

[0026] The Gram angle field algorithm is used to encode the time-segmented sequences corresponding to the power feature map, the response feature map, and the fixation point heatmap, respectively. The resulting encoding matrices are then fused to obtain the B-mode image.

[0027] Preferably, the process of obtaining the execution capability assessment model includes:

[0028] Acquire prefrontal cortex EEG data, task execution response data, and eye movement data from historical users during executive function training, along with corresponding attention concentration indices, execution response indices, and eye movement fixation concentration indices. Process the prefrontal cortex EEG data, task execution response data, and eye movement data to obtain multiple monomodal images, and fuse these multiple monomodal images to obtain multimodal images. Use the attention concentration indices, execution response indices, and eye movement fixation concentration indices as ground truth values ​​for the multimodal images, annotate the multimodal images, and use the multimodal images and their annotated ground truth values ​​as a sample data set.

[0029] A sample dataset is composed of several sample data points, which is used to train a neural network. The trained neural network is then used as a performance evaluation model.

[0030] According to another aspect of this application, a training system for improving execution capabilities is provided, wherein the system comprises:

[0031] The data acquisition module is used to collect prefrontal EEG data, task execution response data, and eye movement data when the user processes each task during executive ability training. The content of the task is determined according to the task parameters, which include at least: task type, difficulty level, several visual stimulus parameters, and several auditory stimulus parameters.

[0032] The feature map extraction module is used to determine a power feature map based on the prefrontal EEG data, a response feature map based on the task execution response data, and a fixation point heatmap based on the eye movement data and screen parameters.

[0033] The feature map processing module is used to perform synchronous sliding truncation processing on the power feature map, the response feature map and the gaze point heatmap respectively to obtain the corresponding synchronous map segment sequence, process all synchronous map segment sequences to obtain multiple single-modal images, and fuse the multiple single-modal images to obtain a multimodal image;

[0034] The index generation module is used to input the multimodal image into the execution capability assessment model to obtain the attention concentration index, execution response index, and eye movement fixation concentration index corresponding to the multimodal image.

[0035] The dual feedback adjustment module is used to adjust the visual stimulus parameters and / or auditory stimulus parameters of the current task in real time according to the attention concentration index, the eye movement fixation concentration index and preset rules, and to determine whether the execution response index meets the execution response index threshold. If it does not meet the threshold, the task type and / or difficulty level of the next task are adjusted to achieve real-time dual feedback execution ability training within and between tasks.

[0036] Compared with existing technologies, this application provides a training method, system, and device for improving executive function. The method includes: during executive function training, collecting prefrontal cortex EEG data, task execution response data, and eye movement data of the user when processing each task, wherein the task content is determined based on task parameters, which at least include: task type, difficulty level, visual stimulus parameters, and auditory stimulus parameters; determining a power feature map based on the prefrontal cortex EEG data, determining a response feature map based on the task execution response data, and determining a fixation heatmap based on the eye movement data and screen parameters; and simultaneously performing sliding truncation processing on the power feature map, the response feature map, and the fixation heatmap to obtain corresponding... The time-map is segmented into sequences, and all segments of the same time-map are processed to obtain multiple unimodal images. These unimodal images are then fused to obtain multimodal images. These multimodal images are input into an executive ability assessment model to obtain attention concentration indicators, executive response indicators, and eye-tracking fixation concentration indicators corresponding to the multimodal images. Based on the attention concentration indicators, eye-tracking fixation concentration indicators, and preset rules, the visual and / or auditory stimulus parameters of the current task are adjusted in real time. It is also determined whether the executive response indicators meet the executive response indicator threshold. If not, the task type and / or difficulty level of the next task are adjusted to achieve real-time dual feedback for executive ability training within and between tasks. This application can dynamically adjust the current and next training tasks based on the real-time attention concentration, eye-tracking fixation concentration, and executive response indicators obtained by the user during interactive training of brain executive ability, achieving real-time dual feedback during training within and between tasks. This can improve the targeting of training and the accuracy of brain executive ability indicators, and achieve better training results. Attached Figure Description

[0037] Other features, objects, and advantages of this application will become more apparent from the following detailed description of non-limiting embodiments with reference to the accompanying drawings:

[0038] Figure 1 A schematic diagram of a training method for improving execution capability according to one aspect of this application is shown;

[0039] Figure 2 A schematic diagram illustrating the task content at a moment in the current task according to one aspect of this application;

[0040] Figure 3 Showing according to such Figure 2 A schematic diagram of the task content at another moment in the same example of the current task;

[0041] Figure 4 A schematic diagram of a training system for improving execution capabilities according to another aspect of this application is shown;

[0042] Figure 5 A schematic diagram illustrating the effect comparison of a training method for improving execution capability according to one aspect of this application is shown.

[0043] The same or similar reference numerals in the accompanying drawings represent the same or similar parts. Detailed Implementation

[0044] The present application will now be described in further detail with reference to the accompanying drawings.

[0045] In a typical configuration of various embodiments of this application, the method execution entity, each trusted party of the system, and / or each module of the device may include one or more processors (CPU), input / output interfaces, network interfaces, and memory.

[0046] Memory may include non-persistent storage in computer-readable media, such as random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM. Memory is an example of computer-readable media.

[0047] Computer-readable media includes both permanent and non-permanent, removable and non-removable media that can store information by any method or technology. Information can be computer-readable instructions, data structures, modules of programs, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other non-transferable medium that can be used to store information accessible by a computing device. As defined herein, computer-readable media does not include non-transitory computer-readable media, such as modulated data signals and carrier waves.

[0048] To further illustrate the technical means adopted and the effects achieved in this application, the technical solution of this application will be clearly and completely described below in conjunction with the accompanying drawings and preferred embodiments.

[0049] Figure 1 The diagram illustrates a training method for improving execution capabilities according to one aspect of this application, wherein one embodiment of the method includes:

[0050] When performing executive function training, S101 collects prefrontal EEG data, task execution response data, and eye movement data of the user when handling each task. The content of the task is determined according to the task parameters, which include at least: task type, difficulty level, visual stimulus parameters, and auditory stimulus parameters.

[0051] S102 determines a power feature map based on the prefrontal EEG data, determines a response feature map based on the task execution response data, and determines a fixation heatmap based on the eye movement data and screen parameters.

[0052] S103 performs synchronous sliding truncation processing on the power feature map, the response feature map, and the gaze point heatmap respectively to obtain the corresponding time-sharing map segment sequence, processes all time-sharing map segment sequences to obtain multiple single-modal images, and fuses the multiple single-modal images to obtain a multimodal image;

[0053] S104 Input the multimodal image into the execution capability evaluation model to obtain the attention concentration index, execution response index and eye movement fixation concentration index corresponding to the multimodal image;

[0054] S105 adjusts the visual and / or auditory stimulus parameters of the current task in real time according to the attention concentration index, the eye-tracking fixation concentration index, and preset rules, and determines whether the execution response index meets the execution response index threshold. If it does not meet the threshold, the task type and / or difficulty level of the next task are adjusted to achieve real-time dual feedback execution ability training within and between tasks.

[0055] The training method claimed in this application can be used for rehabilitation training of patients with brain executive function or impairments, such as children with ADHD. It can be implemented and / or executed through an application (app) deployed in an interactive device 100, wherein the device 100 can be a smart terminal, computer equipment, and / or cloud with the necessary hardware and software environment. The smart terminal includes, but is not limited to, smartphones, tablets, and smart wearable devices with touch-screen interactive displays; the computer equipment includes, but is not limited to, personal computers, laptops, industrial computers, embedded computers, servers, network hosts, single network servers, or network server clusters; the cloud consists of a large number of computers or network servers based on cloud computing, where cloud computing is a type of distributed computing, consisting of a virtual supercomputer composed of a group of loosely coupled computers.

[0056] The smart terminals, computer devices, and / or cloud mentioned herein are merely examples and not limitations. Other existing or future devices and / or resource platforms that are applicable to this application should also be included within the scope of protection of this application, and are hereby incorporated by reference.

[0057] In this application, caregivers arrange for users (i.e., training subjects, typically sick children) to interact with device 100 for training. Task content is displayed to the user through the device 100's screen and / or speakers, and the device receives the user's touch input and / or voice input. Based on the user's input, indicators that can be used to accurately assess the user's executive function can be obtained in real time. Based on the assessment indicators, the task content of the current and next tasks is dynamically adjusted in real time. This interactive task-based approach provides targeted training for the user's executive function, effectively improving their cognitive abilities and achieving excellent training results.

[0058] In this embodiment, in step S101, when performing executive ability training, the prefrontal EEG data, task execution response data and eye movement data of the user when processing each task are collected. The content of the task is determined according to the task parameters, which include at least: task type, difficulty level, visual stimulus parameters and auditory stimulus parameters.

[0059] When a user trains their brain executive function through device 100, device 100 can collect prefrontal cortex EEG data, task execution response data, and eye movement data when the user processes each task. The task content of each task is determined according to the task parameters, which should include at least: task type, difficulty level, visual stimulus parameters, and auditory stimulus parameters.

[0060] During brain executive function training, visual and / or auditory tasks can be presented to the user through device 100 according to preset rules for the user to process, or they can be presented randomly.

[0061] The task types can include finding differences, grouping, and classifying graphics.

[0062] The difficulty level of the first task can be initially set based on the user's age, gender, etc., while the difficulty level of subsequent tasks can be dynamically adjusted based on the user's performance on the previous task. Different difficulty levels correspond to different task content, which may include the number of distracting objects (or patterns), the style (or category) of distracting objects, the similarity between distracting and target objects, the complexity of the lines in the distracting object patterns, the geometric changes in the distracting object patterns, and a time limit. The movement or rotation direction of the distracting and target objects is random but opposite. The time limit may include a limit on the task response time, which can be limited to no more than 2500ms (task response time typically includes the user's reaction time and processing time for the task content; generally, the higher the task difficulty, the larger this preset value). If the user's response time for the current task exceeds the time limit corresponding to the difficulty level of the current task, it indicates that the user's current brain performance is not yet compatible with the difficulty level of the current task, and the difficulty level of the next task can be adjusted to be lower than the current task.

[0063] Visual stimulus parameters may include the vibration frequency, rotation speed, flashing frequency, and color contrast of an object. Auditory stimulus parameters may include background music rhythm, background cue sounds, and white noise. During the user's processing of the current task (including the first task), the visual and / or auditory stimulus parameters for the current task can be dynamically adjusted based on real-time indicators.

[0064] Optionally, in step S101, the acquisition of prefrontal EEG data, task execution response data, and eye movement data during each task processing by the user includes:

[0065] When the user processes each task

[0066] The user's prefrontal EEG signals were collected by the acquisition points located in the Fp1 and Fp2 brain regions of the prefrontal cortex, and the prefrontal EEG signals were preprocessed to obtain prefrontal EEG data.

[0067] Record the response time of each task successfully processed by the user as task execution response data;

[0068] Data acquired through a depth camera is used as eye-tracking data.

[0069] This system allows users to wear a wearable EEG device, such as a miniature dual-channel wearable EEG device, with acquisition points deployed in at least the Fp1 and Fp2 brain regions of the user's prefrontal cortex. The wearable EEG device can communicate in real-time with device 100 and / or the backend cloud. During brain executive function training, when the user processes each task, the wearable EEG device can accurately acquire and preprocess the user's prefrontal EEG signals from the acquisition points in the Fp1 and Fp2 brain regions of the prefrontal cortex, obtaining the user's prefrontal EEG data. For example, the sampling frequency of the wearable EEG device can be set to 1000Hz. When the user processes each task, the wearable EEG device accurately acquires the user's prefrontal EEG signals, performs preprocessing, and transmits the user's prefrontal EEG data in real-time to device 100 and / or the backend cloud to ensure the real-time performance and stability of subsequent processing of the user's prefrontal EEG data.

[0070] In the field of neuroscience, extensive research has shown that prefrontal EEG signals below 40Hz primarily reflect higher cognitive abilities, such as decision-making, emotion regulation, and attention control. Furthermore, prefrontal EEG signals above 40Hz often contain significant environmental and myoelectric interference, such as electromagnetic interference from surrounding electronic devices and myoelectric signals generated by involuntary, minute muscle tremors. Conversely, signals below 0.5Hz are prone to baseline drift, mainly due to unstable electrode-skin contact and slow fluctuations in the instrument. Baseline drift shifts the baseline of the entire EEG signal, affecting the accurate extraction and analysis of rhythmic characteristics across different frequency bands. This noise contaminates the EEG signal, interfering with the analysis and interpretation of truly effective prefrontal EEG signals originating from brain activity. Therefore, the precisely acquired prefrontal EEG signals from users can be preprocessed using a 0.5-40Hz bandpass filter to remove EEG signals below 0.5Hz and above 40Hz, resulting in prefrontal EEG data within the 0.5-40Hz range, including physiological frequency bands such as delta, theta, alpha, and beta. This helps stabilize the baseline of the EEG signals, improves signal purity, and allows for better focus on meaningful changes in the user's EEG rhythms to assess the user's executive function, thus obtaining more accurate indicators of the user's real-time executive ability. The bandpass filter can be either software-embedded or hardware-embedded; no specific limitation is made here.

[0071] When a user processes each task, they interact with the task through touchscreen, buttons, and other operations. Data such as click frequency, operation trajectory (including operation location), reaction time, and processing time can be collected. The response time of each successfully processed task can be recorded as task execution response data. Each task may include several interactive operations. During task processing, one or more operations may require multiple touchscreen or button presses to obtain the correct result. Therefore, based on the user's click frequency, the response accuracy within the response time of each successfully processed task can be calculated. The response time and corresponding accuracy of each successfully processed task can be used as task execution response data. This user task execution response data can be used to assess core executive function dimensions such as inhibitory control, task memory, cognitive flexibility, and selective attention.

[0072] The task response time can include the user's reaction time and the user's processing time. Successful task processing by the user means that the user's response time for processing the task is within a specified threshold and that the user obtains the correct result for each operation of the task.

[0073] The device 100 can utilize a built-in or external calibrated depth camera, such as a ToF (Time of Flight) depth camera, to sample and capture data in real time, including the duration of the user's gaze, the gaze point, and the gaze point offset, as the user performs each task. This data serves as eye-tracking data. The depth camera calibration includes image distortion correction and pupil center localization. For example, the depth camera's sampling frequency can be set to a value within the range of 40-60Hz. In one application scenario, if the user's gaze point offset exceeds a preset offset threshold for more than a preset time threshold, such as 3 seconds, the device 100 can automatically stop training and prompt the user to refocus before resuming training. The preset offset threshold can be set based on the device 100's screen parameters and / or the distance between the user's eyes and the device 100's screen to avoid collecting invalid eye-tracking data that could affect the accuracy of subsequent evaluation metrics.

[0074] Continuing in this embodiment, in step S102, the device 100 can determine a power feature map based on the prefrontal EEG data, determine a response feature map based on the task execution response data, and determine a fixation point heatmap based on the eye movement data and screen parameters.

[0075] After obtaining the user's prefrontal EEG data, task execution response data, and eye movement data, the device 100 can perform feature extraction processing on the user's prefrontal EEG data, execution response data, and eye movement data respectively to obtain their respective feature maps.

[0076] The device 100 processes the user's prefrontal cortex electroencephalogram (EEG) data to obtain the power spectral density (PSD) of various physiological bands, including delta, theta, alpha, and beta. For example, the classic Welch periodogram method can be used, such as the `pwelch()` method in Matlab or the `scipy.signal.welch()` method in Python. Then, feature extraction and quantification are performed on the obtained PSD of each physiological band to obtain a power feature map. For example, based on the PSD of each physiological band, the proportion of the beta band power relative to the total power can be extracted to determine the corresponding power feature map, which can be used to quantify the attention concentration index in the brain's executive function.

[0077] The device 100 can determine a response feature map based on the user's task execution response data.

[0078] In step S10, device 100 determines a gaze point heatmap based on the gaze point, gaze point offset, and screen parameters in the eye movement data. The screen parameters may include the screen pixel size and screen coordinate system of device 100. In the determined gaze point heatmap, the corresponding area can be marked with a corresponding color based on the gaze point and its offset.

[0079] Optionally, determining the response feature map based on the task execution response data includes any one of the following:

[0080] Based on the response time and corresponding difficulty level of each successfully processed task, a response feature map is determined;

[0081] Based on the response time of each successfully processed task, the response time variability is determined, and based on the response time variability and the corresponding task difficulty level, a response feature map is determined.

[0082] The user's task execution response data includes the response time of each successfully processed task. A two-dimensional response feature map can be determined based on the response time of each successfully processed task and the corresponding task difficulty level. In this map, the horizontal axis represents the difficulty level of the successfully processed task, and the vertical axis represents the response time of that task. Alternatively, the response time variability can be determined based on the response time of each successfully processed task. This variability can be the standard deviation of the response times of all successfully processed tasks, including the current task, or the ratio of the standard deviation to its arithmetic mean. A two-dimensional response feature map can then be determined based on the response time variability of each successfully processed task and the corresponding task difficulty level. In this map, the horizontal axis represents the difficulty level of the currently successfully processed task, and the vertical axis represents the corresponding response time variability.

[0083] The task execution response data includes not only the response time of each task successfully processed by the user, but also the corresponding response accuracy. Furthermore, a two-dimensional response feature map can be determined based on the response accuracy within the response time of each successfully processed task and the difficulty level of the corresponding task. In this two-dimensional response feature map, the horizontal axis represents the difficulty level of the currently successfully processed task, and the vertical axis represents the response accuracy within the response time of that task.

[0084] Furthermore, a three-dimensional response feature map can be determined based on the response time, corresponding response accuracy, and difficulty level of each task successfully processed by the user.

[0085] Continuing in this embodiment, in step S103, the device 100 can perform synchronous sliding truncation processing on the power feature map, the response feature map, and the gaze point heatmap respectively to obtain corresponding time-sharing map segment sequences, process all time-sharing map segment sequences to obtain multiple single-modal images, and fuse the multiple single-modal images to obtain a multimodal image.

[0086] The device 100 can perform synchronous sliding cropping processing on the power feature map, response feature map and gaze point heatmap obtained in step S102, respectively, to obtain the same time map segment sequence corresponding to the power feature map, the same time map segment sequence corresponding to the response feature map and the same time map segment sequence corresponding to the gaze point heatmap, respectively. Then, the three same time map segment sequences are processed to obtain three single-modal images. Finally, the three single-modal images are fused to obtain a multimodal image.

[0087] Optionally, the step of simultaneously sliding and truncating the power feature map, the response feature map, and the gaze point heatmap to obtain the corresponding time-series segmented sequences includes:

[0088] Using a sliding window with a preset window size and step size, starting from the same time point, the power feature map, the response feature map, and the fixation point heatmap are synchronously slidably truncated to obtain the same time map segment sequence corresponding to the power feature map, the same time map segment sequence corresponding to the response feature map, and the same time map segment sequence corresponding to the fixation point heatmap.

[0089] The device 100 can, starting from the same time point, use a sliding window with the same preset window size and step size to synchronously slide and truncate the power feature map, response feature map, and fixation heatmap obtained in step S102, respectively, to obtain a time-series of segmented maps corresponding to the power feature map, the response feature map, and the fixation heatmap. For example, the sliding window has a preset window size of 2500ms and a sliding step size of 200ms (to accurately match the user's neural mechanisms and improve flexibility and targeting, the sliding step size can usually be selected within the range of 100-300ms, depending on the user's actual situation). Using this sliding window, the power feature map, response feature map, and fixation heatmap are synchronously slide and truncated, respectively, to obtain a time-series of segmented maps corresponding to the power feature map, the response feature map, and the fixation heatmap.

[0090] Optionally, the process of processing all time-segmented image sequences to obtain multiple single-modal images includes:

[0091] The Markov transfer field algorithm is used to encode the time-segmented sequences corresponding to the power feature map, the response feature map, and the gaze point heatmap, respectively. The encoding matrices are then fused to obtain the R-modal image.

[0092] The recursive graph algorithm is used to encode the time-segmented sequences of the power feature map, the response feature map, and the gaze point heatmap, respectively. The resulting encoding matrices are then fused to obtain the G-modal image.

[0093] The Gram angle field algorithm is used to encode the time-segmented sequences corresponding to the power feature map, the response feature map, and the fixation point heatmap, respectively. The resulting encoding matrices are then fused to obtain the B-mode image.

[0094] Among them, device 100 can use Markov transition field algorithm, recursive graph algorithm and Gram angle field algorithm respectively to encode each time-segmented sequence of graphs and then fuse them to obtain a single-modal image after processing by each algorithm.

[0095] The Markov transfer field algorithm is employed to encode the time-segmented sequences corresponding to the obtained power feature map and the obtained response feature map, respectively. The dynamic amplitude variation features of these three time-segmented sequences are mapped to encoding matrices of three single-modal images of the same scale. These three encoding matrices are then fused to obtain a single fused matrix, i.e., the R-modal image. For example, the arithmetic mean of the element values ​​at the same position in the three identical-scale encoding matrices is used as the element value at the same position in the matrix corresponding to the fused single-modal image.

[0096] The Markov transition field encoding formula can be expressed as follows:

[0097] MTF i,j =P(s) j |s i (1)

[0098] Among them, MTF i,j P(s) represents the value at position (i,j) in the Markov transition field matrix. j |s i ) is from state s i Transition to state s j The conditional probability of s; i ,s j This is the value at position (i,j) in the segmented sequence of the same time plot.

[0099] If the segmented sequence of the same time map is multidimensional, formula (1) can be applied to each dimension to obtain the corresponding dimension's encoding matrix. Then, these encoding matrices of different dimensions are combined to obtain a single-modality encoding matrix.

[0100] The recursive graph algorithm is employed to encode the time-segmented sequences corresponding to the obtained power feature map and the obtained response feature map, respectively. The nonlinear dynamic amplitude variation features of these three time-segmented sequences are mapped to encoding matrices of three single-modal images of the same scale. These three encoding matrices are then fused to obtain a single fused matrix, i.e., the G-modal image. For example, the arithmetic mean of the element values ​​at the same position in the three identical-scale encoding matrices is used as the element value at the same position in the matrix corresponding to the fused single-modal image.

[0101] The recursive graph encoding formula can be as follows:

[0102] RP i,j =Θ(∈-||x) i -x j ||) (2)

[0103] Among them, RP i,j Let $x$ be the value at position (i,j) in the recursive graph, $Θ(·)$ be the Heaviside step function, $\mathbf{x}$ be the threshold parameter, and $|x$ be the value at position (i,j) in the recursive graph. i -x j || For time point x in the segmented sequence of the same time plot i and x j The Euclidean distance between them.

[0104] RP i,j x represents i and x j The similarity between the two points is determined by Θ(·), which is used to determine whether the similarity between the two points exceeds the threshold. ∈ controls the similarity judgment criteria in the segmented sequence of the same time plot. It can be optimized by trial and error. For example, if ∈ = 0.2, the normalized range of Euclidean distance is [0,1].

[0105] If the segmented sequence of the same time plot is multidimensional, then the Euclidean distance in equation (2) above can be extended to the Euclidean norm, as shown in the following formula:

[0106]

[0107] Where k represents different dimensions of the segmented sequence of the same time plot, x i,k x j,k Representing time point x i and x j The value in the k-th feature dimension.

[0108] The Gram corner field algorithm is used to encode the time-segmented sequences corresponding to the obtained power feature map and the obtained response feature map, respectively. The local temporal variations of these three time-segmented sequences are mapped to encoding matrices of three single-modal images of the same scale. These three encoding matrices are then fused to obtain a single fused matrix, i.e., the B-modal image. For example, the arithmetic mean of the element values ​​at the same position in the three identical-scale encoding matrices is used as the element value at the same position in the matrix corresponding to the fused single-modal image.

[0109] The Gram angle field encoding formula can be summarized as follows:

[0110] GAF i,j =cos(φ i +φ j (4)

[0111] in,

[0112]

[0113] Among them, GAF i,j The value of φ at position (i,j) in the Gram angle field matrix. i ,φ j z represents the angle of time in a segmented sequence of time plots. i / j Let Z be the i / jth data point in the segmented sequence of the same time plot, min(Z) be the minimum value, max(Z) be the maximum value, and arccos(·) be the inverse cosine function.

[0114] GAF i,j Representing time point x i and x j The angular relationship between them, min(Z) and max(Z) make all data points map to the same scale range. Through the above min-max normalization, the time in the segmented sequence of the same time plot is mapped to [-1,1], and then mapped to the angular range [0,π], which is used to construct the Gram angle field.

[0115] If the segmented sequence of the same time map is multidimensional, apply formulas (4) and (5) to each dimension to obtain the corresponding dimension's encoding matrix. Then, fuse these different dimension encoding matrices by concatenation or weighted summation to obtain a single-modality encoding matrix.

[0116] In step S103, fusing the multiple single-modal images to obtain a multimodal image includes fusing the obtained R-modal image, G-modal image, and B-modal image to obtain a multimodal image. The R-modal image, G-modal image, and B-modal image can be combined according to the following formula (6) to obtain the multimodal image.

[0117] Image RGB =(R,G,B) (6)

[0118] Among them, Image RGB The images are multimodal images obtained after fusion. R is the R-mode image obtained after encoding using the Markov transfer field algorithm, G is the G-mode image obtained after encoding using the recursive graph algorithm, and B is the B-mode image obtained after encoding using the Gram angle field algorithm.

[0119] Continuing in this embodiment, in step S104, device 100 can input the multimodal image into the execution capability evaluation model to obtain attention concentration index, execution response index and eye movement fixation concentration index corresponding to the multimodal image.

[0120] The device 100 can input the obtained multimodal image into the execution capability assessment model to obtain the attention concentration index, execution response index and eye movement fixation concentration index corresponding to the multimodal image output by the model.

[0121] Optionally, the acquisition of the execution capability assessment model includes:

[0122] Acquire prefrontal cortex EEG data, task execution response data, and eye movement data from historical users during executive function training, along with corresponding attention concentration indices, execution response indices, and eye movement fixation concentration indices. Process the prefrontal cortex EEG data, task execution response data, and eye movement data to obtain multiple monomodal images, and fuse these multiple monomodal images to obtain multimodal images. Use the attention concentration indices, execution response indices, and eye movement fixation concentration indices as ground truth values ​​for the multimodal images, annotate the multimodal images, and use the multimodal images and their annotated ground truth values ​​as a sample data set.

[0123] A sample dataset is composed of several sample data points, which is used to train a neural network. The trained neural network is then used as a performance evaluation model.

[0124] Before using the training method of this application on a user, the user can refer to the foregoing embodiments and / or optional embodiments, and use the task set for brain executive ability training of this application to train the brain executive ability of a historical user. This involves obtaining the prefrontal EEG data, task execution response data, and eye movement data of the same historical user, as well as obtaining the historical user's actual attention concentration index, execution response index, and eye movement fixation concentration index corresponding to the aforementioned data types. Then, the obtained prefrontal EEG data, task execution response data, and eye movement data of the historical user are processed by feature extraction, encoding, etc., to obtain multiple corresponding monomodal images. These multiple monomodal images are then fused to obtain a multimodal image. The obtained historical user's actual attention concentration index, execution response index, and eye movement fixation concentration index corresponding to the aforementioned data types are then used as the three-branch ground truth values ​​of the multimodal image. The multimodal image is then labeled, and the multimodal image and its labeled ground truth values ​​are used as sample data. Following the above steps, obtain several sample data points to form a sample dataset. The sample dataset should have a sufficient number of sample data points and be diverse. Then, use the sample dataset to train a neural network. Use the trained neural network as an evaluation model for performance capabilities. The specific training, verification, and / or testing methods can adopt existing supervised neural network training, verification, and / or testing methods, which will not be elaborated here.

[0125] The neural network structure may include a residual network and three branches. The residual network may include an input layer, intermediate layers, and output and connection layers. The input layer may include standard convolutional modules, batch normalization, and activation function modules to process the input multimodal image. The intermediate layer may include several residual modules and pooling modules to extract features from the multimodal image and obtain feature vectors. The output and connection layers can process the feature vectors output by the intermediate layers and output them to the three branches. Each branch includes one or more fully connected layers and normalization layers to predict the attention concentration index, execution response index, and eye-tracking fixation index corresponding to the multimodal image based on the output of the residual network; that is, the user's attention concentration index, execution response index, and eye-tracking fixation index during current training.

[0126] In this process, the pixel size of the multimodal image should match the input size of the neural network's input layer, and correspondingly, the pixel size of the unimodal image should also match the input size of the neural network's input layer. For example, if the input size of the neural network's input layer is 224×224, then the resolution of the unimodal image obtained in step S103 should be 224×224. If the unimodal image obtained after processing the segmented sequences of each time-series image is not 224×224 pixels, then the unimodal image should first undergo a lossless size conversion to become a 224×224 unimodal image, and then be fused to obtain a 224×224 multimodal image.

[0127] Continuing in this embodiment, in step S105, the device 100 can adjust the visual stimulus parameters and / or auditory stimulus parameters of the current task in real time according to the attention concentration index, the eye-tracking fixation concentration index and preset rules, and determine whether the execution response index meets the execution response index threshold. If it does not meet the threshold, the device 100 adjusts the task type and / or difficulty level of the next task to achieve real-time dual feedback execution ability training within and between tasks.

[0128] The device 100 can adjust the visual and / or auditory stimulus parameters of the current task in real time based on the attention concentration and eye-tracking attention indicators corresponding to the user's current state, as output by the performance evaluation model, combined with preset rules. This provides real-time positive feedback to the user, enabling real-time reinforcement training within the task. Simultaneously, based on the comparison between the performance response indicators output by the performance evaluation model and preset performance response indicator thresholds, if the thresholds are not met, the device adjusts the task type and / or difficulty level of the next task accordingly. This achieves real-time dual feedback performance training within and between tasks, continuously guiding the training difficulty or pace towards the user's upper limit of brain performance, further reinforcing the performance and achieving better training results.

[0129] Traditional question-based methods for training brain function cannot dynamically adjust the type and content of each question before feedback. Feedback is typically provided after each question is completed, resulting in a relatively long feedback cycle (approximately 1000-2500ms). The user's state may fluctuate multiple times within this cycle, but the system cannot respond to these subtle changes in real time. This leads to training effectiveness relying on external encouragement or the user's willpower. Recent research in neuroscience and other related fields has shown that the brain exhibits detectable neural signal characteristics before attentional decline occurs, such as N2pc (approximately 200ms) in the visual pathway, and N1 (approximately 100ms) and P3 (approximately 300ms) in the auditory pathway. These ERPs (Event-Related Potentials) indicate the rapid processing of sensory stimuli and the regulation of selective attention. Especially among people with ADHD, the prefrontal-parietal-occipital network has insufficient inhibitory control, making them more susceptible to distractions and attention drift in a short period of time. Traditional question bank-based training methods are unable to capture and correct this drift in time, and due to feedback delays, they cannot use this key neural time window to intervene in real time.

[0130] Compared to traditional question-based training methods, this application, in addition to inter-task feedback with a similar feedback cycle, also provides instantaneous visual and / or auditory feedback within the task. By dynamically adjusting visual and / or auditory stimulus parameters within a 100-250 millisecond time window—a time window highly coupled with N2pc (related to visual attention) and N1 and P3 (related to auditory attention)—it innovatively achieves high temporal matching with the neural mechanisms of selective attention in the brain. This allows for timely intervention when the user's attention is about to drift, effectively activating RPE (Reward Predication Error) signals and promoting the neural plasticity adjustment of the prefrontal-parietal-occipital network, thereby enhancing attention maintenance and executive function regulation. Therefore, this application enables instantaneous dynamic adjustment of visual and / or auditory stimuli within the task, with a short feedback cycle (100-250ms), and has the following significant effects:

[0131] 1. Higher feedback frequency and stronger immediate intervention.

[0132] Traditional question-based training methods only provide feedback and adjustments based on the user's performance after each question, resulting in a long cycle and a tendency to miss the brain's reward-sensitive period, leading to a solidified attentional decline. This application, however, can dynamically adjust stimulus parameters within a task, with the adjustment time window highly coupled with relevant ERP components, proactively intervening at the initial stage of attentional fluctuations to prevent attentional breakdown.

[0133] 2. More personalized and precise

[0134] Traditional question bank-based training methods can only make rough adjustments based on average performance over time intervals, making it difficult to accurately adapt to the user's current state. In contrast, this application can capture subtle changes in the user's attention in real time and achieve personalized and precise intervention through closed-loop adjustment via visual and / or auditory dual channels.

[0135] 3. More in line with the cognitive characteristics of ADHD patients

[0136] ADHD patients often have insufficient prefrontal inhibitory function and fragmented attention. Traditional question-based training methods suffer from delayed feedback, making timely correction difficult. However, the instantaneous dynamic feedback mechanism of this application can help users re-anchor their attentional focus by adjusting visual and / or auditory stimulus parameters before the user glances at distracting objects or responds slowly.

[0137] 4. Enhanced training motivation and immersion

[0138] Millisecond-level success / failure feedback instantly activates the user's dopamine reward system, enhancing their sense of control and immersion, creating a positive reinforcement cycle. This avoids the boredom and frustration caused by delayed feedback in traditional question-based training methods.

[0139] An exemplary task content diagram for training brain executive abilities is shown below. Figure 2 , 3 As shown, users initiate training for the current task via touchscreen. Users can interact with the screen or background audio prompts to perform the training. The screen provides real-time updates on relevant information, such as remaining time, score, and continuous attention span. Based on the user's current attention span and eye-tracking attention metrics, the system dynamically adjusts the visual and / or auditory stimulus parameters of the task content in milliseconds to match the user's current brain function, providing timely feedback and greater targeting, thus improving training efficiency and effectiveness.

[0140] In addition, this application can effectively enhance the user's core executive function modules and higher-order executive abilities, such as working memory updating, inhibitory control, and cognitive flexibility, through millisecond-level instantaneous dynamic feedback and adjustment within tasks and adjustment of task difficulty between tasks. This can make up for the shortcomings of traditional question bank-based training methods in simulating complex dynamic cognitive processes.

[0141] For example, suppose the current task is "spot the difference" with a response time limit of 2500ms. After the user performs certain operations on the current task, the acquired prefrontal EEG data, task execution response data, and eye movement data are processed and encoded to obtain multimodal images. These multimodal images are then input into an execution ability assessment model to obtain attention concentration indicators, execution response indicators, and eye-tracking attention indicators corresponding to the user's current state. If the obtained attention concentration indicators do not meet the preset attention concentration indicator threshold and decline by more than 10% in a short period of time, the visual stimulus parameters can be adjusted for the current task. For example, the image vibration frequency in the visual stimulus parameters can be adjusted to increase by 1Hz compared to the original image vibration frequency, so as to create an interference enhancement effect when the user continues to process the current task, thereby strengthening the user's training effect on processing the current task. When the attention concentration indicators recover, the image vibration frequency can be automatically reduced simultaneously to achieve positive feedback reinforcement. If the obtained eye-tracking fixation concentration index drops below the preset eye-tracking fixation concentration index threshold, the visual stimulus parameters can be adjusted. For example, if the user's gaze point deviates from the target area for 200 milliseconds or more, the rotation speed in the adjusted visual stimulus parameters can be increased by 5-10° / s compared to the original rotation speed. This creates an interference enhancement effect while the user continues to process the current task, strengthening the user's training effect in processing the current task and inducing the user's gaze point to focus on the core target area. When the eye-tracking fixation concentration index returns to a stable level, the rotation speed is automatically adjusted back simultaneously, thereby achieving positive feedback reinforcement. If the obtained execution response index is lower than the preset execution response index threshold by more than 10%, the task type and / or difficulty level of the next task can be adjusted. For example, the task type of the next task can be adjusted to "image classification" (which is easier to handle than "spot the difference"), or the task type can be left unchanged, but the difficulty level of the next task can be lowered by one level. If the subsequent real-time execution response metrics are higher than the preset execution response metric threshold by more than 10%, the task type and / or difficulty level of the next task can be adjusted to achieve positive feedback reinforcement.

[0142] In the training methods for improving executive ability described in the above embodiments and / or optional embodiments, the relevant data processing and feedback adjustment can usually be completed within hundreds of milliseconds. This can effectively enhance the ability to instantly perceive and regulate the current state of the user during the brain executive ability training process, thereby achieving real-time dual positive feedback within and between tasks. This makes the training of the user's brain executive ability more targeted, obtains a more accurate current state of the user, and achieves better training results.

[0143] Figure 4The diagram illustrates a training system for improving execution capabilities according to another aspect of this application, wherein, in one embodiment, the system includes:

[0144] The data acquisition module 410 is used to collect prefrontal EEG data, task execution response data and eye movement data when the user is processing each task during executive ability training. The content of the task is determined according to the task parameters, which include at least: task type, difficulty level, several visual stimulus parameters and several auditory stimulus parameters.

[0145] The feature map extraction module 420 is used to determine a power feature map based on the prefrontal EEG data, to determine a response feature map based on the task execution response data, and to determine a fixation point heatmap based on the eye movement data and screen parameters.

[0146] The feature map processing module 430 is used to perform synchronous sliding truncation processing on the power feature map, the response feature map and the gaze point heatmap respectively to obtain the corresponding synchronous map segment sequence, process all synchronous map segment sequences to obtain multiple single-modal images, and fuse the multiple single-modal images to obtain a multimodal image.

[0147] The index generation module 440 is used to input the multimodal image into the execution capability assessment model to obtain the attention concentration index, execution response index and eye movement fixation concentration index corresponding to the multimodal image;

[0148] The dual feedback adjustment module 450 is used to adjust the visual stimulus parameters and / or auditory stimulus parameters of the current task in real time according to the attention concentration index, the eye movement fixation concentration index and preset rules, and to determine whether the execution response index meets the execution response index threshold. If it does not meet the threshold, the task type and / or difficulty level of the next task are adjusted to achieve real-time dual feedback execution ability training within and between tasks.

[0149] The system is deployed in the aforementioned device 100. This system can be used to train the executive functions of children with ADHD (i.e., the user, or the training subject). When the user trains their executive functions, a wearable EEG device worn by the user can collect the user's prefrontal cortex EEG signals in real time, and perform preprocessing such as bandpass filtering to obtain prefrontal cortex EEG data. Additionally, a calibrated depth camera, either built into or external to device 100, can collect the user's eye movement data in real time.

[0150] In this embodiment, when a user performs brain executive function training, the system’s data acquisition module 410 collects the user’s prefrontal EEG data, task execution response data and eye movement data when the user processes each task. The content of the task is determined according to the task parameters, which include at least the following: task type, difficulty level, several visual stimulation parameters and several auditory stimulation parameters.

[0151] Continuing in this embodiment, the feature map extraction module 420 of the system can perform feature extraction processing on the obtained prefrontal EEG data, execution response count, and eye movement data of the user to obtain corresponding feature maps. Specifically, the obtained prefrontal EEG data of the user can be processed to obtain the power spectral density of each physiological band such as δ, θ, α, and β. The power spectral density of each physiological band can be extracted and quantized to obtain a power feature map. The obtained task execution response data of the user can be processed to determine the response feature map. Based on the fixation point, fixation point offset, and screen parameters in the obtained eye movement data, a fixation point heatmap can be determined.

[0152] Continuing in this embodiment, the feature map processing module 430 of the system can perform synchronous sliding truncation processing on the obtained power feature map, response feature map and gaze point heatmap respectively, to obtain the same time map segment sequence corresponding to the power feature map, the same time map segment sequence corresponding to the response feature map and the same time map segment sequence corresponding to the gaze point heatmap respectively. Then, these three same time map segment sequences are processed to obtain three single-modal images. Then, these three single-modal images are fused to obtain a multimodal image.

[0153] Continuing in this embodiment, the obtained multimodal image can be input into the execution capability assessment model through the index generation module 440 of the system to obtain the attention concentration index, execution response index and eye movement fixation concentration index corresponding to the multimodal image output by the model.

[0154] Continuing in this embodiment, the system's dual-feedback adjustment module 450 can adjust the visual and / or auditory stimulus parameters of the current task in real time based on the attention concentration and eye-tracking attention indicators corresponding to the user's current state, as output by the performance evaluation model, and in accordance with preset rules. This provides real-time positive feedback to the user, enabling real-time reinforcement training within the task. Simultaneously, based on the comparison between the performance response indicators output by the performance evaluation model and preset performance response indicator thresholds, if the thresholds are not met, the task type and / or difficulty level of the next task are adjusted accordingly. This achieves real-time dual-feedback performance training within and between tasks, continuously guiding the training difficulty or pace towards the user's upper limit of brain performance, further reinforcing the performance and achieving better training results.

[0155] In this embodiment, any functions or executable method steps that the various components of the system can achieve are the same as those in the foregoing related method embodiments and / or optional embodiments, and will not be repeated here.

[0156] This system can be deployed as an app on Device 100. It collects prefrontal cortex EEG signals using a wearable EEG device, making it ideal for everyday home, school, and caregiving / rehabilitation settings. Supporting iOS and Android platforms, the system provides a user-friendly interface designed specifically for children. Parents, teachers, and caregivers can instantly view training data and results after each training session, including raw data such as prefrontal cortex EEG data, task execution response data, and eye movement data, as well as feature maps, monomodal images, multimodal images, and assessment indicators. It visually presents changes in brain performance indicators related to executive function, such as attention span, execution response, and eye fixation concentration, during the current or historical training sessions, facilitating long-term tracking and scientific management.

[0157] This system, based on wearable EEG devices and smart mobile applications, creates an efficient, portable, and real-time feedback training ecosystem for improving executive function in home, school, care, or rehabilitation settings. Its innovations lie in: employing lightweight EEG sensing technology to capture and analyze users' (e.g., children with ADHD or ADHD patients) EEG activity, executive responses, and eye-tracking data during training, ensuring high-precision monitoring; enabling real-time, stable, and efficient transmission of collected data, automatically synchronizing it to the cloud to guarantee data security and traceability; and most innovatively, providing real-time dual feedback both within and between tasks, allowing for more targeted training. Furthermore, it instantly generates analysis reports after training, including raw EEG data, indicator evaluation results, and personalized training suggestions, helping parents, teachers, or caregivers intuitively understand the user's training progress, conduct long-term tracking and scientific management, and contribute to the systematic improvement of the user's executive function. This system adopts a modular architecture design, combining non-invasive EEG signals, task execution response data, eye-tracking data and other multimodal data acquisition, processing and analysis, as well as dynamic training task generation and intra-task and inter-task dual feedback closed-loop regulation, to build an efficient and personalized positive intervention platform for brain executive ability.

[0158] An exemplary comparative diagram showing the effects of two months of brain executive function training using the training system for improving executive function described in this application is shown below. Figure 5 As shown, this control group used a two-tailed test, n=45, p<0.01. This control group was determined through a multidisciplinary joint clinical study involving experts from pediatrics and psychology departments at two top-tier hospitals, reviewed and approved by the ethics committee, and strictly implemented under their supervision. Participating users underwent baseline assessments using the same standards at the same designated institution before and after the training program. This involved examining and diagnosing commonly used indicators of brain executive function (or ability), such as visual recognition rate, auditory recognition rate, inhibitory index, and attention span, before and after the training program for comparison and evaluation of the system. The training program lasted two months, with 30 minutes of training per day, ensuring at least 30 training sessions were completed within the program. This training frequency and duration were designed to balance the continuity and effectiveness of the training, taking into account the user's daily learning and life rhythm, ensuring good feasibility and compliance. Figure 5 The before-and-after training comparison shown, through the quantitative values ​​of relevant indicators, clearly demonstrates that the training system for improving executive function proposed in this application has a significant effect in improving the user's brain executive function, including targeted training and good results. In particular, it can provide an efficient training approach for the rehabilitation of children with ADHD.

[0159] According to another aspect of this application, a computer-readable medium is also provided, the computer-readable medium storing computer-readable instructions that can be executed by a processor to implement some or all of the foregoing method embodiments and / or optional embodiments.

[0160] It should be noted that the method embodiments and / or optional embodiments in this application do not strictly limit the order of execution of each step, as long as the method embodiments and / or optional embodiments can solve the defects existing in the prior art, achieve the inventive purpose of this application, and obtain beneficial effects. The method embodiments and / or optional embodiments in this application can be implemented in software and / or combinations of software and hardware. The software program involved in this application can be executed by a processor to implement the steps or functions of the above embodiments. Similarly, the software program of this application (including related data structures) can be stored in a computer-readable recording medium.

[0161] Furthermore, part or all of this application can be applied as a computer program product, such as computer program instructions, which, when executed by a computer, can invoke or provide the methods and / or technical solutions according to this application through the operation of the computer. The program instructions invoking the methods of this application may be stored in a fixed or removable recording medium, and / or transmitted via data streams in broadcast or other signal carrying media, and / or stored in the working memory of a computer device operating according to the program instructions.

[0162] According to another aspect of this application, a training device for improving execution capabilities is also provided. The device includes: a memory for storing computer program instructions and a processor for executing the program instructions, wherein when the computer program instructions are executed by the processor, the device is triggered to run part or all of the methods and / or technical solutions of the foregoing embodiments.

[0163] It will be apparent to those skilled in the art that the present invention is not limited to the details of the exemplary embodiments described above, and that the invention can be implemented in other specific forms without departing from its spirit or essential characteristics. Therefore, the embodiments should be considered in all respects as exemplary and non-limiting, and the scope of the invention is defined by the appended claims rather than the foregoing description. Thus, all variations falling within the meaning and scope of equivalents of the claims are intended to be embraced within the present invention. No reference numerals in the claims should be construed as limiting the scope of the claims.

[0164] In this application, when terms such as "upper," "lower," "left," "right," "front," "rear," "top," "bottom," "inner," "outer," "middle," "vertical," "horizontal," "lateral," and "longitudinal" are used, the indicated orientation and / or positional relationship is based on the orientation and / or positional relationship shown in the accompanying drawings. These terms are primarily for the purpose of better describing this application and its embodiments, and are not intended to limit the indicated device, element, or component to having a specific orientation, or to be constructed and operated in a specific orientation. Furthermore, some of the above terms, in addition to indicating orientation or positional relationship, can also be used to indicate other meanings; for example, the term "upper" can also be used in some cases to indicate a certain dependency or connection relationship. Those skilled in the art can understand the specific meaning of these terms in this application according to the specific circumstances.

[0165] Furthermore, the terms "installation," "setup," "equipped with," "connection," and "linked" should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral structure; they can refer to a mechanical connection or an electrical connection; they can refer to a direct connection or an indirect connection via an intermediate medium; and they can refer to an internal connection between two devices, components, or constituent parts. Those skilled in the art can understand the specific meaning of the above terms in this application based on the specific circumstances.

[0166] Furthermore, the terms "first," "second," etc., are primarily used to distinguish different devices, units, modules, elements, circuits, or components (which may be the same or different in specific type and construction), and are not intended to indicate or imply the relative importance, order, and / or quantity of the indicated devices, units, modules, elements, circuits, or components. Unless otherwise stated, "a plurality of" means two or more.

[0167] Furthermore, it is clear that the word "comprising" does not exclude other units or steps, and the singular does not exclude the plural. Multiple units or devices recited in the apparatus claims may also be implemented by a single unit or device through software and / or hardware.

Claims

1. A training method for improving execution capabilities, characterized in that, The method includes: When performing executive function training, the prefrontal cortex EEG data, task execution response data and eye movement data of the user are collected when processing each task. The content of the task is determined according to the task parameters, which include at least: task type, difficulty level, visual stimulus parameters and auditory stimulus parameters. Based on the prefrontal EEG data, a power feature map is determined; based on the task execution response data, a response feature map is determined; and based on the eye movement data and screen parameters, a fixation heatmap is determined. The power feature map, the response feature map, and the gaze point heatmap are simultaneously slidably truncated to obtain corresponding time-series image segments. All time-series image segments are then processed to obtain multiple single-modal images. Finally, the multiple single-modal images are fused to obtain a multimodal image. The multimodal images are input into the execution capability assessment model to obtain attention concentration index, execution response index, and eye movement fixation concentration index corresponding to the multimodal images; Based on the attention concentration index, the eye-tracking fixation concentration index, and preset rules, the visual and / or auditory stimulus parameters of the current task are adjusted in real time, and it is determined whether the execution response index meets the execution response index threshold. If it does not meet the threshold, the task type and / or difficulty level of the next task are adjusted to achieve real-time dual feedback execution ability training within and between tasks.

2. The method according to claim 1, characterized in that, The data collected includes prefrontal cortex EEG data, task execution response data, and eye movement data during each task processing by the user, including: When the user processes each task The user's prefrontal EEG signals were collected by the acquisition points located in the Fp1 and Fp2 brain regions of the prefrontal cortex, and the prefrontal EEG signals were preprocessed to obtain prefrontal EEG data. Record the response time of each task successfully processed by the user as task execution response data; Data acquired through a depth camera is used as eye-tracking data.

3. The method according to claim 2, characterized in that, The step of determining the response feature map based on the task execution response data includes any one of the following: Based on the response time and corresponding difficulty level of each successfully processed task, a response feature map is determined; Based on the response time of each successfully processed task, the response time variability is determined, and based on the response time variability and the corresponding task difficulty level, a response feature map is determined.

4. The method according to claim 1, characterized in that, The process of simultaneously sliding and truncating the power feature map, the response feature map, and the gaze point heatmap to obtain the corresponding time-series segmented sequences includes: Using a sliding window with a preset window size and step size, starting from the same time point, the power feature map, the response feature map, and the fixation point heatmap are synchronously slidably truncated to obtain the same time map segment sequence corresponding to the power feature map, the same time map segment sequence corresponding to the response feature map, and the same time map segment sequence corresponding to the fixation point heatmap.

5. The method according to claim 1, characterized in that, The process of processing all time-segmented image sequences yields multiple single-modal images, including: The Markov transfer field algorithm is used to encode the time-segmented sequences corresponding to the power feature map, the response feature map, and the gaze point heatmap, respectively. The encoding matrices are then fused to obtain the R-modal image. The recursive graph algorithm is used to encode the time-segmented sequences of the power feature map, the response feature map, and the gaze point heatmap, respectively. The resulting encoding matrices are then fused to obtain the G-modal image. The Gram angle field algorithm is used to encode the time-segmented sequences corresponding to the power feature map, the response feature map, and the fixation point heatmap, respectively. The resulting encoding matrices are then fused to obtain the B-mode image.

6. The method according to claim 1, characterized in that, The acquisition of the execution capability assessment model includes: Acquire prefrontal cortex EEG data, task execution response data, and eye movement data from historical users during executive function training, along with corresponding attention concentration indices, execution response indices, and eye movement fixation concentration indices. Process the prefrontal cortex EEG data, task execution response data, and eye movement data to obtain multiple monomodal images, and fuse these multiple monomodal images to obtain multimodal images. Use the attention concentration indices, execution response indices, and eye movement fixation concentration indices as ground truth values ​​for the multimodal images, annotate the multimodal images, and use the multimodal images and their annotated ground truth values ​​as a sample data set. A sample dataset is composed of several sample data points, which is used to train a neural network. The trained neural network is then used as a performance evaluation model.

7. A training system for improving execution capabilities, characterized in that, The system includes: The data acquisition module is used to collect prefrontal EEG data, task execution response data, and eye movement data when the user processes each task during executive ability training. The content of the task is determined according to the task parameters, which include at least: task type, difficulty level, several visual stimulus parameters, and several auditory stimulus parameters. The feature map extraction module is used to determine a power feature map based on the prefrontal EEG data, a response feature map based on the task execution response data, and a fixation point heatmap based on the eye movement data and screen parameters. The feature map processing module is used to perform synchronous sliding truncation processing on the power feature map, the response feature map and the gaze point heatmap respectively to obtain the corresponding synchronous map segment sequence, process all synchronous map segment sequences to obtain multiple single-modal images, and fuse the multiple single-modal images to obtain a multimodal image; The index generation module is used to input the multimodal image into the execution capability assessment model to obtain the attention concentration index, execution response index, and eye movement fixation concentration index corresponding to the multimodal image. The dual feedback adjustment module is used to adjust the visual stimulus parameters and / or auditory stimulus parameters of the current task in real time according to the attention concentration index, the eye movement fixation concentration index and preset rules, and to determine whether the execution response index meets the execution response index threshold. If it does not meet the threshold, the task type and / or difficulty level of the next task are adjusted to achieve real-time dual feedback execution ability training within and between tasks.

8. A computer-readable medium, characterized in that, It stores computer-readable instructions that are executed by a processor to implement part or all of the method as described in any one of claims 1 to 6.

9. A training device for improving execution capabilities, characterized in that, The device includes: One or more processors; and A memory storing computer-readable instructions, which, when executed, cause the processor to perform some or all of the operations of the method as described in any one of claims 1 to 6.