Attention prediction method and device, equipment and storage medium

By collecting and segmenting multimodal physiological signals and using the Transformer model to predict attention and adjust task difficulty, the problem of lack of real-time and objective quantification in ADHD assessment methods is solved, and real-time, accurate assessment and dynamic optimization of ADHD are achieved.

CN120678430APending Publication Date: 2025-09-23SHENZHEN PENGRUI BRAIN SCI & TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510585133.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-06
Publication Date
2025-09-23

AI Technical Summary

Technical Problem

Existing ADHD assessment methods lack real-time and objective quantitative means, and the assessment effect is poor.

Method used

By continuously collecting the multimodal physiological signals of subjects in specific task scenarios, dividing them into windows of preset length, and using a set step size for attention prediction, combined with the self-attention mechanism of the Transformer model for feature fusion and time-weighted smoothing, attention prediction results are generated, and the task difficulty is dynamically adjusted to form a closed-loop feedback mechanism.

Benefits of technology

It achieves real-time, accurate and objective assessment of ADHD, improves the timeliness and compliance of assessment, and ensures quantitative assessment results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120678430A_ABST
    Figure CN120678430A_ABST
Patent Text Reader

Abstract

The invention is suitable for the technical field of human-computer interaction, and provides an attention prediction method, device and equipment and a storage medium, and the method comprises the steps: continuously collecting multi-modal physiological signals of a subject in a specific task scene; segmenting the multi-modal physiological signal into a plurality of windows with preset lengths according to a set step length; performing attention prediction on the multi-modal physiological signal in each window to obtain an attention prediction result; adjusting the task difficulty in the specific task scene based on the size of the attention prediction result; and returning to execute the step of continuously collecting the multi-modal physiological signals of the subject in the specific task scene until the attention prediction execution is finished. According to the scheme, the quantitative evaluation effect of ADHD can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application belongs to the field of human-computer interaction technology, and in particular relates to an attention prediction method, apparatus, device and storage medium. Background Art

[0002] ADHD (Attention-deficit hyperactivity disorder) is a common neurodevelopmental disorder, usually manifested by symptoms such as inattention and hyperactivity.

[0003] Existing evaluation methods mostly rely on questionnaires or short-term observations, lack real-time and objective quantitative means, and have poor evaluation results. Summary of the Invention

[0004] The embodiments of the present application provide an attention prediction method, apparatus, device, and storage medium to address the problem that the existing ADHD assessment methods lack real-time and objective quantification means and have poor assessment results.

[0005] A first aspect of an embodiment of the present application provides an attention prediction method, comprising:

[0006] Continuously collect multimodal physiological signals of subjects in specific task scenarios;

[0007] Segmenting the multimodal physiological signal into a plurality of windows of preset lengths with a set step size;

[0008] Performing attention prediction on the multimodal physiological signal in each of the windows to obtain an attention prediction result; the attention prediction result is used to indicate the attention level of the subject in the specific task scenario;

[0009] Adjusting the difficulty of the task in the specific task scenario based on the size of the attention prediction result;

[0010] Return to the step of continuously collecting the subject's multimodal physiological signals in a specific task scenario until the attention prediction execution is completed.

[0011] A second aspect of an embodiment of the present application provides an attention prediction device, comprising:

[0012] The acquisition module is used to continuously collect multimodal physiological signals of the subject in a specific task scenario;

[0013] a window module, configured to divide the multimodal physiological signal into a plurality of windows of preset lengths with a set step size;

[0014] A prediction module, configured to perform attention prediction on the multimodal physiological signal in each window to obtain an attention prediction result; the attention prediction result is used to indicate the attention level of the subject in the specific task scenario;

[0015] An adjustment module is used to adjust the task difficulty in the specific task scenario based on the size of the attention prediction result; and return to the step of continuously collecting the multimodal physiological signals of the subject in the specific task scenario until the attention prediction execution is completed.

[0016] A third aspect of an embodiment of the present application provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps of the method described in the first aspect when executing the computer program.

[0017] A fourth aspect of an embodiment of the present application provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the method described in the first aspect are implemented.

[0018] The fifth aspect of the present application provides a computer program product, including a computer-readable code, or a non-volatile computer-readable storage medium carrying the computer-readable code. When the computer-readable code is run in an electronic device, the processor in the electronic device executes the steps in the method described in the first aspect above.

[0019] In an embodiment of the present application, the multimodal physiological signals of the subject are continuously collected in a specific task scenario, and the multimodal physiological signals are divided into multiple windows of preset length with a set step size. The attention prediction is performed on the multimodal physiological signals in each window to obtain the attention prediction result. Then, based on the size of the attention prediction result, the task difficulty in the specific task scenario is adjusted until the attention prediction execution is completed, forming a dynamically optimized closed-loop feedback mechanism, integrating data collection, prediction evaluation and feedback adjustment functions into one, and being able to timely adjust the training plan according to the evaluation results of the subject, adjust the task difficulty in the specific task scenario, so as to carry out effective testing of the subject and ensure the quantitative evaluation effect of ADHD. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] Various other advantages and benefits will become apparent to those skilled in the art upon reading the detailed description of the preferred embodiment below. The accompanying drawings are for illustration purposes only and are not to be considered as limiting the present application. The same reference numerals are used throughout the drawings to represent the same components. In the drawings:

[0021] Figure 1This is the process of the attention prediction method in some embodiments of the present application Figure 1 ;

[0022] Figure 2 This is the process of the attention prediction method in some embodiments of the present application Figure 2 ;

[0023] Figure 3 is a module diagram of an attention prediction device in some embodiments of the present application;

[0024] Figure 4 This is a structural diagram of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0025] The following embodiments of the technical solution of the present application will be described in detail with reference to the accompanying drawings. The following embodiments are only used to more clearly illustrate the technical solution of the present application and are therefore only examples and are not intended to limit the scope of protection of the present application.

[0026] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which this application belongs; the terms used herein are only for the purpose of describing specific embodiments and are not intended to limit this application; the terms "including" and "having" and any variations thereof in the specification and claims of this application and the above-mentioned figure descriptions are intended to cover non-exclusive inclusions.

[0027] In the description of the embodiments of this application, the technical terms "first" and "second" are used only to distinguish different objects and should not be understood to indicate or imply relative importance or implicitly specify the quantity, specific order, or primary and secondary relationship of the indicated technical features. In the description of the embodiments of this application, the meaning of "plurality" is more than two, unless otherwise clearly and specifically defined.

[0028] References herein to "embodiments" mean that a particular feature, structure, or characteristic described in connection with the embodiments may be included in at least one embodiment of the present application. The appearance of this phrase in various places in the specification does not necessarily refer to the same embodiment, nor does it constitute an independent or alternative embodiment that is mutually exclusive of other embodiments. It is understood, both explicitly and implicitly, by those skilled in the art that the embodiments described herein may be combined with other embodiments.

[0029] In the description of the embodiments of this application, the term "and / or" is simply a description of the association relationship between associated objects, indicating that three relationships can exist. For example, A and / or B can represent the following three situations: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the character " / " in this document generally indicates that the associated objects are in an "or" relationship.

[0030] In order to illustrate the technical solution described in this application, specific embodiments are provided below.

[0031] In an alternative embodiment, in combination Figure 1 As shown, an attention prediction method provided by an embodiment of the present application includes:

[0032] Step 101: continuously collect multimodal physiological signals of a subject in a specific task scenario.

[0033] Multimodal physiological signals are used to indicate the subject's task performance in a specific task scenario.

[0034] Multimodal physiological signals specifically refer to physiological signals in multiple modal forms. Examples include EEG (Electroencephalogram), HRV (Heart Rate Variability), or other motion-specific signals of the human body, such as head movement data and eye movement signals. These multimodal physiological signals can reflect the subject's attention state and emotional response in different dimensions, allowing them to be integrated into task-related assessment and training.

[0035] In an optional implementation, specific task scenarios can be configured to include static or slightly dynamic tasks.

[0036] Optionally, the setting of static or slightly dynamic tasks can be to design specific task scenarios for the subjects, such as visual tracking tasks, simple calculation tasks, slight movement tasks, etc., and the task duration can be set to 2-5 minutes.

[0037] Among them, static tasks include: visual tracking, color distinction, number games, reaction speed tests, etc., and slightly dynamic tasks include: grabbing randomly falling sticks, classifying small objects, etc.

[0038] In these specific task scenarios, the subjects wear wearable devices to ensure that the sensors are attached and working properly, and use the wearable devices to collect the subjects' multimodal physiological signals to obtain the subjects' EEG signals, HRV signals, motion characteristic signals, etc.

[0039] When collecting EEG signals, the wearable device can be used to perform power spectrum analysis on the subject's EEG signals to extract relevant features indicating concentration or attention in each band of their brain waves.

[0040] When collecting HRV signals, power features such as RMSSD (Root Mean Square of Successive Differences), VLF (Very Low Frequency Power), LF (Low Frequency Power), and HF (High Frequency Power) can be extracted to analyze the subject's stress level.

[0041] In frequency domain analysis, the power spectrum is usually divided into VLF (≤0.04Hz), LF (0.04-0.15Hz) and HF (0.15-0.4Hz). The values ​​of different power bands reflect the activity of the autonomic nervous system and the regulatory ability of the heart, and are used to analyze the stress level of the subjects.

[0042] RMSSD is a metric used in time-domain analysis. It is calculated by first averaging the squares of the differences between adjacent RR intervals and then taking the square root. RMSSD primarily reflects parasympathetic nervous system activity, particularly the influence of the vagus nerve. A higher RMSSD generally indicates better parasympathetic nervous system regulation and may be associated with lower stress levels.

[0043] The RR interval is the time interval between adjacent R waves.

[0044] When collecting motion feature data, the subject's head movement data (such as acceleration and angular velocity) can be collected, based on which the movement frequency (i.e., the number of movements per unit time) and amplitude (the magnitude of acceleration or angular velocity) can be calculated; or the subject's eye movement data can be collected to analyze the duration and number of times the subject fixates on an area of ​​interest. For example, long gazes may indicate concentration, while frequent shifts may indicate distraction. Based on this, motion feature data is obtained.

[0045] Furthermore, preprocessing can be performed on the collected multimodal data. This can be done to remove signal loss, dropouts caused by loose equipment, or environmental interference, as well as non-standard cases. In an optional implementation, wavelet denoising can be used to remove artifacts from EEG signals, while filters can be used to remove baseline drift and high-frequency noise from PPG signals. The data can then be normalized to align the characteristic scales of the physiological signals across modalities.

[0046] In this way, by designing specific tasks and synchronously collecting multimodal physiological signals, the subject's task performance is associated with his or her multimodal physiological signals. By continuously collecting the subject's multimodal physiological signals in real time, the subject's task performance in specific task scenarios can be predicted, providing a multi-dimensional basis for more accurate assessment of ADHD.

[0047] Step 102 : Segment the multimodal physiological signal into a plurality of windows of preset lengths using a set step size.

[0048] The window length refers to the time span or number of samples in each window. For example, if the sampling rate of the signal is fs, the number of samples of the multimodal physiological signal corresponding to the window length T is N = T * fs.

[0049] The step size is the amount by which the window moves in the signal. For example, if the step size is S, then the starting position of the next window is the starting position of the current window plus S samples (or S time units).

[0050] The window length T and the step length S can be set according to the amount of multimodal physiological signal data and the required calculation accuracy of the data.

[0051] Step 103: perform attention prediction on the multimodal physiological signal in each window to obtain an attention prediction result.

[0052] The attention prediction results are used to indicate the subject's attention level in a specific task scenario.

[0053] Optionally, the attention prediction result may be a result of high concentration, medium concentration, low concentration, etc. Attention can be divided into different intervals according to the degree of concentration. When performing attention prediction on the multimodal physiological signals in each window, the concentration interval (i.e., the attention level interval) into which the subject's task performance falls is determined based on the multimodal physiological signals in the window, thereby obtaining the attention prediction result.

[0054] Attention prediction is performed on the multimodal physiological signals in each window. The concentration scores of each physiological signal can be directly determined based on the physiological signals of each modality of the subject, and the overall concentration range is determined by combining these concentration scores to obtain the attention prediction result.

[0055] Alternatively, the multimodal physiological signals in each window can be used to perform attention prediction with the help of a pre-trained attention prediction model to obtain attention prediction results.

[0056] Correspondingly, in an optional embodiment, attention prediction is performed on the multimodal physiological signal in each window to obtain an attention prediction result, including:

[0057] Step 201 : performing feature extraction on the multimodal physiological signal in each window to obtain a first feature vector of the physiological signal of each modality.

[0058] Step 202: Perform feature fusion on the first feature vector to generate a fused first fused feature vector.

[0059] Step 203: Input the first fused feature vector into the attention prediction model to obtain a first attention prediction value.

[0060] Among them, the first attention prediction value is used to indicate the subject's attention level in a specific task scenario corresponding to the multimodal physiological signal in the current window.

[0061] Optionally, the first fused feature vector can be generated by splicing the first feature vectors of the physiological signals of each modality to obtain a spliced ​​feature vector, performing feature fusion on the spliced ​​feature vector based on the self-attention mechanism of the Transformer model to generate a fused feature vector, and inputting the fused feature vector into the fully connected layer of the attention prediction model for classification processing to obtain a first attention prediction value.

[0062] Among them, the implementation of the self-attention mechanism of the Transformer model can be to perform feature fusion processing on the spliced ​​feature vector through the multi-head self-attention layer of the Transformer model to capture the global correlation between the features corresponding to physiological signals of different modalities.

[0063] Optionally, the feature vectors of the physiological signals of each modality include brain wave feature vectors, heart rate variability feature vectors, and motion feature vectors. The generated concatenated feature vector can be:

[0064] X fusion =Concat(X EEG , X HRV , X Motion );

[0065] Among them, X EEG is the brain wave feature vector, X HRV is the heart rate variability feature vector, X Motion is the motion feature vector;

[0066] Then, based on the self-attention mechanism of the Transformer model, the concatenated feature vector is fused to generate a fused feature vector:

[0067]

[0068] Where Q = W Q X fusion ; K=W K X fusion ; V=W V X fusion ; Q, K, V are query matrix, key matrix and value matrix respectively, derived from X fusion Linear projection of .

[0069] Wherein, Z is the fusion feature vector; W Q 、W K 、W V are transformation matrices, which are used to map the concatenated feature vectors to different feature spaces to obtain corresponding linear projection representations; The scaling factor is set.

[0070] This process uses machine learning models to evaluate multimodal physiological signals and predict task performance, breaking through the limitations of traditional methods that rely on questionnaires or short-term observations. It achieves effective real-time tracking of subjects' attention fluctuations, realizes real-time evaluation and prediction of data, and improves the timeliness and objectivity of the evaluation.

[0071] Step 204 : Perform time-weighted smoothing on the first attention prediction value of each window to obtain an attention prediction result corresponding to the multimodal physiological signal.

[0072] Weighted smoothing, as a data processing method, can be used to reduce data noise or fluctuations, resulting in smoother data results. Time-weighted smoothing builds on this approach by taking time into account to determine the time weight of data objects. For example, more recent data may have a higher weight, and vice versa.

[0073] The time weight can be set based on the time sequence of the windows, giving a greater weight to the first attention prediction value of the recent window and a smaller weight to the first attention prediction value of the earlier window. For example, the weights of the first attention prediction values ​​of different windows can be set in a linear decreasing manner, where the weight of the prediction value of the earlier window is smaller and the weight of the prediction value of the closer window is larger.

[0074] For example, if there are five time windows in total, the weight of the most recent window can be set to 1 (maximum weight), and the weight of the most distant window can be set to 0.2 (minimum weight). In this way, the contribution of the most recent window can be highlighted, better reflecting the subject's current attention state and emotional change trend.

[0075] While comprehensively considering the prediction values ​​of multiple windows, more attention is paid to the task performance trend of multimodal physiological signals in the recent window, so that the final attention prediction result is more stable and accurate. It can realize attention prediction by integrating the first attention prediction value of each window and the corresponding time weight, ensuring the accuracy and effectiveness of the final attention prediction result.

[0076] In a time-weighted smoothing calculation process, the multimodal physiological signal detected in real time is divided into windows of length T, and the multimodal physiological signal in a single window is set to S t, the window step size is S. Features are extracted from the multimodal physiological signals in each window and the features are input into the model for prediction.

[0077] For a single window, generate the first attention prediction value y t =f(S t ).

[0078] Where f is the model function; y t is the attention prediction value of window t.

[0079] Perform time-weighted smoothing on the attention prediction values ​​of multiple windows to obtain the attention prediction results:

[0080]

[0081] Among them, ω t =exp(-αt), which is used to represent the time weight; α is the setting coefficient; t is the window number; k is the total number of windows.

[0082] In an optional implementation, the training steps of the attention prediction model include:

[0083] Obtaining a user's multimodal physiological signals and a task performance score corresponding to the multimodal physiological signals in a specific task scenario;

[0084] Determine the attention label of the multimodal physiological signal based on the attention level interval to which the task performance score belongs;

[0085] The attention prediction model is trained based on the model training samples including the attention label and the multimodal physiological signal in combination with a set loss function.

[0086] Among them, the attention label is used to indicate the subject's attention level in a specific task scenario.

[0087] Optionally, the multimodal physiological signals of the subject in a specific task scenario are obtained by acquiring historical multimodal physiological signals or existing multimodal physiological signals in some databases.

[0088] Obtaining the task performance score corresponding to the multimodal physiological signal can be performed by obtaining the task performance score after professionals score the task performance based on the multimodal physiological signal of the subject.

[0089] The task performance scores are mapped to the attention level range to determine the attention labels of the multimodal physiological signals. The task performance data are used to annotate the subject's multimodal physiological signals, and a correspondence is established between the task performance scores and the subject's multimodal physiological signals in the same time period. This enables the attention prediction model to learn the relationship between multimodal physiological signals and task performance, so that after subsequent model training is completed, the attention prediction model can be used to realize attention evaluation and task performance prediction based on multimodal physiological signals.

[0090] Optionally, when determining the attention label of the multimodal physiological signal based on the attention level interval to which the task performance score belongs, the task performance score can be normalized and mapped to the corresponding attention level interval to determine the attention label as high concentration, medium concentration, low concentration, etc.

[0091] For example, normalize the task performance score (such as accuracy, reaction time, etc.) to the interval [0, 1] and set it as P. Define the label according to the attention level interval that the normalized result falls into:

[0092] High concentration: P≥0.8; medium concentration: 0.5≤P<0.8; low concentration: P<0.5.

[0093] After normalizing the task performance scores, the corresponding attention labels are determined, and the attention labels are used to indicate the degree of attention.

[0094] After obtaining the attention label, the data of the multimodal physiological signal is annotated, and then the model training samples are formed based on the annotated multimodal physiological signal to implement the training of the attention prediction model.

[0095] Alternatively, the attention prediction model can be trained directly based on the annotated multimodal physiological signals, or feature fusion processing can be performed on the annotated multimodal physiological signals, and the attention prediction model can be trained using the fused feature vector.

[0096] In an optional embodiment, the attention prediction model is trained based on the model training samples of the attention label and the multimodal physiological signal in combination with a set loss function, including:

[0097] Extract features of multimodal physiological signals in the model training samples to obtain the second eigenvectors of the physiological signals of each modality;

[0098] Performing feature fusion on the second feature vector to generate a fused second fused feature vector;

[0099] Inputting the second fused feature vector into the fully connected layer of the attention prediction model for classification processing to obtain a second attention prediction value;

[0100] Based on the second attention prediction value and the attention label, the attention prediction model is iteratively trained in combination with a set loss function until the number of iterative training of the attention prediction model reaches a threshold or until the attention prediction model converges.

[0101] Among them, the second attention prediction value is used to indicate the subject's attention level in a specific task scenario corresponding to the multimodal physiological signal in the current window.

[0102] The attention prediction model is a classification model that predicts attention based on input data. Optionally, the attention prediction model can be deployed on the device or in the cloud for easy storage and access.

[0103] When extracting features from multimodal physiological signals in model training samples, feature extraction can be performed based on the data types of physiological signals of different modalities.

[0104] The characteristic vectors of physiological signals of each modality include brain wave characteristic vectors, heart rate variability characteristic vectors and motion characteristic vectors.

[0105] For example, the feature vector extracted from the physiological signal of each modality can be expressed as:

[0106] EEG feature vector: X EEG ∈R ne , that is, X EEG is a vector containing ne real elements, belonging to the ne-dimensional real space. The features in the EEG feature vector can include power spectral density, etc.

[0107] HRV eigenvector: X HRV ∈R nh , that is, X HRV is a vector containing nh real elements, belonging to the nh-dimensional real space. The features in the HRV feature vector can include RMSSD, LF / HF, etc.

[0108] Motion feature vector: X Motion ∈R nm , that is, X Motion is a vector containing nm real elements and belongs to nm-dimensional real space. The features in the motion feature vector may include motion amplitude and frequency.

[0109] Feature fusion can be implemented by splicing the feature vectors extracted from the physiological signals of each modality to obtain a fused feature vector; or after splicing the feature vectors extracted from the physiological signals of each modality, the spliced ​​feature vectors can be fused based on the self-attention mechanism of the Transformer model to generate a fused feature vector.

[0110] Optionally, performing feature fusion on the second feature vector to generate a fused second fused feature vector includes:

[0111] Concatenate the feature vectors to obtain the concatenated feature vector:

[0112] X fusion =Concat(X EEG , X HRV , X Motion );

[0113] Among them, X EEG is the brain wave feature vector, X HRV is the heart rate variability feature vector, X Motion is the motion feature vector;

[0114] Based on the self-attention mechanism of the Transformer model, the concatenated feature vectors are subjected to feature fusion to generate the fused feature vector:

[0115]

[0116] Where Q = W Q X fusion ; K=W K X fusion ; V=W V X fusion ;

[0117] Wherein, Z is the fusion feature vector; W Q 、W K 、W V are transformation matrices, which are used to map the concatenated feature vectors to different feature spaces to obtain corresponding linear projection representations; The scaling factor is set.

[0118] In the self-attention mechanism, Q (query), K (key), and V (value) are three sets of representations obtained by linear transformation of input features. They are used to calculate the attention weights between different signals and extract the most relevant information. Q, K, and V are the query matrix, key matrix, and value matrix, respectively, and are derived from X fusion Because EEG, HRV, and motion data in physiological signals from different modalities have different data dimensions and feature distributions, direct correlation calculations cannot be performed. Linear projection allows them to be matched in the same feature space, allowing similarities to be calculated in the same space for data from different modalities. This avoids the high overhead of direct attention calculations on high-dimensional data. Dimensionality reduction can reduce the amount of computation, improve training efficiency, reduce computational complexity, and improve computational stability.

[0119] After projecting Q, K, and V, calculate the attention score QK T The dot product similarity between query and key is calculated; As a scaling factor, it prevents excessive values ​​from affecting gradient stability. Softmax normalizes the weights, making the calculation result a probability distribution. Finally, Z = AV, which generates the fused feature vector.

[0120] The resulting fused feature vector is able to better represent the attention level of ADHD and provide more accurate predictions. Z is the new fused feature, which contains the correlation information of EEG, HRV, and motion signals and is used for the final model classification.

[0121] After generating the fused feature vector, the fused feature vector is input into the fully connected layer of the attention prediction model for classification, and the output result is obtained:

[0122]

[0123] Among them, W and b are weight and bias.

[0124] Optionally, the loss function is set to include a classification loss function term and a feature divergence loss function term.

[0125] In an optional embodiment, the feature vectors of the physiological signals of each modality include brain wave feature vectors, heart rate variability feature vectors, and motion feature vectors. Accordingly, the loss function is set as:

[0126] L=L cls +γL div ;

[0127]

[0128] Among them, L is the loss function, γ is the set weight coefficient, N is the total number of the windows, C is the total number of categories of the attention tags, and y ij is the attention label, is the second attention prediction value.

[0129] Among them, L cls is the classification loss (cross entropy loss); L div This is feature divergence loss, which prevents over-dependence between signal features of different modalities on a single signal. Based on these two function terms, the total loss function L = Lcls + γLdiv is constructed to ensure the effectiveness of model training.

[0130] Optionally, in one embodiment, the method further comprises:

[0131] The actual performance score of the subject's attention in a specific task scenario is obtained; based on the data difference between the actual performance score and the attention prediction result, the model parameters of the attention prediction model are updated to obtain the attention prediction model with updated model parameters.

[0132] During implementation, the parameters of the attention prediction model are updated, which can be:

[0133]

[0134] in, is the attention prediction result output by the model; η is the learning rate; θ is the model parameter of the attention prediction model, and θ' is the updated model parameter.

[0135] The tester may score the subject's actual performance of attention in a specific task scenario based on the subject's actual performance of attention in a specific task scenario.

[0136] The data difference between the actual performance score and the attention prediction result can be used to determine the gap between the model prediction result and the subject's actual performance score. Based on this gap multiplied by the learning rate, the model parameters of the attention prediction model are numerically updated, effectively ensuring the improvement and timely update of the model prediction effect, and ensuring the model prediction performance.

[0137] Step 104: Adjust the task difficulty in a specific task scenario based on the size of the attention prediction result.

[0138] The magnitude of the adjustment of task difficulty in a specific task scenario is positively correlated with the level of attention prediction results.

[0139] Optionally, when the attention prediction result is less than the attention threshold, the task difficulty in a specific task scenario can be reduced. By setting real-time feedback rules and thresholds, when the attention prediction result of the model meets the set conditions, such as setting the attention score threshold θ, if The feedback mechanism is immediately triggered to adjust the task difficulty in real time.

[0140] For example, in a visual tracking task, when the model predicts that the child's attention score is lower than the preset value (such as 6 points), the feedback mechanism is immediately triggered, causing the system to automatically slow down the speed of tracking targets or reduce the number of targets, thereby adjusting the task to an appropriate difficulty in real time to help the subject complete the task better.

[0141] Optionally, when the feedback mechanism is triggered, prompt information such as whether the subject's attention is decreasing or increasing can be given.

[0142] The above implementation process forms a dynamically optimized closed-loop feedback mechanism, integrating data collection, predictive evaluation, and feedback adjustment functions. It can timely adjust the training plan according to the evaluation results of the subjects, adjust the task difficulty in specific task scenarios, and construct a dynamically optimized training and evaluation process to carry out effective testing of subjects, improve the ADHD test results, help improve patient compliance, and ensure the quantitative evaluation effect of ADHD.

[0143] After step 104, the next cycle execution process is performed for each of the above steps, and the process returns to step 101 to continuously collect the multimodal physiological signals of the subject in the specific task scenario until the attention prediction execution is completed.

[0144] The end of the attention prediction execution can be determined by the output of the attention prediction result meeting the set conditions, such as the attention prediction result value being within a set range. Alternatively, it can be determined by the number of cycles of the loop execution process reaching a threshold.

[0145] Furthermore, the method may also include:

[0146] The multimodal physiological signals and attention prediction results of the subjects in specific task scenarios are recorded to form the subjects' attention change trend data.

[0147] This can generate long-term trend reports for parents and doctors to review, improving the efficiency of disease tracking and rehabilitation training for ADHD subjects.

[0148] It should be understood that, although the various steps in the flowcharts involved in the various embodiments described above are displayed in sequence according to the instructions of the arrows, these steps are not necessarily executed in sequence in the order indicated by the arrows. Unless otherwise specified herein, there is no strict order restriction on the execution of these steps, and these steps can be executed in other orders. Moreover, at least a portion of the steps in the flowcharts involved in the various embodiments described above can include multiple steps or multiple stages, and these steps or stages are not necessarily executed and completed at the same time, but can be executed at different times, and the execution order of these steps or stages is not necessarily to be carried out in sequence, but can be executed in turn or alternately with other steps or at least a portion of steps or stages in other steps.

[0149] Based on the same inventive concept, the embodiments of the present application further provide an attention prediction device. The attention prediction device provided in the embodiments of the present application can implement each process of the embodiments of the above-mentioned attention prediction method and can achieve the same technical effects. Therefore, the specific limitations in one or more embodiments of the attention prediction device provided below can refer to the limitations of the attention prediction method above. To avoid repetition, they are not further described here.

[0150] In this embodiment, the computing side can be divided into functional modules according to the above method. For example, each function can be divided into functional modules, or two or more functions can be integrated into one processing module.

[0151] See also Figure 3 , Figure 3 This is a module diagram of an attention prediction device provided in an embodiment of the present application. For the sake of convenience, only the parts related to the embodiment of the present application are shown.

[0152] The attention prediction device 300 includes:

[0153] The acquisition module 301 is used to continuously acquire multimodal physiological signals of the subject in a specific task scenario;

[0154] A window module 302 is configured to segment the multimodal physiological signal into a plurality of windows of preset lengths using a set step size;

[0155] A prediction module 303 is configured to perform attention prediction on the multimodal physiological signal in each window to obtain an attention prediction result; the attention prediction result is used to indicate the attention level of the subject in the specific task scenario;

[0156] The adjustment module 304 is used to adjust the task difficulty in the specific task scenario based on the size of the attention prediction result; and return to the step of continuously collecting the multimodal physiological signals of the subject in the specific task scenario until the attention prediction execution is completed.

[0157] Optionally, the prediction module 303 is specifically configured to:

[0158] Performing feature extraction on the multimodal physiological signal in each of the windows to obtain a first feature vector of the physiological signal of each modality;

[0159] Performing feature fusion on the first feature vector to generate a fused first fused feature vector;

[0160] Inputting the first fused feature vector into an attention prediction model to obtain a first attention prediction value;

[0161] Time-weighted smoothing is performed on the first attention prediction value of each of the windows to obtain the attention prediction result corresponding to the multimodal physiological signal.

[0162] Optionally, the device further comprises:

[0163] Model training module, used for:

[0164] Obtaining a multimodal physiological signal of the user in the specific task scenario and a task performance score corresponding to the multimodal physiological signal;

[0165] Determining an attention tag for the multimodal physiological signal based on the attention level interval to which the task performance score belongs; the attention tag is used to indicate the subject's attention level in the specific task scenario;

[0166] The attention prediction model is trained based on the model training samples including the attention label and the multimodal physiological signal in combination with a set loss function.

[0167] Optionally, the model training module is specifically used to:

[0168] Performing feature extraction on the multimodal physiological signals in the model training samples to obtain second feature vectors of the physiological signals of each modality;

[0169] Performing feature fusion on the second feature vector to generate a fused second fused feature vector;

[0170] Inputting the second fused feature vector into the fully connected layer of the attention prediction model for classification processing to obtain a second attention prediction value;

[0171] Based on the second attention prediction value and the attention label, the attention prediction model is iteratively trained in combination with the set loss function until the number of iterative training of the attention prediction model reaches a threshold or until the attention prediction model converges.

[0172] Optionally, the feature vectors of the physiological signals of each modality include brain wave feature vectors, heart rate variability feature vectors, and motion feature vectors; and the model training module is more specifically used to:

[0173] The feature vectors are concatenated to obtain a concatenated feature vector:

[0174] X fusion =Concat(X EEG , X HRV , X Motion );

[0175] Among them, X EEG is the brain wave feature vector, X HRV is the heart rate variability feature vector, X Motion is the motion feature vector;

[0176] Based on the self-attention mechanism of the Transformer model, the concatenated feature vectors are subjected to feature fusion to generate the fused feature vector:

[0177]

[0178] Where Q = W Q X fusion ; K=W K X fusion ; V=W V X fusion ;

[0179] Wherein, Z is the fusion feature vector; W Q 、W K 、W V are transformation matrices, which are used to map the concatenated feature vectors to different feature spaces to obtain corresponding linear projection representations; The scaling factor is set.

[0180] Optionally, the characteristic vectors of the physiological signals of each modality include brain wave characteristic vectors, heart rate variability characteristic vectors and motion characteristic vectors; and the set loss function is:

[0181] L=L cls +γL div ;

[0182]

[0183] Among them, L is the loss function, γ is the set weight coefficient, N is the total number of the windows, C is the total number of categories of the attention tags, and y ij is the attention label, is the second attention prediction value.

[0184] Optionally, the device further comprises:

[0185] Update modules for:

[0186] Obtaining a score of the subject's actual attention performance in the specific task scenario;

[0187] Based on the data difference between the actual performance score and the attention prediction result, the model parameters of the attention prediction model are updated to obtain the attention prediction model after the model parameters are updated.

[0188] The above integrated modules can be implemented in the form of hardware. It should be noted that the division of modules in this embodiment is schematic and is only a logical function division. In actual implementation, there may be other division methods.

[0189] It should be noted that the relevant contents of each step involved in the above method embodiment can be referred to the functional description of the corresponding functional module and will not be repeated here.

[0190] In one embodiment, Figure 4 As shown, an electronic device is provided. The electronic device 4 of this embodiment includes: at least one processor 400 ( Figure 4 Only one is shown), a memory 401 and a computer program 402 stored in the memory 401 and executable on the at least one processor 400, wherein the processor 400 implements the steps of any of the above-mentioned method embodiments when executing the computer program 402.

[0191] The electronic device 4 may be an electronic atomization device, such as an HNB device. The electronic device 4 may include, but is not limited to, a processor 400 and a memory 401. Those skilled in the art will understand that Figure 4 It is only an example of the electronic device 4 and does not constitute a limitation of the electronic device 4. It may include more or fewer components than shown in the figure, or a combination of certain components, or different components. For example, the electronic device may also include input and output devices, network access devices, buses, etc.

[0192] The processor 400 may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or any conventional processor.

[0193] The memory 401 may be an internal storage unit of the electronic device 4, such as a hard disk or memory of the electronic device 4. The memory 401 may also be an external storage device of the electronic device 4, such as a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. equipped on the electronic device 4. Furthermore, the memory 401 may include both an internal storage unit of the electronic device 4 and an external storage device. The memory 401 is used to store the computer program and other programs and data required by the electronic device. The memory 401 may also be used to temporarily store data that has been output or is about to be output.

[0194] Those skilled in the art can clearly understand that, for the convenience and brevity of description, only the division of the above-mentioned functional units and modules is used as an example for illustration. In actual applications, the above-mentioned functions can be distributed and completed by different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiment can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of software functional units. In addition, the specific names of the functional units and modules are only for the convenience of distinguishing each other, and are not used to limit the scope of protection of this application. The specific working process of the units and modules in the above-mentioned system can refer to the corresponding process in the aforementioned method embodiment, and will not be repeated here.

[0195] In the above embodiments, the description of each embodiment has its own focus. For parts that are not described or recorded in detail in a certain embodiment, reference can be made to the relevant description of other embodiments.

[0196] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0197] In the embodiments provided in the present application, it should be understood that the disclosed devices / electronic devices and methods can be implemented in other ways. For example, the device / electronic device embodiments described above are merely schematic. For example, the division of the modules or units is merely a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.

[0198] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.

[0199] In addition, the functional units in the various embodiments of the present application may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.

[0200] If the integrated module / unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the present application implements all or part of the process in the above-mentioned embodiment method, and can also be completed by instructing the relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium, and when the computer program is executed by the processor, it can implement the steps of the above-mentioned various method embodiments. Among them, the computer program includes computer program code, and the computer program code can be in source code form, object code form, executable file or some intermediate form. The computer-readable medium may include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electric carrier signal, telecommunication signal and software distribution medium. It should be noted that the content contained in the computer-readable medium can be appropriately increased or decreased according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, computer-readable media do not include electric carrier signals and telecommunication signals.

[0201] The present application implements all or part of the processes in the above-mentioned embodiment methods, and may also be implemented through a computer program product. When the computer program product runs on an electronic device, the electronic device can implement the steps in the above-mentioned method embodiments when executing the computer program product.

[0202] The above-described embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. These modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present application, and should all be included in the scope of protection of the present application.

Claims

1. A method for predicting attention, characterized in that: include: Continuously collect multimodal physiological signals of subjects in specific task scenarios; Segmenting the multimodal physiological signal into a plurality of windows of preset lengths with a set step size; Performing attention prediction on the multimodal physiological signal in each of the windows to obtain an attention prediction result; the attention prediction result is used to indicate the attention level of the subject in the specific task scenario; Adjusting the difficulty of the task in the specific task scenario based on the size of the attention prediction result; Return to the step of continuously collecting the subject's multimodal physiological signals in a specific task scenario until the attention prediction execution is completed.

2. The method according to claim 1, characterized in that The performing attention prediction on the multimodal physiological signal in each of the windows to obtain an attention prediction result includes: Performing feature extraction on the multimodal physiological signal in each of the windows to obtain a first feature vector of the physiological signal of each modality; Performing feature fusion on the first feature vector to generate a fused first fused feature vector; Inputting the first fused feature vector into an attention prediction model to obtain a first attention prediction value; Time-weighted smoothing is performed on the first attention prediction value of each of the windows to obtain the attention prediction result corresponding to the multimodal physiological signal.

3. The method according to claim 2, characterized in that The training steps of the attention prediction model include: Obtaining a multimodal physiological signal of the user in the specific task scenario and a task performance score corresponding to the multimodal physiological signal; Determining an attention tag for the multimodal physiological signal based on the attention level interval to which the task performance score belongs; the attention tag is used to indicate the subject's attention level in the specific task scenario; The attention prediction model is trained based on the model training samples including the attention label and the multimodal physiological signal in combination with a set loss function.

4. The method according to claim 3, characterized in that The training of the attention prediction model based on the model training samples including the attention label and the multimodal physiological signal in combination with a set loss function includes: Performing feature extraction on the multimodal physiological signals in the model training samples to obtain second feature vectors of the physiological signals of each modality; Performing feature fusion on the second feature vector to generate a fused second fused feature vector; Inputting the second fused feature vector into the fully connected layer of the attention prediction model for classification processing to obtain a second attention prediction value; Based on the second attention prediction value and the attention label, the attention prediction model is iteratively trained in combination with the set loss function until the number of iterative training of the attention prediction model reaches a threshold or until the attention prediction model converges.

5. The method according to claim 3, characterized in that The characteristic vectors of the physiological signals of each modality include brain wave characteristic vectors, heart rate variability characteristic vectors and motion characteristic vectors; The performing feature fusion on the second feature vector to generate a fused second fused feature vector includes: The feature vectors are concatenated to obtain a concatenated feature vector: X fusion =Concat(X EEG ,X HRV ,X Motion ); Among them, X EEG is the brain wave feature vector, X HRV is the heart rate variability feature vector, X Motion is the motion feature vector; Based on the self-attention mechanism of the Transformer model, the concatenated feature vectors are subjected to feature fusion to generate the fused feature vector: where Q = W Q X fusion ; K = W K X fusion ; V = W V X fusion ; Wherein, Z is the fusion feature vector; W Q 、W K 、W V are transformation matrices, which are used to map the concatenated feature vectors to different feature spaces to obtain corresponding linear projection representations; The scaling factor is set.

6. The method according to claim 5, characterized in that The characteristic vectors of the physiological signals of each modality include brain wave characteristic vectors, heart rate variability characteristic vectors and motion characteristic vectors; the set loss function is: L=L cls +γL div ; Among them, L is the loss function, γ is the set weight coefficient, N is the total number of the windows, C is the total number of categories of the attention tags, and y ij is the attention label, is the second attention prediction value.

7. The method according to claim 2, characterized in that Also includes: Obtaining a score of the subject's actual attention performance in the specific task scenario; Based on the data difference between the actual performance score and the attention prediction result, the model parameters of the attention prediction model are updated to obtain the attention prediction model after the model parameters are updated.

8. An attention prediction device, characterized in that include: The acquisition module is used to continuously collect multimodal physiological signals of the subject in a specific task scenario; a window module, configured to divide the multimodal physiological signal into a plurality of windows of preset lengths with a set step size; A prediction module, configured to perform attention prediction on the multimodal physiological signal in each window to obtain an attention prediction result; the attention prediction result is used to indicate the attention level of the subject in the specific task scenario; An adjustment module is used to adjust the task difficulty in the specific task scenario based on the size of the attention prediction result; and return to the step of continuously collecting the multimodal physiological signals of the subject in the specific task scenario until the attention prediction execution is completed.

9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 7 are implemented.

10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.