Multi-modal human physiological data classification model and training method therefor, multi-modal human physiological data classification method, and device
By adopting the multi-headed self-attention module and a fusion expert system in the multimodal human physiological data classification model, the problem of data loss affecting the performance of the model is solved, and more accurate and robust classification results are achieved.
Patent Information
- Application Number
- PCT/CN2024/133264
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-26
- Filing Date
- 2024-11-20
- Publication Date
- 2025-05-30
AI Technical Summary
During the human physiological data collection process, data is missing due to poor sensor contact or equipment interference, which in turn affects the accuracy of the classification results of the model.
A classification model of multimodal human physiological data is proposed, including multi-headed self-attention module, normalization module, fusion expert system and decision-making module. This model can maintain good performance in the absence of data by extracting and fusion of multimodal synchronous data.
Through the fusion and feature extraction of multimodal data, the robustness of the model and the accuracy of classification results are improved, and good performance can be maintained when data is missing.
Smart Images

Figure CN2024133264_30052025_PF_FP_ABST
Abstract
Description
Classification model, training and classification method and equipment for multimodal human physiological data
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS
[0002] This application claims the Chinese patent application filed with the China Patent Office on November 30, 2023, with application number 202311635615.7 and invention name “Classification model, training and classification method and device for multimodal human physiological data”, and the Chinese patent application filed with the China Patent Office on November 23, 2023, with application number 202311576945.3 and invention name “Training method, sample data generation and prediction method for physiological state prediction model”, and the Chinese patent application filed with the China Patent Office on November 23, 2023, with application number 202311576945.3 and invention name “Training method, sample data generation and prediction method for physiological state prediction model”. The priority of the Chinese patent application filed with the China Patent Office on December 22, application number 202311790144.7, invention name “Physiological signal processing method, device and electronic device based on human intelligence”, and the priority of the Chinese patent application filed with the China Patent Office on December 26, 2023, application number 202311810607.1, invention name “Physiological signal processing method, device and electronic device based on human intelligence”, all contents of which are incorporated by reference in this application. Technical Field
[0003] The present invention relates to the fields of ergonomics and artificial intelligence technology, and in particular to a classification model, training and classification method, and equipment for multimodal human physiological data. Background Art
[0004] In human factors-related research and application scenarios, people often hope to use neural network models to classify the user's current physiological or psychological state based on human physiological data (such as electroencephalogram, electrocardiogram, etc.), for example, to determine whether the user is currently in a happy or angry state.
[0005] However, during the physiological data collection process, the collected physiological data may be missing due to poor contact of sensors or electrodes, user movement, interference from other equipment, etc., which may lead to deviations in the classification results of the model. Summary of the Invention
[0006] This application proposes a classification model, training and classification method and equipment for multimodal human physiological data.
[0007] In a first aspect of the present application, a classification model for multimodal human physiological data is proposed, wherein the classification model comprises: a multi-headed self-attention module, a normalization module, a fusion expert system, and a decision module;
[0008] The fusion expert system includes: an EEG expert subsystem, an ECG expert subsystem, an electrodermal expert subsystem and a multimodal synchronous fusion expert subsystem;
[0009] The multi-head self-attention module is used to extract features from the input multimodal synchronized data and output them to the normalization module; the multimodal synchronized data includes: physiological indicators of the user; the physiological indicators include: electroencephalogram (EEG) data, electrocardiogram (ECG) data and electrodermal activity (EDA) data;
[0010] The normalization module is used to generate normalized feature data based on the extracted feature data and output the normalized feature data to the fusion expert system;
[0011] The EEG expert subsystem is configured to perform a classification task based on the normalized feature data corresponding to the user's EEG data to obtain a first classification result;
[0012] The ECG expert subsystem is configured to perform a classification task based on the normalized feature data corresponding to the user's ECG data to obtain a second classification result;
[0013] The skin electrical expert subsystem is configured to perform a classification task based on the normalized feature data corresponding to the user's skin electrical data to obtain a third classification result;
[0014] The multimodal synchronous fusion expert subsystem is used to perform a classification task according to the normalized feature data corresponding to the multimodal synchronous data to obtain a fourth classification result;
[0015] The decision module is used to calculate a final classification result based on the first classification result, the second classification result, the third classification result and the fourth classification result, as well as the weight corresponding to each classification result.
[0016] In a second aspect of the present application, a training method for a multimodal human physiological data classification model is proposed, which is applicable to the multimodal human physiological data classification model described above, and the training method comprises:
[0017] Using the EEG training set to train the EEG expert subsystem and the multi-head self-attention module, and freezing the weights of the EKG expert subsystem, the GEP expert subsystem, and the multimodal synchronous fusion expert subsystem;
[0018] The ECG expert subsystem is trained using the ECG training set, and the weights of the multi-head self-attention module, the EEG expert subsystem, the electrodermal expert subsystem, and the multimodal synchronous fusion expert subsystem are frozen;
[0019] The electrodermal training set is used to train the electrodermal expert subsystem, and the weights of the multi-head self-attention module, the EEG expert subsystem, the ECG expert subsystem, and the multimodal synchronous fusion expert subsystem are frozen;
[0020] The classification model is trained using a multimodal random coding training set, thereby adjusting the weights of the multi-head self-attention module, the EEG expert subsystem, the ECG expert subsystem, the electrodermal expert subsystem, and the multimodal synchronous fusion expert subsystem.
[0021] In a third aspect of the present application, a classification method for multimodal human physiological data is proposed, wherein the method utilizes the classification model for multimodal human physiological data as described above to perform a classification task.
[0022] In a fourth aspect of the present application, a computer-readable storage device is provided, storing a computer program that can be loaded by a processor and execute the method described above.
[0023] This application has the following beneficial effects:
[0024] This application has made improvements based on the original VLMo (Vision Language pretrained Model) to make the classification model more adaptable to the physiological data of the human body. It can not only make full use of the information of each modality (EEG data, ECG data and skin conduction data) but also realize multimodal fusion of data, so that the performance of the model does not only depend on the quality of a certain modality data, thereby improving the robustness of the model.
[0025] In order to ensure that the classification model maintains good performance in the case of missing data, the preprocessed data is artificially masked during the model training phase to simulate the situation of missing data, so that the trained model can generalize the masked part, which is the missing data part in actual situation, and thus obtain the most reliable results.
[0026] Considering that EEG signals contain information in three dimensions: time, frequency, and space, this application converts EEG signals into frequency domain information and then obtains quasi-multispectral images based on different leads according to the positions of each electrode point when the EEG signals were collected. This quasi-multispectral image is then processed in blocks, and EEG data as model input data is generated based on the block data. This method fully exploits the information contained in the multidimensional space and improves the efficiency of data use. BRIEF DESCRIPTION OF THE DRAWINGS
[0027] FIG1 is a schematic diagram of the main components of an embodiment of a classification model for multimodal human physiological data of the present application;
[0028] FIG2 is a schematic diagram of the main steps of an embodiment of a training method for a multimodal human physiological data classification model of the present application;
[0029] FIG3 is a schematic diagram of a system architecture provided in an embodiment of the present application;
[0030] FIG4 is a flow chart of a method for generating EEG signal sample data according to an embodiment of the present application;
[0031] FIG5 is a flow chart of a method for training a physiological state prediction model according to an embodiment of the present application;
[0032] FIG6 is a schematic diagram of a physiological state prediction model provided in an embodiment of the present application;
[0033] FIG7 is a schematic diagram of the structure of a recognition model provided in an embodiment of the present application;
[0034] FIG8 is a schematic diagram of the structure of a channel attention network provided by an embodiment of the present application;
[0035] FIG9 is a schematic diagram of the structure of a spatial attention network provided by an embodiment of the present application;
[0036] FIG10 is a schematic diagram of the structure of another recognition model provided in an embodiment of the present application;
[0037] FIG11 is a schematic diagram of the structure of a physiological signal recognition network provided in an embodiment of the present application;
[0038] FIG12 is a schematic diagram of a time window provided according to an embodiment of the present application;
[0039] FIG13 is a schematic diagram of the structure of a neural network model provided according to an embodiment of the present application;
[0040] FIG14 is a schematic diagram of an original electrocardiogram signal provided according to an embodiment of the present application;
[0041] FIG15 is a schematic diagram of a target electrocardiogram signal provided according to an embodiment of the present application;
[0042] FIG16 is a schematic diagram of another target electrocardiogram signal provided according to an embodiment of the present application;
[0043] FIG17 is a schematic structural diagram of an electronic device provided according to an embodiment of the present application. DETAILED DESCRIPTION
[0044] The preferred embodiments of the present application are described below with reference to the accompanying drawings. Those skilled in the art should understand that these embodiments are only used to explain the technical principles of the present application and are not intended to limit the scope of protection of the present application.
[0045] To make the purpose, technical solutions, and advantages of the embodiments of this application more clear, the technical solutions in the embodiments of this application will be clearly and completely described below in conjunction with the drawings in the embodiments of this application. Obviously, the described embodiments are part of the embodiments of this application, not all of the embodiments. Based on the embodiments of this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0046] It should be noted that, in the description of this application, the terms "first" and "second" are merely for the convenience of description, and do not indicate or imply the relative importance of the devices, elements or parameters, and therefore should not be understood as limiting this application. In addition, the term "and / or" in this application is merely a description of the association relationship between associated objects, indicating that three relationships may exist. For example, A and / or B can represent: A exists alone, A and B exist at the same time, and B exists alone. In addition, the character " / " in this article, unless otherwise specified, generally indicates that the related objects before and after are in an "or" relationship.
[0047] Example 1
[0048] Microsoft proposed a unified vision language model VLMo (Vision Language pretrained Model). VLMo is equivalent to a hybrid expert model. Its FFN (Feed Forward Network) part has three modes: the visual mode for image encoding (V-FFN), the language mode for text encoding (L-FFN), and the visual language mode for image-text fusion (VL-FFN).
[0049] Figure 1 is a schematic diagram of the main components of an embodiment of the multimodal human physiological data classification model of the present application. As shown in Figure 1, the classification model of this embodiment includes: a multi-head self-attention module 10, a normalization module 20, a fusion expert system 30, and a decision module 40.
[0050] The fusion expert system 30 includes an EEG expert subsystem 31, an ECG expert subsystem 32, an EGD expert subsystem 33, and a multimodal synchronization fusion expert subsystem 34. The multimodal synchronization data includes the user's physiological indicators, which include EEG data, ECG data, and EGD data, all collected within the same time period.
[0051] In this embodiment, the multi-head self-attention module 10 is used to extract features from the input multimodal synchronous data and output it to the normalization module 20.
[0052] In this embodiment, the normalization module 20 is used to normalize the extracted feature data and output the normalized feature data to the fusion expert system 30 .
[0053] In this embodiment, the EEG expert subsystem 31 is used to perform classification tasks based on the normalized feature data corresponding to the user's EEG data to obtain a first classification result; the ECG expert subsystem 32 is used to perform classification tasks based on the normalized feature data corresponding to the user's ECG data to obtain a second classification result; the skin electrodermal expert subsystem 33 is used to perform classification tasks based on the normalized feature data corresponding to the user's skin electrodermal data to obtain a third classification result; the multimodal synchronous fusion expert subsystem 34 is used to perform classification tasks based on the normalized feature data corresponding to the multimodal synchronous data to obtain a fourth classification result.
[0054] In this embodiment, the decision module 40 is used to calculate the final classification result according to the first classification result, the second classification result, the third classification result and the fourth classification result, as well as the weight corresponding to each classification result.
[0055] Specifically, the final classification result is calculated according to the following formula (1): G = W1*f1+W2*f2+W3*f3+W4*f4 (1)
[0056] Among them, G represents the final classification result, f1, f2, f3 and f4 represent the first classification result, second classification result, third classification result and fourth classification result respectively, and W1, W2, W3 and W4 all represent weights.
[0057] FIG2 is a schematic diagram of the main steps of an embodiment of a training method for a multimodal human physiological data classification model of the present application. The training method of this embodiment is applicable to the classification model of the multimodal human physiological data described above. As shown in FIG2 , the training method of this embodiment includes steps S10-S40:
[0058] Step S10: Use the EEG training set to train the EEG expert subsystem and the multi-head self-attention module, and freeze the weights of the ECG expert subsystem, the skin electrodensis expert subsystem, and the multimodal synchronous fusion expert subsystem.
[0059] Step S20: Use the ECG training set to train the ECG expert subsystem, and freeze the weights of the multi-head self-attention module, the EEG expert subsystem, the electrodermal expert subsystem, and the multimodal synchronous fusion expert subsystem.
[0060] Step S30: Use the skin electrodermal training set to train the skin electrodermal expert subsystem, and freeze the weights of the multi-head self-attention module, the EEG expert subsystem, the ECG expert subsystem, and the multimodal synchronous fusion expert subsystem.
[0061] Step S40: Use the multimodal random coding training set to train the classification model, thereby adjusting the weights of the multi-head self-attention module, the EEG expert subsystem, the ECG expert subsystem, the skin electrode expert subsystem, and the multimodal synchronous fusion expert subsystem.
[0062] In an optional embodiment, the method for training the classification model further includes steps S1-S4 of preprocessing the data:
[0063] Step S1: Generate EEG data containing time domain information, frequency domain information, and spatial information based on the EEG signal of each user and the corresponding electrode point position, and then obtain an EEG training set.
[0064] Step S2: De-noise the ECG signal of each user and align it with the EEG signal of the user based on time information to generate an ECG training set.
[0065] Because for a certain user, only the EEG signals, ECG signals and skin conduction signals collected at the same time have relevant meaning, so these data must be aligned in time.
[0066] Step S3: Denoise the skin electrodermal signal of each user and align it with the EEG signal of the user according to time information, thereby generating a skin electrodermal training set.
[0067] Step S4: Perform random masking on the EEG training set, the ECG training set, and the electrodermal training set to simulate data missing conditions, and generate a multimodal randomly coded training set.
[0068] Preferably, the above step S1 may specifically include steps S11-S14:
[0069] Step S11: Perform fast Fourier transform on the EEG signal of each user to extract frequency domain information.
[0070] Step S12: selecting data of a frequency band of interest from the frequency domain information.
[0071] For example, for cognitive load tasks, we can use data from the theta (4-7Hz), alpha (8-13Hz), and beta (13-30Hz) frequency bands, which are closely related to memory.
[0072] Step S13: obtaining quasi-multispectral images based on different leads according to the data of the frequency band of interest and the electrode point positions when collecting the EEG signal.
[0073] Step S14: performing block processing on the quasi-multispectral image and generating the user's EEG data based on the block data.
[0074] The EEG signals collected by an EEG instrument are time series data. Conventional methods transform them into the frequency domain, thereby using information in both the time and frequency domains. This approach ignores the spatial information contained in the EEG signals. To fully utilize information from various dimensions, this application performs special processing on the EEG signals, thereby obtaining information in three dimensions: time, frequency, and space, thereby improving information utilization.
[0075] In a preferred embodiment, the above step S11 may further include:
[0076] (1) The electrode point positions when collecting the EEG signal are projected from three-dimensional space to a two-dimensional surface to obtain two-dimensional position information of each electrode point. Specifically, the electrode point positions when collecting the EEG signal can be converted from three-dimensional space to a two-dimensional surface using azimuth equidistant projection in a Cartesian coordinate system, so that the converted lead data have a spatial topological relationship.
[0077] EEG signals are time series collected by an EEG cap worn on the head. The distribution of the EEG cap's electrodes on the scalp has a spatial topological structure, so the collected EEG signals contain both temporal and spatial information. To convert the spatially distributed activity map into a two-dimensional image, the positions of the electrode points need to be projected from three-dimensional space onto a two-dimensional surface. To achieve this, we use azimuthal equidistant projection in a Cartesian coordinate system. Compared to other projection methods, this projection method not only converts coordinate points from three dimensions to two dimensions, but also preserves the distance relationship between three-dimensional points, resulting in a spatial topological relationship between each lead in the converted EEG data.
[0078] (2) Calculate the distance between each electrode point and other surrounding electrode points based on the two-dimensional position information.
[0079] (3) For the data of each lead in the frequency band of interest, corresponding weights are set for the data of other leads according to the calculated distance, thereby obtaining a multispectral image based on different leads.
[0080] Optionally, the above step S14 may further include:
[0081] (1) Divide the multispectral image into blocks, N = H * W / P 2 small blocks of the same size; where H, W, C, and P represent the height, width, number of channels, and size of the small block, respectively.
[0082] (2) Flatten the image in each patch into a vector, obtain the patch embedding through linear projection, and add a [I_CLS] token as the position code of the patch to obtain the block data corresponding to each patch. The position code is used to record the position of the patch in the image during the block division.
[0083] (3) Generate the user's EEG data based on the block data.
[0084] In this application, the multispectral image is processed as a picture, and the width and height of the picture represent the spatial distribution of cerebral cortical activity. In this embodiment, the EEG data is processed into a 32*32 grid structure (P=32).
[0085] Although the various steps in the above embodiment are described in the above-mentioned order, those skilled in the art will understand that in order to achieve the effect of this embodiment, different steps do not have to be executed in such an order. They can be executed simultaneously (in parallel) or in a reverse order. These simple changes are within the scope of protection of this application.
[0086] Furthermore, based on the above classification model, the present application also provides an embodiment of a classification method for multimodal human physiological data. In this embodiment, the classification task is performed using the above-described classification model for multimodal human physiological data.
[0087] Before performing the classification task, it is also necessary to refer to the preprocessing steps S1-S3 in the above training method embodiment to preprocess the user's EEG signals, ECG signals and skin electrode signals respectively, that is, to generate EEG data containing time domain information, frequency domain information and spatial information based on the user's EEG signals and the corresponding electrode point positions, and to denoise and align the ECG signals and skin electrode signals.
[0088] Furthermore, the present application also provides an embodiment of a computer-readable storage device. The storage device of this embodiment stores a computer program that can be loaded by a processor and execute the method described above.
[0089] The computer-readable storage device may include: a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, and other devices that can store program codes.
[0090] Example 2
[0091] In some examples, in order to determine the physiological state of the human body under the control of EEG signals, a physiological state prediction model can be used to predict the physiological state of the human body based on EEG signals. EEG signals can be obtained based on electroencephalograms (EEGs). Physiological state prediction models include, for example, deep learning models. Specifically, end-to-end deep learning models can be used to predict physiological states. Although end-to-end deep learning models can automatically learn some features with better distinguishability, they often fit to some non-important features, causing the model to partially collapse onto some bad features, making the model's prediction accuracy for physiological states not high and the prediction effect poor.
[0092] In view of this, the embodiments of the present application provide an optimized method for generating EEG signal sample data, a method for training a physiological state prediction model, and a physiological state prediction method. The system architecture shown in Figure 3 can be used to implement at least one of the method for generating EEG signal sample data, the method for training a physiological state prediction model, and the physiological state prediction method. As shown in Figure 3, the system architecture 100 includes an execution device 110, a training device 120, a database 130, a client device 140, a data storage system 150, a data acquisition device 160, and a sample data generation device 170. The execution device 110 includes a computing module 111 and an I / O interface 112, and the computing module 111 stores a trained physiological state prediction model 101.
[0093] The data acquisition device 160 is used to collect raw EEG signal sample data and store it in the database 130 .
[0094] The sample data generating device 170 performs data processing based on the original EEG signal sample data stored in the database 130 to obtain EEG signal priori features, and generates target EEG signal sample data based on the EEG signal priori features.
[0095] The training device 120 stores the physiological state prediction model to be trained. The training device 120 performs auxiliary training on the physiological state prediction model to be trained based on the target EEG signal sample data. The general training method adopts the deep learning method, among which auxiliary learning, transfer learning, reinforcement learning, and popular learning are some network branches added in the training stage (i.e., the training stage), which can better complete the training model task in the neural network model. The auxiliary training mentioned in this application adopts the auxiliary learning training method, and the physiological state auxiliary prediction model can be obtained after the auxiliary training. The training device 120 also trains the physiological state auxiliary prediction model based on the original EEG signal sample data stored in the database 130 to obtain the trained physiological state prediction model 101. The trained physiological state prediction model 101 obtained by the training device 120 can be applied to different systems or devices, for example, to the execution device 110.
[0096] The execution device 110 is configured with an I / O interface 112, through which data is exchanged with external devices, such as a client device 140. A user can input raw EEG signal data into the I / O interface 112 via the client device 140. This raw EEG signal data can be collected by the data acquisition device 160 and stored in the database 130. The data can then be confirmed by the user or processed accordingly before being input into the execution device 110. Furthermore, the user can also input instructions to the I / O interface 112 via the client device 140, instructing the execution device 110 to perform data processing operations.
[0097] The execution device 110 can call data, code, etc. in the data storage system 150 , and can also store data, instructions, etc. in the data storage system 150 .
[0098] The calculation module 111 stores the trained physiological state prediction model 101 . The calculation module 111 uses the trained physiological state prediction model 101 to predict the physiological state based on the input original EEG signal data.
[0099] The prior feature data mentioned in this application is obtained based on the original EEG signal sample data, that is, in general, the prior features can be extracted from the original EEG signal sample data.
[0100] Next, the I / O interface 112 may return the prediction result to the client device 140 to provide it to the user.
[0101] In the case shown in FIG1 , the user can manually specify the raw EEG signal data to be input into the execution device 110, for example, by performing an input operation in the interface provided by the I / O interface 112. In another case, the client device 140 can automatically input the raw EEG signal sample data into the I / O interface 112. For example, the client device 140 automatically inputs the raw EEG signal sample data after obtaining the user's authorization, and the user can set the corresponding permissions in the client device 140. The user can view the predicted results of the physiological state output by the execution device 110 on the client device 140, and the specific presentation form can be a specific method such as display, sound, action, etc. The client device 140 can also store the predicted results of the predicted physiological state in the database 130.
[0102] It is worth noting that Figure 3 is only a schematic diagram of a system architecture provided by an embodiment of the present application. The positional relationship between the devices, components, modules, etc. shown in the figure does not constitute any limitation. For example, in Figure 1, the data storage system 150 is an external memory relative to the execution device 110. In other cases, the data storage system 150 can also be placed in the execution device 110.
[0103] FIG4 is a flow chart of a method for generating EEG signal sample data provided in an embodiment of the present application.
[0104] As shown in FIG4 , the method 200 for generating EEG signal sample data provided in an embodiment of the present application includes, for example, steps S210 - S220 .
[0105] Step S210 , processing the original EEG signal sample data to obtain EEG signal prior features.
[0106] For example, the raw EEG signal sample data can be EEG signal data collected based on a brain-computer interface, or data obtained by preprocessing the collected EEG signal data. Preprocessing includes bandpass filtering the collected EEG signal data, removing redundant data, and other processing. Bandpass filtering the collected EEG signal data includes removing information with excessively high or low frequencies in the EEG signal data.
[0107] For example, EEG signal prior features represent prior knowledge about the original EEG signal sample data. These features typically include important or highly discriminative information within the original EEG signal sample data. Incorporating prior knowledge into model training can optimize model performance, reduce training difficulty, and improve generalization.
[0108] Step S220: generating target EEG signal sample data based on the EEG signal priori features.
[0109] Exemplarily, the target EEG signal sample data is used to assist in training a physiological state prediction model to be trained to obtain a physiological state auxiliary prediction model. The physiological state prediction model to be trained includes, for example, a deep learning model, specifically a convolutional neural network (CNN) model or other models based on the CNN model.
[0110] After the physiological state auxiliary prediction model is trained, it can be further trained based on the original EEG signal sample data to obtain a trained physiological state prediction model. The trained physiological state prediction model can be used to predict physiological states, including motor imagery states, emotional states, etc.
[0111] In an embodiment of the present application, EEG signal prior features are obtained based on the original EEG signal sample data, and target EEG signal sample data is generated based on the EEG signal prior features. Since the target EEG signal sample data contains prior information, auxiliary training of the physiological state prediction model to be trained based on the target EEG signal sample data can improve the prediction accuracy of the model. After the physiological state auxiliary prediction model is trained, the physiological state auxiliary prediction model is further trained based on the original EEG signal sample data to obtain a trained physiological state prediction model, which further improves the prediction accuracy of the model. It can be seen that adding EEG signal prior features for auxiliary training when training the model can optimize the performance of the model, reduce the training difficulty of the model, improve the generalization ability of the model, and improve the prediction accuracy of the model.
[0112] It is understood that the physiological state prediction model in Example 2 of the present application has the same function and implementation as the EEG expert subsystem in Example 1 above. The training of the physiological state prediction model in Example 2 can also be regarded as training the EEG expert subsystem, that is, the physiological state prediction model in Example 2 can be replaced by the EEG expert subsystem. In addition, in this case, the EEG training set used to train the EEG expert subsystem is equivalent to including the original EEG signal sample data and the target EEG signal sample data.
[0113] Next, the process of generating target EEG signal sample data based on EEG signal prior features will be further described. The target EEG signal sample data includes first target EEG signal sample data and second target EEG signal sample data.
[0114] In one example, first target EEG signal sample data can be constructed based on EEG signal prior features. For example, the EEG signal prior features can be directly used as the first target EEG signal sample data. Or when the data format of the first target EEG signal sample data is predefined as a preset data format, the EEG signal prior features can be subjected to related processing such as disassembly or transformation according to the preset data format to construct the first target EEG signal sample data. It can be understood that directly constructing the first target EEG signal sample data for auxiliary training of the model based on the EEG signal prior features allows the model to focus on prior knowledge that is conducive to classification, thereby improving the physiological state auxiliary prediction model's ability to process important information in a targeted manner.
[0115] In another example, the original EEG signal sample data and the EEG signal prior features can be merged to obtain the second target EEG signal sample data. For example, the original EEG signal sample data and the EEG signal prior features can be decomposed and fused to obtain the second target EEG signal sample data. Or the original EEG signal sample data and the EEG signal prior features can be spliced to obtain the second target EEG signal sample data. For example, when the original EEG signal sample data is vector or matrix data, the EEG signal prior features can be spliced in a specific position of the vector or matrix data, and the specific position includes the head position, the tail position, and any position in the middle. In some cases, the specific position preferably the tail position can improve the training effect of the model and the prediction accuracy of the model. It can be understood that the original EEG signal sample data and the EEG signal prior features are merged to obtain the second target EEG signal sample data for auxiliary training of the model, so that the model pays attention to the prior knowledge while taking into account the original data, thereby improving the generalization of the physiological state auxiliary prediction model.
[0116] FIG5 is a flow chart of a method for training a physiological state prediction model according to an embodiment of the present application.
[0117] As shown in FIG5 , the training method 300 of the physiological state prediction model provided in an embodiment of the present application includes, for example, steps S310 - S330 .
[0118] Step S310: Acquire target EEG signal sample data.
[0119] Exemplarily, the target EEG signal sample data is determined based on EEG signal prior features, and the EEG signal prior features are obtained based on original EEG signal sample data.
[0120] Step S320 : Based on the target EEG signal sample data, auxiliary training is performed on the physiological state prediction model to be trained to obtain an auxiliary physiological state prediction model.
[0121] Step S330 : training the physiological state auxiliary prediction model based on the original EEG signal sample data to obtain a trained physiological state prediction model.
[0122] It can be understood that the embodiments of the present application obtain EEG signal prior features based on the original EEG signal sample data, and obtain target EEG signal sample data based on the EEG signal prior features. Since the target EEG signal sample data contains prior information, the physiological state prediction model to be trained is assisted in training based on the target EEG signal sample data, thereby improving the prediction accuracy of the model. After the physiological state auxiliary prediction model is trained, the physiological state auxiliary prediction model is further trained based on the original EEG signal sample data to obtain a trained physiological state prediction model, thereby further improving the prediction accuracy of the model. In addition, the embodiments of the present application can directly use the original EEG signal sample data as the input of the model, avoiding the tedious process of preprocessing the sample data and the process of manual feature selection, thereby saving time and labor costs.
[0123] The embodiment of the present application trains the model twice based on the auxiliary learning method. The training of the model based on the auxiliary learning method enables the model to learn more detailed classification weights, and at the same time enables the model to pay more attention to features that are conducive to classification. During the first training, the physiological state auxiliary prediction model is obtained based on the target EEG signal sample data with prior knowledge. The first training enables the model to learn the ability to predict the physiological state based on prior knowledge. After obtaining the physiological state auxiliary prediction model, the physiological state auxiliary prediction model is directly migrated for the second training. During the second training, the physiological state auxiliary prediction model is trained based on the original EEG signal sample data. Since the physiological state auxiliary prediction model has undergone the first training, the relevant parameters of the model can be fine-tuned during the second training, thereby further improving the prediction accuracy of the model while improving the training efficiency.
[0124] Next, the process of obtaining target EEG signal sample data will be further described. The target EEG signal sample data includes first target EEG signal sample data and second target EEG signal sample data.
[0125] In one example, the original EEG signal sample data can be processed to obtain EEG signal prior features, and then the first target EEG signal sample data can be constructed based on the EEG signal prior features. For example, the EEG signal prior features can be directly used as the first target EEG signal sample data, or the EEG signal prior features can be disassembled or transformed according to a predefined preset data format to construct the first target EEG signal sample data. For details, please refer to the above description and will not be repeated here. It can be understood that the first target EEG signal sample data for auxiliary training model is constructed directly based on the EEG signal prior features and the model is trained, so that the model focuses on the prior knowledge that is conducive to classification, and improves the physiological state auxiliary prediction model's targeted processing ability of important information.
[0126] In another example, the original EEG signal sample data can be processed to obtain EEG signal prior features, and then the original EEG signal sample data and the EEG signal prior features can be merged to obtain second target EEG signal sample data. For example, the original EEG signal sample data and the EEG signal prior features can be decomposed and fused to obtain the second target EEG signal sample data, or the original EEG signal sample data and the EEG signal prior features can be spliced to obtain the second target EEG signal sample data. For details, please refer to the above description and will not be repeated here. It can be understood that the original EEG signal sample data and the EEG signal prior features are merged to obtain the second target EEG signal sample data for auxiliary training model and model training is performed, so that the model pays attention to the prior knowledge while taking into account the original data, thereby improving the generalization of the physiological state auxiliary prediction model. The original EEG signal sample data and the EEG signal prior features are concatenated to train a classification model with a set number of labels (for example, the prior features pre-set labels include four: left hand, right hand, happy, and sad, and the corresponding original EEG signal sample classification includes two labels: motor imagery and emotion). After data concatenation, the six-category training model can be deeply trained. Labels are marked and classified when physiological signals are collected. Labels such as 0, 1, 2, 3, 4, and 5 can be used to correspond to different physiological states.
[0127] In another embodiment of the present application, appropriate EEG signal prior features may be determined by feature engineering to perform model training.
[0128] For example, the original EEG signal sample data can be processed according to the preset characteristic index to obtain the characteristic value corresponding to the preset characteristic index, and the characteristic value corresponding to the preset characteristic index is used as the prior feature of the EEG signal.
[0129] The preset characteristic indicators may include, for example, (δ+θ+α+β) total energy, δ wave absolute energy, θ wave absolute energy, α wave absolute energy, β wave absolute energy, Renyi entropy, wavelet transform, wavelet absolute mean, Fourier transform average coefficient, minimum index, maximum index, average index, variance index, standard deviation index, differential median index, spectrum maximum peak index, etc. The preset characteristic indicators can be set according to actual conditions, and examples are not given here one by one. Taking the preset characteristic indicator as the minimum index as an example, the original EEG signal sample data is processed to obtain the minimum value of the original EEG signal sample data, and the minimum value is the characteristic value corresponding to the minimum index. Taking the preset characteristic indicator as the spectrum maximum peak index as an example, the original EEG signal sample data is converted from the time domain to the frequency domain to obtain the spectrum, and the spectrum maximum peak is calculated as the characteristic value corresponding to the spectrum maximum peak index. Then, the minimum value and the spectrum maximum peak are used as the prior features of the EEG signal.
[0130] Exemplarily, suitable preset feature indicators can be determined by feature engineering. For example, multiple groups of feature indicators to be verified are determined from multiple candidate feature indicators, and each group of feature indicators to be verified includes at least one candidate feature indicator. Multiple candidate feature indicators include, for example, a minimum value indicator, a maximum value indicator, an average value indicator, a variance indicator, a standard deviation indicator, a differential median indicator, a spectrum maximum peak indicator, and the like. Taking the first group of feature indicators to be verified including 10 feature indicators such as the minimum value indicator, the maximum value indicator, and the average value indicator as an example, similarly, the second group of feature indicators to be verified includes 10 feature indicators such as the average value indicator, the variance indicator, and the standard deviation indicator, and the third group of feature indicators to be verified includes 10 feature indicators such as the standard deviation indicator, the differential median indicator, and the spectrum maximum peak indicator.
[0131] After obtaining multiple sets of feature indicators to be verified, each set of feature indicators to be verified is verified to obtain a verification result corresponding to each set of feature indicators to be verified. The verification result represents the prediction accuracy of the physiological state prediction model used for indicator verification based on the feature values corresponding to each set of feature indicators to be verified. The physiological state prediction model used for indicator verification may include the physiological state prediction model to be trained, the physiological state auxiliary prediction model, the trained physiological state prediction model, or other physiological state prediction models.
[0132] Then, based on the verification results, at least one group corresponding to a high model prediction accuracy is selected from multiple groups of feature indicators to be verified, preferably a group corresponding to the highest model prediction accuracy, and the feature indicators included in the selected group of feature indicators to be verified are used as preset feature indicators. The indicator verification process can be completed before model training. After the preset feature indicators are obtained through the indicator verification process, the original EEG signal sample data is processed based on the preset feature indicators to obtain the EEG signal prior features. Alternatively, the indicator verification process can also be carried out simultaneously during the model training process. For example, multiple model trainings can be carried out, and each model training is based on the EEG signal prior features obtained by different preset feature indicators. Finally, at least one group of feature indicators to be verified with the best training effect is used as the preset feature indicators.
[0133] Next, the training process of the physiological state prediction model to be predicted and the physiological state auxiliary prediction model will be further explained.
[0134] Regarding the training process of the physiological state prediction model to be predicted, in one example, when the first target EEG signal sample data is obtained based on the prior features of the EEG signal, when training the physiological state prediction model to be trained, the first target EEG signal sample data can be input into the physiological state prediction model to be trained for prediction to obtain a first sub-category prediction result representing the physiological state. Then, based on the first sub-category prediction result, the model parameters of the physiological state prediction model to be trained are reversely adjusted to obtain a physiological state auxiliary prediction model. For example, the first target EEG signal sample data includes a first sample label. Based on the loss value between the first sub-category prediction result and the first sample label, the model parameters of the physiological state prediction model to be trained can be reversely adjusted to obtain a physiological state auxiliary prediction model.
[0135] Exemplarily, the first target EEG signal sample data is used to control the physiological state, which includes major categories of physiological states such as motor imagery state and emotional state. The motor imagery state can be further subdivided into smaller categories such as left-hand movement state and right-hand movement state. The emotional state can be further subdivided into smaller categories such as happy state and sad state. The first target EEG signal sample data is used to control hand movements and emotions, which include left-hand movement and right-hand movement, and emotions include happy and sad. Therefore, the motor imagery state and emotional state can be predicted based on the first target EEG signal sample data. It can be understood that in addition to the motor imagery state and emotional state, the physiological state can also include other physiological states according to actual needs; the motor imagery state can also include other motor imagery states according to actual needs in addition to the left-hand movement state and right-hand movement state; and the emotional state can also include other emotional states according to actual needs in addition to the happy state and sad state.
[0136] The physiological state prediction model to be trained is, for example, a four-class classification model, which is used to predict the physiological state based on the first target EEG signal sample data to obtain a first sub-classification prediction result. The first sub-classification prediction result may include, for example, left hand movement state, right hand movement state, happy state, and sad state. Specifically, the probability of the left hand movement state, the probability of the right hand movement state, the probability of the happy state, and the probability of the sad state may be output.
[0137] Regarding the training process of the physiological state prediction model to be predicted, in another example, when the original EEG signal sample data and the EEG signal prior features are merged to obtain the second target EEG signal sample data, when training the physiological state prediction model to be trained, the second target EEG signal sample data is input into the physiological state prediction model to be trained for prediction, and a second detailed classification prediction result and a first coarse classification prediction result representing the physiological state are obtained. Then, based on the second detailed classification prediction result and the first coarse classification prediction result, the model parameters of the physiological state prediction model to be trained are reversely adjusted to obtain a physiological state auxiliary prediction model. For example, the second target EEG signal sample data includes a second sample label. Based on the loss value between the second detailed classification prediction result and the first coarse classification prediction result and the second sample label, the model parameters of the physiological state prediction model to be trained can be reversely adjusted to obtain a physiological state auxiliary prediction model.
[0138] Exemplarily, the physiological state prediction model to be trained is, for example, a six-category model, which is used to predict the physiological state based on the second target EEG signal sample data to obtain a second detailed classification prediction result and a first coarse classification prediction result. The second detailed classification prediction result includes, for example, left hand movement state, right hand movement state, happy state, and sad state. The first coarse classification prediction result includes motor imagery state and emotional state. Specifically, the probability of the left hand movement state, the probability of the right hand movement state, the probability of the happy state, the probability of the sad state, the probability of the motor imagery state, and the probability of the emotional state can be output.
[0139] Regarding the training process of the physiological state auxiliary prediction model, after the physiological state auxiliary prediction model is obtained through auxiliary training, the original EEG signal sample data is further input into the physiological state auxiliary prediction model for prediction to obtain a second coarse classification prediction result representing the physiological state. Then, based on the second coarse classification prediction result, the model parameters of the physiological state auxiliary prediction model are reversely adjusted to obtain a trained physiological state prediction model. For example, if the original EEG signal sample data includes a third sample label, the model parameters of the physiological state auxiliary prediction model can be reversely adjusted based on the loss value between the second coarse classification prediction result and the third sample label to obtain a trained physiological state prediction model.
[0140] For example, the original EEG signal sample data includes, for example, EEG signal data for controlling the occurrence of the left hand movement state, EEG signal data for controlling the occurrence of the right hand movement state, EEG signal data for controlling the occurrence of the happy state, and EEG signal data for controlling the occurrence of the sad state. The physiological state auxiliary prediction model obtained by training based on the first target EEG signal sample data is, for example, a four-category model, and the physiological state auxiliary prediction model obtained by training based on the second target EEG signal sample data is, for example, a six-category model. The physiological state auxiliary prediction model has the ability to process EEG signal data for controlling the occurrence of the left hand movement state, the right hand movement state, the happy state, and the sad state. When performing the second model training, the four-category or six-category physiological state auxiliary prediction model can be set as a two-category model to output two classification results. The two-category physiological state auxiliary prediction model is used to perform physiological state prediction based on the original EEG signal sample data to obtain a second coarse classification prediction result. The second coarse classification prediction result includes, for example, a motor imagery state and an emotional state, and specifically can output the probability of the motor imagery state and the probability of the emotional state.
[0141] It can be understood that the EEG signal prior features contain important information or information with high discrimination, so they are more accurate in predicting small categories of physiological states. After obtaining the first target EEG signal sample data based on the EEG signal prior features, the first target EEG signal sample data is used to train a physiological state auxiliary prediction model for finely classifying physiological states, which can improve the model's fine classification prediction accuracy. After obtaining the physiological state auxiliary prediction model for fine classification, a physiological state prediction model for coarse classification of physiological states is trained based on the original EEG signal sample data on this basis, and coarse classification is performed on the basis of fine classification, thereby improving the accuracy of coarse classification.
[0142] It can be understood that after the original EEG signal sample data and the EEG signal prior features are merged to obtain the second target EEG signal sample data, the second target EEG signal sample data is used to train a physiological state auxiliary prediction model for simultaneous fine classification and coarse classification, and coarse classification is taken into account on the basis of fine classification, so that the model has both relatively accurate fine classification capabilities and preliminary coarse classification capabilities. On this basis, the physiological state prediction model for coarse classification is further trained based on the original EEG signal sample data, further strengthening the model's coarse classification capability and improving the model's coarse classification accuracy.
[0143] It can be understood that the embodiments of the present application propose a method of inputting prior knowledge into the model for training, and it is feasible to combine prior knowledge with original EEG signal sample data and input the model, which further proves that model training based on auxiliary learning methods is feasible and the trained model has good prediction effect. The model after auxiliary training based on prior features of EEG signals has a good classification effect in identifying motor imagery states and emotional states.
[0144] To facilitate understanding of the training process of the physiological state prediction model of this application, FIG4 illustrates a convolutional neural network model as an example. It should be understood that the model structure shown in FIG4 and the examples of data size or data dimension used below for ease of understanding are for reference only, and this application does not impose specific limitations on the model structure, data size, or data dimension.
[0145] FIG6 is a schematic diagram of a physiological state prediction model provided in an embodiment of the present application.
[0146] As shown in Figure 6, the physiological state prediction model 400 can be a physiological state prediction model to be trained or a physiological state auxiliary prediction model. Compared with other models that can usually only be applied to the physiological state prediction of a single brain-computer interface (BCI) paradigm, the physiological state prediction model 400 of the embodiment of the present application can be applied to the physiological state prediction of multiple BCI paradigms. The BCI paradigm represents a way of conducting data acquisition experiments based on the brain-computer interface, such as collecting the EEG signal process corresponding to the motor imagination state as one paradigm, and collecting the EEG signal process corresponding to the emotional state as another paradigm. The physiological state prediction model 400 includes at least a temporal convolution layer 420, a spatial convolution layer 430, and a separable convolution layer 440 connected in sequence. Of course, the physiological state prediction model 400 can also include an input layer 410 and an output layer 450.
[0147] When the physiological state prediction model 400 is a physiological state prediction model to be trained, the temporal convolution layer 420 is used to extract the first time feature data of the target EEG signal sample data (including the first target EEG signal sample data and the second target EEG signal sample data) in the time dimension. When the physiological state prediction model 400 is an auxiliary physiological state prediction model, the temporal convolution layer 420 is used to extract the first time feature data of the original EEG signal sample data in the time dimension. The spatial convolution layer 430 is used to extract the spatial feature data of the first time feature data in the spatial dimension. The separable convolution layer 440 is used to extract the second time feature data of the spatial feature data in the time dimension, and to perform feature fusion on the second time feature data.
[0148] The following describes an exemplary process of model training.
[0149] First, the subject's EEG signals can be collected multiple times under different physiological conditions to obtain raw EEG signal sample data. For example, a single experiment can yield a set of raw EEG signal sample data, which can include EEG signal data corresponding to left-hand movement, right-hand movement, happiness, and sadness. The collected data can also be preprocessed and used as the raw EEG signal sample data. Preprocessing includes bandpass filtering and removing redundant data.
[0150] Then, based on the original EEG signal sample data, an EEG signal prior feature is obtained, and based on the EEG signal prior feature, the target EEG signal sample data (including the first target EEG signal sample data and the second target EEG signal sample data) is obtained. The target EEG signal sample data can be divided into a training set, a validation set, and a test set. For example, a target EEG signal sample data includes 10 seconds of EEG signal data collected, and the 10 seconds of EEG signal data is divided into three data sets, and the three data sets include 8 seconds of data, 1 second of data, and 1 second of data, respectively. The 8 seconds of data are used as a training set, the 1 second of data are used as a validation set to verify whether the model is overfitting or underfitting, and the other 1 second of data are used as a test set to test the prediction accuracy of the trained model. The physiological state prediction model 400 of the embodiment of the present application includes a compact deep learning model. After the model is trained based on a training set of target EEG signal sample data (including first target EEG signal sample data and second target EEG signal sample data) or original EEG signal sample data, the trained model is used to classify the validation set and the test set, so that the model can accurately classify the physiological state.
[0151] During the model training process, the original EEG signal sample data is taken as an example with a data size of [22,990]. Taking the second target EEG signal sample data as an example, the original EEG signal sample data and the EEG signal prior features are merged and processed, and the data size of the obtained second target EEG signal sample data is, for example, [22,1000], where 22 represents the number of channels, and the number of channels is the number of electrodes used to collect the original EEG signal sample data. Each channel collects 990 data points within 10 seconds, for example.
[0152] First, the second target EEG signal sample data is input into the model through the input layer 410 .
[0153] Secondly, feature extraction is performed on the second target EEG signal sample data through the time convolution layer 420. For example, feature extraction is performed through a convolution kernel with a data size of [1, 64] and then F1 filters (not shown in the figure) are used to extract F1 types of time dimension information to obtain the first time feature data with a data size of [F1, 22, 1000] in the time dimension, where F1 is a preset integer value.
[0154] Then, the first time feature data is input into the spatial convolution layer 430 for feature extraction. The spatial convolution layer 430 performs spatial feature extraction through a convolution kernel with a data size of [22, 1], and then connects to D*F1 spatial filters (not shown in the figure) to extract spatial dimension information. The data size of the extracted spatial feature data is, for example, [D*F1, 1, 1000], where D is a preset integer value.
[0155] Next, the spatial feature data is input into the separable convolution layer 440 to learn the temporal features of each channel and fuse or integrate the data of multiple channels. The separable convolution layer 440 includes, for example, a temporal convolution kernel with a data size of [1, 16] and a normal convolution kernel with a data size of [1, 1]. The temporal convolution kernel with a data size of [1, 16] can be used to extract features from the spatial feature data to obtain second temporal feature data for each channel, and the normal convolution kernel with a data size of [1, 1] can be used to perform feature fusion on the second temporal feature data of multiple channels to achieve fusion or integration of the feature data of multiple channels to obtain channel fusion feature data for multiple channels.
[0156] Finally, the feature data output by the separable convolution layer 440 is transmitted to the output layer 450, and the physiological state prediction result is output through the output layer 450. The output layer 450 may include a fully connected layer, which classifies the learned deep features to predict the current physiological state.
[0157] It can be understood that the physiological state prediction model extracts temporal features, spatial features and channel fusion features through the temporal convolution layer, the spatial convolution layer and the separable convolution layer, which improves the diversity of features and makes the extracted features contain deep and important information, thereby improving the prediction accuracy of the physiological state prediction model.
[0158] Through the embodiments of this application, the physiological state prediction model learns the weights of motor imagery and affective states and classifies physiological states through a fully connected layer. Using an auxiliary learning approach, the model is trained on more refined subcategories such as left-hand movement, right-hand movement, happiness, and sadness, making it more accurate and reliable to identify the larger categories of motor imagery and affective states.
[0159] Optionally, the physiological state prediction method 500 provided in the embodiment of the present application includes, for example, the following steps.
[0160] Step 1: Obtain the original EEG signal data.
[0161] Step 2: Input the original EEG signal data into the trained physiological state prediction model for prediction to obtain the physiological state prediction result.
[0162] Among them, the trained physiological state prediction model can be trained by the training method of the physiological state prediction model mentioned above, which will not be repeated here. The predicted results of the predicted physiological state include, for example, motor imagery state and emotional state, specifically the probability of motor imagery state and the probability of emotional state. The trained physiological state prediction model is trained based on the prior knowledge of EEG signals. Therefore, the trained physiological state prediction model is used to predict the physiological state based on the original EEG signal data, so that the accuracy of the prediction result is higher.
[0163] Optionally, another physiological state prediction method 600 provided in an embodiment of the present application includes, for example, steps.
[0164] Step 1: Obtain the original EEG signal data.
[0165] Step 2: Input the original EEG signal data into the trained physiological state prediction model for prediction to obtain the physiological state prediction result.
[0166] The trained physiological state prediction model includes at least one of a trained first physiological state prediction model and a trained second physiological state prediction model. For example, one of the trained first physiological state prediction model and the trained second physiological state prediction model with higher prediction accuracy can be selected to predict the original EEG signal data.
[0167] The trained first physiological state prediction model is trained based on the first target EEG signal sample data, and the trained second physiological state prediction model is trained based on the second target EEG signal sample data. The first target EEG signal sample data is constructed based on EEG signal prior features, which are obtained by processing the original EEG signal sample data. The second target EEG signal sample data is obtained by combining the original EEG signal sample data and the EEG signal prior features.
[0168] Exemplarily, the trained first physiological state prediction model is obtained by the following steps: auxiliary training is performed on the first physiological state prediction model to be trained based on the first target EEG signal sample data to obtain a first physiological state auxiliary prediction model; and the first physiological state auxiliary prediction model is trained based on the original EEG signal sample data to obtain a trained first physiological state prediction model. The model training process can be referred to above and will not be repeated here.
[0169] Exemplarily, the trained second physiological state prediction model is obtained by the following steps: auxiliary training of the second physiological state prediction model to be trained based on the second target EEG signal sample data to obtain a second physiological state auxiliary prediction model; and training the second physiological state auxiliary prediction model based on the original EEG signal sample data to obtain a trained second physiological state prediction model. The model training process can be referred to above and will not be repeated here.
[0170] The device for generating EEG signal sample data provided in the embodiments of the present application may include: a processing module and a generating module.
[0171] Exemplarily, the processing module is used to process original EEG signal sample data to obtain EEG signal prior features.
[0172] Exemplarily, the generation module is used to generate target EEG signal sample data based on prior features of the EEG signal, wherein the target EEG signal sample data is used to assist in training the physiological state prediction model to be trained to obtain an auxiliary physiological state prediction model, and the original EEG signal sample data is used to train the auxiliary physiological state prediction model to obtain a trained physiological state prediction model.
[0173] Exemplarily, the generation module is used to perform at least one of the following: constructing first target EEG signal sample data based on EEG signal prior features; merging the original EEG signal sample data and the EEG signal prior features to obtain second target EEG signal sample data.
[0174] Exemplarily, the processing module is specifically used to: process the original EEG signal sample data according to a preset characteristic index, and obtain a characteristic value corresponding to the preset characteristic index as a priori feature of the EEG signal.
[0175] Exemplarily, the device for generating EEG signal sample data also includes: a first model training module, used to input the first target EEG signal sample data into the physiological state prediction model to be trained for prediction, and obtain a first sub-classification prediction result representing the physiological state; based on the first sub-classification prediction result, reversely adjust the model parameters of the physiological state prediction model to be trained to obtain a physiological state auxiliary prediction model.
[0176] Exemplarily, the first sub-category prediction result includes at least one of the following: left hand movement state, right hand movement state, happy state, sad state.
[0177] Exemplarily, the first model training module is also used to: input the second target EEG signal sample data into the physiological state prediction model to be trained for prediction, and obtain a second sub-classification prediction result and a first coarse classification prediction result representing the physiological state; based on the second sub-classification prediction result and the first coarse classification prediction result, reversely adjust the model parameters of the physiological state prediction model to be trained to obtain a physiological state auxiliary prediction model.
[0178] Exemplarily, the second detailed classification prediction result includes at least one of the following: left hand movement state, right hand movement state, happy state, sad state; the first coarse classification prediction result includes at least one of the following: movement imagination state, emotional state.
[0179] Exemplarily, the first model training module is also used to: input the original EEG signal sample data into the physiological state auxiliary prediction model for prediction to obtain a second coarse classification prediction result that characterizes the physiological state; based on the second coarse classification prediction result, reversely adjust the model parameters of the physiological state auxiliary prediction model to obtain a trained physiological state prediction model.
[0180] Exemplarily, the second coarse classification prediction result includes at least one of the following: motor imagery state, emotional state.
[0181] Exemplarily, the original EEG signal sample data includes at least one of the following: EEG signal data for controlling the occurrence of the left hand movement state, EEG signal data for controlling the occurrence of the right hand movement state, EEG signal data for controlling the occurrence of the happy state, and EEG signal data for controlling the occurrence of the sad state.
[0182] Exemplarily, the processing module is also used to: determine multiple groups of feature indicators to be verified from multiple candidate feature indicators, wherein each group of feature indicators to be verified includes at least one candidate feature indicator; verify the multiple groups of feature indicators to be verified respectively to obtain verification results corresponding to each group of feature indicators to be verified, wherein the verification results represent the prediction accuracy of the physiological state prediction model used for indicator verification based on the characteristic values corresponding to each group of feature indicators to be verified for physiological state prediction; based on the verification results, select at least one group from the multiple groups of feature indicators to be verified as the preset feature indicator.
[0183] Exemplarily, the physiological state prediction model or physiological state auxiliary prediction model to be trained includes: a temporal convolution layer, used to extract first time feature data of the target EEG signal sample data or the original EEG signal sample data in the time dimension; a spatial convolution layer, used to extract spatial feature data of the first time feature data in the spatial dimension; a separable convolution layer, used to extract second time feature data of the spatial feature data in the time dimension, and perform feature fusion on the second time feature data.
[0184] It can be understood that the specific implementation process of the device for generating EEG signal sample data can refer to the implementation process of the method for generating EEG signal sample data above, and will not be repeated here.
[0185] The training device for the physiological state prediction model provided in the embodiment of the present application includes: a first acquisition module, a first training module, and a second training module.
[0186] Exemplarily, the first acquisition module is used to acquire target EEG signal sample data, wherein the target EEG signal sample data is determined based on EEG signal prior features, and the EEG signal prior features are obtained based on original EEG signal sample data.
[0187] Exemplarily, the first training module is used to perform auxiliary training on the physiological state prediction model to be trained based on the target EEG signal sample data to obtain the physiological state auxiliary prediction model.
[0188] Exemplarily, the second training module is used to train the physiological state auxiliary prediction model based on the original EEG signal sample data to obtain a trained physiological state prediction model.
[0189] Exemplarily, the first acquisition module is specifically configured to: process original EEG signal sample data to obtain EEG signal prior features; and construct first target EEG signal sample data according to the EEG signal prior features.
[0190] Exemplarily, the first acquisition module is specifically used to: process the original EEG signal sample data to obtain EEG signal prior features; and merge the original EEG signal sample data and the EEG signal prior features to obtain second target EEG signal sample data.
[0191] Exemplarily, processing the original EEG signal sample data to obtain the EEG signal prior features includes: processing the original EEG signal sample data according to preset feature indicators to obtain feature values corresponding to the preset feature indicators as the EEG signal prior features.
[0192] Exemplarily, the first training module is specifically used to: input the first target EEG signal sample data into the physiological state prediction model to be trained for prediction, and obtain a first sub-classification prediction result representing the physiological state; based on the first sub-classification prediction result, reversely adjust the model parameters of the physiological state prediction model to be trained to obtain a physiological state auxiliary prediction model.
[0193] Exemplarily, the first sub-category prediction result includes at least one of the following: left hand movement state, right hand movement state, happy state, sad state.
[0194] Exemplarily, the first training module is specifically used to: input the second target EEG signal sample data into the physiological state prediction model to be trained for prediction, and obtain a second detailed classification prediction result and a first coarse classification prediction result representing the physiological state; based on the second detailed classification prediction result and the first coarse classification prediction result, reversely adjust the model parameters of the physiological state prediction model to be trained to obtain a physiological state auxiliary prediction model.
[0195] Exemplarily, the second detailed classification prediction result includes at least one of the following: left hand movement state, right hand movement state, happy state, sad state; the first coarse classification prediction result includes at least one of the following: movement imagination state, emotional state.
[0196] Exemplarily, the second training module is specifically used to: input the original EEG signal sample data into the physiological state auxiliary prediction model for prediction to obtain a second coarse classification prediction result that characterizes the physiological state; based on the second coarse classification prediction result, reversely adjust the model parameters of the physiological state auxiliary prediction model to obtain a trained physiological state prediction model.
[0197] Exemplarily, the second coarse classification prediction result includes at least one of the following: motor imagery state, emotional state.
[0198] Exemplarily, the original EEG signal sample data includes at least one of the following: EEG signal data for controlling the occurrence of the left hand movement state, EEG signal data for controlling the occurrence of the right hand movement state, EEG signal data for controlling the occurrence of the happy state, and EEG signal data for controlling the occurrence of the sad state.
[0199] Exemplarily, the preset characteristic indicators are obtained in the following manner: determining multiple groups of characteristic indicators to be verified from multiple candidate characteristic indicators, wherein each group of characteristic indicators to be verified includes at least one candidate characteristic indicator; verifying the multiple groups of characteristic indicators to be verified respectively to obtain verification results corresponding to each group of characteristic indicators to be verified, wherein the verification results represent the prediction accuracy of the physiological state prediction model used for indicator verification based on the characteristic values corresponding to each group of characteristic indicators to be verified for physiological state prediction; based on the verification results, selecting at least one group from the multiple groups of characteristic indicators to be verified as the preset characteristic indicator.
[0200] Exemplarily, the physiological state prediction model or physiological state auxiliary prediction model to be trained includes: a temporal convolution layer, used to extract first time feature data of the target EEG signal sample data or the original EEG signal sample data in the time dimension; a spatial convolution layer, used to extract spatial feature data of the first time feature data in the spatial dimension; a separable convolution layer, used to extract second time feature data of the spatial feature data in the time dimension, and perform feature fusion on the second time feature data.
[0201] It can be understood that the specific implementation process of the training device for the physiological state prediction model can refer to the implementation process of the training method for the physiological state prediction model mentioned above, and will not be repeated here.
[0202] The physiological state prediction device provided in the embodiment of the present application includes: a second acquisition module and a first prediction module.
[0203] Exemplarily, the second acquisition module is used to acquire original EEG signal data.
[0204] Exemplarily, the first prediction module is used to input the original EEG signal data into a trained physiological state prediction model for prediction, thereby obtaining a prediction result of the physiological state.
[0205] Exemplarily, the physiological state prediction device also includes a second model training module, which is specifically used to: obtain target EEG signal sample data, wherein the target EEG signal sample data is determined based on EEG signal prior features, and the EEG signal prior features are obtained based on the original EEG signal sample data; based on the target EEG signal sample data, perform auxiliary training on the physiological state prediction model to be trained to obtain a physiological state auxiliary prediction model; based on the original EEG signal sample data, perform training on the physiological state auxiliary prediction model to obtain a trained physiological state prediction model.
[0206] Exemplarily, the physiological state prediction device further includes a first sample acquisition module, which is used to: process the original EEG signal sample data to obtain EEG signal prior features; and construct first target EEG signal sample data based on the EEG signal prior features.
[0207] Exemplarily, the first sample acquisition module is further used to: process the original EEG signal sample data to obtain EEG signal prior features; and merge the original EEG signal sample data and the EEG signal prior features to obtain second target EEG signal sample data.
[0208] Exemplarily, the first sample acquisition module is further configured to: process original EEG signal sample data according to preset characteristic indicators, and obtain characteristic values corresponding to the preset characteristic indicators as prior features of the EEG signal.
[0209] Exemplarily, the second model training module is also used to: input the first target EEG signal sample data into the physiological state prediction model to be trained for prediction, and obtain a first subclassification prediction result representing the physiological state; based on the first subclassification prediction result, reversely adjust the model parameters of the physiological state prediction model to be trained to obtain a physiological state auxiliary prediction model.
[0210] Exemplarily, the first sub-category prediction result includes at least one of the following: left hand movement state, right hand movement state, happy state, sad state.
[0211] Exemplarily, the second model training module is also used to: input the second target EEG signal sample data into the physiological state prediction model to be trained for prediction, and obtain a second sub-classification prediction result and a first coarse classification prediction result representing the physiological state; based on the second sub-classification prediction result and the first coarse classification prediction result, reversely adjust the model parameters of the physiological state prediction model to be trained to obtain a physiological state auxiliary prediction model.
[0212] Exemplarily, the second detailed classification prediction result includes at least one of the following: left hand movement state, right hand movement state, happy state, sad state; the first coarse classification prediction result includes at least one of the following: movement imagination state, emotional state.
[0213] Exemplarily, the second model training module is also used to: input the original EEG signal sample data into the physiological state auxiliary prediction model for prediction to obtain a second coarse classification prediction result that characterizes the physiological state; based on the second coarse classification prediction result, reversely adjust the model parameters of the physiological state auxiliary prediction model to obtain a trained physiological state prediction model.
[0214] Exemplarily, the second coarse classification prediction result includes at least one of the following: motor imagery state, emotional state.
[0215] Exemplarily, the original EEG signal sample data includes at least one of the following: EEG signal data for controlling the occurrence of the left hand movement state, EEG signal data for controlling the occurrence of the right hand movement state, EEG signal data for controlling the occurrence of the happy state, and EEG signal data for controlling the occurrence of the sad state.
[0216] Exemplarily, the first sample acquisition module is also used to: determine multiple groups of feature indicators to be verified from multiple candidate feature indicators, wherein each group of feature indicators to be verified includes at least one candidate feature indicator; verify the multiple groups of feature indicators to be verified respectively to obtain verification results corresponding to each group of feature indicators to be verified, wherein the verification results represent the prediction accuracy of the physiological state prediction model used for indicator verification based on the characteristic values corresponding to each group of feature indicators to be verified for physiological state prediction; based on the verification results, select at least one group from the multiple groups of feature indicators to be verified as the preset feature indicator.
[0217] Exemplarily, the physiological state prediction model or physiological state auxiliary prediction model to be trained includes: a temporal convolution layer, used to extract first time feature data of the target EEG signal sample data or the original EEG signal sample data in the time dimension; a spatial convolution layer, used to extract spatial feature data of the first time feature data in the spatial dimension; a separable convolution layer, used to extract second time feature data of the spatial feature data in the time dimension, and perform feature fusion on the second time feature data.
[0218] It can be understood that the specific implementation process of the physiological state prediction device can refer to the implementation process of the physiological state prediction method above, and will not be repeated here.
[0219] Another physiological state prediction device provided in an embodiment of the present application includes: a third acquisition module and a second prediction module.
[0220] Exemplarily, the third acquisition module is used to acquire original EEG signal data.
[0221] Exemplarily, the second prediction module is used to input the original EEG signal data into a trained physiological state prediction model for prediction, thereby obtaining a prediction result of the physiological state. The trained physiological state prediction model includes at least one of a trained first physiological state prediction model and a trained second physiological state prediction model, wherein the trained first physiological state prediction model is trained based on first target EEG signal sample data, and the trained second physiological state prediction model is trained based on second target EEG signal sample data. The first target EEG signal sample data is constructed based on EEG signal prior features, which are obtained by processing the original EEG signal sample data. The second target EEG signal sample data is obtained by merging the original EEG signal sample data and the EEG signal prior features.
[0222] Exemplarily, another physiological state prediction device also includes: a first physiological state prediction model training module, which is used to assist in training the first physiological state prediction model to be trained based on the first target EEG signal sample data to obtain a first physiological state auxiliary prediction model; based on the original EEG signal sample data, the first physiological state auxiliary prediction model is trained to obtain a trained first physiological state prediction model.
[0223] Exemplarily, another physiological state prediction device also includes: a second physiological state prediction model training module, which is used to assist in training the second physiological state prediction model to be trained based on the second target EEG signal sample data to obtain a second physiological state auxiliary prediction model; based on the original EEG signal sample data, the second physiological state auxiliary prediction model is trained to obtain a trained second physiological state prediction model.
[0224] Exemplarily, another physiological state prediction device further includes a second sample acquisition module, which is used to: process the original EEG signal sample data to obtain EEG signal prior features; and construct first target EEG signal sample data based on the EEG signal prior features.
[0225] Exemplarily, the second sample acquisition module is further used to: process the original EEG signal sample data to obtain EEG signal prior features; and merge the original EEG signal sample data and the EEG signal prior features to obtain second target EEG signal sample data.
[0226] Exemplarily, the second sample acquisition module is further configured to process original EEG signal sample data according to preset characteristic indicators, and obtain characteristic values corresponding to the preset characteristic indicators as prior features of the EEG signal.
[0227] Exemplarily, the first physiological state prediction model training module is also used to: input the first target EEG signal sample data into the first physiological state prediction model to be trained for prediction, and obtain a first subclassification prediction result representing the physiological state; based on the first subclassification prediction result, reversely adjust the model parameters of the first physiological state prediction model to be trained to obtain a first physiological state auxiliary prediction model.
[0228] Exemplarily, the first sub-category prediction result includes at least one of the following: left hand movement state, right hand movement state, happy state, sad state.
[0229] Exemplarily, the second physiological state prediction model training module is also used to: input the second target EEG signal sample data into the second physiological state prediction model to be trained for prediction, and obtain a second sub-classification prediction result and a first coarse classification prediction result representing the physiological state; based on the second sub-classification prediction result and the first coarse classification prediction result, reversely adjust the model parameters of the second physiological state prediction model to be trained to obtain a second physiological state auxiliary prediction model.
[0230] Exemplarily, the second detailed classification prediction result includes at least one of the following: left hand movement state, right hand movement state, happy state, sad state; the first coarse classification prediction result includes at least one of the following: movement imagination state, emotional state.
[0231] Exemplarily, the first physiological state prediction model training module or the second physiological state prediction model training module is also used to: input the original EEG signal sample data into the first physiological state auxiliary prediction model or the second physiological state auxiliary prediction model for prediction to obtain a second coarse classification prediction result characterizing the physiological state; based on the second coarse classification prediction result, reversely adjust the model parameters of the first physiological state auxiliary prediction model or the second physiological state auxiliary prediction model to obtain a trained first physiological state prediction model or a trained second physiological state prediction model.
[0232] Exemplarily, the second coarse classification prediction result includes at least one of the following: motor imagery state, emotional state.
[0233] Exemplarily, the original EEG signal sample data includes at least one of the following: EEG signal data for controlling the occurrence of the left hand movement state, EEG signal data for controlling the occurrence of the right hand movement state, EEG signal data for controlling the occurrence of the happy state, and EEG signal data for controlling the occurrence of the sad state.
[0234] Exemplarily, the second sample acquisition module is also used to: determine multiple groups of feature indicators to be verified from multiple candidate feature indicators, wherein each group of feature indicators to be verified includes at least one candidate feature indicator; verify the multiple groups of feature indicators to be verified respectively to obtain verification results corresponding to each group of feature indicators to be verified, wherein the verification results characterize the prediction accuracy of the physiological state prediction model used for indicator verification based on the characteristic values corresponding to each group of feature indicators to be verified for physiological state prediction; based on the verification results, select at least one group from the multiple groups of feature indicators to be verified as the preset feature indicator.
[0235] Exemplarily, the first physiological state prediction model to be trained, the second physiological state prediction model to be trained, the first physiological state auxiliary prediction model or the second physiological state auxiliary prediction model include: a temporal convolution layer, a spatial convolution layer and a separable convolution layer. The temporal convolution layer of the first physiological state prediction model to be trained is used to extract the first temporal feature data of the first target EEG signal sample data in the temporal dimension, the temporal convolution layer of the second physiological state prediction model to be trained is used to extract the first temporal feature data of the second target EEG signal sample data in the temporal dimension, the first physiological state auxiliary prediction model or the second physiological state auxiliary prediction model is used to extract the first temporal feature data of the original EEG signal sample data in the temporal dimension; the spatial convolution layer is used to extract the spatial feature data of the first temporal feature data in the spatial dimension; the separable convolution layer is used to extract the second temporal feature data of the spatial feature data in the temporal dimension, and to perform feature fusion on the second temporal feature data.
[0236] It can be understood that the specific implementation process of another physiological state prediction device can refer to the implementation process of another physiological state prediction method mentioned above, and will not be repeated here.
[0237] Example 3
[0238] Currently, users can monitor their health status using physiological data acquisition devices combined with analysis devices. For example, the physiological data acquisition device can collect the user's physiological signals and upload them to the analysis device. The analysis device can then analyze and process the physiological signals to determine the user's health status and provide feedback to the user. The physiological data acquisition device can be used to collect at least one of the following physiological signals: EEG signals, respiratory signals, heart rate signals, electrodermal signals, skin temperature signals, and electromyographic signals.
[0239] Since different types of physiological signals are processed differently, the analysis device needs to first identify the type of physiological signal after receiving the physiological signal uploaded by the physiological acquisition device, and then accurately analyze the type to obtain the user's health status.
[0240] Regarding the embodiment of the first embodiment of the present application, in which a classification task is required to be performed on a classification model of multimodal human physiological data, the data is collected through a physiological acquisition device combined with an analysis device to monitor a type of health condition. Therefore, before the classification task is performed, the type of physiological data can be identified according to the technical solution of this embodiment. In addition, the physiological signal in this embodiment of the present application also refers to the physiological data in the first embodiment of the present application.
[0241] The present application provides a method for processing physiological signals based on human factors intelligence, which is applied to an electronic device. Optionally, the electronic device can be a terminal device or a server. The server can be a single server, or a server cluster consisting of several servers, or a cloud computing service center. The method can include the following steps:
[0242] Step 1: Perform weighted processing on the input physiological signal through the channel attention network to obtain a first weighted physiological signal.
[0243] The number of physiological signal channels is not compressed during the process of the channel attention network processing, that is, the number of physiological signal channels does not change during this process. This ensures that the channel attention network can capture cross-channel interaction information while reducing the computational complexity of the channel attention network during the process of processing physiological signals.
[0244] The physiological signal may be sent from a physiological acquisition device to an electronic device, or may be pre-stored in the electronic device. Optionally, the physiological signal may be a single type of signal. For example, the physiological signal may be an electroencephalogram (EEG) signal, an electromyogram (EMG) signal, or a skin electrode signal. Alternatively, the physiological signal may include multiple types of signals, such as an EEG signal, an EMG signal, and a skin electrode signal.
[0245] In the process of processing a physiological signal through a channel attention network to obtain a first weighted physiological signal, the electronic device can first determine a target channel weight vector of the physiological signal, and then perform weighted processing on the physiological signal based on the target channel weight vector to obtain the first weighted physiological signal.
[0246] Step 2: The first weighted physiological signal is weighted processed by the spatial attention network to obtain a second weighted physiological signal.
[0247] After the electronic device obtains the first weighted physiological signal, it can input the first weighted physiological signal into the spatial attention network so that the spatial attention network performs weighted processing on the first weighted physiological signal to obtain the second weighted physiological signal.
[0248] It can be understood that, in the process of processing the physiological signal through the spatial attention network to obtain the second weighted physiological signal, the electronic device can first determine the target spatial weight vector, and then perform weighted processing on the first weighted physiological signal based on the target spatial weight vector to obtain the first weighted physiological signal.
[0249] Step 3: Obtain the type of the physiological signal based on the second weighted physiological signal.
[0250] In an embodiment of the present application, the electronic device may input the physiological signal into a physiological signal recognition network to obtain the type of the bio-signal, which may be EEG, EMG, ECG, GD, skin temperature, or blood flow.
[0251] In summary, the embodiments of the present application provide a method for processing physiological signals based on human factors intelligence. This method can perform weighted processing on physiological signals through a channel attention network to obtain a first weighted physiological signal, and perform weighted processing on the first weighted physiological signal through a spatial attention network to obtain a second weighted physiological signal. Then, based on the second weighted physiological signal, the type of physiological signal can be identified. Since the number of channels of the physiological signal is not compressed during the process of processing the physiological signal by the channel attention network, the computational complexity of the channel attention network is reduced, thereby improving the efficiency of identifying the type of physiological signal.
[0252] In an embodiment of the present application, an electronic device is configured with a recognition model, and the electronic device can process the physiological signal through the recognition model to obtain the type of the physiological signal. Referring to Figure 7, Figure 7 shows a structural schematic diagram of a recognition model, which includes: a channel attention network 201 and a spatial attention network 202 connected in sequence. Referring to Figure 16, the channel attention network includes: a first sub-network 2011 and a second sub-network 2012 in parallel. The process of the electronic device performing step one may include:
[0253] Step S1: Input the physiological signal into the first sub-network to obtain a first channel weight vector.
[0254] During the process of processing the physiological signal by the first sub-network, the number of physiological signal channels is not compressed. That is, the number of physiological signal channels does not change during the process. Since the number of channels is not compressed, there is no need to perform a channel number restoration operation, thereby reducing the computational complexity of the first sub-network in determining the first channel weight vector.
[0255] In an embodiment of the present application, referring to FIG8 , the first subnetwork includes a global maximum pooling layer and a first convolutional layer connected in sequence. The electronic device may compress the physiological signal in the spatial dimension using the global maximum pooling layer to obtain a first compressed physiological signal. The electronic device may then convolve the first compressed physiological signal using the first convolutional layer to obtain a first channel weight vector of the physiological signal. Optionally, the first convolutional layer is a one-dimensional convolutional layer.
[0256] In an embodiment of the present application, after obtaining the first compressed physiological signal, the electronic device may use the transpose() function to transpose the first compressed physiological signal so that the first convolutional layer can perform convolution processing on the first compressed physiological signal. Then, after obtaining the output result of the first convolutional layer, the electronic device may again use the transpose() function to transpose the output result to obtain the first channel weight vector.
[0257] Step S2: Input the physiological signal into the second sub-network to obtain a second channel weight vector.
[0258] During the second sub-network's processing of the physiological signal, the number of physiological signal channels is not compressed. That is, the number of physiological signal channels remains unchanged during the second sub-network's processing. Because the number of channels is not compressed, there is no need to restore the number of channels, thereby reducing the computational complexity of the second sub-network in determining the second channel weight vector.
[0259] As shown in Figure 8, the second sub-network 2012 may include a global average pooling layer and a second convolutional layer connected in sequence. The electronic device may compress the physiological signal input to the channel attention network in the spatial dimension using the global average pooling layer to obtain a second compressed physiological signal. The electronic device may then convolve the second compressed physiological signal using the second convolutional layer to obtain a second channel weight vector.
[0260] It is understood that after obtaining the second compressed physiological signal, the electronic device can use the transpose() function to transpose the second compressed physiological signal so that the second convolutional layer can perform convolution processing on the first compressed physiological signal. Then, after obtaining the output result of the second convolutional layer, the electronic device can again use the transpose() function to transpose the output result to obtain the second channel weight vector.
[0261] Step S3: Obtain a first weighted physiological signal based on the first channel weight vector, the second channel weight vector, and the physiological signal.
[0262] In an embodiment of the present application, as shown in FIG8 , the electronic device may perform a weighted summation on the first channel weight vector and the second channel weight vector to obtain a target channel weight vector. The electronic device may then use the target channel weight vector to perform weighted processing on the physiological signal to obtain a first weighted physiological signal. For example, the electronic device may multiply the target channel weight vector by the physiological signal to obtain the first weighted physiological signal.
[0263] It is understandable that the weight of the first channel weight vector and the weight of the second channel weight vector can both be pre-stored by the electronic device. For example, the weight of the first channel weight vector and the weight of the second channel weight vector can both be 1. In this case, the electronic device can perform an element-wise sum operation on the first channel weight vector and the second channel weight vector, and then use an activation function to operate on the weight vector obtained after the element-wise summation to obtain the target channel weight vector.
[0264] The activation function may be a sigmoid() function.
[0265] FIG9 is a schematic diagram of the structure of a spatial attention network provided by an embodiment of the present application. As can be seen from FIG9, the spatial attention network 202 may include: a pooling network 2021 and a convolution block 2022. Based on this, the electronic device processes the first weighted physiological signal through the spatial attention network to obtain the second weighted physiological signal. The process may include:
[0266] Step A1: compress the first weighted physiological signal in the channel dimension based on the pooling network to obtain a third compressed physiological signal.
[0267] Continuing with FIG9 , the pooling network 2021 may include a maximum pooling layer and an average pooling layer. The electronic device may compress the first weighted physiological signal along the channel dimension using the maximum pooling layer to obtain a first sub-compressed physiological signal, and may also compress the first weighted physiological signal along the channel dimension using the average pooling layer to obtain a second sub-compressed physiological signal. The electronic device may then concatenate the first sub-compressed physiological signal with the second sub-compressed physiological signal to obtain a third compressed physiological signal.
[0268] For example, the electronic device may concatenate the first compressed sub-physiological signal and the second compressed sub-physiological signal using a concat() function to obtain a third compressed physiological signal.
[0269] Step A2: performing convolution processing on the third compressed physiological signal through a convolution block to obtain a target spatial weight vector.
[0270] In an embodiment of the present application, the convolution block may include: connecting a third convolution layer and a first batch normalization layer (BN) in sequence. The electronic device may perform convolution processing on the third compressed physiological signal through the third convolution layer to obtain an initial spatial weight vector. Then, the electronic device may perform normalization processing on the initial spatial weight vector through the first batch normalization layer to obtain the normalized initial spatial weight vector. Afterwards, the electronic device may use an activation function to operate on the initial spatial weight vector after batch normalization to obtain a target spatial weight vector.
[0271] Since convolution blocks are used for signal processing in the spatial attention network, the problems of overfitting and gradient disappearance can be effectively reduced.
[0272] It is understandable that the activation function may be a mish() function.
[0273] Step A3: Use the target space weight vector to perform weighted processing on the first weighted physiological signal to obtain a second weighted physiological signal.
[0274] After obtaining the target spatial weight vector, the electronic device can use the target spatial weight vector to perform weighted processing on the first weighted physiological signal to obtain a second weighted physiological signal. For example, the electronic device can multiply the first weighted physiological signal by the target spatial weight vector to obtain the second weighted physiological signal.
[0275] Figure 10 is a schematic diagram of the structure of another recognition model provided in an embodiment of the present application. Referring to Figure 10 , the model may also include: a physiological signal recognition network 203 connected to the spatial attention network 202 . The electronic device may input the second weighted physiological signal into the physiological signal recognition network to determine the type of the physiological signal. Optionally, the physiological signal recognition network may be a convolutional neural network.
[0276] Referring to Figure 11, the physiological signal recognition network 203 may include: a first signal recognition sub-network 2031, a second signal recognition sub-network 2032, and a third signal recognition sub-network 2033, which are connected in sequence. The first signal recognition sub-network 2031 includes: a fourth convolutional layer, a fifth convolutional layer, and a first average pooling layer, which are connected in sequence. The second signal recognition sub-network 2032 includes: a sixth convolutional layer, a seventh convolutional layer, and a second average pooling layer, which are connected in sequence. The third signal recognition sub-network 2033 includes: an eighth convolutional layer, a ninth convolutional layer, and a second batch normalization layer, which are connected in sequence.
[0277] The size of the convolution kernel of the fourth convolution layer is different from the size of the convolution kernel of the fifth convolution layer. For example, the size of the convolution kernel of the fourth convolution layer can be 3*3, and the size of the convolution kernel of the fifth convolution layer can be 1*1. The size of the convolution kernel of the sixth convolution layer is different from the size of the convolution kernel of the seventh convolution layer. For example, the size of the convolution kernel of the sixth convolution layer can be 1*4, and the size of the convolution kernel of the seventh convolution layer can be 1*1. The size of the convolution kernel of the first average pooling layer is different from the size of the convolution kernel of the second average pooling layer. For example, the size of the convolution kernel of the first average pooling layer can be 2*4, and the size of the convolution kernel of the second average pooling layer can be 2*2. The size of the convolution kernel of the eighth convolution layer is different from the size of the convolution kernel of the ninth convolution layer. For example, the size of the convolution kernel of the eighth convolution layer can be 2*2, and the size of the convolution kernel of the ninth convolution layer can be 1*1.
[0278] The electronic device may sequentially perform convolution processing on the second weighted physiological signal through the fourth convolution layer and the fifth convolution layer to obtain a first processing result. The electronic device then inputs the first processing result into the first average pooling layer. The first average pooling layer may perform compression processing on the first processing result to obtain a second processing result.
[0279] The electronic device may then perform convolution processing on the second processing result in sequence through the sixth convolution layer and the seventh convolution layer to obtain a third processing result. The electronic device may then input the third processing result into the second average pooling layer. The second average pooling layer may perform compression processing on the third processing result to obtain a fourth processing result.
[0280] The electronic device can then perform convolution processing on the fourth processing result through the eighth and ninth convolutional layers, obtaining a fifth processing result, and input the fifth processing result into the second batch normalization layer. The second batch normalization layer can then perform normalization processing on the fifth processing result to obtain the type of the physiological signal.
[0281] It is understood that the order of the steps of the human factor intelligent physiological signal processing method provided in the embodiments of the present application can be adjusted appropriately, and the steps can be increased or decreased accordingly. Any person skilled in the art can easily conceive of a modified method within the scope of the technology disclosed in this application, and the modified method should be included in the scope of protection of this application, so it will not be described in detail.
[0282] In summary, the embodiments of the present application provide a method for processing physiological signals based on human factors intelligence. This method can perform weighted processing on physiological signals through a channel attention network to obtain a first weighted physiological signal, and perform weighted processing on the first weighted physiological signal through a spatial attention network to obtain a second weighted physiological signal. Then, based on the second weighted physiological signal, the type of physiological signal can be identified. Since the number of channels of the physiological signal is not compressed during the process of processing the physiological signal by the channel attention network, the computational complexity of the channel attention network is reduced, thereby improving the efficiency of identifying the type of physiological signal.
[0283] The present application also provides a human factors-based intelligent physiological signal processing device, which includes:
[0284] The first weighted processing module is used to perform weighted processing on the input physiological signal through the channel attention network to obtain a first weighted physiological signal, wherein the number of channels of the physiological signal is not compressed during the process of the channel attention network processing the physiological signal.
[0285] The second weighted processing module is used to perform weighted processing on the first weighted physiological signal through a spatial attention network to obtain a second weighted physiological signal.
[0286] The identification module is configured to obtain a type of the physiological signal based on the second weighted physiological signal.
[0287] Optionally, the channel attention network includes: a first sub-network and a second sub-network. The first weighted processing module can be used to:
[0288] Inputting the physiological signal into the first sub-network to obtain a first channel weight vector, wherein the number of channels of the physiological signal is not compressed during the process of the first sub-network processing the physiological signal;
[0289] Inputting the physiological signal into the second sub-network to obtain a second channel weight vector, wherein the number of channels of the physiological signal is not compressed during the process of processing the physiological signal by the second sub-network;
[0290] A first weighted physiological signal is obtained based on the first channel weight vector, the second channel weight vector, and the physiological signal.
[0291] Optionally, the first sub-network includes: a global maximum pooling layer and a first convolutional layer connected in sequence. The first weighted processing module can be used to:
[0292] The physiological signal is compressed in the spatial dimension by a global maximum pooling layer to obtain a first compressed physiological signal;
[0293] The first compressed physiological signal is convolved by the first convolutional layer to obtain a first channel weight vector of the physiological signal.
[0294] Optionally, the second sub-network includes: a global average pooling layer and a second convolutional layer connected in sequence. The first weighted processing module can be used to:
[0295] The physiological signal is compressed in the spatial dimension through a global average pooling layer to obtain a second compressed physiological signal;
[0296] The second compressed physiological signal is convolved by the second convolutional layer to obtain a second channel weight vector of the physiological signal.
[0297] Optionally, the first weighted processing module may be used to:
[0298] Perform weighted summation on the first channel weight vector and the second channel weight vector to obtain the target channel weight vector;
[0299] The target channel weight vector is used to perform weighted processing on the physiological signal to obtain a first weighted physiological signal.
[0300] Optionally, the spatial attention network includes: a pooling network and a convolution block. The second weighted processing module can be used to:
[0301] compressing the first weighted physiological signal in a channel dimension based on a pooling network to obtain a third compressed physiological signal;
[0302] performing convolution processing on the third compressed physiological signal through a convolution block to obtain a target spatial weight vector;
[0303] The target space weight vector is used to perform weighted processing on the first weighted physiological signal to obtain a second weighted physiological signal.
[0304] Optionally, the pooling network includes: a maximum pooling layer and an average pooling layer. The second weighted processing module can be used to:
[0305] compressing the first weighted physiological signal in the channel dimension through a maximum pooling layer to obtain a first sub-compressed physiological signal;
[0306] Compressing the first weighted physiological signal in the channel dimension through an average pooling layer to obtain a second sub-compressed physiological signal;
[0307] The first sub-compressed physiological signal and the second sub-compressed physiological signal are concatenated to obtain a third compressed physiological signal.
[0308] Optionally, the identification module can be used to:
[0309] The second weighted physiological signal is processed by a physiological signal recognition network to obtain the type of the physiological signal.
[0310] In summary, the embodiments of the present application provide a human-based intelligent physiological signal processing device, which can perform weighted processing on physiological signals through a channel attention network to obtain a first weighted physiological signal, and perform weighted processing on the first weighted physiological signal through a spatial attention network to obtain a second weighted physiological signal. Then, based on the second weighted physiological signal, the type of physiological signal can be identified. Since the number of channels of the physiological signal is not compressed during the process of processing the physiological signal by the channel attention network, the computational complexity of the channel attention network is reduced, thereby improving the efficiency of identifying the type of physiological signal.
[0311] Example 4
[0312] The following describes, with reference to the accompanying drawings, a method and device for processing physiological signals based on human factors intelligence, a method and device for extracting HRV features, an electronic device, and a computer-readable storage medium provided in embodiments of the present application.
[0313] An embodiment of the present application provides a method for processing physiological signals based on human intelligence, which may include the following steps.
[0314] Step 1: Acquire the original physiological signal and perform slicing processing on the original physiological signal to obtain multiple physiological signal segments.
[0315] To extract more accurate features from raw physiological signals and provide a more accurate benchmark for medical diagnosis, the raw physiological signals can be purified to remove noise. However, if the raw physiological signals are long, the purification effect may not be very good. Therefore, the raw physiological signals can be sliced and then purified on the sliced physiological signal segments.
[0316] In the embodiment of the present application, the original physiological signal may include an electrocardiogram signal and an electroencephalogram signal. For example, the original physiological signal may be an ECG (electrocardiogram) signal or an EEG (electroencephalogram) signal.
[0317] According to one embodiment of the present application, a specific implementation of slicing the original physiological signal may include: segmenting the original physiological signal using time windows, wherein at least adjacent time windows overlap, and each time window remains the same, that is, each time window has the same time length.
[0318] In an embodiment of the present application, the sampling frequency can be set according to the actual duration of the original physiological signal, that is, the duration of the time window can be set to ensure that the original physiological signal can be evenly divided. In addition, compared with signal segmentation using a non-overlapping method, overlapping time windows can obtain more data; and because the waveform of the original physiological signal may have phase differences, each time the waveform of the original physiological signal is obtained, it does not necessarily start from the P wave or T wave position of the original physiological signal, but may start from any position of the original physiological signal. Using adjacent time windows with overlapping can more accurately simulate the original physiological signal and improve the accuracy of subsequent processing.
[0319] In a specific implementation, for a raw physiological signal with a time duration of T seconds and k data points, the time window can be set to L seconds, the delay duration to o seconds, and the number of data points within o seconds to y. The raw physiological signal is then segmented into multiple physiological signal segments based on this time window. Each physiological signal segment has a time duration of L seconds and contains m data points. Here, L is less than T, o is less than L, m is less than k, and y is less than m.
[0320] FIG12 is a schematic diagram of a time window provided according to an embodiment of the present application. In FIG12 , time windows t0 and t1 are adjacent, and the delay duration is o seconds.
[0321] As an example, each data in the original physiological signal having a time length of T seconds and including k data is defined as {x0, x1, ..., x k-1}, and divide it into n physiological signal segments with a time window of L, each physiological signal segment is defined as t0, t1, ...t n-1 , and t0={x0,x1,...,x m-1},t1={x y , x y+1 ,...,x y+m-1},...,t n-1 ={x k-m , x k-m+1 ,...,x k-1}.
[0322] In an embodiment of the present application, the acquired original physiological signal is segmented using time windows of the same duration to obtain multiple physiological signal segments of shorter duration with overlapping time windows, which facilitates subsequent purification processing.
[0323] Step 2: Input multiple physiological signal segments into a pre-trained neural network model for purification to obtain the target physiological signal, wherein the neural network model is trained based on the true value of the original physiological signal as a label.
[0324] In the embodiment of the present application, the true value is a purified physiological signal obtained by bandpass filtering the original physiological signal, removing the noise in the original physiological signal. In a specific implementation, each physiological signal segment can be bandpass filtered based on the measured physiological signal value of the physiological signal segment, and the bandpass frequency range is (Hr-5, Hr+5). The signal obtained after bandpass filtering is used as the true value. Where Hr is the physiological signal value.
[0325] As an example, bandpass filtering of physiological signal segments can be achieved by the following formula (1): i =FIR(t i )i=0,1,2,...,n-1 (1)
[0326] In formula (1), FIR represents the signal filter used, which can be an ideal bandpass, Butterworth filtering, IIR filtering or other signal filtering methods. The frequency band range is Hr-5 to Hr+5. Hr represents the true physiological signal value of the physiological signal segment ti. Fti is the signal after bandpass filtering of the physiological signal segment ti, that is, the true value corresponding to the physiological signal segment ti.
[0327] In some embodiments, if the true physiological signal value cannot be determined, a narrower frequency band range cannot be used for filtering processing. A common physiological signal value range can be selected as the frequency band range, and a wider frequency band range can be used to consider the physiological signal value ranges of different age groups and different situations.
[0328] As an example, taking the original physiological signal as an original electrocardiogram signal, the physiological signal value may be a heart rate value, the physiological signal value range may be a heart rate range, and the heart rate range is a normal heart rate range for a person, such as [40, 200], where 40 means a person's heart rate is 40 beats / minute, and 200 means a person's heart rate is 200 beats / minute. Because the heart rate range of athletes and newborn babies needs to be considered, a wider heart rate range can be selected as the frequency band range.
[0329] In some embodiments, the bandpass filtered signal can be used as the true value, or label, during the neural network model training process, and multiple physiological signal segments can be used as training data for the neural network model. The neural network model is iteratively trained based on the training data and labels to obtain a trained neural network model. In a specific implementation, multiple training data can be input into the neural network model to obtain multiple predicted signals. Based on each predicted signal and the corresponding label, the parameters of the neural network model are adjusted using a loss function until the true value and the predicted signal are substantially the same, thereby obtaining a trained neural network model.
[0330] The loss function can be any loss function such as MES (Mean Squared Error) or cross entropy loss function.
[0331] In an embodiment of the present application, after the neural network model is pre-trained, multiple physiological signal segments can be input into the pre-trained neural network model for purification processing, and its specific implementation may include: performing downsampling operations on multiple physiological signal segments based on the Focus downsampling network structure to obtain a sampling data group; performing splicing operations on the sampling data group to obtain first spliced data; performing time filtering operations on the spliced data based on multiple filtering intervals to obtain multiple filtered data; performing splicing operations on multiple filtered data to obtain second spliced data; encoding and decoding the second spliced data to obtain the target physiological signal.
[0332] In some embodiments, the neural network model can be any model that can achieve signal purification (filtering). Exemplarily, the neural network model can include at least one of an EEGNet (Electroencephalogram) model, a GAN (Generative Adversarial Nets) model, and a Transformer model. In addition, the neural network model can include a Focus downsampling unit, a time filtering unit, an encoding unit, a decoding unit, and two splicing units, and the Focus downsampling unit adopts a Focus downsampling structure for downsampling physiological signal segments, the time filtering unit is used to perform a time filtering operation on the first spliced data, and the encoding unit and the decoding unit are used to encode and decode the second spliced data.
[0333] In a specific implementation, multiple physiological signal segments can be input into the Focus downsampling unit for downsampling operations respectively, and a sampling data group corresponding to each physiological signal segment can be obtained. Then, the sampling data group corresponding to the same physiological signal segment is spliced by the first splicing unit to obtain the first spliced data corresponding to each physiological signal segment. The first spliced data is input into the time filtering unit, and time filtering processing of different sizes is performed in multiple different filtering intervals to obtain multiple filtered data corresponding to each physiological signal segment. Then, the multiple filtered data corresponding to the same physiological signal segment is spliced by the second splicing unit to obtain the second spliced data corresponding to each physiological signal segment. The second spliced data is input into the encoding unit for encoding processing, and the encoding result is input into the decoding unit for decoding processing to obtain the target physiological signal segment corresponding to each physiological signal segment. By splicing multiple target physiological signal segments, the target physiological signal, that is, the signal after the original physiological signal is purified, can be obtained.
[0334] As an example, the downsampling operation of the Focus downsampling unit is to intermittently sample the physiological signal segment. Assume that the physiological signal segment is defined as T = [t0, t1, ..., t n], if sampling is performed at time interval 1, the downsampling value sequence is T0 = [t0, t2, t4, ...], T1 = [t1, t3, t5, ...], then the Focus downsampling unit includes two groups of downsampling value sequences T0 and T1, if sampling is performed at time interval 2, T0 = [t0, t3, t6, ...], T1 = [t1, t4, t7, ...], T2 = [t2, t5, t8, ...], then the Focus downsampling unit includes three groups of downsampling value sequences T0, T1 and T2, and the same applies to other time intervals. In other words, the Focus downsampling unit may include at least two downsampling value sequences, and for each physiological signal segment, the at least two downsampling value sequences are used respectively to obtain at least two sampling data groups corresponding to each physiological signal segment. Then, the at least two sampling data groups corresponding to the same physiological signal segment are spliced to obtain the first spliced data corresponding to each physiological signal segment.
[0335] As an example, the temporal filtering unit includes multiple filter kernels of different sizes, which are used to perform spatial filtering of different sizes on the input first spliced data in the temporal dimension, similar to the convolution filtering of two-dimensional images in a neural network model. In addition, the larger the filter kernel, the better the effect of extracting low-frequency signals, and the smaller the filter kernel, the better the effect of extracting high-frequency signals. In actual use, the size of the filter kernel can be set or the filter kernel used can be selected according to needs. For example, assuming that the temporal filtering unit includes three filter kernels, the filter kernel of k0 is [1,64], the filter kernel of k1 is [1,128], and the filter kernel of k2 is [1,256]. In addition, for the same physiological signal segment, the number of filtered data obtained after processing by the temporal filtering unit is the same as the number of filter kernels. The second splicing unit uses a set of spatial filters to perform a splicing operation on the multiple filtered data corresponding to the same physiological signal segment. Specifically, different weights are assigned to the multiple filtered data to obtain the second spliced data corresponding to the physiological signal segment.
[0336] As an example, the encoding unit may include a convolution layer, a Bn (Batch Normalization) layer, a pooling layer, and an activation layer, and the decoding unit may include a deconvolution layer, a Bn layer, a depooling layer, and an activation layer. In a specific implementation, the second spliced data is input into the encoding unit, and feature extraction is first performed through the convolution layer to obtain a first result. The first result is input into the Bn layer to normalize the distribution of data in the first result to obtain a second result. The second result is input into the pooling layer for downsampling processing to obtain a third result. The third result is then input into the activation layer to obtain an encoding result. The encoding result is input into the decoding unit, and upsampling is first performed through the deconvolution layer to obtain a fourth result. The fourth result is input into the Bn layer to normalize the distribution of data in the fourth result to obtain a fifth result. The fifth result is input into the depooling layer for upsampling processing to obtain a sixth result. The sixth result is then input into the activation layer to obtain a target physiological signal segment.
[0337] For example, FIG13 is a structural diagram of a neural network model provided according to an embodiment of the present application. The neural network model includes 6 layers, layer0 is the input of the neural network model, that is, the original physiological signal (divided into multiple physiological signal segments), layer1 is the Focus downsampling unit, including four downsampling value sequences of Focus0, Focus1, Focus2 and Focus3, which are represented by sequences 0 to 3 in FIG13, respectively, and can downsample the signal without losing data information. Layer2 is the first splicing unit for layer 1. The output of er1 is spliced, layer3 is the temporal filtering unit, including three filter kernels k0, k1 and k2, which are used to perform temporal filtering of different sizes on the output of layer2, layer4 is the second splicing unit, and a set of spatial filters is used to splice the output of layer3, layer5 is the encoding unit, including the convolution layer, Bn layer, pooling layer and activation layer, layer6 is the decoding unit, including the deconvolution layer, Bn layer, depooling layer and activation layer, layer7 is the output of the neural network model, that is, the target physiological signal, which is represented by layers 0 to 7 in Figure 13 respectively.
[0338] Taking the original physiological signal as an electrocardiogram signal as an example, an electrocardiogram signal processing method provided in an embodiment of the present application can be to obtain an electrocardiogram signal with a time length of T, divide the electrocardiogram into n electrocardiogram slice signals ti (i=0,1,...n-1) with overlapping time windows, and use a neural network model to purify the n electrocardiogram slice signals to obtain a target electrocardiogram signal.
[0339] In an embodiment of the present application, the original physiological signal is sliced, and the physiological signal segments obtained by slicing are input into a neural network model for purification processing, that is, the noise in the original physiological signal is filtered out to obtain a target physiological signal. The target physiological signal has a higher signal-to-noise ratio and can be used to extract higher-quality features, thereby improving the accuracy of feature extraction and providing a more accurate judgment basis for medical diagnosis. Moreover, slicing the original physiological signal before purification and purifying the physiological signal segments can improve the purification accuracy and achieve a better denoising effect.
[0340] It can be understood that the purification processing of the original physiological signal in the embodiment of the present application can also be called denoising processing. The two are essentially the same in technology. For example, in Example 1, the electrocardiogram signal of each user is denoised, and the processing method of the following steps can be referred to.
[0341] An HRV feature extraction method provided in an embodiment of the present application may include the following steps.
[0342] Step 1: Obtain the original electrocardiogram signal and segment the original electrocardiogram signal to obtain multiple overlapping electrocardiogram slice signals.
[0343] Step 2: Input multiple ECG slice signals into a pre-trained neural network model for purification processing to obtain a target ECG signal, wherein the neural network model is trained based on the purified ECG signal of the original ECG signal as a label, and the purified ECG signal is obtained by bandpass filtering the original ECG signal.
[0344] It should be noted that the implementation process of steps one and two is similar to that in the aforementioned embodiment, except that the original physiological signal is replaced with the original electrocardiogram signal. Therefore, the specific implementation process of steps one and two can refer to the relevant description of the aforementioned embodiment, and this embodiment will not repeat it.
[0345] For example, referring to Figures 14 and 15, Figure 14 is a schematic diagram of an original electrocardiogram signal provided according to an embodiment of the present application, and Figure 15 is a schematic diagram of a target electrocardiogram signal provided according to an embodiment of the present application. The electrocardiogram slice signal input to the neural network model is the ECG signal with burrs in Figure 14, and the target electrocardiogram signal output by the neural network model is the relatively smooth ECG signal in Figure 15. When training the neural network model, it is necessary to give a real desired waveform, and the collected ECG signal is often in the form of Figure 14. Therefore, while collecting ECG data, the heart rate value at that time can be recorded in real time through a pulse oximeter device or an ECG data acquisition device with heart rate measurement. According to the real-time heart rate value, the ECG signal of Figure 14 is band-pass filtered by a relatively narrow band-pass filter to obtain the ECG signal of Figure 15. During the model training process, the ECG signal of Figure 14 can be used as the input of the neural network model, and the ECG signal of Figure 15 can be used as the label (i.e., the true value) to train the neural network model, so that the trained neural network model can purify the ECG signal similar to Figure 14 to obtain the ECG signal similar to Figure 15.
[0346] Step 3: Determine the peak position and trough position of the target ECG signal, and determine the HRV feature based on the peak position and trough position.
[0347] Among them, HRV (Heart Rate Variability) is the time difference between heartbeats. It is an important indicator reflecting the regulation of the autonomic nervous system and one of the important indicators for assessing heart health.
[0348] In a specific implementation, the peak distance between each two adjacent peaks or the trough distance between each two adjacent troughs can be determined based on the peak position and trough position of the target electrocardiogram signal. An HRV signal can be formed based on the peak distance or trough distance. Various HRV features such as time domain features, frequency domain features, and Poincare features can then be extracted from the HRV signal to assist in medical diagnosis.
[0349] For example, see Figure 16, which is a schematic diagram of another target electrocardiogram signal provided according to an embodiment of the present application. First, find the position of the peaks and troughs in the ECG signal, calculate the time length between adjacent peaks (in ms), or calculate the time length between adjacent troughs. The time signal formed by the time length of every two adjacent peaks or the time length of every two adjacent troughs is the HRV signal. A variety of HRV features can be extracted from the HRV signal, such as time domain features (there are more than a dozen), frequency domain features (about 7 to 8 commonly used), Poincare features, and other features.
[0350] The HRV feature extraction method provided in the embodiment of the present application obtains an original electrocardiogram signal, and segments the original electrocardiogram signal to obtain multiple overlapping electrocardiogram slice signals; the multiple electrocardiogram slice signals are input into a pre-trained neural network model for purification processing to obtain a target electrocardiogram signal, wherein the neural network model is trained based on the purified electrocardiogram signal of the original electrocardiogram signal as a label, and the purified electrocardiogram signal is obtained by bandpass filtering the original electrocardiogram signal; the peak position and trough position of the target electrocardiogram signal are determined, and the HRV feature is determined based on the peak position and trough position. The above method purifies the original electrocardiogram signal through a pre-trained neural network model, filters out the noise in the original electrocardiogram signal, and obtains a purer target electrocardiogram signal. Then, the HRV feature is extracted based on the target electrocardiogram signal, which can reduce the interference of noise, improve the accuracy of feature extraction, and thus improve the accuracy of medical diagnosis. Moreover, slicing the original electrocardiogram signal before purification and purifying the electrocardiogram slice signal can improve the purification accuracy and achieve better denoising effect.
[0351] It should be noted that the above embodiment uses an electrocardiogram (ECG) signal as an example to illustrate HRV feature extraction. If the original physiological signal is an EEG signal, the heart rate value can be replaced with the frequency band of the EEG signal, such as the Delta band (bandpass less than 4 Hz), the Theta oscillation band (bandpass 4-7 Hz), the Alpha band (bandpass 7-12 Hz), the Beta band (bandpass 12-30 Hz), and the Gamma band (bandpass 30-50 Hz or higher). Furthermore, waveforms in different frequency bands have similar functions on the human body. Therefore, the EEG signal can also be processed as described above, and a neural network model capable of purifying the EEG signal can be trained. This neural network model can then be used to extract the original EEG signal to obtain a target EEG signal. Features related to medical diagnosis can then be extracted based on the target EEG signal to provide a basis for medical diagnosis. If the original physiological signal is an EEG signal, the specific implementation can refer to the relevant description of the above embodiment, and this embodiment will not be repeated here.
[0352] An embodiment of the present application provides a physiological signal processing device based on human factors intelligence, which may include:
[0353] A first acquisition module is used to acquire original physiological signals;
[0354] The first segmentation module is used to slice the original physiological signal to obtain multiple physiological signal segments;
[0355] The first processing module is used to input multiple physiological signal segments into a pre-trained neural network model for purification processing to obtain target physiological signals, wherein the neural network model is trained based on the true value of the original physiological signal as a label.
[0356] According to one embodiment of the present application, the first segmentation module is further configured to:
[0357] The original physiological signal is segmented using time windows, wherein at least adjacent time windows overlap.
[0358] According to one embodiment of the present application, each time window remains the same.
[0359] According to one embodiment of the present application, the true value is obtained based on band-pass filtering of the original physiological signal.
[0360] According to one embodiment of the present application, the first processing module is further configured to:
[0361] Based on the Focus downsampling network structure, multiple physiological signal segments are downsampled to obtain a sampling data group;
[0362] Performing a splicing operation on the sample data group to obtain first spliced data;
[0363] Performing time filtering operations on the spliced data based on multiple filtering intervals to obtain multiple filtered data;
[0364] Performing a splicing operation on the plurality of filtered data to obtain second spliced data;
[0365] The second spliced data is encoded and decoded to obtain a target physiological signal.
[0366] According to one embodiment of the present application, the original physiological signal includes an electrocardiogram signal and an electroencephalogram signal.
[0367] According to one embodiment of the present application, the neural network model includes at least one of an EEGNet model, a GAN model, and a Transformer model.
[0368] By applying the human intelligence-based physiological signal processing method provided in the embodiments of this application, the raw physiological signals are purified using a pre-trained neural network model, filtering out noise from the raw physiological signals to obtain a purer target physiological signal. Feature extraction based on this target physiological signal can reduce noise interference, improve the accuracy of feature extraction, and thus enhance the accuracy of medical diagnosis. Furthermore, slicing the raw physiological signals before purification and purifying the physiological signal segments can improve purification accuracy and achieve better denoising effects.
[0369] The above is a schematic diagram of a physiological signal processing device based on human intelligence according to an embodiment of the present application. It should be noted that the technical solution of the physiological signal processing device based on human intelligence and the technical solution of the physiological signal processing method based on human intelligence are based on the same concept. For details not described in detail in the technical solution of the physiological signal processing device based on human intelligence, please refer to the description of the technical solution of the physiological signal processing method based on human intelligence.
[0370] The present application provides an HRV feature extraction device, which may include:
[0371] The second acquisition module is used to acquire the original electrocardiogram signal;
[0372] A second segmentation module is used to segment the original electrocardiogram signal to obtain a plurality of overlapping electrocardiogram slice signals;
[0373] a second processing module, configured to input the plurality of electrocardiogram slice signals into a pre-trained neural network model for purification processing to obtain a target electrocardiogram signal, wherein the neural network model is trained based on the purified electrocardiogram signal of the original electrocardiogram signal as a label, and the purified electrocardiogram signal is obtained by bandpass filtering the original electrocardiogram signal;
[0374] The determination module is used to determine the peak position and the trough position of the target electrocardiogram signal, and determine the HRV characteristics based on the peak position and the trough position.
[0375] According to one embodiment of the present application, the second segmentation module is further configured to:
[0376] The original electrocardiogram signal is segmented using time windows, wherein at least adjacent time windows overlap.
[0377] According to one embodiment of the present application, the second processing module is further configured to:
[0378] Based on the Focus downsampling network structure, downsampling operations are performed on multiple ECG slice signals to obtain a sampling data group;
[0379] Performing a splicing operation on the sample data group to obtain first spliced data;
[0380] Performing time filtering operations on the spliced data based on multiple filtering intervals to obtain multiple filtered data;
[0381] Performing a splicing operation on the plurality of filtered data to obtain second spliced data;
[0382] The second spliced data is encoded and decoded to obtain a target electrocardiogram signal.
[0383] By applying the HRV feature extraction method provided in the embodiments of this application, the raw ECG signal is purified using a pre-trained neural network model, filtering out noise from the raw ECG signal to obtain a purer target ECG signal. Extracting HRV features based on this target ECG signal can reduce noise interference, improve the accuracy of feature extraction, and thereby enhance the accuracy of medical diagnosis. Furthermore, slicing the raw ECG signal before purification and then purifying the ECG slice signal can improve purification accuracy and achieve better denoising effects.
[0384] The above is a schematic diagram of an HRV feature extraction device according to an embodiment of the present application. It should be noted that the technical solution of this HRV feature extraction device and the technical solution of the aforementioned HRV feature extraction method are based on the same concept. For details not described in detail in the technical solution of the HRV feature extraction device, please refer to the description of the technical solution of the aforementioned HRV feature extraction method.
[0385] Figure 17 is a schematic diagram of the structure of an electronic device provided according to an embodiment of the present application. The electronic device 3100 includes: a memory 3101, a processor 3102, and a computer program stored in the memory 3101 and executable on the processor 3102. When the processor 3102 executes the computer program, it implements the training method of the multimodal human physiological data classification model and the classification method of multimodal human physiological data as shown in the above-mentioned embodiment 1, as well as the method for generating EEG signal sample data and the training method for the physiological state prediction model as shown in the embodiment 2, as well as the human factor intelligence-based physiological signal processing method as shown in the embodiment 3, and the human factor intelligence-based physiological signal processing method or HRV feature extraction method as shown in the embodiment 4.
[0386] The present application also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the training method of a multimodal human physiological data classification model and the classification method of multimodal human physiological data as shown in Example 1, as well as the method for generating EEG signal sample data and the training method for a physiological state prediction model as shown in Example 2, as well as the human-based intelligent physiological signal processing method as shown in Example 3, and the human-based intelligent physiological signal processing method or HRV feature extraction method as shown in Example 4.
[0387] One embodiment of the present application provides a computer program product, which includes instructions. When the instructions are executed by a processor of a computer device, the computer device is enabled to execute the training method of the multimodal human physiological data classification model and the classification method of the multimodal human physiological data shown in the above-mentioned embodiment 1, as well as the method for generating EEG signal sample data and the training method for the physiological state prediction model shown in embodiment 2, as well as the human-based intelligent physiological signal processing method shown in embodiment 3, and the human-based intelligent physiological signal processing method or HRV feature extraction method shown in embodiment 4.
[0388] Those skilled in the art should be able to appreciate that the method steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of electronic hardware and software, the composition and steps of each example have been generally described in terms of function in the above description. Whether these functions are performed in electronic hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art may use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.
[0389] Thus far, the technical solutions of the present application have been described in conjunction with the preferred embodiments shown in the accompanying drawings. However, it will be readily understood by those skilled in the art that the scope of protection of the present application is clearly not limited to these specific embodiments. Without departing from the principles of the present application, those skilled in the art may make equivalent changes or substitutions to the relevant technical features, and the technical solutions after such changes or substitutions will fall within the scope of protection of the present application.
Claims
1. A classification model for multimodal human physiological data, the classification model comprising: Multi-head self-attention module, normalization module, fusion expert system and decision module; The fusion expert system includes: an EEG expert subsystem, an ECG expert subsystem, an electrodermal expert subsystem and a multi-modal synchronous fusion expert subsystem; The multi-head self-attention module is used to extract features from the input multimodal synchronous data and output them to the normalization module; the multimodal synchronous data includes: physiological indicators of the user; the physiological indicators include: EEG data, ECG data and skin conductance data; The normalization module is used to generate normalized feature data according to the extracted feature data, and output the normalized feature data to the fusion expert system; The EEG expert subsystem is used to perform a classification task according to the normalized feature data corresponding to the EEG data of the user to obtain a first classification result; The ECG expert subsystem is used to perform a classification task according to the normalized feature data corresponding to the ECG data of the user to obtain a second classification result; The skin electrical expert subsystem is used to perform a classification task according to the normalized feature data corresponding to the skin electrical data of the user to obtain a third classification result; The multimodal synchronous fusion expert subsystem is used to perform a classification task according to the normalized feature data corresponding to the multimodal synchronous data to obtain a fourth classification result; The decision module is used to calculate a final classification result based on the first classification result, the second classification result, the third classification result and the fourth classification result, as well as the weight corresponding to each classification result.
2. The classification model of multimodal human physiological data according to claim 1, wherein: The step of "calculating a final classification result according to the first classification result, the second classification result, the third classification result, the fourth classification result and the weight corresponding to each classification result" includes: The final classification result is calculated according to the following formula: G=W1*f1+W2*f2+W3*f3+W4*f4 Among them, G represents the final classification result, f1, f2, f3 and f4 represent the first classification result, the second classification result, the third classification result and the fourth classification result respectively, and W1, W2, W3 and W4 all represent weights.
3. A training method for a multimodal human physiological data classification model, applicable to the multimodal human physiological data classification model as claimed in any one of claims 1 to 2, the training method comprising: Using the EEG training set to train the EEG expert subsystem and the multi-head self-attention module, and freezing the weights of the EKG expert subsystem, the GEP expert subsystem, and the multimodal synchronous fusion expert subsystem; The ECG expert subsystem is trained using the ECG training set, and the weights of the multi-head self-attention module, the EEG expert subsystem, the electrodermal expert subsystem, and the multimodal synchronous fusion expert subsystem are frozen; The electrodermal expert subsystem is trained using the electrodermal training set, and the weights of the multi-head self-attention module, the EEG expert subsystem, the ECG expert subsystem, and the multimodal synchronous fusion expert subsystem are frozen; The classification model is trained using a multimodal random coding training set, thereby adjusting the weights of the multi-head self-attention module, the EEG expert subsystem, the ECG expert subsystem, the skin electricity expert subsystem, and the multimodal synchronous fusion expert subsystem.
4. The training method of the multimodal human physiological data classification model according to claim 3, wherein: The EEG training set includes original EEG signal sample data and target EEG signal sample data. The EEG expert subsystem is trained using the EEG training set, including: Acquire target EEG signal sample data, wherein the target EEG signal sample data is determined based on EEG signal prior features, and the EEG signal prior features are obtained based on original EEG signal sample data; Based on the target EEG signal sample data, auxiliary training is performed on the EEG expert subsystem to be trained to obtain the EEG expert auxiliary subsystem; and Based on the original EEG signal sample data, the EEG expert auxiliary subsystem is trained to obtain a trained EEG expert subsystem.
5. The training method of the multimodal human physiological data classification model according to claim 4, wherein: The step of obtaining target EEG signal sample data includes: Processing the original EEG signal sample data to obtain the EEG signal prior features; and Construct first target EEG signal sample data according to the EEG signal priori features.
6. The training method of the multimodal human physiological data classification model according to claim 4, wherein: The step of obtaining target EEG signal sample data includes: Processing the original EEG signal sample data to obtain the EEG signal prior features; and The original EEG signal sample data and the EEG signal prior features are combined and processed to obtain the second target EEG signal sample data.
7. The training method for a multimodal human physiological data classification model according to claim 5 or 6, wherein: The processing of the original EEG signal sample data to obtain the EEG signal prior features includes: According to the preset characteristic index, the original EEG signal sample data is processed to obtain a characteristic value corresponding to the preset characteristic index as the prior feature of the EEG signal.
8. The training method of the multimodal human physiological data classification model according to claim 5, wherein: The step of performing auxiliary training on the EEG expert subsystem to be trained based on the target EEG signal sample data to obtain the EEG expert auxiliary subsystem includes: Inputting the first target EEG signal sample data into the EEG expert subsystem to be trained for prediction, and obtaining a first sub-classification prediction result representing a physiological state; and Based on the first sub-classification prediction result, the model parameters of the EEG expert subsystem to be trained are reversely adjusted to obtain the EEG expert auxiliary subsystem.
9. The training method of the multimodal human physiological data classification model according to claim 6, wherein: The step of performing auxiliary training on the EEG expert subsystem to be trained based on the target EEG signal sample data to obtain the EEG expert auxiliary subsystem includes: Inputting the second target EEG signal sample data into the EEG expert subsystem to be trained for prediction, and obtaining a second detailed classification prediction result and a first coarse classification prediction result representing a physiological state; and Based on the second detailed classification prediction result and the first coarse classification prediction result, the model parameters of the EEG expert subsystem to be trained are reversely adjusted to obtain the EEG expert auxiliary subsystem.
10. The training method for a multimodal human physiological data classification model according to claim 8 or 9, wherein: The step of training the EEG expert auxiliary subsystem based on the original EEG signal sample data to obtain a trained EEG expert subsystem includes: Inputting the original EEG signal sample data into the EEG expert assistance subsystem for prediction to obtain a second coarse classification prediction result representing the physiological state; and Based on the second coarse classification prediction result, the model parameters of the EEG expert auxiliary subsystem are reversely adjusted to obtain the trained EEG expert subsystem.
11. The training method of the multimodal human physiological data classification model according to claim 3, wherein: The method further comprises: Generate EEG data including time domain information, frequency domain information and spatial information according to the EEG signal of each user and the corresponding electrode point position, thereby obtaining the EEG training set; De-noising the ECG signal of each user, and aligning it with the EEG signal of the user according to time information, thereby generating the ECG training set; De-noising the electrodermal signal of each user, and aligning it with the electroencephalogram signal of the user according to time information, thereby generating the electrodermal training set; The EEG training set, the ECG training set and the EDG training set are randomly masked to simulate data missing conditions, thereby generating the multimodal randomly coded training set.
12. The training method of the multimodal human physiological data classification model according to claim 11, wherein: The step of "generating EEG data containing time domain information, frequency domain information and spatial information according to the EEG signal of each user and the corresponding electrode point position, thereby obtaining the EEG training set" includes: Perform fast Fourier transform on each user’s EEG signal to extract frequency domain information; Selecting data of a frequency band of interest from the frequency domain information; According to the data of the frequency band of interest and the electrode point positions when collecting the EEG signal, a quasi-multispectral image based on different leads is obtained; The multi-spectral image is processed into blocks, and the EEG data of the user is generated according to the block data.
13. The training method of the multimodal human physiological data classification model according to claim 12, wherein: The step of "obtaining a quasi-multispectral image based on different leads according to the data of the frequency band of interest and the electrode point positions when collecting the EEG signal" includes: Projecting the electrode point positions when collecting the EEG signal from the three-dimensional space to a two-dimensional surface to obtain the two-dimensional position information of each electrode point; Calculate the distance between each electrode point and other surrounding electrode points according to the two-dimensional position information; For the data of each lead in the data of the frequency band of interest, corresponding weights are set for the data of other leads according to the distance, so as to obtain quasi-multispectral images based on different leads.
14. The training method of the multimodal human physiological data classification model according to claim 12, wherein: The step of "block processing of the multi-spectral image and generating the EEG data of the user according to the block data" includes: The multispectral image is divided into blocks as a picture, and is divided into N = H * W / P 2 small blocks of the same size; where H, W, C, and P represent the height, width, number of channels, and size of the small block of the image, respectively; Flatten the image in each small block into a vector, obtain the patch embedding through linear projection, and add a [I_CLS] token as the position code of the small block, so as to obtain the block data corresponding to each small block; Generate the user's EEG data according to the block data.
15. The training method of the multimodal human physiological data classification model according to claim 13, wherein: The step of "projecting the electrode point positions when collecting the EEG signal from the three-dimensional space to the two-dimensional surface to obtain the two-dimensional position information of each electrode point" includes: The electrode point positions when collecting the electroencephalogram signal are converted from three-dimensional space to two-dimensional surface by using azimuth equidistant projection in a Cartesian coordinate system, so that the converted lead data have a spatial topological relationship.
16. The training method of the multimodal human physiological data classification model according to claim 11, wherein: The denoising process of the electrocardiogram signal of each user comprises: Acquire an original electrocardiogram signal, and perform slicing processing on the original electrocardiogram signal to obtain a plurality of electrocardiogram signal segments; The multiple ECG signal segments are input into a pre-trained neural network model for denoising to obtain a target ECG signal, wherein the neural network model is trained based on the true value of the original ECG signal as a label.
17. The training method of the multimodal human physiological data classification model according to claim 16, wherein: The slicing process of the original electrocardiogram signal includes: The original electrocardiogram signal is segmented using time windows, wherein at least adjacent time windows overlap.
18. The training method of the multimodal human physiological data classification model according to claim 17, wherein: The step of inputting the plurality of ECG signal segments into a pre-trained neural network model for denoising includes: Performing downsampling operations on the multiple ECG signal segments based on the Focus downsampling network structure to obtain a sampling data group; Performing a splicing operation on the sampled data group to obtain first spliced data; Based on multiple filtering intervals, the spliced data are respectively subjected to time filtering operations to obtain multiple filtered data; Performing a splicing operation on the plurality of filtered data to obtain second spliced data; The second spliced data is encoded and decoded to obtain the target electrocardiogram signal.
19. A method for classifying multimodal human physiological data, the method utilizing the classification model for multimodal human physiological data as described in any one of claims 1-2 to perform a classification task.
20. The method for classifying multimodal human physiological data according to claim 19, wherein: Before using the classification model of the multimodal human physiological data to perform a classification task, the method includes: Performing weighted processing on the input physiological data through a channel attention network to obtain first weighted physiological data, wherein the number of channels of the physiological data is not compressed during the process of the channel attention network processing the physiological data; Performing weighted processing on the first weighted physiological data through a spatial attention network to obtain second weighted physiological data; Based on the second weighted physiological data, a type of the physiological data is obtained.
21. The method for classifying multimodal human physiological data according to claim 20, wherein: The channel attention network includes: a first sub-network and a second sub-network; The input physiological data is weighted by the channel attention network to obtain first weighted physiological data, including: Inputting the physiological data into the first sub-network to obtain a first channel weight vector, wherein the number of channels of the physiological data is not compressed during the process of the first sub-network processing the physiological data; Inputting the physiological data into the second sub-network to obtain a second channel weight vector, wherein the number of channels of the physiological data is not compressed during the process of the second sub-network processing the physiological data; The first weighted physiological data is obtained based on the first channel weight vector, the second channel weight vector and the physiological data.
22. The method for classifying multimodal human physiological data according to claim 21, wherein: The obtaining the first weighted physiological data based on the first channel weight vector, the second channel weight vector and the physiological data comprises: Performing weighted summation on the first channel weight vector and the second channel weight vector to obtain a target channel weight vector; The target channel weight vector is used to perform weighted processing on the physiological data to obtain first weighted physiological data.
23. The method according to any one of claims 20 to 22, wherein: The spatial attention network includes: a pooling network and a convolution block; weighted processing is performed on the first weighted physiological data through the spatial attention network to obtain second weighted physiological data, including: Based on the pooling network, compressing the first weighted physiological data in the channel dimension to obtain third compressed physiological data; Performing convolution processing on the third compressed physiological data through the convolution block to obtain a target space weight vector; The first weighted physiological data is weighted by using the target space weight vector to obtain second weighted physiological data.
24. The method according to any one of claims 20 to 22, wherein: The obtaining the type of the physiological data based on the second weighted physiological data includes: The second weighted physiological data is processed by a physiological data recognition network to obtain the type of the physiological data.
25. A computer readable storage device, wherein: A computer program is stored which can be loaded by a processor and execute the method according to any one of claims 3 to 24.
Citation Information
Patent Citations
Multi-modal emotion recognition method based on self-attention mechanism
CN111553295A
Identity recognition method based on electroencephalogram signal class multispectral image sequence
CN114139573A
Psychological scale confidence assessment method and system based on multi-modal physiological data
CN115299947A
Multi-modal fusion attention assessment method and system based on VR, and storage medium
CN115329818A
Small sample sentiment analysis method and system based on masking language model and ladder network
CN116384407A
Cited By
Object 6D pose estimation method and system based on attention mechanism
CN120318325A
Hyperspectral and laser radar classification method based on explicit interaction and adaptive alignment
CN120656008A
Rehabilitation evaluation method and system based on motor imagery and storage medium
CN120954692A
Attention interpretable electroencephalogram decoding method based on adaptive fuzzy convolution and TSK guidance
CN121167432A
Autistic child evaluation system and device based on electroencephalogram large model
CN121313116A