Multi-mode postoperative child pain identification method

By constructing a multimodal network model, combining EEG, ECG and blood oxygen saturation data, and using frequency domain streaming neural network, time domain streaming neural network and CNN-GRU neural network, the problem of insufficient accuracy of traditional pain assessment methods was solved, and more accurate postoperative pain identification in children was achieved.

CN120661081APending Publication Date: 2025-09-19SOUTH CHINA UNIV OF TECH
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510663554.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-22
Publication Date
2025-09-19

AI Technical Summary

Technical Problem

In the existing technology, traditional pain assessment methods have the problems of strong subjectivity and insufficient accuracy in children's pain management. The accuracy of single-modal physiological signals for pain identification is not high, and multimodal data fusion methods need to be improved in feature extraction and model construction.

Method used

A multimodal network model was adopted to combine EEG, ECG and blood oxygen saturation data. By constructing a frequency domain streaming neural network, a time domain streaming neural network, a blood oxygen saturation CNN-GRU neural network and an ECG CNN-GRU neural network, Dense convolution blocks and Attention blocks were used for feature extraction and fusion, and an MLP feature fusion layer was used for pain classification.

Benefits of technology

It improves the accuracy and reliability of pain recognition, comprehensively characterizes the characteristics of pain-related EEG signals, and uses dynamic feature fusion to improve recognition accuracy, providing a more accurate postoperative pain management solution for children.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120661081A_ABST
    Figure CN120661081A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-mode postoperative child pain identification method. The method aims at improving the accuracy of child postoperative pain recognition by fusing ECG (electrocardiogram), SPO2 (blood oxygen) and EEG (electroencephalogram) multi-mode data. The method comprises the following steps of: firstly, respectively acquiring and preprocessing each modal data, and converting EEG (electroencephalogram) into two types of two-dimensional frames; then, feature extraction is automatically conducted on the data of all the modes through specific neural networks including the frequency domain flow neural network, the time domain flow neural network, the oxyhemoglobin saturation CNN-GRU neural network and the electrocardio CNN-GRU neural network; and finally, inputting the extracted features into an MLP feature fusion layer to carry out pain identification. Compared with a traditional single-mode identification method, the postoperative child pain related characteristics can be more comprehensively captured, the pain identification accuracy is effectively improved, and powerful support is provided for medical staff to accurately evaluate the postoperative child pain degree in time.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of medical data processing and pattern recognition, and particularly relates to a multimodal postoperative pain recognition method for children. Background Art

[0002] Postoperative pain management for children is of great significance to their recovery. However, due to children's limited cognitive and expressive abilities, traditional pain assessment methods are subjective and inaccurate. In recent years, the use of physiological signals for pain recognition has become a research direction. EEG signals reflect brain neural activity, ECG signals reflect heart function, and blood oxygen saturation reflects blood oxygen content, all of which are related to pain. However, the accuracy of existing single-modality physiological signals for pain recognition is not high. Although multimodal data fusion methods have advantages, they need to be improved in terms of feature extraction and model construction. There is an urgent need to solve these problems and provide a more accurate solution for postoperative pain recognition in children. Summary of the Invention

[0003] The main purpose of the present invention is to overcome the shortcomings and deficiencies of the existing technology and provide a multimodal method for identifying pain in children after surgery.

[0004] In order to achieve the above object, the present invention adopts the following technical solutions:

[0005] A multimodal method for identifying pain in children after surgery, comprising the following steps:

[0006] S1. Data collection steps: Collect EEG data, ECG data, and blood oxygen saturation data of children after surgery;

[0007] S2, feature extraction step: using the collected EEG data to construct a two-dimensional frame to obtain the time domain features of the EEG; calculating the differential entropy of multiple frequency intervals of the collected EEG data, using the obtained differential entropy to construct a two-dimensional frame to obtain the frequency features of the EEG;

[0008] S3. Multimodal network model construction steps: The multimodal network model includes four parallel connections: a frequency domain streaming neural network, a time domain streaming neural network, a blood oxygen saturation CNN-GRU neural network, an electrocardiogram CNN-GRU neural network, and an MLP feature fusion layer. The outputs of the four neural networks are concatenated and passed through the MLP feature fusion layer for pain classification.

[0009] S4. Model training step: The time domain features and frequency domain features of the training samples are obtained through the feature extraction step, and the time domain features and frequency domain features of the EEG data are used as input together with the ECG data and blood oxygen saturation data to train the multimodal network model.

[0010] S5. Pain recognition step for the sample to be tested: obtain time domain features and frequency domain features from the EEG data of the sample to be tested through the feature extraction step, input the time domain features and frequency domain features of the EEG data of the sample to be tested, together with the ECG and blood oxygen saturation data, into the trained multimodal network model to obtain the pain recognition result of the sample to be tested.

[0011] Furthermore, the process of step S1 is as follows:

[0012] At the same time, EEG data, ECG data, and blood oxygen saturation data of postoperative children for the same duration t were collected. Among them, EEG data were collected using a 32-channel EEG acquisition device, and the position of the electrodes was in accordance with the standard electrode placement method 10-20 system electrode placement method specified by the International Electroencephalography Society. Hereinafter, EEG is referred to as EEG, ECG as ECG, and blood oxygen saturation as SPO2. Medical staff gave each research subject's pain level as the data label based on the clinical pain assessment scale when collecting data. Synchronous acquisition and standard electrode placement methods were used to ensure the spatiotemporal consistency of multimodal signals and provide high-quality input for subsequent feature fusion. The 10-20 system electrode placement method covers the main functional areas of the brain.

[0013] Furthermore, the process of step S2 is as follows:

[0014] S2.1. Define a 7×7 two-dimensional matrix. The positions of the 32 electrodes collected by EEG are mapped to the two-dimensional matrix as follows: the first row, 3-5 columns are the voltage values ​​of the 01, 0z, and 02 channels respectively; the second row, 2-6 columns are the voltage values ​​of the P7, P3, Pz, P4, and P8 channels respectively; the third row, 1, 2, 3, 5, 6, and 7 columns are the voltage values ​​of the TP9, CP5, CP1, CP2, CP6, and TP10 channels respectively; the fourth row, 2, 3, 5, and 6 columns are the voltage values ​​of the T7, C3, C4, and T8 channels respectively; the fifth row, 1, Columns 2, 3, 5, 6, and 7 are the voltage values ​​of the FT9, FC5, FC1, FC2, FC6, and FT10 channels, respectively; columns 2-6 of the sixth row are the voltage values ​​of the F7, F3, Fz, F4, and F8 channels, respectively; columns 2-4 of the seventh row are the voltage values ​​of the FP1, Fpz, and FP2 channels, respectively. Adjacent matrix positions correspond to electrode signals in adjacent areas of the cerebral cortex, preserving the spatial relationship between brain regions and improving feature expression capabilities. This provides structured input for subsequent convolutional networks to extract spatial features and fully utilizes the spatial distribution information of the signal.

[0015] S2.2. Fill the 32 channel voltage values ​​at each sampling point of the EEG data into a two-dimensional matrix according to the above mapping relationship to form a two-dimensional frame. Fill 0.0 where there is no corresponding electrode in the two-dimensional matrix. Apply the minimum and maximum normalization method to the two-dimensional frame, mapping each element value in the two-dimensional frame to the range of 0-1, eliminating the influence of voltage differences between channels and accelerating model convergence. Apply the bicubic interpolation method to the two-dimensional frame, expanding the two-dimensional matrix from 7×7 to m×m, improving the resolution of the feature matrix and retaining more detailed information. Arrange the two-dimensional frames obtained at each sampling point into a three-dimensional matrix to obtain the time domain characteristics of the EEG;

[0016] S2.3. Define five pain-related EEG frequency intervals: 1-4 Hz, 4-8 Hz, 8-13 Hz, 13-30 Hz, and 30-45 Hz. This helps to explore the different physiological significance of different frequency components of EEG signals during pain.

[0017] For each EEG channel signal in the EEG data, the power spectral density is first calculated. The time domain signal is converted into a frequency domain signal through fast Fourier transform to obtain the power spectral density value at each frequency point. Then, the area under the power spectral density curve in the specified frequency band is calculated according to the specified EEG frequency interval. Finally, the differential entropy features of each frequency interval of the channel are obtained according to the formula DE = 0.5 × log (2 × π × e × P), where DE represents differential entropy, P represents probability density function, and e represents a natural constant. The above steps effectively capture the changes in EEG rhythms under different pain states by calculating multi-band differential entropy features. The larger the differential entropy value, the higher the uncertainty and complexity of the frequency band signal, which enhances the distinguishability of the feature.

[0018] The differential entropy features of the specific frequency interval of each channel are filled into a two-dimensional matrix according to the above mapping relationship to form a two-dimensional frame. 0.0 is filled in the places where there are no corresponding electrodes in the two-dimensional matrix. The two-dimensional frame is interpolated using the bicubic interpolation method, and the two-dimensional matrix is ​​expanded from 7×7 to m×m. The two-dimensional frames obtained from each sampling point are arranged into a three-dimensional matrix to obtain the frequency domain features of the EEG. The three-dimensional matrix structure is adapted to deep learning architectures such as CNN, laying the foundation for subsequent multimodal fusion and pain classification.

[0019] Furthermore, the process of step S3 is as follows:

[0020] S3.1. Construct a frequency domain streaming neural network. The network is connected from input to output in the following order: 1 convolution layer, 1 dense convolution block, 1 attention block, 1 batch normalization layer, 1 Relu layer and 1 global average pooling layer. Suppose the frequency domain feature of the input is x freq After each layer operation, the output y is obtained freq , that is, y freq=GAP(ReLU(BN(Attention(DenseConv(Conv(x freq ))))) , where Conv() represents the convolution operation function, DenseConv() represents the dense convolution block operation function, Attention() represents the attention block operation function, ReLU() represents the linear rectification activation function, and GAP() represents the global average pooling function. Basic features are extracted through convolution, and dense convolution blocks extract deeper features. The attention mechanism is then used to dynamically enhance key information and suppress noise, thereby significantly improving the representation ability and generalization performance of the model. In addition, the attention mechanism can calculate weights through global or local context, establish associations between long-distance features, and make up for the defect that convolution operations can only capture local information.

[0021] S3.2. Construct a time-domain stream neural network. The network is connected from input to output in the following order: 1 convolution layer, 1 dense convolution block, 1 attention block, 1 batch normalization layer, 1 ReLU layer and 1 global average pooling layer. Suppose the input time domain feature is x time After each layer operation, the output y is obtained time , that is, y time =GAP(ReLU(BN(Attention(DenseConv(Conv(x time )))));

[0022] S3.3. Construct a CNN-GRU neural network for blood oxygen saturation. The network structure is as follows: from the input layer to the output layer, it is connected in sequence as follows: 3 convolution blocks and 1 GRU layer. The structure of the convolution block is connected in sequence as follows: 1 one-dimensional convolution layer, 1 ReLU layer, 1 BN layer, and 1 maximum pooling layer. Suppose the input blood oxygen data is x spo2 , the output is y spo2 , that is, y spo2 =GRU(ConvBlock3(ConvBlock2(ConvBlock1(x spo2 )))),where ConvBlock i () represents the i-th convolution block operation function, i = 1, 2, 3, GRU() represents the gated recurrent unit layer operation function. Compared with EEG data, which has multiple channels, blood oxygen saturation data belongs to single-channel data, so a relatively simple CNN+GRU hybrid architecture is used. Local features, such as the morphological features of the blood oxygen waveform, are extracted through the convolution block, and then the GRU is used to model the temporal dependency, such as the blood oxygen change trend, to achieve multi-scale feature fusion. In addition, the convolution and pooling operations of the three convolution blocks are used to gradually compress the feature dimension and enhance the feature expression ability.

[0023] S3.4. Construct an ECG CNN-GRU neural network. The network structure is as follows: from input to output, it is connected in sequence as follows: 3 convolution blocks and 1 GRU layer. The structure of the convolution block is connected in sequence as follows: 1 one-dimensional convolution layer, 1 ReLU layer, 1 BN layer, and 1 maximum pooling layer. Suppose the input ECG data is x ecg , the output is y ecg , that is, y ecg =GRU(ConvBlock3(ConvBlock2(ConvBlock1(x ecg )))), ECG data and blood oxygen saturation data are both single-channel data, so the same architecture is used;

[0024] S3.5, the output feature vector y of the frequency domain flow neural network, time domain flow neural network, blood oxygen saturation CNN-GRU neural network, and electrocardiogram CNN-GRU neural network freq 、y time 、y spo2 、y ecg Splicing is performed on the feature dimension to obtain the spliced ​​feature vector y concat =[y freq ;y time ;y spo2 ;y ecg ], and then y concat Input to the MLP feature fusion layer, where the MLP feature fusion layer contains 4 hidden layers, and the number of neurons in each layer is 512, 256, 128, and 64 respectively. Let the input of the j=1, 2, 3, and 4 hidden layers be q j , the output is h j , then h j =ReLU(Dropout(FC j (q j ))) Among them FC j () represents the fully connected operation function of the jth fully connected layer. It uses multi-layer MLP to gradually reduce the dimension and explore the nonlinear relationship between features. It realizes feature abstraction from low-order to high-order through a hierarchical structure. It introduces Dropout to suppress overfitting and ReLU activation function to enhance nonlinear expression ability, so that the model can learn complex inter-modal dependencies.

[0025] The number of output neurons in the MLP feature fusion layer is the same as the number of pain classification categories C. When C = 2, the sigmoid activation function is used to convert the output into a probability distribution, that is, S = σ(h4); when C>2, the softmax activation function is used, that is, S = softmax(h4), and S is used to represent the probability of the sample belonging to each pain category.

[0026] Furthermore, the Dense convolution block adopts a densely connected architecture, consisting of L convolutional layers connected sequentially. Each convolutional layer receives the output of all previous layers, and through dense connections, a "feedforward + feedback" feature flow is achieved. The shallow basic features and deep abstract features are integrated, which improves feature diversity and feature reuse rate, effectively alleviating overfitting.

[0027] Assume the initial input feature tensor is Where B represents the batch size, H, W, and D represent the height, width, and depth of the feature respectively, and C0 represents the initial number of channels.

[0028] For the l=1,2,…,L convolutional layer, its input feature tensor X l It is the concatenation of the output feature tensors of all previous layers in the channel dimension, that is: Where k represents the number of channels output by each convolutional layer, and [·;·] represents the concatenation operation in the channel dimension.

[0029] The specific operation of the lth convolutional layer is as follows:

[0030] First, perform two three-dimensional convolution operations to extract features:

[0031] The first convolution uses a convolution kernel with a shape of (3,3,1) and the number of convolution kernels is k. The convolution kernel is The output of the first convolution is: Where Conv3D() represents a three-dimensional convolution operation;

[0032] The second convolution uses a convolution kernel with a shape of (1,1,3) and the number of convolution kernels is also k. The convolution kernel is The output of the second convolution is: Using two 3D convolutions has fewer parameters than using one 3D convolution, which makes training easier and speeds up operation.

[0033] Finally, the input X of the convolutional layer l And the result of convolution Splicing along the channel axis to get the output of the convolution layer

[0034] After L layers of convolutional layers, the final output of the Dense convolution block is X L+1 .

[0035] Furthermore, the data input by the Attention block undergoes channel mean calculation, spatial attention calculation, and deep attention calculation in sequence.

[0036] Assume the initial input feature tensor is Where B1 is the batch size, H1, W1, D1 are the height, width and depth of the feature respectively, and C1 is the initial number of channels;

[0037] First, for X in Perform channel mean calculation, the result is X mean , reflecting the overall activity level of each channel. Then, deep attention calculation and spatial attention calculation are performed in parallel:

[0038] Deep attention calculation: for X mean , use an average pooling layer to pool the depth dimension, set the pooling size to [1,1,5], set the average pooling operation to AvgPool(), and get the tensor X after pooling pool =AvgPool(X mean ), then X pool Flattened to a one-dimensional vector X through the Flatten layer flat , and then X flat Input the fully connected layer to get the result X fc , for X fc Apply the sigmoid activation function σ() to get X sig =σ(X fc ), and finally perform a reshape operation to convert X sig Reshape into deep attention weights

[0039] Spatial attention calculation: For X mean , use the average pooling layer to pool its height and width dimensions. The pooling size of the average pooling layer is [H, W, 1]. After the average pooling operation, the tensor X is obtained. pool_s =AvgPool(X mean ), then it is flattened by the Flatten layer, processed by the fully connected layer, activated by the sigmoid function, and reshaped to obtain the spatial attention weight Strengthening the characteristic expression of pain-related brain regions;

[0040] The attention calculation in the above steps is divided into two parts: deep attention and spatial attention. Deep attention focuses on the key frequency ranges and important sampling points of the EEG data, allowing the model to better capture dynamic changes in frequency and time, while spatial attention focuses on important brain regions, allowing the model to better capture differences between brain regions.

[0041] Then, the initial input feature tensor is X in First, multiply the element by element with the deep attention empty weight d to obtain the feature C after deep attention weighting D =X0⊙d, where ⊙ represents element-by-element multiplication;

[0042] Finally, C D Multiply element-wise with the spatial attention weight s to get the output C S =C D ⊙s, where Through two feature weightings, adaptive enhancement of the joint features of frequency band brain areas and time brain areas is achieved, while noise is suppressed.

[0043] Furthermore, the global average pooling layer performs a global average pooling operation on the input data, compressing the height, width, and depth dimensions of the data to 1. The compressed feature vector retains the global statistical information of each channel, avoiding the parameter redundancy problem of the fully connected layer, while significantly reducing the computational cost, speeding up the model training and operation, and improving the generalization ability of the model.

[0044] Let the global average pooling operation function be GAP(), and let the input tensor be Where B2 is the batch size, H2, W2, D2 are the height, width and depth of the feature respectively, and C2 is the initial number of channels; cin Perform global average pooling operation to obtain the output of the global average pooling layer It can be expressed as: X out =GAP(X cin ).

[0045] Furthermore, the process of step S4 is as follows:

[0046] S4.1. Using stratified sampling, divide the multimodal data into training and test sets according to the pain classification labels. Ensure that the proportion of each pain category in the training and test sets is consistent, so that the model can evenly learn the characteristics of different pain levels, which helps to solve the problem of data imbalance. The EEG data in the training set data is subjected to feature extraction steps to obtain time domain features and frequency domain features, which are input into the multimodal model together with the ECG and SPO2 data.

[0047] S4.2. Use an early stopping strategy during training. Evaluate the model performance on the validation set after a certain number of training rounds. Stop training if the loss value on the validation set does not decrease for i consecutive rounds. This prevents the model from overfitting to the noise and special patterns of the training data, prevents the model from overfitting, and speeds up the training of the model.

[0048] Compared with the prior art, the present invention has the following advantages and beneficial effects:

[0049] 1. Electrode mapping and matrix construction enhance the structured expression of EEG features: By defining a 7×7 two-dimensional matrix to spatially map 32 EEG electrodes, the spatial distribution of the EEG signals is converted into structured matrix data. This mapping method simulates the physiological topology of the cerebral cortex, concentrating closely located electrodes in adjacent positions in the matrix. Compared to directly using raw EEG signals, this structured data enables the dense convolution blocks and three-dimensional convolution layers in the frequency-domain and time-domain streaming neural networks to more efficiently extract spatial features, thereby enhancing the ability to express pain-related EEG features.

[0050] 2. Collaborative extraction of time domain and frequency domain features to comprehensively characterize the pain characteristics of EEG signals: The time domain features and frequency domain features of EEG are constructed separately to achieve a multi-dimensional characterization of EEG signals. The time domain features retain the dynamic characteristics of EEG signals changing over time; the frequency domain features focus on the association between different frequency components and pain. The two features complement each other and comprehensively characterize the pain-related EEG signal characteristics from the two perspectives of temporal dynamic changes and frequency component characteristics. Compared with a single time domain or frequency domain feature analysis, it can more completely present the changing patterns of EEG signals under pain conditions, thereby improving the accuracy and reliability of pain recognition.

[0051] 3. Multi-dimensional feature extraction enhances information representation: The constructed multimodal network model consists of four parallel branches: a frequency-domain streaming neural network, a time-domain streaming neural network, a blood oxygen saturation CNN-GRU neural network, and an electrocardiogram (ECG) CNN-GRU neural network. The frequency-domain streaming neural network and the time-domain streaming neural network, through a combination of dense convolution blocks and an attention block, perform multi-scale feature extraction and key feature enhancement on the signal, leveraging the synergistic effect of three-dimensional convolution and attention mechanisms. In the dense convolution block, each layer receives features from all previous layers and extracts features through a combination of two three-dimensional convolution kernels, achieving feature reuse and information accumulation. The attention block adaptively focuses on key regions and time sequences through spatial and temporal attention calculations, enhancing the representation of effective information. The blood oxygen saturation and ECG CNN-GRU neural networks utilize CNN convolution blocks to extract local features, and then use GRU to capture the temporal dependencies of sequential data. This effectively processes temporal trends in blood oxygen and ECG data and mines deep features. Compared with a single modality, this multi-dimensional feature extraction method can obtain more comprehensive information from multiple aspects such as frequency domain, time domain, and physiological indicators, and construct a richer and more accurate pain feature representation.

[0052] 4. Dynamic feature fusion improves recognition accuracy: The outputs of the four neural networks are spliced ​​and fused through the MLP feature fusion layer, effectively integrating the complementary information of different modalities. This allows the signal features captured by the frequency domain and time domain streaming neural networks to be combined with the physiological status characteristics reflected by blood oxygen saturation and electrocardiogram data, avoiding misjudgments caused by information limitations of a single modality. At the same time, it explores the potential physiological correlations between the data, which not only helps to improve the accuracy of pain recognition, but also provides a new perspective for medical research, helping doctors to better understand the intrinsic connection between pain and physiological signals, and promoting the development of postoperative pain management. BRIEF DESCRIPTION OF THE DRAWINGS

[0053] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0054] Figure 1 This is a flow chart of a multimodal postoperative pain recognition method for children disclosed in the present invention;

[0055] Figure 2 is a schematic diagram of the electrode positions in a 7×7 grid in the present invention;

[0056] Figure 3 It is an architectural diagram of the multimodal fusion model constructed by the present invention. DETAILED DESCRIPTION

[0057] In order to enable those skilled in the art to better understand the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments in the present invention, all other embodiments obtained by those skilled in the art without creative work are within the scope of protection of the present invention.

[0058] References to "embodiments" in this application mean that a particular feature, structure, or characteristic described in connection with the embodiment may be included in at least one embodiment of the application. The appearance of this phrase in various places in the specification does not necessarily refer to the same embodiment, nor does it constitute an independent or alternative embodiment that is mutually exclusive of other embodiments. It is understood, both explicitly and implicitly, by those skilled in the art that the embodiments described in this application may be combined with other embodiments.

[0059] Example 1

[0060] Figure 1A flowchart of a multimodal method for identifying pain in children after surgery is disclosed. This embodiment specifically discloses a multimodal method for identifying pain in children after surgery, comprising the following steps:

[0061] S1. Data collection steps: Collect EEG data, ECG data, and blood oxygen saturation data of children after surgery;

[0062] The specific implementation process of step S1 is as follows:

[0063] Forty children aged 1-10 years were enrolled in the Guangzhou Women and Children's Medical Center. EEG, ECG, and oxygen saturation data were collected simultaneously for 120 seconds using a 500Hz sampling rate electrocardiogram (ECG) device, a 60Hz sampling rate oxygen saturation (O2) monitor, and a BrainVisionactiCHamp Plus 32-channel EEG signal acquisition system with a sampling rate of 64Hz. Medical staff assigned each participant a pain rating based on a clinical pain assessment scale (0-10). Data were segmented into 4-second segments, resulting in a total of 1200 samples.

[0064] S2, feature extraction step: using the collected EEG data to construct a two-dimensional frame to obtain the time domain features of the EEG; calculating the differential entropy of multiple frequency intervals of the collected EEG data, using the obtained differential entropy to construct a two-dimensional frame to obtain the frequency features of the EEG;

[0065] The specific implementation process of step S2 is as follows:

[0066] S2.1. Define a 7×7 two-dimensional matrix. The positions of the 32 electrodes collected by EEG are mapped to the two-dimensional matrix as follows: the first row, 3-5 columns are the voltage values ​​of the 01, 0z, and 02 channels, respectively; the second row, 2-6 columns are the voltage values ​​of the P7, P3, Pz, P4, and P8 channels, respectively; the third row, 1, 2, 3, 5, 6, and 7 columns are the voltage values ​​of the TP9, CP5, CP1, CP2, CP6, and TP10 channels, respectively; the fourth row, 2, 3, 5, and 6 columns are the voltage values ​​of the T7, C3, C4, and T8 channels, respectively; the fifth row, 1, 2, 3, 5, 6, and 7 columns are the voltage values ​​of the FT9, FC5, FC1, FC2, FC6, and FT10 channels, respectively; the sixth row, 2-6 columns are the voltage values ​​of the F7, F3, Fz, F4, and F8 channels, respectively; the seventh row, 2-4 columns are the voltage values ​​of the FP1, Fpz, and FP2 channels, respectively.

[0067] S2.2, EEG time domain feature extraction: For each sample, according to the above mapping relationship (such as Figure 2As shown in the figure, the voltage values ​​of 32 channels at each sampling point of the sample are filled into the matrix to form a two-dimensional frame. 0.0 is filled in the places where there is no corresponding electrode in the two-dimensional matrix. The minimum and maximum normalization method is used for the two-dimensional frame to map the element value to the range of 0 to 1. Then, the matrix is ​​expanded from 7×7 to 14×14 by the bicubic interpolation method. The two-dimensional frames of each sampling point are arranged into a three-dimensional matrix to obtain the time domain characteristics of the EEG.

[0068] S2.3. EEG frequency domain feature extraction: Define the five EEG frequency intervals related to pain as 1-4Hz, 4-8Hz, 8-13Hz, 13-30Hz, and 30-45Hz. For each EEG channel signal, use the Welch method to calculate the power spectral density and convert it into a frequency domain signal through fast Fourier transform. Calculate the area under the power spectral density curve in the specified frequency band, and then use the formula DE = 0.5×log(2×π×e×P) to obtain the differential entropy features of each channel in different frequency intervals, where DE represents differential entropy, P represents probability density function, and e represents a natural constant. Fill the differential entropy features into a 7×7 two-dimensional matrix according to the electrode mapping rules to form a two-dimensional frame. Use the bicubic interpolation method to interpolate the two-dimensional frame, expand the two-dimensional matrix from 7×7 to m×m, and arrange the two-dimensional frames of each feature interval into a three-dimensional matrix to obtain the frequency domain features of EEG.

[0069] S3. Steps for building a multimodal network model: The multimodal network model includes four parallel connections: a frequency domain flow neural network, a time domain flow neural network, a blood oxygen saturation CNN-GRU neural network, an electrocardiogram CNN-GRU neural network, and an MLP feature fusion layer. The outputs of the four neural networks are spliced ​​and then passed through the MLP feature fusion layer for pain classification. The model architecture is shown in the figure below. Figure 3 As shown;

[0070] The specific implementation process of step S3 is as follows:

[0071] S3.1. Build a frequency-domain streaming neural network, which consists of the following connections from input to output: 1 convolutional layer, 1 dense convolution block, 1 attention block, 1 batch normalization layer, 1 ReLU layer, and 1 global average pooling layer.

[0072] S3.1.1. The Dense convolution block uses a densely connected architecture consisting of four sequentially connected convolutional layers. Each convolutional layer receives the output of all previous layers. Each layer contains two convolution operations. The first convolution uses a convolution kernel of shape (3,3,1) and the number of convolution kernels is 12. The second convolution uses a convolution kernel of shape (1,1,3) and the number of convolution kernels is also 12.

[0073] S3.1.2, In the Attention block, first calculate the channel mean of the data to obtain the result X mean , and then perform parallel spatial attention calculation and deep attention calculation;

[0074] Deep attention calculation, first use an average pooling layer to X mean The depth dimension is pooled with the pooling size set to [1, 1, 5]. Then, it is flattened by the Flatten layer, processed by the fully connected layer, activated by the sigmoid function, and reshaped to obtain the deep attention weight d.

[0075] Spatial attention calculation, first use an average pooling layer to X mean Pooling is performed in height and width dimensions, and the average pooling layer pooling size is [H, W, 1], where H is X mean The value of the height dimension, W is X mean The value of the width dimension is then flattened by the Flatten layer, processed by the fully connected layer, activated by the sigmoid function, and reshaped to obtain the spatial attention weight s;

[0076] The initial input data is first multiplied element-by-element by the deep attention weight d to obtain the feature weighted by the deep attention, and the feature is then multiplied element-by-element by the spatial attention weight s to obtain the output;

[0077] S3.1.3. The global average pooling layer performs a global average pooling operation on the input data, compressing the height, width, and depth dimensions of the data to 1.

[0078] S3.2. Build a time-domain streaming neural network. The network is connected from input to output in the following order: 1 convolutional layer, 1 dense convolution block, 1 attention block, 1 batch normalization layer, 1 ReLU layer, and 1 global average pooling layer. The dense convolution block, attention block, and global average pooling layer are consistent with the above settings.

[0079] S3.3. Construct a blood oxygen saturation CNN-GRU neural network with the following structure: from the input layer to the output layer, it is connected in sequence: three convolutional blocks and one GRU layer. The convolutional block structure is connected in sequence: one one-dimensional convolutional layer, one ReLU layer, one BN layer, and one maximum pooling layer.

[0080] S3.4. Construct an ECG CNN-GRU neural network. The network structure is as follows: from input to output, it is connected in sequence: 3 convolution blocks and 1 GRU layer. The structure of the convolution block is connected in sequence: 1 one-dimensional convolution layer, 1 ReLU layer, 1 BN layer, and 1 maximum pooling layer.

[0081] S3.5. Concatenate the output feature vectors of the frequency domain flow neural network, time domain flow neural network, blood oxygen saturation CNN-GRU neural network, and electrocardiogram CNN-GRU neural network in the feature dimension to obtain the concatenated feature vector y concat , and then y concat The input is sent to the MLP feature fusion layer, where the MLP feature fusion layer contains 4 hidden layers, and the number of neurons in each layer is 512, 256, 128, and 64 respectively. The number of output neurons in the MLP feature fusion layer and the number of pain classification categories are set to 2, and the sigmoid activation function is used to output the probability of the sample belonging to each pain category.

[0082] S4. Model training step: The time domain features and frequency domain features of the training samples are obtained through the feature extraction step, and the time domain features and frequency domain features of the EEG data are used as input together with the ECG data and blood oxygen saturation data to train the multimodal network model.

[0083] The specific implementation process of step S4 is as follows:

[0084] S4.1. Using a stratified sampling method, the collected multimodal data were divided into a training set and a test set in a ratio of 7:3. After feature extraction, the EEG data of the training set were input into the multimodal model together with the ECG data and SPO2 data.

[0085] S4.2. During the training process, an early stopping strategy is adopted, and the model performance is evaluated on the validation set every 5 rounds. If the loss value on the validation set does not decrease for 3 consecutive rounds, the training is stopped.

[0086] S5. Pain recognition step for the sample to be tested: obtain time domain features and frequency domain features from the EEG data of the sample to be tested through the feature extraction step, input the time domain features and frequency domain features of the EEG data of the sample to be tested, together with the ECG and blood oxygen saturation data, into the trained multimodal network model to obtain the pain recognition result of the sample to be tested.

[0087] The specific implementation process of step S5 is as follows:

[0088] Ten additional postoperative children were selected as test samples. Four seconds of EEG, ECG, and blood oxygen saturation data were collected from these children. Following the aforementioned feature extraction steps, EEG time and frequency domain features were derived. These features, along with the ECG and blood oxygen saturation data, were input into the trained model to generate pain recognition results for each test sample. The pain level assigned by medical staff based on a clinical pain assessment scale was used as the standard, with samples above level 3 considered painful. The model achieved an 80% accuracy rate in identifying pain.

[0089] Example 2

[0090] Based on the multimodal postoperative pain recognition method for children disclosed in the above embodiment 1, this embodiment specifically discloses another multimodal postoperative pain recognition method for children, including the following steps:

[0091] S1. Data collection step: refer to the data collection step in Example 1.

[0092] S2. Feature extraction step: construct a two-dimensional frame using the collected EEG data to obtain the time domain features of the EEG; calculate the differential entropy of multiple frequency intervals for the collected EEG data, use the obtained differential entropy to construct a two-dimensional frame, and obtain the frequency features of the EEG.

[0093] The specific implementation process of step S2 is as follows:

[0094] S2.1. EEG time domain feature extraction: Refer to step S2.1 in Example 1 to form a two-dimensional frame and normalize it. Then, use the bicubic interpolation method to expand the matrix from 7×7 to 32×32. Arrange the two-dimensional frames of each sampling point into a three-dimensional matrix to obtain the EEG time domain features:

[0095] S2.2. EEG frequency domain feature extraction: Refer to step S2.2 in Example 1 to calculate the differential entropy feature, map it into a two-dimensional frame, and normalize it. Bicubic interpolation is used to expand the matrix from 7×7 to 32×32, and the two-dimensional frames of each feature interval are arranged into a three-dimensional matrix to obtain the EEG frequency domain features:

[0096] S3. Multimodal network model construction step: Referring to the multimodal network model construction step in Example 1, the number of sequentially connected convolutional layers of the dense convolution blocks in the frequency domain stream neural network and the time domain stream neural network is 6, the number of output channels of each convolutional layer is 24, the number of pain classification categories is set to two categories: no pain and pain, and the output layer of the MLP feature fusion layer adopts the sigmoid activation function;

[0097] S4. Model training step: refer to the model training step in Example 1.

[0098] S5. Select another 10 postoperative children as test samples, collect their EEG, ECG and blood oxygen saturation data for 4 seconds, and obtain EEG time domain and frequency domain features according to the above feature extraction steps. Input the trained model together with the ECG and blood oxygen saturation data to obtain the pain recognition results of each test sample. The pain level of each research subject given by medical staff based on the clinical pain assessment scale is used as the standard. Samples with a level greater than 3 can be considered as pain samples. The accuracy of the model recognition results reaches 90%. In the recognition process, the model can quickly output results to meet the needs of clinical real-time monitoring.

[0099] The technical features of the above embodiments can be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0100] The above embodiments are preferred implementation modes of the present invention, but the implementation modes of the present invention are not limited to the above embodiments. Any other changes, modifications, substitutions, combinations, and simplifications that do not deviate from the spirit and principles of the present invention should be considered as equivalent replacement methods and are included in the scope of protection of the present invention.

Claims

1. A multimodal method for identifying pain in children after surgery, characterized in that: The method for identifying pain in children after surgery comprises the following steps: S1. Data collection steps: Collect EEG data, ECG data, and blood oxygen saturation data of children after surgery; S2, feature extraction step: using the collected EEG data to construct a two-dimensional frame to obtain the time domain features of the EEG; calculating the differential entropy of multiple frequency intervals of the collected EEG data, using the obtained differential entropy to construct a two-dimensional frame to obtain the frequency features of the EEG; S3. Multimodal network model construction steps: The multimodal network model includes four parallel connections: a frequency domain streaming neural network, a time domain streaming neural network, a blood oxygen saturation CNN-GRU neural network, an electrocardiogram CNN-GRU neural network, and an MLP feature fusion layer. The outputs of the four neural networks are concatenated and passed through the MLP feature fusion layer for pain classification. S4. Model training step: The time domain features and frequency domain features of the training samples are obtained through the feature extraction step, and the time domain features and frequency domain features of the EEG data are used as input together with the ECG data and blood oxygen saturation data to train the multimodal network model. S5. Pain recognition step for the sample to be tested: obtain time domain features and frequency domain features from the EEG data of the sample to be tested through the feature extraction step, input the time domain features and frequency domain features of the EEG data of the sample to be tested, together with the ECG and blood oxygen saturation data, into the trained multimodal network model to obtain the pain recognition result of the sample to be tested.

2. A multimodal postoperative pain recognition method for children according to claim 1, characterized in that: The process of step S1 is as follows: At the same time, we began to collect EEG data, ECG data, and blood oxygen saturation data of children after surgery for the same duration t. A 32-channel EEG data acquisition device was used for EEG data collection, and the electrode positions were based on the standard electrode placement method 10-20 system electrode placement method specified by the International Electroencephalography Society. Hereinafter, EEG is referred to as EEG, ECG as ECG, and blood oxygen saturation as SPO2. Medical staff gave each subject's pain level as a data label based on the clinical pain assessment scale during data collection.

3. A multimodal postoperative pain recognition method for children according to claim 1, characterized in that: The process of step S2 is as follows: S2.

1. Define a 7×7 two-dimensional matrix. The positions of the 32 electrodes collected by EEG are mapped to the two-dimensional matrix as follows: the first row, 3-5 columns are the voltage values ​​of the 01, 0z, and 02 channels, respectively; the second row, 2-6 columns are the voltage values ​​of the P7, P3, Pz, P4, and P8 channels, respectively; the third row, 1, 2, 3, 5, 6, and 7 columns are the voltage values ​​of the TP9, CP5, CP1, CP2, CP6, and TP10 channels, respectively; the fourth row, 2, 3, 5, and 6 columns are the voltage values ​​of the T7, C3, C4, and T8 channels, respectively; the fifth row, 1, 2, 3, 5, 6, and 7 columns are the voltage values ​​of the FT9, FC5, FC1, FC2, FC6, and FT10 channels, respectively; the sixth row, 2-6 columns are the voltage values ​​of the F7, F3, Fz, F4, and F8 channels, respectively; the seventh row, 2-4 columns are the voltage values ​​of the FP1, Fpz, and FP2 channels, respectively. S2.

2. Fill the 32 channel voltage values ​​at each sampling point of the EEG data into a two-dimensional matrix according to the above mapping relationship to form a two-dimensional frame. Fill the positions without corresponding electrodes in the two-dimensional matrix with 0.

0. Apply the minimum and maximum normalization method to the two-dimensional frame, map each element value in the two-dimensional frame to the range of 0-1, and interpolate the two-dimensional frame using the bicubic interpolation method to expand the two-dimensional matrix from 7×7 to m×m. Arrange the two-dimensional frames obtained at each sampling point into a three-dimensional matrix to obtain the time domain characteristics of the EEG; S2.

3. Define five EEG frequency intervals related to pain: 1-4 Hz, 4-8 Hz, 8-13 Hz, 13-30 Hz, and 30-45 Hz. For each EEG channel signal of the EEG data, the power spectral density is first calculated. The time domain signal is converted into a frequency domain signal by fast Fourier transform to obtain the power spectral density value of each frequency point. Then, the area under the power spectral density curve in the specified frequency band is calculated according to the specified EEG frequency interval. Finally, the differential entropy characteristics of each frequency interval of the channel are obtained according to the formula DE = 0.5 × log (2 × π × e × P), where DE represents differential entropy, P represents probability density function, and e represents a natural constant. The differential entropy features of the specific frequency interval of each channel are filled into the two-dimensional matrix according to the above mapping relationship to form a two-dimensional frame. 0.0 is filled in the places without corresponding electrodes in the two-dimensional matrix. The two-dimensional frame is interpolated using the bicubic interpolation method to expand the two-dimensional matrix from 7×7 to m×m. The two-dimensional frames obtained at each sampling point are arranged into a three-dimensional matrix to obtain the frequency domain features of the EEG.

4. A multimodal postoperative pain recognition method for children according to claim 1, characterized in that: The process of step S3 is as follows: S3.

1. Construct a frequency domain streaming neural network. The network is connected from input to output in the following order: 1 convolution layer, 1 dense convolution block, 1 attention block, 1 batch normalization layer, 1 Relu layer and 1 global average pooling layer. Suppose the frequency domain feature of the input is x freq After each layer operation, the output y is obtained freq , that is, y freq =GAP(ReLU(BN(Attention(DenseConv(Conv(x freq ))))),Conv() represents the convolution operation function, DenseConv() represents the dense convolution block operation function, Attention() represents the attention block operation function, ReLU() represents the linear rectification activation function, and GAP() represents the global average pooling function; S3.

2. Construct a time-domain stream neural network. The network is connected from input to output in the following order: 1 convolution layer, 1 dense convolution block, 1 attention block, 1 batch normalization layer, 1 ReLU layer and 1 global average pooling layer. Suppose the input time domain feature is x time After each layer operation, the output y is obtained time , that is, y time =GAP(ReLU(BN(Attention(denseConv(Conv(x time ))))); S3.

3. Construct a CNN-GRU neural network for blood oxygen saturation. The network structure is as follows: from the input layer to the output layer, it is connected in sequence as follows: 3 convolution blocks and 1 GRU layer. The structure of the convolution block is connected in sequence as follows: 1 one-dimensional convolution layer, 1 ReLU layer, 1 BN layer, and 1 maximum pooling layer. Suppose the input blood oxygen data is x spo2 , the output is y spo2 , that is, y spo2 =GRU(ConvBlock3(ConvBlock2(ConvBlock1(x spo2 )))),where ConvBlock i () represents the i-th convolutional block operation function, i = 1, 2, 3, GRU() represents the gated recurrent unit layer operation function; S3.

4. Construct an ECG CNN-GRU neural network. The network structure is as follows: from input to output, it is connected in sequence as follows: 3 convolution blocks and 1 GRU layer. The structure of the convolution block is connected in sequence as follows: 1 one-dimensional convolution layer, 1 ReLU layer, 1 BN layer, and 1 maximum pooling layer. Suppose the input ECG data is x ecg , the output is y ecg , that is, y ecg =GRU(ConvBlock3(ConvBlock2(ConvBlock1(x ecg )))); S3.5, the output feature vector y of the frequency domain flow neural network, time domain flow neural network, blood oxygen saturation CNN-GRU neural network, and electrocardiogram CNN-GRU neural network freq 、y time 、y spo2 、y ecg Splicing is performed on the feature dimension to obtain the spliced ​​feature vector y concat =[y freq ;y time ;y spo2 ;y ecg ], and then y concat Input to the MLP feature fusion layer, where the MLP feature fusion layer contains 4 hidden layers, and the number of neurons in each layer is 512, 256, 128, and 64 respectively. Let the input of the j=1, 2, 3, and 4 hidden layers be q j , the output is h j , then h j =ReLU(Dropout(FC j (q j ))) Among them FC j () represents the fully connected operation function of the jth fully connected layer; The number of output neurons in the MLP feature fusion layer is the same as the number of pain classification categories C. When C = 2, the sigmoid activation function is used to convert the output into a probability distribution, that is, S = σ(h4); when C>2, the softmax activation function is used, that is, S = softmax(h4), and S is used to represent the probability of the sample belonging to each pain category.

5. A multimodal postoperative pain recognition method for children according to claim 4, characterized in that: The Dense convolution block adopts a densely connected architecture, which consists of L convolution layers connected sequentially, and each convolution layer receives the output of all previous layers; Assume the initial input feature tensor is Where B represents the batch size, H, W, and D represent the height, width, and depth of the feature respectively, and h0 represents the initial number of channels. For the l=1,2,…,L convolutional layer, its input feature tensor X l It is the concatenation of the output feature tensors of all previous layers in the channel dimension, that is: Where k represents the number of channels output by each convolutional layer, and [·;·] represents the concatenation operation in the channel dimension. The specific operation of the lth convolutional layer is as follows: First, perform two three-dimensional convolution operations to extract features: The first convolution uses a convolution kernel with a shape of (3,3,1) and the number of convolution kernels is k. The convolution kernel is The output of the first convolution is: Where Conv3D() represents a three-dimensional convolution operation; The second convolution uses a convolution kernel with a shape of (1,1,3) and the number of convolution kernels is also k. The convolution kernel is The output of the second convolution is: Finally, the input X of the convolutional layer l And the result of convolution Splicing along the channel axis to get the output of the convolution layer After L layers of convolutional layers, the final output of the Dense convolution block is X L+1 .

6. A multimodal postoperative pain recognition method for children according to claim 4, characterized in that: The data input by the Attention block is sequentially subjected to channel mean calculation, spatial attention calculation, and depth attention calculation. Assume the initial input feature tensor is Where B1 is the batch size, H1, W1, D1 are the height, width and depth of the feature respectively, and C1 is the initial number of channels; First, for X in Perform channel mean calculation, the result is X mean , then, perform deep attention calculation and spatial attention calculation in parallel: Deep attention calculation: for X mean , use an average pooling layer to pool the depth dimension, set the pooling size to [1,1,5], set the average pooling operation to AvgPool(), and get the tensor X after pooling pool =AvgPool(X mean ), then X pool Flattened to a one-dimensional vector X through the Flatten layer flat , and then X flat Input the fully connected layer to get the result X fc , for X fc Apply the sigmoid activation function σ() to get X sig =σ(X fc ), and finally perform a reshape operation to convert X sig Reshape into deep attention weights Spatial attention calculation: For X mean , use the average pooling layer to pool its height and width dimensions. The pooling size of the average pooling layer is [H, W, 1]. After the average pooling operation, the tensor X is obtained. pool_s =AvgPool(X mean ), then it is flattened by the Flatten layer, processed by the fully connected layer, activated by the sigmoid function, and reshaped to obtain the spatial attention weight Then, the initial input feature tensor is X in First, multiply the element by element with the deep attention empty weight d to obtain the feature C after deep attention weighting D =X0⊙d, where ⊙ represents element-by-element multiplication; Finally, C D Multiply element-wise with the spatial attention weight s to get the output C S =C D ⊙s, where 7. A multimodal postoperative pain recognition method for children according to claim 4, characterized in that: The global average pooling layer performs a global average pooling operation on the input data, compressing the height, width, and depth dimensions of the data to 1. Let the global average pooling operation function be GAP(), and let the input tensor be Where B2 is the batch size, H2, W2, D2 are the height, width and depth of the feature respectively, and C2 is the initial number of channels; cin Perform global average pooling operation to obtain the output of the global average pooling layer It can be expressed as: X out =GAP(X cin ).

8. A multimodal postoperative pain recognition method for children according to claim 1, characterized in that: The process of step S4 is as follows: S4.

1. Using a stratified sampling method, divide the multimodal data into training and test sets according to the pain classification labels. Extract the EEG data from the training set to obtain time-domain and frequency-domain features, which are then input into the multimodal model along with the ECG and SPO2 data. S4.

2. Use an early stopping strategy during training. Evaluate the model performance on the validation set after a certain number of training rounds. If the loss value on the validation set does not decrease for i consecutive rounds, stop training to prevent the model from overfitting.

Citation Information

Cited By

  • Electroencephalogram pain positioning device, training method and equipment

    CN121730846A

  • Electroencephalographic pain localization device, training method and apparatus

    CN121730846B