A heterogeneous graph convolution epilepsy classification model construction method, epilepsy monitoring device, equipment and medium

By constructing a heterogeneous map convolutional epilepsy classification model, fusing facial and motor characteristics, and introducing a modal coordinated learning mechanism, it solves the problem that children's epilepsy is difficult to identify and monitor in complex environments, and achieves high-accurate epilepsy monitoring.

CN119516284BActive Publication Date: 2025-05-06CAPITAL NORMAL UNIVERSITY +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510065904.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-16
Publication Date
2025-05-06
Estimated Expiration
2045-01-16

AI Technical Summary

Technical Problem

The prior art is difficult to accurately identify and monitor epilepsy in childhood, especially in complex hospital environments, and the identification complexity and accuracy challenges exist for monitoring based on video data.

Method used

A heterogeneous graph convolutional epilepsy classification model is constructed, and the adaptive graph convolutional network model (F2AGCN network model) that fuses facial and motion, combined with a modal coordinated learning mechanism, extracts bone motion characteristics and facial features in video data to realize the classification and monitoring of epilepsy seizures.

Benefits of technology

It improves the accuracy of identification and monitoring efficiency of epilepsy, reduces the burden on experts, and provides effective guarantees for children's life safety.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119516284B_ABST
    Figure CN119516284B_ABST
Patent Text Reader

Abstract

The present application discloses a method for constructing a heterogeneous graph convolution epilepsy classification model, an epilepsy monitoring device, equipment and medium, and relates to the field of intelligent medical devices. The method first constructs a feature extraction module; then constructs an adaptive graph convolution network model that integrates face and motion, namely, an F2AGCN network model; then uses a modal coordination learning mechanism to train the F2AGCN network model, and obtains the trained F2AGCN network model as a heterogeneous graph convolution epilepsy classification model. The present application constructs an adaptive graph convolution network model that integrates facial and motion information for epilepsy classification, and introduces a modal coordination learning mechanism to force the model to learn feature representations that maintain semantic consistency between skeletal and facial modalities, ensuring that the fused features maintain semantic consistency between different modes, thereby improving the effectiveness of the integration, and the heterogeneous graph convolution epilepsy classification model constructed can directly realize the classification and monitoring of epilepsy based on video data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of medical devices, and in particular to a method for constructing a heterogeneous graph convolution epilepsy classification model, an epilepsy monitoring device, equipment and medium. Background Art

[0002] Epilepsy is a chronic disease in which neurons in the brain suddenly discharge abnormally, resulting in temporary brain dysfunction. Most epileptic seizures are accompanied by obvious behavioral changes, such as body tremors, rigidity, eye deviations, and facial distortions, but some changes are very subtle, such as brief absences or cessation of movement, and require careful observation to discern. These visual behavioral changes are key indicators for epilepsy doctors to evaluate epileptic seizures. For children, the complex visual background environment of the hospital coupled with its smaller target volume can affect the observation and identification of epileptic seizures. In addition, children's movements may be more active and diverse, making it increasingly complex and challenging to accurately extract and analyze epileptic seizure features from video clips. However, when faced with a large number of cases, experts often need to invest a lot of time and effort to identify epileptic seizures. Therefore, a more accurate epilepsy monitoring device is proposed, which aims to effectively reduce the burden on experts and assist in diagnosis and treatment. When a patient has an epileptic seizure, it can be discovered in time through the monitoring device and diagnosed in time, which can provide effective protection for the patient's life safety.

[0003] At present, classification methods based on deep learning have shown excellent performance in various fields. Neural networks can overcome these challenges by automatically learning more robust data distribution change characteristics from training data. In the field of epilepsy classification based on deep learning, it is mainly divided into epilepsy classification research based on electroencephalogram (EEG) data and epilepsy classification research based on visual data.

[0004] Epilepsy research based on deep learning mainly uses electroencephalogram (EEG) data, and EEG-based detection has made significant progress. Deep learning models such as Convolutional Neural Networks (CNN), Recurrent Neural Networks (RNN), and Long-Short Term Memory (LSTM) are widely used to process epilepsy EEG data. At the same time, there are many open-source epilepsy EEG datasets available for expert annotation. However, the use of EEG to diagnose epileptic seizures requires specialized equipment, which may limit the patient's activities. On the other hand, video data, as a non-invasive source of information, shows great potential in the detection of diseases such as Parkinson's disease, Alzheimer's disease, and spinal muscular atrophy. However, the visual manifestations of epileptic seizures are complex and varied, and it is difficult to distinguish the visual features of different seizure types. Some epilepsy types even lack specific visual features. In addition, clinical surveillance videos are shot under different conditions, and epileptic seizures usually occur suddenly and for a short time. At present, it is still a difficult point to monitor epileptic seizures based on video data. Summary of the invention

[0005] The purpose of this application is to provide a heterogeneous graph convolution epilepsy classification model construction method, epilepsy monitoring device, equipment and medium to achieve epileptic seizure monitoring based on video data, and overcome the defect of electroencephalogram monitoring that restricts patient activities.

[0006] To achieve the above objectives, this application provides the following solutions.

[0007] In a first aspect, the present application provides a method for constructing a heterogeneous graph convolution epilepsy classification model, comprising the following steps.

[0008] Construct a feature extraction module; the feature extraction module is used to extract the skeletal motion features and facial features of the target person in the video data.

[0009] Construct an F2AGCN network model; the F2AGCN network model is an adaptive graph convolutional network model that integrates face and motion; the F2AGCN network model includes: a first average pooling layer, a second average pooling layer, a first fully connected layer, a second fully connected layer, a fusion layer, and three FFDE modules connected in sequence; the FFDE module is used to perform deep feature extraction on skeletal motion features and facial features to obtain skeletal motion deep features and facial deep features; the first average pooling layer and the second average pooling layer are both connected to the last FFDE module of the three FFDE modules connected in sequence, the first fully connected layer is connected to the first average pooling layer, the second fully connected layer is connected to the second average pooling layer, and the first average pooling layer and the second average pooling layer are both connected to the fusion layer; the first average pooling layer and the first fully connected layer are used to perform epilepsy classification based on the skeletal motion deep features to obtain motion classification results, the second average pooling layer and the second fully connected layer are used to perform epilepsy classification based on the facial deep features to obtain facial classification results, and the fusion layer is used to fuse the motion classification results and the facial classification results to obtain a comprehensive classification result.

[0010] The F2AGCN network model is trained using a modality coordination learning mechanism to obtain a trained adaptive graph convolutional network model as a heterogeneous graph convolutional epilepsy classification model.

[0011] In a second aspect, the present application provides an epilepsy monitoring device, which includes: a camera, a server and a terminal.

[0012] The camera is connected to the server, the server is connected to the terminal, a heterogeneous graph convolutional epilepsy classification model is provided in the server, and the heterogeneous graph convolutional epilepsy classification model is obtained by using the above-mentioned heterogeneous graph convolutional epilepsy classification model construction method.

[0013] The camera is used to monitor the target person and obtain video data.

[0014] The server is used to perform epilepsy classification based on the video data using the heterogeneous graph convolution epilepsy classification model to obtain a comprehensive classification result.

[0015] The terminal is used to receive and display the comprehensive classification result.

[0016] In a third aspect, the present application provides a computer device, comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the above-mentioned heterogeneous graph convolution epilepsy classification model construction method.

[0017] In a fourth aspect, the present application provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the above-mentioned heterogeneous graph convolution epilepsy classification model construction method.

[0018] According to the specific embodiments provided in this application, this application has the following technical effects.

[0019] The present application provides a method for constructing a heterogeneous graph convolution epilepsy classification model, an epilepsy monitoring device, equipment and medium. First, a feature extraction module is constructed; then an adaptive graph convolution network model that integrates face and motion, namely, an F2AGCN network model, is constructed; then the F2AGCN network model is trained using a modality coordination learning mechanism to obtain the trained F2AGCN network model as a heterogeneous graph convolution epilepsy classification model. The present application constructs an adaptive graph convolution network model that integrates face and motion information for epilepsy classification, and introduces a modality coordination learning mechanism to force the model to learn feature representations that maintain semantic consistency between skeletal and facial modalities, ensuring that the fused features maintain semantic consistency between different modes, thereby improving the effectiveness of the integration. The heterogeneous graph convolution epilepsy classification model constructed can directly realize the classification and monitoring of epilepsy based on video data. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the drawings required for use in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0021] Figure 1 A schematic flowchart of a method for constructing a heterogeneous graph convolution epilepsy classification model provided in one embodiment of the present application.

[0022] Figure 2 A schematic diagram of the structure of an AGCN Block provided in one embodiment of the present application.

[0023] Figure 3 A schematic diagram of the structure of an adaptive graph convolutional network model that integrates face and motion provided in one embodiment of the present application.

[0024] Figure 4 A schematic diagram of the structure of a convolutional block attention module provided in one embodiment of the present application.

[0025] Figure 5 A bone distribution map provided for an embodiment of the present application.

[0026] Figure 6 This is a graph showing the accuracy of each algorithm training process provided in an embodiment of the present application.

[0027] Figure 7 This is a Loss curve diagram of each algorithm training process provided in an embodiment of the present application.

[0028] Figure 8 A schematic diagram of the structure of a computer device provided in one embodiment of the present application. DETAILED DESCRIPTION

[0029] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.

[0030] In order to make the above-mentioned objects, features and advantages of the present application more obvious and easy to understand, the present application is further described in detail below with reference to the accompanying drawings and specific implementation methods.

[0031] In an exemplary embodiment, Figure 1 As shown, a method for constructing a heterogeneous graph convolution epilepsy classification model is provided, including the following steps 101 to 103.

[0032] Step 101, constructing a feature extraction module; the feature extraction module is used to extract the skeletal motion features and facial features of the target person in the video data.

[0033] Step 102, constructing an F2AGCN network model; the F2AGCN network model is an adaptive graph convolutional network model that integrates face and motion; the F2AGCN network model includes: a first average pooling layer, a second average pooling layer, a first fully connected layer, a second fully connected layer, a fusion layer, and three FFDE modules connected in sequence; the FFDE module is used to perform deep feature extraction on skeletal motion features and facial features to obtain skeletal motion deep features and facial deep features; the first average pooling layer and the second average pooling layer are both connected to the last FFDE module of the three FFDE modules connected in sequence, the first fully connected layer is connected to the first average pooling layer, the second fully connected layer is connected to the second average pooling layer, and the first average pooling layer and the second average pooling layer are both connected to the fusion layer; the first average pooling layer and the first fully connected layer are used to perform epilepsy classification based on the skeletal motion deep features to obtain a motion classification result, the second average pooling layer and the second fully connected layer are used to perform epilepsy classification based on the facial deep features to obtain a facial classification result, and the fusion layer is used to fuse the motion classification result and the facial classification result to obtain a comprehensive classification result.

[0034] Step 103, using a modality coordination learning mechanism to train the F2AGCN network model, and obtaining a trained F2AGCN network model as a heterogeneous graph convolution epilepsy classification model.

[0035] The above steps 101 to 103 are implemented, an adaptive graph convolutional network model that fuses facial and motion information is used for epilepsy classification, and a modality coordination learning mechanism is introduced to force the model to learn feature representations that maintain semantic consistency between skeletal and facial modalities, ensuring that the fused features maintain semantic consistency between different modes, thereby improving the effectiveness of the integration. The heterogeneous graph convolutional epilepsy classification model constructed can directly realize the classification and monitoring of epilepsy based on video data.

[0036] In another exemplary embodiment, the video data in the above step 101 includes not only the activities of the target person (in the embodiment of the present application, the target person is an epileptic patient, hereinafter referred to as the patient), but also the activities of the patient's family members, resulting in a large amount of redundancy. In order to improve the accuracy of the algorithm and speed up the processing speed, it is crucial to minimize this redundancy. For the extraction of motion features, the use of skeleton point data can effectively represent the most motion features with the least information. By using skeleton point data, the patient's motion information can be extracted from the video without processing the entire video frame. In addition, the head coordinates of the skeleton point data can provide information about the patient's facial area, which is convenient for the extraction of facial features.

[0037] The feature extraction module in this application extracts motion and facial features by first inputting video data into a human pose estimation (HPE) algorithm to extract the patient's bone point information. The HPE algorithm uses a pre-trained human pose estimation model to perform human pose estimation.

[0038] At the same time, the algorithm uses the skeleton point information of the patient's head to locate the patient's face while passing the skeleton point information to the subsequent algorithm, thereby obtaining facial features. The face positioning process extracts all facial coordinate data from the skeleton point data corresponding to the video, which can be expressed as a set This is the positioning of the patient's facial area, and the face positioning algorithm formula is as follows.

[0039] (1)

[0040] (2)

[0041] In formula (1) and formula (2), and Respectively represent the cropping coordinates of the facial contour along the horizontal (left and right) and vertical (up and down) dimensions, , and are the horizontal coordinates of the first, second and i-th bone points on the target person’s face, i is the number of bone points on the target person’s face, , , They are the ordinates of the first, second, and i-th bone points on the target person’s face, respectively. and denote the minimum and maximum functions respectively, and They are the height and width of the video data respectively.

[0042] In the formula, use and The function finds the four coordinate points closest to the facial contour among the facial feature points to determine the facial area in the video. Considering that the input is a surveillance video from a hospital, and the hospital environment is often complex, the facial feature points obtained by the HPE algorithm may not be very accurate, and may even suffer from information loss due to partial occlusion and interference. Therefore, when determining the facial area range, the formula uses the MIN and MAX functions to obtain the maximum and minimum values ​​of the facial area range. Then the width and height are increased to expand the facial area to ensure the integrity of the patient's facial information.

[0043] After being processed by the face positioning algorithm, the obtained facial feature data and the corresponding bone point motion feature data are input into the F2AGCN network model together.

[0044] Before constructing the adaptive graph convolutional network model that integrates face and motion in the above steps, a model that can use visual data to classify epilepsy was studied. There are two main types of algorithms for epilepsy classification, one is a feature classification algorithm, and the other is a temporal classification algorithm.

[0045] 1. Feature classification algorithm.

[0046] When using visual data to classify epilepsy, although epileptic seizures do not originate from abnormal movements of the patient, most epileptic seizures manifest as abnormal features of the face, hands, body parts, or the entire body.

[0047] Quantitative analysis of early epileptic seizures has been done by attaching infrared reflective markers to key points on the body or using cameras with color and depth streams to assess patient motion. These methods suffer from poor video quality due to occlusion by bed sheets or clinical staff, differences in lighting and posture, or compression artifacts and out-of-focus details, making the algorithms unreliable.

[0048] In recent years, with the advancement of computer technology, the development of deep learning technology and the improvement of graphics card computing power, it has become possible to solve video classification problems using deep learning technology. Deep learning models have considerable potential in video recognition and classification, and can detect and identify signs or characteristics of certain diseases, including epilepsy.

[0049] Some studies, such as the algorithm proposed by Ahmedt-Aristizabal, extract motion features of each frame of the video through a convolutional neural network (CNN), then use a recurrent neural network to extract information in the aggregated feature segments, and finally perform classification through a fully connected layer.

[0050] This method has some problems. First, in order to reduce the amount of calculation, the algorithm cuts with the bed as the center after determining the position of the bed in the first frame, which has limitations. The displacement of the bed will seriously affect the accuracy of the algorithm. Second, the CNN network training is used to extract motion features on medical data sets. Medical data sets are generally small data sets, which leads to limitations in the extracted motion features.

[0051] 2. Time series classification algorithm.

[0052] With the development of deep learning technology, more and more time series classification models have been proposed, such as LSTM, Seq2Seq (Sequence to Sequence), Transformer (a model used to process sequence data in machine learning), etc. At the same time, there are also spatiotemporal sequence models that extract temporal and spatial domain features from sequences, such as ST-GCN (Spatial Temporal Graph Convolutional Networks), AGCN (Adaptive Graph Convolutional Networks), etc. These models have achieved good accuracy in a wide range of motion recognition and classification fields.

[0053] The skeleton graph used in ST-GCN is heuristically predefined and only represents the physical structure of the human body. Its graph convolution formula can be expressed as Equation (3).

[0054] (3)

[0055] In formula (3), For output, For input, is the kernel size in the spatial dimension. represents the adjacency matrix, Indicates a mask. When there is an element with value 0 in Ineffective. Therefore, it cannot guarantee optimality for motion recognition tasks. For example, the relationship between two hands is crucial for recognizing gestures such as “clapping” and “reading”.

[0056] In order to solve the problems in ST-GCN, Shi Lei et al. improved the graph convolution in AGCN, as shown in formula (4):

[0057] (4)

[0058] In formula (4), the adjacency matrix is ​​given by , , It consists of three parts. Represents the physical structure of the human body. is a parameterized optimization learned entirely from the training data The adjacency matrix not only shows the existence of connections between nodes, but also the strength of the connections between nodes. refers to a unique graph, learned for each sample, representing a unique configuration.

[0059] Lei Shi et al. designed AGCN Block (Adaptive Graph Convolutional Neural Networks Block, hereinafter referred to as AGCN Block or AGCN block or AGCN module) as the basic module of AGCN network by improving the adjacency matrix in graph convolution. Figure 2 As shown in Figure 2, the AGCN block consists of two modules: a graph convolution module and a temporal convolution module. In the graph convolution module, the adjacency matrix and its weights defined in equation (4) are used to extract spatial features. In the temporal convolution module, convolution kernels are used to extract temporal features.

[0060] The spatial and temporal convolutions are followed by a batch normalization (BN) layer and a ReLU (Rectified Linear Units) activation function. In addition, there is a Dropout layer between the two modules to enhance robustness, and a residual connection is used to maintain training stability when increasing the output dimension.

[0061] Furthermore, the present application also studies the Convolutional Block Attention Module (CBAM, hereinafter referred to as CBAM, CBAM block or CBAM module), which is an attention mechanism used in various computer vision tasks, such as image classification, object detection and semantic segmentation. In image classification tasks, CBAM helps the model pay more attention to the key areas of the image, thereby improving the classification accuracy. CBAM consists of two main parts: a channel attention module and a spatial attention module. These modules respectively pay attention to the channel dimension and the spatial dimension, thereby improving the model's ability to extract relevant image features.

[0062] The channel attention module adaptively selects and weights useful channel information in the feature map by considering the global context information of each channel, and calculates its importance weight. In contrast, the spatial attention module adaptively selects and weights useful spatial information in the feature map by considering the feature response of each spatial position, and calculates its importance weight.

[0063] CBAM integrates channel attention and spatial attention to generate the final attention feature map. First, the spatial attention module is applied to the input feature map to obtain the spatial attention feature map. Then, the channel attention feature map is element-wise multiplied with the original input feature map to weight each channel. Then, the spatial attention module is applied to the weighted feature map to generate the spatial attention feature map. Finally, the spatial attention feature map is element-wise multiplied with the weighted feature map to obtain the final attention feature map.

[0064] By adding CBAM, CNN can better model features and enhance its expressiveness and generalization capabilities. CBAM is compatible with various CNN architectures and shows strong performance in tasks such as image classification, object detection, and image segmentation.

[0065] This application proposes a temporal classification model that integrates facial and motion features, namely the F2AGCN network model, by analyzing the motion and facial features during epileptic seizures and various temporal classification algorithms in recent years. This model is improved on the basis of the AGCN model. The embodiment of this application uses skeleton data to locate the facial area in the video data through the coordinates of the patient's facial skeleton points to obtain facial data. The F2AGCN network model uses facial features and motion features as input, integrates and processes them, and completes the classification of epilepsy.

[0066] In another exemplary embodiment, Figure 3 (a) in the figure is the overall structure diagram of the adaptive graph convolutional network model that integrates face and motion, as shown in Figure 3As shown in (a), the input video is processed by the feature extraction module for human pose estimation (HPE) and face localization to obtain skeletal motion features and facial features. Then, these data are processed and deep feature extracted by three FFDE (Fusion Facial Downsampling and Feature Extraction Module) modules, where the three FFDE modules correspond to Figure 3 FFDE Block 1, FFDEBlock 2, and FFDE Block 3 in (a). The input and output information of the three FFDE modules are shown in Table 1, where the header represents the sequentially connected input sequence from the three FFDE modules. The first and second rows describe the input and output dimensions of the motion features, and the two comma-separated numbers in the table represent the time, channel, and vector length, respectively. The third and fourth rows represent the output and output dimensions of the facial features, and the three comma-separated numbers in the table represent the time, channel, and width and height of the facial features, respectively.

[0067] Table 1 Input and output of FFED module

[0068]

[0069] The FFDE module in the F2AGCN network model includes the F2AGCN block and the AGCN block. The F2AGCN block is obtained by improving the AGCN block. The improvement method is: the two-dimensional convolutional layer in the AGCN block is replaced by the convolutional block attention module.

[0070] Specifically, Figure 3 As shown in (a) in the figure, the FFDE module mainly consists of feature processing modules based on F2AGCN blocks, AGCN blocks and various convolutional layers. In each FFDE module, the skeletal motion features and facial features are simultaneously input into the F2AGCN block. In the F2AGCN block, the facial features and skeletal motion features extract semantic information in the spatial and channel dimensions, respectively, and then are fused into a heterogeneous graph structure for temporal feature extraction. The output heterogeneous graph structure separates the skeletal and facial features, and the motion features are continuously input into N AGCN blocks. In three similar FFDE modules, the values ​​of N are 3, 2 and 2, respectively. At the same time, the facial features are reconstructed from a one-dimensional feature vector into a two-dimensional feature matrix, which is then downsampled and input into the next FFDE module.

[0071] After three FFDE modules, the motion and facial features are input into the average pooling and fully connected layers respectively. Finally, the two branches are fused by element-wise multiplication of their respective class probabilities.

[0072] The F2AGCN block is introduced in detail below.

[0073] Figure 3 (b) in the figure is a schematic diagram of the structure of the F2AGCN block. Figure 3 As shown in (b) in Figure 2, the F2AGCN block is an improved version of the AGCN block. The skeleton motion features are processed by the GCN module to extract deeper motion features. The convolutional block attention module captures the expression features in each frame of the image while reducing the size of each channel.

[0074] like Figure 3 As shown in (b), the F2AGCN block includes: a GCN module, a convolutional block attention module, a connection module, a temporal convolution module and a post-processing module; the GCN module is used to perform deep feature extraction on skeletal motion features to obtain spatial skeletal motion deep features; the convolutional block attention module is used to perform deep feature extraction on facial features to obtain spatial facial depth features; the connection module is used to connect the spatial skeletal motion deep features and the spatial facial depth features to obtain a facial-motion heterogeneous graph structure; the temporal convolution module is used to perform temporal feature extraction on the facial-motion heterogeneous graph structure to obtain facial-motion deep features; the post-processing module is used to separate the facial-motion depth features to obtain skeletal motion depth features and facial depth features.

[0075] Despite the size reduction, facial features still have a higher parameter weight per channel in all FFDE modules compared to motion features. For example, in FFDE module 3, facial features are 25, while the vector length of each channel of the motion feature is only 25. Therefore, in the subsequent fusion process of facial features and motion features, the motion features are ignored due to the imbalance of information.

[0076] To address the above imbalance, this application proposes a "hierarchical attention fusion model". This model uses the attention layer in the graph convolution to enhance the relationship between skeleton points and introduces a convolutional block attention module after the convolution layer of the face branch, such as Figure 4 As shown, Figure 4 The channel attention module and spatial attention module in correspond to Figure 3(b) The channel and spatial attention modules of the face branch in the figure use learnable parameters in the channel and spatial attention modules to emphasize the key spatial and channel information in facial features, suppressing irrelevant details while retaining useful data. Facial features are first input into the channel attention module for maximum and average pooling in the channel dimension, then added after passing through the Shared MLP (shared multi-layer perceptron), and multiplied with the original input after passing through the Sigmoid activation function; in the spatial attention module, the shared multi-layer perceptron is replaced by the feature concatenation operation in the channel dimension, and the spatial relationship features are further extracted through the convolution layer. This method maintains a balanced information flow between motion and facial features, and effectively integrates the key features of sparse and dense data.

[0077] After the processed facial features pass through the convolutional layer and the convolutional block attention module, they are flattened from the original two-dimensional format to a one-dimensional sequence. After this conversion, the skeleton point features and the flattened facial features are converted to the same dimension and spatially connected to create a face-motion heterogeneous graph structure. This face-motion heterogeneous graph structure is jointly input into the TCN (temporal convolution) module to extract temporal features. After post-processing in the TCN module, the face-motion heterogeneous graph structure is separated again into separate facial and motion components, and then output separately.

[0078] In another exemplary embodiment, during the model training process, the present application provides a modality coordination learning mechanism that utilizes both the mean square error (MSE) and cross entropy loss functions, and uses hyperparameter control to balance the two losses. This method forces the model to learn feature representations that maintain semantic consistency between skeleton and facial morphology. This ensures that the fused features maintain semantic consistency between different modalities, thereby improving integration efficiency.

[0079] In the training of the algorithm model parameters, different loss functions are used in the face feature processing branch and the motion feature processing branch. The expression of the loss function is shown in formula (5).

[0080] (5)

[0081] in, is the total loss function, is the cross entropy loss function, is the MSE loss function, is the result of motion classification, is the face classification result, is a hyperparameter used to adjust the relative importance of the two branch losses. The contribution of different branches to the overall loss of the model can be controlled.

[0082] The cross entropy loss function is usually defined as the distance between the true probability distribution and the predicted probability distribution. This design enables the loss function to effectively penalize classification errors and promote the model to learn more accurate classification boundaries. Therefore, the model uses the cross entropy loss function in the action feature branch to distinguish the boundary between epileptic seizures and normal actions.

[0083] The MSE loss function has good mathematical properties in problems with continuous output space, and its gradient is relatively stable, which helps the model to converge to the optimal solution stably. Therefore, the MSE loss function is used in the facial feature branch to learn continuous facial seizure features, thereby assisting the action branch to obtain better classification results.

[0084] In order to illustrate the effect of the above model, it is verified in another exemplary embodiment.

[0085] First, we obtain the data set, introduce the characteristics of the experimental data used in the experiment, and the process of converting the experimental data into the data set. Then, we give the adaptive graph convolutional network model that integrates face and motion for ablation experiments, and compare the ablation results with other mainstream algorithms to prove the effectiveness of the adaptive graph convolutional network model that integrates face and motion. The experiment is conducted on a GPU server in the Baidu PaddlePaddle AI Studio environment, using the open source Baidu PaddlePaddle platform, framework version 2.2.2, and Python version 3.7.

[0086] The data studied in the examples of this application come from Beijing Children's Hospital affiliated to Capital Medical University. Epilepsy experts annotated the videos and obtained 38 seizure videos of 30 patients and corresponding label data such as video frame number, seizure time, and seizure type. The examples of this application use monitoring videos of epileptic seizures in patients in a clinical environment. The average duration of each video is about 2 minutes, and the frame rate is 25 frames per second. Due to the suddenness of epileptic seizures, the videos are all clinical videos, which include not only the activities of patients, but also the activities of patients' families and medical staff. The actions of different patients in different videos are also different.

[0087] Using the pre-trained human posture estimation model, the epilepsy video data is processed to extract the patient's skeleton information. The OpenPose toolbox is used to extract skeleton information from the video. OpenPose is a deep learning-based library used to accurately detect and estimate the key points and posture information of the human body from images or videos to obtain skeleton information. The skeleton distribution diagram is shown in Figure 5As shown, 0-24 are nose, neck, right shoulder, right elbow, right wrist, left shoulder, left elbow, left wrist, middle hip, right hip, right knee, right ankle, left hip, left knee, left ankle, right eye, left eye, right ear, left ear, left big toe, left little toe, left heel, right big toe, right little toe, right heel.

[0088] In epilepsy videos, since there are not only patients present but also family members and medical staff, the human pose estimation model extracts the skeleton point information of multiple individuals in the video.

[0089] In order to accurately select the patient's skeletal point information from multiple individuals in the frame output by OpenPose, the skeletal points of the patient in the first frame can be manually selected. Subsequently, the skeletal points of the patient in the next frame can be determined by calculating the straight-line distance between the skeletal points of each individual and the skeletal points of the current frame. This process can obtain all the skeletal point information of the patient in the entire video. In order to prevent OpenPose from being unable to identify the skeletal points of the patient in one frame, resulting in the selection of the skeletal points of another individual closest to the patient's position in the previous frame, a distance threshold can be used to determine whether the skeletal point information belongs to the patient.

[0090] In the process of extracting skeleton point information, due to the complexity of the monitoring video environment, some segments may occlude the patient's body, resulting in inaccurate or even lost skeleton point information. Through manual annotation, the skeleton point information in these segments can be corrected. Severely occluded segments can be excluded from the dataset to ensure the availability of the dataset data.

[0091] Through OpenPose processing and skeleton point information correction, the skeleton information of the patient in the video is extracted. The video includes two states: epileptic seizure and the patient's normal state. Therefore, based on the expert's annotation of the corresponding epileptic seizure time in the video, the skeleton output and the corresponding video clips are cropped to obtain epileptic seizure and normal data, and assigned binary classification labels.

[0092] Finally, the dataset is divided into training set, validation set and test set in a ratio of 7:2:1. The dataset length is 210 frames, including 105 frames of epilepsy data and 126 frames of normal data.

[0093] In the F2AGCN network, in addition to adding facial feature input, a convolutional block attention module (CBAM) and an improved MSE (mean square error) loss function are also integrated. Next, an ablation experiment will be conducted on the epilepsy dataset to demonstrate the enhanced effects of these improvements.

[0094] Table 2 Ablation experiment results

[0095]

[0096] As shown in Table 2, the first column in Table 2 indicates whether the CBAM module is used in the adaptive graph convolutional network model for fusion of face and motion, and the second column indicates whether the cross entropy loss or MSE is used in the face branch. The experiment evaluates four performance indicators: precision, recall, F1 score, and accuracy. The introduction of the convolutional attention module improves the precision and recall of the F2AGCN network model, but the F1 score decreases slightly.

[0097] In addition, combining the convolutional attention module and using the mean square error loss function in the face branch instead of the cross entropy loss function can comprehensively improve all four indicators. This proves the enhancement effect of the convolutional attention module and the mean square error loss function on the F2AGCN network.

[0098] In order to demonstrate the effect of this algorithm on facial feature fusion, it is compared with five time series algorithms: LSTM, Transformer, ST-GCN, AGCN, and CTR-GCN (Channel-wise Topology Refinement GraphConvolution) whose input is only skeleton sequence data. The epoch of model training is set to 100, and the Adam optimizer is used in the process. The learning rate is set to 0.001, the weight decay is set to 0.0001, and the loss function uses the cross entropy function. The accuracy curve and loss curve of each algorithm training process in the experiment are plotted, as shown in the figure. Figure 6 and Figure 7 shown.

[0099] Figure 6 and Figure 7 The accuracy and loss curves of each algorithm model in 0-100 rounds of training are shown. It can be seen that in 80-100 rounds of training, the fitting effect of the F2AGCN algorithm is slightly higher than that of other algorithms, which reflects that the addition of facial features has a certain improvement effect on the fitting of the algorithm. After 100 rounds of training, the experimental results are shown in Table 3.

[0100] Table 3 Comparative experimental results

[0101]

[0102] The data in Table 3 include the training models of various algorithms on the epilepsy classification training set and their classification results on the test set. The first five rows show the test results of five time series classification algorithms using only skeleton data. The sixth row shows the model trained only on facial data. Here, the algorithm uses ResNe18 to extract facial features, then uses LSTM for temporal feature extraction, and finally completes the classification. The seventh row uses an algorithm based on the model in the sixth row, adding a motion feature extraction branch. ResNet18+LSTM is used to classify facial data, and AGCN is used to classify skeleton data. The final results of the two branches are combined by element-by-element multiplication to produce the final result. The last row shows the classification performance of the AGCN algorithm proposed in this application on the test set.

[0103] In Table 3, the proposed adaptive graph convolutional network model that integrates face and motion is distinguished from other comparison algorithms by input data, and is evaluated using precision, recall, F1 score, and accuracy metrics to present its test results. The results show that both skeleton data and facial data can classify epileptic seizures, but facial data alone is less effective. The dual-stream algorithm ResNet18+LSTM+AGCN that uses both facial and skeleton data for classification outperforms the ResNet18+LSTM algorithm and the AGCN algorithm in all evaluation indicators, proving that combining facial features improves the accuracy of the epilepsy classification algorithm.

[0104] The F2AGCN network model in this application outperforms other algorithms in all four evaluation indicators, indicating that it has high accuracy in epileptic seizure classification. The experiment also verified the effectiveness of the heterogeneous graph structure proposed in this application in facial movement integration, modality coordination learning mechanism and hierarchical attention fusion mechanism.

[0105] Epilepsy patients face the risk of injury or even death during seizures. However, due to the suddenness of epileptic seizures, visual real-time monitoring poses many challenges. Recognition becomes more challenging for children who may show smaller targets, more active movements, and diverse changes. This application proposes a F2AGCN network model, which classifies epilepsy by analyzing the motion and facial features during epileptic seizures. On the basis of the AGCN block, an improved F2AGCN block is introduced to unify the skeleton and facial features into a facial motion heterogeneous graph structure, with skeleton points and facial pixels as nodes in the graph. By adding an attention layer to the graph convolution and integrating the CBAM attention module, F2AGCN improves the integration of sparse and dense data from motion and facial modalities. In addition, a modality coordination learning mechanism is introduced to maintain the semantic consistency between the two types of features using an enhanced loss function. This study uses a pediatric epileptic seizure dataset provided by Beijing Children's Hospital to create a classification dataset. The effectiveness and superiority of the F2AGCN algorithm are verified through ablation experiments on the dataset and comparative experiments with various algorithms. In addition, the study confirms that integrating facial features significantly improves the accuracy of the algorithm in epilepsy classification.

[0106] The heterogeneous graph convolution epilepsy classification model constructed in the embodiment of the present application has the following advantages.

[0107] 1. Detect epileptic seizure information based on epilepsy motion feature recognition. Considering the sparsity of skeleton point data and the density of facial image data, a heterogeneous graph structure is designed to unify skeleton points and facial pixels into nodes in the graph, each with different attributes and connectivity.

[0108] 2. Using the action recognition AGCN algorithm, an epilepsy classification algorithm F2AGCN that integrates facial features and skeletal motion features is introduced. A multi-scale attention mechanism is used to focus on relevant information at different spatial and semantic levels. A hierarchical attention fusion mechanism is proposed, which applies attention mechanisms at different abstract levels of skeletal structure and facial region, promoting the refined integration of sparse and dense data.

[0109] 3. A modality coordination learning mechanism is introduced by designing a specific loss function to force the model to learn feature representations that maintain semantic consistency between skeleton and facial modalities. This ensures that the fused features maintain semantic consistency between different modalities, thereby improving the effectiveness of the integration.

[0110] In an exemplary embodiment, the heterogeneous graph convolutional epilepsy classification model constructed by the above method embodiment is applied to an epilepsy monitoring device, which includes: a camera, a server and a terminal; the camera is connected to the server, the server is connected to the terminal, the server is provided with a heterogeneous graph convolutional epilepsy classification model, and the heterogeneous graph convolutional epilepsy classification model is obtained by the heterogeneous graph convolutional epilepsy classification model construction method of the above embodiment; the camera is used to monitor the target person and obtain video data; the server is used to classify epilepsy based on the video data using the heterogeneous graph convolutional epilepsy classification model to obtain a comprehensive classification result; the terminal is used to receive and display the comprehensive classification result.

[0111] In an exemplary embodiment, a computer device is provided. The computer device may be a server or a terminal. The internal structure diagram thereof may be as follows: Figure 8 As shown. The computer device includes a processor, a memory, an input / output interface (Input / Output, referred to as I / O) and a communication interface. Among them, the processor, the memory and the input / output interface are connected through a system bus, and the communication interface is connected to the system bus through the input / output interface. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The input / output interface of the computer device is used to exchange information between the processor and an external device. The communication interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, a method for constructing a heterogeneous graph convolution epilepsy classification model is implemented.

[0112] Those skilled in the art will understand that Figure 8 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine certain components, or have a different arrangement of components.

[0113] In an exemplary embodiment, a computer device is provided, including a memory and a processor, wherein a computer program is stored in the memory, and the processor implements the steps in the above-mentioned method embodiments when executing the computer program.

[0114] In an exemplary embodiment, a computer-readable storage medium is provided, storing a computer program, and when the computer program is executed by a processor, the steps in the above method embodiments are implemented.

[0115] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties.

[0116] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to the memory, database or other medium used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM may be in various forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM).

[0117] The database involved in each embodiment provided in this application may include at least one of a relational database and a non-relational database. The non-relational database may include a distributed database based on blockchain, etc., but is not limited thereto. The processor involved in each embodiment provided in this application may be a general-purpose processor, a central processing unit, a graphics processor, a digital signal processor, a programmable logic device, a data processing logic device based on quantum computing, etc., but is not limited thereto.

[0118] The technical features of the above embodiments may be arbitrarily combined. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0119] This article uses specific examples to illustrate the principles and implementation methods of this application. The description of the above embodiments is only used to help understand the method and core ideas of this application. At the same time, for those skilled in the art, according to the ideas of this application, there will be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as limiting this application.

Claims

1. A method for constructing a heterogeneous graph convolution epilepsy classification model, characterized in that: include: Construct feature extraction module; The feature extraction module is used to extract the skeletal motion features and facial features of the target person in the video data; Construct F2AGCN network model; The F2AGCN network model is an adaptive graph convolutional network model that integrates face and motion, and the F2AGCN network model includes: a first average pooling layer, a second average pooling layer, a first fully connected layer, a second fully connected layer, a fusion layer, and three FFDE modules connected in sequence; the FFDE module is used to perform deep feature extraction on skeletal motion features and facial features to obtain skeletal motion deep features and facial deep features; the first average pooling layer and the second average pooling layer are both connected to the last FFDE module of the three FFDE modules connected in sequence; the first average pooling layer and the first fully connected layer are used to perform epilepsy classification based on the skeletal motion deep features to obtain a motion classification result, the second average pooling layer and the second fully connected layer are used to perform epilepsy classification based on the facial deep features to obtain a facial classification result, and the fusion layer is used to fuse the motion classification result and the facial classification result to obtain a comprehensive classification result; The F2AGCN network model is trained using a modality coordination learning mechanism to obtain a trained F2AGCN network model as a heterogeneous graph convolution epilepsy classification model; The FFDE module includes an F2AGCN block and an AGCN block; The F2AGCN block is obtained by improving the AGCN block by replacing the two-dimensional convolutional layer in the AGCN block with a convolutional block attention module; The F2AGCN block includes: a GCN module, a convolutional block attention module, a connection module, a time convolution module and a post-processing module; The GCN module is used to extract deep features of the skeletal motion features to obtain spatial skeletal motion deep features; The convolutional block attention module is used to extract deep features of facial features to obtain spatial facial deep features; The connection module is used to connect the spatial skeletal motion depth features and the spatial facial depth features to obtain a facial-motion heterogeneous graph structure; The temporal convolution module is used to extract temporal features from the face-motion heterogeneous graph structure to obtain face-motion depth features; The post-processing module is used to separate the face-motion depth feature to obtain the skeleton motion depth feature and the face depth feature.

2. The method for constructing a heterogeneous graph convolutional epilepsy classification model according to claim 1, characterized in that: In terms of extracting the skeletal motion features and facial features of the target person in the video data, the feature extraction module is specifically used to: The human body posture estimation algorithm is used to extract the skeleton point information of the n-th frame video image in the video data; n=1,2,...,N, where N is the number of frames of the video data; The distance between each skeleton point is calculated using the skeleton point information of the nth frame video image; Determine the target person's skeleton point information according to the distance between each skeleton point; Based on the target person's skeleton point information, human body posture estimation is performed to obtain the target person's skeleton motion characteristics; The face of the target person is located based on the skeleton point information of the target person, and the facial contour of the target person is obtained; The facial features of the target person in the nth frame of the video image are extracted based on the facial contour of the target person.

3. The method for constructing a heterogeneous graph convolutional epilepsy classification model according to claim 2, characterized in that: The formula for locating the face of the target person based on the skeleton point information of the target person and obtaining the facial contour of the target person is: ; ; in, and Represent the cropping coordinates of the facial contour along the horizontal and vertical dimensions, , and are the horizontal coordinates of the first, second and i-th bone points on the target person’s face, i is the number of bone points on the target person’s face, , , are the ordinates of the first, second, and i-th bone points on the target person’s face, and They are the height and width of the video data respectively.

4. The method for constructing a heterogeneous graph convolutional epilepsy classification model according to claim 1, characterized in that: The convolutional block attention module includes a channel attention module and a spatial attention module connected in sequence.

5. The method for constructing a heterogeneous graph convolutional epilepsy classification model according to claim 1, characterized in that: The loss function of training the F2AGCN network model using the modal coordination learning mechanism is: ; in, is the total loss function, is the cross entropy loss function, is the MSE loss function, is the result of motion classification, is the face classification result, is a hyperparameter used to adjust the relative importance of the two branch losses.

6. An epilepsy monitoring device, characterized in that: The device comprises: a camera, a server and a terminal; The camera is connected to the server, the server is connected to the terminal, a heterogeneous graph convolutional epilepsy classification model is provided in the server, and the heterogeneous graph convolutional epilepsy classification model is obtained by using the heterogeneous graph convolutional epilepsy classification model construction method according to any one of claims 1 to 5; The camera is used to monitor the target person and obtain video data; The server is used to perform epilepsy classification based on the video data using the heterogeneous graph convolution epilepsy classification model to obtain a comprehensive classification result; The terminal is used to receive and display the comprehensive classification result.

7. A computer device comprising: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the heterogeneous graph convolution epilepsy classification model construction method according to any one of claims 1 to 5.

8. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method for constructing a heterogeneous graph convolution epilepsy classification model described in any one of claims 1 to 5 is implemented.