Face and posture feature collaborative epilepsy classification model construction method, epilepsy monitoring device, equipment, medium and product

By constructing an epilepsy classification model that coordinates facial and posture characteristics, using video data to realize the monitoring and classification of epilepsy, the problems of EEG-based equipment limitation and lack of video data sets in the prior art are solved, and high-precision epilepsy classification and monitoring are achieved.

CN120015341AActive Publication Date: 2025-05-16BEIJING CHILDRENS HOSPITAL AFFILIATED TO CAPITAL MEDICAL UNIV +1

Patent Information

Application Number
CN202510072429.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-16
Publication Date
2025-05-16
Estimated Expiration
2045-01-16

AI Technical Summary

Technical Problem

In the prior art, the EEG-based epilepsy classification method has the problem of equipment restricting patient activities, and the lack of annotated open source epilepsy video data sets, resulting in less application of visual data in epilepsy classification research.

Method used

A method for constructing epilepsy classification model that coordinates facial and pose features is proposed. By constructing a feature extraction module and an adaptive graph convolutional network model (FAGCN network model) that fuses facial and poses, monitoring and classification of epilepsy seizures is realized with video data.

Benefits of technology

Epilepsy classification and monitoring based on video data is realized, the defects in EEG monitoring that limits patient activity are overcome, and classification accuracy is improved by fusing facial and posture characteristics.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120015341A_ABST
    Figure CN120015341A_ABST
Patent Text Reader

Abstract

The invention discloses a face and posture feature collaborative epilepsy classification model construction method, an epilepsy monitoring device, equipment, a medium and a product, and relates to the field of intelligent medical instruments. The method comprises the steps of firstly constructing a feature extraction module; then constructing an adaptive graph convolutional network model fusing face and posture information, namely an FAGCN network model; the method comprises the steps that firstly, an epilepsy data set is constructed, an FAGCN network model is trained through the epilepsy data set, a trained FAGCN network model is obtained, then a feature extraction module and the trained FAGCN network model are connected, and an epilepsy classification model with the cooperative face and posture features is obtained. According to the method, the adaptive graph convolutional network model fusing the face and posture information is constructed and is used for epilepsy classification, and epilepsy classification and monitoring can be directly realized based on video data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of medical devices, and in particular to a method for constructing an epilepsy classification model that coordinates facial and posture features, an epilepsy monitoring device, equipment, medium and product. Background Art

[0002] Epilepsy is a chronic disease in which brain neurons suddenly discharge abnormally, leading to transient brain dysfunction. According to the latest epidemiological data in China, the overall prevalence of epilepsy in China is 7.0‰. It is estimated that there are about 9 million epilepsy patients in China, and about 400,000 new epilepsy patients are added each year. Epilepsy has become the second most common disease in neurology in China, second only to headache.

[0003] Involuntary body stiffness or spasms caused by epileptic seizures, as well as visual information such as facial loss and hand bending, are one of the important ways for epilepsy experts to identify epilepsy. However, faced with a large number of cases, experts often need to spend a lot of time and energy to identify epileptic seizures. Therefore, proposing a more accurate epilepsy classification algorithm for patient detection management and diagnostic decision-making is an effective way to reduce the burden on experts and assist expert diagnosis and treatment. When a patient has an attack, the epilepsy classification algorithm can be used through monitoring and other facilities to detect and obtain diagnosis and treatment in a timely manner, providing effective protection for the patient's health.

[0004] At present, classification methods based on deep learning have shown extremely excellent performance in various fields. Neural networks can overcome these challenges by automatically learning more robust features of changes in data distribution from training data. In the field of epilepsy classification based on deep learning, it is mainly divided into epilepsy classification research based on electroencephalogram (EEG) data and epilepsy classification research based on visual data.

[0005] Most of the epilepsy research based on deep learning is based on EEG. At present, some significant progress has been made in the use of EEG for detection. Research has widely used deep learning models such as Convolutional Neural Networks (CNN), Recurrent Neural Networks (RNN) and Long-Short Term Memory (LSTM) to process epilepsy EEG data. At the same time, there are many open source epilepsy EEG datasets annotated by experts, but using EEG to judge the onset of patients requires special equipment, which will limit the patient's movement. At the same time, video data, as a non-invasive source of information, shows great potential in disease detection, such as Parkinson's disease, Alzheimer's disease and spinal muscular atrophy. However, due to the complex and changeable visual manifestations of epileptic seizures, the video visual features of different types of epileptic seizures are not clearly distinguished, and some types of epileptic seizures have no specific visual features; in addition, the shooting conditions of clinical monitoring videos vary greatly, and epileptic seizures usually occur suddenly and stop suddenly. There is currently a lack of annotated open source epilepsy video datasets, and there are few related studies. Summary of the invention

[0006] The purpose of this application is to provide a method for constructing an epilepsy classification model that coordinates facial and posture features, an epilepsy monitoring device, equipment, medium and product, so as to realize epileptic seizure monitoring based on video data and overcome the defect that electroencephalogram monitoring restricts patient activities.

[0007] To achieve the above objectives, this application provides the following solutions.

[0008] In a first aspect, the present application provides a method for constructing an epilepsy classification model based on the coordination of facial and posture features, comprising:

[0009] Constructing a feature extraction module; the feature extraction module is used to extract the skeleton point feature sequence and the facial feature sequence of the target person in the video data;

[0010] Constructing a FAGCN network model; the FAGCN network model is an adaptive graph convolutional network model that integrates face and posture, and the FAGCN network model includes: 10 FAGCN blocks connected in sequence and 1 fully connected layer, wherein the 1st, 5th and 8th FAGCN blocks are all first FAGCN blocks, the 1st FAGCN block is connected to the 5th FAGCN block through a Reshape function and a first convolution and tiling module, the 5th FAGCN block is connected to the 8th FAGCN block through a Reshape function and a first convolution and tiling module, and the remaining 7 FAGCN blocks are all second FAGCN blocks; the first FAGCN block is used to process a skeleton point feature sequence and a facial feature sequence, and the second FAGCN block is used to process a skeleton feature sequence;

[0011] Construct epilepsy dataset;

[0012] Based on the epilepsy data set, the FAGCN network model is trained to obtain a trained FAGCN network model;

[0013] The feature extraction module is connected to the trained FAGCN network model to obtain an epilepsy classification model that coordinates facial and posture features.

[0014] In a second aspect, the present application provides an epilepsy monitoring device, the device comprising: a camera, a server and a terminal;

[0015] The camera is connected to the server, the server is connected to the terminal, the server is provided with an epilepsy classification model coordinated by facial and posture features, and the epilepsy classification model coordinated by facial and posture features is obtained by using the above-mentioned epilepsy classification model construction method coordinated by facial and posture features;

[0016] The camera is used to monitor the target person and obtain video data;

[0017] The server is used to perform epilepsy recognition based on the video data by using the epilepsy classification model coordinated with the facial and posture features to obtain a recognition result;

[0018] The terminal is used to receive and display the recognition result.

[0019] In a third aspect, the present application provides a computer device, comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the above-mentioned method for constructing an epilepsy classification model by coordinating facial and posture features.

[0020] In a fourth aspect, the present application provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the above-mentioned method for constructing an epilepsy classification model by coordinating facial and posture features.

[0021] In a fifth aspect, the present application provides a computer program product, which, when executed by a processor, implements the above-mentioned method for constructing an epilepsy classification model by coordinating facial and posture features.

[0022] According to the specific embodiments provided in this application, this application has the following technical effects.

[0023] The present application provides a method for constructing an epilepsy classification model that coordinates facial and posture features, an epilepsy monitoring device, equipment, medium and product. First, a feature extraction module is constructed; then an adaptive graph convolutional network model that integrates facial and posture is constructed, namely, a FAGCN network model; then an epilepsy data set is constructed, and the FAGCN network model is trained using the epilepsy data set to obtain a trained FAGCN network model, and then the feature extraction module and the trained FAGCN network model are connected to obtain an epilepsy classification model that coordinates facial and posture features. The present application constructs an adaptive graph convolutional network model that integrates facial and posture information for epilepsy classification, which can directly realize the classification and monitoring of epilepsy based on video data. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the drawings required for use in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0025] Figure 1 A flowchart of a method for constructing an epilepsy classification model based on the coordination of facial and posture features provided in one embodiment of the present application;

[0026] Figure 2 A schematic diagram of the structure of a FAGCN network model provided in an embodiment of the present application;

[0027] Figure 3 A schematic diagram of skeleton information provided in an embodiment of the present application;

[0028] Figure 4 The accuracy curve diagram of each algorithm training process provided in an embodiment of the present application;

[0029] Figure 5 A Loss curve diagram of each algorithm training process provided in an embodiment of the present application;

[0030] Figure 6 A schematic diagram of the structure of a computer device provided in one embodiment of the present application. DETAILED DESCRIPTION

[0031] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.

[0032] In order to make the above-mentioned objects, features and advantages of the present application more obvious and easy to understand, the present application is further described in detail below with reference to the accompanying drawings and specific implementation methods.

[0033] In an exemplary embodiment, Figure 1 As shown, a method for constructing an epilepsy classification model by coordinating facial and posture features is provided, including the following steps 101 to 105.

[0034] Step 101, constructing a feature extraction module; the feature extraction module is used to extract the skeleton point feature sequence and the facial feature sequence of the target person in the video data.

[0035] Step 102, constructing a FAGCN network model; the FAGCN network model is an adaptive graph convolutional network model that integrates facial and posture information, and the FAGCN network model includes: 10 FAGCN blocks connected in sequence and 1 fully connected layer, wherein the 1st, 5th and 8th FAGCN blocks are all first FAGCN blocks, the 1st FAGCN block is connected to the 5th FAGCN block through a Reshape function and a first convolution and tiling module, the 5th FAGCN block is connected to the 8th FAGCN block through a Reshape function and a first convolution and tiling module, and the remaining 7 FAGCN blocks are all second FAGCN blocks; the first FAGCN block is used to process a skeleton point feature sequence and a facial feature sequence, and the second FAGCN block is used to process a skeleton feature sequence.

[0036] Step 103: construct an epilepsy dataset.

[0037] Step 104: training the FAGCN network model based on the epilepsy dataset to obtain a trained FAGCN network model.

[0038] Step 105, connecting the feature extraction module and the trained FAGCN network model to obtain an epilepsy classification model that coordinates facial and posture features.

[0039] By implementing the above steps 101 to 105, an adaptive graph convolutional network model integrating facial and posture information is constructed for epilepsy classification, which can directly realize epilepsy classification and monitoring based on video data.

[0040] Before building the adaptive graph convolutional network model that fuses facial and posture information in the above steps, a model that can use visual data for epilepsy classification was studied.

[0041] In recent years, with the improvement of computer technology, the development of deep learning technology and the improvement of graphics card computing power, it has become possible to use deep learning technology to solve video classification problems. Deep learning models have considerable potential in video recognition and classification, and can be used to detect and identify signs or characteristics of certain diseases, including epilepsy. Some studies, such as the algorithm proposed by Ahmedt-Aristizabal D et al., extract the motion features of each frame of the video through a neural convolutional network (CNN), and then use a recurrent neural network to extract information from the aggregated feature fragments, and finally classify them through a fully connected layer. This method has some problems. First, in order to reduce the amount of calculation, the algorithm determines the position of the bed in the first frame and then crops it with the bed as the center. This has limitations, and the displacement of the bed will seriously affect the accuracy of the algorithm. Second, the use of CNN network training to extract motion features on medical data sets, which are generally small, leads to limitations in the extracted motion features.

[0042] In subsequent studies, such as P′erez- The epilepsy research algorithm proposed by F et al. does not use the motion model trained on the current medical dataset. Instead, it uses the CNN model STCNN trained on a large-scale natural human motion dataset to extract motion sequences, and then uses RNN to predict and classify the sequences, achieving an accuracy of about 90%.

[0043] The above method is an effective framework for studying the identification and classification of diseases based on visual data. However, in the study of epilepsy classification, facial information during epileptic seizures is not used, and the sequence classification models such as RNN used at the same time are relatively basic algorithms. With the development of deep learning technology, more time series classification models have been proposed, such as LSTM, Seq2Seq, Transformer, etc.; there are also spatiotemporal sequence models that extract features from sequences in the time domain and spatial domain, such as ST-GCN, AGCN, etc. These models have achieved good accuracy in a wide range of fields of action recognition and classification.

[0044] By analyzing the movements and facial features during epileptic seizures and various time series classification algorithms in recent years, the embodiment of the present application is improved based on the AGCN model. On the basis of using the skeleton point motion data, the facial area in the video data is located by the coordinates of the patient's facial skeleton points, so that facial features and motion features are used as input, and the processing and calculation are integrated in the algorithm to complete the classification of epilepsy.

[0045] At present, most of the epilepsy recognition algorithms based on deep learning are based on EEG datasets. However, for patients, especially children, the detection equipment will greatly restrict their actions and cause many inconveniences. However, open source epilepsy datasets based on vision are extremely rare. Therefore, it is necessary to obtain visual data of epileptic seizures and use a series of methods to process the data into datasets that can be used for deep learning experiments.

[0046] In another exemplary embodiment, the epilepsy data set construction in the above step 103 includes three steps: data screening, skeleton point extraction, and data processing, which are described in detail below.

[0047] 1. Data screening.

[0048] In order to ensure the accuracy and confidence of subsequent skeleton point extraction, and thus affect the accuracy of classification, it is necessary to screen the video data of epileptic seizures. The video data studied was extracted from video EEG examinations in clinical diagnosis and treatment activities. There are cases where objects block the patient's body, or family members and medical staff interfere with the patient during the seizure, which hinders the normal recognition of epileptic seizure movements, so they need to be eliminated.

[0049] 2. Skeleton point extraction.

[0050] Since epilepsy is a sudden disease, the video data is the synchronous monitoring video of the video EEG examination during clinical diagnosis and treatment activities. This makes the data environment complex and diverse, and the surrounding medical staff and family members will become interference factors. For this situation where the patient target is small and there is a lot of redundant data, the redundant information can be reduced by extracting the patient's skeleton point information. The pre-trained human skeleton point recognition model is used to process the epilepsy data and extract the patient's skeleton point information.

[0051] However, in the process of using the pre-trained human skeleton point recognition model to process epilepsy data, due to the presence of the patient's family members, the skeleton points in each frame extracted are interfered by other people, so it is necessary to accurately extract the patient's skeleton point information from the skeleton points of multiple people. The skeleton point information of the patient in the first frame can be manually selected, and then the Euclidean distance between the coordinates of each person's skeleton point in the next frame and the coordinates of the patient's skeleton point in this frame is calculated to obtain the patient's skeleton point information in the next frame. Due to the complex data environment, there is a situation where the patient's skeleton point is not recognized in a certain frame. By calculating the Euclidean distance, the wrong skeleton point information may be selected. Therefore, the threshold can be increased to determine whether the Euclidean distance exceeds the threshold to prevent the extraction of wrong skeleton point information.

[0052] 3. Organize epilepsy data.

[0053] Since the video contains both epileptic seizures and normal states of the patient, the output skeleton points need to be intercepted according to the corresponding seizure time in the video annotated by the epilepsy expert, and the corresponding video data is sliced ​​according to the skeleton point data to facilitate the acquisition of facial data corresponding to the skeleton points. The epilepsy data is intercepted as 210 frames per data according to the expert annotation, so as to obtain the seizure and normal data, and the skeleton point data, video data, and corresponding labels are finally organized into a data set. The data set is randomly divided into training set, prediction set, and test set in a ratio of 7:2:1 to obtain the epilepsy data set.

[0054] Epilepsy patient video data includes clinical videos of patients and their families or / and medical staff. In order to obtain the patient's facial video data and input it into the subsequent algorithm steps, the epilepsy patient video data needs to be processed by the facial positioning algorithm. The facial positioning algorithm obtains all facial coordinate data from the skeletal point data corresponding to the video, which can be expressed as a set FS = {(x 1 ,y 1 ),(x 2 ,y 2 ),…(x i ,y i )}, and then locate the patient's face area, and the positioning algorithm formula is as follows:

[0055]

[0056]

[0057] Where FAX and FAY represent the cropping coordinates of the facial contour along the horizontal and vertical dimensions, respectively. 1 、x 2 and x i are the horizontal coordinates of the first, second and i-th bone points on the target person’s face, y 1 ,y2 ,y i are the ordinates of the first, second, and i-th bone points on the target person’s face, respectively. h and V w They are the height and width of the video data respectively.

[0058] In formulas (1) and (2), MIN and MAX are minimum and maximum value functions.

[0059] After being processed by the facial positioning algorithm, the obtained facial area data is convolved and flattened and then input into the subsequent FAGCN block together with the corresponding bone point action data.

[0060] In another exemplary embodiment, in terms of extracting the skeleton point feature sequence and facial feature sequence of the target person in the video data, the above-mentioned feature extraction module is specifically used to: use a human posture estimation algorithm to extract the skeleton point information of each person in the n-th frame video image in the video data; n = 1, 2, ..., N, N is the number of frames of the video data; calculate the Euclidean distance between the skeleton point information of each person in the n-th frame video image and the skeleton point information of the target person in the n-1-th frame video image; determine the skeleton point information of the person with the smallest Euclidean distance in the n-th frame video image as the target person in the n-th frame video image skeleton point information of the target person in the n-th frame video image; human body posture estimation is performed based on the skeleton point information of the target person in the n-th frame video image, and the skeleton point features of the target person in the n-th frame video image are obtained; the face of the target person is located based on the skeleton point information of the target person in the n-th frame video image, and the facial contour of the target person in the n-th frame video image is obtained; the facial features of the target person in the n-th frame video image are extracted based on the facial contour of the target person in the n-th frame video image; based on the skeleton point features and facial features of the target person in each frame video image in the video data, a skeleton point feature sequence and a facial feature sequence of the target person are constructed.

[0061] In another exemplary embodiment, the overall framework of the above FAGCN network model is as follows: Figure 2 As shown, Figure 2 (a) is a schematic diagram of the input structure of the FAGCN network model. Figure 2 (b) is a schematic diagram of the composition structure of the FAGCN network model. Figure 2 (c) in the figure is a schematic diagram of the overall structure of the FAGCN network model.

[0062] like Figure 2As shown in (c) in the figure, FAGCN mainly includes 10 FAGCN blocks that are improved based on the AGCN block. Each FAGCN block contains a graph convolution module (GCN) and a temporal convolution module (TCN). In the graph convolution module GCN, the input data is first downsampled by convolution conv1 to extract features, and then the data dimension is adjusted using the Reshape and Transpose functions, so that the neighborhood of the skeleton points in the data is determined by the subsequent convolution conv2 and weighted sum calculation is performed to highlight the important features in the skeleton action. The temporal convolution module (TCN) contains two normalization functions (BN) and an activation function (RELU) as well as a temporal convolution. The convolution uses a convolution kernel of size 9×1 to perform calculations in the temporal dimension to extract continuous action features in the skeleton point sequence. The number of input and output channels of each FAGCN block in the algorithm is shown in Table 1.

[0063] Table 1 Number of input and output channels of FAGCN block

[0064]

[0065] Among the 10 FAGCN blocks, they can be divided into two categories according to different inputs. The first category is the 1st, 5th, and 8th FAGCN blocks represented by blue (i.e. Figure 2 The second category is the FAGCN block represented by green (i.e. Figure 2 The green block in the Figure 2 Taking Block 1 in (a) as an example, the blue FAGCN block, i.e., the first FAGCN block, is introduced. In these three first FAGCN blocks, the input includes not only the skeleton feature sequence but also the facial feature sequence. The skeleton feature sequence is processed by the graph convolution module; the facial feature sequence is converted into a one-dimensional sequence through convolution and flattening. The two outputs are converted into the same dimension and concatenated and input into the temporal convolution module to extract temporal features. After being processed by the temporal convolution module, the two feature sequences are split, the skeleton feature sequence is input into the subsequent multiple green FAGCN blocks, and the facial feature sequence is converted into a two-dimensional feature sequence again through the Reshape function, and after convolution and flattening, it is jumped and input into the 5th or 8th FAGCN block. The green FAGCN block inputs only the skeleton feature sequence, and residual convolution is added to improve the stability of the algorithm.

[0066] Finally, the skeleton point feature sequence processed by 10 FAGCN blocks and the facial feature sequence processed by 3 FAGCN blocks are input into the fully connected layer to obtain the final classification result.

[0067] Therefore, in an embodiment of the present application, the first FAGCN block includes: a first graph convolution module, a second convolution and tiling module, a connection module, a first time convolution module and a first post-processing module; the first graph convolution module and the second convolution and tiling module are both connected to the connection module, the connection module is connected to the first time convolution module, and the first time convolution module is connected to the first post-processing module; the first graph convolution module is used to extract deep features from the target bone point feature sequence to obtain a spatial target bone point deep feature sequence; the target bone point feature sequence is a bone point feature sequence or a bone point deep feature sequence output by a second FAGCN block located before the first FAGCN block; the second convolution module The tiling module is used to extract deep features from the target facial feature sequence to obtain a spatial target facial depth feature sequence; the target facial feature sequence is a facial feature sequence or a facial depth feature sequence output by the first FAGCN block located before the first FAGCN block; the connection module is used to connect the spatial target bone point depth feature sequence and the spatial target facial depth feature sequence to obtain a facial-posture heterogeneous graph structure sequence; the first temporal convolution module is used to extract temporal features from the facial-posture heterogeneous graph structure sequence to obtain a facial-posture depth feature sequence; the first post-processing module is used to separate the facial-posture depth feature sequence to obtain a bone point depth feature sequence and a facial depth feature sequence.

[0068] The second FAGCN block includes a second graph convolution module and a second temporal convolution module connected in sequence; the second graph convolution module is used to perform deep feature extraction on the target bone point feature sequence to obtain a spatial target bone point deep feature sequence; the target bone point feature sequence is a bone point deep feature sequence output by the first FAGCN block or the second FAGCN block located before the second FAGCN block; the second temporal convolution module is used to perform temporal feature extraction on the spatial target bone point deep feature sequence to obtain a bone point deep feature sequence.

[0069] In order to illustrate the effect of the model constructed in this application, an experimental verification was carried out, and the characteristics of the experimental data used in the experiment and the processing process of processing the experimental data into a data set were introduced. Then, the experimental results of the model of this application were displayed and compared with other mainstream algorithms to prove the effectiveness of the FAGCN network model, as follows.

[0070] The data used in the study of the examples of this application are derived from clinical data from Beijing Children's Hospital affiliated to Capital Medical University. After being labeled by epilepsy specialists, a total of 38 seizure video records of 30 patients were collected, including clinical information such as video frame number, seizure time node and seizure type. These video materials are all collected from synchronized monitoring videos of patients undergoing video EEG examinations, and the average length of a single video is about 2 minutes. In view of the sudden nature of epileptic seizures, in addition to recording the patient's performance, the video also includes activity pictures of family members and medical staff during clinical treatment, and different patients have different movement characteristics during seizures.

[0071] In order to extract the patient's kinematic features, the present embodiment uses a pre-trained human skeleton point recognition model to analyze the epileptic seizure video. Specifically, the OpenPose deep learning human posture estimation tool is used, which can accurately detect and evaluate the key points and posture information of the human body from images or videos. Figure 3 As shown in the figure, the human skeleton consists of 25 key points (numbered 0-24), corresponding to: nose, neck, right shoulder, right elbow, right wrist, left shoulder, left elbow, left wrist, middle hip, right hip, right knee, right ankle, left hip, left knee, left ankle, right eye, left eye, right ear, left big toe, left little toe, left heel, right big toe, right little toe and right heel. The skeleton point information is obtained through OpenPose processing, and the video clips are divided into two categories: attack period and non-attack period according to the attack time nodes marked by experts, and a binary classification dataset is constructed accordingly.

[0072] The experiment was conducted on a GPU server in the Baidu Paddle AI Studio environment, using the open source Baidu Paddle Paddle open platform, with a framework version of 2.2.2 and a Python version of 3.7. The dataset used was the self-made epilepsy classification dataset mentioned above. In order to demonstrate the effect of integrating facial features in this algorithm, it was compared with five time series algorithms LSTM, Transformer, ST-GCN, AGCN, and CTR-GCN, whose input was only bone point sequence data. The epoch of algorithm training was set to 100, and the Adam optimizer was used in the process. The learning rate was set to 0.001, the weight decay was set to 0.0001, and the loss function used the cross entropy function.

[0073] The training process of each algorithm in the experiment draws the Accuarcy and Loss curves, such as Figure 4 , Figure 5 shown.

[0074] Figure 4 , Figure 5The Accuarcy and Loss curves of each algorithm model in 0-100 rounds of training are shown in Figure 2. It can be seen that the fitting effect of the FAGCN algorithm in 80-100 rounds of training is slightly higher than that of other algorithms, which shows that adding facial features has a certain improvement effect on algorithm fitting. After 100 rounds of training, the experimental results are shown in Table 2.

[0075] The data in Table 2 are the training models of six algorithms on the epilepsy classification training set and the classification results on the test set. In Table 2, the algorithm FAGCN of this application is distinguished from the five comparison algorithms based on the input data, and the algorithm test results are displayed using four evaluation indicators: Precision, Recall, F1 Score, and Accuarcy. The best results are marked in bold. It can be seen that the FAGCN algorithm of this application exceeds the other five algorithms in all four evaluation indicators, proving that FAGCN has a high-precision classification effect on epileptic seizures. It also proves that adding facial features can assist action features in improving the accuracy of epilepsy classification algorithms.

[0076] Table 2 Experimental results

[0077]

[0078]

[0079] Epileptic seizures often occur suddenly and stop suddenly. If they cannot be discovered and treated in time, they often lead to injuries or even death of patients. As a non-invasive monitoring method, real-time observation of whether patients are ill through video monitoring can effectively reduce human and material resources. The embodiment of the present application analyzes the motion characteristics and facial features during epileptic seizures, and proposes an algorithm FAGCN that is improved based on AGCN and integrates facial features and skeletal point motion features to classify epilepsy. An epilepsy classification dataset was created by extracting synchronized video data from video EEG monitoring of epileptic patients, and experiments were conducted on this dataset using FAGCN and multiple comparison algorithms, which proved the effectiveness and superiority of the FAGCN algorithm. It also proved that adding facial features can effectively improve the algorithm's classification accuracy for epilepsy.

[0080] In an exemplary embodiment, the epilepsy classification model of coordinated facial and posture features constructed in the above method embodiment is applied to an epilepsy monitoring device, which includes: a camera, a server and a terminal; the camera is connected to the server, and the server is connected to the terminal, and the server is provided with an epilepsy classification model of coordinated facial and posture features, and the epilepsy classification model of coordinated facial and posture features is obtained by using the epilepsy classification model of coordinated facial and posture features construction method of the above embodiment; the camera is used to monitor the target person and obtain video data; the server is used to perform epilepsy recognition based on the video data using the epilepsy classification model of coordinated facial and posture features to obtain a recognition result; the terminal is used to receive and display the recognition result.

[0081] In an exemplary embodiment, a computer device is provided. The computer device may be a server or a terminal. The internal structure diagram thereof may be as follows: Figure 6 As shown. The computer device includes a processor, a memory, an input / output interface (Input / Output, referred to as I / O) and a communication interface. The processor, the memory and the input / output interface are connected through a system bus, and the communication interface is connected to the system bus through the input / output interface. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The input / output interface of the computer device is used to exchange information between the processor and an external device. The communication interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, a method for constructing an epilepsy classification model that cooperates with facial and posture features is implemented.

[0082] Those skilled in the art will understand that Figure 6 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine certain components, or have a different arrangement of components.

[0083] In an exemplary embodiment, a computer device is provided, including a memory and a processor, wherein a computer program is stored in the memory, and the processor implements the steps in the above-mentioned method embodiments when executing the computer program.

[0084] In an exemplary embodiment, a computer-readable storage medium is provided, storing a computer program, and when the computer program is executed by a processor, the steps in the above method embodiments are implemented.

[0085] In an exemplary embodiment, a computer program product is provided, which, when executed by a processor, implements the steps in the above method embodiments.

[0086] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties.

[0087] Those of ordinary skill in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to the memory, database or other medium used in the embodiments provided in the present application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM may be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM).

[0088] The database involved in each embodiment provided in this application may include at least one of a relational database and a non-relational database. The non-relational database may include a distributed database based on blockchain, etc., but is not limited thereto. The processor involved in each embodiment provided in this application may be a general-purpose processor, a central processing unit, a graphics processor, a digital signal processor, a programmable logic device, a data processing logic device based on quantum computing, etc., but is not limited thereto.

[0089] The technical features of the above embodiments may be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0090] This article uses specific examples to illustrate the principles and implementation methods of this application. The description of the above embodiments is only used to help understand the method and core ideas of this application. At the same time, for those skilled in the art, according to the ideas of this application, there will be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as limiting this application.

Claims

1. A method for constructing an epilepsy classification model based on the synergy of facial and posture features, characterized in that: include: Construct feature extraction module; The feature extraction module is used to extract the skeleton feature sequence and facial feature sequence of the target person in the video data; Constructing a FAGCN network model; the FAGCN network model is an adaptive graph convolutional network model that integrates facial and posture features, and the FAGCN network model includes: 10 FAGCN blocks connected in sequence and 1 fully connected layer, wherein the 1st, 5th and 8th FAGCN blocks are all first FAGCN blocks, the 1st FAGCN block is connected to the 5th FAGCN block through a Reshape function and a first convolution and tiling module, the 5th FAGCN block is connected to the 8th FAGCN block through a Reshape function and a first convolution and tiling module, and the remaining 7 FAGCN blocks are all second FAGCN blocks; the first FAGCN block is used to process a skeleton point feature sequence and a facial feature sequence, and the second FAGCN block is used to process a skeleton feature sequence; Construct epilepsy dataset; Based on the epilepsy data set, the FAGCN network model is trained to obtain a trained FAGCN network model; The feature extraction module is connected to the trained FAGCN network model to obtain an epilepsy classification model that coordinates facial and posture features.

2. The method for constructing an epilepsy classification model based on facial and posture features according to claim 1, characterized in that: In terms of extracting the skeleton point feature sequence and facial feature sequence of the target person in the video data, the feature extraction module is specifically used for: The human body posture estimation algorithm is used to extract the skeleton point information of each person in the n-th frame video image in the video data; n=1, 2, ..., N, where N is the number of frames of the video data; Calculate the Euclidean distance between the skeleton point information of each person in the nth frame of the video image and the skeleton point information of the target person in the n-1th frame of the video image; Determine the skeleton point information of the person with the smallest Euclidean distance in the n-th frame of the video image as the skeleton point information of the target person in the n-th frame of the video image; Based on the skeleton point information of the target person in the n-th frame video image, the human body posture is estimated to obtain the skeleton point features of the target person in the n-th frame video image; The face of the target person is located based on the skeleton point information of the target person in the n-th frame of the video image, and the facial contour of the target person in the n-th frame of the video image is obtained; Extracting facial features of a target person in the n-th frame of the video image based on the facial contour of the target person in the n-th frame of the video image; Based on the skeleton point features and facial features of the target person in each frame of video image in the video data, a skeleton point feature sequence and a facial feature sequence of the target person are constructed.

3. The method for constructing an epilepsy classification model based on facial and posture features according to claim 2, characterized in that: The formula for extracting the facial features of the target person in the n-th frame of the video image based on the facial contour of the target person in the n-th frame of the video image is: Among them, FAX and FAY represent the cropping coordinates of the facial contour along the horizontal and vertical dimensions, respectively, and x1, x2 and x i are the horizontal coordinates of the first, second, and i-th bone points on the target person’s face, y1, y2, y i are the ordinates of the first, second, and i-th bone points on the target person’s face, respectively. h and V w They are the height and width of the video data respectively.

4. The method for constructing an epilepsy classification model based on facial and posture features according to claim 1, characterized in that: The first FAGCN block includes: a first graph convolution module, a second convolution and tiling module, a connection module, a first time convolution module and a first post-processing module; the first graph convolution module and the second convolution and tiling module are both connected to the connection module, the connection module is connected to the first time convolution module, and the first time convolution module is connected to the first post-processing module; The first graph convolution module is used to extract deep features from the target bone point feature sequence to obtain a spatial target bone point deep feature sequence; the target bone point feature sequence is a bone point feature sequence or a bone point deep feature sequence output by a second FAGCN block located before the first FAGCN block; The second convolution and tiling module is used to extract deep features from the target facial feature sequence to obtain a spatial target facial deep feature sequence; the target facial feature sequence is a facial feature sequence or a facial deep feature sequence output by a first FAGCN block located before the first FAGCN block; The connection module is used to connect the target skeleton point depth feature sequence in space and the target facial depth feature sequence in space to obtain a facial-posture heterogeneous graph structure sequence; The first time convolution module is used to extract temporal features from the face-posture heterogeneous graph structure sequence to obtain a face-posture deep feature sequence; The first post-processing module is used to separate the face-posture depth feature sequence to obtain a skeleton point depth feature sequence and a face depth feature sequence.

5. The method for constructing an epilepsy classification model based on facial and posture features according to claim 1, characterized in that: The second FAGCN block includes a second graph convolution module and a second temporal convolution module connected in sequence; The second graph convolution module is used to extract deep features from the target bone point feature sequence to obtain a spatial target bone point deep feature sequence; the target bone point feature sequence is a bone point deep feature sequence output by the first FAGCN block or the second FAGCN block located before the second FAGCN block; The second time convolution module is used to extract time features from the spatial target skeleton point depth feature sequence to obtain the skeleton point depth feature sequence.

6. The method for constructing an epilepsy classification model based on facial and posture features according to claim 1, characterized in that: The constructing of the epilepsy data set specifically includes: Obtain a video dataset of epilepsy patients; Filter the video data in the video data set, remove the video data in which the target person is blocked, and obtain the filtered video data set; A feature extraction module is used to extract a skeleton feature sequence and a facial feature sequence of a target person in each video data in the screened video data set; The epilepsy data set is constructed by marking the skeleton point feature sequence and the facial feature sequence of the target person in each video data as to whether the person has an epileptic seizure.

7. An epilepsy monitoring device, characterized in that: The device comprises: a camera, a server and a terminal; The camera is connected to the server, the server is connected to the terminal, the server is provided with an epilepsy classification model coordinated by facial and posture features, and the epilepsy classification model coordinated by facial and posture features is obtained by using the method for constructing an epilepsy classification model coordinated by facial and posture features according to any one of claims 1 to 6; The camera is used to monitor the target person and obtain video data; The server is used to perform epilepsy recognition based on the video data by using the epilepsy classification model coordinated with the facial and posture features to obtain a recognition result; The terminal is used to receive and display the recognition result.

8. A computer device comprising: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the method for constructing an epilepsy classification model based on the coordination of facial and posture features as described in any one of claims 1 to 6.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method for constructing an epilepsy classification model by coordinating facial and posture features according to any one of claims 1 to 6 is implemented.

10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the method for constructing an epilepsy classification model by coordinating facial and posture features according to any one of claims 1 to 6 is implemented.

Citation Information

Patent Citations

  • Systems and Methods for Optimizing Pose Estimation

    US20190171871A1

  • Pose estimation-based pedestrian fall action recognition method and device

    WO2023082882A1

Cited By

  • Real-time epilepsy behavior detection and analysis method and system

    CN121015145A

  • Epilepsy prediction system based on multi-modal biological image and image data processing method

    CN121565504A

  • Epilepsy prediction system based on multi-modal biological images and image data processing method

    CN121565504B