A method and system for multi-modal non-contact detection of seizures in children
Patent Information
- Application Number
- CN202610738101.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-05-27
- Publication Date
- 2026-08-21
AI Technical Summary
[0007]本申请提供了一种儿童癫痫发作多模态非接触检测方法及系统,主要在于改进深度学习模型,以儿童癫痫发作时在阵挛期中雷达信号的特异性表现为基准,改进模型的三通道并行特征提取模块和多模态癫痫发作识别模块,这些针对癫痫发作生理特点的结构和算法改进,有效提升了模型对不同发作类型的识别精度和鲁棒性,解决现有技术对癫痫发作的识别检测精度不足的问题
[0057] In summary, this embodiment of the application performs modal decomposition on the radar signals collected by non-contact millimeter-wave radar for detecting children, obtaining three modal component signals of body movement, respiration, and heart rate, along with temperature and acoustic event information collected by auxiliary sensors. Using each component signal and information as input, a three-channel parallel feature extraction module extracts the time-spectrum map in a deep learning model, and convolution yields feature vectors for the three components and auxiliary feature vectors, serving as four feature vectors related to epileptic seizures. A causal-aware temporal attention mechanism is introduced to weightedly fuse these four feature vectors. Multi-scale temporal network modeling is used to model and weightedly fuse the features. Epileptic seizure characteristics are detected through windows at different time scales, and multi-level hierarchical classification is used for identification, outputting the epileptic seizure type and confidence level. Finally, based on the epileptic seizure information output by the model, an alarm is output when the alarm conditions are met. Therefore, this embodiment of the application achieves non-contact real-time detection and classification of all types of epileptic seizures in children. Based on the characteristics of epileptic seizure manifestations, the detection model is improved to achieve multi-modal detection and improve the detection accuracy of epileptic seizures.
Smart Images

Figure CN122604305A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the fields of medical electronics and artificial intelligence technology, and in particular to a multimodal non-contact detection method and system for childhood epilepsy seizures. Background Technology
[0002] Epilepsy is one of the most common neurological disorders in children, affecting approximately 10.5 million children under the age of 15 worldwide with active epilepsy. The International League Against Epilepsy (ILAE) 2017 classification categorizes epileptic seizures into three main types: generalized seizures (including tonic-clonic, myoclonic, and absence seizures), focal seizures, and seizures of unknown origin. Childhood epilepsy not only directly impacts neurodevelopment but also poses a fatal risk of sudden epileptic death (SUDEP), especially generalized tonic-clonic seizures that occur unsupervised at night.
[0003] The current clinical gold standard—continuous video electroencephalography (cEEG) monitoring—has the following systemic drawbacks: it requires professional neurophysiologists to operate and interpret the equipment; electrode attachment can cause skin damage and discomfort to children (especially infants); the equipment is expensive and is usually only available in specialized hospitals; and it cannot enable long-term home monitoring.
[0004] Regarding existing non-EEG epilepsy detection technologies, wearable devices such as Empatica EpiMonitor and NightWatch have been clinically validated, but they can only detect generalized tonic-clonic seizures (GTCS) and are ineffective against absence seizures, focal seizures, and other non-motor seizures. Furthermore, wearability compliance is a significant challenge for young children. Video AI monitoring systems (such as Nelli) can detect various motor seizures, but they have inherent drawbacks such as privacy intrusion, light dependence, and sensitivity to occlusion. Mattress sensors (such as Emfit), while non-contact, have a sensitivity of only about 70% and cannot distinguish between seizure types.
[0005] Breakthroughs have been made in the application of radar technology for epilepsy detection. Fearns et al. (2022 American Epilepsy Association Annual Meeting) used a 60GHz millimeter-wave radar (TI IWR6843ISK) to detect 83 seizures in 25 patients in an epilepsy monitoring unit, and collaborated with Boston Children's Hospital on a clinical study of pediatric epilepsy detection based on ultra-wideband radar. However, these studies have not yet developed a complete systemic technical solution, particularly lacking detection methods for all types of seizures (such as absence seizures and other non-motor seizures), age-adaptive mechanisms for children, and a hospital-home dual-scenario deployment architecture.
[0006] It is evident that existing detection methods for childhood epilepsy seizures have not yet achieved non-contact, real-time detection and classification of all types of childhood epilepsy seizures, and the accuracy of epilepsy seizure identification and detection is insufficient. Summary of the Invention
[0007] This application provides a multimodal non-contact detection method and system for childhood epilepsy seizures. The main improvement lies in the deep learning model, which is based on the specific performance of radar signals during the clonic phase of a childhood epilepsy seizure. The model's three-channel parallel feature extraction module and multimodal epilepsy seizure recognition module are improved. These structural and algorithmic improvements, which are tailored to the physiological characteristics of epilepsy seizures, effectively enhance the model's recognition accuracy and robustness for different seizure types, and solve the problem of insufficient recognition accuracy of epilepsy seizures in existing technologies.
[0008] In a first aspect, this application provides a multimodal non-contact detection method for childhood epileptic seizures, including:
[0009] The reflected echo signal of the child collected by the millimeter-wave radar is decomposed into three modal components to obtain auxiliary information collected by the auxiliary sensor. The three modal components include body motion component, respiratory component and heart rate component.
[0010] A deep learning model consisting of a three-channel parallel feature extraction module and a multimodal epileptic seizure recognition module was constructed.
[0011] In the deep learning model, the three-channel parallel feature extraction module is used to extract the component feature vectors corresponding to each component from the three modal components, and to extract auxiliary feature vectors from the auxiliary information;
[0012] In the multimodal epileptic seizure recognition module, an attention mechanism is used to perform weighted fusion based on the component feature vectors and auxiliary feature vectors, and epileptic seizure information is output through multi-scale temporal network modeling.
[0013] When the epileptic seizure information meets the alarm conditions, the corresponding output mode is selected to output the alarm information.
[0014] Optionally, a deep learning model is constructed, consisting of a three-channel parallel feature extraction module and a multimodal epileptic seizure recognition module, including:
[0015] A three-channel parallel feature extraction module is constructed using a frequency trajectory adaptive deformable convolutional network and a one-dimensional convolutional neural network;
[0016] A multimodal epileptic seizure recognition module was constructed by introducing a causal-perceptive temporal attention mechanism and a multi-scale temporal network.
[0017] A deep learning model is formed by using the three-channel parallel feature extraction module and the multimodal epilepsy seizure recognition module.
[0018] Optionally, a three-channel parallel feature extraction module is constructed using a frequency trajectory adaptive deformable convolutional network and a one-dimensional convolutional neural network, including:
[0019] Analyze the frequency trajectory changes of radar-collected signals during clonic episodes in children and the characteristic changes during myoclonic seizures;
[0020] Based on the frequency trajectory change, an offset prediction sub-network is introduced into the original deformable convolutional network on the basis of the sampling position of the standard two-dimensional convolution, to construct a frequency trajectory adaptive deformable convolutional network.
[0021] A three-channel parallel feature extraction module is constructed, the deformable convolutional network is introduced into the body motion channel, the receptive field of the convolution is updated, and a multi-scale transient detection branch is introduced based on the feature changes.
[0022] A one-dimensional convolutional neural network is introduced into the three-channel parallel feature extraction module;
[0023] This invention introduces a causal-aware temporal attention mechanism and a multi-scale temporal network to construct a multimodal epileptic seizure recognition module. The module includes: constructing an improved causal-aware temporal attention mechanism based on the temporal causal propagation law of the signals corresponding to the three modal components during an epileptic seizure, wherein the temporal causal propagation law is used to determine the order constraints for weight calculation; constructing a multi-scale temporal network based on a feature capture mechanism with different time windows; constructing the multimodal epileptic seizure recognition module based on the temporal attention mechanism, the multi-scale temporal network, and a hierarchical classification head; training the multimodal epileptic seizure recognition module with preset sample data, setting a focus adjustment factor using a hierarchical aware focus loss function, and applying consistency penalties to the logical conflict prediction results of different levels of output using KL divergence, updating parameters until the module training is complete.
[0024] Optionally, the three-channel parallel feature extraction module is used to extract component feature vectors corresponding to the three modal components from the three modal components respectively, and auxiliary feature vectors are extracted from the auxiliary information, including:
[0025] After the three modal components are processed by bandpass filtering, they are combined with the auxiliary information and input into the three-channel parallel feature extraction module.
[0026] Perform feature extraction, transform the body motion components to generate a temporal spectrogram, and extract body motion feature vectors based on the deformable convolutional network;
[0027] For the respiratory component and heart rate component, the corresponding respiratory feature vector and heart rate feature vector are extracted using the one-dimensional convolutional neural network;
[0028] The auxiliary feature vector is extracted from the temperature change distribution information and acoustic event information of the auxiliary information using the one-dimensional convolutional neural network.
[0029] The body motion feature vector, the respiratory feature vector, and the heart rate feature vector together constitute the component feature vector.
[0030] Optionally, the motion components are transformed to generate a time-spectrum image, and the motion feature vector is extracted based on the deformable convolutional network, including:
[0031] The body motion component is subjected to short-time Fourier transform processing to generate a time-frequency spectrum, and a spectral fingerprint database is constructed. The fingerprint database serves as a benchmark reference for feature extraction and analysis.
[0032] Based on the time-spectrum map, the body motion feature map is extracted through the deformable convolutional network, and the broadband transient impact feature is extracted through short-time wavelet transform;
[0033] According to preset weights, the body motion feature map and the transient impact feature are weighted and residually fused to obtain the body motion feature vector.
[0034] Optionally, based on the temporal spectrogram, the body motion feature map is extracted through the deformable convolutional network, including:
[0035] Based on the time-spectrum diagram, a dynamic offset field in the frequency axis direction is generated through the offset prediction subnetwork of the deformable convolutional network, so that the convolutional receptive field extends along the frequency trajectory of the clonic period, and the offset field is related to the frequency trajectory in the horizontal direction.
[0036] Based on the offset field, bilinear interpolation is performed on the sampling points of the standard two-dimensional convolution to extract the motion feature map;
[0037] Specifically, broadband transient impact features are extracted through short-time wavelet transform, including: for short-time transient features of myoclonic seizures, the multi-scale transient detection branch of the three-channel parallel feature extraction module performs multi-scale short-time wavelet scattering transform on the body motion component within a preset time window to extract broadband transient impact features.
[0038] Optionally, in the multimodal epileptic seizure recognition module, a weighted fusion based on the component feature vectors and auxiliary feature vectors is performed through an attention mechanism, including:
[0039] In the multimodal epileptic seizure recognition module, the input component feature vector and auxiliary feature vector are concatenated to obtain a fused feature vector;
[0040] A causal-aware temporal attention mechanism is adopted. Based on the fused feature vector, a one-way mask is applied in the time dimension to constrain the flow of the component feature vectors in sequence, so that the calculation of attention weights follows the order constraint.
[0041] Bidirectional global attention is retained in both the channel and spatial dimensions, and the fused feature vector is weighted to obtain weighted fused features.
[0042] Optionally, epileptic seizure information can be output through multi-scale temporal network modeling, including:
[0043] In the multi-scale temporal network, sliding windows at different time scales are modeled, and the weighted fusion features are analyzed to detect the epileptic seizure characteristics under different time windows.
[0044] Based on the seizure characteristics, a multi-level hierarchical classification is performed, and the epileptic seizure type and confidence level are output as epileptic seizure information.
[0045] Specifically, based on the epileptic seizure information, loss analysis is performed using hierarchical perception focus loss and KL divergence, and the module is updated.
[0046] Optionally, after outputting the epileptic seizure information, the following may also be included:
[0047] The signals corresponding to the three modal components were subjected to energy drop detection, inter-respiratory period coefficient of variation detection, and heart rate change amplitude detection, respectively, to obtain the detection results;
[0048] When the test results meet the preset conditions, it is determined that an absence seizure has occurred;
[0049] For the absence seizure, the confidence level of the epileptic seizure signal is updated, and the updated confidence level is used to improve the reliability of the alarm determination.
[0050] Before performing mode decomposition on the reflected echo signals of children acquired by millimeter-wave radar, the process includes: selecting a corresponding normal range template and seizure feature template based on the child's age parameters; and using the normal range template and seizure feature template, adaptively adjusting the parameters of mode decomposition, the normalization factor of the spectrum in the three-channel parallel feature extraction module, and the classification threshold of seizures in the multimodal epileptic seizure recognition module, respectively.
[0051] Secondly, this application provides a multimodal non-contact detection system for childhood epilepsy seizures, comprising:
[0052] The perception layer is used to acquire the reflected echo signal of the child collected by the millimeter-wave radar, as well as to acquire auxiliary information collected by auxiliary sensors.
[0053] The signal processing layer is used to perform mode decomposition on the reflected echo signal to obtain three-mode components, which include body motion component, respiratory component and heart rate component.
[0054] The feature extraction layer is used to construct a deep learning model consisting of a three-channel parallel feature extraction module and a multimodal epileptic seizure recognition module. In the deep learning model, the three-channel parallel feature extraction module is used to extract the component feature vectors corresponding to each component from the three modal components and to extract auxiliary feature vectors from the auxiliary information.
[0055] The fusion recognition layer is used in the multimodal epileptic seizure recognition module to perform weighted fusion based on the component feature vector and auxiliary feature vector through an attention mechanism, and to output epileptic seizure information through multi-scale temporal network modeling.
[0056] An output layer is applied to select the corresponding output mode and output alarm information when the epileptic seizure information meets the alarm conditions.
[0057] In summary, this embodiment of the application performs modal decomposition on the radar signals collected by non-contact millimeter-wave radar for detecting children, obtaining three modal component signals of body movement, respiration, and heart rate, along with temperature and acoustic event information collected by auxiliary sensors. Using each component signal and information as input, a three-channel parallel feature extraction module extracts the time-spectrum map in a deep learning model, and convolution yields feature vectors for the three components and auxiliary feature vectors, serving as four feature vectors related to epileptic seizures. A causal-aware temporal attention mechanism is introduced to weightedly fuse these four feature vectors. Multi-scale temporal network modeling is used to model and weightedly fuse the features. Epileptic seizure characteristics are detected through windows at different time scales, and multi-level hierarchical classification is used for identification, outputting the epileptic seizure type and confidence level. Finally, based on the epileptic seizure information output by the model, an alarm is output when the alarm conditions are met. Therefore, this embodiment of the application achieves non-contact real-time detection and classification of all types of epileptic seizures in children. Based on the characteristics of epileptic seizure manifestations, the detection model is improved to achieve multi-modal detection and improve the detection accuracy of epileptic seizures. Attached Figure Description
[0058] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.
[0059] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0060] Figure 1 A flowchart illustrating a multimodal non-contact detection method for childhood epilepsy seizures provided in this application embodiment;
[0061] Figure 2 This is an example diagram of an FMCW radar signal processing and three-channel separation process provided in this application;
[0062] Figure 3 This is an example diagram of a deep learning seizure recognition model architecture provided in this application;
[0063] Figure 4 This is a schematic diagram of Micro-Doppler characteristic signals for different types of epileptic seizures, provided as an example in this application;
[0064] Figure 5 This application provides an example of a hospital + home dual-scenario deployment diagram;
[0065] Figure 6 This is a schematic diagram of an age-adaptive epileptic seizure feature template system provided as an example in this application;
[0066] Figure 7 This is a schematic diagram of the overall system architecture for multimodal non-contact detection of childhood epilepsy seizures, provided in an optional embodiment of this application. Detailed Implementation
[0067] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0068] To facilitate understanding of the embodiments of this application, further explanations and descriptions will be provided below in conjunction with the accompanying drawings and specific embodiments. These embodiments do not constitute a limitation on the embodiments of this application.
[0069] Figure 1 The flowchart illustrates a multimodal non-contact detection method for childhood epilepsy seizures provided in this application embodiment. This method is applicable to real-time, non-contact, multi-dimensional intelligent monitoring and early warning of all types of epilepsy seizures in both hospital and home settings. For example, to detect all types of epilepsy seizures in children aged 0-14 years, the method may specifically include the following steps:
[0070] Step 110: Perform mode decomposition on the reflected echo signal of the child collected by the millimeter-wave radar to obtain three mode components, and acquire auxiliary information collected by auxiliary sensors.
[0071] The three modal components include body motion component, respiratory component, and heart rate component.
[0072] For example, a millimeter-wave radar sensor can be used for non-contact physiological signal detection and acquisition. The millimeter-wave radar sensor can operate in the 60 GHz ISM band and includes a frequency modulated continuous wave (FMCW) radar chip, configured to transmit frequency modulated continuous wave signals to children aged 0-14 years at a distance of 0.5 to 3 meters and receive reflected echoes.
[0073] The acquired reflected echo signal can be processed in four stages to separate the three-mode components. Specifically, MTI (Moving Target Indication) static clutter suppression, range FFT (Fast Fourier Transform), and Doppler FFT are performed sequentially to obtain a range-Doppler map. The phase signal is extracted from the target range gate and separated into three channels—body motion, respiration, and heart rate—using seizure-guided adaptive variational mode decomposition (VMD), resulting in three-mode components. The interpretable three modes include body motion, respiration, and heart rate. The VMD parameters can be adaptively adjusted according to real-time signal characteristics; for example, when a suspected pre-seizure signal is detected, the VMD parameters can be dynamically adjusted to optimize the signal separation accuracy during the seizure.
[0074] As an example, in conjunction with reference Figure 2 The diagram illustrates the radar signal processing and three-channel separation process in detail:
[0075] First, static environmental echoes are eliminated through MTI clutter suppression.
[0076] Then, a Range-Doppler Map is constructed using distance FFT and Doppler FFT.
[0077] Subsequently, phase extraction processing is performed, and sub-millimeter-level displacement detection is achieved by utilizing the relationship between body surface displacement and phase change. The body surface displacement can be expressed as... The phase change relationship can be expressed as: .
[0078] Next, mode decomposition is performed, decomposing the phase signal into three channels—body movement, respiration, and heart rate—through adaptive VMD signal separation. This embodiment improves the VMD signal separation process and proposes an adaptive VMD strategy based on seizure guidance. Specifically, the VMD parameters are determined by the number of modes. and penalty factor Composition, using standard VMD parameters (such as...) during normal monitoring. , When a suspected pre-seizure warning sign is detected (such as a sudden increase in body energy, abnormal breathing, or abnormal heart rate), dynamic adjustments are made. and The value is used to optimize the accuracy of three-channel signal separation during an attack.
[0079] To efficiently and accurately acquire reflected echo signals and improve the accuracy of subsequent epileptic seizure identification, this example provides relevant radar chip selection and system parameters for reference only. Please refer to Table 1 below for details:
[0080] parameter TI IWR6843ISK Infineon BGT60TR13C TI IWR1843 Recommended configuration frequency 60-64GHz 58-63.5GHz 76-81GHz 60GHz ISM bandwidth 4GHz 5.5GHz 5GHz 4GHz TX / RX 3TX / 4RX 1TX / 3RX 3TX / 4RX 3TX / 4RX Distance resolution ~3.75cm ~2.7cm ~3.0cm 3.75cm Frame rate ≤20fps Configurable ≤20fps ≥20fps Detection distance 0.5-3m 0.3-1.5m 0.5-3m 0.5-3m
[0081] Table 1
[0082] Among them, the TI IWR6843ISK is the preferred choice. This chipset has been clinically validated for effective detection of GTCS seizures, and its 3TX / 4RX MIMO configuration provides better spatial resolution, which is beneficial for unilateral motor recognition of focal seizures.
[0083] To achieve multimodal epileptic seizure detection, this embodiment introduces an auxiliary sensor, consisting of a low-resolution infrared thermal imaging array and a MEMS microphone array, configured for non-contact detection of changes in body surface temperature distribution and acoustic events, respectively, using the temperature change distribution and acoustic events as auxiliary information. This auxiliary information can be combined with reflected echo signals to form a multi-source signal.
[0084] This embodiment's non-contact detection solution is completely contactless and child-friendly, requiring no sensors to be worn or attached to children. It addresses the pain point of poor compliance with wearable devices for young children, making it especially suitable for infants aged 0-3 years. Electromagnetic safety is ensured: the 60GHz radar has a power density of approximately 0.013 W / m² at a distance of 1 meter, only about 1 / 700th of the ICNIRP safety limit, ensuring safety for children.
[0085] Furthermore, through multi-sensor fusion enhancement, radar + infrared + audio three-modal fusion can be achieved, overcoming the limitation of a single sensor's insufficient ability to detect non-motor seizures.
[0086] Understandably, this application can also be used to address issues including: how to achieve real-time detection and classification of all types of epileptic seizures in children aged 0-14 years using non-contact methods, covering generalized seizures (tonic-clonic, myoclonic, absence seizures, etc.) and focal seizures under the ILAE 2017 classification; how to improve the detection capability of non-motor seizures (especially absence seizures) through multi-sensor fusion; how to establish an age-adaptive detection model covering infancy to school age, adapting to the differences in physiological parameters and seizure characteristics of children of different ages; and how to achieve flexible deployment of the same system in both hospital ICU and home bedroom scenarios. Detailed explanations of each step and optional embodiments will follow.
[0087] Step 120: Construct a deep learning model consisting of a three-channel parallel feature extraction module and a multimodal epileptic seizure recognition module.
[0088] In the specific implementation, refer to Figure 3 As shown, the deep learning model employs a hybrid architecture combining four-channel input, causal-aware temporal attention mechanism fusion, and multi-scale temporal modeling as the detection model. It primarily consists of a three-channel parallel feature extraction module (hereinafter referred to as the parallel feature extraction module) and a multimodal epileptic seizure recognition module (hereinafter referred to as the epileptic seizure recognition module). By analyzing the frequency changes of childhood epileptic seizures during the clonic and myoclonic phases, the three-channel parallel feature extraction module was improved. Furthermore, the multimodal epileptic seizure recognition module was improved based on the temporal causal propagation patterns of body movement, respiration, and heart rate signals during childhood epileptic seizures. This embodiment, through structural and algorithmic improvements tailored to the physiological characteristics of epileptic seizures, effectively enhances the model's recognition accuracy and robustness for different seizure types.
[0089] Optionally, this embodiment constructs a deep learning model consisting of a three-channel parallel feature extraction module and a multimodal epileptic seizure recognition module, which may include: constructing a three-channel parallel feature extraction module using a frequency trajectory adaptive deformable convolutional network and a one-dimensional convolutional neural network; introducing a causal-aware temporal attention mechanism and a multi-scale temporal network to construct a multimodal epileptic seizure recognition module; and using the three-channel parallel feature extraction module and the multimodal epileptic seizure recognition module to form a deep learning model.
[0090] In practical implementation, the above-mentioned frequency trajectory adaptive deformable convolutional network and one-dimensional convolutional neural network are used to construct a three-channel parallel feature extraction module, which can specifically include: analyzing the frequency trajectory changes of radar-acquired signals during clonic episodes and the feature changes during myoclonic seizures; based on the frequency trajectory changes, an offset prediction sub-network is introduced into the original deformable convolutional network on the basis of the sampling position of the standard two-dimensional convolution to construct a frequency trajectory adaptive deformable convolutional network; a three-channel parallel feature extraction module is constructed, the deformable convolutional network is introduced into the body movement channel to update the convolutional receptive field, and a multi-scale transient detection branch is introduced based on the feature changes; a one-dimensional convolutional neural network is introduced into the three-channel parallel feature extraction module.
[0091] In practical implementation, the aforementioned introduction of a causal-aware temporal attention mechanism and a multi-scale temporal network to construct a multimodal epileptic seizure recognition module can specifically include: constructing an improved causal-aware temporal attention mechanism based on the temporal causal propagation law of the signals corresponding to the three modal components during an epileptic seizure, wherein the temporal causal propagation law is used to determine the order constraints for weight calculation; constructing a multi-scale temporal network based on a feature capture mechanism for different time windows; constructing a multimodal epileptic seizure recognition module based on the temporal attention mechanism and the multi-scale temporal network, combined with a hierarchical classification head; training the multimodal epileptic seizure recognition module with preset sample data, setting a focus adjustment factor using a hierarchical aware focus loss function, and applying consistency penalties to the logical conflict prediction results of different levels of output using KL divergence, updating parameters until the module training is complete.
[0092] The following details the structural improvements to the three-channel parallel feature extraction module and the multimodal epileptic seizure recognition module:
[0093] Specifically, refer to Figure 3 As shown in the figure, this embodiment analyzes and finds that tonic-clonic seizures in children have a typical frequency sliding trajectory that decreases from 4-8Hz to 1Hz during the clonic phase, and myoclonic seizures are characterized by broadband transient impacts with extremely short durations. Traditional convolutional networks with fixed receptive fields cannot adapt to these specific manifestations, resulting in limited feature extraction effects.
[0094] Therefore, to improve feature extraction performance and address the limited receptive field issue inherent in traditional fixed convolutional networks, this embodiment introduces a frequency trajectory adaptive deformable convolutional network for the body motion channel in the three-channel parallel feature extraction module. Furthermore, an offset prediction sub-network is introduced to generate the offset field, extending the receptive field of the convolutional network along the frequency-decreasing trajectory of the clonic phase. Simultaneously, a multi-scale transient detection branch is introduced into the three-channel parallel feature extraction module to capture myoclonic impulses. In addition, a one-dimensional convolutional neural network is introduced for feature capture of the remaining modal components and auxiliary information. Thus, this embodiment combines the frequency trajectory adaptive deformable convolutional network, the offset prediction sub-network, the multi-scale transient detection branch, and the one-dimensional convolutional neural network to improve the three-channel parallel feature extraction module. This allows the module to analyze the complex changing features present during epileptic seizures and extract relevant features from other modal components and auxiliary information, efficiently capturing features of different seizure types and effectively improving the accuracy and robustness of subsequent seizure type identification.
[0095] In the multimodal fusion section, based on the temporal causal propagation patterns of body movement, respiration, and heart rate signals during seizures, the existing Cross-Modal Attention mechanism is improved into a causal-aware temporal attention mechanism (hereinafter referred to as the temporal attention mechanism) to constrain the flow of component feature vectors. Simultaneously, a multi-scale temporal modeling network and an ILAE hierarchical classification head are introduced to identify and output the seizure type and corresponding confidence level. The seizure types cover generalized and focal seizures under the ILAE 2017 classification criteria.
[0096] To further improve the model's recognition accuracy and robustness, this embodiment incorporates a hierarchical perceptual focus loss calculation unit within the epilepsy seizure recognition module. This unit includes improved loss functions such as the hierarchical perceptual focus loss function and KL divergence. During training, based on the output results, the hierarchical perceptual focus loss function is used to set focus adjustment factors, calculating loss values for the top-level binary classification of generalized and focal seizures, as well as for the fine-grained classification of each subtype. The loss weight for the top-level classification is higher than that for the fine-grained classification, and the focus adjustment factors for each classification head are set independently, with the adjustment factor corresponding to the seizure category having a higher value than that for the normal category.
[0097] KL divergence can apply a consistency penalty to logical conflict prediction results at different levels. When there is an ILAE-level logical conflict between the top-level classification result and the fine-grained classification result, an additional hierarchical consistency penalty term based on the KL divergence of the probability distributions of the two classifier heads is added. The initial weight coefficient of the penalty term is 0.5 and is dynamically updated according to the proportion of hierarchical conflict samples during the training cycle. The parameters are continuously updated during the training of each loss function sub-model until the model training is completed, and then applied to the multimodal non-contact detection scenario of childhood epilepsy seizures.
[0098] Therefore, through the above-mentioned structural and algorithmic improvements targeting the physiological characteristics of epileptic seizures, the detection model in this embodiment can effectively achieve multimodal detection and improve the detection accuracy of epileptic seizures.
[0099] Step 130: In the deep learning model, the three-channel parallel feature extraction module is used to extract the component feature vectors corresponding to each component from the three modal components, and to extract auxiliary feature vectors from the auxiliary information.
[0100] In the specific implementation, the three modal components and auxiliary information are input into the deep learning model, which then calls the parallel feature extraction module. Within this module, the body motion component is transformed to generate a Micro-Doppler time-frequency spectrogram, and a frequency trajectory adaptive deformable convolutional network combined with a multi-scale transient detection branch is used to extract body motion feature vectors. For the respiratory and heart rate components, a one-dimensional convolutional neural network is used to extract feature vectors. For the infrared (including temperature distribution) and audio signals (including sound events) in the auxiliary information, a one-dimensional convolutional neural network is used to extract auxiliary feature vectors. Finally, the component feature vectors and auxiliary feature vectors corresponding to the three components are obtained, forming a four-way feature vector.
[0101] Optionally, the above-described method of using the three-channel parallel feature extraction module to extract component feature vectors corresponding to the three modal components from the three modal components, and extracting auxiliary feature vectors from the auxiliary information, may include the following sub-steps:
[0102] Sub-step 1301: After the three-mode components are processed by bandpass filtering, they are combined with the auxiliary information and input into the three-channel parallel feature extraction module.
[0103] In practical implementation, bandpass filtering can be uniformly applied to all three modal components, referring to... Figure 2 As shown, the body movement channel can be processed by a 0.5-10Hz bandpass filter; the breathing channel can be processed by a 0.1-1.5Hz bandpass filter; and the heart rate channel can be processed by a 1.5-4.0Hz bandpass filter.
[0104] The components processed by bandpass filtering can be combined with auxiliary information and input into the parallel feature extraction module.
[0105] Sub-step 1302 involves performing feature extraction, transforming the motion components to generate a temporal spectrum, and extracting motion feature vectors based on the deformable convolutional network.
[0106] In this embodiment, the body motion component is subjected to a Short-Time Fourier Transform (STFT) to generate a Micro-Doppler time-frequency spectrum, where the STFT transformation parameters can be set. The time-frequency spectrum is then used to extract features from the frequency trajectory adaptive deformable convolutional network to obtain the body motion feature vector.
[0107] In one optional embodiment, transforming the body motion components to generate a time-spectrum image and extracting body motion feature vectors based on the deformable convolutional network may include: performing a short-time Fourier transform on the body motion components to generate a time-spectrum image, constructing a spectral fingerprint database, the fingerprint database serving as a benchmark reference for feature extraction analysis; extracting a body motion feature map based on the time-spectrum image using the deformable convolutional network, and extracting broadband transient impact features using a short-time wavelet transform; and performing weighted residual fusion of the body motion feature map and the transient impact features according to preset weights to obtain a body motion feature vector.
[0108] In practical implementation, this embodiment extracts body motion feature maps based on the time-spectrum map using the deformable convolutional network. Specifically, it may include: generating a dynamic offset field along the frequency axis direction using the offset prediction subnetwork of the deformable convolutional network based on the time-spectrum map, so that the receptive field of the convolution extends along the frequency trajectory of the clonic period, and the offset field is related to the frequency trajectory in the horizontal direction; and performing bilinear interpolation correction on the sampling points of the standard two-dimensional convolution based on the offset field to extract the body motion feature map.
[0109] In practical implementation, this embodiment extracts broadband transient impact features through short-time wavelet transform. Specifically, it can include: for the short-time transient features of myoclonic seizures, the multi-scale transient detection branch of the three-channel parallel feature extraction module performs multi-scale short-time wavelet scattering transform on the body motion component within a preset time window to extract broadband transient impact features.
[0110] The following combination Figure 2 , 3 And 4. A detailed explanation of the process for extracting the body motion feature vector in this application:
[0111] Specifically, a 128×128 Micro-Doppler time spectrum can be generated through STFT processing. The STFT transformation parameters can be: window length of 256 points and overlap of 75%.
[0112] By analyzing the time-spectral graphs, a Micro-Doppler spectral fingerprint database specific to different seizure types can be established. This includes tonic-clonic seizures characterized by a rhythmic modulation band decreasing from 4-8 Hz to 1 Hz during the clonic phase; myoclonic seizures characterized by broadband transient pulses lasting less than 100 milliseconds; focal motor seizures characterized by spatially asymmetrical unilateral Micro-Doppler energy characteristics; and absence seizures characterized by a sudden drop in motor signal energy exceeding 80% leading to cessation of movement. Data from the spectral fingerprint database can be combined to analyze the characteristic manifestations of childhood epileptic seizures, serving as a reference benchmark for feature extraction and model improvement.
[0113] Next, using the time-frequency spectrum as a reference, a frequency trajectory adaptive deformable convolutional network is employed to extract body motion feature maps. When extracting body motion feature maps, this deformable convolutional network, based on the sampling positions of standard two-dimensional convolution, generates a dynamic offset field along the frequency axis direction through an offset prediction subnetwork. This offset field can be adaptively calculated based on the local gradient of the frequency trajectory in the current time-frequency spectrum, causing the convolutional receptive field to dynamically extend along the frequency trajectory direction that decreases from 4-8 Hz to 1 Hz during the clonic phase of a tonic-clonic seizure.
[0114] Furthermore, the offset prediction subnetwork takes the time-spectrum map as input, and after two layers of convolution operations, outputs a two-dimensional dynamic offset field matching the size of the backbone convolution kernel. The offset value at each position in this offset field corresponds to the slope estimate of the frequency trajectory relative to the horizontal direction at that position. The deformable convolution performs bilinear interpolation correction on each sampling point based on the offset field to complete the convolution operation and outputs a volumetric feature map. Notably, the offset prediction subnetwork and the backbone deformable convolution network are jointly trained end-to-end, eliminating the need for manual labeling of the frequency trajectory.
[0115] While acquiring the body motion feature map, the multi-scale transient detection branch set in the feature extraction module can simultaneously extract broadband transient impact features from the body motion components. Specifically, this branch can extract broadband transient impact features with a duration of less than 100 milliseconds from the body motion phase signal using a short-time wavelet scattering transform with a time window of no more than 30 milliseconds.
[0116] Finally, the transient features and feature maps are fused using a weighted residual method. The initial value of the fusion weight coefficient can be zero and is adaptively adjusted during training. The fused body motion feature vector has a dimension of 256.
[0117] Reference Figure 4 The example shown illustrates the improved concept of feature extraction in this application and explains the correlation between the features extracted by this application and actual epileptic seizures. The following is a detailed description of the Micro-Doppler feature signals for different types of epileptic seizures:
[0118] The Micro-Doppler signature consists of characteristic features of generalized tonic-clonic seizures (GTCS) and myoclonic seizures.
[0119] GTCS is the most characteristic seizure type. The tonic phase is characterized by generalized muscle rigidity lasting 10-20 seconds, with a broadband sustained energy on the time-spectral map. This is followed by the clonic phase, characterized by rhythmic limb jerks with an initial frequency of 4-8 Hz, gradually decreasing in frequency to approximately 1 Hz before ceasing, lasting from 30 seconds to 2 minutes. This frequency decrease pattern creates a distinctive "frequency sliding" characteristic on the radar time-spectral map, which is the core basis for GTCS detection.
[0120] Myoclonic seizures are characterized by extremely brief (less than 100 milliseconds) sudden muscle twitches, appearing as broadband transient pulses on a Micro-Doppler spectrum. This type of seizure requires a frame rate of at least 20 fps for reliable capture. Absence seizures are characterized by loss of consciousness accompanied by behavioral arrest, a sudden drop in motor signal energy exceeding 80%, and are accompanied by abrupt changes in respiratory pattern and subtle changes in heart rate, forming the "autonomic micro-feature triangle" defined in this invention. Focal motor seizures are characterized by rhythmic twitching of one limb, which can be identified in multi-channel radar systems through spatially asymmetrical Micro-Doppler energy distribution.
[0121] Sub-step 1303: Extract the corresponding respiratory feature vector and heart rate feature vector from the respiratory component and heart rate component using the one-dimensional convolutional neural network.
[0122] Sub-step 1304: Extract auxiliary feature vectors from the temperature change distribution information and acoustic event information of the auxiliary information through the one-dimensional convolutional neural network.
[0123] The body motion feature vector, the respiratory feature vector, and the heart rate feature vector together constitute the component feature vector.
[0124] A unified explanation is provided for sub-steps 1303-1304:
[0125] In this embodiment, the parallel feature extraction module extracts relevant feature vectors from the respiratory component, heart rate component, and auxiliary information using a one-dimensional convolutional neural network. Specifically, respiratory feature vectors such as respiratory rate, amplitude, regularity, and pause events can be extracted from the respiratory component; heart rate feature vectors such as heart rate, HRV index, and heart rate mutation events can be extracted from the heart rate component. Actual research has shown that during an epileptic seizure, the respiratory rate decreases by an average of 18%, while the heart rate increases by approximately 6%, and these heart rate changes can occur approximately 13 seconds before the onset of the EEG seizure.
[0126] Among them, the respiratory feature vector is 128-dimensional, the heart rate feature vector is 128-dimensional, and the auxiliary feature vector is 64-dimensional.
[0127] Step 140: In the multimodal epileptic seizure recognition module, the component feature vector and auxiliary feature vector are weighted and fused through an attention mechanism, and epileptic seizure information is output through multi-scale temporal network modeling.
[0128] In related technologies, some existing techniques introduce attention mechanisms, such as Cross-Modal Attention, into multimodal information classification. Through an 8-head Cross-Modal Attention mechanism, it can learn the interaction relationships between modalities, such as interaction weights, to achieve weighted fusion of feature vectors. However, existing Cross-Modal Attention mechanisms are not suitable for addressing the characteristics of tonic-clonic and myoclonic manifestations present in epileptic seizures.
[0129] To improve the accuracy of epileptic seizure recognition and classification, this embodiment improves the Cross-Modal Attention mechanism by proposing a causal-aware temporal attention mechanism to enhance the seizure recognition module. Specifically, after concatenating the extracted four feature vectors, a causal constraint mask is applied through the causal-aware temporal attention mechanism, ensuring that the calculation of attention weights follows a specific order. The weighted fusion of each feature vector is then combined to obtain the corresponding weighted fused features.
[0130] The weighted fusion features are output to a multi-scale temporal network for modeling, capturing epileptic seizures. After hierarchical classification, the seizure type and corresponding confidence score (referred to as confidence score) are output as epileptic seizure information. The seizure type may include, but is not limited to: tonic-clonic seizures, myoclonic seizures, absence seizures, and focal motor seizures.
[0131] In an optional embodiment, the weighted fusion based on the component feature vectors and auxiliary feature vectors through an attention mechanism in the multimodal epileptic seizure recognition module may include: concatenating the input component feature vectors and auxiliary feature vectors to obtain a fused feature vector; employing a causal-aware temporal attention mechanism, applying a one-way mask in the time dimension based on the fused feature vector to constrain the flow of the component feature vectors in sequence, so that the calculation of attention weights follows the sequence constraint; retaining bidirectional global attention in the channel and spatial dimensions, and performing weighted processing on the fused feature vector to obtain a weighted fused feature.
[0132] In an optional embodiment, the above-mentioned modeling of epileptic seizure information through a multi-scale temporal network can specifically include: modeling sliding windows at different time scales in the multi-scale temporal network, analyzing the weighted fusion features, and detecting seizure characteristics under different time windows; performing multi-level hierarchical classification based on the seizure characteristics, and outputting the epileptic seizure type and confidence level as epileptic seizure information; wherein, based on the epileptic seizure information, loss analysis is performed using hierarchical perceptual focus loss and KL divergence, and the module is updated.
[0133] The following combination Figure 4 The process of identifying and classifying epileptic seizures by the seizure recognition module is described in detail below:
[0134] In this embodiment, the four feature vectors are concatenated to obtain a 576-dimensional fusion vector, which is then input into the causal perception temporal attention mechanism.
[0135] The causal-aware temporal attention mechanism processes the fusion vector in the temporal, channel, and spatial dimensions. Specifically, a unidirectional causal constraint mask is applied in the temporal dimension. The application reference corresponds to the physiological propagation sequence of various signals during an epileptic seizure, from abnormal body movement to changes in respiration, and then to changes in heart rate. This ensures that the mask forces the calculation of attention weights to follow the temporal order of body movement features first, respiration features in the middle, and heart rate features last. Bidirectional global attention is retained in the channel and spatial dimensions, meaning the unidirectional mask is applied only in the temporal direction. Finally, the weighted fusion features obtained after weighted fusion are input into the multi-scale temporal model.
[0136] The multi-scale temporal model employs multi-scale temporal modeling, with its temporal modeling layer simultaneously using sliding windows (referred to as time windows) of at least three different time scales to identify epileptic seizures based on weighted fusion features. For example, short time windows (e.g., a 5-second short window) can be used to capture short-duration seizures such as myoclonus and absence seizures; medium time windows (e.g., a 30-second medium window) can be used to capture focal seizures; and long time windows (e.g., a 120-second long window) can be used to capture the complete process of tonic-clonic seizures. The epileptic seizure features detected by multi-scale temporal modeling can be used for classification.
[0137] To achieve accurate seizure classification, this embodiment introduces an ILAE hierarchical classification head. This head employs a three-level structure, outputting the corresponding seizure type and confidence level. The three-level structure includes: Level 1 distinguishing between generalized seizures, focal seizures, and seizures of unknown origin; Level 2 distinguishing between motor initiation and non-motor initiation; and Level 3 identifying specific seizure subtypes.
[0138] Furthermore, in practical applications, loss analysis can be performed using hierarchical perceptual focus loss and KL divergence to continuously update model parameters. This includes setting focus adjustment factors independently for each classifier head and, in IOLAE hierarchical classification, when logically conflicting prediction results occur, incorporating a hierarchical consistency penalty term based on KL divergence into the result.
[0139] In summary, this application makes improvements to the three-channel parallel feature extraction module and the multimodal fusion recognition module, including:
[0140] The main body motion channel employs frequency trajectory adaptive deformable convolution, while the offset prediction subnetwork generates a dynamic offset field based on the local gradient of the time-frequency map, extending the receptive field along the frequency-decreasing trajectory of the clonic phase. Simultaneously, a multi-scale transient detection branch is set up, using a 10-30 millisecond short-window wavelet scattering transform to capture myoclonic impulses, which are then fused with the main feature using a weighted residual with an initial weight of 0.1.
[0141] The multimodal fusion component replaces Cross-Modal Attention with causal-aware temporal attention. A unidirectional mask is applied in the temporal dimension, constraining information flow in the order of body movement → respiration → heart rate, while maintaining bidirectional global attention in the channel and spatial dimensions. During training, a hierarchical-aware focus loss is used, with weights applied to top-level and fine-grained classifications, and KL divergence is used to penalize hierarchical conflicts. In actual testing, these improvements significantly enhance the model's sensitivity to frequency slippage and transient events, resulting in better hierarchical classification consistency.
[0142] Furthermore, after outputting the epileptic seizure information, this embodiment may further include: performing energy drop detection, respiratory interval variation coefficient detection, and heart rate change amplitude detection on the signals corresponding to the three modal components respectively to obtain detection results; when the detection results meet preset conditions, determining that an absence seizure has occurred; and updating the confidence level of the epileptic seizure signal for the absence seizure, with the updated confidence level used to improve the reliability of the alarm determination.
[0143] Furthermore, based on the improvements made to each module in the above model, this embodiment specifically proposes a method for detecting the autonomic nervous system micro-feature triangle of absence seizures to further improve the accuracy of epileptic seizure detection and recognition (see [reference]). Figure 4 -c). This autonomic nervous system micro-feature triangle detection method includes auxiliary identification of absence seizures.
[0144] The criteria for auxiliary identification of absence seizures are: simultaneous detection of a sudden drop in motor energy, a sudden change in breathing pattern, and a slight change in heart rate in the body movement channel. When all three occur simultaneously within a preset time window, it is identified as an absence seizure.
[0145] For example, this can be achieved through simultaneous detection of the following three conditions: ① A sudden drop in body motion signal energy exceeding 80% (amplitude of arrest); ② A sudden increase in the coefficient of variation (CV) between respiratory intervals exceeding a preset threshold within a 10-second window (abrupt change in respiratory pattern); ③ A change in heart rate exceeding 3% or a sudden change in HRV (micro-change in heart rate). When all three conditions are met simultaneously within the same 10-second time window, the detection result is considered to meet the preset conditions, and the system can determine it as a suspected absence seizure. Since these three conditions rarely occur simultaneously under normal circumstances, this method has high theoretical specificity.
[0146] Therefore, this embodiment achieves full-type seizure coverage. By combining spectral fingerprinting and autonomic nerve micro-feature triangles, non-contact detection of all types of epileptic seizures, from tonic-clonic to absence seizures, can be realized.
[0147] Step 150: When the epileptic seizure information meets the alarm conditions, select the corresponding output mode to output the alarm information.
[0148] In the specific implementation, when the confidence score of an epileptic seizure exceeds a preset threshold, the epileptic seizure information is determined to meet the alarm conditions, and an alarm can then be triggered. The output mode is adaptively selected based on the deployment scenario, including either hospital mode or home mode. Specifically, when outputting alarm information, in hospital mode, the alarm is pushed to the HIS system and nurse station; in home mode, the alarm is pushed to the cloud platform and parents' mobile terminals, and video recording is triggered.
[0149] Reference Figure 5 The installation and deployment example shown in this embodiment provides the following installation solutions applicable to different modes:
[0150] The hospital version is installed on the ceiling or bracket, 1.5-3 meters away from the child. It integrates all sensors (radar + infrared + microphone) and interfaces with the hospital information system (HIS) via Ethernet or WiFi. It supports centralized monitoring of multiple beds, video EEG linkage verification, and nurse station alarm.
[0151] The home version offers two forms: a wall-mounted all-in-one unit that integrates all sensors and connects to the cloud platform via WiFi / 4G, with an estimated retail price of 1,500-2,500 yuan; and a desktop portable device that only contains a radar sensor, with the model quantized to below 2MB using INT8, supporting 4-8 hours of operation with a built-in battery, with an estimated retail price of 800-1,200 yuan.
[0152] Therefore, this embodiment enables flexible deployment in two scenarios, with the same core technology supporting both hospital ICU and home bedroom scenarios, filling the market gap for nighttime home epilepsy monitoring.
[0153] Furthermore, based on the identification and detection of epileptic seizures, this embodiment also proposes a method for analyzing epileptic seizure trends. The data obtained from the above steps can be used as long-term continuous monitoring data. Based on the long-term continuous monitoring data, the method analyzes indicators such as the frequency trend of epileptic seizures, nocturnal apnea events, and post-seizure heart rate recovery time to generate a sudden epileptic death risk score, providing a basis for clinical intervention.
[0154] In practical implementation, this embodiment is specifically designed for detecting epileptic seizures in children. Based on the aforementioned model improvements, further applicability improvements are proposed, primarily using the child's age as a benchmark to establish a template and refine the detection parameters. Specifically, before performing mode decomposition on the continuous wave signal of the child acquired by millimeter-wave radar, the process may further include: selecting a corresponding normal range template and seizure feature template based on the child's age parameters; and using the normal range template and seizure feature template to adaptively adjust the parameters of mode decomposition, the normalization factor of the spectrum in the three-channel parallel feature extraction module, and the classification threshold of epileptic seizures in the multimodal epileptic seizure recognition module, respectively.
[0155] Specifically, refer to Figure 6 As shown, there are significant differences in body shape, physiological parameters, and seizure characteristics among children of different ages. The age-adaptive mechanism of this invention includes:
[0156] Normal range template for physiological parameters: Baselines for respiratory rate and heart rate differ across age groups, requiring corresponding adjustments to the threshold for abnormality assessment. For example, children aged 0-14 years can be divided into four age groups for parameter adaptation, constructing templates as follows: 0-1 year old infancy: respiratory rate 30-60 breaths / min, heart rate 100-160 bpm, VMD modality number K=3; 1-3 year old toddlers: respiratory rate 24-40 breaths / min, heart rate 90-150 bpm, VMD modality number K=3-4; 3-6 year old preschool children: respiratory rate 22-34 breaths / min, heart rate 80-120 bpm, VMD modality number K=4; 6-14 year old school children: respiratory rate 18-30 breaths / min, heart rate 60-100 bpm, VMD modality number K=4-5.
[0157] Based on the input child's age parameters, the system selects the normal range template for physiological parameters and the epileptic characteristic template for the corresponding age group, and adaptively adjusts the VMD parameters, classification threshold, and Micro-Doppler spectral normalization factor, including:
[0158] VMD parameter adaptation: K=3 is used for simple movement patterns in infancy, and K=4-5 is used for complex movement patterns in school-age children.
[0159] Adaptive classification thresholds: Confidence thresholds for each seizure type are independently calibrated by age group.
[0160] Micro-Doppler spectral normalization: standardizes the amplitude of motion according to the body size scaling factor.
[0161] Therefore, this embodiment achieves age-adaptive capabilities, covering different age groups, and utilizes an adaptive parameter system to ensure that children of different ages obtain optimal detection performance.
[0162] In summary, the embodiments of this application mainly improve the deep learning model. Based on the specific performance of radar signals during the clonic phase of childhood epileptic seizures, the physiological characteristics of epileptic seizures are analyzed. Starting from the model structure and algorithm, the feature extraction module and the multimodal epileptic seizure recognition module of the model are improved respectively, which effectively improves the model's recognition accuracy and robustness for different seizure types and solves the problem of insufficient recognition and detection accuracy of epileptic seizures in the existing technology.
[0163] It should be noted that, for the sake of simplicity, the method embodiments are described as a series of actions. However, those skilled in the art should know that the embodiments of this application are not limited to the described order of actions, because according to the embodiments of this application, some steps may be performed in other orders or simultaneously.
[0164] like Figure 7 As shown in the illustration, this application also provides a multimodal non-contact detection system 700 for childhood epilepsy seizures, comprising:
[0165] The perception layer is used to acquire the reflected echo signal of the child collected by the millimeter-wave radar, as well as to acquire auxiliary information collected by auxiliary sensors.
[0166] The signal processing layer is used to perform mode decomposition on the reflected echo signal to obtain three-mode components, which include body motion component, respiratory component and heart rate component.
[0167] The feature extraction layer is used to construct a deep learning model consisting of a three-channel parallel feature extraction module and a multimodal epileptic seizure recognition module. In the deep learning model, the three-channel parallel feature extraction module is used to extract the component feature vectors corresponding to each component from the three modal components and to extract auxiliary feature vectors from the auxiliary information.
[0168] The fusion recognition layer is used in the multimodal epileptic seizure recognition module to perform weighted fusion based on the component feature vector and auxiliary feature vector through an attention mechanism, and to output epileptic seizure information through multi-scale temporal network modeling.
[0169] An output layer is applied to select the corresponding output mode and output alarm information when the epileptic seizure information meets the alarm conditions.
[0170] The perception layer comprises four modules: a millimeter-wave FMCW radar sensor module, which can be installed in children's activity areas to perform non-contact detection of children at a frame rate of no less than 20fps within a distance of 0.5-3 meters, collecting reflected echo signals; a low-resolution infrared thermal imaging array, using an 8×8 pixel thermopile array (such as Panasonic AMG8833), to non-contactly detect changes in body surface temperature distribution, assisting in the detection of abnormal body temperature during seizures and acquiring temperature distribution data; a MEMS microphone array, using a dual-microphone configuration to detect seizure-related acoustic events (such as fall impact sounds, abnormal breathing sounds, crying sounds, etc.); and an age adaptation module, which stores the normal range of physiological parameters and seizure characteristic templates corresponding to children of different ages.
[0171] The signal processing layer may include a signal processing module. The feature extraction layer may include a three-channel parallel feature extraction module. The fusion recognition layer may include a multimodal fusion seizure recognition module. The application output layer may include an alarm and output module.
[0172] Furthermore, the application output layer may also include a SUDEP risk warning module for generating a sudden epileptic death risk score.
[0173] Furthermore, in practical applications, this non-contact detection system supports two deployment models: hospital version and home version, corresponding to two output modes.
[0174] Taking hospital-based deployment as an example:
[0175] Reference Figure 5 As shown, a hospital-grade system was deployed in the Pediatric Epilepsy Monitoring Center of Hospital X. The radar module can be installed on the ceiling above the patient's bed, approximately 2 meters away. The infrared thermal imaging array and dual microphones are housed within the same module housing. The system interfaces with the hospital's information system via PoE Ethernet. Each bed is equipped with an independent radar module, and the real-time monitoring status of all beds is centrally displayed on a large screen at the nurses' station. When a suspected seizure event is detected, the system simultaneously sends an audible and visual alarm to the nurses' station, pushes a message to the on-duty doctor's mobile phone, and automatically triggers recording by the bedside video camera. The system output is compared and validated with the synchronously recorded video EEG data for continuous model optimization.
[0176] Taking home-based deployment as an example:
[0177] Home deployments can include wall-mounted all-in-one units and desktop portable units.
[0178] For example, a wall-mounted integrated device can be deployed in the bedroom of a child with epilepsy. The device is installed on the side or opposite wall of the child's bed, 1-2 meters away from the child. The device integrates radar, infrared, and microphone sensors, and resembles a smart speaker in appearance, so it will not frighten the child. The device connects to the cloud platform via the home's wireless network (WiFi). During daily monitoring, the system runs a full-version model (8MB) at the edge, performing real-time seizure detection. When a seizure is detected, an alarm notification (including seizure type, confidence level, start time, etc.) is immediately pushed to the parent's mobile application (APP), and simultaneously triggers the device's built-in camera to record a 30-second video clip for later review. All monitoring data is encrypted and uploaded to the cloud platform, allowing the attending physician to remotely view seizure logs and long-term trend analysis reports.
[0179] For the desktop portable example, the device is a budget-friendly solution, containing only a radar sensor module, approximately 8×8×5 cm in size and weighing no more than 200 grams. The device can be placed on a bedside table, 0.5-1.5 meters away from the child. It runs a simplified model quantized with INT8 (less than 2MB), deployed on an MCU via the TensorFlow Lite Micro framework, supporting USB-C power supply and a 3000mAh built-in lithium battery (approximately 4-8 hours of independent operation). Alarm push notifications are sent via BLE connection to a parent's mobile app. This version primarily targets GTCS and tonic-clonic seizures, with an expected detection sensitivity greater than 90%.
[0180] For example, this application provides a comparison of parameters for two scenarios as shown in Table 2 below for reference:
[0181] parameter Hospital version Home version (wall-mounted) Home version (portable) Detection distance 1.5-3m 1-2.5m 0.5-1.5m Installation method Ceiling / Helmet Wall-mounted / ceiling-mounted Desktop placement Power supply method Power over electricity (POE) AC adapter USB-C / Battery Communication methods Ethernet / WiFi WiFi / 4G WiFi / BLE sensor Radar + IR + Mic Radar + IR + Mic Radar only Model size Full version (8MB) Full version (8MB) Simplified version (<2MB) Estimated Costs ¥3000-5000 ¥1500-2500 ¥800-1200
[0182] Table 2
[0183] Electromagnetic safety: 60GHz EIRP 160mW, power density at 1m ~0.013W / m2, far below the ICNIRP limit of 10W / m2 (safety margin >700 times).
[0184] In practical applications, the expected performance indicators can be referred to in Table 3 below:
[0185] index Expected value Reference GTCS Sensitivity >93% Fearns et al. (2022) Radar Detection Verification GTCS specificity >85% Based on NightWatch clinical data Myoclonus sensitivity >80% Multimodal fusion enhancement Sensitivity of Absence Detection >70% The micro-feature triangle of the autonomic nervous system (to be verified) False alarm rate <0.5 times / hour Refer to the ILAE-IFCN guidelines and standards Detection delay <15 seconds Superior to wearable devices Model size (full version) <8MB Hospital version / Wall-mounted home version Model size (simplified version) <2MB TinyML Portable Version Inference delay <100ms / frame Real-time edge processing
[0186] Table 3
[0187] It should be noted that the system for multimodal non-contact detection of childhood epilepsy seizures provided in this application can execute the method provided in any embodiment of this application, and has the corresponding functions and beneficial effects of executing the multimodal non-contact detection method for childhood epilepsy seizures.
[0188] It should be noted that, in this document, relational terms such as "first" and "second" are used merely to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
[0189] The above description is merely a specific embodiment of this application, enabling those skilled in the art to understand or implement this application. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of this application. Therefore, this application is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features claimed herein.
Claims
1. A multimodal non-contact detection method for childhood epileptic seizures, characterized in that, include: The reflected echo signal of the child collected by the millimeter-wave radar is decomposed into three modal components to obtain auxiliary information collected by the auxiliary sensor. The three modal components include body motion component, respiratory component and heart rate component. A deep learning model consisting of a three-channel parallel feature extraction module and a multimodal epileptic seizure recognition module was constructed. In the deep learning model, the three-channel parallel feature extraction module is used to extract the component feature vectors corresponding to each component from the three modal components, and to extract auxiliary feature vectors from the auxiliary information; In the multimodal epileptic seizure recognition module, an attention mechanism is used to perform weighted fusion based on the component feature vectors and auxiliary feature vectors, and epileptic seizure information is output through multi-scale temporal network modeling. When the epileptic seizure information meets the alarm conditions, the corresponding output mode is selected to output the alarm information.
2. The method according to claim 1, characterized in that, A deep learning model consisting of a three-channel parallel feature extraction module and a multimodal epileptic seizure recognition module is constructed, including: A three-channel parallel feature extraction module is constructed using a frequency trajectory adaptive deformable convolutional network and a one-dimensional convolutional neural network; A multimodal epileptic seizure recognition module was constructed by introducing a causal-perceptive temporal attention mechanism and a multi-scale temporal network. A deep learning model is formed by using the three-channel parallel feature extraction module and the multimodal epileptic seizure recognition module.
3. The method according to claim 2, characterized in that, A three-channel parallel feature extraction module is constructed using a frequency trajectory adaptive deformable convolutional network and a one-dimensional convolutional neural network, including: Analyze the frequency trajectory changes of radar-collected signals during clonic episodes in children and the characteristic changes during myoclonic seizures; Based on the frequency trajectory change, an offset prediction sub-network is introduced into the original deformable convolutional network on the basis of the sampling position of the standard two-dimensional convolution, to construct a frequency trajectory adaptive deformable convolutional network. A three-channel parallel feature extraction module is constructed, the deformable convolutional network is introduced into the body motion channel, the receptive field of the convolution is updated, and a multi-scale transient detection branch is introduced based on the feature changes. A one-dimensional convolutional neural network is introduced into the three-channel parallel feature extraction module; This invention introduces a causal-aware temporal attention mechanism and a multi-scale temporal network to construct a multimodal epileptic seizure recognition module. This includes: constructing an improved causal-aware temporal attention mechanism based on the temporal causal propagation law of the signals corresponding to the three modal components during an epileptic seizure, whereby the temporal causal propagation law is used to determine the order constraints for weight calculation; constructing a multi-scale temporal network based on a feature capture mechanism with different time windows; constructing the multimodal epileptic seizure recognition module based on the temporal attention mechanism, the multi-scale temporal network, and a hierarchical classification head; training the multimodal epileptic seizure recognition module with preset sample data, setting a focus adjustment factor using a hierarchical aware focus loss function, and applying consistency penalties to the logical conflict prediction results of different levels of output using KL divergence, updating parameters until the module training is complete.
4. The method according to claim 2, characterized in that, Using the three-channel parallel feature extraction module, component feature vectors corresponding to the three modal components are extracted from the three modal components respectively, and auxiliary feature vectors are extracted from the auxiliary information, including: After the three modal components are processed by bandpass filtering, they are combined with the auxiliary information and input into the three-channel parallel feature extraction module. Perform feature extraction, transform the body motion components to generate a temporal spectrogram, and extract body motion feature vectors based on the deformable convolutional network; For the respiratory component and heart rate component, the corresponding respiratory feature vector and heart rate feature vector are extracted using the one-dimensional convolutional neural network; The auxiliary feature vector is extracted from the temperature change distribution information and acoustic event information of the auxiliary information using the one-dimensional convolutional neural network. The body motion feature vector, the respiratory feature vector, and the heart rate feature vector together constitute the component feature vector.
5. The method according to claim 4, characterized in that, The motion components are transformed to generate a time-spectrum image, and the motion feature vector is extracted based on the deformable convolutional network, including: The body motion component is subjected to short-time Fourier transform processing to generate a time-frequency spectrum, and a spectral fingerprint database is constructed. The fingerprint database serves as a benchmark reference for feature extraction and analysis. Based on the time-spectrum map, the body motion feature map is extracted through the deformable convolutional network, and the broadband transient impact feature is extracted through short-time wavelet transform; According to preset weights, the body motion feature map and the transient impact feature are weighted and residually fused to obtain the body motion feature vector.
6. The method according to claim 5, characterized in that, Based on the time-spectral map, the body motion feature map is extracted through the deformable convolutional network, including: Based on the time-spectrum diagram, a dynamic offset field in the frequency axis direction is generated through the offset prediction subnetwork of the deformable convolutional network, so that the convolutional receptive field extends along the frequency trajectory of the clonic period, and the offset field is related to the frequency trajectory in the horizontal direction. Based on the offset field, bilinear interpolation is performed on the sampling points of the standard two-dimensional convolution to extract the motion feature map; Specifically, broadband transient impact features are extracted through short-time wavelet transform, including: for short-time transient features of myoclonic seizures, the multi-scale transient detection branch of the three-channel parallel feature extraction module performs multi-scale short-time wavelet scattering transform on the body motion component within a preset time window to extract broadband transient impact features.
7. The method according to claim 1, characterized in that, In the multimodal epileptic seizure recognition module, a weighted fusion based on the component feature vectors and auxiliary feature vectors is performed through an attention mechanism, including: In the multimodal epileptic seizure recognition module, the input component feature vector and auxiliary feature vector are concatenated to obtain a fused feature vector; A causal-aware temporal attention mechanism is adopted. Based on the fused feature vector, a one-way mask is applied in the time dimension to constrain the flow of the component feature vectors in sequence, so that the calculation of attention weights follows the order constraint. Bidirectional global attention is retained in both the channel and spatial dimensions, and the fused feature vector is weighted to obtain weighted fused features.
8. The method according to claim 7, characterized in that, By modeling epileptic seizures using multi-scale temporal networks, information on the epileptic seizures is output, including: In the multi-scale temporal network, sliding windows at different time scales are modeled, and the weighted fusion features are analyzed to detect the epileptic seizure characteristics under different time windows. Based on the seizure characteristics, a multi-level hierarchical classification is performed, and the epileptic seizure type and confidence level are output as epileptic seizure information. Specifically, based on the epileptic seizure information, loss analysis is performed using hierarchical perception focus loss and KL divergence, and the module is updated.
9. The method according to claim 1, characterized in that, After outputting information about epileptic seizures, the following is also included: The signals corresponding to the three modal components were subjected to energy drop detection, inter-respiratory period coefficient of variation detection, and heart rate change amplitude detection, respectively, to obtain the detection results; When the test results meet the preset conditions, it is determined that an absence seizure has occurred; For the absence seizure, the confidence level of the epileptic seizure signal is updated, and the updated confidence level is used to improve the reliability of the alarm determination. Before performing mode decomposition on the reflected echo signals of children acquired by millimeter-wave radar, the process includes: selecting a corresponding normal range template and seizure feature template based on the child's age parameters; and using the normal range template and seizure feature template, adaptively adjusting the parameters of mode decomposition, the normalization factor of the spectrum in the three-channel parallel feature extraction module, and the classification threshold of seizures in the multimodal epileptic seizure recognition module, respectively.
10. A multimodal non-contact detection system for childhood epilepsy seizures, characterized in that, include: The perception layer is used to acquire the reflected echo signal of the child collected by the millimeter-wave radar, as well as to acquire auxiliary information collected by auxiliary sensors. The signal processing layer is used to perform mode decomposition on the reflected echo signal to obtain three-mode components, which include body motion component, respiratory component and heart rate component. The feature extraction layer is used to construct a deep learning model consisting of a three-channel parallel feature extraction module and a multimodal epileptic seizure recognition module. In the deep learning model, the three-channel parallel feature extraction module is used to extract the component feature vectors corresponding to each component from the three modal components and to extract auxiliary feature vectors from the auxiliary information. The fusion recognition layer is used in the multimodal epileptic seizure recognition module to perform weighted fusion based on the component feature vector and auxiliary feature vector through an attention mechanism, and to output epileptic seizure information through multi-scale temporal network modeling. An output layer is applied to select the corresponding output mode and output alarm information when the epileptic seizure information meets the alarm conditions.