Intelligent electrocardiogram abnormity prediction method and system based on video
By collecting facial information through video and using the FANMAMBA network model and ECG anomaly detection large model, the operational complexity of traditional ECG detection and the accuracy of rPPG signals are solved, and non-contact, efficient and accurate ECG anomaly prediction is achieved, which is suitable for health monitoring and family health management.
Patent Information
- Application Number
- CN202510801330.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-16
- Publication Date
- 2025-09-23
- Estimated Expiration
- 2045-06-16
AI Technical Summary
Existing ECG detection technology has the disadvantages of complex operation, high cost, reliance on contact electrodes which may lead to skin allergy risks, long detection time, and low detection efficiency, making it impossible to achieve long-term automated monitoring. In addition, existing rPPG technology has poor signal extraction accuracy in complex environments, affecting the reliability and practicality of the ECG detection system.
Facial information is collected through video, and the FANMAMBA network model and ECG anomaly detection large model are used to achieve non-contact ECG anomaly prediction. Combined with Fourier analysis and multi-scale feature fusion, high-quality ECG signals are generated and anomaly detection is performed.
It achieves efficient and accurate ECG signal monitoring, can extract cardiac abnormality features under non-contact conditions, reduces the amount of calculation, improves the real-time and accuracy of detection, and is suitable for health monitoring and family health management.
Smart Images

Figure CN120678442A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of biomedical signal processing technology, and specifically relates to an intelligent method and system for predicting electrocardiogram (ECG) anomalies based on video. This technology can be widely applied in fields such as health monitoring and engineering applications. Background Art
[0002] In healthcare, ECG testing is an important tool for monitoring heart disease. Traditional testing requires electrodes to be attached to the patient's body surface. This contact-based testing is not only complex and requires specialized personnel to perform, but also severely restricts the patient's freedom of movement.
[0003] In addition, traditional testing equipment is expensive and difficult to maintain, hindering its popularization in home health monitoring scenarios.
[0004] In recent years, portable contact ECG monitoring devices, such as smartwatches and handheld ECG devices, have become increasingly popular. While these devices enable the collection of ECG information anytime, anywhere, their practical application still requires direct contact between electrodes and the skin. This not only poses the risk of skin allergies, but also presents a cumbersome electrode attachment process, making long-term, automated monitoring difficult. Furthermore, existing ECG monitoring technology relies on electrode attachment, resulting in complex operation and lengthy testing times, limiting the effectiveness of large-scale health screening.
[0005] As a non-contact method for detecting physiological signals, rPPG technology has garnered widespread attention in the health monitoring field in recent years. Currently, studies have successfully utilized rPPG technology to acquire basic physiological signals such as heart rate and respiratory rate. However, the technology for accurately converting rPPG signals into ECG signals is still in the exploratory stage, and a mature and reliable solution has yet to be developed. Existing rPPG extraction algorithms are sensitive to environmental factors. In complex scenarios such as unstable lighting conditions and motion interference, the accuracy of rPPG signal extraction can be severely affected, leading to errors in the subsequent ECG signal generation, significantly reducing the reliability and practicality of ECG detection systems based on rPPG technology.
[0006] Currently, ECG abnormality monitoring relies primarily on manual interpretation by physicians, which is inefficient and subject to subjective factors. The consistency and accuracy of monitoring results are difficult to guarantee. Although machine learning and deep learning technologies have driven the development of automated monitoring models, these models require high-quality ECG signals as input to complete monitoring. Furthermore, the lack of labeled datasets makes model training prone to overfitting, limiting their clinical application.
[0007] Given the limitations of portable contact ECG detection technology, as well as the shortcomings of existing rPPG-based ECG generation and ECG anomaly monitoring technologies, video-based intelligent methods for predicting ECG anomalies offer a solution. This approach breaks through the constraints of traditional detection methods. Through non-contact video acquisition and intelligent algorithms, it enables anomaly monitoring using only video data. This approach effectively improves the convenience, accuracy, and cost-effectiveness of ECG detection and monitoring, meeting the urgent needs of clinical and home health monitoring.
[0008] Patent document CN111667921A discloses an AI-powered ECG system for detecting atrial fibrillation and arrhythmia. This system includes an electrocardiogram (ECG) device, a smartphone, and mobile internet. The ECG device transmits detected or monitored ECG signals to the smartphone via Bluetooth or WiFi. The smartphone then transmits the ECG signals to the cloud via a mobile data network, sending reports back to the patient, their family, and their designated physician. The cloud platform includes a server, database, user management system, and AI-powered ECG automatic monitoring software for atrial fibrillation and arrhythmia. This solution fails to address the technical challenges of batch screening and data verification.
[0009] This solution requires access to an ECG monitor and ECG reports for intelligent analysis, and cannot implement intelligent detection without ECG data modalities. This problem urgently needs to be addressed. Summary of the Invention
[0010] In view of the defects in the prior art, the purpose of the present invention is to provide an intelligent method and system for predicting ECG abnormalities based on video.
[0011] The video-based intelligent method for predicting electrocardiographic abnormalities provided by the present invention includes:
[0012] Step S1: Collect video data, lock facial information in the video data and extract rPPG signals;
[0013] Step S2: constructing a FANMAMBA network model, inputting the rPPG signal into the FANMAMBA network model to generate an ECG signal;
[0014] Step S3: Input the rPPG signal and ECG signal into the ECG anomaly detection model to obtain the fusion features of the signals, and then generate an anomaly prediction result.
[0015] Preferably, the step S1 includes:
[0016] Step S1.1: Lock the coordinates of the face area in the video data through face recognition;
[0017] Step S1.2: cropping the face image according to the face region coordinates to obtain an image block, and scaling the image block to obtain a scaled image block;
[0018] Step S1.3: Convert the scaled image block from the RGB color space to the YCbCr color space to obtain a converted image;
[0019] Step S1.4: normalizing the converted image to obtain a normalized image block, thereby obtaining an image sequence as input data;
[0020] Step S1.5: Processing the image sequence through CNN to extract and obtain a feature map corresponding to the rPPG data, thereby obtaining an rPPG signal;
[0021] The resolution of the video data is greater than or equal to 720p, and the frame rate of the video data is greater than or equal to 30fps.
[0022] Preferably, the step S2 includes:
[0023] Step S2.1: Fourier transform the rPPG signal, decouple it to obtain the transformation result, then express the periodic components of the transformation result through trigonometric functions, and process the non-periodic components of the transformation result through activation functions to obtain input data;
[0024] Step S2.2: Let the dimension of the input data be D, and then divide it into P small blocks, and process each of the small blocks using different scales to obtain a processing result;
[0025] Step S2.3: Mapping the processing results into ECG data through a predictor;
[0026] In step S2.1, the mathematical expression of the input data is:
[0027]
[0028] in, represents the input data, x represents the rPPG signal; W1, W2, and B2 are a learnable weight matrix, another learnable weight matrix, and another learnable weight matrix, respectively; σ represents the activation function; [||] indicates that the combination method is cascaded along the first dimension of the data; where (B2+W2x) represents the non-periodic component; W1x represents the periodic component;
[0029] In step S2.2, each of the small blocks is processed at different scales to obtain a processing result, including:
[0030] Step S2.2.1: Let the dimension of the input data be D, and then divide it into P small blocks; project the small blocks into different spaces to obtain projections Uj 、V j With Z j ;
[0031] Step S2.2.2: Capture the hidden state through SSM and output the preliminary results;
[0032] Step S2.2.3: One-dimensionally convolve the preliminary result, activate it through the Silu activation function, perform a merging operation through the routing matrix R, and perform weighted summation to obtain the processing result;
[0033] In step S2.3, the ECG data is expressed mathematically as follows:
[0034] E=W pred M out +b pred
[0035] Where, E represents ECG data; W pred Represents the weight matrix of the fully connected layer of the model; M out Indicates the processing result; b pred Represents the bias vector of the fully connected layer of the model, with dimension m; m represents the length of the ECG data.
[0036] Preferably, in step S2.2.1, the projection U j 、V j With Z j The mathematical expressions from top to bottom are:
[0037] U j =W u X j +b u
[0038] V j =W v X j +b v
[0039] Z j =W z X j +b z
[0040] Among them, W u 、W v With W z The projection U j The weight matrix, projection V j The weight matrix and projection Z j The weight matrix of b u 、b v with b z The projection Uj The bias vector, projection V j The bias vector and projection Z j The bias vector of
[0041] In step S2.2.2, the mathematical expression of the hidden state is:
[0042]
[0043] in, represents the hidden state of the jth block at the tth moment; ρ(A) represents a linear operator based on the matrix A, where the matrix A is a structured matrix; σ represents the activation function; and Represents U j and V j The value at time t;
[0044] The mathematical expression of the preliminary results is:
[0045]
[0046] Among them, Y i Indicates preliminary results, is the hidden state at the last moment, W y Represents a weight matrix, b y Represents a bias vector;
[0047] In step S2.2.3, the one-dimensional convolution of the preliminary result is expressed as follows:
[0048]
[0049] in, Represents the result of one-dimensional convolution; Represents a one-dimensional convolution kernel; Conv1D represents a one-dimensional convolution; Represents the convolution bias;
[0050] The activation is performed by Silu activation function, and the mathematical expression is:
[0051]
[0052] in, Represents the activation result of the activation function, Silu represents the Silu activation function;
[0053] The weighted sum of the merging operation is performed to obtain the processing result, which is expressed as follows:
[0054]
[0055] Among them, M oat represents the processing result; P and M represent the dimensions of the routing matrix R, respectively, wherein the dimension of the routing matrix R is P×M; R jk represents the element of the routing matrix R, i.e., the probability that the jth patch is assigned to the kth scale; It is the result of processing the j-th block at the k-th scale.
[0056] Preferably, in step S3, inputting the rPPG signal and the ECG signal into a large ECG abnormality detection model comprises:
[0057] Step A1: extracting rPPG signals and ECG signals respectively to obtain rPPG physiological characteristics and ECG physiological characteristics;
[0058] Step A2: feature fusion of the rPPG physiological features and ECG physiological features, and let the model learn ICU sudden death data to obtain a weight vector, thereby generating a knowledge transfer framework for the ECG anomaly detection model;
[0059] Step A3: Using the knowledge transfer framework, the ECG anomaly detection model captures abnormal patterns in the ICU sudden death data and outputs abnormality prediction results; the abnormal patterns include: premature ventricular contractions and atrial fibrillation;
[0060] The ECG anomaly detection model is learned and trained using ICU sudden death data.
[0061] Preferably, in step A1, the mathematical expression for extracting the rPPG physiological characteristics is:
[0062]
[0063] Among them, Φ rPPG (X) represents the rPPG feature vector; A sys represents the systolic waveform amplitude; the T sys represents the duration of the systolic period; T dias represents the duration of diastole; Indicates the rising waveform speed, that is, the slope; Indicates the relative intensity of dicrotic wave; T peak represents the time interval between adjacent peaks; SDPP represents the standard deviation of pulse intervals; RMSSD represents the root mean square of the difference between adjacent pulse intervals; Indicates the ratio of low-frequency to high-frequency power;
[0064] The mathematical expression for ECG physiological feature extraction is:
[0065]
[0066] Among them, ΦECG (X) represents the ECG feature vector; A P Indicates the P wave amplitude; A QRS A represents the QRS complex amplitude; T represents the T wave amplitude; ΔST represents the ST segment deviation; T PR represents the PR interval; T QRS Indicates the duration of the QRS wave; T QT represents the QT interval; represents the corrected QT interval; represents the relative change of adjacent RR intervals; σ RR represents the standard deviation of the RR interval.
[0067] Preferably, in step A2, the feature fusion of the rPPG physiological feature and the ECG physiological feature is expressed as follows:
[0068]
[0069] Among them, Φ fusion Represents the fused feature vector; represents the feature connection operation; Φ cross represents the cross eigenvector; α rPPG ,α ECG ,α cross Represents the weight vector of the PPG feature group, ECG feature group and cross feature group;
[0070] The model is trained on ICU sudden death data to obtain the weight vector, i.e., the weight vector of the PPG feature group, ECG feature group, and cross feature group. The mathematical expression is:
[0071] [α rPPG ,α ECG ,α cross ]=σ(W g ·[Φ rPPG (X),Φ ECG (X),Φ cross ]+b g )
[0072] Among them, σ represents the activation function, that is, the sigmoid activation function; W g represents the gating weight matrix; b g represents the gate bias vector; the symbol · represents the dot product;
[0073] The mathematical expression of the knowledge transfer framework is:
[0074]
[0075] Where: L contrast represents the knowledge transfer framework; zi Represents the feature representation of the current sample; represents the jth positive sample; represents the kth negative sample, that is, the feature representation of the normal ECG feature; sim represents the cosine similarity function, τ represents the temperature parameter, M represents the number of positive samples, K represents the number of negative samples; K represents the total number of negative samples; M represents the total number of positive samples; exp represents the natural exponential function.
[0076] Preferably, the training and optimization of the ECG abnormality detection model has a mathematical expression of the optimization objective as follows:
[0077] L total =αL cls +βL contrast +δL reg
[0078] Among them, L total Indicates the total loss; L cls represents the classification loss; L contrast represents contrastive learning loss; L reg represents the regularization loss.
[0079] Preferably, the ECG abnormality detection model is mathematically expressed as:
[0080] P(risk|X)=softmax(W cls ·h graph +b cls )
[0081] Where: P(risk|X) represents the probability that the input feature sequence X has the risk of cardiac abnormality; W cls with b cls Represent the weight matrix and bias vector of the model classifier respectively; h graph Represents the output features of the graph neural network layer of the model; softmax represents the softmax activation function.
[0082] Preferably, a graph G is obtained based on the fused feature vectors through a graph neural network, thereby generating an anomaly prediction result;
[0083] The mathematical expression of the graph G is:
[0084] G=(V,E)
[0085] Among them, V represents the feature set; E represents the feature relationship set;
[0086] The node feature h of the graph G v The update formula is:
[0087]
[0088] in, Represents the update result, N(v) is the set of neighbor nodes of node v, It is the aggregated neighbor node feature, AGGREGATE and UPDATE represent the aggregation function and update function respectively; represents the feature representation of node u in the lth layer; where l represents the layer index of the graph neural network; u represents the neighboring node connected to node v, representing other waveform features that are physiologically related to the current waveform feature; represents the feature representation of node v in layer l.
[0089] According to the present invention, a video-based intelligent system for predicting electrocardiogram abnormalities is provided, comprising:
[0090] Module M1: collects video data, locks facial information in the video data and extracts rPPG signals;
[0091] Module M2: constructing a FANMAMBA network model, inputting the rPPG signal into the FANMAMBA network model, and generating an ECG signal;
[0092] Module M3: Input the rPPG and ECG signals into the ECG anomaly detection model to obtain the fusion features of the signals and generate an anomaly prediction result.
[0093] Compared with the prior art, the present invention has the following beneficial effects:
[0094] 1. The non-contact monitoring method of the present invention not only surpasses traditional contact technology, but also realizes efficient and accurate ECG signal monitoring, is widely applicable to health monitoring, and provides a new solution for heart health management.
[0095] 2. The present invention incorporates a method of generating ECG from rPPG, which can extract deeper abnormal cardiac state characteristics under non-contact conditions and achieve non-contact accurate identification
[0096] 3. The FANMAMBA method provided by the present invention combines the Fourier analysis network to extract deep periodic features and applies the mamba architecture to reduce the amount of calculation while maintaining high-precision conversion, solving the bottlenecks of existing methods in terms of computational efficiency and real-time performance.
[0097] 4. The present invention uses ICU case information data to fine-tune the large model and adopts comparative transfer learning to enhance the model's ability to recognize various signs of sudden death, thereby improving the accuracy of anomaly detection. BRIEF DESCRIPTION OF THE DRAWINGS
[0098] Other features, objects and advantages of the present invention will become more apparent upon reading the detailed description of non-limiting embodiments with reference to the following drawings:
[0099] Figure 1 A schematic diagram of the system architecture provided by the present invention;
[0100] Figure 2 This is a flow chart of rPPG signal extraction provided by the present invention;
[0101] Figure 3 A schematic diagram of the FANMAMBA model structure provided by the present invention;
[0102] Figure 4 This is a flow chart of ECG abnormality detection provided by the present invention. DETAILED DESCRIPTION
[0103] The present invention will be described in detail below with reference to specific embodiments. The following examples will help those skilled in the art to further understand the present invention, but are not intended to limit the present invention in any form. It should be noted that, for those skilled in the art, several changes and improvements can be made without departing from the scope of the present invention. These all fall within the scope of protection of the present invention.
[0104] The present invention aims to provide a system and method for efficiently and accurately extracting rPPG data through video, generating ECG data, and thus realizing non-contact abnormality warning, so as to overcome the many limitations of traditional ECG detection methods and realize more convenient and reliable heart health monitoring.
[0105] The present invention relates to the field of biomedical signal processing, and more specifically, to a non-contact electrocardiogram (ECG) signal detection and analysis method. This method uses a camera to contactlessly capture video data from a monitored individual, extracts rPPG signals using video processing techniques, and converts them into ECG data using the FANMAMBA deep learning model. The generated ECG data is then analyzed using a large electrocardiogram (ECG) model to detect and analyze anomalies, generating detection and analysis data that assists doctors in further diagnosing the patient's condition.
[0106] The present invention extracts remote pulse wave data, i.e., Remote Photoplethysmography (rPPG), through video, and then uses a deep learning model to generate electrocardiogram data, i.e., Electrocardiogram (ECG). The ECG anomaly detection model is then used to perform intelligent early warning on the ECG, thereby constructing a video-based ECG anomaly warning system.
[0107] According to the present invention, a video-based intelligent method for predicting electrocardiographic abnormalities includes:
[0108] S1: Collect video data to ensure that the face is complete and unobstructed, and the surrounding lighting environment is stable;
[0109] S2: Through face recognition, lock facial information and extract rPPG signals;
[0110] S3: Builds a FANMAMBA network model based on the Fourier series analysis network and period information, takes the rPPG obtained in S1 as input, extracts and learns the features of the rPPG signal, and thus generates a high-quality ECG signal;
[0111] S4: Using the ECG anomaly detection large model, the ECG signal and rPPG signal generated in S3 are taken as input to extract deep features, thereby achieving the result of video-based ECG anomaly prediction.
[0112] Specifically, video data acquisition utilizes a high-resolution, low-noise, and stable frame-rate camera to ensure clear capture of subtle optical changes in the subject's face or skin surface, regardless of ambient lighting conditions. The camera's position and angle are optimized to capture the optimal imaging area and minimize signal deviations caused by angular variations.
[0113] In addition, to reduce ambient light interference, an ambient light sensor can be equipped to monitor the ambient light intensity and color temperature in real time, and automatically adjust camera parameters such as exposure time and gain, to ensure the stability of the collected video data quality and lay the foundation for the subsequent accurate extraction of rPPG signals.
[0114] Specifically, rPPG data extraction: Multi-scale feature fusion and spatiotemporal attention mechanisms are combined to extract rPPG data. The MTCNN algorithm is used to perform face recognition and data preprocessing on video frames, including cropping, scaling, color space conversion from RGB to YCbCr, and normalization.
[0115] A deep learning model was constructed, consisting of a multi-layer convolutional neural network (CNN) and a Transformer. The CNN extracts features through convolution operations at different scales, while the Transformer uses a multi-head attention mechanism to capture inter-frame dependencies. The labeled dataset was partitioned into training, validation, and test sets according to a specific ratio. Hyperparameters such as the initial learning rate, batch size, and number of training rounds were set for training. The model was optimized using a loss function consisting of a time-domain loss and a frequency-domain loss based on an FFT energy constraint, resulting in a purer and more stable rPPG signal. The time-domain loss includes mean squared error and Pearson correlation coefficient.
[0116] Specifically, ECG data generation: construct the FANMAMBA model, the full name of which is Fourier Analysis Networks & MAMBA model, which integrates the Fourier analysis network, namely Fourier Analysis Networks, referred to as FAN, the hardware perception algorithm and the simplified state space model, namely Structural Sequence Model, referred to as SSM architecture.
[0117] The input rPPG signal sequence undergoes two linear transformations. The transformed results are then processed with trigonometric functions and activation functions before being concatenated to obtain fused features. The fused features are then divided into multiple small blocks through a block-by-block operation. Each block is processed by a Mamba unit, followed by a one-dimensional convolution and activation function at different scales. A multi-scale router combines the results of the different scales based on a learnable routing matrix, and finally a fully connected layer-based predictor maps them to ECG data. The model is trained using a large amount of accurately labeled rPPG-ECG data pairs to improve computational efficiency and real-time performance. The Mamba unit processing includes projection operations, recursive calculations based on SSM, and output calculations.
[0118] Specifically, ECG anomaly detection involves inputting the generated ECG signal into a large ECG anomaly detection model built on the Transformer architecture to extract deep features such as morphology, rhythm, and frequency. This model integrates a graph neural network for cross-channel feature fusion and uses a sliding window mechanism for continuous and real-time monitoring. The final output includes normal / abnormal classification, anomaly type identification, and risk level prediction. Identified anomaly types include premature ventricular contractions and atrial fibrillation.
[0119] First, video data acquisition: Use a high-definition camera with a resolution of at least 720p to ensure that it can clearly capture subtle optical changes in the monitored person's face or skin surface. The ambient lighting for acquisition should be uniform and stable to avoid degradation of video image quality due to lighting issues.
[0120] The video frame rate is set to 30fps, that is, 30 frames of image are collected per second. Assuming the acquisition time is T seconds and the total number of frames is N, the acquired video frame sequence can be expressed as Among them, F t Represents the t-th frame image.
[0121] Specifically, a higher frame rate can ensure the smoothness of data acquisition, reduce motion blur, provide rich time series information for subsequent accurate extraction of rPPG signals, and improve data accuracy.
[0122] Second, rPPG data extraction: face recognition and data preprocessing: the MTCNN algorithm, namely the Multi-task Cascaded Convolutional Networks algorithm, is used to perform real-time face recognition on the collected video frames.
[0123] The MTCNN algorithm uses three cascaded convolutional neural networks to perform rough positioning, precise positioning, and key point detection of the face area. t ,After MTCNN algorithm processing, the position coordinates of the face area (x t ,y t ,w t ,h t ), where (x t ,y t ) is the coordinate of the upper left corner, w t With h t The width and height respectively.
[0124] Cropping and scaling: Cropping the face according to the face area coordinates to obtain an image block of size a×a Where a is the set image block size parameter. Bilinear interpolation is used to scale it so that the image block has consistency at different scales. The scaled image block is recorded as
[0125] Color space conversion: convert the scaled image blocks Converting from RGB color space to YCbCr color space facilitates the separation of luminance and chrominance information to highlight the low-frequency components of the rPPG signal, remove high-frequency noise, and adjust the frequency distribution, thereby improving the quality and stability of the rPPG signal and increasing the accuracy of its frequency component extraction.
[0126] Normalization: Perform normalization on the converted image blocks, map the pixel values to the range of [-1, 1], and form a new sequence of normalized image blocks As input for subsequent steps.
[0127] Model Construction and Training: Preprocessed facial video images are passed through a multi-layer convolutional neural network (CNN) to extract local features related to rPPG and output a feature map of size N×D, where N is the number of frames and D is the feature dimension. The feature sequence (N×D) output by the CNN is used as the input to a Transformer, and positional encoding is added to preserve temporal information. A multi-layer Transformer Encoder is used to model the dependencies between frames using a self-attention mechanism.
[0128] During training, the time domain loss LOSS is used time and frequency domain loss LOSS freq Combined as the loss function L.
[0129] 1. Time domain loss LOSS time It consists of two parts: mean square error (MSE) and Pearson correlation coefficient.
[0130] Mean square error, or MSE: The mean square error is used to measure the predicted value The average square error with the true value y in the time domain is:
[0131]
[0132] Where n is the number of samples, is the predicted value of the i-th sample, y i is the true value of the i-th sample. MSE can reflect the degree of deviation between the predicted value and the true value. The greater the deviation, the greater the MSE value.
[0133] Pearson correlation coefficient: The Pearson correlation coefficient is used to measure the linear correlation between the predicted value and the true value. Its formula is:
[0134]
[0135] in, is the mean of the predicted values, is the mean of the true values.
[0136] Combining the mean square error and the loss based on the Pearson correlation coefficient, we get the time domain loss LOSS time for:
[0137] LOSS time =α×MSE+(1-α)×(1-r)
[0138] Among them, α is a hyperparameter used to balance the weights of MSE and Pearson correlation coefficient loss terms.
[0139] 2. Frequency domain loss LOSS freq Frequency domain loss uses FFT energy constraint. First, the predicted value And the true value y are respectively subjected to fast Fourier transform, i.e. FFT, to obtain their representation in the frequency domain and Y(f).
[0140] FFT energy constraint: Calculate the energy of the predicted value and the true value in the frequency domain. The formulas are:
[0141]
[0142] Frequency domain loss LOSS freq Defined as:
[0143]
[0144] This loss function measures the relative difference in frequency domain energy between the predicted value and the true value. The greater the difference, the greater the loss value.
[0145] 3. Total loss function L; the time domain loss LOSS time and frequency domain loss LOSS freq Combined, we get the total loss function L, whose mathematical expression is:
[0146] L=LOSS time +σLOSS freq
[0147] Among them, σ is a hyperparameter used to balance the weights of time domain loss and frequency domain loss. By adjusting hyperparameters such as α and σ, the performance of the model in the time domain and frequency domain can be optimized to make the model fit the data better. Then, the gradient of the loss function with respect to the model parameter θ is calculated by the back propagation algorithm. Update the model using the Adam optimizer.
[0148] (3) ECG data generation: The complete technical solution for generating ECG data based on rPPG signals includes two core modules: the Fourier analysis network (FAN) and the adaptive multi-scale Mamba module.
[0149] 1. Fourier Analysis Networks, or FAN for short.
[0150] Assume that the input rPPG signal sequence is X=[x(t1),x(t2),…,x(t n )], in order to extract the expressiveness of enhancing the periodicity modeling from the rPPG signal, FAN was used to model the periodicity of the rPPG signal.
[0151] First, the time series, i.e., the rPPG signal series, is decoupled into two parts, periodic and non-periodic components, through Fourier transform, which helps to better analyze and understand the different components in the time series data.
[0152] The periodic components are represented by trigonometric functions. The periodicity and fluctuation of trigonometric functions can help the model capture the periodic features in the rPPG signal, which are related to the periodicity of heartbeats.
[0153] The Gaussian Error Linear Unit activation function, namely Gaussian Error Linear Units, or GELU activation function for short, is applied to the non-periodic components.
[0154] The GELU activation function has nonlinear characteristics and can introduce the nonlinear expression ability of the model, allowing the model to learn more complex functional relationships.
[0155] In summary, the input data is processed in the following way:
[0156]
[0157] Here, [·||·] indicates that the combination is cascaded along the first dimension of the data. W1, W2, and B2 are learnable weight matrices. σ represents the activation function, which can further enhance its expressiveness in modeling periodicity.
[0158] 2. Adaptive Multi-Scale block;
[0159] Multi-scale Mamba block. Mamba is an efficient architecture for sequence modeling. Its core idea is to use the structured state space model (SSM) to capture long-range dependencies in the sequence. In the multi-scale Mamba block, this scheme processes the input φ(x) at different scales to capture the feature information of different scales in the rPPG signal.
[0160] Block operation: Block the input data φ(x) into different small blocks (patches). Assume that the dimension of the input data φ(x) is D, and it is divided into P small blocks, and the size of each small block is The jth small block is represented by X j , j = 1, 2, ..., P. P represents the number of divided small blocks.
[0161] The purpose of the blocking operation is to localize the input data, enabling the model to extract detailed features in different local regions, thereby capturing features at different scales. Smaller blocks can capture local details of the signal, while larger blocks can capture more macroscopic features. This is the innovation: capturing features at different scales.
[0162] Multi-scale processing: Each small block X j It needs to be processed at different scales. k (k=1,2,…,M), there is a corresponding processing flow, combining detail features and global features to improve the integrity of feature extraction, thereby improving the generation quality. Taking a scale as an example, the processing process is as follows:
[0163] Mamba unit processing: The core of the Mamba unit is based on the recursive structure of SSM. j , its processing process is represented by the following steps:
[0164] Specifically, an adaptive multi-scale SSM structure was introduced for the first time. Through overlapping blocking strategies and dynamic routing mechanisms, the features of different time scales were simultaneously captured, solving the computational complexity problem of traditional methods in long time series modeling. At the same time, the routing matrix was used to adaptively assign processing weights to different scales, reconstructing the P-QRS-T waveform structure in ECG more accurately than existing methods, significantly improving the generation quality and computational efficiency.
[0165] 1. Projection operation: First, X j Projected into different spaces through three linear transformations, we get U j 、V j and Z j :
[0166] U j =W u patchX j +b u
[0167] V j =W v patchX j +b v
[0168] Z j =W z patchX j +b z
[0169] Among them, W u 、W v and W z is the learnable weight matrix, b u 、b v and b z is the bias vector.
[0170] 2. SSM recursion: Mamba uses SSM to capture the long-range dependencies of rPPG sequences. is the hidden state at time t, and its update formula is:
[0171]
[0172] Among them, ρ(A) is a linear operator based on the matrix A, σ is the activation function, and U j and Vj The value at the tth moment; the matrix A is a structured matrix whose structure enables the model to process long sequences efficiently.
[0173] 3. Output calculation: final output Y j is obtained by another linear transformation:
[0174]
[0175] in, is the hidden state at the last moment, W y is the weight matrix, b y is the bias vector. Specifically, W y Represents a weight matrix, b y Represents a bias vector.
[0176] Specifically, the output calculation step is a direct continuation of the SSM recursion. It maps the long-term temporal dependency information captured by the SSM recursion, that is, the final hidden state h_T, to the target feature space through the linear transformation Y = W_yh_T + b_y, realizing the conversion from state space to feature representation, and providing the necessary feature representation for subsequent multi-scale feature fusion.
[0177] One-dimensional convolution and activation function: After being processed by the Mamba unit, the output Y j Perform a one-dimensional convolution operation:
[0178]
[0179] in, is a one-dimensional convolution kernel, is the convolution bias. One-dimensional convolution can further capture local patterns and sequence information in small block features. The size and step size of the convolution kernel can be adjusted according to different scales to meet the requirements of feature extraction at different scales. Then, the Sigmoid-weighted Linear Unit (Silu) activation function is used:
[0180] Silu(x)=x·σ(x)
[0181] Among them, σ(x) is the sigmoid function; so
[0182] Multi-Scale Router: A multi-scale router is responsible for managing data processing paths at different scales. This is achieved through a learnable routing matrix R, where the elements R jkrepresents the probability of the jth small block being assigned to the kth scale. The dimension of the routing matrix R is P×M. After routing, the results of processing at different scales are merged to obtain the output M of the adaptive multi-scale module. out The routing matrix R has a dimension of P×M because it is responsible for determining the processing weight of each of the P small blocks at each of the M scales.
[0183] The merging operation can be performed in a weighted summation manner:
[0184]
[0185] in, It is the result of processing the j-th block at the k-th scale.
[0186] Predictor: The predictor converts the output M of the adaptive multi-scale module into out Mapped to ECG data, the mathematical expression is:
[0187] E=[E(t ′ 1),E(t ′ 2),…,E(t ′ m )]
[0188] Among them, t j ′ represents the time point of ECG data, m is the length of ECG data. Let the weight matrix of the fully connected layer be W pred , the dimension is ( It's M out dimension), the bias vector of the fully connected layer is b pred , the dimension is m, then the output of the predictor is:
[0189] E=W pred M out +b pred
[0190] This fully connected layer maps the features extracted by the adaptive multi-scale module to the space of the electrocardiogram data, thereby generating ECG data. Where E represents the generated ECG data, and the superscript “′” represents the time point of the ECG data;
[0191] Specifically, the superscript "'" represents a time point, which is used to distinguish the time point of the ECG data from the time point of the original rPPG signal. In other words, "'" indicates a time point in the ECG data sequence, not a time point in the original rPPG signal.
[0192] 4. Model training: Use a large amount of accurately labeled rPPG-ECG data pairs (N is the number of data pairs) for training. The mean square error loss function L is used to measure the ECG data predicted by the model. The difference between the real ECG data E is:
[0193]
[0194] Use the Adam optimizer to update the model parameters. The Adam optimizer combines the advantages of AdaGrad and RMSProp and can adaptively adjust the learning rate of each parameter. Its update formula is as follows:
[0195] First, calculate the first moment estimate m of the gradient t and the second moment estimate v t :
[0196] m t =β1m t-1 +(1-β1)g t
[0197]
[0198] Among them, g t is the gradient of the loss function with respect to the parameter θ at step t, and β1 and β2 are the decay rates, which are set to 0.9 and 0.999 respectively.
[0199] Then the first-order moment estimate and the second-order moment estimate are bias corrected:
[0200]
[0201] Finally update the parameter θ:
[0202]
[0203] Where α is the learning rate and ∈ is a small constant used to avoid division by zero errors, set to 10 -8 .
[0204] (IV) ECG Anomaly Detection: This phase builds a cardiac anomaly warning system based on the rPPG signals extracted and ECG data generated in the first two phases. This system uses transfer learning to transfer knowledge from sudden death ECG data in the hospital ICU to remote monitoring scenarios, enabling accurate early warning of potential cardiac anomalies.
[0205] The system architecture consists of three main modules, including:
[0206] Dual feature extraction module: extracts multiple physiological features from rPPG and ECG signals simultaneously;
[0207] Knowledge transfer learning module: transfer abnormal pattern recognition capabilities from ICU sudden death data;
[0208] Abnormal risk assessment module: integrates multiple features for risk judgment and early warning.
[0209] 1. Data preparation and preprocessing:
[0210] First, we collected public arrhythmia datasets: We selected the MIT-BIH Arrhythmia Database, the CPSC2018 dataset, and the European ST-T Database. For each dataset, we recorded detailed metadata, including the data source, acquisition equipment, sampling frequency, and annotation method.
[0211] Second, organize abnormal ECG data of ICU patients: extract the patient's ECG data from the ICU information system to ensure that the data timestamp is accurate. For missing or damaged data, data interpolation method will be used.
[0212] The core innovation of this system is the sudden death data from ICU patients, which contains clinically significant precursor signals of sudden death and provides the model with real-world abnormal pattern learning samples. This is the innovation point.
[0213] Third, noise removal, namely wavelet transform filtering method: In the wavelet transform filtering process, the Daubechies wavelet basis function is selected, which has good time-frequency localization characteristics and can effectively separate noise and useful information in the signal.
[0214] According to the characteristics of rPPG and ECG signals, the number of decomposition layers is set to 5. The specific steps are as follows:
[0215] First, the original signal is decomposed by wavelet to obtain the approximate coefficients and detail coefficients at different scales; then, the detail coefficients are threshold processed using the soft threshold method. The threshold calculation formula is: Where σ is the standard deviation of noise and N is the signal length; finally, wavelet reconstruction is performed on the processed coefficients to obtain the denoised signal.
[0216] Fourth, data standardization: Through the downsampling method, the sampling rate is unified to 128Hz. Then, the sliding window method is used, n data points are set as a window, and sliding calculation is performed on the entire data set. For the data in each window, its mean μ and standard deviation σ are calculated, and then the data in the window are calculated according to the formula Standardization is performed to ensure that the data has a uniform scale and distribution over different time periods.
[0217] 2. Feature extraction module:
[0218] (1) Extract multiple physiological features from rPPG signals to capture signal patterns that may indicate cardiac abnormalities, especially the precursors to sudden death:
[0219]
[0220] Among them, A sys It represents the amplitude of the systolic waveform, which reflects the strength of cardiac contraction and is related to ventricular ejection volume. A too low value may indicate a decrease in cardiac output. sys It indicates the duration of systole, reflects the myocardial contraction speed, and is related to cardiac contractile function. A decrease may indicate weakened myocardial function. dias Indicates the duration of diastole, reflects myocardial diastolic function, and is related to cardiac filling capacity. Abnormalities may indicate diastolic dysfunction;
[0221] Indicates the speed of the rising waveform, that is, the slope, which reflects the elasticity of the vascular wall and peripheral resistance and is related to arterial compliance. A slowdown may indicate vascular sclerosis; A dic Indicates the dicrotic wave amplitude, reflecting the reflection intensity of the pulse wave in peripheral blood vessels. It is related to vascular resistance and compliance. An increase may indicate a decrease in vascular compliance. Indicates the relative intensity of the dicrotic wave, reflecting the ratio of the peripheral vascular reflected wave to the forward wave. It is related to arterial stiffness. An increase may indicate vascular aging or hypertension. peak It represents the time interval between adjacent peaks, corresponding to the cardiac cycle, that is, the inverse of the heart rate, reflecting the basic rhythm of cardiac electrical activity. Irregularity may indicate arrhythmia. SDPP represents the standard deviation of pulse intervals, reflecting heart rate variability and related to the function of the autonomic nervous system. A decrease may indicate a weakening of the autonomic nervous system regulation function. RMSSD represents the root mean square of the difference between adjacent pulse intervals, reflecting short-term heart rate variability and mainly related to parasympathetic nerve activity. A decrease may indicate a decrease in parasympathetic nerve tone. It represents the power ratio of low frequency to high frequency, reflects the balance between sympathetic and parasympathetic nerves, and is related to the regulatory function of the autonomic nervous system. An increase may indicate an increase in sympathetic nerve tension.
[0222] (2) ECG feature extraction formula:
[0223]
[0224] Among them, A P Indicates the P wave amplitude, which reflects the atrial depolarization process and is related to the intensity of atrial electrical activity. An increase may indicate atrial hypertrophy; A QRS Indicates the amplitude of the QRS complex wave, reflecting the ventricular depolarization process, and is related to the ventricular muscle mass and electrical activity intensity. An increase may indicate ventricular hypertrophy; A TT represents the amplitude of the T wave, which reflects the ventricular repolarization process and is related to the electrical activity of ventricular repolarization. Abnormalities may indicate myocardial ischemia or electrolyte imbalance. ΔST represents the ST segment deviation, which reflects the period between ventricular depolarization and repolarization and is closely related to the myocardial blood supply status. Elevation or depression may indicate myocardial ischemia or damage. PR PR interval, which reflects the time from the onset of atrial excitation to the onset of ventricular excitation, is related to the atrioventricular conduction time. Prolongation may indicate first-degree atrioventricular block. QRS Indicates the duration of the QRS wave, reflecting the time required for ventricular depolarization and related to the intraventricular conduction velocity. Widening may indicate intraventricular conduction block; T QT QT interval reflects the total time of ventricular depolarization and repolarization and is related to the complete process of ventricular electrical activity. Prolongation may indicate abnormal repolarization and increase the risk of arrhythmia. It stands for Corrected QT Interval, corrects the effect of heart rate on the QT interval, provides a standardized measurement of repolarization time, and prolongation is a risk marker for a variety of fatal arrhythmias; It indicates the relative change of adjacent RR intervals, reflects the dynamic changes of heart rate, and is related to the stability of sinus rhythm. An increase may indicate unstable heart rhythm; RR It represents the standard deviation of the RR interval, reflects heart rate variability, and is related to the autonomic nervous system regulation function. A decrease in it is an independent predictor of a high risk of sudden death.
[0225] (3) Feature fusion and weight learning formula:
[0226]
[0227] Among them, Φ fusion represents the fused feature vector; Φ rPPG (X) represents the rPPG feature vector; Φ ECG (X) represents the ECG feature vector; Φ cross represents the cross eigenvector in, It represents the ratio of ventricular electrical activity time to atrial cycle, reflecting the coordination of cardiac electrical activity; ρ rPPG-ECG represents the correlation coefficient, that is, the correlation between the rPPG signal and the ECG signal; α rPPG ,α ECG ,α cross Represents the weight vector of each feature group; Represents a feature connection operation.
[0228] (4) The weight vector is obtained by learning from ICU sudden death data:
[0229] [α rPPG ,α ECG ,α cross]=σ(W g ·[Φ rPPG (X),Φ ECG (X),Φ cross ]+b g )
[0230] Where: W g represents the gating weight matrix, b g represents the gate bias vector, and σ represents the sigmoid activation function.
[0231] 3. ECG abnormality warning model architecture
[0232] 3.1. ECG Model Infrastructure
[0233] This paper uses a large model for ECG anomaly detection based on a multi-layer Transformer structure, which has strong temporal modeling and attention focusing capabilities. The structure is as follows:
[0234] Input layer: Fusion feature Φ fusion Embedded into a 128-dimensional vector, and the sine and cosine position vectors are introduced to retain time information;
[0235] Enhanced Transformer encoder: Consists of 6 Transformer layers, each with 8 attention sublayers and a feedforward network;
[0236] Attention Gating Unit (AGU): Enhances the perception of changes in P wave, QRS complex, and T wave morphology;
[0237] Spatiotemporal graph neural network layer: captures the relationship between ECG waveform structures. The features are represented as a graph G = (V, E), where nodes v∈V represent features and edges e∈E represent feature relationships. Node features h are updated through model message passing. v , the formula is:
[0238]
[0239] Among them, N(v) is the set of neighbor nodes of node v, is the aggregated neighbor node feature, AGGREGATE and UPDATE are the aggregation and update functions respectively.
[0240] Bayesian uncertainty quantification layer: provides a reliable prediction confidence assessment, the formula is
[0241]
[0242] Among them, p(y|x) represents the given input feature sequence x, that is, the fusion feature Φ fusionPredict the probability of cardiac abnormality risk category y, y∈{0,1}, where 0 indicates normal and 1 indicates abnormal risk; W t represents the network parameters after random dropout in the t-th forward propagation, T represents the number of Monte Carlo sampling, which is set to 20, and U(x) represents the uncertainty of the prediction, which is calculated based on information entropy.
[0243] Binary classification output layer: This layer determines whether there is a risk of cardiac abnormality and generates a warning signal based on the risk score and uncertainty. The fully connected layer maps the features into binary classification results. The formula is:
[0244] P(risk|X)=softmax(W cls ·h graph +b cls )
[0245] Among them, W cls and b cls are the weight matrix and bias vector of the classifier, h graph It is the output feature of the graph neural network layer. P(risk|X) represents the probability that the input feature sequence X has the risk of cardiac abnormality, and its value range is [0,1].
[0246] 3.2 Knowledge Transfer Framework Based on ICU Sudden Death Data:
[0247] (1) Contrastive transfer learning: To address the diversity of sudden death patterns, a multi-positive contrastive learning framework is designed:
[0248]
[0249] Where: z i Represents the feature representation of the current sample, that is, the feature vector processed by the feature extractor; represents the jth positive sample, which is also the feature representation of the sudden death precursor ECG feature; represents the kth negative sample, that is, the feature representation of the normal ECG feature; sim(·,·) represents the cosine similarity function, which calculates the similarity of two feature vectors; τ represents the temperature parameter, which controls the smoothness of the feature distribution; M represents the number of positive samples; K represents the number of negative samples.
[0250] This contrastive learning framework can effectively capture various abnormal patterns in ICU sudden death data and enhance the model's ability to identify various precursors to sudden death.
[0251] 3.3 Loss Function and Optimization Strategy
[0252] (1) Loss function;
[0253] In view of the particularity of the ECG anomaly detection task, a comprehensive optimization objective is designed:
[0254] L total =αL cls +βL contrast +δL reg
[0255] Where: L cls Represents the classification loss, using weighted cross entropy:
[0256]
[0257] Among them, w i is the sample weight, which is used to solve the problem of class imbalance; y i is the true label, where 0 = normal and 1 = abnormal; p i is the predicted abnormal risk probability;
[0258] L contrast represents the contrastive learning loss, as described in 3.2, L reg Represents regularization loss to prevent overfitting:
[0259]
[0260] Among them, the first term is L2 regularization, the second term is the smooth constraint of the adjacent layer parameters α, β, δ are the weight coefficients for balancing the loss terms
[0261] (2) Layer freezing and fine-tuning strategy. Specifically, a layered optimization strategy is designed based on the sensitivity of different network layers to ICU sudden death knowledge:
[0262] First, the pre-training stage: use the public ECG dataset to train the entire model.
[0263] Second, the knowledge transfer stage:
[0264] Shallow feature extraction part: completely frozen.
[0265] Intermediate feature transformation part: low learning rate fine-tuning.
[0266] Deep classification part: comprehensive update with high learning rate.
[0267] Third, fine-tuning stage: only fine-tune the classification head and keep the feature extraction part unchanged.
[0268] This hierarchical strategy ensures that the model can retain common ECG features and fully learn high-level feature patterns unique to sudden death.
[0269] 4. Model evaluation: Use multiple evaluation metrics to comprehensively measure model performance:
[0270] (1) Basic classification indicators:
[0271] Sensitivity (sensitivity), the mathematical expression is:
[0272]
[0273] Among them, Sensitivity represents sensitivity; TP represents true positive, that is, the number of abnormal heart rhythm cases correctly identified; FN represents false negative, that is, the number of abnormal heart rhythm cases that are not identified.
[0274] Specificity, mathematical expression is:
[0275]
[0276] Among them, TN represents true negatives, that is, the number of normal heart rhythm cases correctly identified; FP represents false positives, that is, the number of normal cases mistakenly identified as abnormal heart rhythm.
[0277] F1 score, mathematical expression is:
[0278]
[0279] Among them, Precision represents the accuracy rate, that is, the proportion of true abnormalities among predicted abnormal heart rhythms; Recall represents the recall rate, that is, the proportion of all abnormal heart rhythms that are correctly identified.
[0280] AUC-ROC: Area under the receiver operating characteristic curve
[0281] (2) Time-related indicators:
[0282] Time-to-Alert: The time interval from when the warning is triggered to when the abnormal event occurs
[0283] False Alert Duration Ratio: The ratio of false alarm duration to total monitoring time
[0284] (3) Uncertainty assessment:
[0285] Expected Calibration Error: The deviation between the predicted probability and the actual probability
[0286] Decision Reliability: The accuracy of high-confidence predictions
[0287] The experiment adopted five-fold cross validation, using 60% data for training, 20% for validation, and 20% for testing in each training. The results were repeated five times and the average value was taken to reduce the influence of randomness.
[0288] In addition, the receiver operating characteristic curve, i.e., ROC curve, and precision-recall curve, i.e., PR curve, are plotted to intuitively evaluate the performance of the model.
[0289] The present invention also provides a video-based ECG anomaly prediction intelligent system, which can be implemented by executing the process steps of the video-based ECG anomaly prediction intelligent method. That is, those skilled in the art can understand the video-based ECG anomaly prediction intelligent method as a preferred implementation of the video-based ECG anomaly prediction intelligent system.
[0290] According to the present invention, a video-based intelligent system for predicting electrocardiogram abnormalities is provided, comprising:
[0291] Module M1: collects video data, locks facial information in the video data and extracts rPPG signals;
[0292] Module M2: constructing a FANMAMBA network model, inputting the rPPG signal into the FANMAMBA network model, and generating an ECG signal;
[0293] Module M3: Input the rPPG and ECG signals into the ECG anomaly detection model to obtain the fusion features of the signals and generate an anomaly prediction result.
[0294] The dual feature extraction module is used to extract multiple physiological features from rPPG and ECG signals simultaneously, corresponding to module M1 and module M2;
[0295] The knowledge transfer learning module and the abnormal risk assessment module are used to transfer abnormal pattern recognition capabilities from ICU sudden death data and integrate multiple features for risk judgment and early warning, corresponding to module M3.
[0296] Those skilled in the art will appreciate that, in addition to implementing the system and its various devices, modules, and units provided by the present invention in purely computer-readable program code, it is entirely possible to implement the same functions of the system and its various devices, modules, and units provided by the present invention in the form of logic gates, switches, application-specific integrated circuits, programmable logic controllers, and embedded microcontrollers by logically programming the method steps. Therefore, the system and its various devices, modules, and units provided by the present invention can be considered a hardware component, and the devices, modules, and units included therein for implementing various functions can also be considered as structures within the hardware component; the devices, modules, and units for implementing various functions can also be considered as both software modules implementing the method and structures within the hardware component.
[0297] The above describes specific embodiments of the present invention. It should be understood that the present invention is not limited to the specific embodiments described above, and those skilled in the art may make various changes or modifications within the scope of the claims, which do not affect the essence of the present invention. The embodiments of this application and the features in the embodiments may be combined with each other in any manner unless there is a conflict.
Claims
1. A video-based intelligent method for predicting ECG abnormalities, characterized in that: include: Step S1: Collect video data, lock facial information in the video data and extract rPPG signals; Step S2: constructing a FANMAMBA network model, inputting the rPPG signal into the FANMAMBA network model to generate an ECG signal; Step S3: Input the rPPG signal and ECG signal into the ECG anomaly detection model to obtain the fusion features of the signals, and then generate an anomaly prediction result.
2. The video-based intelligent method for predicting ECG anomalies according to claim 1, characterized in that: In the step S1, it includes: Step S1.1: Lock the coordinates of the face area in the video data through face recognition; Step S1.2: cropping the face image according to the face region coordinates to obtain an image block, and scaling the image block to obtain a scaled image block; Step S1.3: Convert the scaled image block from the RGB color space to the YCbCr color space to obtain a converted image; Step S1.4: normalizing the converted image to obtain a normalized image block, thereby obtaining an image sequence as input data; Step S1.5: Processing the image sequence through CNN to extract and obtain a feature map corresponding to the rPPG data, thereby obtaining an rPPG signal; The resolution of the video data is greater than or equal to 720p, and the frame rate of the video data is greater than or equal to 30fps.
3. The video-based intelligent method for predicting ECG anomalies according to claim 1, characterized in that: In the step S2, it includes: Step S2.1: Fourier transform the rPPG signal, decouple it to obtain the transformation result, then express the periodic components of the transformation result through trigonometric functions, and process the non-periodic components of the transformation result through activation functions to obtain input data; Step S2.2: Let the dimension of the input data be D, and then divide it into P small blocks, and process each of the small blocks using different scales to obtain a processing result; Step S2.3: Mapping the processing results into ECG data through a predictor; In step S2.1, the mathematical expression of the input data is: in, represents the input data, x represents the rPPG signal; W1, W2, and B2 are a learnable weight matrix, another learnable weight matrix, and another learnable weight matrix, respectively; σ represents the activation function; [||] indicates that the combination method is cascaded along the first dimension of the data; where (B2+W2x) represents the non-periodic component; W1x represents the periodic component; In step S2.2, each of the small blocks is processed at different scales to obtain a processing result, including: Step S2.2.1: Let the dimension of the input data be D, and then divide it into P small blocks; project the small blocks into different spaces to obtain projections U j 、V j With Z j ; Step S2.2.2: Capture the hidden state through SSM and output the preliminary results; Step S2.2.3: One-dimensionally convolve the preliminary result, activate it through the Silu activation function, perform a merging operation through the routing matrix R, and perform weighted summation to obtain the processing result; In step S2.3, the ECG data is expressed mathematically as follows: E=W pred M out +b pred Where, E represents ECG data; W pred Represents the weight matrix of the fully connected layer of the model; M out Indicates the processing result; b pred Represents the bias vector of the fully connected layer of the model, with dimension m; m represents the length of the ECG data.
4. The video-based intelligent method for predicting ECG anomalies according to claim 3, characterized in that: In the step S2.2.1, the projection U j 、V j With Z j The mathematical expressions from top to bottom are: U j =W u X j +b u V j =W v X j +b v Z j =W z X j +b z Among them, W u 、W v With W z The projection U j The weight matrix, projection V j The weight matrix and projection Z j The weight matrix of b u 、b v with b z The projection U j The bias vector, projection V j The bias vector and projection Z j The bias vector of In step S2.2.2, the mathematical expression of the hidden state is: in, represents the hidden state of the jth block at the tth moment; ρ(A) represents a linear operator based on the matrix A, where the matrix A is a structured matrix; σ represents the activation function; and Represents U j and V j The value at time t; The mathematical expression of the preliminary results is: Among them, Y j Indicates preliminary results, is the hidden state at the last moment, W y Represents a weight matrix, b y Represents a bias vector; In step S2.2.3, the one-dimensional convolution of the preliminary result is expressed as follows: in, Represents the result of one-dimensional convolution; Represents a one-dimensional convolution kernel; Conv1D represents a one-dimensional convolution; Represents the convolution bias; The activation is performed by Silu activation function, and the mathematical expression is: in, Represents the activation result of the activation function, Silu represents the Silu activation function; The weighted sum of the merging operation is performed to obtain the processing result, which is expressed as follows: Among them, M out represents the processing result; P and M represent the dimensions of the routing matrix R, respectively, wherein the dimension of the routing matrix R is P×M; R jk represents the element of the routing matrix R, i.e., the probability that the jth patch is assigned to the kth scale; It is the result of processing the j-th block at the k-th scale.
5. The video-based intelligent method for predicting ECG anomalies according to claim 3, characterized in that: In step S3, the rPPG signal and the ECG signal are input into a large ECG abnormality detection model, including: Step A1: extracting rPPG signals and ECG signals respectively to obtain rPPG physiological characteristics and ECG physiological characteristics; Step A2: feature fusion of the rPPG physiological features and ECG physiological features, and let the model learn ICU sudden death data to obtain a weight vector, thereby generating a knowledge transfer framework for the ECG anomaly detection model; Step A3: Using the knowledge transfer framework, the ECG anomaly detection model captures abnormal patterns in the ICU sudden death data and outputs abnormality prediction results; the abnormal patterns include: premature ventricular contractions and atrial fibrillation; The ECG anomaly detection model is learned and trained using ICU sudden death data.
6. The video-based intelligent method for predicting ECG anomalies according to claim 5, characterized in that: In step A1, the mathematical expression for extracting rPPG physiological characteristics is: Among them, Φ rPPG (X) represents the rPPG feature vector; A sys represents the systolic waveform amplitude; the T sys represents the duration of the systolic period; T dias represents the duration of diastole; Indicates the rising waveform speed, that is, the slope; Indicates the relative intensity of dicrotic wave; T peak represents the time interval between adjacent peaks; SDPP represents the standard deviation of pulse intervals; RMSSD represents the root mean square of the difference between adjacent pulse intervals; Indicates the ratio of low-frequency to high-frequency power; The mathematical expression for ECG physiological feature extraction is: Among them, Φ ECG (X) represents the ECG feature vector; A P Indicates the P wave amplitude; A QRS A represents the QRS complex amplitude; T represents the T wave amplitude; ΔST represents the ST segment deviation; T PR represents the PR interval; T QRS Indicates the duration of the QRS wave; T QT represents the QT interval; represents the corrected QT interval; represents the relative change of adjacent RR intervals; σ RR represents the standard deviation of the RR interval.
7. The video-based intelligent method for predicting ECG anomalies according to claim 5, characterized in that: In step A2, the rPPG physiological features and ECG physiological features are fused, and the mathematical expression is: Among them, Φ fusion Represents the fused feature vector; represents the feature connection operation; Φ cross represents the cross eigenvector; α rPPG ,α ECG ,α cross Represents the weight vector of the PPG feature group, ECG feature group and cross feature group; The model is trained on ICU sudden death data to obtain the weight vector, i.e., the weight vector of the PPG feature group, ECG feature group, and cross feature group. The mathematical expression is: [a rPPG ,a ECG ,a cross ]=σ(W g ·[Φ rPPG (X),F ECG (X),F cross ]+b g ) Among them, σ represents the activation function, that is, the sigmoid activation function; W g represents the gating weight matrix; b g represents the gate bias vector; the symbol · represents the dot product; The mathematical expression of the knowledge transfer framework is: Where: L contrast represents the knowledge transfer framework; z i Represents the feature representation of the current sample; represents the jth positive sample; represents the kth negative sample, that is, the feature representation of the normal ECG feature; sim represents the cosine similarity function, τ represents the temperature parameter, M represents the number of positive samples, K represents the number of negative samples; K represents the total number of negative samples; M represents the total number of positive samples; exp represents the natural exponential function.
8. The video-based intelligent method for predicting ECG anomalies according to claim 5, characterized in that: The mathematical expression of the optimization target of training and optimizing the ECG anomaly detection model is: L total =αL cls +βL contrast +δL reg Among them, L total Indicates the total loss; L cls represents the classification loss; L contrast represents contrastive learning loss; L reg represents the regularization loss; The ECG abnormality detection model is mathematically expressed as: P(risk|X)=softmax(W cls ·h graph +b cls ) Where: P(risk|X) represents the probability that the input feature sequence X has the risk of cardiac abnormality; W cls with b cls Represent the weight matrix and bias vector of the model classifier respectively; h graph Represents the output features of the graph neural network layer of the model; softmax represents the softmax activation function.
9. The video-based intelligent method for predicting ECG anomalies according to claim 6, characterized in that: Based on the fused feature vector, the graph G is obtained through the graph neural network, and then the anomaly prediction result is generated; The mathematical expression of the graph G is: G=(V,E) Among them, V represents the feature set; E represents the feature relationship set; The node feature h of the graph G v The update formula is: in, Represents the update result, N(v) is the set of neighbor nodes of node v, It is the aggregated neighbor node feature, AGGREGATE and UPDATE represent the aggregation function and update function respectively; represents the feature representation of node u in the lth layer, where l represents the layer index of the graph neural network; u represents the neighboring nodes connected to node v, representing other waveform features that are physiologically related to the current waveform feature; represents the feature representation of node v in layer l.
10. A video-based intelligent system for predicting ECG abnormalities, characterized in that: include: Module M1: collects video data, locks facial information in the video data and extracts rPPG signals; Module M2: constructing a FANMAMBA network model, inputting the rPPG signal into the FANMAMBA network model, and generating an ECG signal; Module M3: Input the rPPG and ECG signals into the ECG anomaly detection model to obtain the fusion features of the signals and generate an anomaly prediction result.
Citation Information
Patent Citations
Artificial intelligence ECG atrial fibrillation and arrhythmia detection system
CN111667921A
End-to-end non-contact atrial fibrillation automatic detection system and method based on vPPG signal
CN112587153A
Remote heart rate detection method based on rPPG signal
CN116994310A
Multi-modal fusion algorithm for electrocardiosignal anomaly detection
CN118520279A
Non-contact heart rate monitoring method based on palm video and related equipment
CN118948237A
Cited By
Lightweight PPG biological feature recognition method based on SSM and feature refining
CN121685995A
A lightweight PPG biometric identification method based on SSM and feature refinement
CN121685995B