ECG signal generation method and device based on multi-source pulse wave signals, equipment and medium
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- RUIJIN HOSPITAL AFFILIATED TO SHANGHAI JIAO TONG UNIV SCHOOL OF MEDICINE
- Filing Date
- 2026-05-12
- Publication Date
- 2026-08-07
AI Technical Summary
[0002]当前,随着医疗健康与人工智能技术的深度融合,非接触式生理信号监测与ECG生成需求日益迫切,传统单一模态的生理信号采集与处理方式存在明显局限,一方面,单一的脉搏波提取方式易受干扰,难以全面捕捉人体生理特征,且不同模态信号的特征空间不统一,导致跨模态信息传递存在壁垒;另一方面,目前多采用单一模态训练或单一信号提取方式,缺乏对多源脉搏波信号的整合与优化,也未针对ECG生成的核心需求设计专属训练流程,使得脉搏波特征与ECG特征的映射不够精准,ECG解码器的通用适配性不足,同时缺乏对特征转换模块的系统训练,导致多源脉搏波信号的优势无法充分发挥,难以实现非接触式ECG信号的精准生成,此外,目前缺乏对多模态特征的协同优化与跨域适配训练,易出现特征空间错位、生成误差较大等问题,无法满足临床及日常应用中对ECG信号生成准确性、稳定性的需求
[0014]本申请有益效果为:本申请分别利用rPPG波形预测网络、POS算法、CHROM算法对采集的当前视频序列进行波形提取,以得到多源脉搏波信号;其中,所述多源脉搏波信号为rPPG信号、POS信号、CHROM信号;采用脉搏波信号编码器对所述多源脉搏波信号进行特征编码,得到各脉搏波特征;其中,所述脉搏波信号编码器为基于跨模态特征对齐机制预训练后的编码器;将各所述脉搏波特征进行加和平均聚合,以得到聚合特征;将所述聚合特征输入特征转换器,转换得到适配ECG特征空间的rPPG2ECG特征,并利用预训练的ECG解码器对所述rPPG2ECG特征进行解码,以生成ECG信号。由此可见,本申请通过分别利用rPPG波形预测网络、POS算法、CHROM算法从当前视频序列中提取rPPG信号、POS信号、CHROM信号形成多源脉搏波信号,可充分整合不同提取方式下脉搏波信号的优势,弥补单一脉搏波信号提取易受干扰、表征不全面的缺陷,提升脉搏波信号的多样性与可靠性;采用经跨模态特征对齐机制预训练后的脉搏波信号编码器对多源脉搏波信号进行特征编码,能够使各脉搏波特征处于统一的特征空间,消除不同模态脉搏波特征之间的域差异,提升特征编码的准确性与鲁棒性;通过对各脉搏波特征进行加和平均聚合,可有效融合多源脉搏波特征的有效信息,抑制单一特征的噪声干扰,增强聚合特征对生理信息的表征能力;将聚合特征输入特征转换器转换得到适配ECG特征空间的rPPG2ECG特征,能够实现脉搏波特征向ECG特征空间的精准映射,解决跨模态特征不兼容的问题,再利用预训练的ECG解码器对rPPG2ECG特征进行解码生成ECG信号,可显著提升ECG信号生成的准确性、稳定性与鲁棒性,降低ECG信号生成过程中的误差,同时无需依赖接触式生理信号采集设备,实现ECG信号的非接触式生成,拓展ECG信号生成的应用场景,提升使用便捷性,且整个流程依托多阶段预训练优化后的模型模块,进一步保障了ECG信号生成的高效性与可靠性。
Smart Images

Figure CN122528033A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of electrocardiogram (ECG) signal generation technology, and in particular to ECG signal generation methods, apparatus, devices, and media based on multi-source pulse wave signals. Background Technology
[0002] Currently, with the deep integration of medical and health technologies and artificial intelligence, the demand for non-contact physiological signal monitoring and ECG generation is becoming increasingly urgent. Traditional single-modality physiological signal acquisition and processing methods have significant limitations. On the one hand, single pulse wave extraction methods are easily interfered with and cannot fully capture human physiological characteristics. Moreover, the feature spaces of different modal signals are not uniform, resulting in barriers to cross-modal information transmission. On the other hand, most current methods use single-modality training or single signal extraction, lacking integration and optimization of multi-source pulse wave signals. They also lack dedicated training processes designed for the core needs of ECG generation, resulting in inaccurate mapping between pulse wave features and ECG features, insufficient universal adaptability of ECG decoders, and a lack of systematic training for feature conversion modules. This prevents the full utilization of the advantages of multi-source pulse wave signals, making it difficult to achieve accurate non-contact ECG signal generation. Furthermore, the current lack of collaborative optimization and cross-domain adaptation training for multi-modal features easily leads to problems such as feature space misalignment and large generation errors, failing to meet the requirements for accuracy and stability of ECG signal generation in clinical and daily applications.
[0003] In summary, how to achieve unified alignment of cross-modal features of multi-source pulse waves and accurately complete the cross-domain conversion and high-quality generation of electrocardiogram signals is a problem to be solved in this field. Summary of the Invention
[0004] In view of this, the purpose of this invention is to provide a method, apparatus, device, and medium for generating ECG signals based on multi-source pulse wave signals, achieving unified alignment of cross-modal features of multi-source pulse waves and accurately completing cross-domain conversion and high-quality generation of ECG signals. The specific solution is as follows: In a first aspect, this application discloses a method for generating ECG signals based on multi-source pulse wave signals, including: The rPPG waveform prediction network, POS algorithm, and CHROM algorithm are used to extract waveforms from the acquired current video sequence to obtain multi-source pulse wave signals; wherein, the multi-source pulse wave signals are rPPG signals, POS signals, and CHROM signals. The multi-source pulse wave signal is feature-encoded using a pulse wave signal encoder to obtain each pulse wave feature; wherein, the pulse wave signal encoder is an encoder pre-trained based on a cross-modal feature alignment mechanism; The pulse wave features described are summed, averaged, and aggregated to obtain aggregated features; The aggregated features are input into the feature converter to obtain rPPG2ECG features adapted to the ECG feature space, and the rPPG2ECG features are decoded using a pre-trained ECG decoder to generate an ECG signal.
[0005] Optionally, before extracting waveforms from the acquired current video sequence using the rPPG waveform prediction network, POS algorithm, and CHROM algorithm respectively, the method further includes: Collect historical video sequences and corresponding historical PPG waveforms; The initial rPPG waveform prediction network constructed is determined as the current rPPG waveform prediction network; The historical video sequence is input into the current rPPG waveform prediction network to obtain the predicted rPPG waveform; The mean square error between the predicted rPPG waveform and the historical PPG waveform is determined as the first waveform reconstruction loss value. The parameters of the current rPPG waveform prediction network are updated based on the first waveform reconstruction loss value to obtain the final rPPG waveform prediction network.
[0006] Optionally, the pulse wave signal encoder includes an rPPG encoder, a POS encoder, and a CHROM encoder; the step of using the pulse wave signal encoder to perform feature encoding on the multi-source pulse wave signal includes: The rPPG signal, POS signal, and CHROM signal are respectively encoded using an rPPG encoder, a POS encoder, and a CHROM encoder to obtain the first rPPG feature, the first POS feature, and the first CHROM feature. Accordingly, the summing and averaging of the pulse wave features includes: The first rPPG feature, the first POS feature, and the first CHROM feature are summed and averaged.
[0007] Optionally, before using a pulse wave signal encoder to perform feature encoding on the multi-source pulse wave signal to obtain each pulse wave feature, the method further includes: The collected historical raw physiological signals of various types are randomly masked to obtain the masked signals; wherein, the historical raw physiological signals of various types include historical multi-source pulse wave signals and historical ECG signals; Construct the current encoder and current decoder respectively adapted to rPPG signal, POS signal, CHROM signal and ECG signal; The current encoder and the current decoder are trained and their parameters are updated using the masked signal and the historical ECG signal to obtain the trained encoder and the trained decoder. The trained encoder is optimized based on a cross-modal feature alignment mechanism to obtain the final rPPG encoder, the final POS encoder, and the final CHROM encoder.
[0008] Optionally, the encoder includes an rPPG encoder, a POS encoder, a CHROM encoder, and an ECG encoder, and the decoder includes an rPPG decoder, a POS decoder, a CHROM decoder, and an ECG decoder; the step of training and updating the parameters of the current encoder and the current decoder using the masked signal and the historical ECG signal to obtain the trained encoder and the trained decoder includes: Each masked signal is input into the corresponding current encoder for feature encoding to obtain various masked physiological features. The various masked physiological features are input into the corresponding current decoder for signal reconstruction, and various reconstructed waveform signals are output. The original real physiological waveforms corresponding to the various reconstructed waveform signals are also obtained. The mean square error between the reconstructed waveform signal and the original real physiological waveform is determined as the second waveform reconstruction loss value; Based on the second waveform reconstruction loss value, backpropagation is performed and the network parameters of the current encoder and the current decoder are iteratively updated to obtain the trained encoder and the trained decoder.
[0009] Optionally, the optimization of the trained encoder based on the cross-modal feature alignment mechanism to obtain the final rPPG encoder, the final POS encoder, and the final CHROM encoder includes: The historical multi-class raw physiological signals are respectively input into the trained encoder to obtain the second rPPG feature, the second POS feature, the second CHROM feature, and the first ECG feature; The second rPPG feature, the second POS feature, and the first CHROM feature are summed, averaged, and aggregated to obtain the first historical aggregated feature. The second rPPG feature, second POS feature, second CHROM feature, first ECG feature, and first historical aggregation feature corresponding to the same historical video sequence are used as positive sample pairs, and the second rPPG feature, second POS feature, second CHROM feature, first ECG feature, and first historical aggregation feature corresponding to different historical video sequences are used as negative samples to construct a contrastive learning sample set. The information noise contrast estimation loss function is used to calculate the sample similarity results of the contrast learning sample set respectively. Based on the sample similarity results, the alignment difference of different modal features is quantified to obtain the contrast learning loss value. Based on the contrastive learning loss value, backpropagation is performed and the network parameters of the trained encoder are iteratively optimized to obtain the final rPPG encoder, the final POS encoder, and the final CHROM encoder.
[0010] Optionally, before inputting the aggregated features into the feature converter to obtain rPPG2ECG features adapted to the ECG feature space, the method further includes: The final rPPG encoder, the final POS encoder, the final CHROM encoder, and the trained ECG encoder were used to encode the historical multi-source pulse wave signal and the historical ECG signal to obtain the third rPPG feature, the third POS feature, the third CHROM feature, and the second ECG feature. The third rPPG feature, the third POS feature, and the third CHROM feature are summed, averaged, and aggregated to obtain the second historical aggregated feature; The current feature converter can be constructed using any one of the following networks: multilayer perceptron, convolutional neural network, or Transformer architecture. The second historical aggregated feature is input into the current feature converter to obtain the historical rPPG2ECG feature; The historical rPPG2ECG features and the second ECG features are respectively input into the trained ECG encoder to generate the first reconstructed waveform and the second reconstructed waveform. The mean square error between the historical rPPG2ECG feature and the second ECG feature is determined as the feature alignment loss value; Obtain the original true ECG waveform corresponding to the second ECG feature, and determine the first mean square error between the original true ECG waveform and the second reconstructed waveform, the second mean square error between the original true ECG waveform and the first reconstructed waveform, and the third mean square error between the first reconstructed waveform and the second reconstructed waveform. The average value of the first mean square error, the second mean square error and the third mean square error is determined as the waveform alignment loss value. Based on the feature alignment loss value and the waveform alignment loss value, backpropagation is performed and the network parameters of the current feature converter and the trained ECG encoder are iteratively updated to obtain the pre-trained feature converter and the pre-trained ECG decoder.
[0011] Secondly, this application discloses an ECG signal generation device based on multi-source pulse wave signals, comprising: The waveform extraction module is used to extract waveforms from the acquired current video sequence using the rPPG waveform prediction network, the POS algorithm, and the CHROM algorithm, respectively, to obtain multi-source pulse wave signals; wherein, the multi-source pulse wave signals are rPPG signals, POS signals, and CHROM signals; The feature encoding module is used to encode the features of the multi-source pulse wave signal using a pulse wave signal encoder to obtain each pulse wave feature; wherein, the pulse wave signal encoder is an encoder pre-trained based on a cross-modal feature alignment mechanism; An aggregation module is used to sum and average the pulse wave features to obtain aggregated features; The signal generation module is used to input the aggregated features into the feature converter, convert them into rPPG2ECG features adapted to the ECG feature space, and use a pre-trained ECG decoder to decode the rPPG2ECG features to generate an ECG signal.
[0012] Thirdly, this application discloses an electronic device, including: Memory, used to store computer programs; A processor is configured to execute the computer program to implement the steps of the aforementioned disclosed ECG signal generation method based on multi-source pulse wave signals.
[0013] Fourthly, this application discloses a computer-readable storage medium for storing a computer program; wherein, when the computer program is executed by a processor, it implements the steps of the aforementioned disclosed ECG signal generation method based on multi-source pulse wave signals.
[0014] The beneficial effects of this application are as follows: This application utilizes an rPPG waveform prediction network, a POS algorithm, and a CHROM algorithm to extract waveforms from the acquired current video sequence to obtain multi-source pulse wave signals; wherein, the multi-source pulse wave signals are rPPG signals, POS signals, and CHROM signals; a pulse wave signal encoder is used to perform feature encoding on the multi-source pulse wave signals to obtain each pulse wave feature; wherein, the pulse wave signal encoder is an encoder pre-trained based on a cross-modal feature alignment mechanism; the pulse wave features are summed, averaged, and aggregated to obtain aggregated features; the aggregated features are input into a feature converter to obtain rPPG2ECG features adapted to the ECG feature space, and a pre-trained ECG decoder is used to decode the rPPG2ECG features to generate an ECG signal. Therefore, this application utilizes an rPPG waveform prediction network, POS algorithm, and CHROM algorithm to extract rPPG, POS, and CHROM signals from the current video sequence to form a multi-source pulse wave signal. This fully integrates the advantages of pulse wave signals extracted using different methods, overcomes the shortcomings of single pulse wave signal extraction being susceptible to interference and having incomplete representation, and improves the diversity and reliability of the pulse wave signal. The use of a pulse wave signal encoder pre-trained with a cross-modal feature alignment mechanism for feature encoding of the multi-source pulse wave signal ensures that each pulse wave feature is in a unified feature space, eliminates domain differences between different modal pulse wave features, and improves the accuracy and robustness of feature encoding. By summing and averaging the pulse wave features, the effective information of the multi-source pulse wave features can be effectively fused, suppressing the influence of single features. This process reduces noise interference and enhances the ability of aggregated features to represent physiological information. The aggregated features are input to a feature converter to obtain rPPG2ECG features adapted to the ECG feature space, enabling accurate mapping of pulse wave features to the ECG feature space and resolving the incompatibility issue of cross-modal features. A pre-trained ECG decoder is then used to decode the rPPG2ECG features and generate ECG signals, significantly improving the accuracy, stability, and robustness of ECG signal generation and reducing errors during the generation process. Furthermore, it eliminates the need for contact-based physiological signal acquisition equipment, enabling non-contact ECG signal generation, expanding application scenarios, and improving ease of use. The entire process relies on multi-stage pre-trained and optimized model modules, further ensuring the efficiency and reliability of ECG signal generation. Attached Figure Description
[0015] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on the provided drawings without creative effort.
[0016] Figure 1 This is a flowchart of an ECG signal generation method based on multi-source pulse wave signals disclosed in this application; Figure 2 This is a schematic diagram of a specific rPPG waveform prediction pre-training method disclosed in this application; Figure 3 This is a schematic diagram of a specific feature learning pre-training method disclosed in this application; Figure 4 This is a schematic diagram of a specific cross-modal feature alignment pre-training method disclosed in this application; Figure 5 This is a schematic diagram of a specific rPPG to ECG waveform reconstruction training disclosed in this application; Figure 6 This is a schematic diagram illustrating a specific rPPG to ECG waveform generation and prediction method disclosed in this application; Figure 7 This is a schematic diagram of the structure of an ECG signal generation device based on multi-source pulse wave signals disclosed in this application; Figure 8 This is a structural diagram of an electronic device disclosed in this application. Detailed Implementation
[0017] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of the present invention.
[0018] Cardiovascular disease is the leading cause of death worldwide. Its early symptoms are often insidious, and some arrhythmias are paroxysmal. Therefore, real-time, continuous monitoring of cardiac physiological status, as well as early detection and risk screening for cardiovascular diseases, are crucial for reducing mortality and improving patient prognosis. Electrocardiography (ECG), as a core technology for capturing cardiac electrophysiological activity, accurately reflects the temporal characteristics and morphological changes of myocardial depolarization and repolarization. It is the gold standard for clinical diagnosis of various cardiovascular diseases such as atrial fibrillation, myocardial ischemia, and arrhythmias, and is also an important basis for cardiovascular health assessment. However, traditional ECG signal acquisition relies on specialized contact-based testing equipment and requires professional operation. The acquisition process is significantly limited by the scenario and equipment, making it difficult to achieve contactless, long-term continuous monitoring in environments such as home, bedside, and public spaces. It also cannot meet the convenient needs of routine health screening and monitoring of special populations (the elderly, children, burn patients, etc.).
[0019] Photoplethysmography (PPG) signals reflect the hemodynamic characteristics of the cardiovascular system by capturing changes in blood perfusion. Its non-invasive nature and ease of acquisition have led to its widespread integration into various wearable devices, making it a crucial technology for convenient cardiovascular monitoring, potentially replacing ECG. However, traditional contact-based PPG signal acquisition still requires direct contact with the skin, which can cause skin irritation and discomfort with prolonged use. Furthermore, the stability of contact-based acquisition can be compromised in certain medical monitoring and exercise monitoring scenarios. Additionally, PPG signals only reflect changes in blood flow and lack direct cardiac electrophysiological characteristics, thus failing to replace ECG for accurate diagnosis of cardiovascular diseases. The emergence of remote photoplethysmography (rPPG) technology breaks through the limitations of contact-based physiological signal acquisition. This technology can extract pulse wave-related physiological features non-contactly using facial video sequences, eliminating the need for direct skin contact. Signal acquisition can be achieved using common hardware such as ordinary cameras, smartphones, and smart monitoring devices. The acquisition process is convenient and non-invasive, adaptable to various scenarios including home monitoring, bedside medical monitoring, and public health screenings, making it an important direction for non-contact cardiovascular health monitoring.
[0020] However, rPPG signals are highly susceptible to interference from complex environmental factors such as changes in lighting, facial movement, and shooting posture. Single rPPG signals predicted directly by deep learning lack robustness and are prone to waveform distortion and feature shift in non-ideal acquisition scenarios. Traditional optical physiological measurement methods such as Plane Orthogonal-to-Skin (POS) and Chrominance-based RPG (CHROM) possess inherent robustness to environmental interference and can effectively suppress the influence of factors such as lighting and motion on pulse wave signals. They are important technical means to improve rPPG signal quality. However, these traditional robust algorithms have not yet been combined with deep learning for applications in rPPG signal enhancement and ECG signal conversion. Existing rPPG to ECG conversion schemes rely solely on rPPG signals predicted by a single network as input, without incorporating robust features such as POS and CHROM to enhance the signal. Furthermore, they lack effective aggregation and cross-modal alignment mechanisms for multi-source rPPG features, making it impossible to maintain the stability and fidelity of signal conversion in complex environments. The generated ECG signals are easily distorted by environmental interference, making it difficult to reproduce clinically valuable electrophysiological characteristics and failing to meet the practical needs of contactless cardiovascular monitoring in complex scenarios.
[0021] Against this backdrop, how to integrate traditional robust optical physiology algorithms with deep learning technology, combining their anti-interference characteristics with the feature learning capabilities of deep networks, to construct a multi-source rPPG feature aggregation-enhanced rPPG-to-ECG waveform generation method, improve the robustness and fidelity of signal conversion in complex environments, and achieve contactless ECG generation with both environmental adaptability and diagnostic reference value, has become an urgent problem to be solved in the field of cardiovascular health monitoring technology. This technological breakthrough can significantly improve the usability of rPPG-to-ECG solutions in real-world complex scenarios, providing more reliable technical support for continuous cardiovascular health monitoring in home, mobile, and public settings. It has significant clinical and practical value for expanding the application scope of contactless cardiac monitoring and improving the efficiency of early screening for cardiovascular diseases.
[0022] To this end, this application provides an ECG signal generation scheme based on multi-source pulse wave signals, which realizes unified alignment of cross-modal features of multi-source pulse waves and accurately completes cross-domain conversion and high-quality generation of ECG signals.
[0023] See Figure 1 As shown in the figure, this application discloses a method for generating ECG signals based on multi-source pulse wave signals, including: Step S11: The current video sequence is extracted using the rPPG waveform prediction network, POS algorithm, and CHROM algorithm to obtain multi-source pulse wave signals; wherein the multi-source pulse wave signals are rPPG signals, POS signals, and CHROM signals.
[0024] The rPPG (remote photoplethysmography) waveform prediction network, the POS (Plane Orthogonal to Skin) algorithm, and the CHROM (Chrominance) algorithm are used to extract waveforms from the acquired current video sequence, resulting in a multi-source pulse wave signal composed of rPPG, POS, and CHROM signals. In other words, relying on mature algorithms and a trained rPPG waveform prediction network, rPPG, POS, and CHROM waveforms are directly and synchronously acquired from the input face video, completing the synchronous acquisition and output of multiple types of raw pulse wave signals.
[0025] In this embodiment, before extracting waveforms from the acquired current video sequence using the rPPG waveform prediction network, POS algorithm, and CHROM algorithm respectively, the method further includes: acquiring historical video sequences and historical PPG waveforms corresponding to the historical video sequences; determining the constructed initial rPPG waveform prediction network as the current rPPG waveform prediction network; inputting the historical video sequences into the current rPPG waveform prediction network to obtain predicted rPPG waveforms; determining the mean square error between the predicted rPPG waveform and the historical PPG waveforms as a first waveform reconstruction loss value; and updating the parameters of the current rPPG waveform prediction network based on the first waveform reconstruction loss value to obtain the final rPPG waveform prediction network.
[0026] like Figure 2 As shown, before extracting the waveform from the video sequence, historical video sequences and matching historical PPG waveforms are pre-acquired. An initial rPPG waveform prediction network is constructed using Physformer, Transformer-based, or convolutional neural network structures. The historical video sequence is input to obtain the predicted rPPG waveform. The first waveform reconstruction loss value is constructed based on the mean square error between the predicted waveform and the real historical PPG waveform. The network parameters are iteratively updated based on the loss value to complete the training and optimization of the rPPG waveform prediction network. In other words, the network learning is supervised by real labeled historical physiological data, and the model parameters are iterated by constraining the waveform reconstruction error, thereby continuously improving the accuracy and generalization ability of the rPPG waveform prediction network for extracting physiological waveforms from face videos.
[0027] Step S12: Use a pulse wave signal encoder to perform feature encoding on the multi-source pulse wave signal to obtain each pulse wave feature; wherein, the pulse wave signal encoder is an encoder pre-trained based on a cross-modal feature alignment mechanism.
[0028] In this embodiment, the pulse wave signal encoder includes an rPPG encoder, a POS encoder, and a CHROM encoder; the step of using the pulse wave signal encoder to perform feature encoding on the multi-source pulse wave signal includes: using the rPPG encoder, POS encoder, and CHROM encoder to perform feature encoding on the rPPG signal, POS signal, and CHROM signal respectively to obtain a first rPPG feature, a first POS feature, and a first CHROM feature.
[0029] A pulse wave signal encoder, pre-trained using a cross-modal feature alignment mechanism, is employed to perform feature encoding on multi-source pulse wave signals and output corresponding pulse wave features. This pulse wave signal encoder specifically includes three independent encoding modules: an rPPG encoder, a POS encoder, and a CHROM encoder. In the actual encoding process, the rPPG encoder, POS encoder, and CHROM encoder are used to perform dedicated feature extraction and encoding operations on the rPPG signal, POS signal, and CHROM signal, respectively, thereby obtaining the first rPPG feature, the first POS feature, and the first CHROM feature. Relying on the differentiated dedicated encoders to adapt to the signal distribution characteristics of different types of pulse wave signals, combined with the cross-modal alignment capability formed in the pre-training stage, it is ensured that all types of encoded features have a unified feature space representation basis, providing standardized and effective feature information for subsequent multi-feature fusion and cross-domain conversion.
[0030] In this embodiment, before using a pulse wave signal encoder to perform feature encoding on the multi-source pulse wave signal to obtain each pulse wave feature, the method further includes: performing random masking processing on the collected historical multi-type raw physiological signals to obtain masked signals; wherein, the historical multi-type raw physiological signals include historical multi-source pulse wave signals and historical ECG signals; constructing current encoders and current decoders adapted to rPPG signals, POS signals, CHROM signals, and ECG signals respectively; using the masked signals and the historical ECG signals to train and update the parameters of the current encoders and current decoders to obtain trained encoders and trained decoders; optimizing the trained encoder based on a cross-modal feature alignment mechanism to obtain the final rPPG encoder, the final POS encoder, and the final CHROM encoder.
[0031] Before completing the feature encoding of multi-source pulse wave signals and acquiring various pulse wave features through the pulse wave signal encoder, multiple types of raw physiological signals, including historical multi-source pulse wave signals and historical ECG signals, are pre-collected. Random masking processing is performed on the collected historical raw physiological signals to generate masked signals. At the same time, dedicated encoder and decoder network structures adapted to rPPG signals, POS signals, CHROM signals, and ECG signals are built. The masked physiological signals are combined with historical ECG signals to complete the supervised training and parameter iterative update of the initial encoder and decoder, obtaining a trained encoder and decoder with basic waveform reconstruction capabilities. Then, relying on the cross-modal feature alignment mechanism, the encoders that have completed basic training are further optimized and adjusted to reduce the feature domain differences between different physiological signal modalities. Finally, the parameters of the rPPG encoder, POS encoder, and CHROM encoder are converged and the performance is solidified, resulting in three types of final encoders that can be adapted to cross-modal fusion tasks.
[0032] In this embodiment, the encoder includes an rPPG encoder, a POS encoder, a CHROM encoder, and an ECG encoder, and the decoder includes an rPPG decoder, a POS decoder, a CHROM decoder, and an ECG decoder.
[0033] like Figure 3 As shown, the entire network architecture includes four independent encoders and four corresponding decoders. The rPPG encoder is specifically used for deep feature extraction of remote photoplethysmography pulse wave signals. The POS encoder is used to parse the implicit features of the pulse signal output by the skin orthogonal plane algorithm. The CHROM encoder is responsible for mining the effective physiological representation of the chromatic pulse wave signal. The ECG encoder focuses on completing the feature encoding mapping of the standard ECG signal. The rPPG decoder, POS decoder, and CHROM decoder can respectively restore and reconstruct the original pulse wave waveform under their respective modalities. The ECG decoder can realize the decoding output of ECG features to complete ECG waveform. Each encoder and decoder is matched and adapted to the corresponding physiological signal type, and undertakes the encoding representation and waveform reconstruction tasks of different modal physiological signals. This provides a complete network foundation module for subsequent mask self-supervised pre-training, cross-modal feature alignment, and ECG signal generation.
[0034] In this embodiment, the step of training and updating the current encoder and decoder using the masked signal and the historical ECG signal to obtain the trained encoder and decoder includes: inputting each of the masked signals into the corresponding current encoder for feature encoding to obtain various masked physiological features; inputting each of the masked physiological features into the corresponding current decoder for signal reconstruction, outputting various reconstructed waveform signals, and obtaining the original real physiological waveforms corresponding to each of the reconstructed waveform signals; determining the mean square error between the reconstructed waveform signal and the original real physiological waveform as the second waveform reconstruction loss value; and backpropagating and iteratively updating the network parameters of the current encoder and decoder based on the second waveform reconstruction loss value to obtain the trained encoder and decoder.
[0035] Each masked signal is input into the corresponding current encoder to complete feature encoding, generating various masked physiological features. The masked physiological features are then sent to the corresponding decoder to perform signal reconstruction, outputting various reconstructed waveform signals and matching and retrieving their corresponding original real physiological waveforms. The mean square error between the reconstructed waveform signal and the original real physiological waveform is calculated and set as the second waveform reconstruction loss value. Backpropagation is performed based on this loss value to iteratively update the network parameters of all current encoders and decoders, completing the synchronous training of multiple encoding and decoding modules, and thus obtaining the trained encoder and the trained decoder. In other words, the rPPG waveform, POS waveform, CHROM waveform, and ECG waveform after random masking are sequentially input into their respective independent encoders to complete feature extraction. Various encoders can adopt basic architectures such as Transformer, convolutional neural network, or recurrent neural network. The POS encoder and CHROM encoder reuse the unified network structure of the rPPG encoder. The encoded features are reconstructed by the matched decoder. The POS decoder and CHROM decoder adopt the same structural design as the rPPG decoder. The complete waveform data is reconstructed based on the mask signal. The reconstruction loss is constructed by calculating the mean square error between the reconstructed waveform and the original real waveform without masking. The parameters of each encoder and decoder are continuously updated by backpropagation with this loss as a constraint. The pre-training optimization of all encoders and decoders is completed by relying on the self-supervised waveform reconstruction method.
[0036] In this embodiment, optimizing the trained encoder based on the cross-modal feature alignment mechanism to obtain the final rPPG encoder, the final POS encoder, and the final CHROM encoder includes: inputting the historical multi-class raw physiological signals into the trained encoder to obtain the second rPPG feature, the second POS feature, the second CHROM feature, and the first ECG feature; summing and averaging the second rPPG feature, the second POS feature, and the first CHROM feature to obtain the first historical aggregated feature; and using the second rPPG feature, the second POS feature, the second CHROM feature, and the first ECG feature corresponding to the same historical video sequence to obtain the first historical aggregated feature. The first historical aggregation feature is used as a positive sample pair, and the second rPPG feature, second POS feature, second CHROM feature, first ECG feature, and first historical aggregation feature corresponding to different historical video sequences are used as negative samples to construct a contrastive learning sample set. The information noise contrastive estimation loss function is used to calculate the sample similarity results of the contrastive learning sample set, and the alignment difference of different modal features is quantified based on the sample similarity results to obtain the contrastive learning loss value. Based on the contrastive learning loss value, backpropagation is performed and the network parameters of the trained encoder are iteratively optimized to obtain the final rPPG encoder, the final POS encoder, and the final CHROM encoder.
[0037] The encoder, which has undergone basic training, is input one by one with various historical raw physiological signals. The encoder outputs the second rPPG feature, the second POS feature, the second CHROM feature, and the first ECG feature. The second rPPG feature, the second POS feature, and the second CHROM feature are summed, averaged, and aggregated to generate the first historical aggregated feature. The single-modal features and historical aggregated features matched by the same historical video sequence are selected as positive sample pairs, and the features corresponding to different historical video sequences are selected as negative samples. This constructs a complete contrastive learning sample set. The sample similarity is calculated by using the information noise contrastive estimation loss function. The feature alignment difference between different physiological signal modalities is quantified, and the contrastive learning loss value is obtained. The gradient information is backpropagated based on the loss value, and the network parameters of each trained encoder are iteratively optimized. Finally, the rPPG encoder, POS encoder, and CHROM encoder with solidified parameters and cross-modal alignment are obtained.
[0038] like Figure 4 In other words, based on the pre-training of waveform mask reconstruction, the weights of each encoder after training are used as the parameter initialization benchmark. The rPPG encoder, POS encoder, CHROM encoder, and ECG encoder are used to encode the features of the corresponding waveforms respectively. Then, the features output by the three types of pulse wave encoders are summed and averaged to generate global aggregated features that can represent the physiological information of the face video. A positive and negative sample system is built based on multi-class single-modal features and aggregated features. Multi-modal features from the same sample source are set as positive samples, and heterogeneous features from different sample sources are set as negative samples. Information noise contrast estimation loss is used to construct contrastive learning training constraints to measure the distribution differences of different waveform features. The network parameters of each encoder are continuously updated. The feature encoding alignment of multi-modal waveforms is completed with the help of contrastive learning mechanism, realizing cross-modal optimization of each pulse wave encoder and completing the cross-modal feature alignment pre-training process.
[0039] Step S13: Sum and average the pulse wave features to obtain aggregated features.
[0040] In this embodiment, the summing and averaging of the pulse wave features includes summing and averaging the first rPPG feature, the first POS feature, and the first CHROM feature.
[0041] After completing the independent feature encoding of multi-source pulse wave signals and sequentially obtaining the first rPPG feature, the first POS feature, and the first CHROM feature, a unified fusion and aggregation operation needs to be performed on the three different modalities of pulse wave features. Specifically, the aggregation method is to add the first rPPG feature, the first POS feature, and the first CHROM feature dimension by dimension and then calculate the average value. By integrating the effective and complementary physiological representation information in multiple types of pulse wave features through the aggregation operation of addition and averaging, the weight ratio of features obtained by different extraction algorithms is balanced, the random noise and modality-specific interference of single signal features are weakened, and the overall physiological correlation features after the fusion of multi-source pulse wave signals are preserved, forming a fusion feature with stronger robustness and global representation ability. This provides a stable and unified fusion feature input basis for subsequent cross-feature space conversion and ECG signal decoding generation.
[0042] Step S14: Input the aggregated features into the feature converter to obtain rPPG2ECG features adapted to the ECG feature space, and use the pre-trained ECG decoder to decode the rPPG2ECG features to generate an ECG signal.
[0043] In this embodiment, before inputting the aggregated features into the feature converter to obtain rPPG2ECG features adapted to the ECG feature space, the method further includes: using the final rPPG encoder, the final POS encoder, the final CHROM encoder, and the trained ECG encoder to perform feature encoding on the historical multi-source pulse wave signal and the historical ECG signal, respectively, to obtain the third rPPG feature, the third POS feature, the third CHROM feature, and the second ECG feature; summing and averaging the third rPPG feature, the third POS feature, and the third CHROM feature to obtain the second historical aggregated feature; constructing the current feature converter using any one of the following networks: multilayer perceptron, convolutional neural network, or Transformer architecture; inputting the second historical aggregated feature into the current feature converter to obtain the historical rPPG2ECG feature; and combining the historical rPPG2ECG feature with the second ECG feature space. CG features are input into the trained ECG encoder to generate a first reconstructed waveform and a second reconstructed waveform. The mean square error between the historical rPPG2ECG features and the second ECG features is determined as the feature alignment loss value. The original real ECG waveform corresponding to the second ECG feature is obtained, and the first mean square error between the original real ECG waveform and the second reconstructed waveform, the second mean square error between the original real ECG waveform and the first reconstructed waveform, and the third mean square error between the first reconstructed waveform and the second reconstructed waveform are determined. The average of the first mean square error, the second mean square error, and the third mean square error is determined as the waveform alignment loss value. Based on the feature alignment loss value and the waveform alignment loss value, backpropagation is performed and the network parameters of the current feature converter and the trained ECG encoder are iteratively updated to obtain the pre-trained feature converter and the pre-trained ECG decoder.
[0044] like Figure 5As shown, before inputting the aggregated features into the feature converter to obtain rPPG2ECG features adapted to the ECG feature space, the final rPPG encoder, POS encoder, CHROM encoder, and trained ECG encoder are used to encode the historical multi-source pulse wave signals and historical ECG signals respectively, resulting in the third rPPG feature, third POS feature, third CHROM feature, and second ECG feature. The three types of pulse wave features are then summed, averaged, and aggregated to form the second historical aggregated feature. The current feature converter is built using any network structure from multilayer perceptron, convolutional neural network, or Transformer. The second historical aggregated feature is input into the current feature converter to obtain the historical rPPG2ECG feature. Then, the historical... The historical rPPG2ECG feature and the second ECG feature are input into the trained ECG decoder to generate the first and second reconstructed waveforms. The mean square error of the historical rPPG2ECG feature and the second ECG feature is calculated as the feature alignment loss value. The original real ECG waveform corresponding to the second ECG feature is collected. The mean square error between the original real ECG waveform and the second reconstructed waveform, the first reconstructed waveform, and the mean square error between the two sets of reconstructed waveforms are calculated in sequence. The average of the three types of mean square errors is taken as the waveform alignment loss value. The feature alignment loss value and the waveform alignment loss value are combined to complete backpropagation. The network parameters of the current feature converter and ECG decoder are iteratively optimized, and finally the pre-trained feature converter and pre-trained ECG decoder are obtained.
[0045] This stage involves a specialized training process for rPPG to ECG waveform reconstruction. It uses pre-trained pulse wave encoders (pre-trained and fixed with cross-modal feature alignment) and the previously trained ECG encoder as the parameter initialization basis. Corresponding historical physiological signals are encoded to obtain multi-modal features. These three types of pulse wave features are summed and averaged to generate global historical aggregated features that characterize the physiological information of facial videos. Feature converters are flexibly constructed using basic network structures such as multilayer perceptrons, convolutional neural networks, or Transformers. These feature converters map the aggregated features to the ECG feature space and output historical rPPG2ECG features. The ECG decoder then completes the process. Using the unified network structure from previous iterations and reusing existing training parameters, two reconstructed ECG waveforms are generated based on the converted cross-domain features and the original ECG features, respectively. The feature alignment loss is constructed by calculating the mean square error between the rPPG2ECG features and the standard ECG features. Simultaneously, the waveform alignment loss is obtained by calculating the mean square error between each pair of the real original ECG waveform and the different decoded output waveforms and taking the average. Relying on the dual losses to jointly constrain the iterative update of the model parameters, the dedicated adaptation training of the feature converter and the fine-tuning optimization of the ECG decoder are completed, enabling the overall network to have the core ability to generate high-quality ECG waveforms from multi-source pulse wave features across domains.
[0046] like Figure 6 As shown, in the actual signal generation stage, after the summation, averaging, and aggregation of multi-source pulse wave features are completed to obtain aggregated features, these aggregated features are used as input signals and fed into a pre-trained feature converter. This feature converter has been specifically trained in the early stages and possesses accurate cross-domain feature mapping capabilities. It can flexibly adapt to basic network architectures such as multilayer perceptrons, convolutional neural networks, or Transformers, and can efficiently transform aggregated features from the multi-source pulse wave feature space to the ECG feature space, eliminating the domain differences between the two modal features and outputting rPPG2ECG features adapted to the ECG signal generation task. The rPPG2ECG feature is then input into a pre-trained ECG decoder. This ECG decoder uses the network structure determined in the previous training and the parameters are fixed, giving it a stable ECG feature decoding capability. It can perform in-depth analysis and waveform reconstruction of the input rPPG2ECG feature, accurately decoding and generating a high-quality and stable ECG signal that conforms to physiological characteristics. The entire process does not require additional parameter training and iteration. Relying on the pre-trained and optimized model module, it achieves efficient and accurate cross-domain conversion of multi-source pulse wave features into ECG signals, completing the actual generation of non-contact ECG signals.
[0047] The beneficial effects of this application are as follows: This application utilizes an rPPG waveform prediction network, a POS algorithm, and a CHROM algorithm to extract waveforms from the acquired current video sequence to obtain multi-source pulse wave signals; wherein, the multi-source pulse wave signals are rPPG signals, POS signals, and CHROM signals; a pulse wave signal encoder is used to perform feature encoding on the multi-source pulse wave signals to obtain each pulse wave feature; wherein, the pulse wave signal encoder is an encoder pre-trained based on a cross-modal feature alignment mechanism; the pulse wave features are summed, averaged, and aggregated to obtain aggregated features; the aggregated features are input into a feature converter to obtain rPPG2ECG features adapted to the ECG feature space, and a pre-trained ECG decoder is used to decode the rPPG2ECG features to generate an ECG signal. Therefore, this application utilizes an rPPG waveform prediction network, POS algorithm, and CHROM algorithm to extract rPPG, POS, and CHROM signals from the current video sequence to form a multi-source pulse wave signal. This fully integrates the advantages of pulse wave signals extracted using different methods, overcomes the shortcomings of single pulse wave signal extraction being susceptible to interference and having incomplete representation, and improves the diversity and reliability of the pulse wave signal. The use of a pulse wave signal encoder pre-trained with a cross-modal feature alignment mechanism for feature encoding of the multi-source pulse wave signal ensures that each pulse wave feature is in a unified feature space, eliminates domain differences between different modal pulse wave features, and improves the accuracy and robustness of feature encoding. By summing and averaging the pulse wave features, the effective information of the multi-source pulse wave features can be effectively fused, suppressing the influence of single features. This process reduces noise interference and enhances the ability of aggregated features to represent physiological information. The aggregated features are input to a feature converter to obtain rPPG2ECG features adapted to the ECG feature space, enabling accurate mapping of pulse wave features to the ECG feature space and resolving the incompatibility issue of cross-modal features. A pre-trained ECG decoder is then used to decode the rPPG2ECG features and generate ECG signals, significantly improving the accuracy, stability, and robustness of ECG signal generation and reducing errors during the generation process. Furthermore, it eliminates the need for contact-based physiological signal acquisition equipment, enabling non-contact ECG signal generation, expanding application scenarios, and improving ease of use. The entire process relies on multi-stage pre-trained and optimized model modules, further ensuring the efficiency and reliability of ECG signal generation.
[0048] See Figure 7 As shown in the figure, this application discloses an ECG signal generation device based on multi-source pulse wave signals, comprising: The waveform extraction module 11 is used to extract waveforms from the acquired current video sequence using the rPPG waveform prediction network, the POS algorithm, and the CHROM algorithm, respectively, to obtain multi-source pulse wave signals; wherein, the multi-source pulse wave signals are rPPG signals, POS signals, and CHROM signals; The feature encoding module 12 is used to perform feature encoding on the multi-source pulse wave signal using a pulse wave signal encoder to obtain each pulse wave feature; wherein, the pulse wave signal encoder is an encoder pre-trained based on a cross-modal feature alignment mechanism; The aggregation module 13 is used to sum and average the pulse wave features to obtain aggregated features; The signal generation module 14 is used to input the aggregated features into the feature converter, convert them into rPPG2ECG features adapted to the ECG feature space, and use a pre-trained ECG decoder to decode the rPPG2ECG features to generate an ECG signal.
[0049] This application integrates a deep learning rPPG prediction network with traditional optical physiological measurement algorithms such as POS and CHROM to construct a multi-source signal input system. Leveraging the anti-interference advantages of POS and CHROM algorithms, it suppresses rPPG signal distortion and shift in complex scenarios, improving the stability and reliability of the input signal. Mask reconstruction pre-training is performed separately for four types of signals: rPPG, POS, CHROM, and ECG. Dedicated codecs are provided for each type of signal to enhance the model's feature extraction and waveform reconstruction capabilities, improving the quality of signal feature representation and network generalization. A three-source pulse wave feature summation and averaging aggregation strategy is adopted to generate… It can comprehensively characterize the robust global aggregation features of facial video physiological information and achieve complementary fusion of multi-source information. Through contrastive learning, it achieves deep spatial alignment of multi-source pulse wave features, aggregation features and ECG features, opens up cross-modal mapping channels and reduces conversion bias. It builds a feature converter and a dual-loss constraint ECG waveform reconstruction network, and relies on the joint optimization of feature alignment loss and waveform alignment loss to achieve accurate adaptation of aggregation features to ECG feature space and high-fidelity ECG waveform generation, so that the output waveform has clinical reference value. At the same time, relying on the non-contact advantage, it can be adapted to continuous cardiovascular monitoring in various complex scenarios.
[0050] Furthermore, embodiments of this application also provide an electronic device. Figure 8 This is a structural diagram of an electronic device 20 according to an exemplary embodiment. The content of the diagram should not be construed as limiting the scope of this application.
[0051] Figure 8This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Specifically, it may include: at least one processor 21, at least one memory 22, a power supply 23, a communication interface 24, an input / output interface 25, and a communication bus 26. The memory 22 stores a computer program, which is loaded and executed by the processor 21 to implement the relevant steps in the ECG signal generation method based on multi-source pulse wave signals disclosed in any of the foregoing embodiments.
[0052] In this embodiment, the power supply 23 is used to provide operating voltage for various hardware devices on the electronic device; the communication interface 24 can create a data transmission channel between the electronic device and external devices, and the communication protocol it follows can be any communication protocol applicable to the technical solution of this application, and is not specifically limited here; the input / output interface 25 is used to acquire external input data or output data to the outside world, and its specific interface type can be selected according to specific application needs, and is not specifically limited here.
[0053] The processor 21 may include one or more processing cores, such as a quad-core processor or an octa-core processor. The processor 21 may be implemented using at least one hardware form selected from DSP (Digital Signal Processing), FPGA (Field-Programmable Gate Array), and PLA (Programmable Logic Array). The processor 21 may also include a main processor and a coprocessor. The main processor, also known as a CPU (Central Processing Unit), is used to process data in the wake-up state; the coprocessor is a low-power processor used to process data in the standby state. In some embodiments, the processor 21 may integrate a GPU (Graphics Processing Unit), which is responsible for rendering and drawing the content to be displayed on the screen. In some embodiments, the processor 21 may also include an AI (Artificial Intelligence) processor, which is used to handle computational operations related to machine learning.
[0054] In addition, the memory 22, as a carrier for resource storage, can be a read-only memory, random access memory, disk or optical disk, etc. The resources stored on it include operating system 221, computer program 222 and data 223, etc., and the storage method can be temporary storage or permanent storage.
[0055] The operating system 221 manages and controls the various hardware devices and computer programs 222 on the electronic device to enable the processor 21 to perform calculations and processing on the massive amounts of data 223 in the memory 22. The operating system can be Windows, Unix, Linux, etc. The computer program 222, in addition to including a computer program capable of performing the ECG signal generation method based on multi-source pulse wave signals disclosed in any of the foregoing embodiments, may further include computer programs capable of performing other specific tasks. The data 223 may include data received by the electronic device from external devices, as well as data collected by its own input / output interface 25.
[0056] Furthermore, this application also discloses a computer-readable storage medium for storing a computer program; wherein, when the computer program is executed by a processor, it implements the aforementioned disclosed method for generating ECG signals based on multi-source pulse wave signals. Specific steps of this method can be found in the corresponding content disclosed in the foregoing embodiments, and will not be repeated here.
[0057] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on its differences from other embodiments. Similar or identical parts between embodiments can be referred to interchangeably. For the apparatus disclosed in the embodiments, since it corresponds to the method disclosed in the embodiments, the description is relatively simple; relevant parts can be referred to in the method section.
[0058] Those skilled in the art will further recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application. The steps of the methods or algorithms described in conjunction with the embodiments disclosed herein can be implemented directly in hardware, software modules executed by a processor, or a combination of both. The software module may be located in random access memory (RAM), memory, read-only memory (ROM), electrically programmable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), register, hard disk, removable disk, CD-ROM (Compact Disc Read-Only Memory), or any other form of storage medium known in the art.
[0059] Finally, it should be noted that in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
[0060] The foregoing has provided a detailed description of the ECG signal generation method, apparatus, device, and medium based on multi-source pulse wave signals provided by the present invention. Specific examples have been used to illustrate the principles and implementation methods of the present invention. The descriptions of the above embodiments are only intended to help understand the method and core ideas of the present invention. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of the present invention. Therefore, the content of this specification should not be construed as a limitation of the present invention.
Claims
1. A method for generating ECG signals based on multi-source pulse wave signals, characterized in that, include: The rPPG waveform prediction network, POS algorithm, and CHROM algorithm are used to extract waveforms from the acquired current video sequence to obtain multi-source pulse wave signals; wherein, the multi-source pulse wave signals are rPPG signals, POS signals, and CHROM signals. The multi-source pulse wave signal is feature-encoded using a pulse wave signal encoder to obtain each pulse wave feature; wherein, the pulse wave signal encoder is an encoder pre-trained based on a cross-modal feature alignment mechanism; The pulse wave features described are summed, averaged, and aggregated to obtain aggregated features; The aggregated features are input into the feature converter to obtain rPPG2ECG features adapted to the ECG feature space, and the rPPG2ECG features are decoded using a pre-trained ECG decoder to generate an ECG signal.
2. The ECG signal generation method based on multi-source pulse wave signals according to claim 1, characterized in that, Before extracting waveforms from the acquired current video sequence using the rPPG waveform prediction network, POS algorithm, and CHROM algorithm respectively, the process also includes: Collect historical video sequences and corresponding historical PPG waveforms; The initial rPPG waveform prediction network constructed is determined as the current rPPG waveform prediction network; The historical video sequence is input into the current rPPG waveform prediction network to obtain the predicted rPPG waveform; The mean square error between the predicted rPPG waveform and the historical PPG waveform is determined as the first waveform reconstruction loss value. The parameters of the current rPPG waveform prediction network are updated based on the first waveform reconstruction loss value to obtain the final rPPG waveform prediction network.
3. The ECG signal generation method based on multi-source pulse wave signals according to claim 1, characterized in that, The pulse wave signal encoder includes an rPPG encoder, a POS encoder, and a CHROM encoder; the feature encoding of the multi-source pulse wave signal using the pulse wave signal encoder includes: The rPPG signal, POS signal, and CHROM signal are respectively encoded using an rPPG encoder, a POS encoder, and a CHROM encoder to obtain the first rPPG feature, the first POS feature, and the first CHROM feature. Accordingly, the summing and averaging of the pulse wave features includes: The first rPPG feature, the first POS feature, and the first CHROM feature are summed and averaged.
4. The ECG signal generation method based on multi-source pulse wave signals according to claim 3, characterized in that, Before using a pulse wave signal encoder to perform feature encoding on the multi-source pulse wave signal to obtain the features of each pulse wave, the method further includes: The collected historical raw physiological signals of various types are randomly masked to obtain the masked signals; wherein, the historical raw physiological signals of various types include historical multi-source pulse wave signals and historical ECG signals; Construct the current encoder and current decoder respectively adapted to rPPG signal, POS signal, CHROM signal and ECG signal; The current encoder and the current decoder are trained and their parameters are updated using the masked signal and the historical ECG signal to obtain the trained encoder and the trained decoder. The trained encoder is optimized based on a cross-modal feature alignment mechanism to obtain the final rPPG encoder, the final POS encoder, and the final CHROM encoder.
5. The ECG signal generation method based on multi-source pulse wave signals according to claim 4, characterized in that, The encoder includes an rPPG encoder, a POS encoder, a CHROM encoder, and an ECG encoder; the decoder includes an rPPG decoder, a POS decoder, a CHROM decoder, and an ECG decoder. The step of training and updating the parameters of the current encoder and decoder using the masked signal and the historical ECG signal to obtain the trained encoder and decoder includes: Each masked signal is input into the corresponding current encoder for feature encoding to obtain various masked physiological features. The various masked physiological features are input into the corresponding current decoder for signal reconstruction, and various reconstructed waveform signals are output. The original real physiological waveforms corresponding to the various reconstructed waveform signals are also obtained. The mean square error between the reconstructed waveform signal and the original real physiological waveform is determined as the second waveform reconstruction loss value; Based on the second waveform reconstruction loss value, backpropagation is performed and the network parameters of the current encoder and the current decoder are iteratively updated to obtain the trained encoder and the trained decoder.
6. The ECG signal generation method based on multi-source pulse wave signals according to claim 5, characterized in that, The optimization of the trained encoder based on the cross-modal feature alignment mechanism to obtain the final rPPG encoder, the final POS encoder, and the final CHROM encoder includes: The historical multi-class raw physiological signals are respectively input into the trained encoder to obtain the second rPPG feature, the second POS feature, the second CHROM feature, and the first ECG feature; The second rPPG feature, the second POS feature, and the first CHROM feature are summed, averaged, and aggregated to obtain the first historical aggregated feature. The second rPPG feature, second POS feature, second CHROM feature, first ECG feature, and first historical aggregation feature corresponding to the same historical video sequence are used as positive sample pairs, and the second rPPG feature, second POS feature, second CHROM feature, first ECG feature, and first historical aggregation feature corresponding to different historical video sequences are used as negative samples to construct a contrastive learning sample set. The information noise contrast estimation loss function is used to calculate the sample similarity results of the contrast learning sample set respectively. Based on the sample similarity results, the alignment difference of different modal features is quantified to obtain the contrast learning loss value. Based on the contrastive learning loss value, backpropagation is performed and the network parameters of the trained encoder are iteratively optimized to obtain the final rPPG encoder, the final POS encoder, and the final CHROM encoder.
7. The ECG signal generation method based on multi-source pulse wave signals according to claim 6, characterized in that, Before inputting the aggregated features into the feature converter to obtain rPPG2ECG features adapted to the ECG feature space, the method further includes: The final rPPG encoder, the final POS encoder, the final CHROM encoder, and the trained ECG encoder were used to encode the historical multi-source pulse wave signal and the historical ECG signal to obtain the third rPPG feature, the third POS feature, the third CHROM feature, and the second ECG feature. The third rPPG feature, the third POS feature, and the third CHROM feature are summed, averaged, and aggregated to obtain the second historical aggregated feature; The current feature converter can be constructed using any one of the following networks: multilayer perceptron, convolutional neural network, or Transformer architecture. The second historical aggregated feature is input into the current feature converter to obtain the historical rPPG2ECG feature; The historical rPPG2ECG features and the second ECG features are respectively input into the trained ECG encoder to generate the first reconstructed waveform and the second reconstructed waveform. The mean square error between the historical rPPG2ECG feature and the second ECG feature is determined as the feature alignment loss value; Obtain the original true ECG waveform corresponding to the second ECG feature, and determine the first mean square error between the original true ECG waveform and the second reconstructed waveform, the second mean square error between the original true ECG waveform and the first reconstructed waveform, and the third mean square error between the first reconstructed waveform and the second reconstructed waveform. The average value of the first mean square error, the second mean square error and the third mean square error is determined as the waveform alignment loss value. Based on the feature alignment loss value and the waveform alignment loss value, backpropagation is performed and the network parameters of the current feature converter and the trained ECG encoder are iteratively updated to obtain the pre-trained feature converter and the pre-trained ECG decoder.
8. An ECG signal generation device based on multi-source pulse wave signals, characterized in that, include: The waveform extraction module is used to extract waveforms from the acquired current video sequence using the rPPG waveform prediction network, the POS algorithm, and the CHROM algorithm, respectively, to obtain multi-source pulse wave signals; wherein, the multi-source pulse wave signals are rPPG signals, POS signals, and CHROM signals; The feature encoding module is used to encode the features of the multi-source pulse wave signal using a pulse wave signal encoder to obtain each pulse wave feature; wherein, the pulse wave signal encoder is an encoder pre-trained based on a cross-modal feature alignment mechanism; An aggregation module is used to sum and average the pulse wave features to obtain aggregated features; The signal generation module is used to input the aggregated features into the feature converter, convert them into rPPG2ECG features adapted to the ECG feature space, and use a pre-trained ECG decoder to decode the rPPG2ECG features to generate an ECG signal.
9. An electronic device, characterized in that, include: Memory, used to store computer programs; A processor for executing the computer program to implement the steps of the ECG signal generation method based on multi-source pulse wave signals as described in any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, Used to store a computer program; wherein, when the computer program is executed by a processor, it implements the steps of the ECG signal generation method based on multi-source pulse wave signals as described in any one of claims 1 to 7.