ECG signal acquisition method and device based on rppg signal, equipment and medium

CN122511580APending Publication Date: 2026-08-04RUIJIN HOSPITAL AFFILIATED TO SHANGHAI JIAO TONG UNIV SCHOOL OF MEDICINE
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
RUIJIN HOSPITAL AFFILIATED TO SHANGHAI JIAO TONG UNIV SCHOOL OF MEDICINE
Filing Date
2026-05-12
Publication Date
2026-08-04

AI Technical Summary

Technical Problem

但传统ECG信号的采集依赖专业的接触式检测设备,且需要专业人员操作,采集过程受场景、设备限制较大,难以实现居家、床旁、公共场景等环境下的无接触、长期连续监测,也无法满足日常健康筛查、特殊人群(老人、儿童、烧伤患者等)监护的便捷化需求

Benefits of technology

[0015]The beneficial effects of this application are as follows: This application obtains the corresponding PPG signal based on the current rPPG signal, and inputs the PPG signal into the PPG encoder to obtain the rPPG conditional code; the initial masked words of the first word, the second word, and the spatial ECG word, which are all in a masked state, are determined as the current masked words; wherein, the first word is a word that meets the first preset low-frequency condition, and the second word is a word that meets the first preset high-frequency condition; using the rPPG conditional code as a constraint, the current masked words are masked and reconstructed using a bidirectional Transformer to obtain the reconstructed words; based on the confidence level of the reconstructed words, the target masked words in the current masked words are masked and demasked to obtain a new current masked word, and then jumps back to the state where the rPPG conditional code is the constraint. The process involves several steps: first, reconstructing the current masked word using a bidirectional Transformer until all masks are removed to obtain the target word; second, inputting the first and second feature words from the target word into a preset low-frequency decoder and a preset high-frequency decoder respectively to obtain target low-frequency and target high-frequency components; third, performing an inverse short-time Fourier transform on the target low-frequency and target high-frequency components to obtain a frequency domain ECG waveform; fourth, the first feature word being a feature word that meets a second preset low-frequency condition and the second feature word being a feature word that meets a second preset high-frequency condition; fifth, inputting the spatial feature words from the target word into an ECG decoder to obtain a spatial domain ECG waveform; and finally, fusing the frequency domain ECG waveform and the spatial domain ECG waveform to obtain the target ECG signal.Therefore, this application obtains the corresponding PPG signal based on the current rPPG signal, obtains the rPPG conditional code by passing the PPG signal through a PPG encoder, and uses this code as a constraint to iteratively reconstruct the first word, the second word, and the spatial ECG word of the initial full mask. The mask is gradually removed based on the confidence level of the reconstructed word until all word reconstruction is completed. This enables accurate recovery of ECG features under the physiological characteristic constraints of the rPPG signal, ensuring a high degree of matching between the reconstruction process and human physiological rhythms, thus improving the authenticity and reliability of ECG features. The phased iterative mask reconstruction method can gradually optimize the word reconstruction accuracy, avoiding the feature ambiguity and error accumulation problems caused by a one-time full reconstruction. Simultaneously, the confidence level-based mask removal strategy prioritizes the retention of high-confidence features, further improving the overall reconstruction effect and efficiency. The reconstructed first frequency... The first and second frequency band feature words are respectively decoded to recover the first and second frequency band components, and then the frequency domain ECG waveform is obtained through inverse short-time Fourier transform. At the same time, the spatial domain feature words are decoded to obtain the spatial domain ECG waveform, realizing the dual-domain synchronous reconstruction of the ECG signal in the frequency domain and the time and spatial domains. This fully utilizes the rhythmic information of the frequency domain component and the morphological information of the spatial domain component, making the ECG signal more closely resemble the real signal in terms of rhythm and waveform. Finally, the frequency domain ECG waveform and the spatial domain ECG waveform are fused, which can complement the defects of single-domain reconstruction and take into account the overall rhythmic characteristics and local detail characteristics of the ECG signal. Based on the non-contact ECG acquisition using rPPG signals, this effectively improves the integrity, accuracy and waveform fidelity of the final target ECG signal, while simplifying the signal conversion link and improving the overall signal acquisition efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122511580A_ABST
    Figure CN122511580A_ABST
Patent Text Reader

Abstract

The application discloses an ECG signal acquisition method and device based on an rPPG signal, equipment and a medium, relates to the technical field of computer software, and comprises the following steps: obtaining a PPG signal based on an rPPG signal, and generating rPPG conditional coding through a PPG encoder. The low-frequency, high-frequency and spatial domain ECG tokens of the mask state are used as the current mask token, the rPPG conditional coding is used as the constraint, and the mask reconstruction is completed through a bidirectional Transformer. The mask is gradually removed according to the confidence, and the reconstruction is circularly performed until all the masks are removed to obtain target tokens. The low-frequency and high-frequency feature tokens in the target tokens are subjected to corresponding decoding and inverse short-time Fourier transform to generate a frequency domain ECG waveform; and the spatial domain feature tokens are subjected to an ECG decoder to obtain a spatial domain ECG waveform. The target ECG signal is obtained by fusing the two waveforms, and a high-precision and high-fidelity ECG signal can be obtained based on the rPPG signal.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of computer software technology, and in particular to a method, apparatus, device, and medium for acquiring ECG signals based on rPPG signals. Background Technology

[0002] Cardiovascular disease is the leading cause of death worldwide. Its early symptoms are often insidious, and some arrhythmias are paroxysmal. Therefore, real-time, continuous monitoring of cardiac physiological status, as well as early detection and risk screening for cardiovascular diseases, are crucial for reducing mortality and improving patient prognosis. Electrocardiography (ECG), as a core technology for capturing cardiac electrophysiological activity, accurately reflects the temporal characteristics and morphological changes of myocardial depolarization and repolarization. It is the gold standard for clinical diagnosis of various cardiovascular diseases such as atrial fibrillation, myocardial ischemia, and arrhythmias, and is also an important basis for cardiovascular health assessment. However, traditional ECG signal acquisition relies on specialized contact-based testing equipment and requires professional operation. The acquisition process is significantly limited by the scenario and equipment, making it difficult to achieve contactless, long-term continuous monitoring in environments such as home, bedside, and public spaces. It also cannot meet the convenient needs of routine health screening and monitoring of special populations (the elderly, children, burn patients, etc.).

[0003] While non-contact monitoring technology based on photoplethysmography (PPG) can alleviate some of the problems, its signals are easily affected by environmental interference and cannot directly reflect the characteristics related to electrocardiographic activity. Existing signal conversion schemes based on remote photoplethysmography (rPPG) mostly adopt a single-mode mapping architecture, performing simple feature fitting only in the time or frequency domain. This results in problems such as insufficient utilization of signal features, high distortion of reconstructed waveforms, and difficulty in synchronizing and preserving rhythm and morphology. Furthermore, the model has weak generalization ability and cannot guarantee signal conversion accuracy under complex interference.

[0004] In summary, how to obtain ECG signals with higher accuracy and fidelity based on rPPG signals is a problem that needs to be solved in this field. Summary of the Invention

[0005] In view of this, the purpose of this invention is to provide a method, apparatus, device, and medium for acquiring ECG signals based on rPPG signals, which achieves higher accuracy and fidelity in acquiring ECG signals. The specific solution is as follows: In a first aspect, this application discloses a method for acquiring ECG signals based on rPPG signals, including: The corresponding PPG signal is obtained based on the current rPPG signal, and the PPG signal is input into the PPG encoder to obtain the rPPG conditional code; The initial masked word of the first word, the second word, and the spatial ECG word, all of which are in a masked state, is determined as the current masked word; wherein, the first word is a word that meets the first preset low-frequency condition, and the second word is a word that meets the first preset high-frequency condition; Using the rPPG conditional encoding as a constraint, and through a bidirectional Transformer, the current masked tokens are reconstructed to obtain the reconstructed tokens; Based on the confidence level of the reconstructed lexical, the target lexical in the current masked lexical is masked to obtain a new current masked lexical. Then, the process jumps back to the step of reconstructing the current masked lexical using the rPPG conditional encoding as a constraint and a bidirectional Transformer, until all masks are removed to obtain the target lexical. The first and second feature words in the target word are respectively input into a preset low-frequency decoder and a preset high-frequency decoder to obtain target low-frequency components and target high-frequency components. The target low-frequency components and the target high-frequency components are subjected to inverse short-time Fourier transform to obtain frequency domain ECG waveforms. The first feature word is a feature word that meets the second preset low-frequency condition and the second feature word is a feature word that meets the second preset high-frequency condition. The spatial feature words in the target words are input into the ECG decoder to obtain the spatial ECG waveform; The frequency domain ECG waveform and the spatial domain ECG waveform are fused to obtain the target ECG signal.

[0006] Optionally, the step of acquiring the corresponding PPG signal based on the current rPPG signal and inputting the PPG signal into the PPG encoder includes: The current rPPG signal is input into the rPPG2PPG encoder to obtain the corresponding PPG signal, and the PPG signal is then input into the PPG encoder. Accordingly, obtaining the rPPG2PPG encoder includes: Collect historical rPPG signals and historical PPG signals corresponding to the historical rPPG signals; The current rPPG2PPG encoder is constructed based on a transformer or a convolutional neural network; The historical rPPG signal and the historical PPG signal are determined as training data; The current rPPG2PPG encoder is iteratively trained using the training data, and updated based on waveform reconstruction loss value and heart rate prediction loss value to obtain the final rPPG2PPG encoder.

[0007] Optionally, the waveform reconstruction loss value is the mean square error of the waveform between the predicted PPG signal output by the current rPPG2PPG encoder and the historical PPG signal, and the heart rate prediction loss value is the absolute heart rate error between the predicted PPG signal and the historical PPG signal extracted by fast Fourier transform.

[0008] Optionally, before acquiring the corresponding PPG signal based on the current rPPG signal and inputting the PPG signal into the PPG encoder, the method further includes: The historical rPPG signal, historical PPG signal, and historical ECG waveform are randomly masked to obtain the masked rPPG signal, masked PPG signal, and masked ECG waveform. The current PPG encoder can be constructed based on any one of the following networks: transformer, convolutional neural network, and recurrent neural network. The historical rPPG signal and the mask rPPG signal are respectively input into the final rPPG2PPG encoder to obtain the first predicted PPG signal and the second predicted PPG signal. The first predicted PPG signal and the second predicted PPG signal are respectively input into the current PPG encoder to obtain the first rPPG feature and the second rPPG feature. The historical PPG signal and the masked PPG signal are respectively input into the current PPG encoder to obtain the first PPG feature and the second PPG feature; The historical ECG waveform and the masked ECG waveform are output to the current ECG encoder to obtain the first ECG feature and the second ECG feature; The mean square error between the first rPPG feature and the second rPPG feature, the mean square error between the first PPG feature and the second PPG feature, and the mean square error between the first ECG feature and the second ECG feature are used to determine the same-modal completion loss value. The first rPPG feature, the first PPG feature, and the first ECG feature are paired up to obtain each feature pair, and the symmetric contrastive learning loss value between each feature in each feature pair is determined. The parameters of the current PPG encoder are updated using the same modal completion loss value and the symmetric contrastive learning loss value to obtain the final PPG encoder. Accordingly, before inputting the spatial feature lexical units in the target lexical unit into the ECG decoder, the method further includes: The current ECG encoder can be constructed based on any one of the following networks: transformer, convolutional neural network, and recurrent neural network. The parameters of the current ECG encoder are updated using the same modal completion loss value and the symmetric contrastive learning loss value to obtain the final ECG encoder.

[0009] Optionally, the bidirectional Transformer includes a preset low-frequency bidirectional Transformer, a preset high-frequency bidirectional Transformer, and a spatial bidirectional Transformer; before performing mask reconstruction on the current masked word using the bidirectional Transformer to obtain the reconstructed word, the method further includes: The target low-frequency component, target high-frequency component, and spatial ECG signal of the historical ECG waveform are encoded using a preset low-frequency encoder, a preset high-frequency encoder, and a final ECG encoder, respectively, to obtain the historical first word, the historical second word, and the historical spatial ECG word. Randomly mask the first historical word, the second historical word, and the historical spatial ECG word respectively to obtain the corresponding first historical mask word, second historical mask word, and historical mask spatial word. The historical rPPG signal is processed sequentially by the final rPPG2PPG encoder and the final PPG encoder to obtain the historical rPPG conditional code. The first word of the historical mask and the conditional encoding of the historical rPPG are input into the current preset low-frequency bidirectional Transformer to obtain the reconstructed first word. The mean square error between the reconstructed first word and the historical first word is determined as the first training loss value. The historical mask second word, the historical first word, and the historical rPPG conditional code are input into the current preset high-frequency bidirectional Transformer to obtain the reconstructed second word. The mean square error between the reconstructed second word and the historical second word is determined as the second training loss value. The historical mask spatial domain lexical and the historical rPPG conditional code are input into the current spatial bidirectional Transformer to obtain the reconstructed spatial domain lexical. The mean square error between the reconstructed spatial domain lexical and the historical spatial ECG lexical is determined as the third training loss value. The parameters of the current preset low-frequency bidirectional Transformer, the current preset high-frequency bidirectional Transformer, and the current spatial bidirectional Transformer are updated using the first training loss value, the second training loss value, and the third training loss value, respectively, to obtain the final preset low-frequency bidirectional Transformer, the final preset high-frequency bidirectional Transformer, and the final spatial bidirectional Transformer.

[0010] Optionally, before inputting the first feature word and the second feature word in the target word into the preset low-frequency decoder and the preset high-frequency decoder respectively, the method further includes: Perform a short-time Fourier transform on the historical ECG waveform and decompose it into a target low-frequency component and a target high-frequency component; Construct the current preset low-frequency encoder and the current preset high-frequency encoder based on any one of the following networks: transformer, convolutional neural network, and recurrent neural network; Construct the current preset low-frequency decoder and the current preset high-frequency decoder based on any one of the following networks: transformer, convolutional neural network, and recurrent neural network; The target low-frequency component of the historical ECG waveform is processed sequentially through the current first-band encoder and the current preset low-frequency decoder to obtain the reconstructed target low-frequency component. The target high-frequency component of the historical ECG waveform is processed sequentially through the current preset high-frequency encoder and the current preset high-frequency decoder to obtain the reconstructed target high-frequency component. The low-frequency component and the high-frequency component of the reconstructed target are subjected to inverse short-time Fourier transform to obtain the frequency domain reconstructed ECG signal; The spatial domain ECG signal of the historical ECG waveform is processed sequentially by an ECG encoder and an ECG decoder to obtain the spatial domain reconstructed ECG signal. The mean square error between the historical ECG waveform and the frequency domain reconstructed ECG signal is determined as the frequency domain reconstruction loss value, the mean square error between the historical ECG waveform and the spatial domain reconstructed ECG signal is determined as the spatial domain reconstruction loss value, and the mean square error between the frequency domain reconstructed ECG signal and the spatial domain reconstructed ECG signal is determined as the dual-domain alignment loss value. The frequency domain reconstruction loss value, the spatial domain reconstruction loss value, and the dual-domain alignment loss value are used to update the parameters of the current preset low-frequency encoder, the current preset low-frequency decoder, the current preset high-frequency encoder, and the current preset high-frequency decoder to obtain the final preset low-frequency encoder, the final preset low-frequency decoder, the final preset high-frequency encoder, and the final preset high-frequency decoder.

[0011] Optionally, the step of removing the mask from the target masked word in the current masked word based on the confidence level of the reconstructed word to obtain a new current masked word includes: Select target reconstructed lexical units with a confidence level greater than a preset confidence threshold from each of the reconstructed lexical units described above; The target mask word corresponding to the target reconstructed word is determined from the current mask word, and the target mask word is replaced with the target reconstructed word to obtain a new current mask word.

[0012] Secondly, this application discloses an ECG signal acquisition device based on rPPG signals, comprising: The signal input module is used to acquire the corresponding PPG signal based on the current rPPG signal and input the PPG signal into the PPG encoder to obtain the rPPG conditional code; The masking module is used to determine the initial masked word of the first word, the second word, and the spatial ECG word, which are all in a masked state, as the current masked word; wherein, the first word is a word that meets the first preset low-frequency condition, and the second word is a word that meets the first preset high-frequency condition; The reconstruction module is used to perform mask reconstruction on the current masked words using the rPPG conditional encoding as a constraint and a bidirectional Transformer to obtain the reconstructed words. The mask removal module is used to remove the mask of the target mask word in the current mask word according to the confidence of the reconstructed word word, so as to obtain a new current mask word word, and jump back to the step of using the rPPG conditional encoding as a constraint and reconstructing the current mask word word through a bidirectional Transformer, until all masks are removed to obtain the target word word; The frequency domain waveform acquisition module is used to input the first feature word and the second feature word in the target word into a preset low-frequency decoder and a preset high-frequency decoder, respectively, to obtain the target low-frequency component and the target high-frequency component, and to perform an inverse short-time Fourier transform on the target low-frequency component and the target high-frequency component to obtain a frequency domain ECG waveform; the first feature word is a feature word that meets the second preset low-frequency condition and the second feature word is a feature word that meets the second preset high-frequency condition; The spatial waveform acquisition module is used to input the spatial feature words in the target word into the ECG decoder to obtain the spatial ECG waveform. The waveform fusion module is used to fuse the frequency domain ECG waveform and the spatial domain ECG waveform to obtain the target ECG signal.

[0013] Thirdly, this application discloses an electronic device, comprising: Memory, used to store computer programs; A processor is configured to execute the computer program to implement the steps of the aforementioned disclosed method for acquiring ECG signals based on rPPG signals.

[0014] Fourthly, this application discloses a computer-readable storage medium for storing a computer program; wherein, when the computer program is executed by a processor, it implements the steps of the aforementioned disclosed ECG signal acquisition method based on rPPG signals.

[0015] The beneficial effects of this application are as follows: This application obtains the corresponding PPG signal based on the current rPPG signal, and inputs the PPG signal into the PPG encoder to obtain the rPPG conditional code; the initial masked words of the first word, the second word, and the spatial ECG word, which are all in a masked state, are determined as the current masked words; wherein, the first word is a word that meets the first preset low-frequency condition, and the second word is a word that meets the first preset high-frequency condition; using the rPPG conditional code as a constraint, the current masked words are masked and reconstructed using a bidirectional Transformer to obtain the reconstructed words; based on the confidence level of the reconstructed words, the target masked words in the current masked words are masked and demasked to obtain a new current masked word, and then jumps back to the state where the rPPG conditional code is the constraint. The process involves several steps: first, reconstructing the current masked word using a bidirectional Transformer until all masks are removed to obtain the target word; second, inputting the first and second feature words from the target word into a preset low-frequency decoder and a preset high-frequency decoder respectively to obtain target low-frequency and target high-frequency components; third, performing an inverse short-time Fourier transform on the target low-frequency and target high-frequency components to obtain a frequency domain ECG waveform; fourth, the first feature word being a feature word that meets a second preset low-frequency condition and the second feature word being a feature word that meets a second preset high-frequency condition; fifth, inputting the spatial feature words from the target word into an ECG decoder to obtain a spatial domain ECG waveform; and finally, fusing the frequency domain ECG waveform and the spatial domain ECG waveform to obtain the target ECG signal.Therefore, this application obtains the corresponding PPG signal based on the current rPPG signal, obtains the rPPG conditional code by passing the PPG signal through a PPG encoder, and uses this code as a constraint to iteratively reconstruct the first word, the second word, and the spatial ECG word of the initial full mask. The mask is gradually removed based on the confidence level of the reconstructed word until all word reconstruction is completed. This enables accurate recovery of ECG features under the physiological characteristic constraints of the rPPG signal, ensuring a high degree of matching between the reconstruction process and human physiological rhythms, thus improving the authenticity and reliability of ECG features. The phased iterative mask reconstruction method can gradually optimize the word reconstruction accuracy, avoiding the feature ambiguity and error accumulation problems caused by a one-time full reconstruction. Simultaneously, the confidence level-based mask removal strategy prioritizes the retention of high-confidence features, further improving the overall reconstruction effect and efficiency. The reconstructed first frequency... The first and second frequency band feature words are respectively decoded to recover the first and second frequency band components, and then the frequency domain ECG waveform is obtained through inverse short-time Fourier transform. At the same time, the spatial domain feature words are decoded to obtain the spatial domain ECG waveform, realizing the dual-domain synchronous reconstruction of the ECG signal in the frequency domain and the time and spatial domains. This fully utilizes the rhythmic information of the frequency domain component and the morphological information of the spatial domain component, making the ECG signal more closely resemble the real signal in terms of rhythm and waveform. Finally, the frequency domain ECG waveform and the spatial domain ECG waveform are fused, which can complement the defects of single-domain reconstruction and take into account the overall rhythmic characteristics and local detail characteristics of the ECG signal. Based on the non-contact ECG acquisition using rPPG signals, this effectively improves the integrity, accuracy and waveform fidelity of the final target ECG signal, while simplifying the signal conversion link and improving the overall signal acquisition efficiency. Attached Figure Description

[0016] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on the provided drawings without creative effort.

[0017] Figure 1 This is a flowchart of an ECG signal acquisition method based on rPPG signals disclosed in this application; Figure 2 This is a schematic diagram of a specific rPPG2PPG encoder training disclosed in this application; Figure 3 This is a schematic diagram of a specific PPG encoder and ECG encoder training disclosed in this application; Figure 4 This is a schematic diagram of a specific bidirectional Transformer training method disclosed in this application; Figure 5 This is a schematic diagram of a specific dual-domain encoder-decoder training method disclosed in this application; Figure 6 This is a specific ECG fusion diagram disclosed in this application; Figure 7 This is a schematic diagram of the structure of an ECG signal acquisition device based on rPPG signal disclosed in this application; Figure 8 This is a structural diagram of an electronic device disclosed in this application. Detailed Implementation

[0018] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of the present invention.

[0019] Cardiovascular disease is the leading cause of death worldwide. Its early symptoms are often insidious, and some arrhythmias are paroxysmal. Therefore, real-time, continuous monitoring of cardiac physiological status, as well as early detection and risk screening for cardiovascular diseases, are crucial for reducing mortality and improving patient prognosis. Electrocardiography (ECG), as a core technology for capturing cardiac electrophysiological activity, accurately reflects the temporal characteristics and morphological changes of myocardial depolarization and repolarization. It is the gold standard for clinical diagnosis of various cardiovascular diseases such as atrial fibrillation, myocardial ischemia, and arrhythmias, and is also an important basis for cardiovascular health assessment. However, traditional ECG signal acquisition relies on specialized contact-based testing equipment and requires professional operation. The acquisition process is significantly limited by the scenario and equipment, making it difficult to achieve contactless, long-term continuous monitoring in environments such as home, bedside, and public spaces. It also cannot meet the convenient needs of routine health screening and monitoring of special populations (the elderly, children, burn patients, etc.).

[0020] Photoplethysmography (PPG) signals reflect the hemodynamic characteristics of the cardiovascular system by capturing changes in blood perfusion. Its non-invasive nature and ease of acquisition have led to its widespread integration into various wearable devices, making it a crucial technology for convenient cardiovascular monitoring, potentially replacing ECG. However, traditional contact-based PPG signal acquisition still requires direct contact with the skin, which can cause skin irritation and discomfort with prolonged use. Furthermore, the stability of contact-based acquisition can be compromised in certain medical monitoring and exercise monitoring scenarios. Additionally, PPG signals only reflect changes in blood flow and lack direct cardiac electrophysiological characteristics, thus failing to replace ECG for accurate diagnosis of cardiovascular diseases. The emergence of remote photoplethysmography (rPPG) technology breaks through the limitations of contact-based physiological signal acquisition. This technology can extract pulse wave-related physiological features non-contactly using facial video sequences, eliminating the need for direct skin contact. Signal acquisition can be achieved using common hardware such as ordinary cameras, smartphones, and smart monitoring devices. The acquisition process is convenient and non-invasive, adaptable to various scenarios including home monitoring, bedside medical monitoring, and public health screenings, making it an important direction for non-contact cardiovascular health monitoring.

[0021] However, rPPG signals are significantly affected by external factors such as changes in lighting, facial movements, and shooting posture, resulting in insufficient robustness of signal characteristics. Furthermore, rPPG signals essentially only reflect hemodynamic changes in the cardiovascular system and lack the direct cardiac electrophysiological labeling found in ECG, making them unsuitable for direct analysis and precise diagnosis of cardiovascular diseases. Current research on the conversion of physiological signals to ECG signals largely focuses on the conversion from contact-based PPG to ECG. Efficient conversion technologies for non-contact rPPG signals to ECG signals still have many shortcomings and have not yet formed mature solutions: on the one hand, there is a lack of effective fusion and spatial alignment of multimodal physiological characteristics of rPPG, PPG, and ECG, making it impossible to achieve feature-level precise mapping from rPPG signals to ECG signals; on the other hand, existing conversion schemes struggle to balance signal fidelity and robustness, resulting in ECG signals with significant deviations from real ECG signals in morphological characteristics and temporal dynamics, failing to reproduce clinically valuable electrophysiological characteristics and thus failing to meet the practical needs of clinical auxiliary diagnosis and daily health monitoring.

[0022] Achieving high-fidelity and robust conversion from facial rPPG signals to ECG signals, enabling contactless rPPG signals to possess the clinical electrophysiological diagnostic value of ECG signals, and combining the advantages of contactless and convenient rPPG acquisition with the precise diagnostic characteristics of ECG, thus overcoming the limitations of traditional cardiovascular monitoring equipment and scenarios, and realizing home-based, contactless, and continuous cardiovascular health monitoring, has become an urgent problem to be solved in the field of cardiovascular health monitoring technology. The realization of this technology will not only provide a convenient and non-invasive detection method for routine cardiovascular health screening, but will also play an important role in resource-limited medical environments, monitoring of special populations, and bedside continuous monitoring scenarios. It has significant practical application value and clinical significance in improving the efficiency of early detection of cardiovascular diseases and lowering the diagnostic threshold.

[0023] Therefore, this application provides an ECG signal acquisition scheme based on rPPG signals, which acquires ECG signals with higher accuracy and fidelity based on rPPG signals.

[0024] See Figure 1 As shown in the figure, this application discloses a method for acquiring ECG signals based on rPPG signals, including: Step S11: Obtain the corresponding PPG signal based on the current rPPG signal, and input the PPG signal into the PPG encoder to obtain the rPPG conditional code.

[0025] In this embodiment, obtaining the rPPG2PPG encoder includes: acquiring historical rPPG signals and historical PPG signals corresponding to the historical rPPG signals; constructing the current rPPG2PPG encoder based on a converter or convolutional neural network; determining the historical rPPG signals and historical PPG signals as training data; iteratively training the current rPPG2PPG encoder using the training data, and updating the current rPPG2PPG encoder based on waveform reconstruction loss value and heart rate prediction loss value to obtain the final rPPG2PPG encoder.

[0026] like Figure 2As shown, obtaining an rPPG2PPG encoder requires completing a pre-training process for rPPG to PPG waveform prediction. First, historical rPPG signals and their corresponding historical PPG signals are collected as training data. Then, mature networks such as Physformer, or custom architectures based on Transformer or convolutional neural networks, are used to construct an initial rPPG2PPG encoder. The aforementioned historical rPPG and PPG signals are input as training samples into the initial encoder. Iterative training is performed using the training data, with waveform reconstruction loss and heart rate prediction loss forming a dual loss constraint. The encoder parameters are continuously updated and optimized until the model converges, ultimately yielding an rPPG2PPG encoder that can accurately convert rPPG signals to PPG signals.

[0027] In this embodiment, the waveform reconstruction loss value is the mean square error of the waveform between the predicted PPG signal output by the current rPPG2PPG encoder and the historical PPG signal, and the heart rate prediction loss value is the absolute heart rate error between the predicted PPG signal and the historical PPG signal extracted by fast Fourier transform.

[0028] The waveform reconstruction loss is the mean square error between the predicted PPG signal output by the current rPPG2PPG encoder and the historical PPG signals in the training samples, used to constrain the morphological fit of the converted PPG waveform. The heart rate prediction loss is obtained by extracting the heart rate values ​​of the predicted PPG signal and the historical PPG signal through Fast Fourier Transform, and calculating the absolute error between the two heart rate values. This is used to ensure the accuracy of the heart rate value of the converted PPG signal. The dual loss can overcome the heart rate deviation problem caused by single waveform loss optimization.

[0029] In this embodiment, the step of obtaining the corresponding PPG signal based on the current rPPG signal and inputting the PPG signal into the PPG encoder includes: inputting the current rPPG signal into the rPPG2PPG encoder to obtain the corresponding PPG signal, and inputting the PPG signal into the PPG encoder.

[0030] The operation of obtaining the corresponding PPG signal based on the current rPPG signal and inputting it into the PPG encoder is as follows: The current rPPG signal to be processed is input into the trained rPPG2PPG encoder, and the corresponding PPG signal is obtained by the encoder conversion; then the PPG signal is input into the PPG encoder composed of Transformer, convolutional neural network or recurrent neural network to complete the feature extraction of PPG signal, and then obtain the rPPG conditional code for subsequent ECG feature reconstruction constraints.

[0031] In this embodiment, before obtaining the corresponding PPG signal based on the current rPPG signal and inputting the PPG signal into the PPG encoder, the method further includes: performing random masking processing on the historical rPPG signal, the historical PPG signal, and the historical ECG waveform to obtain a masked rPPG signal, a masked PPG signal, and a masked ECG waveform; constructing the current PPG encoder based on any one of a converter, a convolutional neural network, or a recurrent neural network; inputting the historical rPPG signal and the masked rPPG signal into the final rPPG2PPG encoder to obtain a first predicted PPG signal and a second predicted PPG signal, and inputting the first predicted PPG signal and the second predicted PPG signal into the current PPG encoder to obtain a first rPPG feature and a second rPPG feature; and inputting the historical rPPG signal and the masked PPG signal into the current PPG encoder to obtain a first rPPG feature and a second rPPG feature. The current PPG encoder is input with the input signal to obtain the first PPG feature and the second PPG feature. The historical ECG waveform and the masked ECG waveform are output to the current ECG encoder to obtain the first ECG feature and the second ECG feature. The mean square error between the first rPPG feature and the second rPPG feature, the mean square error between the first PPG feature and the second PPG feature, and the mean square error between the first ECG feature and the second ECG feature are used to determine the same-modal completion loss value. The first rPPG feature, the first PPG feature, and the first ECG feature are paired in pairs to obtain each feature pair, and the symmetric contrast learning loss value between each feature in each feature pair is determined. The parameters of the current PPG encoder are updated using the same-modal completion loss value and the symmetric contrast learning loss value to obtain the final PPG encoder.

[0032] like Figure 3 As shown, before acquiring the corresponding PPG signal based on the current rPPG signal and inputting the PPG signal into the PPG encoder, it is necessary to obtain the final PPG encoder and the final ECG encoder. The process of acquiring the PPG encoder is explained in detail below: First, random masking is performed on the historical rPPG signal, historical PPG signal, and historical ECG waveform to obtain masked rPPG signal, masked PPG signal, and masked ECG waveform. Then, a PPG encoder to be trained is built using any architecture from Transformer, Convolutional Neural Network (CNN), or Recurrent Neural Network (RNN). In other words, before conducting multi-task cross-modal feature alignment training, the rPPG signal, PPG waveform, and ECG waveform corresponding to the input face video sequence are randomly masked to obtain the three types of masked signals. A PPG encoder for PPG feature encoding is then constructed using neural network structures such as Transformer, CNN, or RNN. Under preset authorization conditions, a face video sequence is acquired using a face video sequence acquisition device, and the corresponding rPPG signal is extracted from the face video sequence. It should be noted that the preset authorization condition is that the acquisition device has pre-obtained authorization from each user in the current frame for face video sequence acquisition and recognition operations.

[0033] Next, the historical rPPG signal and the masked rPPG signal are input into the trained rPPG2PPG encoder to obtain the first predicted PPG signal and the second predicted PPG signal, respectively. These two signals are then input into the current PPG encoder to output the first rPPG feature and the second rPPG feature. Similarly, the historical PPG signal and the masked PPG signal are input into the current PPG encoder to obtain the first PPG feature and the second PPG feature. The historical ECG waveform and the masked ECG waveform are input into the current ECG encoder to obtain the first ECG feature and the second ECG feature. The mean square error between the first and second rPPG features, the first and second PPG features, and the first and second ECG features is used as the in-modal completion loss value. In other words, using a fixed-parameter rPPG2PPG encoder and a PPG encoder, the original and masked rPPG converted PPG features, the original and masked PPG features, and the original and masked ECG features are extracted, respectively. The mean square error between the original features and the masked features is calculated within the same modality and used as the completion loss for mask reconstruction, thus achieving in-modal feature mask reconstruction constraints.

[0034] Then, the first rPPG feature, the first PPG feature, and the first ECG feature are paired to form feature pairs, and the symmetric contrastive learning loss value between each feature pair is calculated. Combining the in-modal completion loss value and the symmetric contrastive learning loss value, the parameters of the current PPG encoder are updated to obtain the final PPG encoder. In other words, a symmetric contrastive learning task is constructed with pairwise pairings of rPPG, PPG, and ECG modalities, and the inter-modal alignment loss is calculated. The PPG encoder parameters are iteratively updated using the in-modal mask reconstruction completion loss and the cross-modal contrastive learning alignment loss as joint constraints, achieving deep spatial alignment of the three modalities and obtaining the final usable PPG encoder.

[0035] Step S12: Determine the initial masked word of the first word, the second word, and the spatial ECG word, all of which are in a masked state, as the current masked word; wherein, the first word is a word that meets the first preset low-frequency condition, and the second word is a word that meets the first preset high-frequency condition.

[0036] Understandably, after obtaining the rPPG conditional encoding, before starting the mask reconstruction process, it is necessary to first define the initial mask term set. Specifically, all initial mask terms, including the first term, the second term, and the spatial ECG term, are uniformly determined as the mask terms to be processed in this round. The first term must meet the first preset low-frequency condition, meaning the feature component it represents corresponds to low-frequency features in the physiological signal; similarly, the second term must meet the first preset high-frequency condition, meaning the feature component it represents corresponds to high-frequency features in the physiological signal. Through precise definition of different frequency domain feature terms and initial mask state settings, a clear initial processing object and feature range are provided for the subsequent bidirectional Transformer mask reconstruction operation based on rPPG conditional constraints, ensuring that the reconstruction process focuses on the orderly recovery of features in different frequency bands and spatial ECG features.

[0037] Step S13: Using the rPPG conditional encoding as a constraint, and through a bidirectional Transformer, perform mask reconstruction on the current masked tokens to obtain the reconstructed tokens.

[0038] In this embodiment, the bidirectional Transformer includes a preset low-frequency bidirectional Transformer, a preset high-frequency bidirectional Transformer, and a spatial bidirectional Transformer. Before performing mask reconstruction on the current masked word using the bidirectional Transformer to obtain the reconstructed word, the method further includes: encoding the target low-frequency component, target high-frequency component, and spatial ECG signal of the historical ECG waveform using a preset low-frequency encoder, a preset high-frequency encoder, and a final ECG encoder, respectively, to obtain a historical first word, a historical second word, and a historical spatial ECG word; randomly masking the historical first word, the historical second word, and the historical spatial ECG word, respectively, to obtain the corresponding historical mask first word, historical mask second word, and historical mask spatial word; processing the historical rPPG signal sequentially through a final rPPG2PPG encoder and a final PPG encoder to obtain historical rPPG conditional encoding; and inputting the historical mask first word and the historical rPPG conditional encoding into the current preset low-frequency bidirectional Transformer to obtain the reconstructed first word. The mean squared error between the reconstructed first lexical unit and the historical first lexical unit is determined as the first training loss value. The historical mask second lexical unit, the historical first lexical unit, and the historical rPPG conditional code are input into the current preset high-frequency bidirectional Transformer to obtain the reconstructed second lexical unit. The mean squared error between the reconstructed second lexical unit and the historical second lexical unit is determined as the second training loss value. The historical mask spatial domain lexical unit and the historical rPPG conditional code are input into the current spatial domain bidirectional Transformer to obtain the reconstructed spatial domain lexical unit. The mean squared error between the reconstructed spatial domain lexical unit and the historical spatial domain ECG lexical unit is determined as the third training loss value. The parameters of the current preset low-frequency bidirectional Transformer, the current preset high-frequency bidirectional Transformer, and the current spatial domain bidirectional Transformer are updated using the first training loss value, the second training loss value, and the third training loss value, respectively, to obtain the final preset low-frequency bidirectional Transformer, the final preset high-frequency bidirectional Transformer, and the final spatial domain bidirectional Transformer.

[0039] like Figure 4As shown, the bidirectional Transformer includes a preset low-frequency bidirectional Transformer, a preset high-frequency bidirectional Transformer, and a spatial bidirectional Transformer. Before performing mask reconstruction using the bidirectional Transformer, the preset low-frequency encoder, the preset high-frequency encoder, and the final ECG encoder are used to encode the target low-frequency component, the target high-frequency component, and the spatial ECG signal of the historical ECG waveform to obtain the first historical word, the second historical word, and the historical spatial ECG word. Then, the three are randomly masked to obtain the corresponding historical mask words. In other words, the bidirectional Transformer used in this embodiment includes three types: LF bidirectional Transformer, HF bidirectional Transformer, and spatial bidirectional Transformer. Before conducting mask reconstruction training, the pre-trained LF encoder (i.e., the preset low-frequency encoder), HF encoder (i.e., the preset high-frequency encoder), and ECG encoder are used to encode the low-frequency component, the high-frequency component, and the spatial signal of the historical ECG signal into LF tokens, HF tokens, and spatial tokens. Then, the three types of historical tokens are randomly masked to obtain the corresponding mask tokens.

[0040] Next, the historical rPPG signal is processed by the final rPPG2PPG encoder and the final PPG encoder to obtain the historical rPPG conditional code; the first word of the historical mask and the historical rPPG conditional code are input into a preset low-frequency bidirectional Transformer, and the first training loss value is obtained with mean square error; the second word of the historical mask, the first word of the history, and the historical rPPG conditional code are input into a preset high-frequency bidirectional Transformer, and the second training loss value is obtained with mean square error; the spatial word of the historical mask and the historical rPPG conditional code are input into a spatial bidirectional Transformer, and the third training loss value is obtained with mean square error; the corresponding bidirectional Transformer parameters are updated using the three types of loss values ​​respectively to obtain the final three types of bidirectional Transformers. In other words, the historical rPPG signal is processed by an rPPG2PPG encoder and a PPG encoder with fixed parameters to obtain the historical rPPG conditional code; an LF bidirectional Transformer is trained with the mask LFtoken and the rPPG conditional code, and the mean square error between the reconstructed LFtoken and the original LFtoken is calculated as the loss; an HF bidirectional Transformer is trained with the mask HFtoken, the original LFtoken, and the rPPG conditional code, and the corresponding mean square error is calculated as the loss; a spatial bidirectional Transformer is trained with the mask spatial token and the rPPG conditional code, and the corresponding mean square error is calculated as the loss; the parameters of the three types of bidirectional Transformers are iteratively updated using the three types of losses respectively, and finally, three types of bidirectional Transformers that can be used for ECG feature reconstruction are obtained.

[0041] After obtaining the pre-trained preset low-frequency bidirectional Transformer, preset high-frequency bidirectional Transformer, and spatial bidirectional Transformer, rPPG conditional coding is used as a constraint. The bidirectional Transformer is used to perform mask reconstruction on the current masked words, thereby outputting the reconstructed words. That is, rPPG conditional coding is used as a physiological feature constraint. LF bidirectional Transformer, HF bidirectional Transformer, and spatial bidirectional Transformer are used to perform feature recovery and reconstruction on the low-frequency words, high-frequency words, and spatial ECG words currently in the masked state, respectively, to obtain the complete reconstructed words.

[0042] Step S14: Based on the confidence level of the reconstructed lexicon, remove the mask from the target mask lexicon in the current mask lexicon to obtain a new current mask lexicon. Then, jump back to the step of reconstructing the current mask lexicon using the rPPG conditional encoding as a constraint and a bidirectional Transformer, until all masks are removed to obtain the target lexicon.

[0043] In this embodiment, the step of removing the target masking word in the current masking word based on the confidence level of the reconstructed word to obtain a new current masking word includes: selecting target reconstructed word from each of the reconstructed word and finding that the confidence level is greater than a preset confidence threshold; determining the target masking word corresponding to the target reconstructed word from the current masking word and replacing the target masking word with the target reconstructed word to obtain a new current masking word.

[0044] Based on the confidence level of the reconstructed words, a masking operation is performed on the current masked words. First, target reconstructed words with confidence levels higher than a preset confidence threshold are selected from all reconstructed words. Then, the corresponding target masked word is located in the current masked words, and this target masked word is replaced with the target reconstructed word that meets the confidence threshold, thus obtaining the updated current masked word. Subsequently, using the new current masked word as input, the process jumps back to the steps of mask reconstruction using rPPG conditional coding as constraints and bidirectional Transformer. The cyclical process of mask reconstruction, high-confidence word selection, masking, and word replacement is repeated. An iterative masking strategy is used to gradually recover word features until all masks in the current masked word are completely removed, finally obtaining a complete, maskless target word, providing an accurate and reliable feature foundation for subsequent frequency and spatial ECG waveform reconstruction.

[0045] Step S15: Input the first feature word and the second feature word in the target word into the preset low-frequency decoder and the preset high-frequency decoder respectively to obtain the target low-frequency component and the target high-frequency component. Perform inverse short-time Fourier transform on the target low-frequency component and the target high-frequency component to obtain the frequency domain ECG waveform; the first feature word is a feature word that meets the second preset low-frequency condition and the second feature word is a feature word that meets the second preset high-frequency condition.

[0046] In this embodiment, before inputting the first and second feature words in the target word into the preset low-frequency decoder and preset high-frequency decoder respectively, the method further includes: performing a short-time Fourier transform on the historical ECG waveform and decomposing it into a target low-frequency component and a target high-frequency component; constructing a current preset low-frequency encoder and a current preset high-frequency encoder based on any one of a transformer, convolutional neural network, or recurrent neural network; constructing a current preset low-frequency decoder and a current preset high-frequency decoder based on any one of a transformer, convolutional neural network, or recurrent neural network; processing the target low-frequency component of the historical ECG waveform sequentially through the current first-band encoder and the current preset low-frequency decoder to obtain the reconstructed target low-frequency component; processing the target high-frequency component of the historical ECG waveform sequentially through the current preset high-frequency encoder and the current preset high-frequency decoder to obtain the reconstructed target high-frequency component; and processing the reconstructed target low-frequency component and the reconstructed target high-frequency component. The high-frequency components undergo inverse short-time Fourier transform to obtain the frequency-domain reconstructed ECG signal. The spatial-domain ECG signal of the historical ECG waveform is sequentially processed by an ECG encoder and an ECG decoder to obtain the spatial-domain reconstructed ECG signal. The mean square error between the historical ECG waveform and the frequency-domain reconstructed ECG signal is determined as the frequency-domain reconstruction loss value, the mean square error between the historical ECG waveform and the spatial-domain reconstructed ECG signal is determined as the spatial-domain reconstruction loss value, and the mean square error between the frequency-domain reconstructed ECG signal and the spatial-domain reconstructed ECG signal is determined as the dual-domain alignment loss value. The parameters of the current preset low-frequency encoder, the current preset low-frequency decoder, the current preset high-frequency encoder, and the current preset high-frequency decoder are updated using the frequency-domain reconstruction loss value, the spatial-domain reconstruction loss value, and the dual-domain alignment loss value to obtain the final preset low-frequency encoder, the final preset low-frequency decoder, the final preset high-frequency encoder, and the final preset high-frequency decoder.

[0047] First, before inputting the first and second feature words in the target word into the preset low-frequency decoder and preset high-frequency decoder respectively, a short-time Fourier transform is performed on the historical ECG waveform to decompose it into the target low-frequency component and the target high-frequency component. Then, using any one of Transformer, convolutional neural network, or recurrent neural network, the current preset low-frequency encoder, preset high-frequency encoder, preset low-frequency decoder, and preset high-frequency decoder are constructed respectively.

[0048] Secondly, the target low-frequency component of the historical ECG waveform is processed by the current preset low-frequency encoder and preset low-frequency decoder to obtain the reconstructed target low-frequency component, and the target high-frequency component is processed by the current preset high-frequency encoder and preset high-frequency decoder to obtain the reconstructed target high-frequency component; inverse short-time Fourier transform is performed on both to obtain the frequency domain reconstructed ECG signal; the spatial domain ECG signal of the historical ECG waveform is processed by the ECG encoder and ECG decoder to obtain the spatial domain reconstructed ECG signal; the mean square error between the historical ECG waveform and the frequency domain reconstructed ECG signal is used as the frequency domain reconstruction loss value, the mean square error between the historical ECG waveform and the spatial domain reconstructed ECG signal is used as the spatial domain reconstruction loss value, and the mean square error between the frequency domain reconstructed ECG signal and the spatial domain reconstructed ECG signal is used as the dual-domain alignment loss value.

[0049] Next, the frequency domain reconstruction loss, spatial domain reconstruction loss, and dual-domain alignment loss are used to jointly update the parameters of the current preset low-frequency encoder, preset low-frequency decoder, preset high-frequency encoder, and preset high-frequency decoder. The training is iterated until the model converges, and the final preset low-frequency encoder, preset low-frequency decoder, preset high-frequency encoder, and preset high-frequency decoder are obtained.

[0050] like Figure 5As shown, a VQ-VAE (Vector Quantized Variational Autoencoder) architecture is adopted to simultaneously perform compression and reconstruction learning on the input ECG signal in both the spatial and frequency domains. After training, an LF encoder, LF decoder, LF codebook, HF encoder, HF decoder, HF codebook, ECG decoder, and spatial codebook are obtained. First, the input ECG signal is decomposed in the frequency domain using Short Time Fourier Transform (STFT) to obtain the LF low-frequency component and the HF high-frequency component. The LF encoder, HF encoder, LF decoder, HF decoder, and ECG decoder are all constructed using networks based on Transformer, convolutional neural networks, or recurrent neural networks. The LF encoder is used to encode the features of the LF component of the ECG signal, the HF encoder is used to encode the features of the HF component of the ECG signal, the LF decoder is used to decode the LF component features and recover the LF component, the HF decoder is used to decode the HF component features and recover the HF component, and the ECG decoder is used to decode the ECG signal features. The LF codebook is used for quantization mapping of the LF component features, the HF codebook is used for quantization mapping of the HF component features, and the spatial codebook is used for quantization mapping of the spatial ECG features. Based on the LF and HF components obtained from decoding, frequency domain ECG signal reconstruction is completed through inverse short-time Fourier transform (iSTFT). During training, frequency domain reconstruction loss, spatial domain reconstruction loss, and dual-domain alignment loss are introduced. The frequency domain reconstruction loss is the mean square error between the frequency domain encoded and decoded ECG signal and the input ECG signal. The spatial domain reconstruction loss is the mean square error between the spatial domain encoded and decoded ECG signal and the input ECG signal. The dual-domain alignment loss is the mean square error between the spatial domain reconstructed ECG signal and the frequency domain reconstructed ECG signal, so as to achieve high-fidelity compressed reconstruction of ECG signal under the joint constraints of spatial and frequency domains.

[0051] In this embodiment, the first feature word that meets the second preset low-frequency condition and the second feature word that meets the second preset high-frequency condition in the target word obtained after iterative mask reconstruction are respectively input into the preset low-frequency decoder and preset high-frequency decoder that have been trained. After decoding, the corresponding target low-frequency component and target high-frequency component are obtained. Then, the inverse short-time Fourier transform is performed on the target low-frequency component and the target high-frequency component to restore the frequency domain features to the time domain signal, and finally the complete frequency domain ECG waveform is obtained, which provides a reliable signal foundation in the frequency domain dimension for subsequent fusion with the spatial domain ECG waveform to generate the final target ECG signal.

[0052] Step S16: Input the spatial feature words in the target words into the ECG decoder to obtain the spatial ECG waveform.

[0053] In this embodiment, before inputting the spatial feature lexical units in the target lexical units into the ECG decoder, the method further includes: constructing the current ECG encoder based on any one of the following networks: transformer, convolutional neural network, and recurrent neural network; and updating the parameters of the current ECG encoder using the same modality completion loss value and the symmetric contrastive learning loss value to obtain the final ECG encoder.

[0054] First, an initial ECG encoder is built based on any one of the network structures of Transformer, Convolutional Neural Network, or Recurrent Neural Network. Then, the ECG encoder is iteratively updated with parameters by using the joint constraints of the same modality completion loss value and the symmetric contrastive learning loss value to obtain the final ECG encoder with stable feature extraction and modality space alignment.

[0055] Subsequently, the spatial feature words in the target words after the iterative mask reconstruction are input into the trained ECG decoder. The spatial ECG waveform is obtained by decoding and reconstruction, which provides the spatial dimension signal basis for the subsequent fusion of spatial and frequency domain ECG waveforms.

[0056] Step S17: Fuse the frequency domain ECG waveform and the spatial domain ECG waveform to obtain the target ECG signal.

[0057] like Figure 6 As shown, the frequency domain ECG waveform obtained by inverse short-time Fourier transform is weighted and averaged with the spatial domain ECG waveform obtained by ECG decoder. The features of the spatial domain reconstructed ECG waveform and the frequency domain reconstructed ECG waveform are integrated, and the complete spatial domain shape and accurate frequency domain timing characteristics are fully combined to obtain a target ECG signal with high fidelity and high robustness and clinical reference value, thus realizing the accurate conversion from rPPG signal to ECG signal.

[0058] The beneficial effects of this application are as follows: This application obtains the corresponding PPG signal based on the current rPPG signal, and inputs the PPG signal into the PPG encoder to obtain the rPPG conditional code; the initial masked words of the first word, the second word, and the spatial ECG word, which are all in a masked state, are determined as the current masked words; wherein, the first word is a word that meets the first preset low-frequency condition, and the second word is a word that meets the first preset high-frequency condition; using the rPPG conditional code as a constraint, the current masked words are masked and reconstructed using a bidirectional Transformer to obtain the reconstructed words; based on the confidence level of the reconstructed words, the target masked words in the current masked words are masked and demasked to obtain a new current masked word, and then jumps back to the state where the rPPG conditional code is the constraint. The process involves several steps: first, reconstructing the current masked word using a bidirectional Transformer until all masks are removed to obtain the target word; second, inputting the first and second feature words from the target word into a preset low-frequency decoder and a preset high-frequency decoder respectively to obtain target low-frequency and target high-frequency components; third, performing an inverse short-time Fourier transform on the target low-frequency and target high-frequency components to obtain a frequency domain ECG waveform; fourth, the first feature word being a feature word that meets a second preset low-frequency condition and the second feature word being a feature word that meets a second preset high-frequency condition; fifth, inputting the spatial feature words from the target word into an ECG decoder to obtain a spatial domain ECG waveform; and finally, fusing the frequency domain ECG waveform and the spatial domain ECG waveform to obtain the target ECG signal.Therefore, this application obtains the corresponding PPG signal based on the current rPPG signal, obtains the rPPG conditional code by passing the PPG signal through a PPG encoder, and uses this code as a constraint to iteratively reconstruct the first word, the second word, and the spatial ECG word of the initial full mask. The mask is gradually removed based on the confidence level of the reconstructed word until all word reconstruction is completed. This enables accurate recovery of ECG features under the physiological characteristic constraints of the rPPG signal, ensuring a high degree of matching between the reconstruction process and human physiological rhythms, thus improving the authenticity and reliability of ECG features. The phased iterative mask reconstruction method can gradually optimize the word reconstruction accuracy, avoiding the feature ambiguity and error accumulation problems caused by a one-time full reconstruction. Simultaneously, the confidence level-based mask removal strategy prioritizes the retention of high-confidence features, further improving the overall reconstruction effect and efficiency. The reconstructed first frequency... The first and second frequency band feature words are respectively decoded to recover the first and second frequency band components, and then the frequency domain ECG waveform is obtained through inverse short-time Fourier transform. At the same time, the spatial domain feature words are decoded to obtain the spatial domain ECG waveform, realizing the dual-domain synchronous reconstruction of the ECG signal in the frequency domain and the time and spatial domains. This fully utilizes the rhythmic information of the frequency domain component and the morphological information of the spatial domain component, making the ECG signal more closely resemble the real signal in terms of rhythm and waveform. Finally, the frequency domain ECG waveform and the spatial domain ECG waveform are fused, which can complement the defects of single-domain reconstruction and take into account the overall rhythmic characteristics and local detail characteristics of the ECG signal. Based on the non-contact ECG acquisition using rPPG signals, this effectively improves the integrity, accuracy and waveform fidelity of the final target ECG signal, while simplifying the signal conversion link and improving the overall signal acquisition efficiency.

[0059] See Figure 7 As shown in the figure, this application discloses an ECG signal acquisition device based on rPPG signals, comprising: The signal input module 11 is used to obtain the corresponding PPG signal based on the current rPPG signal and input the PPG signal into the PPG encoder to obtain the rPPG conditional code; The masking module 12 is used to determine the initial masked word of the first word, the second word, and the spatial ECG word, which are all in a masked state, as the current masked word; wherein, the first word is a word that meets the first preset low-frequency condition, and the second word is a word that meets the first preset high-frequency condition; The reconstruction module 13 is used to perform mask reconstruction on the current masked words using the rPPG conditional encoding as a constraint and through a bidirectional Transformer to obtain the reconstructed words. The mask removal module 14 is used to remove the mask of the target mask word in the current mask word according to the confidence of the reconstructed word word, so as to obtain a new current mask word word, and jump back to the step of using the rPPG conditional encoding as a constraint and reconstructing the current mask word word through a bidirectional Transformer, until all masks are removed to obtain the target word word. The frequency domain waveform acquisition module 15 is used to input the first feature word and the second feature word in the target word into a preset low-frequency decoder and a preset high-frequency decoder, respectively, to obtain the target low-frequency component and the target high-frequency component, and to perform an inverse short-time Fourier transform on the target low-frequency component and the target high-frequency component to obtain a frequency domain ECG waveform; the first feature word is a feature word that meets the second preset low-frequency condition and the second feature word is a feature word that meets the second preset high-frequency condition; The spatial waveform acquisition module 16 is used to input the spatial feature words in the target word into the ECG decoder to obtain the spatial ECG waveform. The waveform fusion module 17 is used to fuse the frequency domain ECG waveform and the spatial domain ECG waveform to obtain the target ECG signal.

[0060] Furthermore, embodiments of this application also provide an electronic device. Figure 8 This is a structural diagram of an electronic device 20 according to an exemplary embodiment. The content of the diagram should not be construed as limiting the scope of this application.

[0061] Figure 8 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Specifically, it may include: at least one processor 21, at least one memory 22, a power supply 23, a communication interface 24, an input / output interface 25, and a communication bus 26. The memory 22 stores a computer program, which is loaded and executed by the processor 21 to implement the relevant steps in the ECG signal acquisition method based on rPPG signals disclosed in any of the foregoing embodiments.

[0062] In this embodiment, the power supply 23 is used to provide operating voltage for various hardware devices on the electronic device; the communication interface 24 can create a data transmission channel between the electronic device and external devices, and the communication protocol it follows can be any communication protocol applicable to the technical solution of this application, and is not specifically limited here; the input / output interface 25 is used to acquire external input data or output data to the outside world, and its specific interface type can be selected according to specific application needs, and is not specifically limited here.

[0063] The processor 21 may include one or more processing cores, such as a quad-core processor or an octa-core processor. The processor 21 may be implemented using at least one hardware form selected from DSP (Digital Signal Processing), FPGA (Field-Programmable Gate Array), and PLA (Programmable Logic Array). The processor 21 may also include a main processor and a coprocessor. The main processor, also known as a CPU (Central Processing Unit), is used to process data in the wake-up state; the coprocessor is a low-power processor used to process data in the standby state. In some embodiments, the processor 21 may integrate a GPU (Graphics Processing Unit), which is responsible for rendering and drawing the content to be displayed on the screen. In some embodiments, the processor 21 may also include an AI (Artificial Intelligence) processor, which is used to handle computational operations related to machine learning.

[0064] In addition, the memory 22, as a carrier for resource storage, can be a read-only memory, random access memory, disk or optical disk, etc. The resources stored on it include operating system 221, computer program 222 and data 223, etc., and the storage method can be temporary storage or permanent storage.

[0065] The operating system 221 manages and controls the various hardware devices and computer programs 222 on the electronic device to enable the processor 21 to perform calculations and processing on the massive amounts of data 223 in the memory 22. The operating system can be Windows, Unix, Linux, etc. The computer program 222, in addition to including a computer program capable of performing the ECG signal acquisition method based on rPPG signals disclosed in any of the foregoing embodiments, may further include computer programs capable of performing other specific tasks. The data 223 may include data received by the electronic device from external devices, as well as data collected by its own input / output interface 25.

[0066] Furthermore, this application also discloses a computer-readable storage medium for storing a computer program; wherein, when the computer program is executed by a processor, it implements the aforementioned disclosed method for acquiring ECG signals based on rPPG signals. Specific steps of this method can be found in the corresponding content disclosed in the foregoing embodiments, and will not be repeated here.

[0067] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on its differences from other embodiments. Similar or identical parts between embodiments can be referred to interchangeably. For the apparatus disclosed in the embodiments, since it corresponds to the method disclosed in the embodiments, the description is relatively simple; relevant parts can be referred to in the method section.

[0068] Those skilled in the art will further recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application. The steps of the methods or algorithms described in conjunction with the embodiments disclosed herein can be implemented directly in hardware, software modules executed by a processor, or a combination of both. The software module may be located in random access memory (RAM), memory, read-only memory (ROM), electrically programmable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), register, hard disk, removable disk, CD-ROM (Compact Disc Read-Only Memory), or any other form of storage medium known in the art.

[0069] Finally, it should be noted that in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0070] The foregoing has provided a detailed description of the ECG signal acquisition method, apparatus, device, and medium based on rPPG signals provided by the present invention. Specific examples have been used to illustrate the principles and implementation methods of the present invention. The descriptions of the above embodiments are only intended to help understand the method and core ideas of the present invention. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of the present invention. Therefore, the content of this specification should not be construed as a limitation of the present invention.

Claims

1. A method for acquiring ECG signals based on rPPG signals, characterized in that, include: The corresponding PPG signal is obtained based on the current rPPG signal, and the PPG signal is input into the PPG encoder to obtain the rPPG conditional code; The initial masked word of the first word, the second word, and the spatial ECG word, all of which are in a masked state, is determined as the current masked word; wherein, the first word is a word that meets the first preset low-frequency condition, and the second word is a word that meets the first preset high-frequency condition; Using the rPPG conditional encoding as a constraint, and through a bidirectional Transformer, the current masked tokens are reconstructed to obtain the reconstructed tokens; Based on the confidence level of the reconstructed lexical, the target lexical in the current masked lexical is masked to obtain a new current masked lexical. Then, the process jumps back to the step of reconstructing the current masked lexical using the rPPG conditional encoding as a constraint and a bidirectional Transformer, until all masks are removed to obtain the target lexical. The first and second feature words in the target word are respectively input into a preset low-frequency decoder and a preset high-frequency decoder to obtain target low-frequency components and target high-frequency components. The target low-frequency components and the target high-frequency components are subjected to inverse short-time Fourier transform to obtain frequency domain ECG waveforms. The first feature word is a feature word that meets the second preset low-frequency condition and the second feature word is a feature word that meets the second preset high-frequency condition. The spatial feature words in the target words are input into the ECG decoder to obtain the spatial ECG waveform; The frequency domain ECG waveform and the spatial domain ECG waveform are fused to obtain the target ECG signal.

2. The ECG signal acquisition method based on rPPG signal according to claim 1, characterized in that, The step of acquiring the corresponding PPG signal based on the current rPPG signal and inputting the PPG signal into the PPG encoder includes: The current rPPG signal is input into the rPPG2PPG encoder to obtain the corresponding PPG signal, and the PPG signal is then input into the PPG encoder. Accordingly, obtaining the rPPG2PPG encoder includes: Collect historical rPPG signals and historical PPG signals corresponding to the historical rPPG signals; The current rPPG2PPG encoder is constructed based on a transformer or a convolutional neural network; The historical rPPG signal and the historical PPG signal are determined as training data; The current rPPG2PPG encoder is iteratively trained using the training data, and updated based on waveform reconstruction loss value and heart rate prediction loss value to obtain the final rPPG2PPG encoder.

3. The ECG signal acquisition method based on rPPG signal according to claim 2, characterized in that, The waveform reconstruction loss value is the mean square error of the waveform between the predicted PPG signal output by the current rPPG2PPG encoder and the historical PPG signal, and the heart rate prediction loss value is the absolute heart rate error between the predicted PPG signal and the historical PPG signal extracted by fast Fourier transform.

4. The ECG signal acquisition method based on rPPG signal according to claim 2, characterized in that, Before acquiring the corresponding PPG signal based on the current rPPG signal and inputting the PPG signal into the PPG encoder, the method further includes: The historical rPPG signal, historical PPG signal, and historical ECG waveform are randomly masked to obtain the masked rPPG signal, masked PPG signal, and masked ECG waveform. The current PPG encoder can be constructed based on any one of the following networks: transformer, convolutional neural network, and recurrent neural network. The historical rPPG signal and the mask rPPG signal are respectively input into the final rPPG2PPG encoder to obtain the first predicted PPG signal and the second predicted PPG signal. The first predicted PPG signal and the second predicted PPG signal are respectively input into the current PPG encoder to obtain the first rPPG feature and the second rPPG feature. The historical PPG signal and the masked PPG signal are respectively input into the current PPG encoder to obtain the first PPG feature and the second PPG feature; The historical ECG waveform and the masked ECG waveform are output to the current ECG encoder to obtain the first ECG feature and the second ECG feature; The mean square error between the first rPPG feature and the second rPPG feature, the mean square error between the first PPG feature and the second PPG feature, and the mean square error between the first ECG feature and the second ECG feature are used to determine the same-modal completion loss value. The first rPPG feature, the first PPG feature, and the first ECG feature are paired up to obtain each feature pair, and the symmetric contrastive learning loss value between each feature in each feature pair is determined. The parameters of the current PPG encoder are updated using the same modal completion loss value and the symmetric contrastive learning loss value to obtain the final PPG encoder. Accordingly, before inputting the spatial feature lexical units in the target lexical unit into the ECG decoder, the process further includes: The current ECG encoder can be constructed based on any one of the following networks: transformer, convolutional neural network, and recurrent neural network. The parameters of the current ECG encoder are updated using the same modal completion loss value and the symmetric contrastive learning loss value to obtain the final ECG encoder.

5. The ECG signal acquisition method based on rPPG signal according to claim 4, characterized in that, The bidirectional Transformer includes a preset low-frequency bidirectional Transformer, a preset high-frequency bidirectional Transformer, and a spatial bidirectional Transformer; before performing mask reconstruction on the current masked words using the bidirectional Transformer to obtain the reconstructed words, the process further includes: The target low-frequency component, target high-frequency component, and spatial ECG signal of the historical ECG waveform are encoded using a preset low-frequency encoder, a preset high-frequency encoder, and a final ECG encoder, respectively, to obtain the historical first word, the historical second word, and the historical spatial ECG word. Randomly mask the first historical word, the second historical word, and the historical spatial ECG word respectively to obtain the corresponding first historical mask word, second historical mask word, and historical mask spatial word. The historical rPPG signal is processed sequentially by the final rPPG2PPG encoder and the final PPG encoder to obtain the historical rPPG conditional code. The first word of the historical mask and the conditional encoding of the historical rPPG are input into the current preset low-frequency bidirectional Transformer to obtain the reconstructed first word. The mean square error between the reconstructed first word and the historical first word is determined as the first training loss value. The historical mask second word, the historical first word, and the historical rPPG conditional code are input into the current preset high-frequency bidirectional Transformer to obtain the reconstructed second word. The mean square error between the reconstructed second word and the historical second word is determined as the second training loss value. The historical mask spatial domain lexical and the historical rPPG conditional code are input into the current spatial bidirectional Transformer to obtain the reconstructed spatial domain lexical. The mean square error between the reconstructed spatial domain lexical and the historical spatial ECG lexical is determined as the third training loss value. The parameters of the current preset low-frequency bidirectional Transformer, the current preset high-frequency bidirectional Transformer, and the current spatial bidirectional Transformer are updated using the first training loss value, the second training loss value, and the third training loss value, respectively, to obtain the final preset low-frequency bidirectional Transformer, the final preset high-frequency bidirectional Transformer, and the final spatial bidirectional Transformer.

6. The ECG signal acquisition method based on rPPG signal according to claim 5, characterized in that, Before inputting the first feature word and the second feature word in the target word into the preset low-frequency decoder and the preset high-frequency decoder respectively, the method further includes: Perform a short-time Fourier transform on the historical ECG waveform and decompose it into a target low-frequency component and a target high-frequency component; Construct the current preset low-frequency encoder and the current preset high-frequency encoder based on any one of the following networks: transformer, convolutional neural network, and recurrent neural network; Construct the current preset low-frequency decoder and the current preset high-frequency decoder based on any one of the following networks: transformer, convolutional neural network, and recurrent neural network; The target low-frequency component of the historical ECG waveform is processed sequentially through the current first-band encoder and the current preset low-frequency decoder to obtain the reconstructed target low-frequency component. The target high-frequency component of the historical ECG waveform is processed sequentially through the current preset high-frequency encoder and the current preset high-frequency decoder to obtain the reconstructed target high-frequency component. The low-frequency component and the high-frequency component of the reconstructed target are subjected to inverse short-time Fourier transform to obtain the frequency domain reconstructed ECG signal; The spatial domain ECG signal of the historical ECG waveform is processed sequentially by an ECG encoder and an ECG decoder to obtain the spatial domain reconstructed ECG signal. The mean square error between the historical ECG waveform and the frequency domain reconstructed ECG signal is determined as the frequency domain reconstruction loss value, the mean square error between the historical ECG waveform and the spatial domain reconstructed ECG signal is determined as the spatial domain reconstruction loss value, and the mean square error between the frequency domain reconstructed ECG signal and the spatial domain reconstructed ECG signal is determined as the dual-domain alignment loss value. The frequency domain reconstruction loss value, the spatial domain reconstruction loss value, and the dual-domain alignment loss value are used to update the parameters of the current preset low-frequency encoder, the current preset low-frequency decoder, the current preset high-frequency encoder, and the current preset high-frequency decoder to obtain the final preset low-frequency encoder, the final preset low-frequency decoder, the final preset high-frequency encoder, and the final preset high-frequency decoder.

7. The ECG signal acquisition method based on rPPG signal according to any one of claims 1 to 6, characterized in that, The step of removing the mask from the target masked word in the current masked word based on the confidence level of the reconstructed word to obtain a new current masked word includes: Select target reconstructed lexical units with a confidence level greater than a preset confidence threshold from each of the reconstructed lexical units described above; The target mask word corresponding to the target reconstructed word is determined from the current mask word, and the target mask word is replaced with the target reconstructed word to obtain a new current mask word.

8. An ECG signal acquisition device based on rPPG signals, characterized in that, include: The signal input module is used to acquire the corresponding PPG signal based on the current rPPG signal and input the PPG signal into the PPG encoder to obtain the rPPG conditional code; The masking module is used to determine the initial masked word of the first word, the second word, and the spatial ECG word, which are all in a masked state, as the current masked word; wherein, the first word is a word that meets the first preset low-frequency condition, and the second word is a word that meets the first preset high-frequency condition; The reconstruction module is used to perform mask reconstruction on the current masked words using the rPPG conditional encoding as a constraint and a bidirectional Transformer to obtain the reconstructed words. The mask removal module is used to remove the mask of the target mask word in the current mask word according to the confidence of the reconstructed word word, so as to obtain a new current mask word word, and jump back to the step of using the rPPG conditional encoding as a constraint and reconstructing the current mask word word through a bidirectional Transformer, until all masks are removed to obtain the target word word; The frequency domain waveform acquisition module is used to input the first feature word and the second feature word in the target word into a preset low-frequency decoder and a preset high-frequency decoder, respectively, to obtain the target low-frequency component and the target high-frequency component, and to perform an inverse short-time Fourier transform on the target low-frequency component and the target high-frequency component to obtain a frequency domain ECG waveform; the first feature word is a feature word that meets the second preset low-frequency condition and the second feature word is a feature word that meets the second preset high-frequency condition; The spatial waveform acquisition module is used to input the spatial feature words in the target word into the ECG decoder to obtain the spatial ECG waveform. The waveform fusion module is used to fuse the frequency domain ECG waveform and the spatial domain ECG waveform to obtain the target ECG signal.

9. An electronic device, characterized in that, include: Memory, used to store computer programs; A processor for executing the computer program to implement the steps of the ECG signal acquisition method based on rPPG signals as described in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, Used to store a computer program; wherein, when the computer program is executed by a processor, it implements the steps of the ECG signal acquisition method based on the rPPG signal as described in any one of claims 1 to 7.