A face video physiological rhythm reconstruction method and related device
Patent Information
- Application Number
- CN202611083572.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-21
- Publication Date
- 2026-09-22
AI Technical Summary
[0005]本申请实施例的主要目的在于提出一种人脸视频生理节律重建方法、电子设备、存储介质及程序产品,旨在解决现有人脸视频生理信号编辑方法中,因缺乏一维生理时序信号与二维人脸空间区域的显式对齐机制,导致生理变化作用区域不明确、背景区域被错误调制、生成视频空间准确性不足的技术问题
[0027]本申请实施例至少包括以下有益效果:本申请提供一种人脸视频生理节律重建方法、装置、电子设备、存储介质及程序产品,该方案通过将原始人脸视频与原始生理时序信号时间同步后,构造目标生理编辑指令得到目标生理控制信号,再利用人脸空间先验信息将该一维目标生理控制信号映射为空间化生理条件特征,并以分层注入方式输入条件扩散模型的多尺度去噪网络生成编辑后人脸视频,由此实现了一维生理时序信号与二维人脸空间区域的显式对齐,使目标生理节律能够精准作用于面部有效区域(如皮肤区域、关键点区域等),避免了背景、头发、衣物等非生理响应区域被错误调制,解决了现有技术中因缺乏跨域空间对齐机制而导致的生理变化作用区域不明确、空间表达错位的技术问题,同时通过分层条件扩散注入使生理条件在多尺度去噪过程中持续影响视频生成,显著提高了生成视频的条件可控性、空间准确性和时序稳定性。
Smart Images

Figure CN122799062A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of facial video editing technology, and in particular to a method and related equipment for facial video physiological rhythm reconstruction. Background Technology
[0002] Remote photoplethysmography (rPPG) can capture subtle skin color fluctuations caused by cyclical changes in blood volume in ordinary facial videos, thereby estimating physiological indicators such as heart rate, respiratory rhythm, and heart rate variability. Because these signals are associated with individual physiological states, emotional arousal levels, stress responses, and fatigue, facial videos not only contain information about identity, expressions, and movements, but also implicit physiological rhythm information that can be mined by computational models. In recent years, researchers have begun exploring the editing and modification of rPPG physiological signals implicit in facial videos to serve applications such as physiological privacy protection, data augmentation, and physiological state simulation.
[0003] Existing facial video physiological signal editing techniques mainly fall into two categories. The first is based on conditional generative adversarial networks (GANs), represented by Phys-EdiGAN. This method embeds the target rPPG signal into the facial video through an adversarial game between the generator and discriminator. Its key technology lies in the generator learning the implicit mapping between the target rPPG signal and video frames. The second is based on 3D rendering, represented by a related patent from the University of California (US20250241548A1). This method generates facial videos containing the target rPPG waveform by rendering a 3D deformed facial model, UV physiological map modulation, lighting models, and camera models. Additionally, some schemes employ frame-level sinusoidal modulation to remove or encrypt the rPPG signal in specific facial regions.
[0004] However, the aforementioned existing technical solutions generally suffer from a common core technical deficiency: rPPG physiological signals are one-dimensional time series, while face videos are complex data consisting of two-dimensional space and one-dimensional time. Existing methods lack an effective mechanism for explicitly aligning the one-dimensional physiological time series signal with the two-dimensional facial spatial region. Specifically, methods based on conditional generative adversarial networks typically concatenate the target rPPG signal as a global conditional vector at the generator input. The physiological signal globally affects the entire frame image, lacking the ability to differentiate and control different facial regions, resulting in unclear areas of physiological change and potential mismodulation of background regions. While 3D rendering-based methods map physiological changes onto the 3D face surface through UV mapping, this method relies on explicit 3D deformation models and rendering pipelines. Its UV space deviates from the actual 2D facial spatial distribution of a specific person in the input video and struggles to handle existing face videos under different poses, expressions, and lighting conditions. Frame-level modulation methods use a fixed sine wave for global modulation, completely lacking spatial selectivity. Summary of the Invention
[0005] The main objective of this application is to propose a method, electronic device, storage medium, and program product for reconstructing physiological rhythms in facial videos. This aims to solve the technical problems in existing facial video physiological signal editing methods, which lack an explicit alignment mechanism between one-dimensional physiological temporal signals and two-dimensional facial spatial regions, resulting in unclear areas of physiological change action, incorrect modulation of background regions, and insufficient accuracy of generated video space.
[0006] To achieve the above objectives, one aspect of this application proposes a method for reconstructing the physiological rhythm of facial videos, the method comprising: Acquire the original face video, the original physiological time sequence signal, and the target editing task information, and perform time synchronization and video segmentation on the original face video and the original physiological time sequence signal; Physiological editing instructions are constructed based on the target editing task information, and target physiological control signals are obtained based on the physiological editing instructions. Using prior information about facial space, the target physiological control signal is mapped into spatialized physiological condition features; The spatialized physiological features are input into the multi-scale denoising network of the conditional diffusion model in a hierarchical conditional injection manner to generate an edited face video. Output the edited face video.
[0007] In some embodiments, in the video segmentation step, the original face video is segmented into multiple video segments according to a preset video segment length and sliding step, and the effective face region is located in each segmented video segment.
[0008] In some embodiments, the face spatial prior information is obtained through at least one of the following steps: Perform face detection on each frame of the original face video to generate a face mask that identifies the pixel region where the face is located; or, Perform skin region segmentation on each frame of the original face video to generate a skin region probability map that identifies the facial skin regions; or, Perform facial landmark detection on each frame of the original face video to generate a facial landmark heatmap that identifies the location and distribution of landmarks; or, Perform facial semantic segmentation on each frame of the original facial video to generate a facial semantic segmentation map.
[0009] In some embodiments, in the step of constructing physiological editing instructions based on the target editing task information, the editing operator type and corresponding adjustable control parameters are determined according to the target editing task. The editing operator type is selected from one or more of signal replacement, amplitude modulation, phase shift, frequency modulation, signal interpolation, or regional modulation. The adjustable control parameters include one or more of the target physiological signal, amplitude modulation coefficient, phase offset, frequency scaling coefficient, interpolation weight, local regional modulation weight, or editing intensity parameter.
[0010] In some embodiments, before mapping the target physiological control signal to spatialized physiological condition features, the method further includes the following steps: The target physiological control signal is input into the physiological prior modeling module, the target physiological control signal is differentially enhanced, and the signal difference between adjacent time steps is calculated to highlight local dynamic changes; Local feature encoding is performed on the differentially enhanced signal to obtain intermediate time-series features; By performing forward and reverse state scans using a state-space model-based periodic temporal modeling unit, long-range dependencies and periodic rhythms in the intermediate temporal features are captured, and physiological prior features are output.
[0011] In some embodiments, the state space model may be one or more of the following: a selective state space model, a bidirectional state space model, a Mamba module, an S4 module, an S5 module, a linear recursive state model, or a gated recursive state model.
[0012] In some embodiments, in the local feature encoding step, local features are filtered on the differentially enhanced signal by one-dimensional convolution or a lightweight temporal mapping function, and temporal location encoding information is superimposed.
[0013] In some embodiments, in the forward and reverse state scanning steps, state recursion is performed along the forward and reverse time axes, respectively. The forward scan output and the reverse scan output are fused and superimposed with periodic position embedding or frequency domain embedding to obtain the physiological prior features.
[0014] In some embodiments, in the step of mapping the target physiological control signal to spatialized physiological condition features, for each of the multiple scale levels of the conditional diffusion model, the face spatial prior and the face video intermediate features aligned with the scale of that layer are obtained respectively. The target physiological control signal or physiological prior features are then mapped across domains with the face spatial prior and the face video intermediate features using a spatial adaptation function to obtain the spatialized physiological condition features corresponding to that layer.
[0015] In some embodiments, the spatial adaptation function is implemented through a cross-modal attention mechanism: the target physiological control signal or physiological prior features are linearly transformed to generate a query vector; the intermediate features of the face video are linearly transformed to generate a key vector and a value vector; the attention weights of the query vector and the key vector are calculated; the value vector is weighted and summed using the attention weights to obtain the attention output; then the attention output is fused element-wise with the face spatial prior to obtain the spatialized physiological condition features of this layer.
[0016] In some embodiments, the spatial adaptation function is implemented through one of the following: spatial gating mechanism, deformable convolutional region weighting, facial key point map mapping, graph neural network region propagation, or conditional normalization mapping.
[0017] In some embodiments, in the hierarchical conditional injection step, the spatialized physiological conditional features corresponding to each scale level are injected into the encoding layer, bottleneck layer, decoding layer, residual layer, or attention layer of the corresponding scale of the conditional diffusion model.
[0018] In some embodiments, the hierarchical conditional injection is implemented through at least one of cross-attention injection, conditional normalization injection, FiLM modulation, AdaIN modulation, gated residual injection, Control branch injection, or multi-scale conditional token injection.
[0019] In some embodiments, generating the edited face video includes: The original face video or its latent space representation is subjected to diffusion forward noise addition processing to obtain noisy video features; Under the constraints of the spatialized physiological condition features, progressive reverse denoising sampling is performed. In each reverse denoising sampling step, the current denoising feature is modulated by the spatialized physiological condition features to gradually recover the edited face video without noise.
[0020] In some embodiments, the conditional diffusion model is one of the following: pixel space video diffusion model, latent space video diffusion model, three-dimensional U-Net diffusion model, spatiotemporal Transformer diffusion model, or DiT-type diffusion model.
[0021] In some embodiments, before outputting the edited face video, the method further includes the following steps: Estimated physiological signals are re-extracted from the edited face video using an independent physiological signal extractor or physiological signal estimation network; The estimated physiological signal is compared with the target physiological control signal in terms of time-domain waveform consistency, frequency-domain spectrum consistency, and heart rate estimation consistency. A physiological consistency score is generated based on the results of time-domain waveform consistency comparison, frequency-domain spectrum consistency comparison, and heart rate estimate consistency comparison.
[0022] In some embodiments, the method further includes a multi-objective joint training step, the multi-objective joint training step comprising: Obtain a training dataset, which includes original face video samples, original physiological time-series signal samples, and target physiological signal labels; The training dataset is input into the conditional diffusion model for training. During the training process, the diffusion noise prediction loss, video reconstruction loss, physiological consistency loss, temporal continuity loss, and identity structure preservation loss are optimized simultaneously.
[0023] In some embodiments, the video reconstruction loss constrains visual reconstruction quality by calculating the difference between the generated video and the target video in pixel space or perceptual feature space; the temporal continuity loss constrains inter-frame temporal smoothness by calculating the optical flow alignment error between adjacent frames; and the identity structure preservation loss constrains identity structure preservation by calculating the distance between the generated video and the original video in identity feature vector space or the geometric consistency of facial key points.
[0024] To achieve the above objectives, another aspect of this application provides an electronic device, which includes a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the method described above.
[0025] To achieve the above objectives, another aspect of the embodiments of this application proposes a computer-readable storage medium storing a computer program that, when executed by a processor, implements the method described above.
[0026] To achieve the above objectives, another aspect of the embodiments of this application proposes a computer program product, including a computer program that, when executed by a processor, implements the method described above.
[0027] The embodiments of this application include at least the following beneficial effects: This application provides a method, apparatus, electronic device, storage medium, and program product for reconstructing physiological rhythms in facial videos. This solution synchronizes the original facial video with the original physiological temporal signal, constructs a target physiological editing instruction to obtain a target physiological control signal, and then uses prior information about facial space to map the one-dimensional target physiological control signal into spatialized physiological condition features. The edited facial video is generated by inputting the feature into a multi-scale denoising network of a conditional diffusion model using a hierarchical injection method. This achieves explicit alignment between the one-dimensional physiological temporal signal and the two-dimensional facial spatial region, enabling the target physiological rhythm to accurately act on the effective facial region (such as the skin region, key point region, etc.), avoiding the incorrect modulation of non-physiological response regions such as background, hair, and clothing. This solves the technical problem in the prior art where the physiological change action area is unclear and the spatial expression is misaligned due to the lack of a cross-domain spatial alignment mechanism. At the same time, the hierarchical conditional diffusion injection allows physiological conditions to continuously affect video generation during the multi-scale denoising process, significantly improving the conditional controllability, spatial accuracy, and temporal stability of the generated video. Attached Figure Description
[0028] Figure 1 This is a flowchart of a method for reconstructing the physiological rhythm of a face video provided in an embodiment of this application; Figure 2 This is a schematic diagram of the physiological editing instruction construction module in an embodiment of this application; Figure 3 This is a flowchart of a method for controllable reconstruction of physiological rhythms of facial videos based on conditional diffusion, provided in an embodiment of this application. Figure 4 This is a schematic diagram of the state-space physiological cycle prior modeling module in an embodiment of this application; Figure 5 This is a schematic diagram of face spatial adaptation and hierarchical conditional diffusion injection in the embodiments of this application; Figure 6 This is a schematic diagram of the closed-loop verification of physiological consistency after generation in the embodiments of this application, showing the process of extracting the estimated rPPG signal from the reconstructed face video in reverse and comparing it with the target physiological signal in terms of correlation, frequency domain consistency, heart rate error, phase error and periodic stability; Figure 7 This is a schematic diagram of the structure of a facial video physiological rhythm reconstruction system provided in an embodiment of this application; Figure 8 This is a schematic diagram of the hardware structure of the electronic device provided in the embodiments of this application. Detailed Implementation
[0029] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative of this application and are not intended to limit it. In the following description, when referring to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with those of this application; they are merely examples of apparatuses and methods consistent with some aspects of the embodiments of this application as detailed in the appended claims.
[0030] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used herein is for the purpose of describing embodiments of this application only and is not intended to limit this application.
[0031] Existing remote physiological sensing technologies mostly focus on extracting rPPG signals or estimating heart rate from videos, while video generation and editing technologies pay more attention to explicit visual attributes such as facial expressions, movements, identity, lighting, or style. Regarding how to inject target physiological rhythms as controllable conditions into facial videos, so that the generated videos present preset physiological signals while maintaining visual naturalness, identity structure, and temporal continuity, current technologies still lack stable and verifiable solutions.
[0032] Some existing methods attempt to modify physiological signals in face videos through pixel perturbation, generative adversarial networks, or latent space editing to achieve privacy protection or target signal replacement. However, these approaches typically suffer from problems such as insufficient modeling of the target physiological temporal conditions, lack of explicit alignment between one-dimensional physiological signals and two-dimensional facial spatial regions, limited stability of adversarial training, weak cross-temporal consistency, and insufficient verifiability of the edited physiological signals.
[0033] Therefore, it is necessary to propose a face video physiological rhythm controllable reconstruction technology that is different from traditional generative adversarial physiological editing. By explicitly constructing physiological editing instructions, extracting the periodic prior of the target physiological signal using a state space model, mapping the one-dimensional physiological prior to the effective face region, and injecting multi-scale conditions during the diffusion model reverse denoising process, a unified visual quality, physiological consistency and temporal stability can be achieved.
[0034] Furthermore, the University of California patent US20250241548A1 proposes a method for synthesizing face videos based on 3D face modeling and UV physiological map modulation. Its main process includes: receiving an input image; encoding the input image into a UV albedo map, a 3D mesh, an illumination model, and a camera model; decomposing the UV albedo map to obtain a UV physiological map or a UV blood map; modifying the UV physiological map according to the target's rPPG signal; and then combining the illumination model and camera model to render a synthesized face video containing changes in the target's pulse. The key technical aspect of this approach lies in the synthetic rendering chain based on 3D face modeling, UV mapping, and biophysical skin reflection or blood map modulation.
[0035] The aforementioned existing technologies can generate face videos with target rPPG waveforms, but their core relies on 3D deformable face models, UV physiological maps, blood map modulation, lighting models, camera models, and rendering processes. They are mainly used to generate synthetic data, improve the training effect of rPPG detection models, or alleviate data distribution bias. The technical problem of this type of method lies in how to generate face videos with biophysical interpretations from a single reference image or face model, rather than achieving multi-scale conditional injection and controllable reconstruction of target physiological rhythms through a diffusion-based reverse denoising process within existing face video spatiotemporal sequences.
[0036] Furthermore, existing methods for editing physiological signals in facial videos primarily employ conditional generative adversarial networks (GANs) to modify rPPG information in the video. Their key technical focus typically lies in learning the implicit mapping between the target rPPG signal and video frames through a generator, and maintaining facial consistency in the edited video through identity preservation constraints. In view of this, this application proposes a method, electronic device, storage medium, and program product for reconstructing physiological rhythms in facial videos based on conditional diffusion. This application does not use an adversarial game between the generator and discriminator as the core generation mechanism, but instead employs a stepwise reverse denoising process of a conditional diffusion model to reconstruct physiological rhythms in facial videos. Thus, this application forms a five-feature collaborative link: "physiological editing instruction construction, state-space physiological prior modeling, facial space adaptation, hierarchical conditional diffusion injection, and post-generation physiological consistency closed-loop verification," which clearly defines the technical boundaries with synthetic video generation methods based on 3D UV physiological map rendering and GAN-based physiological privacy editing methods.
[0037] The differences between this application and the aforementioned University of California patent are as follows: First, this application does not require UV albedo map, 3D mesh, illumination model, camera model, or explicit blood map rendering as necessary steps. Instead, it uses the original face video clip and physiological editing instructions as input to directly reconstruct physiological rhythms in the video feature space or latent space diffusion feature space. Second, this application extracts the dynamic prior of the target physiological signal through differential enhancement and periodic modeling based on the state space model, rather than linearly modulating it through the UV blood map and rPPG scaling factor. Third, this application maps the one-dimensional physiological prior to the effective face region or multi-scale video feature layer through a face space prior adapter, rather than fixing physiological changes to the UV mapping space. Fourth, this application introduces physiological priors layer by layer during the multi-scale denoising process of the diffusion model through a layered conditional diffusion injection mechanism, rather than generating the final frame through the 3D rendering pipeline. Fifth, this application performs closed-loop verification by comparing the generated rPPG with the target signal, thereby ensuring the verifiability of the implicit physiological rhythms in the edited video.
[0038] This application aims to solve the aforementioned problems existing in the prior art and has the following inventive objectives: (1) Solve the problem that face video synthesis methods based on 3D face models, UV physiological maps and explicit rendering links are difficult to directly reconstruct the physiological rhythm in existing face videos in a controllable manner; (2) To address the problem that existing methods, when using the target rPPG signal directly as the generation condition, fail to adequately express the periodicity of physiological signals, local dynamic changes, and long-range dependencies; (3) Solve the problem that one-dimensional physiological temporal features are difficult to accurately align with two-dimensional face spatial regions, facial skin regions and multi-scale video features; (4) Solve the problems of limited training stability, insufficient cross-frame continuity and weak verifiability of generated results in adversarial generation-based physiological editing methods; (5) Establish a state-space physiological prior modeling mechanism, and obtain highly reliable physiological prior features suitable for use in diffusion models through differential enhancement, selective state recursion and periodic rhythm extraction; (6) Establish a feature adaptation mechanism guided by facial spatial priors, and orient physiological priors to the effective facial region, skin region, key point region or spatial attention region; (7) Establish a hierarchical conditional diffusion injection mechanism so that spatialized physiological conditions can be injected into the multi-scale coding layer, decoding layer, residual layer or denoising layer of the diffusion model; (8) Establish a closed-loop verification mechanism for physiological consistency after generation. By re-extracting the rPPG signal from the edited video and comparing it with the target physiological signal, the verifiability and credibility of the physiological editing results can be improved.
[0039] The facial video physiological rhythm reconstruction method provided in this application relates to the interdisciplinary fields of computer vision, generative artificial intelligence, remote physiological perception, emotion computing, facial video editing, and trusted multimodal modeling. The facial video physiological rhythm reconstruction method provided in this application can be applied to a terminal, a server, or software running on a terminal or server. In some embodiments, the terminal can be a smartphone, tablet, laptop, desktop computer, smart speaker, smartwatch, or in-vehicle terminal, but is not limited to these. The server can be configured as an independent physical server, a server cluster composed of multiple physical servers, or a distributed system. It can also be configured as a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms. The server can also be a node server in a blockchain network. The software can be an application implementing the facial video physiological rhythm reconstruction method, but is not limited to the above forms.
[0040] This application can be used in a wide variety of general-purpose or special-purpose computer system environments or configurations. Examples include: personal computers, server computers, handheld or portable devices, tablet devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics, network PCs, minicomputers, mainframe computers, and distributed computing environments including any of the above systems or devices. This application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc., that perform specific tasks or implement specific abstract data types. This application can also be practiced in distributed computing environments where tasks are performed by remote processing devices connected via a communication network. In distributed computing environments, program modules can reside in local and remote computer storage media, including storage devices.
[0041] like Figure 1 As shown, this embodiment provides a method for reconstructing the physiological rhythm of facial videos, specifically including the following steps: Step S101: Obtain the original face video, the original physiological time sequence signal, and the target editing task information, and perform time synchronization and video segmentation on the original face video and the original physiological time sequence signal.
[0042] In one implementation, the video segmentation step involves dividing the original face video into multiple video segments according to a preset video segment length and sliding step size, and locating the effective face region in each segmented video segment.
[0043] Step S102: Construct physiological editing instructions based on the target editing task information, and obtain target physiological control signals based on the physiological editing instructions.
[0044] In one implementation, in the step of constructing physiological editing instructions based on the target editing task information, the editing operator type and corresponding adjustable control parameters are determined according to the target editing task. The editing operator type is selected from one or more of signal replacement, amplitude modulation, phase shift, frequency modulation, signal interpolation, or regional modulation. The adjustable control parameters include one or more of the target physiological signal, amplitude modulation coefficient, phase offset, frequency scaling coefficient, interpolation weight, local regional modulation weight, or editing intensity parameter.
[0045] Step S103: Using prior information about the face space, the target physiological control signal is mapped into spatialized physiological condition features.
[0046] In some embodiments, the face space prior information is obtained through at least one of the following steps: performing face detection on each frame of the original face video to generate a face mask identifying the pixel region where the face is located; or performing skin region segmentation on each frame of the original face video to generate a skin region probability map identifying the facial skin region; or performing facial key point detection on each frame of the original face video to generate a facial key point heatmap identifying the key point location distribution; or performing face semantic segmentation on each frame of the original face video to generate a face semantic segmentation map.
[0047] In some embodiments, before mapping the target physiological control signal to spatialized physiological condition features, the method further includes the following steps: inputting the target physiological control signal into a physiological prior modeling module; firstly, performing differential enhancement on the target physiological control signal and calculating the signal difference between adjacent time steps to highlight local dynamic changes; then, performing local feature encoding on the differentially enhanced signal to obtain intermediate temporal features; finally, performing forward and reverse state scanning through a state-space model-based periodic temporal modeling unit to capture long-range dependencies and periodic rhythms in the intermediate temporal features and outputting physiological prior features.
[0048] Specifically, the state space model adopts one or more of the following: selective state space model, bidirectional state space model, Mamba module, S4 module, S5 module, linear recursive state model, or gated recursive state model.
[0049] Specifically, in the local feature encoding step, local features are filtered on the differentially enhanced signal through one-dimensional convolution or a lightweight temporal mapping function, and temporal location encoding information is superimposed.
[0050] In the forward and reverse state scanning steps, state recursion is performed along the forward and reverse time axes, respectively. The forward scan output and the reverse scan output are fused and superimposed with periodic position embedding or frequency domain embedding to obtain the physiological prior features.
[0051] In one embodiment, in the step of mapping the target physiological control signal to spatialized physiological condition features, for each of the multiple scale levels of the conditional diffusion model, the face spatial prior and the face video intermediate features aligned with the scale of that layer are obtained respectively. The target physiological control signal or physiological prior features are then mapped across domains with the face spatial prior and the face video intermediate features using a spatial adaptation function to obtain the spatialized physiological condition features corresponding to that layer.
[0052] Optionally, the spatial adaptation function is implemented through a cross-modal attention mechanism: the target physiological control signal or physiological prior features are linearly transformed to generate a query vector; the intermediate features of the face video are linearly transformed to generate a key vector and a value vector; the attention weights of the query vector and the key vector are calculated; the value vector is weighted and summed using the attention weights to obtain the attention output; then the attention output is fused element-wise with the face spatial prior to obtain the spatialized physiological condition features of this layer.
[0053] The spatial adaptation function is implemented through one of the following: spatial gating mechanism, deformable convolution region weighting, facial key point map mapping, graph neural network region propagation, or conditional normalization mapping.
[0054] Step S104: Input the spatialized physiological condition features into the multi-scale denoising network of the conditional diffusion model in a hierarchical conditional injection manner to generate the edited face video.
[0055] Specifically, the hierarchical conditional injection is implemented through at least one of the following methods: cross-attention injection, conditional normalization injection, FiLM modulation, AdaIN modulation, gated residual injection, Control branch injection, or multi-scale conditional token injection.
[0056] Optionally, the conditional diffusion model may be one of the following: pixel space video diffusion model, latent space video diffusion model, three-dimensional U-Net diffusion model, spatiotemporal Transformer diffusion model, or DiT-type diffusion model.
[0057] As one implementation method, generating the edited face video includes: The original face video or its latent space representation is subjected to diffusion forward noise addition processing to obtain noisy video features; Under the constraints of the spatialized physiological condition features, progressive reverse denoising sampling is performed. In each reverse denoising sampling step, the current denoising feature is modulated by the spatialized physiological condition features to gradually recover the edited face video without noise.
[0058] Step S105: Output the edited face video.
[0059] In some embodiments, before outputting the edited face video, the method further includes the following steps: re-extracting estimated physiological signals from the edited face video using an independent physiological signal extractor or physiological signal estimation network; comparing the estimated physiological signals with the target physiological control signals in terms of time-domain waveform consistency, frequency-domain spectrum consistency, and heart rate estimation consistency; and generating a physiological consistency score based on the results of the time-domain waveform consistency comparison, frequency-domain spectrum consistency comparison, and heart rate estimation consistency comparison.
[0060] In some embodiments, the method further includes a multi-objective joint training step: obtaining a training dataset, the training dataset including original face video samples, original physiological time-series signal samples and target physiological signal labels; inputting the training dataset into a conditional diffusion model for training, and simultaneously optimizing diffusion noise prediction loss, video reconstruction loss, physiological consistency loss, temporal continuity loss and identity structure preservation loss during the training process.
[0061] Specifically, the video reconstruction loss constrains the visual reconstruction quality by calculating the difference between the generated video and the target video in pixel space or perceptual feature space; the temporal continuity loss constrains the inter-frame temporal smoothness by calculating the optical flow alignment error between adjacent frames; and the identity structure preservation loss constrains identity structure preservation by calculating the distance between the generated video and the original video in identity feature vector space or the geometric consistency of facial key points.
[0062] In one embodiment, this application extends physiological signal editing from a single target waveform input to parameterizable physiological editing instructions. See also Figure 2 According to Table 1, the physiological editing instructions include at least one or more of the following: target physiological signal, amplitude modulation coefficient, phase offset, frequency scaling coefficient, interpolation weight, local region modulation weight, or editing intensity parameter.
[0063] Table 1
[0064] Physiological editing instructions This can be formally represented as: ,in Represents the set of edit operators. This represents the adjustable control parameter corresponding to the operator.
[0065] The family of edit operators is defined as follows: ,in For signal replacement, It is amplitude modulation. For phase translation, For frequency modulation, For signal interpolation, It is a region modulation.
[0066] Different editing operators can be uniformly written as: ,in Control the editing intensity. Control the response of local areas.
[0067] in, Indicates the target physiological signal, Represents the amplitude modulation coefficient. This indicates the phase offset. Indicates the frequency scaling factor. Indicates the interpolation weights. This represents the physiological modulation weight of the r-th facial region. Indicates the editing intensity parameter; This indicates the global editing control signal generated when facial regions are not distinguished. This represents the regional editing control signal applied to the r-th facial region. This represents the original physiological time sequence signal that is synchronized with the original face video.
[0068] The solutions of the embodiments of this application will be described in detail below with reference to the accompanying drawings and specific application examples.
[0069] like Figure 3 As shown, this embodiment provides a method for controllable reconstruction of physiological rhythms in facial videos based on conditional diffusion, specifically including the following steps: Step S1: Multimodal data acquisition, time synchronization and face spatial prior extraction.
[0070] Obtain the original face video primitive physiological time-series signals And target editing task information. This includes timestamp alignment, sampling rate unification, abnormal frame removal, video segmentation, and effective face region localization for facial videos and physiological signals.
[0071] Video segmentation can be represented as:
[0072] in, Indicates the length of a video segment. Indicates the sliding step size. T This indicates the total number of frames in the original video after time synchronization and preprocessing. K This represents the number of valid video segments obtained by segmenting based on the video segment length L and the sliding step size S. The spatial prior information for the face can consist of one or more of the following: face mask, skin region probability map, facial landmark heatmap, face semantic segmentation map, facial blood vessel distribution prior map, or spatial attention saliency map.
[0073] Step S2: Construction of physiological editing instructions.
[0074] Physiological editing instructions are constructed based on the target editing task. If the target task is to protect physiological privacy, alternative signals different from the original signals can be constructed; if the target task is to simulate physiological states, target waveforms can be constructed based on target heart rate, respiratory rate, or pressure status; if the target task is to augment data, various controllable variants can be generated through amplitude, phase, frequency, and region weights.
[0075] The unified editing control signal can be represented as:
[0076] in This represents the facial region index. When local regions are not distinguished, it is denoted as... .
[0077] By transforming the target physiological state into explicit editing instructions, this application avoids the problem of insufficient control semantics caused by using only a single rPPG waveform as a conditional input, and enables the subsequent diffusion reconstruction process to receive structured, interpretable and adjustable physiological control conditions.
[0078] Step S3: State-space physiological cycle prior modeling.
[0079] See Figure 4 The editing control signal is input into the physiological prior modeling module. This module first performs differential enhancement on the editing control signal to highlight local dynamic changes between adjacent time steps; then it performs local feature filtering through one-dimensional convolution or lightweight temporal mapping; finally, it captures long-range dependencies and periodic rhythms through a state-space model-based periodic temporal modeling unit.
[0080] Differential enhancement and local dynamic coding can be represented as: ,in Indicates time location encoding. This represents a one-dimensional convolution or linear mapping function.
[0081] The selective state-space recursion can be expressed as:
[0082] in, This represents the discretized state transition matrix at time t. This represents the discretized input injection matrix corresponding to time t. This represents the output projection matrix that maps the hidden states to the output space. This represents a skip connection matrix or feedforward matrix that directly maps the current input features to the output space. This represents the hidden state at time t. This represents the state-space output characteristics at time t. This represents the physiological temporal characteristics of the input to the state-space model at time t.
[0083] The generation and discretization of input-dependent parameters can be represented as:
[0084] Here, softplus is the softplus function. This represents the learnable projection matrix used to generate adaptive discrete walk lengths. This represents the learnable projection matrix used to generate the input injection matrix. This represents the learnable projection matrix used to generate the output projection matrix. This represents the corresponding learnable bias term. This represents the adaptive distance walk length at time t. A This represents the state transition matrix in a continuous-time state-space model. I This represents the identity matrix that matches the dimension of the state transition matrix.
[0085] Bidirectional periodic modeling output It can be represented as:
[0086] in and These represent the forward and reverse state scan outputs, respectively. This represents the periodic position or frequency domain embedding, where Z represents the input feature sequence obtained through differential enhancement and local feature encoding. LN represents the learnable output projection matrix used to fuse the forward state scan output, the reverse state scan output, and the input feature sequence.
[0087] In one embodiment, the state-space model includes a selective state-space model, a bidirectional state-space model, a Mamba module, an S4 module, a linear recursive state model, or a combination thereof. Through this module, the target physiological signal is no longer directly input into the generation network as a low-dimensional waveform, but is instead transformed into physiological prior features containing local dynamics, long-range dependencies, and periodic rhythms. .
[0088] Step S4: Feature adaptation guided by face spatial priors.
[0089] Since physiological prior features are one-dimensional or low-dimensional temporal representations, while facial videos are complex data with spatial and temporal dimensions, direct fusion can easily lead to unclear areas of physiological change, incorrect modulation of background areas, or misalignment of local facial physiological expressions. Therefore, this application embodiment includes a facial spatial adaptation module that performs cross-domain mapping between physiological prior features and facial spatial prior information.
[0090] Face spatial adaptation can be represented as:
[0091] in, In the conditional diffusion model, the first... At each scale level, the prior information of the face space used; Indicates the scale level of the diffusion network. Indicates the first Layered facial video features, This represents the face space prior aligned with this scale. This represents the physiological prior feature sequence output by the state-space physiological cycle prior modeling module. The physiological prior feature sequence at least characterizes the local dynamic changes, long-range temporal dependence, and periodic rhythm information of the editing control signal. The spatial adaptation function can be one or more of the following: cross attention, spatial gating, deformable convolution region weighting, facial keypoint map mapping, graph neural network region propagation, or conditional normalization mapping.
[0092] In a preferred implementation, spatial adaptation employs cross-modal attention:
[0093] in, Indicates by the first Query features obtained by projecting layered facial video features. This represents the key features obtained by projecting physiological prior features. This represents the value feature obtained by projecting physiological prior features. Indicates the first Cross-modal attention weights of layered face video features to physiological prior features. , ,and They represent the first The corresponding query projection matrix, key projection matrix, and value projection matrix for each layer. Indicates the first The channel dimension of the layer query feature or key feature is used to scale the attention inner product result.
[0094] Therefore, we obtain the first... Physiological conditions for layered spatialization:
[0095] in, This represents element-wise multiplication. This indicates a learnable bias or residual compensation term. This indicates that the face space is prior. This indicates the characteristics of spatialized physiological conditions.
[0096] Step S5: Layered conditional diffusion reverse denoising and reconstruction.
[0097] See Figure 5 This paper proposes a multi-scale denoising network for a conditional diffusion model by inputting spatialized physiological conditional features into the model, thereby subjecting the diffusion model to the target physiological rhythm during the reverse denoising process. Unlike methods that only concatenate conditions at the input end, this application employs a hierarchical conditional injection mechanism, injecting physiological conditions into multiple scale levels of the video U-Net or latent space video diffusion model, including coding layers, bottleneck layers, decoding layers, residual layers, or attention layers.
[0098] The forward noise addition and noise prediction objective can be expressed as:
[0099] in, This represents the sample or latent space features after noise is added at the t-th diffusion time step. This represents a noise-free video sample or latent space feature corresponding to the original face video. This represents the cumulative signal retention coefficient over the first t diffusion time steps. Indicates a Gaussian distribution. This represents the predicted loss for diffuse noise. This represents the expectation operation performed on noise-free samples, random noise, and diffusion time steps. This represents the Gaussian noise obtained from the sampling. This represents a conditional noise prediction network with parameter θ. This represents the total number of scale levels involved in conditional injection.
[0100] No. Conditional injection at each scale level can be written as: ,in Represents the original denoising features. This represents the updated denoising features after modulation. Represents the gated modulation function. This represents the cross-attention or its equivalent conditional injection operator.
[0101] The inverse denoising distribution can be represented as: Thus, the edited video is obtained through stepwise sampling. .in, This represents the conditional inverse denoising probability distribution characterized by parameter θ, where the probability distribution is conditioned on the current noise sample, the diffusion time step, and the spatialized physiological conditions. This represents the predicted mean of the inverse denoising distribution. This represents the prediction covariance matrix or noise variance parameter of the inverse denoising distribution.
[0102] The conditional injection methods include at least one of cross-attention injection, conditional normalization injection, FiLM modulation, AdaIN modulation, gated residual injection, Control branch injection, or multi-scale conditional token injection. Through hierarchical conditional injection, physiological priors can simultaneously influence low-level skin texture changes, mid-level local region responses, and high-level temporal structure consistency.
[0103] In one embodiment, to ensure that the generated video simultaneously satisfies visual naturalness, identity structure preservation, target physiological consistency, and inter-frame continuity, the present invention employs multi-objective joint training constraints. The training objectives include at least diffuse noise prediction loss, video reconstruction loss, physiological consistency loss, temporal continuity loss, and identity structure preservation loss.
[0104] The joint training objective can be expressed as: .
[0105] It is used to constrain the pixel consistency and perceptual consistency of the effective area of the face.
[0106] ,in Indicates optical flow The inter-frame alignment function is used to constrain temporal continuity.
[0107] It is used to constrain the facial identity structure and the geometric consistency of key points.
[0108] in, This represents the predicted loss for diffuse noise. Indicates the video reconstruction loss. Indicates loss of physiological consistency. Indicates the loss of time continuity. This represents the loss that preserves the identity structure and keypoint geometry. This indicates the editing intensity or parameter regularization term. Indicates weight, This represents a spatial mask or pixel-weighted map corresponding to the effective area of the face, used to highlight facial skin areas and other effective physiological response areas. This represents the feature extraction function used to compute perceptual consistency in videos. This refers to the reconstructed or generated frame of the edited video at time t. This represents the optical flow field or temporal alignment mapping from time t to time t+1. This represents the identity feature extraction function used to extract facial identity representation. The weights represent the geometric consistency constraints of facial key points. This represents a function for facial landmark detection, regression, or geometric structure representation.
[0109] It should be noted that the structure preservation loss is not limited to a specific identity preservation component, but can be achieved by facial structural feature similarity, key point geometric consistency, or facial region reconstruction constraints.
[0110] Step S6: Post-generation physiological consistency closed-loop verification See Figure 6 After obtaining the edited facial video Subsequently, in this embodiment, the estimated physiological signal is re-extracted using an independent rPPG extractor or a physiological signal estimation network. and compare it with the target signal in the physiological editing instructions. Perform a consistency comparison.
[0111]
[0112] in, It is used to measure the temporal consistency between the generated video inverse extracted signal and the target signal. , Indicates the weight.
[0113] in, Represents the correlation coefficient. Represents frequency domain transformation. This represents the heart rate value extracted from the edited video. This represents the target heart rate value. This closed-loop mechanism differs from implicit supervision methods based on the self-evaluation module within an adversarial network; its purpose is to reverse-engineer the generated results to verify whether the target physiological rhythm is accurately represented.
[0114] During the training phase, the closed-loop verification participates in joint optimization as a loss function to guide the model to generate face videos consistent with the target physiological rhythm. During the inference phase, the closed-loop verification can be used as an optional post-verification step to output a physiological consistency score, rather than a necessary step in reconstructing the video generation process.
[0115] Step S7: Output the edited face video, physiological consistency evaluation indicators, and technical report. Final output: edited face video PhysioScore, a comprehensive score for physiological consistency, and various evaluation indicators are provided for users to verify the reconstruction effect.
[0116] Example 1: Target Heart Rate Replacement Scenario This embodiment uses target heart rate replacement as an example to explain in detail the specific application process of this application.
[0117] Input a face video containing the raw rPPG signal Input the target physiological signal corresponding to the target heart rate. System constructs signal replacement instructions. Specifically, the physiological editing instruction construction module constructs instructions based on the original physiological signals. and target physiological signals Generate editing instructions .
[0118] The state-space physiological prior modeling module for target physiological signals Processing steps: First, differential enhancement is performed. This highlights the local dynamic changes in the signal; then, local feature selection is performed through one-dimensional convolution. Finally, long-range dependencies and periodic rhythms are modeled using the Mamba module (a selective state-space model) to obtain physiological prior features. .
[0119] The face spatial adaptation module utilizes face masks. One-dimensional physiological priors Mapped to spatialized physiological condition characteristics Specifically, for each scale level l of the diffuse U-Net, physiological priors are aligned with the video features of the corresponding level through cross-modal attention: And apply face mask constraints: .
[0120] The hierarchical conditional diffusion reconstruction module will spatialize physiological condition features. The U-Net network of the conditional diffusion model is input using a hierarchical injection approach. Conditional injection is performed at the encoding layer, bottleneck layer, and decoding layer of the U-Net. The diffusion model generates edited videos containing the target heart rate rhythm through a stepwise reverse denoising process. .
[0121] The physiological consistency closed-loop verification module after generation starts from... Re-extracting and estimating the rPPG signal ,calculate and Correlation coefficient between Heart rate error This verifies whether the target heart rate is accurately represented. When the correlation coefficient is greater than a preset threshold (e.g., 0.8) and the heart rate error is less than a preset threshold (e.g., 3 BPM), the physiological consistency check is confirmed to be successful.
[0122] The final output is a face video containing the target heart rate after editing. And physiological consistency evaluation report.
[0123] Example 2: Physiological Amplitude Enhancement Scenario This embodiment uses physiological amplitude enhancement as an example to illustrate the application of this application in emotional arousal simulation or training data enhancement scenarios.
[0124] Input a face video containing the raw rPPG signal The system is based on the amplitude modulation coefficient. The raw physiological signals are enhanced. Specifically, the physiological editing instruction construction module generates editing instructions. This indicates that the amplitude of the original physiological signal will be increased by 50%.
[0125] The state-space physiological prior modeling module for the enhanced target signal Differential enhancement and state-space modeling are performed to obtain physiological priors containing stronger physiological cycle characteristics. .
[0126] The face spatial adaptation module maps physiological priors to the skin region (rather than the entire face region), ensuring that amplitude enhancement primarily affects the facial skin area. Specifically, it utilizes a skin region probability map. As a spatial priori: .
[0127] The hierarchical conditional diffusion reconstruction module injects spatialized physiological conditions into the multi-scale layers of the diffusion U-Net to generate amplitude-enhanced videos. This makes the subtle periodic changes in facial skin tone more noticeable in the edited video.
[0128] The closed-loop verification module extracts the rPPG signal from the generated video. And calculate its relationship with the target enhancement signal. Correlation (i.e., 1.5 times the original amplitude) and spectral consistency This verifies the amplitude enhancement effect.
[0129] Output edited face video And physiological consistency evaluation indicators.
[0130] Example 3: Regionalized Physiological Modulation Scenario This embodiment uses regional physiological modulation as an example to illustrate the facial region differential modulation capability of this application.
[0131] Input a face video containing the raw rPPG signal The system sets different region modulation weights for different facial areas such as the cheeks, forehead, and nose.
[0132] Specifically, the physiological editing instruction construction module generates editing instructions that include regional modulation weights. ,in This is a physiological modulation weight vector corresponding to each facial region. For example, setting a weight vector for the cheek region... =1.2 (Enhanced), setting for the forehead area. =0.8 (reduced), set for the nasal alar area. =1.0 (keep).
[0133] The face spatial adaptation module utilizes face semantic segmentation maps. Obtain the spatial location of each facial region and incorporate physiological priors. With regional weight Combined, regionally differentiated spatial physiological conditions are generated: ,in Let be the mask for the r-th facial region.
[0134] The hierarchical conditional diffusion reconstruction module injects regionally differentiated physiological conditions into the diffusion U-Net to generate regionally physiologically modulated face videos. This allows physiological rhythms to primarily act on skin areas such as the cheeks rather than hair, background, clothing, or non-facial areas, thereby reducing ineffective area modulation and visual distortion.
[0135] Closed-loop verification module from rPPG signals were extracted to verify whether different regions exhibited differentiated physiological response intensities.
[0136] Example 4: Dual-signal interpolation scenario This embodiment uses dual-signal interpolation as an example to illustrate the application of this application in the generation of physiological state transitions.
[0137] Input a face video containing the raw rPPG signal and target physiological signals The system constructs a transitional physiological editing signal using interpolation weights λ, which is used to generate facial video clips that continuously change between two physiological states.
[0138] Specifically, the physiological editing instruction construction module generates editing instructions. ,in λ=0.5 represents the intermediate transition state between the initial state and the target state.
[0139] The state-space physiological prior modeling module applies interpolated signals Differential augmentation and state-space modeling are performed to obtain physiological priors. .
[0140] The face spatial adaptation module maps physiological priors to the effective face region, and the hierarchical conditional diffusion reconstruction module generates the interpolated face video. .
[0141] In a multi-frame interpolation scenario, the system can set different interpolation weights for different time steps. ,make The process linearly changes from 0 to 1 over time, thus generating continuous video clips that smoothly transition from the original physiological state to the target physiological state. .
[0142] The closed-loop verification module extracts rPPG signals frame by frame or segment by segment from the generated transition video to verify the smooth transition effect of the physiological state.
[0143] In summary, compared with the prior art, this application has at least the following advantages and beneficial effects: (1) By constructing a physiological editing instruction mechanism, the target physiological state is expanded from a single waveform input to interpretable, adjustable and combinable control parameters, thereby improving the coverage of editing tasks; (2) By using state-space physiological cycle prior modeling, the ability to represent local dynamic changes, long-range dependence and periodic rhythm of physiological signals is enhanced, avoiding the insufficient conditional expression caused by directly inputting low-dimensional rPPG waveforms; (3) By using a feature adapter guided by the prior knowledge of the face space, physiological information can be directed to the effective area of the face, the skin area or the key physiological response area, thereby reducing the interference of the background and non-face areas. (4) By using the hierarchical conditional diffusion injection mechanism, the target physiological prior can be applied to multiple scale levels of the diffusion denoising network, thereby improving the conditional controllability and fine-grained physiological expression ability of the generated video. (5) By generating rPPG and comparing it with the target signal, a closed-loop verification is formed, so that the editing results not only meet the visual quality requirements, but also have verifiable physiological consistency. (6) Improve training stability, temporal continuity and multi-task expansion capability by using a diffusion-based reconstruction path that differs from that of conditional generative adversarial networks; (7) Improve patent exclusivity and anti-circumvention capabilities by using state-space models, spatial priors and condition injection methods to lay out technical feature families; (8) By supporting operations such as replacement, enhancement, reduction, translation, frequency modulation, interpolation and regional modulation, this application can be expanded to include the protection of physiological privacy in face video, emotion computing, health data enhancement and physiological state simulation.
[0144] See Figure 7 This application also provides a facial video physiological rhythm reconstruction system, which can implement the above-mentioned method. The system includes: The multimodal data acquisition and synchronization module is used to acquire the original face video, the original physiological time-series signal and the target editing task information, and to perform time synchronization and video segmentation on the original face video and the original physiological time-series signal. A physiological editing instruction construction module is used to construct physiological editing instructions based on the target editing task information and obtain target physiological control signals based on the physiological editing instructions; A face space adaptation module is used to map the target physiological control signal into spatialized physiological condition features using face space prior information; The hierarchical conditional diffusion reconstruction module is used to input the spatialized physiological condition features into the multi-scale denoising network of the conditional diffusion model in a hierarchical conditional injection manner to generate an edited face video. The output module is used to output the edited face video.
[0145] It is understood that the content of the above method embodiments is applicable to this system embodiment. The specific functions implemented in this system embodiment are the same as those in the above method embodiments, and the beneficial effects achieved are also the same as those achieved in the above method embodiments.
[0146] This application also provides an electronic device, which includes a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the above-described method. This electronic device can be any smart terminal, including tablet computers, in-vehicle computers, etc.
[0147] It is understood that the content of the above method embodiments is applicable to this device embodiment. The specific functions implemented by this device embodiment are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.
[0148] Please see Figure 8 , Figure 8The hardware structure of an electronic device according to another embodiment is illustrated. The electronic device includes: The processor 801 can be implemented using a general-purpose CPU (Central Processing Unit), microprocessor, application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of this application. The memory 802 can be implemented as a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM). The memory 802 can store the operating system and other application programs. When the technical solutions provided in the embodiments of this specification are implemented through software or firmware, the relevant program code is stored in the memory 802 and is called and executed by the processor 801 using the methods described in the embodiments of this application. The 803 input / output interface is used to implement information input and output. The communication interface 804 is used to enable communication and interaction between this device and other devices. Communication can be achieved through wired means (such as USB, Ethernet cable, etc.) or wireless means (such as mobile network, WIFI, Bluetooth, etc.). Bus 805 transmits information between various components of the device (e.g., processor 801, memory 802, input / output interface 803, and communication interface 804); The processor 801, memory 802, input / output interface 803, and communication interface 804 are connected to each other within the device via bus 805.
[0149] This application also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the above-described method.
[0150] It is understood that the content of the above method embodiments is applicable to this storage medium embodiment. The specific functions implemented in this storage medium embodiment are the same as those in the above method embodiments, and the beneficial effects achieved are also the same as those achieved in the above method embodiments.
[0151] Memory, as a non-transitory computer-readable storage medium, can be used to store non-transitory software programs and non-transitory computer-executable programs. Furthermore, memory may include high-speed random access memory, and may also include non-transitory memory, such as at least one disk storage device, flash memory device, or other non-transitory solid-state storage device. In some embodiments, memory may optionally include memory remotely located relative to the processor, and these remote memories can be connected to the processor via a network. Examples of such networks include, but are not limited to, the Internet, intranets, local area networks, mobile communication networks, and combinations thereof.
[0152] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the above-described method.
[0153] It is understood that the content of the above method embodiments is applicable to the embodiments of this program product. The specific functions implemented in the embodiments of this program product are the same as those in the above method embodiments, and the beneficial effects achieved are also the same as those achieved in the above method embodiments. The executable computer program code or "code" used to perform the various embodiments can be written in high-level programming languages such as C, C++, Python, Smalltalk, Java, JavaScript, Visual Basic, Structured Query Language (e.g., Transact-SQL), Perl, or in various other programming languages.
[0154] The embodiments described in this application are for the purpose of more clearly illustrating the technical solutions of the embodiments of this application, and do not constitute a limitation on the technical solutions provided by the embodiments of this application. As those skilled in the art will know, with the evolution of technology and the emergence of new application scenarios, the technical solutions provided by the embodiments of this application are also applicable to similar technical problems.
[0155] Those skilled in the art will understand that the technical solutions shown in the figures do not constitute a limitation on the embodiments of this application, and may include more or fewer steps than shown, or combine certain steps, or different steps.
[0156] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs.
[0157] Those skilled in the art will understand that all or some of the steps in the methods disclosed above, as well as the functional modules / units in the systems and devices, can be implemented as software, firmware, hardware, or suitable combinations thereof.
[0158] The terms “first,” “second,” “third,” “fourth,” etc. (if present) in the specification and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms “comprising” and “having,” and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0159] It should be understood that in this application, "at least one (item)" means one or more, and "more than" means two or more. "And / or" is used to describe the relationship between related objects, indicating that three relationships can exist. For example, "A and / or B" can represent three cases: only A exists, only B exists, and both A and B exist simultaneously, where A and B can be singular or plural. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship. "At least one (item) of the following" or similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one (item) of a, b, or c can represent: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, and c can be single or multiple.
[0160] In the several embodiments provided in this application, it should be understood that the disclosed apparatus and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of the units described above is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.
[0161] The units described above as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0162] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0163] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes multiple instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this application. The aforementioned storage medium includes various media capable of storing programs, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0164] The preferred embodiments of the present application have been described above with reference to the accompanying drawings, but this does not limit the scope of the claims of the present application. Any modifications, equivalent substitutions, and improvements made by those skilled in the art without departing from the scope and substance of the embodiments of the present application shall be within the scope of the claims of the present application.
Claims
1. A method for reconstructing physiological rhythms from facial videos, characterized in that, The method includes the following steps: Acquire the original face video, the original physiological time sequence signal, and the target editing task information, and perform time synchronization and video segmentation on the original face video and the original physiological time sequence signal; Physiological editing instructions are constructed based on the target editing task information, and target physiological control signals are obtained based on the physiological editing instructions. Using prior information about facial space, the target physiological control signal is mapped into spatialized physiological condition features; The spatialized physiological features are input into the multi-scale denoising network of the conditional diffusion model in a hierarchical conditional injection manner to generate an edited face video. Output the edited face video.
2. The method according to claim 1, characterized in that, The prior information about the face space is obtained through at least one of the following steps: Perform face detection on each frame of the original face video to generate a face mask that identifies the pixel region where the face is located; or, Perform skin region segmentation on each frame of the original face video to generate a skin region probability map that identifies the facial skin regions; or, Perform facial landmark detection on each frame of the original face video to generate a facial landmark heatmap that identifies the location and distribution of landmarks; or, Perform facial semantic segmentation on each frame of the original facial video to generate a facial semantic segmentation map.
3. The method according to claim 1, characterized in that, In the step of constructing physiological editing instructions based on the target editing task information, the editing operator type and corresponding adjustable control parameters are determined according to the target editing task. The editing operator type is selected from one or more of signal replacement, amplitude modulation, phase shift, frequency modulation, signal interpolation, or regional modulation. The adjustable control parameters include one or more of the target physiological signal, amplitude modulation coefficient, phase offset, frequency scaling coefficient, interpolation weight, local regional modulation weight, or editing intensity parameter.
4. The method according to claim 1, characterized in that, Before mapping the target physiological control signal into spatialized physiological condition features, the method further includes the following steps: The target physiological control signal is input into the physiological prior modeling module, the target physiological control signal is differentially enhanced, and the signal difference between adjacent time steps is calculated to highlight local dynamic changes; Local feature encoding is performed on the differentially enhanced signal to obtain intermediate time-series features; By performing forward and reverse state scans using a state-space model-based periodic temporal modeling unit, long-range dependencies and periodic rhythms in the intermediate temporal features are captured, and physiological prior features are output.
5. The method according to claim 1, characterized in that, In the step of mapping the target physiological control signal into spatialized physiological condition features, for each layer in the multiple scale levels of the conditional diffusion model, the face spatial prior and the face video intermediate features aligned with the scale of that layer are obtained respectively. The target physiological control signal or physiological prior features are then mapped across domains with the face spatial prior and the face video intermediate features using a spatial adaptation function to obtain the spatialized physiological condition features corresponding to that layer.
6. The method according to claim 1, characterized in that, The generation of the edited face video includes: The original face video or its latent space representation is subjected to diffusion forward noise addition processing to obtain noisy video features; Under the constraints of the spatialized physiological condition features, progressive reverse denoising sampling is performed. In each reverse denoising sampling step, the current denoising feature is modulated by the spatialized physiological condition features to gradually recover the edited face video without noise.
7. The method according to claim 1, characterized in that, Before outputting the edited face video, the method further includes the following steps: Estimated physiological signals are re-extracted from the edited face video using an independent physiological signal extractor or physiological signal estimation network; The estimated physiological signal is compared with the target physiological control signal in terms of time-domain waveform consistency, frequency-domain spectrum consistency, and heart rate estimation consistency. A physiological consistency score is generated based on the results of time-domain waveform consistency comparison, frequency-domain spectrum consistency comparison, and heart rate estimate consistency comparison.
8. The method according to claim 1, characterized in that, The method further includes a multi-objective joint training step, which includes: Obtain a training dataset, which includes original face video samples, original physiological time-series signal samples, and target physiological signal labels; The training dataset is input into the conditional diffusion model for training. During the training process, the diffusion noise prediction loss, video reconstruction loss, physiological consistency loss, temporal continuity loss, and identity structure preservation loss are optimized simultaneously.
9. An electronic device, characterized in that, The electronic device includes a memory and a processor, the memory storing a computer program, and the processor executing the computer program to implement the method according to any one of claims 1 to 8.
10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it implements the method of any one of claims 1 to 8.
Citation Information
Patent Citations
Synthetic Generation of Face Videos with Plethysmograph Physiology
US20250241548A1