Facial video-based blood pressure estimation method, device, equipment and storage medium

CN122805231APending Publication Date: 2026-09-25ATHENAEYES CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611282192.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-08-24
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

[0003]现有技术在获取目标面部视频的血压值时,未考虑各窗口信息量及可靠性的差异,导致低质量窗口的噪声掩盖了高质量窗口的有效特征,从而降低了目标面部视频的血压值的可靠性

Benefits of technology

[0016]本申请实施例有益效果在于以下两方面:

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122805231A_ABST
    Figure CN122805231A_ABST
Patent Text Reader

Abstract

The application relates to the technical field of deep learning and the technical field of data processing, and discloses a blood pressure estimation method and device based on a facial video, equipment and a storage medium. The method comprises the following steps: generating a weight value corresponding to each training window according to the recognition degree of each training window and a weight model of a training stage; generating a total loss of a student model according to the weight value corresponding to each training window, a blood pressure estimation value of each training window, a blood pressure reference value of each training window, a shape distillation loss of the student model and a predefined total loss model, training the student model with the total loss as a target, and saving the student model after the training; and outputting a blood pressure estimation value of each inference window through the student model after the training, and generating a blood pressure value of a target facial video according to the weight value corresponding to each inference window, the blood pressure estimation value of each inference window and a blood pressure estimation model. The application improves the reliability of the blood pressure value of the target facial video.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the fields of deep learning technology and data processing technology, and in particular to a method, apparatus, device and storage medium for blood pressure estimation based on facial video. Background Technology

[0002] Blood pressure is a core indicator for assessing hypertension and cardiovascular risk. Traditional blood pressure measurement relies on cuff-type blood pressure monitors or invasive arterial pressure measurements, which, while highly accurate, are not suitable for high-frequency, remote, non-intrusive, and mobile health management scenarios.

[0003] Existing technologies for acquiring blood pressure values ​​from target facial videos do not consider the differences in information content and reliability among different windows. This leads to noise from low-quality windows masking the effective features of high-quality windows, thus reducing the reliability of blood pressure values ​​from the target facial video. Therefore, how to acquire blood pressure values ​​from target facial videos is a technical problem. Summary of the Invention

[0004] This application provides a method, apparatus, device, and storage medium for estimating blood pressure based on facial video, in order to solve the aforementioned technical problem of obtaining blood pressure values ​​from target facial videos.

[0005] In a first aspect, embodiments of this application provide a blood pressure estimation method based on facial video, applied to an electronic device, the blood pressure estimation method comprising: Acquire the original facial video, perform feature extraction on the original facial video, and generate the photoplethysmography signal of the original facial video. The photoplethysmography (PPG) signal of the original facial video is divided into PPG signals of multiple training windows, and the PPG signal of each training window is input into the student model. The student model outputs the morphological feature score of each training window. The morphological feature score of each training window and the recognition model of the training stage are used to generate the recognition score of each training window. Based on the recognition score of each training window and the weight model of the training stage, the weight value corresponding to each training window is generated. Based on the weight value corresponding to each training window, the blood pressure estimate of each training window, the blood pressure reference value of each training window, the morphological distillation loss of the student model, and the predefined total loss model, the total loss of the student model is generated. With the goal of reducing the total loss, the student model is trained and the trained student model is saved. The photoplethysmography (PPG) signal of the target facial video is divided into PPG signals of multiple inference windows. The PPG signal of each inference window is input into the trained student model. The trained student model outputs the blood pressure estimate of each inference window. Based on the weight value corresponding to each inference window, the blood pressure estimate of each inference window and the blood pressure estimation model, the blood pressure value of the target facial video is generated.

[0006] In one possible implementation of the first aspect, the morphological feature score of each training window is output by the student model. The discrimination score of each training window is then used in conjunction with the discrimination model during the training phase to generate the discrimination score of each training window. Finally, based on the discrimination score of each training window and the weight model during the training phase, a weight value corresponding to each training window is generated, including: The first branch of the student model outputs the label uncertainty of each training window, the second branch of the student model outputs the prediction uncertainty of each training window, the third branch of the student model outputs the blood pressure estimate of each training window, and the fourth branch of the student model outputs the morphological feature score of each training window. The recognition score of each training window is generated by the morphological feature score of each training window and the recognition model during the training phase. Based on the recognition score of each training window, the label uncertainty of each training window, the prediction uncertainty of each training window, and the weight model during the training phase, the weight value corresponding to each training window is generated.

[0007] In one possible implementation of the first aspect, the total loss of the student model is generated based on the weight value corresponding to each training window, the estimated blood pressure value for each training window, the reference blood pressure value for each training window, the morphological distillation loss of the student model, and a predefined total loss model. The student model is then trained with the goal of reducing the total loss, and the trained student model is saved, including: Obtain the annotation information corresponding to the original facial video. The annotation information includes all blood pressure measurements obtained by a cuff blood pressure monitor when the original facial video was captured. All blood pressure measurements are arranged in chronological order to generate a discrete data sequence with time on the horizontal axis and blood pressure on the vertical axis. The reference blood pressure value for each training window is obtained from the discrete data sequence by interpolation. Obtain the feature vectors output by the encoders of the teacher model and the student model. Based on the feature vectors output by the encoders of the teacher model and the student model, and a predefined morphological distillation loss model, generate the morphological distillation loss of the student model. Based on the weight values ​​corresponding to each training window, the blood pressure estimate of each training window, the blood pressure reference value of each training window, the morphological distillation loss of the student model, and a predefined total loss model, generate the total loss of the student model. With the goal of reducing the total loss, train the student model until the total loss is less than the preset loss, then stop training the student model and save the trained student model.

[0008] In one possible implementation of the first aspect, the photoplethysmography (PPG) signal of the target facial video is divided into PPG signals of multiple inference windows. The PPG signal of each inference window is input into a trained student model. The trained student model outputs a blood pressure estimate for each inference window. Based on the weight value corresponding to each inference window, the blood pressure estimate for each inference window, and the blood pressure estimation model, a blood pressure value for the target facial video is generated, including: Acquire a target facial video, perform feature extraction on the target facial video, generate an opto-pole permeable (OPP) signal of the target facial video, and divide the OPP signal of the target facial video into OPP signals of multiple inference windows. The photoplethysmography signal of each inference window is input into the trained student model. The label uncertainty of each inference window is output through the first branch of the trained student model, the prediction uncertainty of each inference window is output through the second branch of the trained student model, the blood pressure estimate of each inference window is output through the third branch of the trained student model, and the morphological feature score of each inference window is output through the fourth branch of the trained student model. The recognition score of each inference window is generated by the morphological feature score of each inference window and the recognition model of the inference stage. Based on the recognition score of each inference window, the label uncertainty of each inference window, the prediction uncertainty of each inference window, and the weight model of the inference stage, the weight value of each inference window is generated. Based on the weight value of each inference window, the blood pressure estimate of each inference window, and the blood pressure estimation model, the blood pressure value of the target facial video is generated.

[0009] In one possible implementation of the first aspect, the discrimination model during the training phase is defined as follows: ; This represents the discriminative power of the w-th training window; the higher the discriminative power of the w-th training window, the more complete the comprehensive features of the w-th training window; the lower the discriminative power of the w-th training window, the less complete the comprehensive features of the w-th training window. w represents the sequence number of the training window; Represents the Sigmoid function; Represents the weight vector; This represents the video signal quality characteristics of the w-th training window, used to reflect the video signal quality of the w-th training window; The larger the value, the clearer the video image in the w-th training window; The smaller the value, the blurrier the video image in the w-th training window; This represents the signal quality characteristics of the w-th training window, used to reflect the photoplethysmography signal quality of the w-th training window; The larger the value, the higher the signal-to-noise ratio of the photoplethysmography signal in the w-th training window; The smaller the value, the lower the signal-to-noise ratio of the photoplethysmography signal in the w-th training window; This represents the morphological feature score of the w-th training window, used to reflect whether the pulse wave features of the w-th training window are complete; The larger the value, the more complete the pulse wave features of the w-th training window; The smaller the value, the less the pulse wave features are missing in the w-th training window; This represents the consistency feature of the w-th training window, used to reflect the degree of consistency between the photoplethysmography signal and the reference signal in the w-th training window. The larger the value, the more consistent the photoplethysmography signal of the w-th training window is with the reference signal; The smaller the value, the less consistent the photoplethysmography signal of the w-th training window is with the reference signal; This represents the uncertainty of the w-th training window; Indicates the penalty weight; The weight model during the training phase is defined as follows: ; This represents the weight value corresponding to the w-th training window; This represents the label uncertainty of the w-th training window; the larger the label uncertainty of the w-th training window, the less reliable the blood pressure reference value of the w-th training window is; the smaller the label uncertainty of the w-th training window, the more reliable the blood pressure reference value of the w-th training window is. This represents the prediction uncertainty for the w-th training window; The larger the value, the lower the confidence level of the student model in the blood pressure estimation results for the w-th training window. The smaller the value, the higher the confidence level of the student model in the blood pressure estimation results for the w-th training window; This represents a preset positive integer, used to avoid the denominator being zero.

[0010] In one possible implementation of the first aspect, the morphological distillation loss model is defined as follows: ; This represents the morphological distillation loss of the student model; w represents the sequence number of the training window; This represents the recognition score of the w-th training window; This represents the set of all training windows; This represents the photoplethysmography signal input to the student model in the w-th training window. The encoder representing the student model; The mapping layer representing the student model; This represents the feature vector output by the encoder of the student model; This represents the projected features obtained after the feature vector output by the encoder of the student model is projected through the mapping layer. This represents the photoplethysmography signal input to the teacher model at the w-th training window; The encoder representing the teacher model; This represents the feature vector output by the encoder of the teacher model; Represents the square of the L2 norm; The total loss model is defined as follows: ; L represents the total training loss of the student model; This represents the weight value corresponding to the w-th training window; This represents the blood pressure estimate for the w-th training window; This represents the blood pressure reference value for the w-th training window; This represents the physiological constraint loss of the student model; This represents the morphological distillation loss of the student model; This represents the classification loss of the student model; , , These represent the first weight, the second weight, and the third weight, respectively.

[0011] In one possible implementation of the first aspect, the blood pressure estimation model is defined as follows: ; This indicates the estimated blood pressure value for the current video. This represents the set of inference windows, which consists of all inference windows obtained from the segmentation of the current video. The weight value corresponding to the z-th inference window indicates that the greater the weight of the blood pressure estimate of the z-th inference window in the fusion process, the greater its impact on the blood pressure estimate of the reference facial video; conversely, the smaller the weight of the blood pressure estimate of the z-th inference window in the fusion process, the smaller its impact on the blood pressure estimate of the reference facial video. This represents the blood pressure estimate for the z-th inference window.

[0012] Secondly, embodiments of this application provide a blood pressure estimation device based on facial video, applied to an electronic device, comprising: The acquisition module is used to acquire the original facial video, perform feature extraction on the original facial video, and generate the photoplethysmography signal of the original facial video. The input module is used to divide the photoplethysmography (PPG) signal of the original facial video into PPG signals of multiple training windows, and input the PPG signal of each training window into the student model. The first generation module is used to output the morphological feature score of each training window through the student model, generate the recognition score of each training window through the morphological feature score of each training window and the recognition model of the training stage, and generate the weight value corresponding to each training window based on the recognition score of each training window and the weight model of the training stage. The second generation module is used to generate the total loss of the student model based on the weight value corresponding to each training window, the blood pressure estimate of each training window, the blood pressure reference value of each training window, the morphological distillation loss of the student model, and the predefined total loss model. With the goal of reducing the total loss, the student model is trained and the trained student model is saved. The third generation module is used to divide the photoplethysmography (PPG) signal of the target facial video into PPG signals of multiple inference windows, input the PPG signal of each inference window into the trained student model, output the blood pressure estimate of each inference window through the trained student model, and generate the blood pressure value of the target facial video based on the weight value corresponding to each inference window, the blood pressure estimate of each inference window and the blood pressure estimation model.

[0013] Thirdly, embodiments of this application provide an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the blood pressure estimation method described in the first aspect above.

[0014] Fourthly, embodiments of this application provide a computer-readable storage medium storing a computer program that, when executed by a processor, implements the blood pressure estimation method described in the first aspect above.

[0015] Fifthly, embodiments of this application provide a computer program product that, when run on an electronic device, causes the electronic device to execute the blood pressure estimation method described in the first aspect.

[0016] The beneficial effects of the embodiments of this application are as follows: Firstly, the photoplethysmography (PPG) signal of the target facial video is divided into PPG signals of multiple inference windows. The PPG signal of each inference window is input into the trained student model. The trained student model outputs the blood pressure estimate of each inference window. Based on the weight value corresponding to each inference window, the blood pressure estimate of each inference window, and the blood pressure estimation model, the blood pressure value of the target facial video is generated. By using the weight value corresponding to each inference window, the high-quality and high-confidence inference window dominates the blood pressure value generation process, effectively avoiding the problem of effective features being masked due to equal weight fusion in existing technologies, and improving the reliability of the blood pressure value of the target facial video. Secondly, by using the weight value corresponding to each inference window, the blood pressure value of the target facial video is no longer subject to drastic deviations due to signal quality fluctuations in individual windows, thus enhancing the stability of the blood pressure value of the target facial video. Attached Figure Description

[0017] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0018] Figure 1 This is a diagram illustrating an application scenario of the blood pressure estimation method provided in the embodiments of this application. Figure 2 This is a schematic flowchart of the blood pressure estimation method provided in the embodiments of this application; Figure 3 A flowchart illustrating the implementation of S202 provided in this application embodiment; Figure 4 A schematic block diagram of a blood pressure estimation device provided in the embodiments of this application; Figure 5 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation

[0019] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application. All other embodiments obtained by those skilled in the art based on the embodiments in this application without inventive effort are within the scope of protection of this application.

[0020] It should be understood that in the description of this application and the appended claims, the terms "first," "second," "third," etc., are used only to distinguish descriptions and should not be construed as indicating or implying relative importance. The terms "comprising," "including," "having," and their variations all mean "including but not limited to," unless otherwise specifically emphasized.

[0021] Furthermore, the technical solutions of the various embodiments can be combined with each other, but only if they are based on the ability of those skilled in the art to implement them. When the combination of technical solutions is contradictory or cannot be implemented, it should be considered that such combination of technical solutions does not exist and is not within the scope of protection claimed in this application.

[0022] The blood pressure estimation method provided in this application can be applied to electronic devices, including but not limited to servers, mobile phones, tablets, vehicle-mounted devices, and laptops. This application does not impose any restrictions on the specific type of electronic device.

[0023] Please see Figure 1 , Figure 1The application scenario diagram of the blood pressure estimation method provided in the embodiments of this application is described in detail below: The electronic device connects to a file storage system, retrieves the original facial video from the file storage system, performs feature extraction on the original facial video, and generates a photoplethysmography signal of the original facial video.

[0024] The original facial videos are pre-processed video sample data, mainly used by student and teacher models for training and learning.

[0025] In this embodiment of the application, the electronic device obtains the original facial video from the file storage system, realizing the standardized access of the original facial video.

[0026] Please see Figure 2 , Figure 2 This is a flowchart illustrating the blood pressure estimation method provided in an embodiment of this application, which can be applied to electronic devices.

[0027] like Figure 2 As shown, the blood pressure estimation method provided in this application includes the following steps, detailed below: S201, acquire the original facial video, perform feature extraction on the original facial video, and generate the photoplethysmography signal of the original facial video; The periodic changes in pixel brightness of the skin region in the original facial video can reflect the periodic changes in blood volume in microvessels. By performing time-domain and frequency-domain analysis on the brightness signal of the region of interest in the original facial video, the photoplethysmography signal of the facial video can be extracted.

[0028] S202, the photoplethysmography signal of the original facial video is divided into photoplethysmography signals of multiple training windows, and the photoplethysmography signal of each training window is input into the student model. The training window refers to the window used by the student model during the training phase.

[0029] S203: Output the morphological feature score of each training window through the student model; generate the recognition score of each training window and the recognition model of the training stage through the morphological feature score of each training window; generate the weight value corresponding to each training window based on the recognition score of each training window and the weight model of the training stage. Specifically, the student model outputs a morphological feature score for each training window. Using this score and the recognition rate model from the training phase, the recognition rate of each training window is generated. Finally, based on the recognition rate and the weight model from the training phase, a weight value corresponding to each training window is generated, including: The first branch of the student model outputs the label uncertainty of each training window, the second branch of the student model outputs the prediction uncertainty of each training window, the third branch of the student model outputs the blood pressure estimate of each training window, and the fourth branch of the student model outputs the morphological feature score of each training window. The recognition score of each training window is generated by the morphological feature score of each training window and the recognition model during the training phase. Based on the recognition score of each training window, the label uncertainty of each training window, the prediction uncertainty of each training window, and the weight model during the training phase, the weight value corresponding to each training window is generated.

[0030] The first branch is a lightweight multilayer perceptron or a one-dimensional temporal convolutional network, used to extract features from the photoplethysmography signal of each input window and output the label uncertainty of each window. The second branch consists of a one-dimensional convolutional layer, a bidirectional long short-term memory network or a Transformer temporal coding layer and a regression head. It performs deep temporal modeling on the photoplethysmography signal of each window and outputs the prediction uncertainty of each window. The third branch consists of a one-dimensional convolutional layer and a regression head, which is used to extract local temporal features and perform blood pressure regression mapping on the photoplethysmography signal of each input window, and output the blood pressure estimate of each window. The fourth branch consists of a one-dimensional convolutional layer and a fully connected classification layer, which is used to extract and classify the waveform morphological features of the photoplethysmography signal of each input window and output the morphological feature score of each window.

[0031] The recognition model during the training phase is defined as follows: ; This represents the discriminative power of the w-th training window; the higher the discriminative power of the w-th training window, the more complete the comprehensive features of the w-th training window; the lower the discriminative power of the w-th training window, the less complete the comprehensive features of the w-th training window. w represents the sequence number of the training window; Represents the Sigmoid function; Represents the weight vector; This represents the video signal quality characteristics of the w-th training window, used to reflect the video signal quality of the w-th training window; The larger the value, the clearer the video image in the w-th training window; The smaller the value, the blurrier the video image in the w-th training window; This represents the signal quality characteristics of the w-th training window, used to reflect the photoplethysmography signal quality of the w-th training window; The larger the value, the higher the signal-to-noise ratio of the photoplethysmography signal in the w-th training window; The smaller the value, the lower the signal-to-noise ratio of the photoplethysmography signal in the w-th training window; This represents the morphological feature score of the w-th training window, used to reflect whether the pulse wave features of the w-th training window are complete; The larger the value, the more complete the pulse wave features of the w-th training window; The smaller the value, the less the pulse wave features are missing in the w-th training window; This represents the consistency feature of the w-th training window, used to reflect the degree of consistency between the photoplethysmography signal and the reference signal in the w-th training window. The larger the value, the more consistent the photoplethysmography signal of the w-th training window is with the reference signal; The smaller the value, the less consistent the photoplethysmography signal of the w-th training window is with the reference signal; This represents the uncertainty of the w-th training window; Indicates the penalty weight; The weight model during the training phase is defined as follows: ; This represents the weight value corresponding to the w-th training window; This represents the label uncertainty of the w-th training window; the larger the label uncertainty of the w-th training window, the less reliable the blood pressure reference value of the w-th training window is; the smaller the label uncertainty of the w-th training window, the more reliable the blood pressure reference value of the w-th training window is. This represents the prediction uncertainty for the w-th training window; The larger the value, the lower the confidence level of the student model in the blood pressure estimation results for the w-th training window. The smaller the value, the higher the confidence level of the student model in the blood pressure estimation results for the w-th training window; This represents a preset positive integer, used to avoid the denominator being zero.

[0032] Pulse wave characteristics include the rising limb, reflected wave, diastolic peak, and pulse width.

[0033] The rising segment refers to the curve segment of the pulse wave from its lowest point to its highest point.

[0034] The diastolic peak is the peak value reached by the dicrotic wave during the diastolic phase.

[0035] Pulse width refers to the total time length that a pulse wave takes from the beginning of one cycle to the beginning of the next cycle.

[0036] S204. Based on the weight value corresponding to each training window, the blood pressure estimate value of each training window, the blood pressure reference value of each training window, the morphological distillation loss of the student model, and the predefined total loss model, generate the total loss of the student model. With the goal of reducing the total loss, train the student model and save the trained student model. Specifically, based on the weight values ​​corresponding to each training window, the estimated blood pressure value for each training window, the reference blood pressure value for each training window, the morphological distillation loss of the student model, and a predefined total loss model, the total loss of the student model is generated. With the goal of reducing the total loss, the student model is trained, and the trained student model is saved, including: Obtain the annotation information corresponding to the original facial video. The annotation information includes all blood pressure measurements obtained by a cuff blood pressure monitor when the original facial video was captured. All blood pressure measurements were arranged chronologically, generating a discrete data sequence with time on the horizontal axis and blood pressure on the vertical axis. Reference blood pressure values ​​for each training window were then extracted from this discrete data sequence using interpolation. This interpolation method avoids errors in the model's learning of mapping relationships caused by temporal misalignment. Compared to directly using the average blood pressure from the entire video as a label, interpolation preserves the natural fluctuations in blood pressure over time rather than smoothing it out entirely. This allows the student model to capture richer dynamic patterns, learning the continuous changes in blood pressure over time, rather than simply fitting a global average. Compared to nearest neighbor assignment, interpolation avoids abrupt, physiologically inaccurate jumps at measurement intervals, resulting in a smoother, more continuous label sequence that better reflects the true changes in blood pressure.

[0037] Specifically, blood pressure reference values ​​for each training window are obtained from the discrete data sequence through interpolation, including: The discrete data sequence is interpolated using linear interpolation to generate the target data sequence, and the blood pressure reference value for each training window is obtained from the target data sequence.

[0038] Obtain the feature vectors output by the encoders of the teacher model and the student model. Based on the feature vectors output by the encoders of the teacher model and the student model, and a predefined morphological distillation loss model, generate the morphological distillation loss of the student model. Based on the weight values ​​corresponding to each training window, the blood pressure estimate of each training window, the blood pressure reference value of each training window, the morphological distillation loss of the student model, and a predefined total loss model, generate the total loss of the student model. With the goal of reducing the total loss, train the student model until the total loss is less than the preset loss, then stop training the student model and save the trained student model.

[0039] The student model is trained until the total loss is less than the preset loss, at which point the training stops and the trained student model is saved. This ensures that the trained student model has stable reasoning ability.

[0040] The speciation distillation loss model is defined as follows: ; This represents the morphological distillation loss of the student model; w represents the sequence number of the training window; This represents the recognition score of the w-th training window; This represents the set of all training windows; This represents the photoplethysmography signal input to the student model in the w-th training window. The encoder representing the student model; The mapping layer representing the student model; This represents the feature vector output by the encoder of the student model; This represents the projected features obtained after the feature vector output by the encoder of the student model is projected through the mapping layer. This represents the photoplethysmography signal input to the teacher model at the w-th training window; The encoder representing the teacher model; This represents the feature vector output by the encoder of the teacher model; Represents the square of the L2 norm; The total loss model is defined as follows: ; L represents the total training loss of the student model; This represents the weight value corresponding to the w-th training window; This represents the blood pressure estimate for the w-th training window; This represents the blood pressure reference value for the w-th training window; This represents the physiological constraint loss of the student model; This represents the morphological distillation loss of the student model; This represents the classification loss of the student model; , , These represent the first weight, the second weight, and the third weight, respectively.

[0041] The blood pressure estimate for the w-th training window includes the systolic blood pressure estimate and the diastolic blood pressure estimate for the w-th training window.

[0042] The blood pressure reference values ​​for the w-th training window include the systolic blood pressure reference value and the diastolic blood pressure reference value for the w-th training window.

[0043] : The systolic blood pressure prediction error for the w-th training window is generated by subtracting the systolic blood pressure reference value from the estimated systolic blood pressure value for the w-th training window; the diastolic blood pressure prediction error for the w-th training window is generated by subtracting the diastolic blood pressure reference value from the estimated diastolic blood pressure value for the w-th training window; and the absolute values ​​of the systolic and diastolic blood pressure prediction errors are added together to generate the absolute blood pressure error for the w-th training window.

[0044] The process of obtaining the physiological constraint loss for the student model is as follows: Obtain the sorting penalty, range penalty, smoothing penalty, and consistency penalty of the student model in the w-th training window. Add the sorting penalty, range penalty, smoothing penalty, and consistency penalty of the student model in the w-th training window to generate the loss of the student model in the w-th training window. Add the losses of the student model in all training windows to generate the physiological constraint loss of the student model. Among them, the sorting penalty term is defined as the penalty value applied when the diastolic blood pressure in the blood pressure estimate of the w-th training window is higher than the systolic blood pressure; the range penalty term is defined as the penalty value applied when the blood pressure estimate of the w-th training window exceeds the preset range; the smoothing penalty term is defined as the penalty value applied when the change in blood pressure estimate between the w-th training window and the adjacent window exceeds the preset change rate threshold; and the consistency penalty term is defined as the penalty value applied when there is a contradiction between the blood pressure estimate and the heart rate in the w-th training window.

[0045] The process of obtaining the morphological distillation loss for the student model is as follows: Obtain the classification penalty term of the student model in the w-th training window, and sum the classification penalty terms of the student model in the w-th training window to generate the classification loss of the student model. The graded penalty term is defined as the penalty applied when the blood pressure estimate of the w-th training window is at a different blood pressure level than the blood pressure level corresponding to the mean of all cuff blood pressure measurements.

[0046] S205, the photoplethysmography (PPG) signal of the target facial video is divided into PPG signals of multiple inference windows. The PPG signal of each inference window is input into the trained student model. The trained student model outputs the blood pressure estimate of each inference window. Based on the weight value corresponding to each inference window, the blood pressure estimate of each inference window and the blood pressure estimation model, the blood pressure value of the target facial video is generated.

[0047] Specifically, the photoplethysmography (PPG) signal of the target facial video is divided into PPG signals for multiple inference windows. The PPG signal of each inference window is input into a trained student model. The trained student model outputs a blood pressure estimate for each inference window. Based on the weight values ​​corresponding to each inference window, the blood pressure estimate for each inference window, and the blood pressure estimation model, the blood pressure value of the target facial video is generated, including: Acquire a target facial video, perform feature extraction on the target facial video, generate an opto-pole permeable (OPP) signal of the target facial video, and divide the OPP signal of the target facial video into OPP signals of multiple inference windows. The photoplethysmography signal of each inference window is input into the trained student model. The label uncertainty of each inference window is output through the first branch of the trained student model, the prediction uncertainty of each inference window is output through the second branch of the trained student model, the blood pressure estimate of each inference window is output through the third branch of the trained student model, and the morphological feature score of each inference window is output through the fourth branch of the trained student model. The recognition score of each inference window is generated by the morphological feature score of each inference window and the recognition model of the inference stage. Based on the recognition score of each inference window, the label uncertainty of each inference window, the prediction uncertainty of each inference window, and the weight model of the inference stage, the weight value of each inference window is generated. Based on the weight value of each inference window, the blood pressure estimate of each inference window, and the blood pressure estimation model, the blood pressure value of the target facial video is generated.

[0048] The blood pressure estimate in the inference window includes the systolic blood pressure estimate and the diastolic blood pressure estimate.

[0049] The blood pressure values ​​for the target facial video include systolic and diastolic blood pressure values.

[0050] The recognition model for the reasoning stage is defined as follows: ; This represents the discriminability of the z-th inference window; the higher the discriminability of the z-th inference window, the more complete its comprehensive features; the lower the discriminability of the z-th inference window, the less complete its comprehensive features. z represents the sequence number of the training window; Represents the Sigmoid function; Represents the weight vector; This represents the video signal quality characteristics of the z-th inference window, used to reflect the video signal quality of the z-th inference window; The larger the value, the clearer the video image of the z-th inference window; The smaller the value, the blurrier the video image in the z-th inference window; This represents the signal quality characteristics of the z-th inference window, used to reflect the photoplethysmography signal quality of the z-th inference window; The larger the value, the higher the signal-to-noise ratio of the photoplethysmography signal in the z-th inference window; The smaller the value, the lower the signal-to-noise ratio of the photoplethysmography signal in the z-th inference window; The morphological feature score of the z-th inference window is used to reflect whether the pulse wave features of the z-th inference window are complete. The larger the value, the more complete the pulse wave characteristics of the z-th inference window; The smaller the value, the less complete the pulse wave characteristics of the z-th inference window; This represents the consistency feature of the z-th inference window, used to reflect the degree of consistency between the photoplethysmography (PPG) recording signal and the reference signal in the z-th inference window. The larger the value, the more consistent the photoplethysmography signal of the z-th inference window is with the reference signal; The smaller the value, the less consistent the photoplethysmography signal of the z-th inference window is with the reference signal; This represents the uncertainty of the z-th inference window; Indicates the penalty weight; The weight model for the inference phase is defined as follows: ; This represents the weight value corresponding to the z-th inference window; This represents the label uncertainty of the z-th inference window; the larger the label uncertainty of the z-th inference window, the less reliable the blood pressure reference value of the z-th inference window is; the smaller the label uncertainty of the z-th inference window, the more reliable the blood pressure reference value of the z-th inference window is. This represents the prediction uncertainty of the z-th inference window; The larger the value, the lower the confidence level of the trained student model in the blood pressure estimation result of the z-th inference window. The smaller the value, the higher the confidence level of the trained student model in the blood pressure estimation result of the z-th inference window; This represents a preset positive integer, used to avoid the denominator being zero.

[0051] The blood pressure estimation model is defined as follows: ; This indicates the estimated blood pressure value for the current video. This represents the set of inference windows, which consists of all inference windows obtained from the segmentation of the current video. The weight value corresponding to the z-th inference window indicates that the greater the weight of the blood pressure estimate of the z-th inference window in the fusion process, the greater its impact on the blood pressure estimate of the reference facial video; conversely, the smaller the weight of the blood pressure estimate of the z-th inference window in the fusion process, the smaller its impact on the blood pressure estimate of the reference facial video. This represents the blood pressure estimate for the z-th inference window.

[0052] Among them, the target facial video refers to the facial video to be used for blood pressure prediction during the reasoning stage.

[0053] The inference window refers to the window used by the trained student model during the inference phase.

[0054] Based on the window length and sliding step size, the photoplethysmography (PPG) signal of the target facial video is divided into multiple inference window PPG signals. For example, if the total length of the PPG signal of the target facial video is 30 seconds, the preset window length is 3 seconds, and the sliding step size is 1 second, then starting from second 0, seconds 0 to 3, 1 to 4, and 2 to 5 are extracted sequentially until the signal ends, resulting in a total of 28 inference window PPG signals.

[0055] Specifically, based on the weight value corresponding to each inference window, the blood pressure estimate of each inference window, and the blood pressure estimation model, the blood pressure value of the target facial video is generated, reducing the stringent requirements on the target facial video. In actual application scenarios of mobile health management, users cannot guarantee that they will remain still and have ideal lighting conditions throughout the shooting process. This application effectively extracts useful information from the target facial video containing noise by using the weight value corresponding to each inference window, reducing the dependence on the quality of the target facial video and making non-contact blood pressure measurement more suitable for everyday real-world scenarios.

[0056] The beneficial effects of the embodiments of this application are as follows: Firstly, the photoplethysmography (PPG) signal of the target facial video is divided into PPG signals of multiple inference windows. The PPG signal of each inference window is input into the trained student model. The trained student model outputs the blood pressure estimate of each inference window. Based on the weight value corresponding to each inference window, the blood pressure estimate of each inference window, and the blood pressure estimation model, the blood pressure value of the target facial video is generated. By using the weight value corresponding to each inference window, the high-quality and high-confidence inference window dominates the blood pressure value generation process, effectively avoiding the problem of effective features being masked due to equal weight fusion in existing technologies, and improving the reliability of the blood pressure value of the target facial video. Secondly, by using the weight value corresponding to each inference window, the blood pressure value of the target facial video is no longer subject to drastic deviations due to signal quality fluctuations in individual windows, thus enhancing the stability of the blood pressure value of the target facial video.

[0057] Please see Figure 3 , Figure 3 The implementation flowchart of S202 provided in the embodiments of this application is described in detail below: S301, obtain the window length and sliding step from the configuration file; S302, based on the window length and sliding step size, divides the photoplethysmography signal of the original facial video into photoplethysmography signals of multiple training windows.

[0058] For example, if the total length of the photoplethysmography (PPG) signal of the original facial video is 30 seconds, the preset window length is 3 seconds, and the sliding step is 1 second, then starting from second 0, the PPG signals of 0 to 3 seconds, 1 to 4 seconds, and 2 to 5 seconds are captured sequentially until the signal ends, resulting in a total of 28 training window PPG signals.

[0059] In this embodiment, the photoplethysmography signal of the original facial video is divided into photoplethysmography signals of multiple training windows according to the window length and sliding step size. Without increasing the additional data acquisition cost, a large number of training samples are generated from the limited original facial video, which effectively improves the generalization ability of the student model.

[0060] For the blood pressure estimation method described in the above embodiments, please refer to [link / reference]. Figure 4 , Figure 4 This is a schematic block diagram of the blood pressure estimation device provided in the embodiments of this application. Figure 4 The blood pressure estimation device 400 shown can be applied to, for example... Figure 1 The application scenario diagram shows electronic devices. The following section uses electronic devices as an example to illustrate this. Figure 4 The blood pressure estimation device 400 shown will be described in detail. The blood pressure estimation device 400 may include an acquisition module 401, an input module 402, a first generation module 403, a second generation module 404, and a third generation module 405.

[0061] The acquisition module 401 is used to acquire the original facial video, perform feature extraction on the original facial video, and generate the photoplethysmography signal of the original facial video. The input module 402 is used to divide the photoplethysmography signal of the original facial video into photoplethysmography signals of multiple training windows, and input the photoplethysmography signal of each training window into the student model. The first generation module 403 is used to output the morphological feature score of each training window through the student model, generate the recognition score of each training window and the recognition model of the training stage, and generate the weight value corresponding to each training window based on the recognition score of each training window and the weight model of the training stage. The second generation module 404 is used to generate the total loss of the student model based on the weight value corresponding to each training window, the blood pressure estimate of each training window, the blood pressure reference value of each training window, the morphological distillation loss of the student model and the predefined total loss model, to train the student model with the goal of reducing the total loss, and to save the trained student model. The third generation module 405 is used to divide the photoplethysmography (PPG) signal of the target facial video into PPG signals of multiple inference windows, input the PPG signal of each inference window into the trained student model, output the blood pressure estimate of each inference window through the trained student model, and generate the blood pressure value of the target facial video based on the weight value corresponding to each inference window, the blood pressure estimate of each inference window and the blood pressure estimation model.

[0062] It should be noted that the various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. The same or similar parts between the various embodiments can be referred to each other.

[0063] The beneficial effects of the embodiments of this application are as follows: Firstly, the photoplethysmography (PPG) signal of the target facial video is divided into PPG signals of multiple inference windows. The PPG signal of each inference window is input into the trained student model. The trained student model outputs the blood pressure estimate of each inference window. Based on the weight value corresponding to each inference window, the blood pressure estimate of each inference window, and the blood pressure estimation model, the blood pressure value of the target facial video is generated. By using the weight value corresponding to each inference window, the high-quality and high-confidence inference window dominates the blood pressure value generation process, effectively avoiding the problem of effective features being masked due to equal weight fusion in existing technologies, and improving the reliability of the blood pressure value of the target facial video. Secondly, by using the weight value corresponding to each inference window, the blood pressure value of the target facial video is no longer subject to drastic deviations due to signal quality fluctuations in individual windows, thus enhancing the stability of the blood pressure value of the target facial video.

[0064] Please see Figure 5 , Figure 5 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application.

[0065] like Figure 5 As shown, Figure 5 The electronic device includes: at least one processor 20, a memory 21, and a computer program 22 stored in the memory 21 and executable on the at least one processor 20, wherein the processor 20 executes the computer program 22 to implement the steps in any of the above method embodiments.

[0066] The electronic device may include, but is not limited to, processor 20 and memory 21. Those skilled in the art will understand that... Figure 5This is merely an example of an electronic device and does not constitute a limitation on electronic devices. It may include more or fewer components than shown in the illustration, or combinations of certain components, or different components. For example, it may also include input / output devices, network access devices, etc.

[0067] The processor 20 is used to run a computer program 22 stored in the memory 21, and performs the following steps when executing the computer program 22: Acquire the original facial video, perform feature extraction on the original facial video, and generate the photoplethysmography signal of the original facial video. The photoplethysmography (PPG) signal of the original facial video is divided into PPG signals of multiple training windows, and the PPG signal of each training window is input into the student model. The student model outputs the morphological feature score of each training window. The morphological feature score of each training window and the recognition model of the training stage are used to generate the recognition score of each training window. Based on the recognition score of each training window and the weight model of the training stage, the weight value corresponding to each training window is generated. Based on the weight value corresponding to each training window, the blood pressure estimate of each training window, the blood pressure reference value of each training window, the morphological distillation loss of the student model, and the predefined total loss model, the total loss of the student model is generated. With the goal of reducing the total loss, the student model is trained and the trained student model is saved. The photoplethysmography (PPG) signal of the target facial video is divided into PPG signals of multiple inference windows. The PPG signal of each inference window is input into the trained student model. The trained student model outputs the blood pressure estimate of each inference window. Based on the weight value corresponding to each inference window, the blood pressure estimate of each inference window and the blood pressure estimation model, the blood pressure value of the target facial video is generated.

[0068] The processor 20 may be a Central Processing Unit (CPU), or it may be other general-purpose processors, digital signal processors, field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or any conventional processor.

[0069] In some embodiments, the memory 21 may be an internal storage unit of the electronic device, such as a hard disk or memory of the electronic device.

[0070] The above are merely preferred embodiments of this application and do not limit the patent scope of this application. Any equivalent structural or procedural transformations made using the content of this application's specification and drawings, or direct or indirect applications in other related technical fields, are similarly included within the patent protection scope of this application.

Claims

1. A blood pressure estimation method based on facial video, characterized in that, The blood pressure estimation method, applied to electronic devices, includes: Acquire the original facial video, perform feature extraction on the original facial video, and generate the photoplethysmography signal of the original facial video. The photoplethysmography (PPG) signal of the original facial video is divided into PPG signals of multiple training windows, and the PPG signal of each training window is input into the student model. The student model outputs the morphological feature score of each training window. The morphological feature score of each training window and the recognition model of the training stage are used to generate the recognition score of each training window. Based on the recognition score of each training window and the weight model of the training stage, the weight value corresponding to each training window is generated. Based on the weight value corresponding to each training window, the blood pressure estimate of each training window, the blood pressure reference value of each training window, the morphological distillation loss of the student model, and the predefined total loss model, the total loss of the student model is generated. With the goal of reducing the total loss, the student model is trained and the trained student model is saved. The photoplethysmography (PPG) signal of the target facial video is divided into PPG signals of multiple inference windows. The PPG signal of each inference window is input into the trained student model. The trained student model outputs the blood pressure estimate of each inference window. Based on the weight value corresponding to each inference window, the blood pressure estimate of each inference window and the blood pressure estimation model, the blood pressure value of the target facial video is generated.

2. The blood pressure estimation method according to claim 1, characterized in that, The student model outputs a morphological feature score for each training window. Using this score and the recognition rate model from the training phase, the recognition rate of each training window is generated. Finally, based on the recognition rate and the weight model from the training phase, a weight value is generated for each training window, including: The first branch of the student model outputs the label uncertainty of each training window, the second branch of the student model outputs the prediction uncertainty of each training window, the third branch of the student model outputs the blood pressure estimate of each training window, and the fourth branch of the student model outputs the morphological feature score of each training window. The recognition score of each training window is generated by the morphological feature score of each training window and the recognition model during the training phase. Based on the recognition score of each training window, the label uncertainty of each training window, the prediction uncertainty of each training window, and the weight model during the training phase, the weight value corresponding to each training window is generated.

3. The blood pressure estimation method according to claim 1, characterized in that, Based on the weight values ​​corresponding to each training window, the estimated blood pressure value for each training window, the reference blood pressure value for each training window, the morphological distillation loss of the student model, and a predefined total loss model, the total loss of the student model is generated. With the goal of reducing the total loss, the student model is trained, and the trained student model is saved, including: Obtain the annotation information corresponding to the original facial video. The annotation information includes all blood pressure measurements obtained by a cuff blood pressure monitor when the original facial video was captured. All blood pressure measurements are arranged in chronological order to generate a discrete data sequence with time on the horizontal axis and blood pressure on the vertical axis. The reference blood pressure value for each training window is obtained from the discrete data sequence by interpolation. Obtain the feature vectors output by the encoders of the teacher model and the student model. Based on the feature vectors output by the encoders of the teacher model and the student model, and a predefined morphological distillation loss model, generate the morphological distillation loss of the student model. Based on the weight values ​​corresponding to each training window, the blood pressure estimate of each training window, the blood pressure reference value of each training window, the morphological distillation loss of the student model, and a predefined total loss model, generate the total loss of the student model. With the goal of reducing the total loss, train the student model until the total loss is less than the preset loss, then stop training the student model and save the trained student model.

4. The blood pressure estimation method according to claim 1, characterized in that, The photoplethysmography (PPG) signal of the target facial video is divided into PPG signals for multiple inference windows. The PPG signal of each inference window is input into a trained student model. The trained student model outputs a blood pressure estimate for each inference window. Based on the weight values ​​corresponding to each inference window, the blood pressure estimate for each inference window, and the blood pressure estimation model, the blood pressure value of the target facial video is generated, including: Acquire a target facial video, perform feature extraction on the target facial video, generate an opto-pole permeable (OPP) signal of the target facial video, and divide the OPP signal of the target facial video into OPP signals of multiple inference windows. The photoplethysmography signal of each inference window is input into the trained student model. The label uncertainty of each inference window is output through the first branch of the trained student model, the prediction uncertainty of each inference window is output through the second branch of the trained student model, the blood pressure estimate of each inference window is output through the third branch of the trained student model, and the morphological feature score of each inference window is output through the fourth branch of the trained student model. The recognition score of each inference window is generated by the morphological feature score of each inference window and the recognition model of the inference stage. Based on the recognition score of each inference window, the label uncertainty of each inference window, the prediction uncertainty of each inference window, and the weight model of the inference stage, the weight value of each inference window is generated. Based on the weight value of each inference window, the blood pressure estimate of each inference window, and the blood pressure estimation model, the blood pressure value of the target facial video is generated.

5. The blood pressure estimation method according to claim 1, characterized in that, The recognition model during the training phase is defined as follows: ; This represents the discriminative power of the w-th training window; the higher the discriminative power of the w-th training window, the more complete the comprehensive features of the w-th training window; the lower the discriminative power of the w-th training window, the less complete the comprehensive features of the w-th training window. w represents the sequence number of the training window; Represents the Sigmoid function; Represents the weight vector; This represents the video signal quality characteristics of the w-th training window, used to reflect the video signal quality of the w-th training window; The larger the value, the clearer the video image in the w-th training window; The smaller the value, the blurrier the video image in the w-th training window; This represents the signal quality characteristics of the w-th training window, used to reflect the photoplethysmography signal quality of the w-th training window; The larger the value, the higher the signal-to-noise ratio of the photoplethysmography signal in the w-th training window; The smaller the value, the lower the signal-to-noise ratio of the photoplethysmography signal in the w-th training window; This represents the morphological feature score of the w-th training window, used to reflect whether the pulse wave features of the w-th training window are complete; The larger the value, the more complete the pulse wave features of the w-th training window; The smaller the value, the less the pulse wave features are missing in the w-th training window; This represents the consistency feature of the w-th training window, used to reflect the degree of consistency between the photoplethysmography signal and the reference signal in the w-th training window. The larger the value, the more consistent the photoplethysmography signal of the w-th training window is with the reference signal; The smaller the value, the less consistent the photoplethysmography signal of the w-th training window is with the reference signal; This represents the uncertainty of the w-th training window; Indicates the penalty weight; The weight model during the training phase is defined as follows: ; This represents the weight value corresponding to the w-th training window; This represents the label uncertainty of the w-th training window; the larger the label uncertainty of the w-th training window, the less reliable the blood pressure reference value of the w-th training window is; the smaller the label uncertainty of the w-th training window, the more reliable the blood pressure reference value of the w-th training window is. This represents the prediction uncertainty for the w-th training window; The larger the value, the lower the confidence level of the student model in the blood pressure estimation results for the w-th training window. The smaller the value, the higher the confidence level of the student model in the blood pressure estimation results for the w-th training window; This represents a preset positive integer, used to avoid the denominator being zero.

6. The blood pressure estimation method according to claim 1, characterized in that, The speciation distillation loss model is defined as follows: ; This represents the morphological distillation loss of the student model; w represents the sequence number of the training window; This represents the recognition score of the w-th training window; This represents the set of all training windows; This represents the photoplethysmography signal input to the student model in the w-th training window. The encoder representing the student model; The mapping layer representing the student model; This represents the feature vector output by the encoder of the student model; This represents the projected features obtained after the feature vector output by the encoder of the student model is projected through the mapping layer. This represents the photoplethysmography signal input to the teacher model at the w-th training window; The encoder representing the teacher model; This represents the feature vector output by the encoder of the teacher model; Represents the square of the L2 norm; The total loss model is defined as follows: ; L represents the total training loss of the student model; This represents the weight value corresponding to the w-th training window; This represents the blood pressure estimate for the w-th training window; This represents the blood pressure reference value for the w-th training window; This represents the physiological constraint loss of the student model; This represents the morphological distillation loss of the student model; This represents the classification loss of the student model; , , , represent the first weight, the second weight, and the third weight, respectively.

7. The blood pressure estimation method according to claim 1, characterized in that, The blood pressure estimation model is defined as follows: ; This indicates the estimated blood pressure value for the current video. This represents the set of inference windows, which consists of all inference windows obtained from the segmentation of the current video. The weight value corresponding to the z-th inference window indicates that the greater the weight of the blood pressure estimate of the z-th inference window in the fusion process, the greater its impact on the blood pressure estimate of the reference facial video; conversely, the smaller the weight of the blood pressure estimate of the z-th inference window in the fusion process, the smaller its impact on the blood pressure estimate of the reference facial video. This represents the blood pressure estimate for the z-th inference window.

8. A blood pressure estimation device based on facial video, characterized in that, Applied to electronic devices, including: The acquisition module is used to acquire the original facial video, perform feature extraction on the original facial video, and generate the photoplethysmography signal of the original facial video. The input module is used to divide the photoplethysmography (PPG) signal of the original facial video into PPG signals of multiple training windows, and input the PPG signal of each training window into the student model. The first generation module is used to output the morphological feature score of each training window through the student model, generate the recognition score of each training window through the morphological feature score of each training window and the recognition model of the training stage, and generate the weight value corresponding to each training window based on the recognition score of each training window and the weight model of the training stage. The second generation module is used to generate the total loss of the student model based on the weight value corresponding to each training window, the blood pressure estimate of each training window, the blood pressure reference value of each training window, the morphological distillation loss of the student model, and the predefined total loss model. With the goal of reducing the total loss, the student model is trained and the trained student model is saved. The third generation module is used to divide the photoplethysmography (PPG) signal of the target facial video into PPG signals of multiple inference windows, input the PPG signal of each inference window into the trained student model, output the blood pressure estimate of each inference window through the trained student model, and generate the blood pressure value of the target facial video based on the weight value corresponding to each inference window, the blood pressure estimate of each inference window and the blood pressure estimation model.

9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the blood pressure estimation method as described in any one of claims 1 to 7.

10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the blood pressure estimation method as described in any one of claims 1 to 7.