Video coding method for optimizing physiological signal retention

By dividing face and non-face areas in video encoding, using convolutional neural network and adaptive code rate control, dynamically adjusting coding parameters, the balance problem of video compression efficiency and physiological signal retention is solved, and more efficient physiological signal retention and video compression are achieved.

CN120264000APending Publication Date: 2025-07-04NANJING UNIV OF SCI & TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510535607.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-27
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

Existing video encoding technologies are difficult to find a balance between compression efficiency and physiological signal retention, resulting in problems such as loss of physiological signals or increased storage costs.

Method used

By dividing face and non-face areas, using convolutional neural networks for detection, combining adaptive bit rate control and polynomial fitting model, encoding parameters are dynamically adjusted to optimize physiological signal retention, and adaptive bit rate control encoder is used to adjust the bit rate allocation of face and background, establish an optimal encoding parameter model, and optimize the video compression process.

Benefits of technology

While ensuring video compression efficiency, it significantly improves the retention quality of physiological signals and reduces video file size and storage costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120264000A_ABST
    Figure CN120264000A_ABST
Patent Text Reader

Abstract

The invention discloses a video coding method for optimizing physiological signal retention. The method comprises the following steps: dividing each frame of an input face video into 64 * 64 CTU blocks, and detecting and marking face and non-face areas by using a CNN face detection model; for the first 50 frames, according to quantization parameters, a compressed video bit rate, physiological signal quality and video subjective quality, an optimal coding parameter set of balance is found through an exhaustion method, and a coding parameter model is established through polynomial fitting; face and background bit rate distribution is adjusted according to video content and physiological signal retention requirements by utilizing an automatic code rate control function of an encoder, face blocks are compressed after Lagrange parameters are calculated and updated according to an encoding parameter model, and background encoding parameters are set for non-face blocks; and repeatedly processing the CTU block, coding the test sequence, and outputting a compressed video which retains more physiological signals.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to video coding technology, in particular to an optimized video coding method for physiological signal retention, which is widely applied in fields such as medical monitoring, intelligent video analysis, and psychophysiological research. Background Art

[0002] Heart rate, as an important physiological indicator reflecting heart health, can help detect abnormal heart activities at an early stage through the analysis of heart rate changes. Therefore, continuous and accurate monitoring of heart rate has important clinical value. Non-contact heart rate monitoring technology, especially based on remote Photoplethysmography (rPPG) technology, has become an effective means for heart rate detection. rPPG analyzes the pulse wave signal captured from facial videos and can achieve non-contact real-time heart rate monitoring. However, due to the large amount of video data, especially in remote monitoring and video storage applications, video compression has become an inevitable processing step.

[0003] In existing video coding technologies, when compressing a face video containing physiological signals (such as rPPG, etc.), it is often difficult to balance the video compression efficiency and the retention of physiological signals. Traditional coding methods may cause the loss or distortion of physiological signals due to excessive compression, affecting subsequent analysis and utilization of physiological signals; and reducing the compression ratio to retain physiological signals will make the video file too large, increasing storage and transmission costs. Therefore, a video coding method that can maximize the retention of physiological signals while ensuring a certain compression efficiency is needed. Summary of the Invention

[0004] The purpose of the present invention is to provide a video coding method for optimizing physiological signal retention, which can perform targeted processing according to the video content and physiological signal requirements during the video compression process, effectively retain more physiological signals, and at the same time optimize the video compression efficiency. The method includes the following steps:

[0005] S01, input a face video, divide each frame into coding tree unit (CTU) blocks of 64×64 size, and use a convolutional neural network (CNN) model to perform face detection and mark the face region and non-face region therein;

[0006] S02, sequentially input the marked face video frames into the encoder. If it is the first 50 frames of the input video, adjust the coding parameter QP of the face region according to the quantization parameter, the quality of the extracted physiological signal, and the subjective quality of the compressed video for different QPs raw Adjust the coding parameter QP of the face region face Use the exhaustive method to obtain an optimal coding parameter set QP that balances the physiological signal quality and the subjective quality of the video optional, use polynomial fitting to fit the original coding parameters and the coding parameters of the adjusted face part to establish an optimal coding parameter model QP face = F(QP raw );

[0007] S03. Use the encoder's automatic bitrate control function to dynamically adjust the bitrate allocation between the face and the background according to the video content and the requirement of retaining physiological signals. Select CTU blocks in sequence. If it is a face block, use the coding parameter model QP face = F(QP raw ) to calculate QP face and update the Lagrange parameter λ face ; if it is not a face block, set the background coding parameter QP bg = QP raw ;

[0008] S04. Repeat the steps of selecting and compressing CTU blocks until all CTU blocks are processed, encode the test sequence, and output a compressed video that retains more physiological signals; Description of the Drawings

[0009] Figure 1 : Flowchart of a video coding method for optimizing the retention of physiological signals provided by the present invention

[0010] Figure 2 : Schematic diagram of Lagrange parameter optimization considering physiological signals in adaptive bitrate control coding

[0011] Figure 3 : Heat map of mean square error of coding tree unit blocks under original bitrate control

[0012] Figure 4 : Heat map of mean square error of coding tree unit blocks considering physiological signals

[0013] Figure 5 : Result comparison between the original compression method and the compression method for retaining physiological signals Detailed Embodiment

[0014] The present invention proposes a video coding method for optimizing the retention of physiological signals. By considering physiological signals and combining the relationship between video content (especially the face area) and compression parameters, the fidelity of physiological signals in the video coding process is optimized while maintaining a high video compression efficiency. The specific implementation steps are as follows:

[0015] Divide the face test video frames into CTU blocks of 64×64, and make division marks on the video background part and the face part area. Use a convolutional neural network (CNN) model to perform face detection, mark the face area and non-face area among them. Let the image be I(x,y), where (x,y) represents the spatial coordinates of the image. The image is divided into CTU blocks of size M×N, and the face area A is marked face and the non-face area A bg That is:

[0016] I(x,y)∈A face ,x∈[0,M],y∈[0,N]

[0017]

[0018] Assume that the total bitrate of the video is R C , and the bitrate of the face area is R face , and the bitrate of the non-face area is R bg , then:

[0019] R total =R face +R bg

[0020] Input the marked face video frames into the encoder in sequence. If they are the first 50 frames of the input video, under the condition of ensuring a constant R total , adjust R face and R bg , and maximize R face on the premise of ensuring the relative stability of the subjective quality evaluation index VMAF (Video Multi-Method Assessment Fusion) of the compressed video. Test different configurations during face video compression, use the exhaustive method to achieve the optimal coding parameter combination that balances the physiological signal quality and the video subjective quality, and then use polynomial fitting for the original coding parameters and the adjusted coding parameters

[0021] Assume that the original quantization parameter is QP raw , set the quantization parameter of the background area as QP bg =QP raw , and use the binomial fitting method to fit QP optional as a piecewise function of QP raw to obtain QP face , where a and b are thresholds determined through experiments (a < b), and f1(x), f2(x), and f3(x) are all cubic polynomial functions, so as to establish the coding parameter model QP face =F(QP raw ), and obtain the quantization parameter QP of the face areaface Calculate according to the following formula:

[0022]

[0023] If there are subsequent frames of the input video, it indicates that an encoding parameter model for preserving physiological signals has been established, and subsequent processing is then performed according to the encoding parameter model for preserving physiological signals.

[0024] Enable CTU layer bitrate control. Initialize the quantization parameter QP value of the current CTU in the original encoder to the QP value of the current slice layer, and then calculate according to the size of the set target bitrate, which can be simplified to the following formula:

[0025] QP = g(f(bpp, MAD) / λ)

[0026] In order to preserve physiological signals, the rPPG signal should be considered when encoding the CTU of the face part. When calculating the quantization parameter of the current face part, according to the physiological signal preservation parameter model in the experiment in S02, the formula after introducing the rPPG factor is as follows:

[0027] QP face = g((QP raw + f(bpp, MAD)) / λ face )

[0028] Compress the test video according to the adjusted QP to obtain the Lagrange parameter that can preserve more physiological signals in video encoding, and encode the test sequence, specifically as follows:

[0029] According to the original formula QP raw = 4.2005lnλ raw + 13.7122, solve inversely to obtain According to the relationship between QP raw and QP face to obtain QP face , and then calculate λ according to the relationship between QP and λ, and update λ face , to obtain the relationship between the original QP and λ face as

[0030] As Figure 2 shown, after considering physiological signals in adaptive bitrate control encoding, the Lagrange parameter rises relatively slowly with the increase of the original QP compared to the background. According to the target bitrate allocation method during encoding:

[0031]

[0032] Since the constant α and β inherit the α and β values of the current frame, and the mean absolute difference per pixel (MADperpixel) remains unchanged when the total coding cost and the total number of pixels in the CTU remain the same, it can be seen that the bits per pixel (bpp) of the CTU blocks in the face part increases relative to the bpp of the original face part, thus retaining more physiological signals.

[0033] The Lagrange parameter λ of the face part face , and the Lagrange parameter for the background region is λ raw . For the CTU blocks in the region containing physiological signals, the updated λ is enabled during adaptive bitrate control coding face . For the background CTUs, λ is calculated using the original QP raw . Repeat the steps of selecting and compressing CTU blocks until all CTU blocks are processed, encode the test sequence, and output the compressed video that retains more physiological signals.

[0034] Assume that the Y component of the YUV image block of the original video at position (i, j) is Y(i, j), and the Y component value of the corresponding block in the compressed video at the same position is . The block size is m×n, and the mean squared error (MSE) of the Y component is calculated as:

[0035]

[0036] Output the optimized compressed video, such as Figure 3 and Figure 4 shown. Through the visual display of the MSE of the compressed video and the uncompressed video, it can be seen that, compared with the original rate control algorithm when the video bitrate is comparable, the MSE of the face part decreases after considering physiological signals, and the compression loss of physiological signals is lower than that of the original compression.

[0037] Figure 5 By comparing the video bitrate, the signal-to-noise ratio (SNR) of the rPPG signal extracted from the encoded video, and the video subjective quality VMAF score between the original compression and the compression method that retains physiological signals, it can be seen that, when the compressed video bitrate and the subjective quality score of the compressed video are similar between the video coding method that retains physiological signals and the original compression method, the signal-to-noise ratio of the physiological signals is higher, and the compressed video retains more physiological signals.

Claims

1. A video coding method for optimizing the preservation of physiological signals, characterized in that, Including the following steps: S01, input a face video, divide each frame into coding tree unit (CTU) blocks of 64×64 size, use a convolutional neural network (CNN) model to perform face detection, and mark the face regions and non-face regions therein; S02, sequentially input the labeled face video frames in the encoder. If it is the first 50 frames of the input video, according to the quantization parameter, the physiological signal quality extracted, and the subjective quality of the compressed video during the video encoding process for different QPs raw Adjust the encoding parameter QP of the face region face , and use the exhaustive method to achieve the optimal encoding parameter set QP that balances the physiological signal quality and the subjective quality of the video optional , use polynomial fitting for the original encoding parameters and the encoding parameters of the adjusted face part to establish the optimal encoding parameter model QP face = F(QP raw ); S03, using the encoder's automatic bitrate control function, dynamically adjusts the bitrate allocation between the face and the background according to the video content and the requirement of retaining physiological signals. Sequentially select CTU blocks. If it is a face block, use the encoding parameter model QP face = F(QP raw ) to calculate QP face and update the Lagrangian parameter λ face ; if it is not a face block, set the background encoding parameter QP bg = QP raw ; S04, repeat the steps of selecting and compressing CTU blocks until all CTU blocks are processed, encode the test sequence, and output a compressed video that retains more physiological signals.

2. The video coding method for optimizing physiological signal retention according to claim 1, characterized in that The S01 includes: dividing each input frame into CTU blocks of 64×64, using a convolutional neural network (CNN) model to perform face detection, and statistically calculating the percentage of the face part in each CTU block, thereby marking the CTU blocks as the face part and the background part.

3. The face video compression method for retaining physiological signals according to claim 1, wherein The encoding parameter model in S02 is specifically as follows: Input the marked face video frames into the encoder in sequence. If it is the first 50 frames of the input video, use the exhaustive method to analyze the relationship between the quantization parameter under compression, the quality of the physiological signals extracted from the video, the subjective quality of the compressed video, and the video bit rate to obtain the optimal QP optional Set, and fit the experimental results into an encoding parameter model based on the reference region; Assume that the original quantization parameter is QP raw , set the quantization parameter of the background area to QP bg = QP raw , fit QP using the binomial fitting method according to the experimental results optional as a piecewise function of QP raw to obtain QP face , where a and b are thresholds determined by experiments (a < b), and f1(x), f2(x), and f3(x) are all cubic polynomial functions, thereby establishing the coding parameter model QP face = F(QP raw ), and obtain the quantization parameter QP of the face area face Calculate according to the following formula: If there are subsequent frames of the input video, it indicates that an encoding parameter model for retaining physiological signals has been established, and subsequent processing is performed according to the encoding parameter model for retaining physiological signals.

4. The method for compressing a face video for retaining a physiological signal according to claim 1, wherein: In the S03, the adjustment of the automatic bitrate control includes: Enable the encoder's automatic bitrate control function and use QP for the background area raw , and use the encoding parameter model QP in S02 for the face part with physiological signals face = F(QP raw ) to recalculate QP face , and further calculate the corresponding Lagrangian encoding parameter λ face , to optimize the retention effect of physiological signals; The Lagrangian parameter for the face part is λ face , and the Lagrangian parameter for the background region is λ raw . For the CTU blocks containing physiological signal regions, the updated λ is enabled during adaptive bitrate control coding face . For background CTUs, QP is used raw to calculate the obtained λ raw . Repeat the steps of selecting and compressing CTU blocks until all CTU blocks are processed, encode the test sequence, and output the compressed video that retains more physiological signals 5. The video encoding method for optimizing physiological signal retention according to claim 1, characterized in that, The physiological signals include but are not limited to heart rate, respiratory rate, and blood oxygen saturation.

6. The video coding method for optimizing physiological signal retention according to claim 1, characterized in that The method is applicable to scenarios of extracting physiological signals based on videos in fields such as telemedicine, health monitoring, and affective computing.