Indoor monitoring video compression method based on characteristic rate distortion coding optimization
By using generative adversarial networks to generate the prospective human body in indoor surveillance video compression, and constructing semantic feature encoding code rate model and video quality model, the problem of high code rate consumption and lack of system modeling in the existing technology is solved, and efficient indoor surveillance video compression is achieved, surpassing the performance of VVC encoding.
Patent Information
- Application Number
- CN202510151847.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-11
- Publication Date
- 2025-06-03
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The prior art has problems such as high bit rate consumption and lack of system modeling in indoor surveillance video compression, resulting in unsatisfactory compression effect in bandwidth-constrained scenarios.
A indoor monitoring compression method based on feature rate distortion coding optimization is proposed. By generating an adversarial network to generate a foreground human body in monitoring video, a semantic feature code rate model and video quality model are constructed, and efficient code rate allocation is achieved under rate distortion constraints.
Improve video compression efficiency, achieve compact indoor surveillance video compression, and objective indicators exceed VVC encoding.
Smart Images

Figure CN120088702A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of indoor surveillance video compression, and particularly to an indoor surveillance compression method based on feature rate - distortion coding optimization. Background Art
[0002] With the increasing demand for indoor security, indoor surveillance cameras have been rapidly deployed in recent years. The explosive growth of video surveillance data brings more severe challenges to data storage and transmission. In response to this challenge, on the one hand, the industry continues to develop traditional video compression algorithms based on pixel correlation, and on the other hand, it begins to gradually explore semantic coding compression algorithms based on generative models.
[0003] Traditional video compression algorithms take pixels and image blocks as the basic processing units. Aiming at the spatial redundancy, temporal redundancy, perceptual redundancy, etc. in videos, guided by information theory, they use a predictive differential hybrid coding framework to achieve effective video compression. In recent years, due to the strong representation and learning capabilities of deep learning, it has been widely applied to video coding. For example, the end - to - end deep video compression (DVC) framework proposed by Lu et al. replaces components such as motion estimation and motion compensation in the traditional architecture with a deep neural network, improving the coding performance. However, this is not specifically designed for surveillance videos and performs mediocrely in surveillance video scenarios. Wu et al. compress the foreground and background of surveillance videos separately, use optical flow for foreground image motion estimation, and share background information between adjacent frames, greatly improving the compression ratio. However, this method does not have a compact representation of the foreground human body and has an unsatisfactory compression effect in bandwidth - limited scenarios. With the development of generative technology, generative videos have inspired us. The generative model only needs highly abstract and compact features to guide and can generate the entire human body image and video. For example, Xia et al. extract the sparse motion patterns of the human body and generate the human appearance of the frame to be encoded through a generative network, achieving better coding quality than HEVC at low bitrates. However, Xia et al. use a fixed number of key - point features and do not consider the problem of generative video rate - distortion. Generative videos rely on semantic feature maps or key - point representations, and these features will occupy a certain bitrate during encoding. Existing research lacks a systematic modeling of the relationship between generative video quality and bitrate consumption. Summary of the Invention
[0004] In order to solve the above - mentioned technical problems existing in the prior art, the present invention proposes an indoor surveillance compression method based on feature rate - distortion coding optimization. This method relies on a generative adversarial network to generate the foreground human body in surveillance videos, constructs a semantic feature coding bitrate model and a generative video quality model, and realizes efficient bitrate allocation under rate - distortion constraints to further improve video compression efficiency. The specific technical solutions are as follows:
[0005] An indoor surveillance video compression method based on feature rate - distortion coding optimization, including background adaptive update, foreground feature rate - distortion coding optimization, and coding reconstruction steps. First, analyze the surveillance video, and use the background extraction algorithm and the person segmentation algorithm to extract the background and the human body respectively; use the feature extraction algorithm to convert the human body into a compact key - point feature; input the compact key - point feature and the reference frame into the generative adversarial network to generate a non - reference frame, and fuse the generated non - reference frame with the background to obtain the reconstructed video.
[0006] Further, the background adaptive update step is specifically as follows:
[0007] The background update takes the time - series frames as input and dynamically adjusts the update in combination with the event - triggering mechanism. The events are divided into periodic events and conditional events. The periodic events trigger the extraction of a new background at a fixed time interval, and the conditional events are triggered based on the detection result or threshold change, and update the reference background with the help of the residual information.
[0008] Further, the conditional events include the SSIM threshold event, the residual fluctuation event, and the occlusion release event.
[0009] Further, the SSIM threshold event is triggered for update when the SSIM between the current background and the reference background is less than the threshold.
[0010] Further, the residual fluctuation event is to update the background when the fluctuation of the residual video frame exceeds the set threshold;
[0011] Let F t ={f 1 , f 2 ,..., f t-1 , f t ,...} be the original video - frame sequence, F t ={f 1 , f 2 ,..., f t-1 , f t ,...} be the encoded video frames, R t ={r 1 , r 2 ,..., r t-1 , r t ,...} be the residual video - frame sequence obtained by subtracting the original frame and the reconstructed video frame, and B t ={b 1 , b 2 ,..., b t-1 , b t ,...} be the video background - frame sequence corresponding to each frame; the residual formula is as follows:
[0012] rt-1 (x, y) = f t-1 -f t-1 (1)
[0013] b t = b t-1 + δr t-1 (x, y)(1 - m t ) (2)
[0014] δ is the residual compensation coefficient, which plays a role in adjusting the size of the residual and is determined by the rate - distortion model.
[0015] Furthermore, the occlusion removal event is to update the background when the occluder in the video is removed; specifically, first calculate the mask change value Δm t (x, y), Δm t (x, y) = 1 indicates that the pixel (x, y) changes from the occluded area to the released area. If the proportion of the released area exceeds a certain threshold θ mask , it is considered that the occluder has been removed and the background update is triggered; let the number of pixels in the released area be S release , and the total number of pixels be S total , then the trigger condition is:[[]]
[0016] Δm t = m t - m t-1 (3)
[0017]
[0018] θ mask is an empirical value, usually set to 0.1; after detecting the removal of the occluder, only the residuals in the released area are compensated to update the background; the update formula is:[[]]
[0019] b t = b t-1 + δr t-1 (x, y)Δm t (5)
[0020] Furthermore, the steps for optimizing the foreground feature rate - distortion coding are as follows:[[]]
[0021] Compact feature extraction: Use the target extraction algorithm to extract the foreground human body sequence from the original video, and select the frontal video frame of the person from the human body sequence as the key frame I k ; Use the mean absolute difference (MAD) of the grayscale image to measure the texture complexity T C of the human body sequence frames, and use the joint point extraction algorithm to extract the joint points of the human body sequence and form the human skeleton sequence;
[0022] Feature rate - distortion modeling: By means of statistical experiments, plot the mapping relationships of the bit rate R, video quality PSNR with the number of key points K and texture complexity T C And the quantization parameter; Based on the experimental results, construct the feature rate - distortion models as shown in formulas (6), (7) and (8), and determine the model parameters through experimental fitting;
[0023] R = F(QP, TC, K)=(a 1 K + b 1 )ln(TC)+c 1 K + d 1 (6)
[0024]
[0025] Optimal selection of encoding parameters: Based on the constructed feature rate - distortion model, solve the optimal encoding parameter configuration under the constraint of minimizing the rate - distortion cost. This process can be expressed as:
[0026]
[0027] Among them, [18, 51] is the value range of the residual quantization parameter Qp, [8, 18] is the value range of the number of joint points K. Convert PSNR to MSE, use MSE to directly measure the distortion D, and realize the generation of the foreground encoding with the best rate - distortion performance through the optimal encoding values of the foreground residual quantization parameter and the foreground skeleton joint points.
[0028] Furthermore, for each frame of video, extract the texture complexity T of the foreground C ; Through the rate - distortion model, input the texture complexity T C and the target bit rate R tar , and obtain the optimal residual quantization parameter Qp and the number of joint points K.
[0029] Furthermore, the feature rate - distortion model predicts the distortion D pred , the optimal residual quantization parameter Qp and the number of joint points K, and after encoding, obtain a set of actual R actual and D actual ; During the encoding process, the parameters will be fine - tuned according to the difference between the actual bit rate and the target bit rate and the difference between the actual distortion and the target distortion. The specific fine - tuning formula is as follows:
[0030]
[0031] After encoding, update the R - D model parameters to make the model more accurate.
[0032] Furthermore, the specific encoding and reconstruction steps are as follows:
[0033] At the decoding end, the generation network takes the characteristics of entropy decoding as the network input, predicts and generates the foreground video frame, and realizes the decoding and reconstruction of the video by fusing the generated foreground video frame with the updated background frame.
[0034] Based on the characteristic of stable background in indoor surveillance videos, where the video background is shared among multiple frames, the present invention provides an adaptive update strategy aimed at reducing the coding bitrate consumption of the video background by adaptively and dynamically adjusting the background. To optimize the rate-distortion coding of foreground features, a rate and distortion model based on foreground feature analysis is proposed. Under the rate-distortion constraint, the optimal coding parameters are determined to achieve efficient foreground compression and reconstruction, enabling the compact compression of indoor surveillance videos with objective metrics exceeding those of VVC coding. BRIEF DESCRIPTION OF THE DRAWINGS
[0035] Figure 1 is a schematic diagram of the framework of the indoor surveillance video compression method based on feature rate-distortion coding optimization of the present invention;
[0036] Figure 2 is the video decoding flowchart;
[0037] Figure 3 is the background adaptive update flowchart;
[0038] Figure 4 Bitrate and PSNR graph;
[0039] Figure 5 Schematic diagram of the bitrate allocation model. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0040] The present invention will be further described below with reference to the accompanying drawings.
[0041] As Figure 1 and Figure 2 shown, the indoor surveillance video compression method based on feature rate-distortion coding optimization of the present invention is based on the low information density characteristic of indoor surveillance videos, and includes background adaptive update, foreground feature rate-distortion coding optimization, coding reconstruction, and image denoising. First, the surveillance video is analyzed, and the background and the human body are extracted respectively using the background extraction algorithm and the person segmentation algorithm. The human body is transformed into compact key-point features using the feature extraction algorithm. The compact key-point features and the reference frame are input into the generative adversarial network to generate non-reference frames, and the generated non-reference frames are fused with the background to obtain the reconstructed video. Since there are some details missing in the reconstructed video, the encoder additionally transmits the residual between the reconstructed video and the original video for optimization. This residual includes the background and human body residuals, and the two are encoded in parallel.
[0042] Background Adaptive Update
[0043] Since stable backgrounds are generally present in indoor surveillance videos in a short period, background sharing of video frames can be achieved within a certain time through background extraction. Considering the volatility of the surveillance video background, a static background image is difficult to replace all video backgrounds. Therefore, dynamic background updates are required. The background update takes time-series frames as input and dynamically adjusts the update in combination with an event-triggering mechanism. Events are divided into periodic events and conditional events. The former triggers the extraction of a new background at fixed time intervals, and the latter is triggered based on detection results or threshold changes and updates the reference background with the help of residual information.
[0044] Event definition:
[0045] Periodic event (time-driven): Every fixed time T, trigger the re-extraction of the background.
[0046] Conditional event (state-driven): Trigger the background update according to the following conditions:
[0047] SSIM threshold event: Trigger the update when the SSIM between the current background and the reference background is less than the threshold.
[0048] Residual fluctuation event: Update the background when the fluctuation of the residual video frame exceeds the set threshold.
[0049] Occlusion removal event: Update the background when the occluder in the video is removed.
[0050] Let F t ={f 1 ,f 2 ,...,f t-1 ,f t ,...} be the original video frame sequence, F t ={f 1 ,f 2 ,...,f t-1 ,f t ,...} be the encoded video frames, R t ={r 1 ,r 2 ,...,r t-1 ,r t ,...} be the residual video frame sequence obtained by subtracting the original frame and the reconstructed video frame, and B t ={b 1 ,b 2 ,...,b t-1 ,b t ,...} be the video background frame sequence corresponding to each frame. As Figure 3 shown, the present invention extracts b c (b c ∈B t ) close to the first frame as the initial background. The background frame b 1and the generated human body image p 1 Reconstruction gives f 1 , and the residual is obtained from f 1 and f 1 by subtraction. The residual contains a large amount of residual information of the human body. The residual of the human body area should not be updated to the background, and only the background without the human body part is updated.
[0051] Over time, a new background needs to be re-extracted every time interval T. The extracted background is the reference background, which is a periodic event. Calculate the SSIM value between the current background and the reference background. When it is less than the threshold σ (usually 0.975), the current background is updated; otherwise, it is not updated. This is the SSIM threshold event. At the same time, calculate the size of the background residual. When it exceeds the threshold, the dynamic background is updated using the residual. This is the residual fluctuation event. Both rely on the residual for update, and the formula is as follows:
[0052] r t-1 (x,y) = f t-1 -f t-1 (1)
[0053] b t = b t-1 +δr t-1 (x,y)(1 - m t ) (2)
[0054] δ is the residual compensation coefficient, which plays a role in adjusting the size of the residual and is determined by the rate-distortion model.
[0055] The occlusion removal event is relatively complex. First, the mask change needs to be detected, and then the trigger condition is set. Specifically, first calculate the mask change value Δm t (x,y) between frame t - 1 and frame t. Δm t (x,y) = 1 indicates that the pixel (x,y) changes from the occluded area to the background (released area). To avoid the influence of noise, it is necessary to count the change area of Δm t (x,y). If the proportion of the released area exceeds a certain threshold θ mask , it is considered that the occluder has been removed and the background update is triggered. Let the number of pixels in the released area be S release , and the total number of pixels be S total , then the trigger condition is:
[0056] Δm t = m t -m t-1 (3)
[0057]
[0058] θ maskIt is an empirical value, usually set to 0.1 (i.e., triggered when the release area exceeds 10% of the total pixels). After detecting the removal of the occluder, only the residuals in the release area are compensated to update the background. The update formula is:
[0059] b t =b t-1 +δr t-1 (x,y)Δm t (5)
[0060] Foreground Feature Rate-Distortion Coding Optimization
[0061] The present invention generates a video sequence by using the extracted key frames and human skeleton sequences and adopting a generative adversarial network architecture. To ensure the quality of the reconstructed foreground, the encoder also encodes the residual information between the original foreground and the generated foreground. Therefore, the key features determining the quality of the reconstructed foreground include the number of skeleton joints, the quantization parameter Qp of the foreground residual, and the texture complexity of the foreground content.
[0062] Compact Feature Extraction: The present invention uses an object extraction algorithm to extract the foreground human sequence from the original video, and selects the frontal video frame of the person from the human sequence as the key frame I k . The average error MAD of the grayscale image is used to measure the texture complexity T of the human sequence frames C , and a joint point extraction algorithm is used to extract the joint points of the human sequence and form a human skeleton sequence.
[0063] Feature Rate-Distortion Modeling: To construct the rate-distortion model of the foreground video frame, the present invention draws through statistical experiments Figure 4 the mapping relationships of the shown bit rate R, video quality PSNR with the number of key points K, texture complexity T C and quantization parameter. When the number of joint points is fixed, the bit rate and the texture complexity approximately follow a logarithmic function relationship, and the video quality and the texture complexity approximately follow a reciprocal function relationship. Based on the experimental results, the feature rate-distortion models as shown in formulas (6), (7) and (8) are constructed, and the model parameters are determined by experimental fitting.
[0064] R = F(QP, TC, K)=(a 1 K + b 1 )ln(TC)+c 1 K + d 1 (6)
[0065]
[0066] Optimal Selection of Coding Parameters: Based on the constructed bit rate and distortion models, the optimal coding parameter configuration is solved under the constraint of minimizing the rate-distortion cost. This process can be expressed as:
[0067]
[0068] Among them, [18, 51] is the value range of the residual quantization parameter, and [8, 18] is the value range of the number of joint points. Convert PSNR to MSE, and use MSE to directly measure the distortion D. By optimizing the foreground residual quantization parameter and the optimal coding value of the foreground skeleton joint points, foreground coding generation with the best rate-distortion performance is achieved.
[0069] For each frame of video, extract the texture complexity T of the foreground C . Through the rate-distortion model, input the texture complexity T C and the target bitrate R tar , the model will give the best residual quantization parameter Qp and the number of joint points K. If the texture complexity is high and the target bitrate is low, predict a relatively low QP value and a moderate K value according to the prior model to maintain a good quality as much as possible while ensuring the bitrate limit; if the target bitrate is high, predict a relatively high QP value and a high K value according to the prior model.
[0070] The model predicts the distortion D pred , the best residual quantization parameter Qp and the number of joint points K, and obtain a set of actual R actual and D actual through encoding. During the encoding process, the parameters will be fine-tuned according to the difference between the actual bitrate and the target bitrate and the difference between the actual distortion and the target distortion to optimize the encoding effect and improve the encoding accuracy and efficiency. The specific fine-tuning formula is as follows:
[0071]
[0072] If the actual bitrate exceeds the target bitrate, increase Qp to improve the compression ratio and reduce the bitrate; if the actual bitrate is lower than the target bitrate, decrease Qp to improve the video quality. If the actual distortion exceeds the target distortion, increase the number of joint points K to improve the encoding accuracy and details; if the distortion is small and the actual bitrate is too high, reduce the number of joint points K to reduce the bitrate. After encoding, update the R-D model parameters to make the model more accurate.
[0073] Video reconstruction: By encoding the key frames, the number of joint points and their coordinates, and the foreground residual information, the encoder realizes the compact feature expression and compression of the foreground video sequence. At the decoding end, the generation network takes the entropy decoding feature as the network input and predicts and generates the foreground video frames. Finally, the decoded video is reconstructed by fusing the generated foreground video frames with the updated background frames.
Claims
1. A method for indoor surveillance video compression based on characteristic rate-distortion coding optimization, characterized in that Including background adaptive update, foreground feature rate distortion coding optimization, and coding reconstruction steps, firstly, the surveillance video is analyzed, and the background and human body are extracted using the background extraction algorithm and the human segmentation algorithm respectively; The feature extraction algorithm is used to transform the human body into compact key point features. The compact key point features and the reference frame are input into the generative adversarial network to generate non-reference frames, and the generated non-reference frames are fused with the background to obtain the reconstructed video.
2. The indoor surveillance video compression method based on characteristic rate-distortion coding optimization as claimed in claim 1, characterized in that: The background adaptive updating steps are as follows: The background update takes time series frames as input and dynamically adjusts the update in combination with the event trigger mechanism. Events are divided into periodic events and conditional events. Periodic events trigger the extraction of new background at fixed time intervals. Conditional events are triggered based on detection results or threshold changes, and the reference background is updated with the help of residual information.
3. The indoor surveillance video compression method based on characteristic rate-distortion coding optimization as claimed in claim 2, characterized in that: The conditional events include SSIM threshold events, residual fluctuation events and occlusion release events.
4. The indoor surveillance video compression method based on characteristic rate-distortion coding optimization as claimed in claim 3, characterized in that: The SSIM threshold event triggers an update when the SSIM between the current background and the reference background is less than a threshold.
5. The indoor surveillance video compression method based on characteristic rate-distortion coding optimization as claimed in claim 3, characterized in that: The residual fluctuation event is to update the background when the fluctuation of the residual video frame exceeds a set threshold; Let F t ={f1,f2,...,f t-1 ,f t ,...} is the original video frame sequence, F t ={f1,f2,...,f t-1 ,f t ,...} is the encoded video frame, R t ={r1,r2,...,r t-1 ,r t ,...} is the residual video frame sequence obtained by subtracting the original frame and the reconstructed video frame, B t ={b1,b2,...,b t-1 ,b t ,...} is the video background frame sequence corresponding to each frame; the residual formula is as follows: r t-1 (x,y)=f t-1 -f t-1 (1) b t =b t-1 +δr t-1 (x,y)(1-m t ) (2) δ is the residual compensation coefficient, which plays the role of adjusting the residual size and is determined by the rate-distortion model.
6. The indoor surveillance video compression method based on characteristic rate-distortion coding optimization as claimed in claim 5, characterized in that: The occlusion release event is when the occluder in the video is removed, and the background is updated; specifically, the mask change value Δm between frame t-1 and frame t is first calculated t (x,y),Δm t (x,y) = 1 means that the pixel (x,y) changes from the blocked area to the released area. If the proportion of the released area exceeds a certain threshold θ mask , it is considered that the occluder has been removed, triggering background update; let the number of pixels in the released area be S release , the total number of pixels is S total , the triggering condition is: Δm t =m t -m t-1 (3) θ mask It is an empirical value, usually set to 0.
1. After the occluder is detected and removed, only the residual in the released area is compensated to update the background. The update formula is: b t =b t-1 +δr t-1 (x,y)Δm t (5) 7. The indoor surveillance video compression method based on characteristic rate-distortion coding optimization as claimed in claim 2, characterized in that: The foreground feature rate-distortion coding optimization steps are as follows: Compact feature extraction: The target extraction algorithm is used to extract the foreground human sequence from the original video, and the front video frame of the person is selected as the key frame from the human sequence. k ; The average error MAD of the grayscale image is used to measure the texture complexity T of the human sequence frame C ,Using the joint point extraction algorithm to extract the joint points of the human body sequence and form the human skeleton sequence; Feature rate-distortion modeling: Plotting the bitrate R and video quality PSNR vs. the number of key points K and texture complexity T through statistical experiments C and the mapping relationship of the quantization parameters; based on the experimental structure, a characteristic rate distortion model such as formulas (6), (7) and (8) is constructed, and the model parameters are determined by experimental fitting; R=F(QP,TC,K)=(a1K+b1)ln(TC)+c1K+d1 (6) Coding parameter optimization selection: Based on the constructed characteristic rate-distortion model, the optimal coding parameter configuration is solved under the constraint of minimizing the rate-distortion cost. The process can be expressed as: Among them, [18,51] is the value range of the residual quantization parameter Qp, [8,18] is the value range of the number of joint points K, PSNR is converted to MSE, and MSE is used to directly measure the distortion D. By selecting the optimal encoding values of the foreground residual quantization parameter and the foreground skeleton joint points, the foreground encoding generation with the best rate-distortion performance is achieved.
8. The indoor surveillance video compression method based on characteristic rate-distortion coding optimization as claimed in claim 7, characterized in that: For each frame of video, extract the texture complexity T of the foreground C ; Through the rate-distortion model, input texture complexity T C and target bit rate R tar , and obtain the optimal residual quantization parameter Qp and the number of joint points K.
9. The indoor surveillance video compression method based on characteristic rate-distortion coding optimization as claimed in claim 7, characterized in that: The characteristic rate-distortion model predicts the distortion D pred , the optimal residual quantization parameter Qp and the number of joint points K, after encoding, a set of actual R actual and D actual During the encoding process, the parameters will be fine-tuned according to the difference between the actual bit rate and the target bit rate, and the difference between the actual distortion and the target distortion. The specific fine-tuning formula is as follows: After encoding, the RD model parameters are updated to make the model more accurate.
10. The indoor surveillance video compression method based on characteristic rate-distortion coding optimization as claimed in claim 1, characterized in that: The encoding reconstruction steps are as follows: At the decoding end, the generative network uses the characteristics of entropy decoding as network input to predict and generate foreground video frames, and decodes and reconstructs the video by fusing the generated foreground video frames with the updated background frames.