Low-complexity visible light video image compression method
Through the multivariate Gaussian model combining differential pulse coding and Hoffman coding, the problem of high energy consumption and computational complexity of digital video image compression system is solved, low-complexity video image compression is achieved, and the compression rate is improved and the computational complexity is reduced. It is suitable for application environments with limited resources.
Patent Information
- Application Number
- CN202510556686.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-29
- Publication Date
- 2025-07-18
AI Technical Summary
Existing digital video image compression systems have major problems in energy consumption and computing complexity, especially in application scenarios with limited computing resources and energy supply, which are difficult to meet the needs.
The multivariate Gaussian model is used to combine differential pulse coding and Hoffman coding methods. By establishing a multivariate Gaussian model, the pixel mean and residual coefficients are calculated, and the Hoffman coding module is used to compress the video image, including a compression algorithm module, a scalar quantization module and a differential pulse coding module, eliminating time domain redundancy and performing Hoffman coding.
It realizes low-complexity video image compression, improves compression rate and reduces calculation complexity, and is suitable for wireless multimedia sensing networks and mobile terminal video acquisition, and improves compression rate distortion performance on average PSNR 5.1dB~8.1dB, and the calculation complexity is only 30% of H.264/AVCInter.
Smart Images

Figure CN120343284A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a method for compressing video images, and particularly to a method for compressing visible light video images with low complexity. Background Art
[0002] Currently, commonly used digital video compression methods usually include a motion estimation link with relatively high computational complexity, such as encoding methods like H.263, H.264, and H.265. These encoding methods cannot meet the requirements of the following application scenarios:
[0003] First, the digital video image compression system has limited computational resources and energy supply;
[0004] Second, the computational complexity and latency of the digital video image compression system increase significantly as the spatial and temporal resolutions of the video increase.
[0005] Therefore, a video acquisition and encoding system with low complexity solves these problems by adopting a lossy compression method based on a multivariate Gaussian model; however, in terms of data compression after the camera obtains visible light video images, the requirements of this digital video image compression system for energy consumption and computational complexity are not considered. How to utilize the existing signal processing theory and combine the distribution characteristics of visible light video images to achieve efficient encoding and compression of digital videos has become a prominent problem in the design of existing digital video image compression systems. Summary of the Invention
[0006] The object of the present invention is to solve the technical problems of large energy consumption and high computational complexity of existing digital video image compression systems, and provide a method for compressing visible light video images with low complexity.
[0007] To achieve the above object, the present invention adopts the following technical solutions:
[0008] A method for compressing visible light video images with low complexity, characterized in that it includes the following steps:
[0009] 1. Establish a multivariate Gaussian model;
[0010] The multivariate Gaussian model includes a compression algorithm module, a scalar quantization module, a differential pulse coding module, and a Huffman coding module; the input end of the compression algorithm module and the scalar quantization module are respectively used to receive visible light video images, the output ends of the compression algorithm module and the scalar quantization module are respectively connected to the input end of the differential pulse coding module, the output end of the differential pulse coding module is connected to the input end of the Huffman coding module, and the output end of the Huffman coding module is used to output the compressed visible light video image;
[0011] 2】Calculate the pixel mean and residual coefficient of each image in the training set, train the compression algorithm module using the residual coefficient, and train the scalar quantization module using the pixel mean;
[0012] 3】Use the camera to obtain a video, calculate the residual coefficient and pixel mean of each frame image in the video, input the residual coefficient into the trained compression algorithm module to obtain the residual coefficient symbol with the minimum error; input the pixel mean into the trained scalar quantization module to obtain the mean symbol; use all the residual coefficient symbols and all the mean symbols as the lossy compression coding bitstream YS;
[0013] 4】Input the bitstream YS into the differential pulse coding module to eliminate the temporal redundancy, and obtain the bitstream YP;
[0014] 5】Use the Huffman coding module to perform Huffman coding on the bitstream YP to obtain the binary compressed bitstream SBS, and complete the compression processing of the video image.
[0015] Further, in step 1】, the compression algorithm module includes K product quantization units and a minimum error selection unit, where K≥2; each product quantization unit includes a Karhunen-Loeve transform module U and p scalar quantizers Q1, where p≥2;
[0016] The input ends of the K product quantization units are respectively connected to the residual coefficients of the input image, the output ends are respectively connected to the input end of the minimum error selection unit, and the output end of the minimum error selection unit is connected to the input end of the differential pulse coding module;
[0017] The scalar quantization module includes a scalar quantizer Q2, the input end of the scalar quantizer Q2 is connected to the pixel mean of the input image, and the output end is connected to the input end of the differential pulse coding module.
[0018] Further, step 2】 is specifically as follows:
[0019] 2.1. Divide the images in the training set into multiple image blocks;
[0020] 2.2. Calculate the pixel mean of each image block, and then subtract the corresponding pixel mean from the pixels of each image block to obtain all the residual coefficients x;
[0021] 2.3. Train the compression algorithm module using all the residual coefficients x;
[0022] 2.3.1. Define that the residual coefficients x of all image blocks follow a Gaussian distribution N, and we can get:
[0023]
[0024] In the formula, w kDenotes the weight of the k-th Gaussian component, where k = 1, 2, …, K, and K represents the number of Gaussian components corresponding to K product quantization units; u k Denotes the pixel mean of the k-th Gaussian component; Σ k Denotes the covariance matrix of the k-th Gaussian component;
[0025] 2.3.2. Obtain the optimal code rate B of each product quantization unit through the following formula k :
[0026]
[0027] In the formula, is the geometric mean of the eigenvalues of the covariance matrix Σ k of the k-th Gaussian component; B tot is the total code rate;
[0028] 2.3.3. Obtain the code rate b of each scalar quantizer Q1 through the following formula i :
[0029]
[0030] In the formula, γ i is the i-th eigenvalue of the covariance matrix Σ k of the k-th Gaussian component, where i = 1, 2, …, p;
[0031] 2.3.4. Obtain the Karhunen-Loeve transform U through the following formula to complete the training of the compression algorithm module;
[0032] U T Σ k U = diag(γ1, γ2,..., γ p )
[0033] In the formula, U T is the transpose of the Karhunen-Loeve transform U; γ1, γ2,..., γ p ∈γ i ; diag(...) represents the eigenvalues;
[0034] 2.4. Input all pixel means into the scalar quantizer Q2 simultaneously, and use the LBG algorithm to perform quantization processing on the pixel means to obtain the codebook of the scalar quantizer Q2, thus completing the training of the scalar quantization module.
[0035] Furthermore, step 3] is specifically as follows:
[0036] 3.1. Use a camera to obtain a video, and calculate the residual coefficients and pixel means of each frame image in the video;
[0037] 3.2. Input all the residual coefficients obtained in step 3.1 into the compression algorithm module, and use the product quantization unit to obtain all the residual coefficient symbols {x q1 , x q2 ,... x qp};
[0038] 3.3. Use the minimum error selection unit to select the residual coefficient symbol with the minimum error from all the residual coefficient symbols {x q1 , x q2 ,... x qp};
[0039] 3.4. Use the scalar quantizer Q2 to quantize the pixel mean of each frame of image respectively to obtain all the mean symbols;
[0040] 3.5. Take all the residual coefficient symbols and all the mean symbols as the lossy compression coding bitstream YS.
[0041] Furthermore, step 3.2 is specifically as follows:
[0042] 3.2.1. Subtract the corresponding pixel mean from all the residual coefficients of the images obtained in step 3.1 to get all the intermediate calculation amounts; then divide each intermediate calculation amount by the Karhunen-Loeve transform module U to get the decorrelated scalars {x1, x2,... x p};
[0043] 3.2.2. Send the elements in the scalars {x1, x2,... x p} into K product quantization units {PVQ1, PVQ2,..., PVQ K} respectively at the same time. The p scalar quantizers Q1 in the corresponding product quantization units quantize the scalars {x1, x2,... x p} respectively to obtain all the residual coefficient symbols {x q1 , x q2 ,... x qp}.
[0044] Furthermore, step 4] is specifically as follows:
[0045] Input all the mean symbols and all the residual coefficient symbols into the differential pulse coding module respectively to eliminate the temporal redundancy and obtain the bitstream YP;
[0046] The elimination of temporal redundancy is to interpolate the mean symbols and residual coefficient symbols at the corresponding positions in the temporally adjacent image blocks respectively to obtain the temporally redundancy-free compression symbols, that is, the bitstream YP.
[0047] Furthermore, step 5] is specifically as follows:
[0048] The bitstream YP of each frame of the video is combined in frame order by using the Huffman coding module, and the combined binary bitstream and the additional information are connected into a binary compressed bitstream SBS to complete the compression process of the video image;
[0049] The additional information includes the horizontal resolution, vertical resolution, and frame rate of the video frame.
[0050] Further, in step 2.3.1, the weight w k of the Gaussian component, the pixel mean u k and the covariance matrix Σ k are all obtained by training with the EM algorithm; in step 2.3.2, the total bit rate B tot takes a value of 64 ≤ B tot ≤ 256.
[0051] Further, in step 3.4, the bit rate A of the mean symbol takes a value of 5 ≤ A ≤ 8.
[0052] Advantages of the present invention:
[0053] 1. The method of the present invention can establish a multivariate Gaussian model based on the distribution characteristics of each frame of the video, and obtain a lossy compression method of the compression algorithm module based on the multivariate Gaussian model according to the obtained multivariate Gaussian model, so as to realize the efficient compression of the video.
[0054] 2. The method of the present invention combines the differential pulse coding module and the Huffman coding module to realize the compression application of visible light video. Compared with the existing methods for visible light video compression, it can average improve the PSNR of the visible light video compression rate distortion performance by 5.1 dB to 8.1 dB, and at the same time its operation complexity is only 30% of the H.264 / AVC Inter compression method.
[0055] 3. For the temporal redundancy of the video image, after using the differential pulse coding module to modulate and remove the redundancy, the Huffman coding module is further used to remove the data redundancy to obtain the binary compressed bitstream SBS, which can be used for subsequent storage and transmission.
[0056] 4. The method of the present invention can simultaneously eliminate the temporal redundancy of consecutive images in the video, reduce the compression calculation amount of the video image, and save the energy and computing resources of the hardware system.
[0057] 5. The method of the present invention can meet the application environments with energy and computational complexity limitations for the video coding system, such as wireless multimedia sensor networks, space video image acquisition, mobile terminal video acquisition, etc. Description of the Drawings
[0058] Figure 1 is a flowchart of an embodiment of the method of the present invention;
[0059] Figure 2 Schematic diagram of training the compression algorithm module and the scalar quantization module in the method embodiment of the present invention;
[0060] Figure 3 Schematic diagram of selecting the symbol of the residual coefficient with the minimum error in the method of the present invention. Detailed implementation manner
[0061] As Figure 1 shown, a low-complexity visible light video image compression method includes the following steps:
[0062] 1. Establish a multivariate Gaussian model; the multivariate Gaussian model includes a compression algorithm module, a scalar quantization module, a differential pulse coding module, and a Huffman coding module; the compression algorithm module includes K product quantization units and a minimum error selection unit, where K≥2; each product quantization unit includes a Karhunen-Loeve transform module and p scalar quantizers Q1, where p≥2; the input end of the compression algorithm module and the scalar quantization module are respectively used to receive visible light video images, the output ends of the compression algorithm module and the scalar quantization module are respectively connected to the input end of the differential pulse coding module, the output end of the differential pulse coding module and the input end of the Huffman coding module are connected, and the output end of the Huffman coding module is used to output the compressed visible light video image; the input ends of the K product quantization units are respectively connected to the residual coefficients of the input image, and the output ends are respectively connected to the input end of the minimum error selection unit, and the output end of the minimum error selection unit is connected to the input end of the differential pulse coding module; the scalar quantization module includes a scalar quantizer Q2, the input end of the scalar quantizer Q2 is connected to the pixel mean value of the input image, and the output end is connected to the input end of the differential pulse coding module;
[0063] 2. As Figure 2 shown, train the compression algorithm module and the scalar quantization module;
[0064] 2.1. Divide the image Y in the training set into 8×8 image blocks; the training set is at least 1000 frames of images;
[0065] 2.2. Calculate the pixel mean value of each image block, and then subtract the corresponding pixel mean value from each pixel of the image block to obtain all residual coefficients x;
[0066] 2.3. Train the compression algorithm module using all the residual coefficients x;
[0067] 2.3.1. Define that the residual coefficients x of all image blocks follow a Gaussian distribution N, and we can get:
[0068]
[0069] In the formula, wk denotes the weight of the k-th Gaussian component, where k = 1, 2, …, K, and K represents the number of Gaussian components corresponding to the K product quantization units; u k denotes the pixel mean of the k-th Gaussian component; Σ k denotes the covariance matrix of the k-th Gaussian component;
[0070] In this embodiment, the weight w k of the Gaussian component, the pixel mean u k and the covariance matrix Σ k are all obtained by training with the EM algorithm.
[0071] 2.3.2. Calculate the optimal code rate B of each product quantization unit through the following formula k :
[0072]
[0073] In the formula, is the geometric mean of the eigenvalues of the covariance matrix Σ k of the k-th Gaussian component; B tot is the total code rate, which is determined by the user according to the application situation. Usually, the total code rate B tot is in bits, and its value is 64 ≤ B tot ≤ 256.
[0074] 2.3.3. Calculate the code rate b of each scalar quantizer Q1 through the following formula i :
[0075]
[0076] In the formula, γ i is the i-th eigenvalue of the covariance matrix Σ k of the k-th Gaussian component, where i = 1, 2, …, p;
[0077] 2.3.4. Calculate the Karhunen-Loeve transform U through the following formula to complete the training of the compression algorithm module;
[0078] U T Σ k U = diag(γ1, γ2,..., γ p )
[0079] In the formula, U T is the transpose of the Karhunen-Loeve transform U; γ1, γ2,..., γ p ∈ γ i ; diag(...) represents the eigenvalues;
[0080] 2.4. Train the scalar quantization module using all pixel means;
[0081] The scalar quantization module includes a scalar quantizer Q2. The mean values of all pixels are simultaneously input into the scalar quantizer Q2 respectively, and the LBG algorithm is used to perform quantization processing on the pixel mean values to obtain the codebook of the scalar quantizer Q2, completing the training of the scalar quantization module;
[0082] 3. Use a camera to obtain a video, calculate the residual coefficients and pixel mean values of each frame image in the video, input the residual coefficients into the trained compression algorithm module to obtain the residual coefficient symbols with the minimum error; input the pixel mean values into the trained scalar quantization module to obtain the mean symbols; use all the residual coefficient symbols and all the mean symbols as the lossy compression coding bitstream YS;
[0083] 3.1. Use a camera to obtain a video and calculate the residual coefficients and pixel mean values of each frame image in the video;
[0084] 3.2. Input all the residual coefficients obtained in step 3.1 into the compression algorithm module, and use the product quantization unit to obtain all the residual coefficient symbols {x q1 , x q2 ,... x qp};
[0085] 3.2.1. Subtract the corresponding pixel mean values from all the residual coefficients of the images obtained in step 3.1 to obtain all the intermediate calculation quantities; then use the Karhunen-Loeve transform module U to divide each intermediate calculation quantity to obtain the decorrelated scalars {x1, x2,... x p};
[0086] 3.2.2. As Figure 3 shown, input the elements in the scalars {x1, x2,... x p} into the data and simultaneously send them into K product quantization units {PVQ1, PVQ2,..., PVQ K} respectively. The p scalar quantizers Q1 in the corresponding product quantization units perform quantization processing on the scalars {x1, x2,... x p} respectively to obtain the output data, that is, all the residual coefficient symbols {x q1 , x q2 ,... x qp};
[0087] 3.3. Use the minimum error selection unit to select the residual coefficient symbol with the minimum error among all the residual coefficient symbols {x q1 , x q2 ,... x qp};
[0088] 3.4. Use a scalar quantizer Q2 to quantize the pixel mean of each frame of the image respectively to obtain all mean symbols; the bit rate of the mean symbols is A bits, and the value of A is 5 ≤ A ≤ 8;
[0089] 3.5. Take all the residual coefficient symbols and all the mean symbols as the lossy compression coding bitstream YS;
[0090] 4. Input the bitstream YS into the differential pulse coding module to eliminate temporal redundancy, and obtain the bitstream YP;
[0091] Input all the mean symbols and all the residual coefficient symbols into the differential pulse coding module respectively to eliminate temporal redundancy, and obtain the bitstream YP; the elimination of temporal redundancy is to interpolate the mean symbols and residual coefficient symbols at the corresponding positions in temporally adjacent image blocks respectively, so as to encode the compressed symbols with fewer bitstreams and obtain compression symbols without temporal redundancy, that is, the bitstream YP;
[0092] 5. Use the Huffman coding module to perform Huffman coding on the bitstream YP to obtain the binary compressed bitstream SBS, and complete the compression process of the video image. Specifically:
[0093] In order to eliminate the redundant data existing in the bitstream YP, use the Huffman coding module to perform entropy coding on the compression symbols without temporal redundancy; use the Huffman coding module to combine the bitstreams YP of each frame of the video in frame order, and connect the combined binary bitstream with the additional information to form a binary compressed bitstream SBS that can be used for storage and transmission, and complete the compression process of the video image. Among them, the additional information includes the horizontal resolution, vertical resolution and frame rate of the video frame.
[0094] The present invention models a multivariate Gaussian model for the distribution characteristics of visible light videos, designs a lossy compression method for the multivariate Gaussian model according to the parameters of the multivariate Gaussian model, is used for the compression of visible light video frames, and adds a scalar quantization module, a differential pulse coding module and a Huffman coding module to constitute an implementation framework of a visible light video compressor, which can reduce the compression complexity of visible light videos, improve the compression efficiency, avoid distortion performance, and meet the requirements of the system for limited energy and computational complexity.
Claims
1. A low-complexity visible light video image compression method, characterized in that, It includes the following steps: 1) Establish a multivariate Gaussian model; The multivariate Gaussian model includes a compression algorithm module, a scalar quantization module, a differential pulse coding module, and a Huffman coding module; the input end of the compression algorithm module and the scalar quantization module are respectively used to receive visible light video images, the output ends of the compression algorithm module and the scalar quantization module are respectively connected to the input end of the differential pulse coding module, the output end of the differential pulse coding module is connected to the input end of the Huffman coding module, and the output end of the Huffman coding module is used to output the compressed visible light video image; 2) Calculate the pixel mean and residual coefficient of each image in the training set, and use the residual coefficient to train the compression algorithm module, and use the pixel mean to train the scalar quantization module; 3) Use the camera to obtain a video, calculate the residual coefficient and pixel mean of each frame image in the video, input the residual coefficient into the trained compression algorithm module, and obtain the residual coefficient symbol with the minimum error; Input the pixel mean into the trained scalar quantization module to obtain the mean symbol; Take all the residual coefficient symbols and all the mean symbols as the lossy compression coding bitstream YS; 4) Input the bitstream YS into the differential pulse coding module to eliminate temporal redundancy, and obtain the bitstream YP; 5) Use the Huffman coding module to perform Huffman coding on the bitstream YP to obtain the binary compressed bitstream SBS, and complete the compression processing of the video image.
2. The low-complexity visible light video image compression method according to claim 1, wherein: In step 1), the compression algorithm module includes K product quantization units and a minimum error selection unit, K≥2; each product quantization unit includes a Karhunen-Loeve transform module U and p scalar quantizers Q1, p≥2; The input ends of the K product quantization units are respectively connected to the residual coefficients of the input image, and the output ends are respectively connected to the input end of the minimum error selection unit, and the output end of the minimum error selection unit is connected to the input end of the differential pulse coding module; The scalar quantization module includes a scalar quantizer Q2, the input end of the scalar quantizer Q2 is connected to the pixel mean of the input image, and the output end is connected to the input end of the differential pulse coding module.
3. The low-complexity visible light video image compression method according to claim 2, wherein, Step 2) is specifically: 2).
1. Divide the images in the training set into multiple image blocks; 2).
2. Calculate the pixel mean of each image block, and then subtract the corresponding pixel mean from the pixels of each image block to obtain all residual coefficients x; 2).
3. Use all the residual coefficients x to train the compression algorithm module; 2).3.
1. Define that the residual coefficients x of all image blocks follow a Gaussian distribution N, and we can get: where w k represents the weight of the k-th Gaussian component, k = 1, 2, …, K, and K represents the number of Gaussian components corresponding to the K product quantization units; u k represents the pixel mean of the k-th Gaussian component; Σ k represents the covariance matrix of the k-th Gaussian component; 2).3.
2. Obtain the optimal bit rate B of each product quantization unit using the following formula k :[[]]END]] In the formula, is the geometric mean of the eigenvalues of the covariance matrix Σ k of the k-th Gaussian component; B tot is the total code rate; 2).3.
3. Calculate the code rate b of each scalar quantizer Q1 using the following formula i :[[]]END]] where γ i is the i-th eigenvalue of the covariance matrix Σ k of the k-th Gaussian component, i = 1, 2, …, p; 2).3.
4. Obtain the Karhunen-Loeve transform U through the following formula to complete the training of the compression algorithm module; U T Σ k U = diag(γ1, γ2,... p ) where, U T is the transpose of the Karhunen-Loeve transform U; γ1, γ2,..., γ p ∈γ i ; diag(...) represents the eigenvalues; 2).
4. Input all the pixel means into the scalar quantizer Q2 at the same time, and use the LBG algorithm to perform quantization processing on the pixel means to obtain the codebook of the scalar quantizer Q2, and complete the training of the scalar quantization module.
4. The low-complexity visible light video image compression method according to claim 3, wherein Step 3) is specifically: 3).
1. Use the camera to obtain a video, and calculate the residual coefficient and pixel mean of each frame image in the video; 3).
2. Input all the residual coefficients obtained in step 3.1 into the compression algorithm module, and use the product quantization unit to obtain all the residual coefficient signs {x q1 , x q2 ,... x qp}; 3).
3. Select the residual coefficient symbol with the minimum error from all the residual coefficient symbols {x q1 , x q2 ,... x qp}; 3).
4. Use a scalar quantizer Q2 to quantize the pixel mean of each frame of the image respectively to obtain all mean symbols; 3).
5. Take all the residual coefficient symbols and all the mean symbols as the lossy compression coding bitstream YS.
5. The low-complexity visible light video image compression method according to claim 4, characterized in that, Step 3.2 is specifically as follows: 3).2.
1. Subtract the pixel mean corresponding to each of the residual coefficients of all the images obtained in step 3.1 from the residual coefficients to obtain all intermediate calculation results; then divide each intermediate calculation result by the Karhunen-Loeve transform module U to obtain decorrelated scalars {x1, x2,... x p}; 3).2.
2. Send the elements in the scalar {x1, x2,... x p} into K product quantization units {PVQ1, PVQ2,..., PVQ K} simultaneously and separately. The p scalar quantizers Q1 in the corresponding product quantization units perform quantization processing on the scalars {x1, x2,... x p} respectively, and all residual coefficient signs {x q1 , x q2 ,... x qp} are obtained.
6. The low-complexity visible light video image compression method according to claim 5, wherein Step 4] is specifically as follows: Input all the mean symbols and all the residual coefficient symbols into the differential pulse coding module respectively to eliminate the temporal redundancy, and obtain the bitstream YP; The elimination of temporal redundancy is to interpolate the mean symbols and residual coefficient symbols at the corresponding positions in the temporally adjacent image blocks respectively to obtain the compression symbols without temporal redundancy, that is, the bitstream YP.
7. The low-complexity visible light video image compression method according to claim 6, wherein Step 5] is specifically as follows: Use the Huffman coding module to combine the bitstream YP of each frame of the video in frame order, and connect the combined binary bitstream with the additional information into a binary compressed bitstream SBS to complete the compression process of the video image; The additional information includes the horizontal resolution, vertical resolution and frame rate of the video frame.
8. The low-complexity visible light video image compression method according to claim 7, characterized in that: In Step 2.3.1, the weight w of the Gaussian component k , the pixel mean u k and the covariance matrix Σ k are all obtained by training with the EM algorithm; In step 2.3.2, the total bit rate B tot has a value of 64 ≤ B tot ≤ 256.
9. The low-complexity visible light video image compression method according to claim 8, characterized in that: In step 3.4, the code rate A of the mean symbol takes a value of 5 ≤ A ≤ 8.