Video lossless compression coding method and system based on spatio-temporal information
By distinguishing the foreground and background of video frames, and using spatiotemporal information for sampling and sparse encoding, the problems of low lossless video compression efficiency and complex calculation in the prior art are solved, and efficient video lossless compression is achieved.
Patent Information
- Application Number
- CN202510511003.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-23
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2045-04-23
AI Technical Summary
When the existing video lossless compression technology removes redundant information in video data, there are prediction errors and complex operations, which affects video recovery and is not high compression efficiency.
By distinguishing the foreground and background of the video frames, a spatiotemporal matrix and an offset matrix are constructed, sampling and selection are performed along the time series, the amplitude of the pixel units not being sampled is retained, and the code is encoded using sparse encoding.
It improves encoding efficiency, achieves a higher compression ratio, and reduces the computational complexity, ensuring the recovery of video quality.
Smart Images

Figure CN120050428A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of video compression, and in particular to a video lossless compression coding method and system based on spatio-temporal information. Background Art
[0002] Video lossless compression requires compressing video images while ensuring video quality, reducing the storage space of video files. Lossless compression of video helps improve storage efficiency and retrieval efficiency, facilitating long-term and large-scale storage, as well as retrieval management of videos. Video lossless compression technology ensures that the effective information in the video is not lost and can be widely applied in the fields of medical records, legal evidence collection, and road monitoring.
[0003] Existing video lossless compression technologies mainly start from removing redundant information in video data and use coding methods to encode and store the effective data in the video to ensure that the compressed video can be restored to its original state. Commonly used coding methods include predictive coding and transform coding. Among them, predictive coding reduces the amount of data by predicting the pixel values in video frames. Due to the existence of prediction errors, the restoration of the video is affected. The transform coding method performs frequency domain transformation on video data, involving complex operations and predictions, and is limited in application. Summary of the Invention
[0004] In view of this, the present invention provides a video lossless compression coding method and system based on spatio-temporal information, integrating spatio-temporal information into the video lossless compression coding method, and using the context correlation of video frames, which is beneficial to improving the coding efficiency.
[0005] To achieve the above objectives, the present invention is implemented through the following technical solutions: A video lossless compression coding method based on spatio-temporal information, comprising the following steps: S1: For each frame image in the video, distinguish the foreground and the background to obtain foreground units; S2: Connect the foreground units in the image to obtain a foreground region; S3: Calculate the pixel unit index where the centroid of the foreground region is located; and based on the pixel unit index where the centroid is located, construct a spatio-temporal matrix of the foreground region, and then calculate the offset speed of the foreground region; then based on the offset speed of the foreground region, construct a foreground region offset matrix along the time series; S4: Based on the foreground region offset matrix, sample and select the images in the video along the time series; S5: In the images obtained by sampling and selecting in step S4, sample and select the pixel units of the foreground region, and retain the amplitudes of the pixel units that are not sampled and selected; S6: Encode the images obtained in step S5 using sparse coding.
[0006] Optionally, in step S1, for each frame image in the video, the foreground and background are distinguished to obtain foreground units, including: S11: For each frame image in the video, a two-dimensional matrix is constructed, and the elements therein are represented as: ; Among them, represents the element with index in the two-dimensional matrix, represents the index of the pixel unit in the horizontal axis direction in the image, represents the index of the pixel unit in the vertical axis direction in the image, represents the amplitude value of the pixel unit; S12: The two-dimensional matrix is scanned to find the peak, and the index of the pixel unit where the peak is located is ; S13: Based on the two-dimensional matrix, the background noise of the image is estimated: ; Among them, represents the estimated value of the image background noise, represents the number of pixel units in the image; S14: The peak obtained in step S12 is filtered to determine the foreground unit: ; Among them, represents that the pixel unit where the peak is located belongs to the foreground unit, represents that the pixel unit where the peak is located belongs to the background unit, represents the detection scale factor.
[0007] Optionally, in step S2, the foreground units in the image are connected to obtain the foreground region, including: Starting from any foreground unit in the image, the breadth-first algorithm is used to connect the foreground units to obtain the foreground region.
[0008] Optionally, in step S3, the index of the pixel unit where the centroid of the foreground region is located is calculated; and based on the index of the pixel unit where the centroid is located, a spatio-temporal matrix of the foreground region is constructed, and then the offset speed of the foreground region is calculated; then based on the offset speed of the foreground region, a foreground region offset matrix is constructed along the time series, including: S31: Calculate the index of the pixel unit where the centroid of the foreground region is located; S32: Based on the index of the pixel unit where the centroid is located, construct a spatio-temporal matrix of the foreground region: ; Among them, represents the The index of the pixel unit where the centroid is located in the horizontal axis direction in the frame image, indicating the The index of the pixel unit where the centroid is located in the vertical axis direction in the frame image, indicating the frame sequence of the image; S33: Calculate the offset velocity of the foreground region: ; wherein, indicating the Offset velocity of the foreground region in the horizontal axis direction in the indicating the Offset velocity of the foreground region in the vertical axis direction in the indicating the The index of the pixel unit where the centroid is located in the horizontal axis direction in the indicating the The index of the pixel unit where the centroid is located in the vertical axis direction in the S34: Construct a foreground region offset matrix along the time series: .
[0009] Optionally, in the step S31, calculating the index of the pixel unit where the centroid of the foreground region is located includes: S311: Calculate the centroid position of the foreground region: ; wherein, indicating the Position of the centroid of the foreground region in the horizontal axis direction in the indicating the Position of the centroid of the foreground region in the vertical axis direction in the indicating the Number of units in the foreground region in the indicating the Index of each unit in the foreground region in the horizontal axis direction in the indicating the Index of each unit in the foreground region in the vertical axis direction in the S312: The index of the pixel unit where the centroid is located is: ; wherein, indicating the The index of the pixel unit where the centroid is located in the horizontal axis direction in the indicating the The index of the pixel unit where the centroid is located in the vertical axis direction in the Indicates rounding to the nearest integer.
[0010] Optionally, in step S4, based on the foreground region offset matrix, sampling and selection of images in the video along the time series includes: Determine the sampled and selected images: ; Wherein, Indicates that the image is sampled and selected, Indicates that the image is not sampled and selected, Indicates the offset threshold for sampling and selection.
[0011] Optionally, in step S5, among the images obtained by sampling and selection in step S4, sampling and selection of pixel units in the foreground region, and retaining the amplitudes of the pixel units not sampled and selected, includes: S51: In the image obtained by sampling and selection in step S4, determine the edge of the foreground region, including: S511: In the image obtained by sampling and selection in step S4, retain the pixel amplitudes of the foreground units, and set the amplitudes of the remaining pixel units to 0; S512: Traverse the image along the horizontal axis, find the minimum horizontal axis index and the maximum horizontal axis index where the amplitude of the pixel unit is not 0, and determine the horizontal axis edge of the target; S513: Traverse the image along the vertical axis, find the minimum vertical axis index and the maximum vertical axis index where the amplitude of the pixel unit is not 0, and determine the vertical axis edge of the target; S52: Starting from the pixel units at the edge of the foreground region, perform continuous extraction at intervals of the sampling factor to obtain the sampled and selected pixel units, set the amplitudes of the sampled and selected pixel units to 0, and retain the amplitudes of the pixel units not sampled and selected; When sampling and selecting pixel units, the sampling factor is set to: ; Wherein, Is the scale factor.
[0012] Optionally, in step S6, sparse coding is used to encode the image obtained in step S5, including: Use sparse coding to encode the image obtained in step S5; ; Wherein, Indicates the sparse coding result, Indicates the index of the pixel unit with a non-zero amplitude in the horizontal axis direction of the image, Indicates the index of the pixel unit with a non-zero amplitude in the vertical axis direction of the image, Denote the frame sequence of the image obtained in step S5. Denote the amplitude of the pixel units in the image whose amplitudes are not 0.
[0013] The present invention also provides a video lossless compression and coding system based on spatio-temporal information, including: Foreground unit recognition module: construct a two-dimensional matrix of the image, find the peaks, estimate the background noise of the image, and determine the foreground units; Foreground region determination module: connect the foreground units to obtain the foreground region; Foreground region offset matrix module: calculate the centroid of the foreground region, construct the spatio-temporal matrix of the foreground region, calculate the offset speed of the foreground region, and construct the foreground region offset matrix; Image sampling module: determine the image selected for sampling; Foreground region sampling module: determine the edge of the foreground region and sample the pixel units inside the foreground region; Sparse coding module: perform sparse coding on the image.
[0014] Beneficial effects: Based on the spatio-temporal information of video data, the present invention effectively realizes the lossless compression of videos in multiple scenarios by using the sampling method; by utilizing the video context correlation, the coding efficiency is improved, and a high compression ratio is obtained with a low computational complexity.
[0015] The method of constructing a matrix based on video images in the present invention is conducive to the accurate representation and management of video data, provides convenience for subsequent compression coding, and simplifies the processing flow; differentiating the foreground and background of video images is conducive to removing redundant background information, allocating more storage resources to foreground information, and realizing the effective compression of foreground information; adopting different sampling methods for the foreground region contour and the interior of the foreground region is conducive to the accurate coding of the target; sampling the video image along the time axis is conducive to increasing the compression ratio.
[0016] On the basis of obtaining the foreground units by combining peak detection with filtering methods, the present invention is conducive to obtaining the overall region of interest through connectivity processing, avoiding repeated operation and processing of the key attention region; by calculating the centroid of the foreground region, it is conducive to sampling the video image according to the motion law of the foreground region, conducive to discriminating redundant information, and realizing lossless compression; the setting of the sampling factor inside the foreground region is proportional to the size of the foreground region, which is conducive to increasing the sampling rate of small regions. Description of the drawings
[0017] Figure 1 It is a schematic flowchart of a video lossless compression and coding method based on spatio-temporal information provided by an embodiment of the present invention.
[0018] Figure 2This is the image obtained by performing connectivity processing on foreground units in the image in step S2 of an embodiment of the present invention.
[0019] Figure 3 This is the image obtained by performing sampling selection processing on the foreground region in step S5 of an embodiment of the present invention. Detailed implementation manners
[0020] The present invention will be further described below with reference to the accompanying drawings, but the present invention is not limited in any way. Any transformation or replacement made based on the teachings of the present invention falls within the protection scope of the present invention.
[0021] Embodiment 1: A video lossless compression coding method based on spatio-temporal information, as Figure 1 shown, includes the following steps: S1: For each frame image in the video, distinguish the foreground and the background to obtain foreground units, including: S11: For each frame image in the video, construct a two-dimensional matrix, and the elements therein are represented as: ; Wherein, represents the element with index in the two-dimensional matrix, represents the index of the pixel unit in the horizontal axis direction of the image, represents the index of the pixel unit in the vertical axis direction of the image, represents the amplitude of the pixel unit; S12: Scan the two-dimensional matrix to find the peaks, and the index of the pixel unit where the peak is located is ; S13: Based on the two-dimensional matrix, estimate the image background noise: ; Wherein, represents the estimated value of the image background noise, represents the number of pixel units in the image; S14: Filter the peaks obtained in step S12 to determine the foreground units: ; Wherein, represents that the pixel unit where the peak is located belongs to the foreground unit, represents that the pixel unit where the peak is located belongs to the background unit, represents the detection scale factor.
[0022] In the embodiments of the present invention, the foreground refers to valuable information, and the background refers to clutter or content that is not the key focus and can be ignored in video compression. For example, in vehicle tracking, the vehicle is the foreground, and the road and roadside environment are the background; in a video of monitoring the flow of people, people are the foreground and the environment is the background.
[0023] S2: Connect the foreground units in the image to obtain a foreground region, including: Starting from any foreground unit in the image, use the breadth-first algorithm to connect the foreground units to obtain a foreground region.
[0024] In the embodiments of the present invention, Figure 2 The figure shows the image obtained by the connectivity processing of the foreground units in the image in step S2 (the face in the image is blurred due to portrait rights); it can be seen that Figure 2 The person in the figure is selected and the overall contour is continuous, and the background area is represented by blurring.
[0025] S3: Calculate the pixel unit index where the centroid of the foreground region is located; and based on the pixel unit index where the centroid is located, construct a spatio-temporal matrix of the foreground region, and then calculate the offset speed of the foreground region; then, based on the offset speed of the foreground region, construct a foreground region offset matrix along the time series, including: S31: Calculate the pixel unit index where the centroid of the foreground region is located, including: S311: Calculate the centroid position of the foreground region: ; Wherein, represents the position of the centroid of the foreground region in the horizontal axis direction in the th frame image, represents the position of the centroid of the foreground region in the vertical axis direction in the th frame image, represents the number of units in the foreground region in the th frame image, represents the index of each unit in the foreground region in the horizontal axis direction in the th frame image, represents the index of each unit in the foreground region in the vertical axis direction in the th frame image; S312: The pixel unit index where the centroid is located is: ; Wherein, represents the index of the pixel unit where the centroid is located in the horizontal axis direction in the th frame image, represents the index of the pixel unit where the centroid is located in the vertical axis direction in the th frame image, Indicates rounding to the nearest integer; S32: Based on the pixel unit index where the centroid is located, construct the spatio-temporal matrix of the foreground region: ; Wherein, Indicates the index of the pixel unit where the centroid is located in the horizontal axis direction in the -th frame image, Indicates the index of the pixel unit where the centroid is located in the vertical axis direction in the -th frame image, Indicates the frame sequence of the image; S33: Calculate the offset velocity of the foreground region: ; Wherein, Indicates the offset velocity of the foreground region in the horizontal axis direction in the -th frame image, Indicates the offset velocity of the foreground region in the vertical axis direction in the -th frame image, Indicates the index of the pixel unit where the centroid is located in the horizontal axis direction in the -th frame image, Indicates the index of the pixel unit where the centroid is located in the vertical axis direction in the -th frame image; S34: Along the time series, construct the foreground region offset matrix: .
[0026] S4: Based on the foreground region offset matrix, along the time series, sample and select the images in the video, including: Determine the sampled and selected images: ; Wherein, Indicates that the image is sampled and selected, Indicates that the image is not sampled and selected, Indicates the sampling selection offset threshold.
[0027] S5: In the images obtained by sampling and selecting in step S4, sample and select the pixel units of the foreground region, and retain the amplitudes of the pixel units that are not sampled and selected, including: S51: In the images obtained by sampling and selecting in step S4, determine the edge of the foreground region, including: S511: In the images obtained by sampling and selecting in step S4, retain the pixel amplitudes of the foreground units, and set the amplitudes of the remaining pixel units to 0; S512: Traverse the image along the horizontal axis, find the minimum horizontal axis index and the maximum horizontal axis index where the amplitude of the pixel unit is not 0, and determine the horizontal axis edge of the target; S513: Traverse the image along the vertical axis, find the minimum vertical axis index and the maximum vertical axis index where the amplitude of the pixel unit is not 0, and determine the vertical axis edge of the target; S52: Use the pixel unit at the edge of the foreground area as the starting unit, continuously extract at intervals of the sampling factor to obtain the sampled pixel units, set the amplitude of the sampled pixel units to 0, and retain the amplitude of the pixel units not sampled; When sampling and selecting pixel units, the sampling factor is set to: ; Wherein, is the scale factor.
[0028] In the embodiment of the present invention, Figure 3 The figure shows the image obtained by sampling and processing the foreground area in step S5 (due to concerns about portrait rights, the faces in the image are blurred); the amplitude of the pixel units in the background area of the figure is 0, and in the foreground area, the amplitude of the sampled pixel units is set to 0; it can be seen that Figure 3 the background area in [figure number] is black and the resolution of the foreground area is reduced.
[0029] S6: Use sparse coding to encode the image obtained in step S5, including: Use sparse coding to encode the image obtained in step S5; ; Wherein, represents the sparse coding result, represents the index of the pixel unit with a non-zero amplitude in the horizontal axis direction in the image, represents the index of the pixel unit with a non-zero amplitude in the vertical axis direction in the image, represents the frame sequence of the image obtained in step S5, represents the amplitude of the pixel unit with a non-zero amplitude in the image.
[0030] Embodiment 2: The present invention also provides a video lossless compression coding system based on spatio-temporal information, including the following six modules: Foreground unit recognition module: Construct a two-dimensional matrix of the image, find the wave peaks, estimate the background noise of the image, and determine the foreground units; Foreground area determination module: Connect the foreground units to obtain the foreground area; Foreground area offset matrix module: Calculate the centroid of the foreground area, construct the spatio-temporal matrix of the foreground area, calculate the offset speed of the foreground area, and construct the foreground area offset matrix; Image Sampling Module: Determine the sampled image; Foreground Region Sampling Module: Determine the foreground region edge and sample the internal pixel units of the foreground region; Sparse Coding Module: Perform sparse coding on the image.
[0031] It should be noted that the serial numbers of the above embodiments of the present invention are only for description and do not represent the superiority or inferiority of the embodiments. And the term "including", "comprising" or any other variant thereof in this article is intended to cover non-exclusive inclusion, so that a process, device, article or method including a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, device, article or method. Without further limitation, an element defined by the statement "including one..." does not exclude the existence of another identical element in the process, device, article or method including the element.
[0032] Through the description of the above embodiments, those skilled in the art can clearly understand that the above embodiment methods can be implemented by means of software plus a necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of the present invention, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. The computer software product is stored in a storage medium as described above (such as ROM / RAM, magnetic disk, optical disc), and includes several instructions to enable a terminal device (which can be a mobile phone, computer, server, or network device, etc.) to execute the methods described in various embodiments of the present invention.
[0033] The above are only the preferred embodiments of the present invention, and do not limit the patent scope of the present invention accordingly. Any equivalent structure or equivalent process transformation made by using the description of the specification and drawings of the present invention, or directly or indirectly applied in other related technical fields, shall be equally included in the patent protection scope of the present invention.
Claims
1. A video lossless compression encoding method based on spatiotemporal information, characterized in that: The method comprises: S1: For each frame of the video, distinguish the foreground and background and obtain the foreground unit; S2: Connect the foreground units in the image to obtain the foreground area; S3: Calculate the pixel unit index where the centroid of the foreground area is located; and based on the pixel unit index where the centroid is located, construct the spatiotemporal matrix of the foreground area, and then calculate the offset speed of the foreground area; and then construct the foreground area offset matrix along the time series based on the offset speed of the foreground area; S4: based on the foreground region offset matrix, the images in the video are sampled along the time sequence; S5: In the image obtained by sampling in step S4, pixel units in the foreground area are sampled and selected, and the amplitudes of pixel units that are not sampled are retained; S6: Encode the image obtained in step S5 using sparse coding.
2. The video lossless compression encoding method based on spatiotemporal information according to claim 1, characterized in that: The step S1 comprises: S11: For each frame of the video, a two-dimensional matrix is constructed, in which the elements are expressed as: ; in, Indicates that the index in the two-dimensional matrix is Elements of Represents the index of the pixel unit in the image in the horizontal direction, Represents the index of the pixel unit in the image in the vertical direction, Represents the amplitude of the pixel unit; S12: Scan the two-dimensional matrix to find the peak. The pixel unit index where the peak is located is ; S13: Estimate image background noise based on a two-dimensional matrix: ; in, represents the estimated value of image background noise, Indicates the number of pixel units in the image; S14: Filter the peaks obtained in step S12 to determine the foreground units: ; in, Indicates that the pixel unit where the peak is located belongs to the foreground unit. The pixel unit where the peak is located belongs to the background unit. Represents the detection scale factor.
3. The video lossless compression encoding method based on spatiotemporal information according to claim 2 is characterized in that: The step S2 comprises: Starting from any foreground unit in the image, the breadth-first algorithm is used to connect the foreground units to obtain the foreground area.
4. The video lossless compression encoding method based on spatiotemporal information according to claim 3 is characterized in that: The step S3 comprises: S31: Calculate the pixel unit index where the centroid of the foreground area is located; S32: Based on the pixel unit index where the centroid is located, construct the spatiotemporal matrix of the foreground area: ; in, Indicates The index of the pixel unit where the centroid of the frame image is located in the horizontal direction, Indicates The index of the pixel unit where the centroid is located in the frame image in the vertical direction, A sequence of frames representing an image; S33: Calculate the offset speed of the foreground area: ; in, Indicates The offset speed of the foreground area in the frame image along the horizontal axis, Indicates The offset speed of the foreground area along the vertical axis in the frame image, Indicates The index of the pixel unit where the centroid of the frame image is located in the horizontal direction, Indicates The index of the pixel unit where the centroid is located in the frame image in the vertical axis direction; S34: Construct the foreground area offset matrix along the time series: 。 5. The video lossless compression encoding method based on spatiotemporal information according to claim 4 is characterized in that: The step S31 comprises: S311: Calculate the centroid position of the foreground area: ; in, Indicates The position of the centroid of the foreground area in the frame image in the horizontal direction, Indicates The position of the centroid of the foreground area in the frame image along the vertical axis. Indicates The number of cells in the foreground area of the frame image, Indicates The index of each unit in the foreground area of the frame image in the horizontal direction, Indicates The index of each unit in the foreground area of the frame image in the vertical direction; S312: The pixel unit index where the centroid is located is: ; in, Indicates The index of the pixel unit where the centroid of the frame image is located in the horizontal direction, Indicates The index of the pixel unit where the centroid is located in the frame image in the vertical direction, Indicates rounding to the nearest integer.
6. The video lossless compression encoding method based on spatiotemporal information according to claim 4 is characterized in that: The step S4 comprises: Determine the images to be sampled: ; in, Indicates that the image is sampled. Indicates that the image is not sampled. Indicates the offset threshold for sampling selection.
7. The video lossless compression encoding method based on spatiotemporal information according to claim 6 is characterized in that: The step S5 comprises: S51: In the image sampled and selected in step S4, the edge of the foreground area is determined, including: S511: in the image sampled and selected in step S4, retain the pixel amplitude of the foreground unit, and set the amplitudes of the remaining pixel units to 0; S512: traverse the image along the horizontal axis, find the minimum horizontal axis index and the maximum horizontal axis index of the pixel unit whose amplitude is not 0, and determine the horizontal axis edge of the target; S513: traverse the image along the vertical axis, find the minimum vertical axis index and the maximum vertical axis index whose pixel unit amplitude is not 0, and determine the vertical axis edge of the target; S52: starting with the pixel unit at the edge of the foreground area, continuously sampling and selecting with the sampling factor as the interval, obtaining the sampled pixel unit, setting the amplitude of the sampled pixel unit to 0, and retaining the amplitude of the pixel unit that is not sampled; When sampling pixel units, the sampling factor is set to: ; in, is the scale factor.
8. The video lossless compression encoding method based on spatiotemporal information according to claim 7, characterized in that: The step S6 comprises: Encoding the image obtained in step S5 by using sparse coding; ; in, represents the sparse coding result, Indicates the index of the pixel unit in the image whose amplitude is not 0 in the horizontal direction. Indicates the index of the pixel unit in the image whose amplitude is not 0 in the vertical direction. represents the frame sequence of the image obtained in step S5, Indicates the amplitude of the pixel unit in the image whose amplitude is not 0.
9. A video lossless compression coding system based on spatiotemporal information, characterized in that: include: Foreground unit recognition module: constructs a two-dimensional image matrix, finds peaks, estimates image background noise, and determines foreground units; Foreground area determination module: connects the foreground units to obtain the foreground area; Foreground area offset matrix module: calculates the centroid of the foreground area, constructs the foreground area space-time matrix, calculates the offset speed of the foreground area, and constructs the offset matrix of the foreground area; Image sampling module: determines the image selected for sampling; Foreground area sampling module: determine the edge of the foreground area and sample the pixel units inside the foreground area; Sparse coding module: performs sparse coding on images; To implement a video lossless compression encoding method based on spatiotemporal information as described in any one of claims 1-8.
Citation Information
Patent Citations
Video coding method, decoding method and related device
CN119299704A
Method and apparatus of 360 degree camera video processing with targeted view
US20200257918A1