Video compression method, electronic device, storage medium and program product
By downsampling the video frame and determining the target reference macroblock and reference vector in the reference frame, the problem of low video compression efficiency in the prior art is solved, and more efficient video compression and lower storage space occupation are achieved.
Patent Information
- Application Number
- CN202510251684.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-04
- Publication Date
- 2025-06-06
AI Technical Summary
Existing video compression techniques require large computing resources when determining the reference image blocks corresponding to each image block, resulting in low video compression efficiency.
By downsampling the current frame, a pre-compressed macroblock is obtained, and the target reference macroblock and reference vector corresponding to each pre-compressed macroblock are determined in the reference frame, thereby determining the target compressed macroblock and finally generating a compressed frame.
Reduces data processing volume, improves video compression efficiency, improves hardware encoder performance, and reduces storage space usage.
Smart Images

Figure CN120111233A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computers, and in particular to a video compression method, electronic equipment, storage medium and program product. Background Art
[0002] When storing videos, there is a lot of similar content between consecutive frames in a video sequence, which will cause temporal redundancy and occupy a lot of storage space.
[0003] In the related art, a video can be compressed by a hardware encoder, and during the video compression process, a corresponding reference frame can be set for each current frame. When compressing the current frame, the current frame can be divided into multiple image blocks, and a reference image block corresponding to each image block is determined in the reference frame, and the current frame is compressed based on each image block and its corresponding reference image block.
[0004] However, in the above process, large computing resources are required to determine the reference image block corresponding to each image block in the reference frame, resulting in low efficiency in video compression. Summary of the invention
[0005] The embodiments of the present application provide a video compression method, an electronic device, a storage medium, and a program product to improve the efficiency of video compression.
[0006] In a first aspect, an embodiment of the present application provides a video compression method, including:
[0007] Acquire a current frame and a reference frame of a target video, wherein the current frame includes a plurality of current macroblocks;
[0008] Down-sampling each current macroblock to obtain multiple pre-compressed macroblocks;
[0009] Determining a target reference macroblock corresponding to each pre-compressed macroblock in the reference frame, and determining a reference vector corresponding to each target reference macroblock, wherein the size of the reference frame is smaller than the size of the current frame;
[0010] For any pre-compressed macroblock, determine a target compressed macroblock corresponding to the pre-compressed macroblock according to the pre-compressed macroblock, a target reference macroblock corresponding to the pre-compressed macroblock, and a reference vector;
[0011] A compressed frame corresponding to the current frame is generated according to the multiple target compressed macroblocks.
[0012] In a possible implementation method, determining a target reference macroblock corresponding to each pre-compressed macroblock in the reference frame includes:
[0013] Determine a macroblock size corresponding to the pre-compressed macroblock, where the macroblock size is used to indicate the number of pixels included in the horizontal direction and the number of pixels included in the vertical direction of the pre-compressed macroblock;
[0014] According to the macroblock size, the reference frame is divided into blocks to obtain a plurality of reference macroblocks, wherein the size of the reference macroblock is the same as the size of the macroblock;
[0015] Among the multiple reference macroblocks, a target reference macroblock corresponding to each pre-compressed macroblock is determined.
[0016] In a possible implementation method, for any pre-compressed macroblock, determining a target reference macroblock corresponding to the pre-compressed macroblock among the multiple reference macroblocks includes:
[0017] Determine a motion vector between the pre-compressed macroblock and each reference macroblock to obtain a plurality of motion vectors;
[0018] For any motion vector, determining a rate-distortion cost corresponding to the motion vector according to the pre-compressed macroblock and a reference macroblock corresponding to the motion vector;
[0019] Determine a reference vector among the multiple motion vectors according to the rate-distortion cost corresponding to each motion vector, wherein the rate-distortion cost corresponding to the reference vector is the smallest;
[0020] The reference macroblock corresponding to the reference vector is determined as the target reference macroblock.
[0021] In a possible implementation method, determining a rate-distortion cost corresponding to the motion vector according to the pre-compressed macroblock and a reference macroblock corresponding to the motion vector includes:
[0022] Determine the sum of absolute value errors corresponding to the motion vector according to the pre-compressed macroblock and the reference macroblock;
[0023] Determining, based on the motion vector, the number of bits required to encode the motion vector;
[0024] Determine a bit rate cost corresponding to the motion vector according to the pre-compressed macroblock, the number of bits and the motion vector;
[0025] The rate-distortion cost is determined according to the bit rate cost and the sum of absolute value errors.
[0026] In a possible implementation method, determining a target compressed macroblock corresponding to the pre-compressed macroblock according to the pre-compressed macroblock, a target reference macroblock corresponding to the pre-compressed macroblock, and a reference vector includes:
[0027] determining a residual between the pre-compressed macroblock and the target reference macroblock;
[0028] Performing encoding processing on the residual to obtain a target residual;
[0029] The target residual and the reference vector are reconstructed to obtain the target compressed macroblock.
[0030] In a possible implementation method, generating a compressed frame corresponding to the current frame according to the multiple target compressed macroblocks includes:
[0031] For any current macroblock, determining a target position of a target compressed macroblock corresponding to the current macroblock according to the position of the current macroblock in the current frame;
[0032] According to the target position of each target compressed macroblock, the multiple target compressed macroblocks are spliced to obtain the compressed frame.
[0033] In a possible implementation method, obtaining a reference frame of the target video includes:
[0034] Acquire the reference frame in a first preset storage space;
[0035] After generating a compressed frame corresponding to the current frame according to the multiple target compressed macroblocks, the method further includes:
[0036] In the first preset storage space, the compressed frame is updated as the reference frame.
[0037] In a second aspect, an embodiment of the present application provides a video compression device, including: an acquisition module, a downsampling module, a first determination module, a second determination module and a generation module, wherein:
[0038] The acquisition module is used to acquire a current frame and a reference frame of a target video, wherein the current frame includes a plurality of current macroblocks;
[0039] The downsampling module is used to perform downsampling processing on each current macroblock to obtain multiple pre-compressed macroblocks;
[0040] The first determination module is used to determine a target reference macroblock corresponding to each pre-compressed macroblock in the reference frame, and determine a reference vector corresponding to each target reference macroblock, wherein the size of the reference frame is smaller than the size of the current frame;
[0041] The second determination module is used to determine, for any pre-compressed macroblock, a target compressed macroblock corresponding to the pre-compressed macroblock according to the pre-compressed macroblock, a target reference macroblock corresponding to the pre-compressed macroblock, and a reference vector;
[0042] The generation module is used to generate a compressed frame corresponding to the current frame according to the multiple target compressed macroblocks.
[0043] In a possible implementation method, the first determining module is specifically configured to:
[0044] Determine a macroblock size corresponding to the pre-compressed macroblock, where the macroblock size is used to indicate the number of pixels included in the horizontal direction and the number of pixels included in the vertical direction of the pre-compressed macroblock;
[0045] According to the macroblock size, the reference frame is divided into blocks to obtain a plurality of reference macroblocks, wherein the size of the reference macroblock is the same as the size of the macroblock;
[0046] Among the multiple reference macroblocks, a target reference macroblock corresponding to each pre-compressed macroblock is determined.
[0047] In a possible implementation method, the first determining module is specifically configured to:
[0048] Determine a motion vector between the pre-compressed macroblock and each reference macroblock to obtain a plurality of motion vectors;
[0049] For any motion vector, determining a rate-distortion cost corresponding to the motion vector according to the pre-compressed macroblock and a reference macroblock corresponding to the motion vector;
[0050] Determine a reference vector among the multiple motion vectors according to the rate-distortion cost corresponding to each motion vector, wherein the rate-distortion cost corresponding to the reference vector is the smallest;
[0051] The reference macroblock corresponding to the reference vector is determined as the target reference macroblock.
[0052] In a possible implementation method, the first determining module is specifically configured to:
[0053] Determine the sum of absolute value errors corresponding to the motion vector according to the pre-compressed macroblock and the reference macroblock;
[0054] Determining, based on the motion vector, the number of bits required to encode the motion vector;
[0055] Determine a bit rate cost corresponding to the motion vector according to the pre-compressed macroblock, the number of bits and the motion vector;
[0056] The rate-distortion cost is determined according to the bit rate cost and the sum of absolute value errors.
[0057] In a possible implementation method, the second determining module is specifically configured to:
[0058] determining a residual between the pre-compressed macroblock and the target reference macroblock;
[0059] Performing encoding processing on the residual to obtain a target residual;
[0060] The target residual and the reference vector are reconstructed to obtain the target compressed macroblock.
[0061] In a possible implementation method, the generating module is specifically used to:
[0062] For any current macroblock, determining a target position of a target compressed macroblock corresponding to the current macroblock according to the position of the current macroblock in the current frame;
[0063] According to the target position of each target compressed macroblock, the multiple target compressed macroblocks are spliced to obtain the compressed frame.
[0064] In a possible implementation method, the acquisition module is specifically used to:
[0065] Acquire the reference frame in a first preset storage space;
[0066] After generating a compressed frame corresponding to the current frame according to the multiple target compressed macroblocks, the method further includes:
[0067] In the first preset storage space, the compressed frame is updated as the reference frame.
[0068] In a third aspect, an embodiment of the present application provides an electronic device, comprising: at least one processor and a memory; the memory stores computer execution instructions; the at least one processor executes the computer execution instructions stored in the memory, so that the at least one processor executes the video compression method described in the first aspect and various possible designs of the first aspect.
[0069] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, in which computer execution instructions are stored. When a processor executes the computer execution instructions, the video compression method described in the first aspect and various possible designs of the first aspect is implemented.
[0070] In a fifth aspect, an embodiment of the present application provides a computer program product, including a computer program. When the computer program is executed by a processor, it implements the video compression method described in the first aspect and various possible designs of the first aspect.
[0071] In a sixth aspect, an embodiment of the present application provides a chip, the chip comprising at least one processor, the processor being used to execute program instructions to implement the video compression method as described in the first aspect and various possible designs of the first aspect.
[0072] The embodiment of the present application provides a video compression method, electronic device, storage medium and program product. When the hardware encoder needs to compress the target video, it can first determine the current frame and reference frame in the target video, and determine multiple current macroblocks in the current frame, and perform downsampling processing on each current macroblock to obtain multiple pre-compressed macroblocks with reduced resolution, and then determine the target reference macroblock corresponding to each pre-compressed macroblock and the reference vector corresponding to each target reference macroblock in the reference frame, so as to determine the target compressed macroblock corresponding to each pre-compressed macroblock, and finally generate a compressed frame after the current frame is compressed based on the multiple target compressed macroblocks. The above method performs motion estimation based on the downsampled current macroblock and the compressed reference frame, reduces the amount of data processing, improves the efficiency of video compression, thereby improving the performance of the hardware encoder and reducing the storage space occupied. BRIEF DESCRIPTION OF THE DRAWINGS
[0073] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present application and, together with the description, serve to explain the principles of the present application.
[0074] Figure 1 A schematic diagram of the system architecture provided for an embodiment of the present application;
[0075] Figure 2 A schematic diagram of a video compression method provided in an embodiment of the present application;
[0076] Figure 3 A schematic diagram of a current frame and multiple current macroblocks provided in an embodiment of the present application;
[0077] Figure 4 A schematic diagram of generating a compressed frame provided in an embodiment of the present application;
[0078] Figure 5 A schematic diagram of a process for determining a target compressed macroblock provided in an embodiment of the present application;
[0079] Figure 6 A schematic diagram of the structure of a video compression device provided in an embodiment of the present application;
[0080] Figure 7 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application.
[0081] The above drawings have shown clear embodiments of the present application, which will be described in more detail later. These drawings and text descriptions are not intended to limit the scope of the present application in any way, but to illustrate the concept of the present application to those skilled in the art by referring to specific embodiments. DETAILED DESCRIPTION
[0082] Exemplary embodiments will be described in detail herein, examples of which are shown in the accompanying drawings. When the following description refers to the drawings, the same numbers in different drawings represent the same or similar elements unless otherwise indicated. The implementations described in the following exemplary embodiments do not represent all implementations consistent with the present application. Instead, they are merely examples of devices and methods consistent with some aspects of the present application as detailed in the appended claims.
[0083] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, storage, use, processing, transmission, provision, disclosure and application of the relevant data comply with the relevant laws, regulations and standards of the relevant countries and regions, take necessary confidentiality measures, do not violate public order and good morals, and provide corresponding operation entrances for users to choose to authorize or refuse.
[0084] In addition, this application involves conducting big data analysis of user information (including but not limited to personal biometrics, identity data, consumption data, asset data, electronic terminal operation data, etc.), and using artificial intelligence technology to make automated decisions, and making technical solutions that have a significant impact on personal rights and interests based on the results of automated decisions. The application provides users with corresponding operation entrances for them to choose to agree or reject the results of automated decisions; if the user chooses to reject, the expert decision-making process will be entered.
[0085] It should be noted that the video compression method, electronic device, storage medium and program product provided in the present application can be used in the computer field, and can also be used in any field other than the computer field. The application field of the video compression method, electronic device, storage medium and program product in the present application is not limited.
[0086] It should be noted that in the embodiments of the present application, certain software, components, models and other existing solutions in the industry may be mentioned, and they should be regarded as exemplary. Their purpose is only to illustrate the feasibility of implementing the technical solution of the present application, but it does not mean that the applicant has or will necessarily use the solution.
[0087] For ease of understanding, the following Figure 1 , the system architecture applicable to the embodiments of the present application is described.
[0088] Figure 1 This is a schematic diagram of the system architecture provided by the embodiment of this application. Figure 1, including a hardware encoder 101. The hardware encoder 101 can be used to compress the original video. Video data can be input into the hardware encoder 101, and the hardware encoder outputs compressed video data after the encoding algorithm. The hardware encoder 101 can be located in an electronic device, for example, the electronic device can refer to a mobile terminal device, a video conferencing device, a monitoring device, etc.
[0089] In the related art, a video can be compressed by a hardware encoder, and during the video compression process, a corresponding reference frame can be set for each current frame. When compressing the current frame, the current frame can be divided into multiple image blocks, and the reference image block corresponding to each image block is determined in the reference frame, and the current frame is compressed based on each image block and its corresponding reference image block. However, in the above process, large computing resources are required to determine the reference image block corresponding to each image block in the reference frame, resulting in low efficiency in video compression. Furthermore, since the hardware encoder requires more resources to achieve video compression, the performance of the hardware encoder is low. In addition, if the size of the reference frame corresponding to any current frame is the same as the size of the current frame, a large storage space is required to store the reference frame, resulting in a waste of storage resources.
[0090] In view of the above technical problems, in an embodiment of the present application, when the hardware encoder needs to compress the target video, it can first determine the current frame and the reference frame in the target video, and determine multiple current macroblocks in the current frame, and perform downsampling processing on each current macroblock to obtain multiple pre-compressed macroblocks with reduced resolution, and then determine the target reference macroblock corresponding to each pre-compressed macroblock and the reference vector corresponding to each target reference macroblock in the reference frame, so as to determine the target compressed macroblock corresponding to each pre-compressed macroblock, and finally generate a compressed frame after the current frame is compressed based on the multiple target compressed macroblocks. The above method performs motion estimation based on the downsampled current macroblock and the compressed reference frame, reduces the amount of data processing, improves the efficiency of video compression, thereby improving the performance of the hardware encoder and reducing storage space occupancy.
[0091] The technical solution of the present application and how the technical solution of the present application solves the above-mentioned technical problems are described in detail below with specific embodiments. The following specific embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments. The embodiments of the present application will be described below in conjunction with the accompanying drawings.
[0092] Figure 2 A schematic diagram of a video compression method provided in an embodiment of the present application. Figure 2 As shown, the method may include the following steps:
[0093] S201. Obtain a current frame and a reference frame of a target video.
[0094] The execution subject of the embodiment of the present application can be a hardware encoder, or a video compression device set in the hardware encoder. The video compression device can be implemented by software, or by a combination of software and hardware.
[0095] The target video may refer to video data that needs to be compressed by the hardware encoder. The target video is composed of multiple frames.
[0096] The current frame may refer to frame data that needs to be compressed in the target video, and the current frame includes multiple current macroblocks. The current frame may be divided into multiple current macroblocks, and the current macroblock may be composed of multiple pixels, for example, the current macroblock may include 8*8 pixels, 16*16 pixels, etc.
[0097] Next, combine Figure 3 , through specific examples, the current frame and the multiple current macroblocks included in the current frame are explained.
[0098] Figure 3 For a schematic diagram of a current frame and multiple current macroblocks provided in an embodiment of the present application, see Figure 3 , including a current frame and multiple current macroblocks, the current frame is divided into 9 current macroblocks, and the macroblock sizes of the 9 current macroblocks are the same.
[0099] The reference frame may refer to a frame before the current frame in the target video, or a frame after the current frame in the target video. The reference frame may refer to compressed frame data stored in the first preset storage space, that is, the size of the reference frame is smaller than the size of the current frame.
[0100] S202: Perform downsampling processing on each current macroblock to obtain multiple pre-compressed macroblocks.
[0101] Downsampling can reduce the resolution of the current macroblock. For example, assuming that the macroblock size of the current macroblock is 16*16 pixels, after downsampling the current macroblock, a pre-compressed macroblock with a macroblock size of 8*8 pixels can be obtained.
[0102] Optionally, the downsampling process may adopt an average pooling method. Specifically, for any current macroblock, the current macroblock may be divided into multiple target areas, and the pixel values in each target area are averaged to obtain a pre-compressed macroblock after the downsampling process.
[0103] For example, assuming that the size of any current macroblock is 16*16 pixels, the current macroblock can be divided into 4*4 target areas, and the average pixel value in each target area is calculated to obtain a pre-compressed macroblock with a size of 8*8 pixels.
[0104] S203: Determine a target reference macroblock corresponding to each pre-compression macroblock in the reference frame, and determine a reference vector corresponding to each target reference macroblock.
[0105] For any pre-compressed macroblock, the target reference macroblock may refer to a macroblock in the reference frame that is closest to the pre-compressed macroblock. The target reference macroblock may be used to predict the content of the pre-compressed macroblock in the current frame.
[0106] The reference vector may refer to a moving direction and distance from a pre-compressed macroblock in a current frame to a target reference macroblock in a reference frame.
[0107] The reference vector may include a horizontal component and a vertical component, wherein the horizontal component may refer to the offset of the target reference macroblock from the pre-compression macroblock in the horizontal direction, and the vertical component may refer to the offset of the target reference macroblock from the pre-compression macroblock in the vertical direction.
[0108] For example, assuming that the coordinate position of any pre-compressed macroblock in the current frame is (x1, y1), and the coordinate position of the corresponding target reference macroblock in the reference frame is (x2, y2), the reference vector can be determined to be (x2-x1, y2-y1).
[0109] The target reference macroblock can be determined in the following manner: determine the macroblock size corresponding to the pre-compressed macroblock, the macroblock size is used to indicate the number of pixels included in the pre-compressed macroblock in the horizontal direction and the number of pixels included in the vertical direction; according to the macroblock size, divide the reference frame into blocks to obtain multiple reference macroblocks, and the size of the reference macroblock is the same as the macroblock size; among the multiple reference macroblocks, determine the target reference macroblock corresponding to each pre-compressed macroblock.
[0110] S204. For any pre-compressed macroblock, determine a target compressed macroblock corresponding to the pre-compressed macroblock according to the pre-compressed macroblock, the target reference macroblock corresponding to the pre-compressed macroblock, and the reference vector.
[0111] The target compressed macroblock may refer to the content of the current frame predicted based on the content in the reference frame.
[0112] The target compressed macroblock may be determined in the following manner: determining a residual between a pre-compressed macroblock and a target reference macroblock; encoding the residual to obtain a target residual; and reconstructing the target residual and a reference vector to obtain a target compressed macroblock.
[0113] The residual may be used to indicate the difference between the pre-compressed macroblock and the target reference macroblock.
[0114] S205. Generate a compressed frame corresponding to the current frame according to multiple target compressed macroblocks.
[0115] The compressed frame may refer to a compressed frame obtained by splicing multiple target compressed macroblocks according to the positions of multiple current macroblocks in the current frame. That is, the compressed frame may refer to a frame obtained by compressing the current frame.
[0116] The compressed frame can be determined in the following manner: for any current macroblock, the target position of the target compressed macroblock corresponding to the current macroblock is determined according to the position of the current macroblock in the current frame; according to the target position of each target compressed macroblock, multiple target compressed macroblocks are spliced to obtain a compressed frame.
[0117] Next, combine Figure 4 , the process of generating compressed frames is explained through specific examples.
[0118] Figure 4 For a schematic diagram of generating a compressed frame provided in an embodiment of the present application, see Figure 4 , including multiple target compressed macroblocks and compressed frames. The multiple target compressed macroblocks are spliced to obtain a compressed frame according to the positions of the corresponding current macroblocks in the current frame.
[0119] After generating the compressed frame corresponding to the current frame, the method further includes updating the compressed frame as a reference frame in the first preset storage space.
[0120] In an embodiment of the present application, when it is necessary to compress a video, the current frame and the reference frame can be first obtained from the target video, and multiple current macroblocks can be determined in the current frame; each current macroblock is downsampled to obtain multiple pre-compressed macroblocks with reduced resolution; the target reference macroblock corresponding to each pre-compressed macroblock is determined in the reference frame, and the reference vector corresponding to each target reference macroblock is determined; for any pre-compressed macroblock, the target compressed macroblock corresponding to the pre-compressed macroblock can be determined according to the pre-compressed macroblock, the target reference macroblock corresponding to the pre-compressed macroblock, and the reference vector; based on the multiple target compressed macroblocks, a compressed frame corresponding to the current frame is generated. In this way, since each current macroblock is downsampled to obtain multiple pre-compressed macroblocks with reduced resolution, the amount of data processed by the hardware encoder is reduced, while ensuring data accuracy, the amount of data processing is reduced, the efficiency of video compression is improved, and the performance of the hardware encoder is improved. In addition, since the compressed frame after compression is updated to the reference frame for storage, the amount of data of the reference frame is reduced, thereby reducing the storage space occupied.
[0121] Based on any of the above embodiments, Figure 5 , the process of determining the target compressed macroblock ( Figure 2 S203-S204 in the embodiment are described in detail.
[0122] Figure 5Schematic diagram of the process of determining the target compressed macroblock provided by the embodiment of the present application. Figure 5 , the method may include:
[0123] S501. For any pre-compressed macroblock, determine a macroblock size corresponding to the pre-compressed macroblock.
[0124] The macroblock size may be used to indicate the number of pixels included in the horizontal direction and the number of pixels included in the vertical direction of the pre-compression macroblock.
[0125] For example, the macroblock size may be 8*8 pixels, that is, the number of pixels included in the horizontal direction is 8, and the number of pixels included in the vertical direction is 8.
[0126] S502: Divide the reference frame into blocks according to the macroblock size to obtain multiple reference macroblocks.
[0127] The size of the reference macroblock is the same as the size of the macroblock. For example, assuming that the size of the pre-compressed macroblock is 8*8 pixels, the size of the reference macroblock is also 8*8 pixels.
[0128] The block processing may refer to dividing the reference frame into a plurality of reference macroblocks, and the size of each reference macroblock is the macroblock size corresponding to the pre-compressed macroblock.
[0129] S503: Determine a motion vector between the pre-compression macroblock and each reference macroblock to obtain multiple motion vectors.
[0130] The motion vector may refer to the direction and distance of movement from the pre-compressed macroblock in the current frame to each reference macroblock in the reference frame.
[0131] The motion vector may include a horizontal component and a vertical component, wherein the horizontal component may refer to the offset of any reference macroblock from the pre-compressed macroblock in the horizontal direction, and the vertical component may refer to the offset of any reference macroblock from the pre-compressed macroblock in the vertical direction.
[0132] For example, assuming that the coordinate position of any pre-compressed macroblock in the current frame is (x1, y1), and the coordinate position of any reference macroblock in the reference frame is (x2, y2), the motion vector can be determined to be (x2-x1, y2-y1).
[0133] S504: For any motion vector, determine the absolute error sum corresponding to the motion vector according to the pre-compressed macroblock and the reference macroblock.
[0134] The sum of absolute value errors can be used to indicate the similarity between the pre-compressed macroblock and the reference macroblock. The smaller the sum of absolute value errors, the smaller the distortion, that is, the higher the similarity between the pre-compressed macroblock and the reference macroblock; the larger the sum of absolute value errors, the greater the distortion, that is, the lower the similarity between the pre-compressed macroblock and the reference macroblock.
[0135] Assume that the pixel value of each pixel position (i, j) in the pre-compressed macroblock is I rec (i, j), the pixel value of each pixel position (i, j) in the reference macroblock is I orig (i, j), the absolute value error and SAD can be expressed as:
[0136]
[0137] S505: Determine the number of bits required for encoding the motion vector according to the motion vector.
[0138] The number of bits required for encoding the motion vector may refer to a certain number of bits allocated by the hardware encoder for encoding the motion vector.
[0139] Assume that the motion vector is V motion , the number of bits after encoding the motion vector is H(V motion ), then the number of bits R(mvd) required to encode the motion vector is:
[0140] R(mvd)=H(V motion )
[0141] Optionally, the motion vector may be encoded by entropy coding.
[0142] S506: Determine a rate-distortion cost according to the absolute value error and the number of bits required for encoding the motion vector.
[0143] The rate-distortion cost can be used to measure the trade-off between compression efficiency and quality loss during the encoding process. The rate-distortion cost can be used to represent the comprehensive cost between the distortion introduced during the frame compression process and the number of bits required. The smaller the rate-distortion cost, the better the frame compression effect; the larger the rate-distortion cost, the worse the frame compression effect.
[0144] Assume that the absolute value error is represented as SAD, the number of bits required to encode the motion vector is represented as R(mvd), and the Lagrangian factor is represented as λ, then the rate distortion cost RDcost is:
[0145] RDcost = SAD + λR (mvd)
[0146] The Lagrangian factor may be used to balance the trade-off between distortion and bit rate. The Lagrangian factor may be a value set in advance by a user. For example, the value of the Lagrangian factor may be 0.1.
[0147] S507 . Determine a reference vector from multiple motion vectors according to the rate-distortion cost corresponding to each motion vector, wherein the rate-distortion cost corresponding to the reference vector is the smallest.
[0148] The reference vector can be determined in the following manner: obtain multiple rate-distortion costs corresponding to each motion vector, arrange the multiple rate-distortion costs in ascending order, determine the smallest rate-distortion cost among the multiple rate-distortion costs as the target rate-distortion cost, determine the motion vector corresponding to the target rate-distortion cost, and determine the motion vector as the reference vector.
[0149] S508: Determine the reference macroblock corresponding to the reference vector as the target reference macroblock.
[0150] S509: Determine a residual between the pre-compression macroblock and the target reference macroblock.
[0151] The residual may refer to the difference in pixel values between a pre-compression macroblock and a target reference macroblock.
[0152] Assume that the pixel value of the pre-compressed macroblock is MB actual , the pixel value of the target reference macroblock is MB prediction , then the residual MB can be determined residual for:
[0153] MB residual = MB actual -MB prediction
[0154] S510: Encode the residual to obtain a target residual.
[0155] The target residual may refer to the difference after the encoding process to reduce the bit rate.
[0156] The encoding process may refer to converting the residual into a smaller data block by quantization or transformation, thereby reducing the bit rate of the residual. For example, the encoding process may refer to discrete cosine transform DCT processing.
[0157] S511 . Reconstruct the target residual and the reference vector to obtain a target compressed macroblock.
[0158] The reconstruction process may refer to using the motion vector and residual information by the decoder to restore the target macroblock.
[0159] The reconstruction process may be performed in the following manner: the decoder performs inverse quantization and inverse transformation processing on the target residual to obtain a decoded residual after decoding; and the decoded residual and the reference vector are combined to obtain a target compressed macroblock.
[0160] Assume that the decoded residual is represented as MB residual , the reference vector is denoted as MV, then the pixel value MB of the target compressed macroblock reconstraction It can be expressed as:
[0161] MB reconstraction= MB residual + MV
[0162] Optionally, after obtaining the target compressed macroblock, loop filtering may be performed on the target compressed macroblock, thereby reducing fast effects and artifacts and improving the accuracy of the target compressed macroblock.
[0163] In an embodiment of the present application, when it is necessary to compress a video, the target compressed macroblock corresponding to each current macroblock in the current frame can be determined first. For any pre-compressed macroblock after downsampling of the current macroblock, the macroblock size corresponding to the pre-compressed macroblock is determined, and according to the macroblock size, the reference frame is divided into blocks to obtain multiple reference macroblocks; the motion vector between the pre-compressed macroblock and each reference macroblock is determined to obtain multiple motion vectors; for any motion vector, the absolute value error and the number of bits required for encoding the motion vector are determined according to the pre-compressed macroblock and the reference macroblock; the rate-distortion cost is determined according to the absolute value error and the number of bits required for encoding the motion vector; according to the rate-distortion cost corresponding to each motion vector, the reference vector with the smallest rate-distortion cost is determined among the multiple motion vectors; the reference macroblock corresponding to the reference vector is determined as the target reference macroblock, the residual between the pre-compressed macroblock and the target reference macroblock is determined, the residual is encoded to obtain the target residual, and the target residual and the reference vector are reconstructed to obtain the target compressed macroblock. In the above process, since each current macroblock is downsampled, multiple pre-compressed macroblocks with reduced resolution are obtained, thereby reducing the amount of data processed by the hardware encoder. While ensuring data accuracy, the amount of data processing is reduced, the efficiency of video compression is improved, and thus the performance of the hardware encoder is improved. In addition, since the compressed frame after compression is updated to a reference frame for storage, the data amount of the reference frame is reduced, thereby reducing the storage space occupancy.
[0164] Figure 6 This is a schematic diagram of the structure of a video compression device provided in an embodiment of the present application. Figure 6 The video compression device 10 includes: an acquisition module 11, a down-sampling module 12, a first determination module 13, a second determination module 14 and a generation module 15, wherein:
[0165] The acquisition module 11 is used to acquire a current frame and a reference frame of a target video, wherein the current frame includes a plurality of current macroblocks;
[0166] The downsampling module 12 is used to perform downsampling processing on each current macroblock to obtain multiple pre-compressed macroblocks;
[0167] The first determination module 13 is used to determine a target reference macroblock corresponding to each pre-compressed macroblock in the reference frame, and to determine a reference vector corresponding to each target reference macroblock, wherein the size of the reference frame is smaller than the size of the current frame;
[0168] The second determination module 14 is used to determine, for any pre-compressed macroblock, a target compressed macroblock corresponding to the pre-compressed macroblock according to the pre-compressed macroblock, the target reference macroblock corresponding to the pre-compressed macroblock and the reference vector;
[0169] The generating module 15 is used to generate a compressed frame corresponding to the current frame according to the multiple target compressed macroblocks.
[0170] A video compression device provided in an embodiment of the present application can execute the technical solution shown in the above method embodiment, and its implementation principle and beneficial effects are similar, which will not be repeated here.
[0171] In a possible implementation method, the first determining module 13 is specifically configured to:
[0172] Determine a macroblock size corresponding to the pre-compressed macroblock, where the macroblock size is used to indicate the number of pixels included in the horizontal direction and the number of pixels included in the vertical direction of the pre-compressed macroblock;
[0173] According to the macroblock size, the reference frame is divided into blocks to obtain a plurality of reference macroblocks, wherein the size of the reference macroblock is the same as the size of the macroblock;
[0174] Among the multiple reference macroblocks, a target reference macroblock corresponding to each pre-compressed macroblock is determined.
[0175] In a possible implementation method, the first determining module 13 is specifically configured to:
[0176] Determine a motion vector between the pre-compressed macroblock and each reference macroblock to obtain a plurality of motion vectors;
[0177] For any motion vector, determining a rate-distortion cost corresponding to the motion vector according to the pre-compressed macroblock and a reference macroblock corresponding to the motion vector;
[0178] Determine a reference vector among the multiple motion vectors according to the rate-distortion cost corresponding to each motion vector, wherein the rate-distortion cost corresponding to the reference vector is the smallest;
[0179] The reference macroblock corresponding to the reference vector is determined as the target reference macroblock.
[0180] In a possible implementation method, the first determining module 13 is specifically configured to:
[0181] Determine the sum of absolute value errors corresponding to the motion vector according to the pre-compressed macroblock and the reference macroblock;
[0182] Determining, based on the motion vector, the number of bits required to encode the motion vector;
[0183] Determine a bit rate cost corresponding to the motion vector according to the pre-compressed macroblock, the number of bits and the motion vector;
[0184] The rate-distortion cost is determined according to the bit rate cost and the sum of absolute value errors.
[0185] In a possible implementation method, the second determining module 14 is specifically configured to:
[0186] determining a residual between the pre-compressed macroblock and the target reference macroblock;
[0187] Performing encoding processing on the residual to obtain a target residual;
[0188] The target residual and the reference vector are reconstructed to obtain the target compressed macroblock.
[0189] In a possible implementation method, the generating module 15 is specifically used to:
[0190] For any current macroblock, determining a target position of a target compressed macroblock corresponding to the current macroblock according to the position of the current macroblock in the current frame;
[0191] According to the target position of each target compressed macroblock, the multiple target compressed macroblocks are spliced to obtain the compressed frame.
[0192] In a possible implementation method, the acquisition module 11 is specifically used to:
[0193] Acquire the reference frame in a first preset storage space;
[0194] After generating a compressed frame corresponding to the current frame according to the multiple target compressed macroblocks, the method further includes:
[0195] In the first preset storage space, the compressed frame is updated as the reference frame.
[0196] A video compression device provided in an embodiment of the present application can execute the technical solution shown in the above method embodiment, and its implementation principle and beneficial effects are similar, which will not be repeated here.
[0197] Figure 7 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present application. Figure 7As shown, the electronic device 20 may include: a transceiver 21 , a processor 22 , and a memory 23 .
[0198] The processor 22 executes the computer execution instructions stored in the memory, so that the processor 22 executes the scheme in the above embodiment. The processor 22 can be a general-purpose processor, including a central processing unit CPU, a network processor (NP), etc.; it can also be a digital signal processor DSP, an application-specific integrated circuit ASIC, a field programmable gate array FPGA or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components.
[0199] The memory 23 is connected to the processor 22 via a system bus and completes communication between them. The memory 23 is used to store computer program instructions.
[0200] The transceiver 21 may be used to obtain tasks to be executed and configuration information of the tasks to be executed.
[0201] The system bus can be a peripheral component interconnect (PCI) bus or an extended industry standard architecture (EISA) bus, etc. The system bus can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, only one thick line is used in the figure, but it does not mean that there is only one bus or one type of bus. The transceiver is used to realize the communication between the database access device and other computers (such as clients, read-write libraries, and read-only libraries). The memory may include random access memory (RAM) and may also include non-volatile memory.
[0202] The electronic device provided in the embodiment of the present application may be the terminal device of the above embodiment.
[0203] An embodiment of the present application also provides a chip for running instructions, which is used to execute the technical solution of the video compression method in the above embodiment.
[0204] An embodiment of the present application further provides a computer-readable storage medium, in which computer instructions are stored. When the computer instructions are executed on a computer, the computer executes the technical solution of the video compression method in the above embodiment.
[0205] An embodiment of the present application also provides a computer program product, which includes a computer program stored in a computer-readable storage medium. At least one processor can read the computer program from the computer-readable storage medium. When at least one processor executes the computer program, the technical solution of the video compression method in the above embodiment can be implemented.
[0206] An embodiment of the present application provides a chip, which includes at least one processor, and the processor is used to execute program instructions to implement the technical solution of the video compression method described in the first aspect and various possible designs of the first aspect.
[0207] In the several embodiments provided in the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are only schematic, for example, the division of modules is only a logical function division, and there may be other division methods in actual implementation, such as multiple modules can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be an indirect coupling or communication connection through some interfaces, devices or modules, which can be electrical, mechanical or other forms.
[0208] The modules described as separate components may or may not be physically separated, and the components shown as modules may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the modules may be selected according to actual needs to implement the solution of this embodiment.
[0209] In addition, each functional module in each embodiment of the present application can be integrated into one processing unit, or each module can exist physically separately, or two or more modules can be integrated into one unit. The above-mentioned module-composed unit can be implemented in the form of hardware or in the form of hardware plus software functional units.
[0210] The above-mentioned integrated module implemented in the form of a software function module can be stored in a computer-readable storage medium. The above-mentioned software function module is stored in a storage medium, including a number of instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) or a processor to perform some steps of the methods of various embodiments of the present application.
[0211] It should be understood that the above processor can be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), etc. A general-purpose processor can be a microprocessor or any conventional processor. The steps of the method disclosed in the invention can be directly embodied as being executed by a hardware processor, or can be executed by a combination of hardware and software modules in the processor.
[0212] The memory may include a high-speed RAM memory, and may also include a non-volatile storage NVM, such as at least one disk memory, and may also be a USB flash drive, a mobile hard disk, a read-only memory, a magnetic disk or an optical disk, etc.
[0213] The bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, the bus in the drawings of this application is not limited to only one bus or one type of bus.
[0214] The above storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk or optical disk. The storage medium can be any available medium that can be accessed by a general or special purpose computer.
[0215] An exemplary storage medium is coupled to a processor so that the processor can read information from the storage medium and write information to the storage medium. Of course, the storage medium can also be a component of the processor. The processor and the storage medium can be located in an application specific integrated circuit (ASIC). Of course, the processor and the storage medium can also exist as discrete components in an electronic control unit or a main control device.
[0216] Those skilled in the art can understand that all or part of the steps of implementing the above-mentioned method embodiments can be completed by hardware related to program instructions. The aforementioned program can be stored in a computer-readable storage medium. When the program is executed, the steps of the above-mentioned method embodiments are executed; and the aforementioned storage medium includes: ROM, RAM, disk or optical disk and other media that can store program codes.
[0217] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit it. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or replace some or all of the technical features therein with equivalents. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present application.
Claims
1. A video compression method, characterized in that: include: Acquire a current frame and a reference frame of a target video, wherein the current frame includes a plurality of current macroblocks; Down-sampling each current macroblock to obtain multiple pre-compressed macroblocks; Determining a target reference macroblock corresponding to each pre-compressed macroblock in the reference frame, and determining a reference vector corresponding to each target reference macroblock, wherein the size of the reference frame is smaller than the size of the current frame; For any pre-compressed macroblock, determine a target compressed macroblock corresponding to the pre-compressed macroblock according to the pre-compressed macroblock, a target reference macroblock corresponding to the pre-compressed macroblock, and a reference vector; A compressed frame corresponding to the current frame is generated according to the multiple target compressed macroblocks.
2. The method according to claim 1, characterized in that Determining a target reference macroblock corresponding to each pre-compressed macroblock in the reference frame includes: Determine a macroblock size corresponding to the pre-compressed macroblock, where the macroblock size is used to indicate the number of pixels included in the horizontal direction and the number of pixels included in the vertical direction of the pre-compressed macroblock; According to the macroblock size, the reference frame is divided into blocks to obtain a plurality of reference macroblocks, wherein the size of the reference macroblock is the same as the size of the macroblock; Among the multiple reference macroblocks, a target reference macroblock corresponding to each pre-compressed macroblock is determined.
3. The method according to claim 2, characterized in that For any pre-compressed macroblock; Determining, among the multiple reference macroblocks, a target reference macroblock corresponding to the pre-compressed macroblock includes: Determine a motion vector between the pre-compressed macroblock and each reference macroblock to obtain a plurality of motion vectors; For any motion vector, determining a rate-distortion cost corresponding to the motion vector according to the pre-compressed macroblock and a reference macroblock corresponding to the motion vector; Determine a reference vector among the multiple motion vectors according to the rate-distortion cost corresponding to each motion vector, wherein the rate-distortion cost corresponding to the reference vector is the smallest; The reference macroblock corresponding to the reference vector is determined as the target reference macroblock.
4. The method according to claim 3, characterized in that Determining a rate-distortion cost corresponding to the motion vector according to the pre-compressed macroblock and a reference macroblock corresponding to the motion vector includes: Determine the sum of absolute value errors corresponding to the motion vector according to the pre-compressed macroblock and the reference macroblock; Determining, based on the motion vector, the number of bits required to encode the motion vector; Determine a bit rate cost corresponding to the motion vector according to the pre-compressed macroblock, the number of bits and the motion vector; The rate-distortion cost is determined according to the bit rate cost and the sum of absolute value errors.
5. The method according to any one of claims 1 to 4, characterized in that: Determining a target compressed macroblock corresponding to the pre-compressed macroblock according to the pre-compressed macroblock, a target reference macroblock corresponding to the pre-compressed macroblock, and a reference vector, comprises: determining a residual between the pre-compressed macroblock and the target reference macroblock; Performing encoding processing on the residual to obtain a target residual; The target residual and the reference vector are reconstructed to obtain the target compressed macroblock.
6. The method according to any one of claims 1 to 5, characterized in that: Generating a compressed frame corresponding to the current frame according to the multiple target compressed macroblocks, comprising: For any current macroblock, determining a target position of a target compressed macroblock corresponding to the current macroblock according to the position of the current macroblock in the current frame; According to the target position of each target compressed macroblock, the multiple target compressed macroblocks are spliced to obtain the compressed frame.
7. The method according to any one of claims 1 to 5, characterized in that: Acquiring a reference frame of the target video includes: Acquire the reference frame in a first preset storage space; After generating a compressed frame corresponding to the current frame according to the multiple target compressed macroblocks, the method further includes: In the first preset storage space, the compressed frame is updated as the reference frame.
8. A video compression device, characterized in that: include: An acquisition module, a down-sampling module, a first determination module, a second determination module and a generation module, wherein: The acquisition module is used to acquire a current frame and a reference frame of a target video, wherein the current frame includes a plurality of current macroblocks; The downsampling module is used to perform downsampling processing on each current macroblock to obtain multiple pre-compressed macroblocks; The first determination module is used to determine a target reference macroblock corresponding to each pre-compressed macroblock in the reference frame, and determine a reference vector corresponding to each target reference macroblock, wherein the size of the reference frame is smaller than the size of the current frame; The second determination module is used to determine, for any pre-compressed macroblock, a target compressed macroblock corresponding to the pre-compressed macroblock according to the pre-compressed macroblock, a target reference macroblock corresponding to the pre-compressed macroblock, and a reference vector; The generation module is used to generate a compressed frame corresponding to the current frame according to the multiple target compressed macroblocks.
9. An electronic device, characterized in that: include: A processor, and a memory communicatively connected to the processor; The memory stores computer-executable instructions; The processor executes the computer-executable instructions stored in the memory to implement the method according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer-executable instructions, which are used to implement the method according to any one of claims 1 to 7 when executed by a processor.
11. A computer program product, characterized in that The invention comprises a computer program, which implements the method according to any one of claims 1 to 7 when being executed by a processor.
12. A chip, comprising at least one processor, wherein the processor is configured to execute program instructions to perform the method according to any one of claims 1 to 7.