Contour-based inter-frame prediction method, system, device and storage medium

Through the contour-based inter-frame prediction method, reference blocks are selected and motion vectors, contour initial points and color values ​​are encoded, which solves the problem of poor encoding performance in contour-changing scenes in the existing technology and achieves reduced temporal redundancy and bit rate savings in video encoding.

CN117979022BActive Publication Date: 2025-09-16UNIV OF SCI & TECH OF CHINA
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410259982.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-03-07
Publication Date
2025-09-16
Estimated Expiration
2044-03-07

AI Technical Summary

Technical Problem

Existing video coding technologies have difficulty in fully utilizing the temporal correlation between frames when processing scenes with contour changes, resulting in poor coding performance.

Method used

Through the contour-based inter-frame prediction method, a reference block with contour is selected, and the motion vector, contour initial point position and color value are encoded to form a code stream. The code stream is then encoded and decoded in combination with the contour information to reconstruct the image block.

Benefits of technology

Effectively reduce the temporal redundancy of video encoding, shorten encoding and decoding time, save video bit rate, and improve encoding performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117979022B_ABST
    Figure CN117979022B_ABST
Patent Text Reader

Abstract

The present invention discloses a contour-based inter-frame prediction method, system, device, and storage medium. These are one-to-one corresponding schemes, which mainly include: selecting a reference block with a contour and encoding the corresponding motion vector; when the contour of the current coding block coincides with the contour of the reference block, encoding a flag to indicate whether the next contour direction of the current coding block is the same as the next contour direction of the reference block; decoding the motion vector to obtain the corresponding reference block; when the contour of the current coding block coincides with the contour of the reference block, decoding a flag to indicate whether the next contour direction is the same as the next contour of the reference block. The above scheme can effectively reduce temporal redundancy in video coding; experiments have shown that the scheme provided by the present invention can reduce encoding and decoding time and save video bit rate compared to existing schemes.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of video coding technology, and in particular to a contour-based inter-frame prediction method, system, device and storage medium. Background Art

[0002] Inter-frame prediction is one of the most important technologies in video coding. It can be used to remove temporal redundancy between adjacent video frames. In almost all contemporary video coding standards, such as H.265 (High Efficiency Video Coding) and H.266 (Versatile Video Coding), block-based motion estimation and motion compensation are widely used for inter-frame prediction. However, for scenes with contour changes, it is difficult to solve them using motion models. Currently, they can only be processed through residual coding methods. There is also a strong temporal correlation between contours. Current residual coding methods cannot fully utilize this temporal correlation, resulting in poor coding performance when the target contour changes.

[0003] In view of this, the present invention is proposed. Summary of the Invention

[0004] The purpose of the present invention is to provide a contour-based inter-frame prediction method, system, device and storage medium, which can reduce the temporal redundancy of video encoding and improve encoding performance.

[0005] The purpose of the present invention is achieved through the following technical solutions:

[0006] A contour-based inter-frame prediction method, comprising:

[0007] Contour-based coding: Select a reference block with a contour and encode the corresponding motion vector; encode the initial point position of the contour of the current coding block; combine the encoded initial point position of the contour and the contour direction of the reference block to encode the contour information of the current coding block; encode the color value of the current coding block; and combine the encoded motion vector, initial point position of the contour, contour information of the current coding block, and color value to form a code stream.

[0008] The contour-based decoding part: decodes the motion vector in the code stream to obtain the corresponding reference block; decodes the initial point position of the contour in the code stream to locate the contour; decodes the contour information in the code stream based on the reference block obtained by decoding and the located contour; decodes the color value in the code stream; and reconstructs the image block by combining the contour information and color value obtained by decoding.

[0009] A contour-based inter-frame prediction system, comprising:

[0010] The contour-based encoding module is used to select a reference block with a contour and encode the corresponding motion vector; encode the initial point position of the contour of the current coding block; combine the encoded initial point position of the contour and the contour direction of the reference block to encode the contour information of the current coding block; encode the color value of the current coding block; and combine the encoded motion vector, initial point position of the contour, contour information of the current coding block, and color value to form a code stream.

[0011] The contour-based decoding module is used to decode the motion vector in the code stream to obtain the corresponding reference block; decode the initial point position of the contour in the code stream and locate the contour; decode the contour information in the code stream based on the decoded reference block and the located contour; decode the color value in the code stream; and reconstruct the image block by combining the decoded contour information and color value.

[0012] A processing device comprising: one or more processors; a memory for storing one or more programs;

[0013] When the one or more programs are executed by the one or more processors, the one or more processors implement the aforementioned method.

[0014] A readable storage medium stores a computer program, which implements the aforementioned method when the computer program is executed by a processor.

[0015] It can be seen from the technical solution provided by the present invention that the contour-based inter-frame prediction method can effectively reduce the temporal redundancy of video encoding; experiments show that the solution provided by the present invention can reduce encoding and decoding time and save video bit rate compared with existing solutions. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0017] Figure 1 A flowchart of a contour-based inter-frame prediction method provided by an embodiment of the present invention;

[0018] Figure 2 A flowchart of the encoding part provided by an embodiment of the present invention;

[0019] Figure 3 A schematic diagram illustrating the principle of a contour-based inter-frame prediction method provided by an embodiment of the present invention;

[0020] Figure 4A flowchart of the decoding part provided by an embodiment of the present invention;

[0021] Figure 5 A schematic diagram of a contour-based inter-frame prediction system provided by an embodiment of the present invention;

[0022] Figure 6 A schematic diagram of a processing device provided in an embodiment of the present invention. DETAILED DESCRIPTION

[0023] The following is a clear and complete description of the technical solutions in the embodiments of the present invention, in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.

[0024] First, the following terms may be used in this article:

[0025] The term “and / or” means that either or both of them can be realized at the same time. For example, X and / or Y includes both “X” or “Y” and “X and Y”.

[0026] The terms "include," "comprises," "contains," "has," or other similar expressions should be interpreted as non-exclusive. For example, "including certain technical features (such as raw materials, components, ingredients, carriers, dosage forms, materials, dimensions, parts, components, mechanisms, devices, steps, procedures, methods, reaction conditions, processing conditions, parameters, algorithms, signals, data, products, or manufactured articles, etc.) should be interpreted as including not only the technical features explicitly listed, but also other technical features known in the art that are not explicitly listed.

[0027] The following describes in detail the contour-based inter-frame prediction method, system, device, and storage medium provided by the present invention. Any information not described in detail in the embodiments of the present invention is prior art known to those skilled in the art. Where specific conditions are not specified in the embodiments of the present invention, the procedures are performed in accordance with conventional conditions in the art or the conditions recommended by the manufacturer. Instruments used in the embodiments of the present invention, where the manufacturer is not specified, are all commercially available conventional products.

[0028] Example 1

[0029] The embodiment of the present invention provides a contour-based inter-frame prediction method, such as Figure 1 As shown, it mainly includes a contour-based encoding part and a contour-based decoding part.

[0030] like Figure 2As shown, the contour-based encoding part mainly includes:

[0031] Step 21: Select a reference block with an outline and encode the corresponding motion vector.

[0032] In this embodiment of the present invention, the coding block that can be selected for inter-frame prediction (referred to as the current coding block) must have only one contour, with different color values ​​(pixel values) on both sides of the contour. A reference block with only one contour must be selected. If multiple blocks meet these conditions, RDO (Rate-Distortion Optimization) can be used to select the optimal block for encoding as the reference block. When using RDO to select a reference block, the bit rate R must be calculated.

[0033] Those skilled in the art will appreciate that RDO compares the values ​​of different blocks using the formula D+λR to select the reference block with the lowest value, where D is the distortion and λ is a hyperparameter. When used for lossless coding, D=0, and R is calculated using a pre-coding method, i.e., only encoding without writing into the bitstream. The specific calculation method can be implemented with reference to conventional technologies and will not be described in detail in the present invention.

[0034] After determining the reference block, the motion vector needs to be encoded. For example, only the previous frame can be used for prediction, and the horizontal and vertical motion vectors of the previous frame can be encoded.

[0035] Furthermore, a contour encoding method must be pre-selected to facilitate encoding the contour information of the current coding block in the subsequent process. Because contour information is used to reconstruct the video, there are many ways to encode contour information. For example, the chain coding method (3OT) can be used, and the Markov model can be used as the context model to efficiently encode the 3OT chain code. The chain code expresses contour information by indicating the contour direction.

[0036] Step 22: Encode the initial point position of the contour of the current coding block.

[0037] The initial point of the contour is used to locate the contour, and then the contour information can be expressed through the contour direction. The initial point of the contour is the first point where the color value changes counterclockwise from the top left vertex.

[0038] Exemplarily, the method for encoding the position of the initial point of the contour may be to directly encode the number of grids in the counterclockwise direction along the boundary of the coding block from the upper left corner.

[0039] Step 23: Encode the contour information of the current coding block based on the encoded contour initial point position and the contour direction of the reference block.

[0040] like Figure 3As shown in the figure, it is a schematic diagram of the principle of contour-based inter-frame prediction. The lines with arrows in the reference block and the current coding block are contours, and two different color values ​​in the block are separated by a contour; the contour direction is represented by a 3OT chain code: 0 means the same as the previous direction; 1 means the direction is different from the previous direction and different from the previous transfer direction (left or right); 2 means the direction is different from the previous direction and the same as the previous transfer direction.

[0041] Figure 3 In the encoding process, the reference block has a certain correlation with the contour of the current coding block. The main process of encoding the contour information of the current coding block is as follows: (1) The contour of the current coding block is located based on the position of the encoded contour initial point; (2) If the contour of the current coding block coincides with the contour of the reference block, a flag is encoded to indicate whether the next contour direction of the current coding block is the same as the next contour direction of the reference block. Specifically: if the encoded flag indicates that the next contour direction of the current coding block is the same as the next contour direction of the reference block, there is no need to encode the next contour direction; if the next contour direction of the current coding block is different from the next contour direction of the reference block, the contour direction is directly encoded. (3) If the contour of the current coding block does not coincide with the contour of the reference block, the contour direction is directly encoded. (4) When encoding the contour information of the current coding block, the contour direction is encoded in directional order from the contour initial point.

[0042] Step 24: Encode the color value of the current coding block.

[0043] To reconstruct the current coding block, it is also necessary to encode the two color values ​​on both sides of the outline. For example, the two color values ​​of the current coding block can be predicted using the color values ​​on both sides of the reference block. The residual between the color value of the reference block on the same side and the predicted color value of the current coding block is encoded to complete the encoding of the color value of the current coding block.

[0044] The motion vector, initial point position of the contour, contour information of the current coding block and color value of the above encoding are combined to form a code stream, which is transmitted to the decoding end for subsequent decoding.

[0045] It should be noted that the serial numbers of steps 21 to 24 above are mainly used to distinguish different steps and do not represent the execution order of the steps. Those skilled in the art can determine the logical relationship between the steps based on the specific content of the steps; for example, steps 21 and 24 can be executed in any order or simultaneously; and steps 22 and 23 have a logical order, that is, step 22 is executed first, and then step 23; the execution order of the steps is not described here, and can be determined by those skilled in the art based on the specific content of the steps.

[0046] like Figure 4 As shown, the contour-based decoding part mainly includes:

[0047] Step 41: Decode the motion vector in the code stream to obtain the corresponding reference block.

[0048] In an embodiment of the present invention, a reference block can be obtained by decoding a motion vector in a bitstream. For example, the motion vectors in the horizontal and vertical directions are decoded to obtain a reference block of a previous frame.

[0049] Step 42: Decode the initial point position of the contour in the code stream to locate the contour.

[0050] The contour initial point can be obtained by decoding the position information in the code stream to locate the contour. For example, the position quantity is decoded and the contour initial point is obtained by following the upper left corner vertex in a counterclockwise direction along the coding block boundary.

[0051] Furthermore, it is necessary to select a contour encoding method corresponding to the encoding part so as to facilitate decoding of the contour information in the code stream in the subsequent process.

[0052] Step 43: Decode the contour information in the code stream according to the reference block obtained by decoding and the located contour.

[0053] The main process is as follows: (1) Based on the reference block obtained by decoding and the positioned contour, determine whether the contour of the current coding block coincides with the contour of the reference block; (2) If the contour of the current coding block coincides with the contour of the reference block, decode a flag from the contour information to indicate whether the next direction of the contour is the same as the next contour direction of the reference block. Specifically: according to the position of the decoding flag bit located by the reference block, if the next contour direction of the current coding block is the same as the next contour direction of the reference block, there is no need to decode the next contour direction; if the next contour direction of the current coding block is different from the next contour direction of the reference block, directly decode the contour direction; (3) If the contour of the current coding block does not coincide with the contour of the reference block, directly decode the contour direction. (4) During decoding, decode the contour direction in directional order from the initial point of the contour to obtain the contour information. Step 44, decode the color value in the code stream.

[0054] Step 45: Reconstruct the current coding block by combining the contour information and color values ​​obtained through decoding.

[0055] Exemplarily, the two color values ​​of the coding block are predicted using the color values ​​on both sides of the reference block, and the decoding residual is added to the predicted value to obtain the color value.

[0056] Similarly, the serial numbers of steps 41 to 45 are used to distinguish different steps and do not represent the execution order of the steps. The execution order of the steps is not described here and can be determined by those skilled in the art based on the specific content of the steps.

[0057] The above describes the encoding and decoding process of a single coding block. The reconstruction of the entire image and related videos can be completed based on the same process.

[0058] It should be noted that the specific encoding and decoding methods involved in the above solutions can be implemented with reference to conventional technologies and will not be elaborated in the present invention.

[0059] The main advantages of the above scheme in the embodiment of the present invention are as follows: since the contours between adjacent frames have a certain similarity, and encoding requires this contour information, the contour of the current frame can be referenced in the previous frame to reduce the bit rate. Therefore, the above scheme provided in the embodiment of the present invention can effectively reduce the time domain redundancy of video encoding, reduce encoding and decoding time, and save video bit rate.

[0060] In order to illustrate the performance of the present invention, relevant test experiments were also carried out for verification.

[0061] Test conditions: 1) Test sequences: The first six sequences from the validation set of the semantic segmentation video dataset VSPW. The sequence frame rate is 15 fps. Table 1 shows the detailed test sequence information. 2) Evaluation metrics: Byte count and encoding / decoding time.

[0062] Table 1: Details of the test sequence

[0063] Test sequence Resolution Frame rate Camera Movement 112 1280x720 136 still 127 1280x720 131 Slowly zoom in 231 1280x720 136 Slow movement 1296 1920x1080 45 still 1643 1920x1080 45 Vigorous exercise 2097 1920x1080 45 Vigorous exercise

[0064] Table 2 shows the results of the present invention and the SCM-7.0 method on a test sequence. SCM-7.0 is the reference software for the screen extension mode of the HEVC video coding standard. Intra represents all intra mode, in which each video frame uses intra-frame coding, meaning only the current video frame can be used as a reference. LDP mode is a low-latency mode, in which each video frame can only use the current or previous video frame as a reference.

[0065] Table 2: Test results

[0066]

[0067]

[0068] As can be seen from the test results shown in Table 2, the encoding and decoding time of the present invention is significantly less than that of the SCM-7.0 method. The encoding performance of the present invention significantly surpasses that of SCM-7.0. Specifically, when compared with SCM-7.0 LDP, the present invention saves 39.38% of the bit rate.

[0069] Through the description of the above embodiments, those skilled in the art will clearly understand that the above embodiments can be implemented through software or by using software plus a necessary general-purpose hardware platform. Based on this understanding, the technical solutions of the above embodiments can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (such as a CD-ROM, a USB flash drive, a mobile hard disk, etc.) and includes a number of instructions for causing a computer device (such as a personal computer, a server, or a network device) to execute the methods described in the various embodiments of the present invention.

[0070] Example 2

[0071] The present invention also provides a contour-based inter-frame prediction system, which is mainly used to implement the method provided in the above embodiment, such as Figure 5 As shown, the system mainly includes:

[0072] The contour-based encoding module is used to select a reference block with a contour and encode the corresponding motion vector; encode the initial point position of the contour of the current coding block; combine the encoded initial point position of the contour and the contour direction of the reference block to encode the contour information of the current coding block; encode the color value of the current coding block; and combine the encoded motion vector, initial point position of the contour, contour information of the current coding block, and color value to form a code stream.

[0073] The contour-based decoding module is used to decode the motion vector in the code stream to obtain the corresponding reference block; decode the initial point position of the contour in the code stream and locate the contour; decode the contour information in the code stream based on the decoded reference block and the located contour; decode the color value in the code stream; and reconstruct the image block by combining the decoded contour information and color value.

[0074] The specific processing details involved in the above two modules have been introduced in the previous embodiment 1, so they will not be repeated here.

[0075] Those skilled in the art will clearly understand that for the convenience and brevity of description, only the division of the above-mentioned functional modules is used as an example. In actual applications, the above-mentioned functions can be assigned to different functional modules as needed, that is, the internal structure of the system can be divided into different functional modules to complete all or part of the functions described above.

[0076] Example 3

[0077] The present invention also provides a processing device, such as Figure 6As shown, it mainly includes: one or more processors; a memory for storing one or more programs; wherein, when the one or more programs are executed by the one or more processors, the one or more processors implement the method provided by the aforementioned embodiment.

[0078] Furthermore, the processing device further includes at least one input device and at least one output device; in the processing device, the processor, memory, input device, and output device are connected via a bus.

[0079] In the embodiment of the present invention, the specific types of the memory, input device, and output device are not limited; for example:

[0080] The input device can be a touch screen, image acquisition device, physical button or mouse;

[0081] The output device may be a display terminal;

[0082] The memory may be a random access memory (RAM) or a non-volatile memory, such as a disk memory.

[0083] Example 4

[0084] The present invention also provides a readable storage medium storing a computer program, which implements the method provided in the above embodiment when the computer program is executed by a processor.

[0085] In the embodiments of the present invention, the computer-readable storage medium may be provided in the aforementioned processing device, for example, as a memory in the processing device. Alternatively, the computer-readable storage medium may be a USB flash drive, a removable hard drive, a read-only memory (ROM), a magnetic disk, or an optical disk, among other media capable of storing program code.

[0086] The above description is merely a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in the present invention should be included in the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be based on the scope of protection of the claims.

Claims

1. A contour-based inter-frame prediction method, characterized in that: include: Contour-based encoding: selects reference blocks with contours and encodes the corresponding motion vectors; Encode the initial point position of the contour of the current coding block; Encode the contour information of the current coding block by combining the encoded contour initial point position and the contour direction of the reference block; Encode the color value of the current coding block; combine the above-encoded motion vector, contour initial point position, contour information of the current coding block and the color value to form a code stream; wherein, encoding the contour information of the current coding block in combination with the encoded contour initial point position and the contour direction of the reference block includes: locating the contour of the current coding block based on the encoded contour initial point position; if the contour of the current coding block coincides with the contour of the reference block, encoding a flag bit to indicate whether the next contour direction of the current coding block is the same as the next contour direction of the reference block; if the encoded flag bit indicates that the next contour direction of the current coding block is the same as the next contour direction of the reference block, then there is no need to encode the next contour direction; if the next contour direction of the current coding block is different from the next contour direction of the reference block, directly encode the contour direction; if the contour of the current coding block does not coincide with the contour of the reference block, directly encode the contour direction; when encoding the contour information of the current coding block, encode the contour direction in directional order from the contour initial point; The contour-based decoding part: decodes the motion vector in the code stream to obtain the corresponding reference block; decodes the initial point position of the contour in the code stream to locate the contour; decodes the contour information in the code stream based on the reference block obtained by decoding and the located contour; decodes the color value in the code stream; and reconstructs the image block by combining the contour information and color value obtained by decoding.

2. The contour-based inter-frame prediction method according to claim 1, wherein: Also includes: In the encoding part, a contour encoding method is preselected to encode the contour information of the current encoding block; In the decoding part, the contour encoding method corresponding to the encoding part is used to decode the contour information in the code stream.

3. The contour-based inter-frame prediction method according to claim 1, wherein: The selecting of the reference block with the contour includes: A reference block with only one contour is selected. When multiple blocks meet the conditions, the block that is optimal for encoding is selected as the reference block through rate-distortion optimization.

4. The contour-based inter-frame prediction method according to claim 1, wherein: The encoding of the initial point position of the contour of the current coding block includes: The initial point of the contour is the first point whose color value changes counterclockwise from the upper left vertex. The encoded contour initial point position is used to locate the contour.

5. The contour-based inter-frame prediction method according to claim 1, wherein: The encoding of the color value of the current coding block includes: The current coding block has only one outline, and the color values ​​on both sides of the outline are different; the color values ​​on both sides of the reference block are used to predict the two color values ​​of the current coding block, and the residual between the color value of the reference block on the same side and the predicted color value of the current coding block is encoded to complete the encoding of the color value of the current coding block.

6. The contour-based inter-frame prediction method according to claim 1, wherein: The contour information in the decoded code stream according to the reference block obtained by decoding and the positioned contour includes: According to the reference block obtained by decoding and the positioned contour, determine whether the contour of the current coding block coincides with the contour of the reference block; If the contour of the current coding block coincides with the contour of the reference block, the position of the decoding flag is located according to the reference block. If the next contour direction of the current coding block is the same as the next contour direction of the reference block, there is no need to decode the next contour direction; if the next contour direction of the current coding block is different from the next contour direction of the reference block, the contour direction is directly decoded; If the contour of the current coding block does not coincide with the contour of the reference block, the contour direction is directly decoded; During decoding, the contour direction is decoded in directional order from the initial point of the contour, and the contour information is obtained by decoding.

7. A contour-based inter-frame prediction system, characterized in that The method for implementing any one of claims 1 to 6 comprises: The contour-based encoding module is used to select a reference block with a contour and encode the corresponding motion vector; encode the initial point position of the contour of the current coding block; combine the encoded initial point position of the contour and the contour direction of the reference block to encode the contour information of the current coding block; encode the color value of the current coding block; and combine the encoded motion vector, initial point position of the contour, contour information of the current coding block, and color value to form a code stream. The contour-based decoding module is used to decode the motion vector in the code stream to obtain the corresponding reference block; decode the initial point position of the contour in the code stream and locate the contour; decode the contour information in the code stream based on the decoded reference block and the located contour; decode the color value in the code stream; and reconstruct the image block by combining the decoded contour information and color value.

8. A processing device, characterized in that include: one or more processors; a memory for storing one or more programs; When the one or more programs are executed by the one or more processors, the one or more processors implement the method according to any one of claims 1 to 6.

9. A readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 6 is implemented.

Citation Information

Patent Citations

  • Image decoding method for performing intra prediction and device thereof, and image encoding method for performing intra prediction and device thereof

    CN107852507A

  • Video encoding method and apparatus and video decoding method and apparatus, using padding technique based on motion prediction

    WO2019135457A1