Video encoding method, video decoding method, computer device and storage medium
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-10-27
- Publication Date
- 2026-08-11
AI Technical Summary
[0003]现有技术的视频编码方法,存在预测效果较低等问题
[0009]上述方案,计算当前帧和知识图像缓存中各个知识图像之间的差异程度,基于所述当前帧和所述各个知识图像之间的差异程度,确定是否将所述当前帧编码成知识图像码流数据,和/或,确认所述当前帧编码过程中的知识图像参考帧,可以更加灵活的适配场景变化,达到节省比特开销,提升预测准确性的效果。
Smart Images

Figure CN117640940B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of video encoding and decoding technology, and in particular to a video encoding method, a video decoding method, a computer device, and a storage medium. Background Technology
[0002] Because video image data is relatively large, it usually needs to be encoded and compressed. The compressed video image data is called a video stream. The video stream can be transmitted to the user's end via wired or wireless network for decoding and viewing. The entire video encoding and compression process can include prediction, transformation, quantization, and encoding.
[0003] Existing video coding methods suffer from problems such as low prediction accuracy. Summary of the Invention
[0004] The main technical problem addressed by this application is to provide a video encoding method, a video decoding method, a computer device, and a storage medium that can improve prediction accuracy.
[0005] To address the aforementioned issues, a first aspect of this application provides a video encoding method, comprising: calculating the degree of difference between a current frame and various knowledge images in a knowledge image cache; determining, based on the degree of difference between the current frame and the various knowledge images, whether to encode the current frame into knowledge image bitstream data; and / or, identifying a knowledge image reference frame during the encoding process of the current frame.
[0006] To address the aforementioned problems, a second aspect of this application provides a video encoding method, comprising: receiving a video stream, wherein the video stream is obtained by an encoding end using the aforementioned video encoding method; and decoding the video stream.
[0007] To address the aforementioned problems, a third aspect of this application provides a computer device comprising a memory and a processor coupled to each other, wherein the memory stores program data and the processor executes the program data to implement any step of the aforementioned video encoding method and video decoding method.
[0008] To address the aforementioned problems, a fourth aspect of this application provides a computer-readable storage medium storing program data executable by a processor, the program data being used to implement any step of the aforementioned video encoding and video decoding methods.
[0009] The above scheme calculates the degree of difference between the current frame and each knowledge image in the knowledge image cache. Based on the degree of difference between the current frame and each knowledge image, it determines whether to encode the current frame into knowledge image bitstream data, and / or confirms the knowledge image reference frame in the current frame encoding process. This can more flexibly adapt to scene changes, achieve the effect of saving bit overhead and improving prediction accuracy.
[0010] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and are not intended to limit this application. Attached Figure Description
[0011] To more clearly illustrate the technical solutions in this application, the accompanying drawings required in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort. Among them:
[0012] Figure 1 This is a schematic diagram of the structure of an embodiment of the video encoding and decoding system of this application;
[0013] Figure 2 It is a schematic diagram of the frame reference relationship in a bitstream containing I-frames and P-frames;
[0014] Figure 3 It is a schematic diagram of the frame reference relationship in a bitstream containing I-frames, P-frames and B-frames;
[0015] Figure 4 This is a schematic diagram of the bitstream structure in one implementation method;
[0016] Figure 5 This is a flowchart illustrating one embodiment of the video encoding method of this application;
[0017] Figure 6 This is a schematic diagram of the bitstream structure of an embodiment of the video encoding method of this application;
[0018] Figure 7 This is a schematic diagram of the bitstream structure of another embodiment of the video encoding method in this application;
[0019] Figure 8 This is a flowchart illustrating one embodiment of the video decoding method of this application;
[0020] Figure 9 This is a schematic diagram of the structure of an embodiment of the computer device of this application;
[0021] Figure 10 This is a schematic diagram of the structure of an embodiment of the computer-readable storage medium of this application. Detailed Implementation
[0022] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of the embodiments. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of this application.
[0023] The terms "first" and "second" in this application are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of indicated technical features. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one of that feature. In the description of this application, "multiple" means at least two, such as two, three, etc., unless otherwise explicitly specified. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or apparatus that includes a series of steps or units is not limited to the listed steps or units, but may optionally include steps or units not listed, or may optionally include other steps or units inherent to such processes, methods, products, or apparatus.
[0024] In this application, the reference to "embodiment" means that a specific feature, structure, or characteristic described in connection with an embodiment may be included in at least one embodiment of this application. The appearance of this phrase in various places throughout the specification does not necessarily refer to the same embodiment, nor is it a separate or alternative embodiment mutually exclusive with other embodiments. It will be explicitly and implicitly understood by those skilled in the art that the embodiments described herein can be combined with other embodiments.
[0025] In this document, the term "and / or" is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, and B existing alone. Additionally, the character " / " generally indicates that the preceding and following related objects have an "or" relationship. Furthermore, "many" in this document means two or more. Moreover, the term "at least one" in this document means any combination of at least two of any one or more of a plurality of objects. For example, including at least one of A, B, and C can mean including any one or more elements selected from the set consisting of A, B, and C.
[0026] This application provides the following embodiments, and each embodiment is described in detail below.
[0027] Please see Figure 1 , Figure 1 This is a schematic diagram of the structure of an embodiment of the video encoding and decoding system of this application.
[0028] The video encoding / decoding system 100 includes an encoding end 101 and a decoding end 102. The encoding end 101 and the decoding end 102 can be computer equipment, electronic equipment, etc., and can be any device with processing capabilities, such as a computer, server, mobile phone, tablet, etc. This application does not impose any limitations on this. The encoding end 101 and the decoding end 102 can communicate with each other and can be used to perform encoding and / or decoding operations on images / videos.
[0029] Encoding end 101 can be used to perform encoding and compression steps for images / videos to obtain video stream data. Encoding end 101 can transmit the video stream data to decoding end 102. Decoding end 102 can receive the video stream data from encoding end 101 and perform decoding and other related steps containing the video stream data, as well as steps related to backend vision tasks, such as image processing and classification.
[0030] A video stream consists of consecutive frames. Each frame is decoded and played in sequence to form a video picture. Common frame types in video streams include I-frames, P-frames, and B-frames.
[0031] I-frames are intra-coded frames, meaning they are independent frames with all their own encoding and decoding information, and can be encoded and decoded independently without referencing other frames. I-frames require complete encoding of all content within the frame, generally resulting in a larger bitstream and lower compression ratio.
[0032] P-frames are inter-frame predictive coded frames, which require previous frames in the display order as reference images for encoding and decoding.
[0033] B-frames are bidirectional inter-frame predictive coded frames, requiring both past and future frames in the display order as reference frames for encoding and decoding.
[0034] Figure 2 This is a schematic diagram of the frame reference relationship in a bitstream containing I-frames and P-frames. Figure 3 This is a schematic diagram of the frame reference relationship in a bitstream containing I-frames, P-frames, and B-frames. Figure 2 and Figure 3 In this context, POC (pic_order_cnt) represents the playback order of video frames, and DOI (decode order index) represents the encoding and decoding order of video frames. Note that the frame reference relationships in a video sequence can have various free combinations. Figure 2 and Figure 3 This is just to illustrate a common reference relationship.
[0035] Furthermore, related technologies (such as the SVAC3 video codec standard) introduce the concept of a knowledge picture (library picture), and in this application, an L-frame represents a knowledge picture. Simultaneously, the concept of a reference library (RL) frame is introduced, which refers to a P-frame or B-frame that serves as a reference frame for a knowledge picture. A knowledge picture is a long-term reference frame, encoded using I-frames. Knowledge pictures are identified using their knowledge picture index (IDX).
[0036] Regarding the structure of knowledge images in the bitstream: In related technologies, knowledge images are encoded using I-frames. However, since the QP of encoded knowledge images is generally small, the encoding is slow, and the bit rate is generally large, adding a whole frame of knowledge image bitstream to the bitstream will cause large bit rate spikes and jitter during decoding. Therefore, related technologies utilize the patch mechanism in the SVAC3 standard to add an encoding and decoding mechanism for knowledge image patch transmission: the knowledge image is divided into multiple patches and interleaved with multiple display images for encoding. Only one patch is encoded at a time and added to the bitstream, ultimately resulting in an encoded output bitstream that interleaves the knowledge base patch bitstream with the display image bitstream. Figure 4 This refers to the frame reference relationship and the position of the knowledge image in the bitstream under the IPPP configuration in related technologies. Of course, in other embodiments, the knowledge image can also be encoded and transmitted as a whole frame, that is, the whole frame of knowledge image can be directly encoded into a bitstream unit.
[0037] The frame in the video that is defaulted to being an L-frame is called the default knowledge image generation position. The default knowledge image generation position is determined based on the following configuration parameters: I-frame interval (how many frames are between two I-frames), knowledge image update interval (how many I-frame intervals are between updating the knowledge image), and patch width and / or patch height. For example... Figure 4 In the diagram, the frames marked by the slashes POC=0 and POC=3 are the default knowledge image generation locations, and the solid arrows represent reference relationships.
[0038] Generally, starting from a default knowledge image generation position, one frame is encoded into a knowledge image frame every default knowledge image generation interval. That is, there is a default knowledge image generation position every default knowledge image generation interval. The default knowledge image generation interval can be equal to the I-frame interval multiplied by the knowledge image update interval. For example, if the first default knowledge image generation position is frame 0, the I-frame interval is 25 frames, and the knowledge image update interval is 2, then the default knowledge image generation interval could be 50 frames. Therefore, frames 0, 50, 100, 150, and so on should be the default knowledge image generation positions. However, in the case of knowledge image fragmented interleaving transmission, the default knowledge image generation interval is adjusted based on the fragmentation of the knowledge images (e.g., the width and height of the knowledge image fragments). Therefore, during video encoding, the default knowledge image generation interval will change based on the actual situation; that is, the default knowledge image generation interval is not constant. Based on the above, it can be concluded that there is one and only one default knowledge image generation position within the default knowledge image generation interval.
[0039] Furthermore, in related technologies, when an image starts referencing a new knowledge image or begins encoding / decoding a new knowledge image, the previous knowledge image is replaced. In other words, all frames in the current sequence can refer to at most one knowledge image, and that knowledge image must be the most recently encoded / decoded knowledge image.
[0040] Although there are related technologies that take into account that multiple frames of knowledge images can be managed in the knowledge image cache, and that previously encoded knowledge images can be referenced in subsequent sequences, this is beneficial for the encoding of monitoring sequences in scenarios such as switching points of PTZ cameras, as it can reduce the number of times similar knowledge images are transmitted and speed up the encoding process.
[0041] The aforementioned knowledge image multi-frame management mechanisms are primarily designed for sequences of regularly changing scenes in surveillance scenarios, where video content undergoes periodic changes. Therefore, a reference knowledge image is pre-set for each frame within a cycle, which can be reused in subsequent cycles. However, this approach has a limited scope, only applicable to situations where the camera rotates at a constant speed or cruises at a fixed speed. In irregular scenarios, the pre-set reference frame list for the current frame becomes less adaptable. Therefore, a more flexible and adaptive method for selecting the reference knowledge image is needed.
[0042] Specifically, when encountering irregular and non-uniform changes in image scenes, neither generating knowledge images at fixed intervals nor manually pre-setting reference knowledge images can meet the requirements of scene adaptation. This can result in generated knowledge image frames being similar to existing knowledge image frames, leading to wasted bit overhead, or the scene changing significantly without new knowledge images being generated, causing a decline in prediction performance. Furthermore, if multiple knowledge image frames exist, deciding which frame to reference for the current frame is also problematic. To address this, this application proposes a video coding method that determines whether to generate knowledge images and which frames to reference for the current frame based on image similarity. This allows for more flexible adaptation to scene changes, saving bit overhead and improving prediction accuracy.
[0043] Please see Figure 5 , Figure 5 This is a flowchart illustrating the first embodiment of the video encoding method of this application. The specific steps of the video encoding method in this embodiment can be executed using the aforementioned encoding terminal. The method may include the following steps:
[0044] S11: Calculate the degree of difference between the current frame and each knowledge image in the knowledge image cache.
[0045] The degree of difference between the current frame and each knowledge image in the knowledge image buffer (LDPB) can be calculated so that, based on the degree of difference between the current frame and each knowledge image, it can be determined whether to encode the current frame into knowledge image bitstream data, and / or to identify the knowledge image reference frame in the current frame encoding process.
[0046] In one feasible approach, the degree of difference between the current frame and the various knowledge images in the knowledge image cache can be directly calculated.
[0047] For example, the degree of difference between the current frame and the knowledge images in the knowledge image cache can be measured by methods such as SAD (sum of absolute differences), SATD (sum of absolute transformed differences), hash or histogram statistics.
[0048] In another possible approach, the similarity between the current frame and each knowledge image in the knowledge image cache can be calculated, and the similarity between the current frame and each knowledge image in the knowledge image cache can be used to characterize the degree of difference between the current frame and each knowledge image in the knowledge image cache. It can be understood that the similarity between the current frame and each knowledge image in the knowledge image cache is negatively correlated with the degree of difference between the current frame and each knowledge image in the knowledge image cache.
[0049] For example, the cosine similarity between the current frame and each knowledge image in the knowledge image cache can be used to measure the degree of similarity between the current frame and each knowledge image in the knowledge image cache.
[0050] The number of knowledge images in the knowledge image cache is unlimited; for example, it can be one or more.
[0051] S12: Based on the degree of difference between the current frame and each knowledge image, determine whether to encode the current frame into knowledge image bitstream data, and / or confirm the knowledge image reference frame in the current frame encoding process.
[0052] After calculating the degree of difference between the current frame and each knowledge image in the knowledge image cache, it is possible to determine whether to encode the current frame into knowledge image bitstream data based on the degree of difference between the current frame and each knowledge image, and / or to confirm the knowledge image reference frame in the current frame encoding process, so as to adapt to scene changes more flexibly, save bit overhead, and improve prediction accuracy.
[0053] In one application scenario, it can be determined whether to encode the current frame into knowledge image bitstream data based on the degree of difference between the current frame and the various knowledge images.
[0054] In this embodiment, when the current frame is located at the default knowledge image generation position, the degree of difference between the current frame and the various knowledge images can be used to determine whether to encode the current frame into knowledge image bitstream data. If, based on the degree of difference between the current frame and the various knowledge images, it is determined that the current frame should not be encoded into knowledge image bitstream data, then it is not encoded into knowledge image bitstream data. In this case, other frames (e.g., other regular image frames or knowledge image frames) can be used as references to encode the current frame, i.e., inter-frame encoding or intra-frame encoding can be performed. If, based on the degree of difference between the current frame and the various knowledge images, it is determined that the current frame should be encoded into knowledge image bitstream data, then the current frame is encoded into knowledge image bitstream data, and after the knowledge image bitstream data is decoded and reconstructed, the knowledge image of this frame can be added to the knowledge image cache to serve as a knowledge image reference frame for subsequent frames. In other embodiments, even if the current frame is not located at the default knowledge image generation position, if the degree of difference between the current frame and the various knowledge images meets a preset condition, it can still be determined whether to encode the current frame into knowledge image bitstream data.
[0055] Furthermore, when the current frame is located at the default knowledge image generation position, if it is determined, based on the degree of difference between the current frame and the various knowledge images, that the current frame will not be encoded into knowledge image bitstream data; when the current frame is located at an intra-frame encoding position, a knowledge image reference frame in the current frame encoding process can be determined based on the degree of difference between the current frame and the various knowledge images, and then inter-frame encoding is performed on the current frame based on the determined knowledge image reference frame, or intra-frame encoding can be performed directly on the current frame, in which case it is not necessary to determine the knowledge image reference frame in the current frame encoding process based on the degree of difference between the current frame and the various knowledge images; when the current frame is located at an inter-frame encoding position, a knowledge image reference frame in the current frame encoding process can be determined based on the degree of difference between the current frame and the various knowledge images, and then inter-frame encoding is performed on the current frame based on the determined knowledge image reference frame.
[0056] Furthermore, when the current frame is located at the default knowledge image generation position, if it is determined, based on the degree of difference between the current frame and the various knowledge images, that the current frame should be encoded into knowledge image bitstream data, the current frame can be encoded into knowledge image bitstream data first, and then the current frame can be encoded a second time. Specifically, when the current frame is located at an intra-frame encoding position, during the second encoding of the current frame, inter-frame encoding can be performed on the current frame using the "knowledge image reconstructed by encoding and decoding the current frame" as a reference. Alternatively, a knowledge image reference frame in the current frame encoding process can be determined based on the degree of difference between the current frame and the various knowledge images, and then inter-frame encoding can be performed on the current frame using the knowledge image reference frame determined by the degree of difference. When the current frame is located at an inter-frame encoding position, during the second encoding of the current frame, a knowledge image reference frame in the current frame encoding process can be determined based on the degree of difference between the current frame and the various knowledge images, and then inter-frame encoding can be performed on the current frame using the knowledge image reference frame determined by the degree of difference.
[0057] Furthermore, if it is confirmed that the current frame will be encoded into knowledge image bitstream data, then if the current frame is encoded into knowledge image bitstream data for display, a second encoding of the current frame is unnecessary. That is, the original image content of the current frame can be encoded only once to obtain the knowledge image bitstream data. Here, "knowledge image for display" refers to the knowledge image output for display at the decoding end; that is, "display" refers to the decoded output used for display.
[0058] If the degree of difference between the current frame and each of the knowledge images meets the preset conditions, the current frame can be confirmed to be encoded into knowledge image bitstream data; otherwise, the current frame will not be encoded into knowledge image bitstream data.
[0059] If the difference between the current frame and any knowledge image is less than a preset threshold, then the difference between the current frame and each of the knowledge images does not meet the preset condition. That is, if the difference between the current frame and any knowledge image is less than the preset threshold, it can be confirmed that the current frame will not be encoded into knowledge image bitstream data; otherwise, the current frame will be encoded into knowledge image bitstream data.
[0060] In one embodiment, the preset threshold may include a first fixed threshold. That is, if the difference between the current frame and any knowledge image is less than the first fixed threshold, it can be determined that the current frame will not be encoded into knowledge image bitstream data; otherwise, the current frame will be encoded into knowledge image bitstream data. In other words, by pre-setting a first fixed threshold, if the calculated difference between the frame at the current default knowledge image generation position and any knowledge image in the knowledge image cache is less than the first fixed threshold, then the current frame (e.g., the frame at the current default knowledge image generation position) will not generate a knowledge image.
[0061] In a specific example, assume the first fixed threshold TH = w * h * 2^6, where w is the width of the image and h is the height of the image. The degree of difference between two images is measured by SAD, which is the difference between each corresponding pixel of the two images, the absolute values of the subtractions, and the summation.
[0062] Let frame X be the current default knowledge image generation location. LDPB contains two knowledge images {L0, L1}. The SAD between X and L0 is SAD0, and the SAD between X and L1 is SAD1. Comparison reveals that SAD0...<TH,SAD1> If TH indicates that X and L0 are very similar, then no knowledge image will be generated for frame X.
[0063] In another embodiment, the preset threshold may include a floating threshold, and the floating threshold corresponds one-to-one with the knowledge images. The floating threshold corresponding to each knowledge image can be determined based on the degree of difference between each knowledge image and the specified position frame corresponding to each knowledge image. For ease of description, the basic noise cost of each knowledge image can be used to represent the degree of difference between each knowledge image and the specified position frame corresponding to each knowledge image. The basic noise cost of each knowledge image is a factor that positively influences the floating threshold corresponding to each knowledge image. Furthermore, the basic noise cost of each knowledge image can be positively correlated with the floating threshold corresponding to each knowledge image. Further, the floating threshold corresponding to each knowledge image can be equal to the second fixed threshold plus the basic noise cost of each knowledge image. Alternatively, the floating threshold corresponding to each knowledge image can be equal to the product of the second fixed threshold and the basic noise cost of each knowledge image. Preferably, the second fixed threshold is less than the first fixed threshold. Furthermore, the second fixed threshold and the first fixed threshold can be set according to actual conditions such as the image size, and are not limited here.
[0064] There are multiple ways to determine the specified location frame corresponding to each knowledge image, which are not limited here.
[0065] For example, a frame at a fixed position can be selected as the specified position frame within the default knowledge image generation interval to which the knowledge image belongs. Here, "fixed position" can mean that the interval between the frame and the first frame in the default knowledge image generation interval to which the knowledge image belongs is fixed. That is, the k-th frame in the default knowledge image generation interval to which the knowledge image belongs can be selected as the specified position frame, where k is a positive constant.
[0066] For example, the frame adjacent to the frame where the POC is located at the default knowledge image generation position can be selected as the specified position frame corresponding to the knowledge image frame at the default knowledge image generation position.
[0067] In this embodiment, it can be determined whether the degree of difference between the current frame and each knowledge image is less than the floating threshold corresponding to each knowledge image; if the degree of difference between the current frame and any knowledge image is less than its corresponding floating threshold, then the current frame is not encoded into knowledge image bitstream data, otherwise the current frame is encoded into knowledge image bitstream data.
[0068] In a specific example, suppose the designated position frame is selected as the second frame Xn in each default knowledge image generation interval, and in this embodiment, a corresponding knowledge image Ln is generated in each default knowledge image generation interval. Figure 6As shown, starting with RL frames and ending with the next RL frame, there is a default knowledge image generation interval. The knowledge image corresponding to the first default knowledge image generation interval is L0, and the specified position frame is X0. The knowledge image corresponding to the second default knowledge image generation interval is L1, and the specified position frame is X1. Let the basic noise cost be measured by SAD, then the basic noise cost CBn corresponding to knowledge image Ln is the SAD between Xn and Ln. Let the frame at the current default knowledge image generation position be Y, the SAD between Y and L0 be SAD0, and the SAD between Y and L1 be SAD1. There are two knowledge images {L0, L1} in LDPB, with the basic noise cost corresponding to L0 being CB0 and CB1. The second fixed threshold th = w*h*2^5. Comparison shows that SAD0 > CB0 + th, and SAD1 > CB1 + th, indicating that Y is dissimilar to all knowledge image frames in LDPB. Therefore, a new knowledge image needs to be generated for frame Y.
[0069] In another embodiment, the preset threshold may include a first fixed threshold and a floating threshold. In this embodiment, it can be determined whether the degree of difference corresponding to each knowledge image is less than the floating threshold corresponding to each knowledge image, and whether the degree of difference corresponding to each knowledge image is less than the first fixed threshold. If the difference between the current frame and a knowledge image is less than the floating threshold and / or less than the first fixed threshold, then the difference between the current frame and the knowledge image is less than the preset threshold. In a specific example, if the difference between the current frame and a knowledge image is less than the floating threshold or less than the first fixed threshold, then the difference between the current frame and the knowledge image is less than the preset threshold, and the current frame is not encoded into knowledge image bitstream data. In another specific example, if the difference between the current frame and a knowledge image is less than the floating threshold and less than the first fixed threshold, then the difference between the current frame and the knowledge image is less than the preset threshold, and the current frame is not encoded into knowledge image bitstream data.
[0070] In the embodiment where the preset threshold includes a floating threshold, it is necessary to determine the floating threshold corresponding to each knowledge image. The floating threshold corresponding to each knowledge image is related to the specified position frame corresponding to each knowledge image. After confirming using the above implementation method that an image frame needs to be encoded into a knowledge image frame, the specified position frame corresponding to the knowledge image frame can be determined based on the above rules. Then, the degree of difference between the knowledge image frame and the specified position frame (i.e., the basic noise cost) is calculated and recorded. This basic noise cost is then used to determine the knowledge image reference frame for the frame to be encoded and / or whether to generate a knowledge image. Specifically, performing the above judgment and operation within each default knowledge image generation interval yields the basic noise cost corresponding to each knowledge image frame in the LDPB.
[0071] Furthermore, in the implementation method of using the similarity between the current frame and each knowledge image in the knowledge image cache to characterize the degree of difference between the current frame and each knowledge image in the knowledge image cache, it is possible to determine whether the similarity between the current frame and each knowledge image meets the requirements. For example, it is possible to determine whether the similarity between the current frame and each knowledge image is less than the similarity threshold. If they are all less than the threshold, it can be confirmed that the current frame will be encoded into knowledge image bitstream data; otherwise, the current frame will not be encoded into knowledge image bitstream data.
[0072] In another application scenario, the reference frame for the knowledge image in the encoding process of the current frame can be identified based on the degree of difference between the current frame and each of the knowledge images. In this way, the decision on which knowledge image to refer to for the current frame is adaptively made based on the similarity of the images, thereby improving prediction accuracy and compression performance.
[0073] Optionally, when the current frame requires reference knowledge images, a predetermined number of knowledge images can be selected from all knowledge images in the knowledge image cache as reference frames for the encoding process of the current frame, based on the degree of difference between the current frame and the various knowledge images. More preferably, the degree of difference between the selected predetermined number of knowledge images and the current frame is less than the degree of difference between the current frame and any other knowledge images (excluding the predetermined number of knowledge images) among all knowledge images; that is, the predetermined number of knowledge images with the smallest degree of difference from the current frame are selected from all knowledge images in the knowledge image cache.
[0074] In one embodiment, it can be stipulated that all inter-frame coded frames (e.g., P-frames and B-frames) must reference knowledge images during encoding. Thus, during the encoding of all inter-frame coded frames, a predetermined number of knowledge images can be selected from all knowledge images in the knowledge image cache as knowledge image reference frames based on the degree of difference between the current frame and the various knowledge images. In other embodiments, for some inter-frame coded frames, knowledge images may not be referenced during encoding. Therefore, the method for determining knowledge image reference frames described above can be omitted for these inter-frame coded frames.
[0075] In related technologies, the Reference Frame Buffer (RPL) needs to be manually configured, and all reference frame information for the current frame, including knowledge image reference frames and ordinary reference frames, needs to be confirmed based on the RPL. Furthermore, the Dependent Page Buffer (DPB) needs to be updated based on the RPL. In cases of irregularity, it cannot flexibly and adaptively select the knowledge image to be referenced, resulting in poor adaptability of the pre-set RPL. Therefore, combining a scheme that selects knowledge image reference frames based on the degree of difference, this application can separate the knowledge image reference frame portion of the RPL and set a separate Reference Frame Buffer (LRPL) for knowledge images to manage the knowledge image frames to be referenced in the current frame. Then, the Dependent Page Buffer (LDPB) is updated based on the LRPL. This allows control over the removal of knowledge images from the knowledge image cache based on the degree of difference between the current frame and the various knowledge images.
[0076] Specifically, a knowledge image list (i.e., LRPL) for managing the knowledge images to be referenced by the current frame can be set. If the number n (n >= 1) of knowledge images in the knowledge image cache is greater than the upper limit LDPB_LIB_NUM (LDPB_LIB_NUM >= 1) of the knowledge image list, add the preset number c (c >= 1) of knowledge images to the knowledge image list; if the knowledge image list is not full, add some of the other knowledge images to the knowledge image list; if the knowledge image list is full, update the knowledge image cache according to the knowledge image list, that is, remove the remaining part of the other knowledge images from the knowledge image cache. Assume that the upper limit LDPB_LIB_NUM of the knowledge image list is 2, the preset number c is 1, and the number n of knowledge images in the knowledge image cache is 3. In this embodiment, the 1 knowledge image with the smallest degree of difference from the current frame in the knowledge image cache can be added to the knowledge image list, and then one of the remaining two knowledge images in the knowledge image cache is selected and added to the knowledge image list. Then, update the knowledge image cache according to the knowledge image list, that is, only the knowledge images in the knowledge image list are retained in the knowledge image cache, that is, the knowledge images that are not added to the knowledge image list among the original 3 knowledge images in the knowledge image cache will be removed from the knowledge image cache. In a specific example, the current frame can have at most m = 1 knowledge image reference frame, and there are n = 3 knowledge images in LDPB: {L0, L1, L2}, corresponding to POC {0, 50, 100} respectively, and the parameter LDPB_LIB_NUM = 2. The SAD between the current frame and L0 is SAD0, the SAD between the current frame and L1 is SAD1, and the SAD between the current frame and L2 is SAD2. Sorting them from small to large gives SAD1 < SAD0 < SAD2, then the constructed LRPL is {L1, L2}. Among them, L1 is used as the knowledge image reference frame for the current frame, and L2 can be used for reference by subsequent frames. Correspondingly, LDPB will also be updated to {L1, L2}, that is, L0 in LDPB will be removed.
[0077] Adding some knowledge images from the other knowledge images to the knowledge image list can be manifested as adding the knowledge image with the smallest difference from the current frame or the knowledge image with the smallest interval from the current frame to the knowledge image list. In one example, all knowledge images in the LDPB can be placed into the LRPL in ascending order of their corresponding difference levels until LDPB_LIB_NUM is filled. In another example, all knowledge images in the LDPB can be placed in ascending order of their corresponding difference levels, and the first c knowledge image frames are placed into the LRPL for reference to the current frame. Then, the remaining frames in the LDPB are filled into the remaining positions of the LRPL one by one in descending order of their index, for reference by subsequent frames. In this example, the index of a knowledge image in the LDPB is negatively correlated with the interval between that knowledge image and the current frame; that is, the knowledge image with the largest index has the smallest interval with the current frame.
[0078] If the number of knowledge images n (n>=1) in the knowledge image cache is less than or equal to the upper limit LDPB_LIB_NUM (LDPB_LIB_NUM>=1) of the knowledge image list, all knowledge images in the knowledge image cache can be added to the knowledge image list. The order of the knowledge images in the knowledge image list is not restricted. However, a preferred approach is to place all knowledge images from the knowledge image cache into the knowledge image list in ascending order of their difference from the current frame. The first preset number of knowledge images are used as a reference for the current frame, and the remaining images are used as a reference for subsequent frames.
[0079] To facilitate the management of all knowledge image reference frames for the current frame based on knowledge images, the upper limit of the knowledge image list, LDPB_LIB_NUM, can be greater than or equal to the aforementioned preset number c. Additionally, the number of knowledge images n in the knowledge image cache can be greater than or equal to the upper limit m of the number of knowledge image reference frames for the current frame. However, in certain special cases, such as at the start of video encoding, the number n of knowledge images in the knowledge image cache can be less than the upper limit m of the number of knowledge image reference frames for the current frame. Thus, if the upper limit m of the number of knowledge image reference frames for the current frame is greater than the number n of knowledge images in the knowledge image cache, the preset number c is the number n of knowledge images in the knowledge image cache; if the upper limit m of the number of knowledge image reference frames for the current frame is less than or equal to the number n of knowledge images in the knowledge image cache, the preset number c is the upper limit m of the number of knowledge image reference frames.
[0080] Optionally, in the video coding method, a syntax can also be transmitted in the video bitstream, and this syntax specifies the maximum length of the LDPB (i.e., the maximum number of knowledge images in the LDPB). For the numerical values of this syntax, the following two methods are included but not limited to:
[0081] 1) Let the maximum value of this syntax be a preset value s, that is, the numerical range of this syntax is 0 to s.
[0082] 2) Let the maximum value of this syntax be a preset value s, and s plus the number of ordinary reference images in the DPB cannot exceed t. This further limits the length of the LDPB and saves storage space.
[0083] In an example, s is set to 7 and t is set to 15. The syntax max_ldpb_size_minus1 can be added to the sequence header to represent the maximum length of the LDPB minus one. The value range of max_ldpb_size_minus1 is set to 0 to 7. Let max_ldpb_size_minus1 + 1 represent the maximum length of the LDPB, that is, the maximum length of the LDPB is 8.
[0084] The following is to better illustrate the video coding method of this application. The following specific video coding embodiments are provided for exemplary illustration:
[0085] In this embodiment, whether to generate a knowledge image at the default knowledge image generation position is comprehensively confirmed through a first fixed threshold and a floating threshold. Among them, the fixed threshold TH = w * h * 2^6. There are 2 knowledge images in the LDPB: {L0, L1}, corresponding to POC {0, 50}, and the corresponding basic noise cost values are: {CB0, CB1}.
[0086] In addition, in this embodiment, the current frame can have at most m = 1 knowledge image reference frame, and the parameter LDPB_LIB_NUM = 2.
[0087] Let the current frame be a frame X at the default knowledge image generation position, with the corresponding POC being 150. The SAD between X and L0 is SAD0, and the SAD between X and L1 is SAD1; the second fixed threshold th = w * h * 2^5. The measurement criterion is that if SADn < TH or SADn < CBn + th, then a new knowledge image is not generated. After comparison, it is found that SAD0 > TH, and SAD0 > CB0 + th, SAD1 > TH, and SAD1 > CB1 + th, indicating that X is not similar to all the knowledge image frames in the LDPB. Then a new knowledge image L2 needs to be generated for the X frame.
[0088] L2 is encoded using intra-frame coding, and after decoding and reconstruction, it is placed in the LDPB. At this time, there are 3 knowledge images in the LDPB: {L0, L1, L2}, and the corresponding basic noise cost values are: {CB0, CB1, 0}.
[0089] Then, encode to the next frame Y after the POC is located after frame X, that is, the current frame is frame Y after frame X. Assume that the frame adjacent to the frame at the default knowledge image generation position is selected as the designated position frame for calculating the basic noise cost value corresponding to the knowledge image. As Figure 7 shown, the frame filled with diagonal lines represents the default knowledge image generation position, and the frame filled with vertical lines is the designated position frame. Then calculate the basic noise cost value: The basic noise cost value CB2 corresponding to L2 is the SAD between Y and L2. At this time, the basic noise cost values corresponding to the knowledge images {L0, L1, L2} in the LDPB can be updated to: {CB0, CB1, CB2}.
[0090] Next, an LRPL needs to be constructed for the current frame Y. The SAD between the current frame Y and L0 is SAD01, the SAD between the current frame Y and L1 is SAD11, and the SAD between the current frame Y and L2 is SAD21. Sorting them from small to large gives SAD11 < SAD01 < SAD21. Then the knowledge image reference frame of the current frame is L1, and the constructed LRPL is {L1, L2}. Among them, L1 is used as the knowledge image reference frame of the current frame, and L2 can be used for reference by subsequent frames.
[0091] Then, update the LDPB according to the LRPL of the current frame Y. After the update, there are only 2 knowledge images in the LDPB: {L1, L2}, and the corresponding basic noise cost values are: {CB1, CB2}.
[0092] After encoding the Y frame, the LRPL information and its reference frame information need to be transmitted to the decoding end. The decoding end analyzes and obtains the LRPL list as {L1, L2}, and analyzes the knowledge image reference information of Y and knows that Y only references L1, and then the Y frame can be decoded and reconstructed.
[0093] Please refer to Figure 8 , Figure 8 which is a schematic flowchart of the first embodiment of the video decoding method of the present application. The specific steps of the video decoding method in this embodiment can be executed by the above-mentioned decoding end. The method may include the following steps:
[0094] S21: Receive a video code stream, where the video code stream is obtained by the encoding end using the above video encoding method.
[0095] The encoding end can transmit video stream data to the decoding end, enabling the decoding end to receive the video stream data. If the encoding end stores the video stream data, it can then be used as a decoding end to perform operations such as decoding, playback, or storage of the video stream data.
[0096] The specific implementation method of this step can be referred to the specific implementation process of the above-mentioned encoding end, and will not be repeated here.
[0097] S22: Decode the video stream.
[0098] Decoding the video stream yields the decoded video data.
[0099] The specific implementation of this embodiment can be referred to the implementation process of the above embodiments, and will not be repeated here.
[0100] Regarding the above embodiments, this application provides a computer device; please refer to [link / reference]. Figure 9 , Figure 9 This is a schematic diagram of the structure of a computer device according to an embodiment of the present application. The computer device 50 includes a memory 51 and a processor 52, wherein the memory 51 and the processor 52 are coupled to each other. The memory 51 stores program data, and the processor 52 executes the program data to implement the steps of any of the above-described video encoding and video decoding methods. The computer device 50 can serve as the encoding end and / or decoding end in the video encoding and decoding system of the above embodiments, executing the steps of any of the above-described video encoding and video decoding methods.
[0101] In this embodiment, processor 52 can also be referred to as CPU (Central Processing Unit). Processor 52 may be an integrated circuit chip with signal processing capabilities. Processor 52 can also be a general-purpose processor, digital signal processor (DSP), application-specific integrated circuit (ASIC), field-programmable gate array (FPGA), or other programmable logic device, discrete gate or transistor logic device, or discrete hardware component. The general-purpose processor can be a microprocessor, or processor 52 can be any conventional processor.
[0102] The methods described in the above embodiments can be implemented as computer programs; therefore, this application proposes a computer-readable storage medium. Please refer to [link to relevant documentation]. Figure 10 , Figure 10 This is a schematic diagram of a computer-readable storage medium according to an embodiment of the present application. The computer-readable storage medium 60 stores program data 61 that can be executed by a processor. The program data 61 can be executed by the processor to implement the steps of any of the above-described video encoding and video decoding methods.
[0103] In this embodiment, the computer-readable storage medium 60 can be a USB flash drive, a portable hard drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk, or a medium that can store program data 61. Alternatively, it can be a server that stores the program data 61. The server can send the stored program data 61 to other devices for execution, or it can run the stored program data 61 itself.
[0104] In the several embodiments provided in this application, it should be understood that the disclosed methods and apparatus can be implemented in other ways. For example, the apparatus implementations described above are merely illustrative. For instance, the division of modules or units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between devices or units may be electrical, mechanical, or other forms.
[0105] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment, depending on actual needs.
[0106] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0107] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause an electronic device (which may be a personal computer, server, or network device, etc.) or processor to execute all or part of the steps of the methods of the various embodiments of this application.
[0108] Obviously, those skilled in the art should understand that the modules or steps of this application described above can be implemented using general-purpose computing devices. They can be centralized on a single computing device or distributed across a network of multiple computing devices. Optionally, they can be implemented using computer-executable program code, and thus stored in a computer-readable storage medium for execution by a computing device, or they can be fabricated as separate integrated circuit modules, or multiple modules or steps can be fabricated as a single integrated circuit module. Therefore, this application is not limited to any particular hardware and software combination.
[0109] The above description is merely an embodiment of this application and does not limit the patent scope of this application. Any equivalent structural or procedural transformations made using the content of this application's specification and drawings, or direct or indirect applications in other related technical fields, are similarly included within the patent protection scope of this application.
Claims
1. A method of video coding, the method comprising: The method includes: Calculate the degree of difference between the current frame and each knowledge image in the knowledge image cache; Based on the degree of difference between the current frame and each knowledge image, determine whether to encode the current frame into knowledge image bitstream data, and / or, confirm the knowledge image reference frame in the current frame encoding process; The step of determining whether to encode the current frame into knowledge image bitstream data based on the degree of difference between the current frame and each knowledge image, and / or confirming the knowledge image reference frame in the current frame encoding process, includes: If the current frame requires reference knowledge images, a preset number of knowledge images are selected from all knowledge images in the knowledge image cache as knowledge image reference frames in the encoding process of the current frame; wherein, the degree of difference between the preset number of knowledge images and the current frame is less than the degree of difference between the current frame and other knowledge images other than the preset number of knowledge images. The knowledge image cache is updated based on the following steps: Set up a knowledge image list to manage the knowledge images referenced in the current frame; If the number of knowledge images in the knowledge image cache is greater than the upper limit of the knowledge image list, the preset number of knowledge images are added to the knowledge image list; If the knowledge image list is not complete, add the knowledge image with the smallest difference from the current frame or the knowledge image with the smallest interval from the current frame to the knowledge image list. If the knowledge image list is filled, update the knowledge image cache according to the knowledge image list; The upper limit of the knowledge image list is greater than or equal to the preset number.
2. The video coding method of claim 1, wherein, The step of determining whether to encode the current frame into knowledge image bitstream data based on the degree of difference between the current frame and each knowledge image, and / or confirming the knowledge image reference frame in the current frame encoding process, includes: If the current frame is a frame at the default knowledge image generation location, in response to the difference between the current frame and any knowledge image being less than a preset threshold, the current frame will not be encoded into knowledge image bitstream data; otherwise, the current frame will be encoded into knowledge image bitstream data.
3. The video coding method of claim 2, wherein, The preset threshold includes a floating threshold, and the floating threshold corresponding to each knowledge image is determined based on the degree of difference between each knowledge image and its corresponding specified position frame; the step of responding to the fact that the degree of difference between the current frame and any knowledge image is less than the preset threshold, not encoding the current frame into knowledge image bitstream data, otherwise encoding the current frame into knowledge image bitstream data, includes: Determine whether the degree of difference between the current frame and each knowledge image is less than the floating threshold corresponding to each knowledge image; If the difference between the current frame and any knowledge image is less than its corresponding floating threshold, then the current frame will not be encoded into knowledge image bitstream data.
4. The video coding method of claim 2, wherein, The preset threshold includes a floating threshold and a first fixed threshold. The floating threshold for each knowledge image is determined based on the degree of difference between each knowledge image and its corresponding specified position frame. The step of not encoding the current frame into knowledge image bitstream data if the degree of difference between the current frame and any knowledge image is less than the preset threshold, and otherwise encoding the current frame into knowledge image bitstream data, includes: Determine whether the degree of difference between the current frame and each knowledge image is less than the floating threshold corresponding to each knowledge image, and determine whether the degree of difference between the current frame and each knowledge image is less than the first fixed threshold; The response that the difference between the current frame and any knowledge image is less than a preset threshold, and therefore the current frame is not encoded into knowledge image bitstream data, or otherwise the current frame is encoded into knowledge image bitstream data, includes: If the difference between the current frame and a knowledge image is less than the floating threshold and / or less than the first fixed threshold, then the difference between the current frame and the knowledge image is less than the preset threshold.
5. The video coding method of claim 4, wherein, The floating threshold is the sum of the second fixed threshold and the degree of difference between each knowledge image and its corresponding specified location frame, wherein the second fixed threshold is less than the first fixed threshold; The first fixed threshold and the second fixed threshold are positively correlated with the size of the image.
6. The video encoding method according to any one of claims 3-5, characterized in that, The interval between the specified position frame corresponding to each knowledge image and the first frame in the default knowledge image generation interval to which each knowledge image belongs is a constant; or... The designated frame corresponding to each knowledge image is the next frame after the frame at the default knowledge image generation position corresponding to each knowledge image, according to the image playback sequence number.
7. The video coding method of claim 1, wherein, If the upper limit of the number of knowledge image reference frames in the current frame is greater than the number of knowledge images in the knowledge image cache, the preset number is the number of knowledge images in the knowledge image cache; if the upper limit of the number of knowledge image reference frames in the current frame is less than or equal to the number of knowledge images in the knowledge image cache, the preset number is the upper limit of the number of knowledge image reference frames.
8. The video coding method of claim 1, wherein, The method further includes: A preset syntax is set in the video bitstream, which is used to indicate the upper limit of the number of knowledge images in the knowledge image cache. The video bitstream is obtained by encoding the video.
9. A method of video decoding, comprising: The method includes: Receive a video stream, wherein the video stream is obtained by the encoding end using the video encoding method according to any one of claims 1 to 8; The video stream is decoded.
10. A computer device, comprising: It includes a memory and a processor coupled to each other, the memory storing program data, and the processor executing the program data to implement the steps of the method of any one of claims 1 to 8, and / or, to implement the steps of the method of claim 9.
11. A computer readable storage medium, characterized in that, The system stores program data that can be executed by a processor, the program data being used to implement the steps of the method according to any one of claims 1 to 8, and / or, to implement the steps of the method according to claim 9.
Citation Information
Patent Citations
Face video coding method, decoding method and device
CN114205585A
Video encoding method, video decoding method, computer device and storage medium
WO2018171447A1