Screen Encoding Method, Screen Decoding Method and Related Devices

By selecting the prediction block of the block to be encoded from the cache list and outputting the code stream data, the problem of limitation of the search range of IBC technology is solved, improving the encoding performance and reducing the code rate.

CN114697666BActive Publication Date: 2025-05-27CAMBRICON TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202011628731.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-12-30
Publication Date
2025-05-27
Estimated Expiration
2040-12-30

AI Technical Summary

Technical Problem

In the existing screen content encoding method, the search range of the intra-block copy (IBC) technology is limited to the encoded area of ​​the current frame, resulting in low encoding efficiency.

Method used

By obtaining the block to be encoded and selecting its prediction block from the cache list, the cache list contains image blocks in at least two frames of images, and the output code stream data includes the residuals of the prediction block and the block to be encoded and the identification of the prediction block.

Benefits of technology

The encoding performance is improved, the code rate is reduced, and the encoding quality is improved by expanding the prediction block selection range of the block to be encoded.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114697666B_ABST
    Figure CN114697666B_ABST
Patent Text Reader

Abstract

The embodiment of the present application discloses a screen encoding method, a screen decoding method and a related device, the method comprising: obtaining a block to be encoded; selecting a prediction block of the block to be encoded from a cache list; wherein the cache list includes image blocks in at least two frames of images; and outputting code stream data, wherein the code stream data includes a residual between the prediction block and the block to be encoded, and an identifier of the prediction block. By adopting the present application, the encoding performance can be improved and the bit rate can be reduced to a certain extent.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of video coding, and in particular to a screen coding method, a screen decoding method, and related devices. Background Art

[0002] In recent years, with the advent of high-definition and ultra-high-definition applications in people's lives, video coding technology has faced more and more challenges. As a new generation of video coding standard, High Efficiency Video Coding (HEVC) requires only 50% of the number of bits for encoding compared to the previous generation coding standard H.264 under the condition of maintaining the same video quality. Therefore, the HEVC video coding method has become one of the current research focuses.

[0003] Screen Content Coding (SCC) is one of the important extended applications of HEVC. SCC is similar to traditional HEVC and still based on a hybrid coding framework. And according to the characteristics of screen content videos, a series of new technologies are introduced respectively in the processes of intra prediction and inter prediction, such as Intra Block Copy (IBC) mode, Palette (PLT) mode, and block search technology based on Hash value.

[0004] However, the IBC technology uses the previously encoded sample block of the current frame as a prediction block to predict specific sample values. It can be seen that the search range of IBC only involves the encoded area of the current frame, and there is less data available for selection, resulting in low coding efficiency. Summary of the Invention

[0005] Embodiments of this application provide a screen coding method, a screen decoding method, and related devices, which can improve the coding performance and reduce the bit rate to a certain extent.

[0006] In a first aspect, embodiments of this application disclose a screen coding method, which includes: obtaining a block to be encoded; selecting a prediction block of the block to be encoded from a cache list, where the cache list includes image blocks in at least two frames of images; outputting bitstream data, where the bitstream data includes the residual between the prediction block and the block to be encoded, and an identifier of the prediction block.

[0007] In a possible implementation, after obtaining the block to be encoded and before selecting the prediction block of the block to be encoded from the cache list, it further includes: performing a matching determination on the image blocks in the cache list and the block to be encoded to obtain a determination result.

[0008] In a possible implementation, when the determination result is matching and similar, select the prediction block of the block to be encoded from the cache list.

[0009] In a possible implementation, when the judgment result is a non - matching similarity, obtain the reconstructed image block of the block to be encoded, and save the reconstructed image block to the cache list.

[0010] In a possible implementation, the block to be encoded belongs to the first frame of the image, and at least two frames of images include the first frame of the image.

[0011] In a second aspect, an embodiment of the present application discloses a screen decoding method, and the method includes: receiving bitstream data, where the bitstream data includes an identifier of a prediction block of the block to be decoded and a residual between the block to be decoded and the prediction block; selecting the prediction block of the block to be decoded from the cache list according to the identifier of the prediction block; where the cache list includes image blocks in at least two frames of images; determining the block to be decoded according to the prediction block and the residual.

[0012] In a possible implementation, after receiving the bitstream data and before selecting the prediction block of the block to be decoded from the cache list according to the identifier of the prediction block, it further includes: if the bitstream data does not include the identifier of the prediction block of the encoded block, decoding to obtain the reconstructed image block to be encoded; saving the reconstructed image block to the cache list.

[0013] In a possible implementation, the block to be encoded belongs to the first frame of the image, and at least two frames of images include the first frame of the image.

[0014] In a third aspect, an embodiment of the present application discloses a video encoder, and the video encoder includes:

[0015] An acquisition unit, configured to acquire a block to be encoded;

[0016] A selection unit, configured to select a prediction block of the block to be encoded from the cache list; where the cache list includes image blocks in at least two frames of images;

[0017] An output unit, configured to output bitstream data, where the bitstream data includes a residual between the prediction block and the block to be encoded and an identifier of the prediction block.

[0018] In a possible implementation, the selection unit is further configured to perform a matching judgment on the image block in the cache list and the block to be encoded to obtain a judgment result.

[0019] In a possible implementation, when the judgment result is a matching similarity, the selection unit is specifically configured to select a prediction block of the block to be encoded from the cache list.

[0020] In a possible implementation, when the judgment result is a non - matching similarity, the selection unit is specifically configured to obtain the reconstructed image block of the block to be encoded and save the reconstructed image block to the cache list.

[0021] In a possible implementation, the block to be encoded belongs to the first frame image, and at least two frame images include the first frame image.

[0022] Fourthly, an embodiment of the present application discloses a video decoder, which includes:

[0023] A receiving unit, configured to receive bitstream data, where the bitstream data includes an identifier of a prediction block of a block to be decoded, and a residual between the block to be decoded and the prediction block;

[0024] A selection unit, configured to select a prediction block of the block to be decoded from a cache list according to the identifier of the prediction block; wherein, the cache list includes image blocks in at least two frame images;

[0025] A determination unit, configured to determine the block to be decoded according to the prediction block and the residual.

[0026] In a possible implementation, the selection unit is further configured to: if the identifier of the prediction block of the encoded block is not included in the bitstream data, decode to obtain a reconstructed image block to be encoded; and save the reconstructed image block into the cache list.

[0027] In a possible implementation, the block to be encoded belongs to the first frame image, and at least two frame images include the first frame image.

[0028] Fifthly, an embodiment of the present application provides a computer-readable storage medium, in which a computer program is stored. When the computer program runs on one or more processors, it executes the method in the embodiment of the first aspect or the second aspect.

[0029] Sixthly, an embodiment of the present application provides a chip system, which includes at least one processor, a memory, and an interface circuit. A computer program is stored in the memory. When the computer program runs on one or more processors, it executes the method in the embodiment of the first aspect or the second aspect.

[0030] In the above method, reconstructed blocks of encoded pixels of multiple frame images are cached in the cache list, and these multiple frame images can come from the same video frame or different video frames. During the encoding process, the selectable range of the prediction block of the block to be encoded is larger, so the selectivity is higher, and the possibility of selecting the prediction block of the block to be encoded is increased, thereby improving the encoding quality and reducing the bit rate. Description of the Drawings

[0031] To more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0032] Figure 1 It is a schematic structural diagram of a video decoding system provided by an embodiment of the present application;

[0033] Figure 2 It is a schematic flowchart of a screen encoding method provided by an embodiment of the present application;

[0034] Figure 3 It is a schematic flowchart of a screen decoding method provided by an embodiment of the present application;

[0035] Figure 4 It is a schematic structural diagram of a video encoder provided by an embodiment of the present application;

[0036] Figure 5 It is a schematic structural diagram of a video decoder provided by an embodiment of the present application;

[0037] Figure 6 It is a schematic structural diagram of a composition processing device provided by an embodiment of the present application

[0038] Figure 7 It is a schematic structural diagram of a board card provided by an embodiment of the present application. Detailed implementation manners

[0039] The following will clearly and completely describe the technical solutions in the embodiments of the present application with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only some embodiments of the present application, rather than all embodiments. Based on the embodiments of the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present application.

[0040] The terms "first", "second", etc. in the specification and claims of the present application and the above drawings are used to distinguish different objects, rather than to describe a specific order. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units is not limited to the listed steps or units, but may optionally further include steps or units not listed, or may optionally further include other steps or units inherent to these processes, methods, products, or devices.

[0041] References to "embodiments" in this specification mean that a particular feature, structure, or characteristic described in connection with the embodiments can be included in at least one embodiment of the present application. The phrase appears in various places in the specification and does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment mutually exclusive of other embodiments. It is explicitly and implicitly understood by those skilled in the art that the embodiments described herein can be combined with other embodiments.

[0042] Some concepts that may be involved in the embodiments of the present application will be briefly introduced below.

[0043] 1. Video coding

[0044] Video coding refers to the method of converting a file in the original video format into another video format file through compression technology. Video is a sequence of continuous images, composed of continuous frames, and one frame is an image. Due to the visual persistence effect of the human eye, when the frame sequence is played at a certain rate, what we see is a video with continuous actions. Since the similarity between consecutive frames is extremely high, in order to facilitate storage and transmission, we need to encode and compress the original video to remove redundancy in the spatial and temporal dimensions.

[0045] In the field of video coding, the terms "Picture", "Frame", or "Image" can be used as synonyms. It can be understood that video coding is performed on the source side and generally includes processing (such as compression) of the original video to reduce the amount of data required to represent the video, so that it can be stored and / or transmitted more efficiently. Video decoding is performed on the destination side and generally includes performing inverse processing relative to the encoder to reconstruct the video.

[0046] 2. Screen content coding

[0047] Screen content coding is one of the important extensions of HEVC. Screen content refers to the content captured by the image display units of various devices (such as computers, mobile terminals, etc.). The screen content of a scene includes computer images and text images, natural videos and images / mixed text images, and computer-generated animated images, etc. Screen content is used in applications such as desktop collaboration, desktop sharing, cloud computing, cloud gaming, remote desktop, and remote display.

[0048] During the process of screen content encoding, each frame of the video is divided according to a quadtree-based partitioning structure. Before the start of the encoding process, each test sequence will be segmented into many groups of pictures (GOPs). During the partitioning process of the coding unit (CU), screen content encoding still adopts a quadtree-based partitioning structure. First, each frame of the video is divided into coding tree units (CTUs) with a size of 64×64, and then each CTU is further divided into CUs, prediction units (PUs), and transform units (TUs) with different sizes.

[0049] Among them, the size distribution of the CU can be iteratively divided from 64×64 to 8×8. At the same time, a CU can be further divided into one or more PUs. The size of the PU can be iteratively divided from 64×64 to 4×4.

[0050] First, Figure 1 is a schematic structural diagram of a video decoding system provided by an embodiment of the present application. As used herein, the term "video decoder" generally refers to both a video encoder and a video decoder. In the embodiments of the present application, the term "video decoding" or "decoding" may generally refer to video encoding or video decoding.

[0051] Please refer to Figure 1 , Figure 1 As shown, the video decoding system includes a source device 10 and a destination device 20. The source device 10 generates encoded video data, so the source device 10 can be referred to as a video encoding device. The destination device 20 can decode the encoded video data generated by the source device 10. Therefore, the destination device 20 can be referred to as a video decoding device. Various implementations of the source device 10 and the destination device 20 or both may include one or more processors and a memory coupled to the one or more processors. The above memory may include, but is not limited to, RAM, ROM, EEPROM, flash memory, or any other medium that can be used to store the desired program code in the form of computer-accessible instructions or data structures.

[0052] The source device 10 and the destination device 20 can be part of or independent units of a video broadcast system, a cable system, a network-based video streaming service, a gaming application and / or service, a multimedia communication system, and / or various other applications and services. The source device 10 and the destination device 20 can include various devices, including desktop computers, mobile computing devices, notebook (e.g., laptop) computers, tablet computers, set-top boxes, telephone handsets such as so-called "smart" phones, televisions, cameras, display devices, digital media players, video game consoles, in-vehicle computers, or the like.

[0053] Although Figure 1 the source device 10 and the destination device 20 are depicted as separate devices, device embodiments can also include both the source device 10 and the destination device 20 simultaneously or include the functionality of both, i.e., the source device 10 or the corresponding functionality and the destination device 20 or the corresponding functionality. In such embodiments, the source device 10 or the corresponding functionality and the destination device 20 or the corresponding functionality can be implemented using the same hardware and / or software, or using separate hardware and / or software, or any combination thereof.

[0054] From Figure 1 it can be seen that the source device 10 includes a video source 101, a video encoder 100, and an output interface 102. In some embodiments, the output interface 102 can include a regulator / demodulator (modem) and / or a transmitter. The video source 101 can include a video capture device (e.g., a camera), a video archive containing previously captured video data, a video feed interface for receiving video data from a video content provider, and / or a computer graphics system for generating video data, or a combination of video data sources.

[0055] The video encoder 100 can encode the video data from the video source 101 to obtain bitstream data. Specifically, the video encoder 100 selects a predicted block of the block to be encoded from a buffer list, the block to be encoded is an image block in any frame image of the video data, and the candidate list includes image blocks in at least two frames of images. In some embodiments, the source device 100 transmits the encoded video data to the destination device 20 via the output interface 102.

[0056] From Figure 1It can be seen that the destination device 20 includes an input interface 201, a video decoder 200, and a display device 202. The input interface 201 includes a receiver and / or a modem. The input interface 201 can receive encoded video data via the link 30. The display device 202 can be integrated with the destination device 20 or can be external to the destination device 20. Generally, the display device 220 displays the decoded video data. The display device 202 can include various display devices, for example, a liquid crystal display (LCD), a plasma display, an organic light emitting diode (OLED) display, or other types of display devices.

[0057] See Figure 2 , Figure 2 is a schematic flowchart of a screen encoding method provided by an embodiment of the present application. Figure 2 The process shown can be executed by Figure 1 the video encoder 100 in Figure 2 It can be understood that the process shown can be executed in various orders and / or occur simultaneously, not limited to the Figure 2 execution order shown. This method can include but is not limited to the following steps:

[0058] Step S201, obtain the block to be encoded.

[0059] Specifically, the video encoder 100 can obtain video data from the video source 101, or after receiving the video data, the video encoder 100 stores the video data in the video data storage unit. The video encoder 100 can divide the video data into several image blocks, and these image blocks can be further divided into smaller blocks, for example, image block segmentation based on a quadtree structure or a binary tree structure. This segmentation can also include segmentation into strips, tiles, or other larger units. The video encoder 100 generally describes the component that encodes the image blocks within the video strip to be encoded. The strip can be divided into multiple image blocks and may be divided into sets of image blocks called tiles.

[0060] An image block may have a fixed or variable size and may vary in size according to different video compression and coding standards. The size of the image block is M×N pixels. "M×N" and "M multiplied by N" can be used interchangeably to refer to the pixel dimensions of the image block in the horizontal and vertical dimensions, that is, having M pixels in the horizontal direction and N pixels in the vertical direction, where M and N represent non-negative integer values. In addition, the block does not necessarily need to have the same number of pixels in the horizontal and vertical directions. For example, here M = N = 4. Of course, the size of the sub-block of the current image block and the size of the reference block can also be 8×8 pixels, 8×4 pixels, or 4×8 pixels, or the smallest prediction block size. The image blocks described in the embodiments of the present application can be understood but not limited to: CU, PU, or TU, etc. According to the regulations of different video compression and coding standards, a CU may contain one or more PUs, or the sizes of the PU and the CU are the same. In addition, the block to be encoded may refer to the image block that needs to be encoded currently.

[0061] After obtaining the block to be encoded by partitioning the image frame, the position (x, y, W, H) of the target in the block to be encoded can be calculated, where W and H are the length and width of the target, and x and y are the coordinates of the target in the current image block. The target can be text, an image, etc. in the block to be encoded. For example, for the text in video data, all the text in the current image frame can be detected by using OCR calculation, and then the current image frame is partitioned to obtain the block to be encoded (i.e., the image block that needs to be encoded currently). At this time, the block to be encoded can be a PU, and then the position (x, y, W, H) of the text in the PU in the current PU is calculated.

[0062] Step S202, obtain the judgment result.

[0063] Specifically, the video encoder 100 performs a matching judgment on the block to be encoded and the image blocks in the cache list to obtain a judgment result. There are N image blocks stored in the cache list, so the block to be encoded needs to be matched with each of the N image blocks to obtain N judgment results. Among them, the matching method can be to perform a subtraction operation on the block to be encoded and the image blocks in the cache list, and the obtained difference can be the judgment result.

[0064] In a possible implementation manner, when the judgment result is matching and similar, step S203 is executed to select the prediction block of the block to be encoded from the cache list. Specifically, the judgment result can be the difference between the block to be encoded and the image blocks in the cache list. The smaller the difference, the more similar the image block in the cache list corresponding to the difference is to the block to be encoded. Therefore, this image block can be used as the prediction block of the block to be encoded.

[0065] In a possible implementation, when the judgment result is a non - matching similarity, step S204 is executed to save the reconstructed image block of the block to be encoded into the cache list. Specifically, the judgment result can be the difference between the block to be encoded and the image blocks in the cache list. The larger the difference, the less similar the image block in the cache list corresponding to the above - mentioned difference is to the block to be encoded. Therefore, there is no predicted block of the block to be encoded in the cache list.

[0066] Step S203: Select the predicted block of the block to be encoded from the cache list.

[0067] Specifically, after the video encoder 100 obtains the block to be encoded, the video encoder can traverse the cache list in a pre - specified order to select the predicted block of the block to be encoded. Among them, the cache list includes image blocks in at least two frames, and the above - mentioned image values are the reconstructed values of the encoded image blocks. That is to say, in the cache list, not only the reconstructed values of the encoded image blocks of the current frame are cached, but also the reconstructed values of the encoded image blocks of the previous frame or several previous frames of the current frame, or the reconstructed values of the encoded image blocks of the image frames of other video data. The embodiments of the present application do not make any restrictions.

[0068] It can be understood that there may be a difference in the sizes between the block to be encoded and the image blocks in the cache list. Therefore, it is necessary to perform size normalization processing on the block to be encoded and the image blocks in the cache list. If the size of the block to be encoded is, for example, 4×4, and the sizes of the image blocks in the cache list are, for example, 8×4, 8×8, the information of the central 4×4 block can be obtained as the prediction information of the block to be encoded. The coordinates of the upper - left vertex of the central 4×4 block relative to the upper - left vertex of the image block in the cache list are ((W / 4) / 2*4, (H / 4) / 2*4), where the division operation is an integer division operation. If M = 8 and N = 4, then the coordinates of the upper - left vertex of the central 4×4 block relative to the upper - left vertex of the image block in the cache list are (4, 0). Optionally, the information of the upper - left 4x4 block of the image block in the cache list can also be obtained as the prediction information of the block to be encoded. The encoder can encode the position coordinate information of intercepting the 4×4 block to indicate the decoder to intercept the predicted block from the predicted block according to the corresponding coordinates, but the present application is not limited thereto.

[0069] Step S204: Save the reconstructed image block of the block to be encoded into the cache list.

[0070] Specifically, if the result of the matching determination is dissimilar matching, that is, there is no predicted block of the block to be encoded in the cache list, the block to be encoded containing the target can be intercepted according to the calculated position (x, y, W, H), and then the reconstructed image block of the block to be encoded can be obtained according to other encoding methods, and the reconstructed image block of the block to be encoded is saved to the cache list, which can be used by the video encoder 100 as a reference block for intra prediction of blocks in subsequent video frames or images.

[0071] Step S205, output bitstream data.

[0072] Specifically, after the video encoder 100 selects the predicted block of the block to be encoded from the cache list, the video encoder 100 forms a residual image block by subtracting the predicted block from the current image block to be encoded. The residual video data in the residual block can be included in one or more TUs, and then the residual video data can be transformed into residual transform coefficients using a transform such as a discrete cosine transform (DCT) or a conceptually similar transform. Further, the residual video data can also be converted from the pixel value domain to the transform domain, such as the frequency domain. When the sum of the transform coefficients is obtained, the video encoder 100 can quantize the transform coefficients to further reduce the bit rate. After quantization, the quantized coefficients can also be entropy encoded. After entropy encoding, the encoded bitstream data can be transmitted to the video decoder 200 or archived for later transmission.

[0073] In a possible implementation, when the determination result is similar matching, that is, when there is a predicted block of the block to be encoded in the cache list, the output bitstream data includes the residual between the predicted block and the block to be encoded, and the identifier of the predicted block.

[0074] In a possible implementation, when the determination result is dissimilar matching, that is, when there is no predicted block of the block to be encoded in the cache list, the output bitstream data can be the reconstructed image block of the block to be encoded.

[0075] It should be noted that the image blocks cached in the cache list and the block to be encoded may not necessarily come from the same video frame. That is, the block to be encoded can be any image frame in the first video frame, the image blocks cached in the cache list can be any image frame in the first video frame, or can be any image frame in the second video frame, and the first video frame is not the second video frame.

[0076] See Figure 3 , Figure 3 is a schematic flowchart of a screen decoding method provided by an embodiment of the present application. Figure 3 The shown process can be executed by Figure 1 the video decoder 100 in. It can be understood that Figure 3 the shown process can be executed and / or occur simultaneously in various orders, not limited toFigure 3 The execution order shown. This method may include but is not limited to the following steps:

[0077] Step S301, receiving bitstream data.

[0078] Specifically, during the decoding process, the video decoder 200 receives bitstream data representing image blocks of an encoded video strip from the video encoder 100 through the input interface 202. The video decoder 200 can also store the bitstream data in the video data storage unit, which can serve as a decoded image buffer unit for storing encoded video data from the encoded video bitstream. Then, the video decoder 100 can parse the bitstream data through the entropy decoding unit. If the parsed bitstream data contains the residual between the block to be decoded and the predicted block, step S302 is executed; if the parsed bitstream data does not contain the residual between the block to be decoded and the predicted block, step S304 is executed.

[0079] It should be noted that the block to be decoded mentioned in the embodiments of the present application is the image block that needs to be decoded currently.

[0080] Step S302, selecting the predicted block of the block to be decoded from the cache list according to the identifier of the predicted block.

[0081] Specifically, the decoder 200 selects the predicted block of the block to be decoded from the cache list according to the identifier of the predicted block. It can be understood that each image block stored in the cache list has its corresponding identifier. There is a cache list for image blocks maintained on the encoding side, and a cache list for image blocks is also maintained on the decoding side in the same way. Therefore, the corresponding image block can be selected from the cache list on the decoding side as the predicted block of the block to be decoded according to the identifier.

[0082] S303, determining the block to be decoded according to the predicted block and the residual.

[0083] Specifically, when the decoder 200 selects the predicted value of the block to be encoded from the cache list, the decoder 200 sums the parsed residual and the predicted value to obtain the reconstructed block, that is, the decoded image block.

[0084] Step S304, saving the reconstructed image block to the cache list.

[0085] Specifically, if the decoder 200 parses the bitstream data and does not obtain an identifier, it means that there is no image block of the block to be encoded in the cache list. Then, the reconstructed image block of the block to be encoded obtained by parsing is saved to the cache list, and the reconstructed image block in the cache list has its corresponding identifier.

[0086] Please refer to Figure 4 , Figure 4FIG. 0 is a schematic structural diagram of a video encoder 400 provided by an embodiment of the present application. The video encoder 400 may be a node or a component in a node, such as a chip or an integrated circuit. As shown in the figure, the video encoder 400 may include an acquisition unit 401, a selection unit 402, and an output unit 403. The descriptions of each unit are as follows:

[0087] The acquisition unit 401 is configured to acquire a block to be encoded.

[0088] The selection unit 402 is configured to select a prediction block of the block to be encoded from a cache list. The cache list includes image blocks in at least two frames of images.

[0089] The output unit 403 is configured to output bitstream data. The bitstream data includes the residual between the prediction block and the block to be encoded, and an identifier of the prediction block.

[0090] In a possible implementation, the selection unit 402 is further configured to: perform a matching determination between the image blocks in the cache list and the block to be encoded to obtain a determination result.

[0091] In a possible implementation, when the determination result is that they are similar in matching, the selection unit 402 is specifically configured to select a prediction block of the block to be encoded from the cache list.

[0092] In a possible implementation, when the determination result is that they are not similar in matching, the acquisition unit 401 is further configured to acquire a reconstructed image block of the block to be encoded and save the reconstructed image block to the cache list.

[0093] In a possible implementation, the block to be encoded belongs to the first frame of image, and the at least two frames of images include the first frame of image.

[0094] Please refer to Figure 5 , Figure 5 FIG. 28 is a schematic structural diagram of a video decoder 500 provided by an embodiment of the present application. The video encoder 500 may be a node or a component in a node, such as a chip or an integrated circuit. As shown in the figure, the video encoder 500 may include a receiving unit 501, a selection unit 502, and a determination unit 503. The descriptions of each unit are as follows:

[0095] The receiving unit 501 is configured to receive bitstream data. The bitstream data includes an identifier of a prediction block of a block to be decoded, and a residual between the block to be decoded and the prediction block.

[0096] A selection unit 502 is configured to select a prediction block of the block to be decoded from a cache list according to an identifier of the prediction block, where the cache list includes image blocks in at least two frames of images.

[0097] A determination unit 503 is configured to determine the block to be decoded according to the prediction block and the residual.

[0098] In a possible implementation, the selection unit 502 is further configured to: if an identifier of a prediction block of the encoded block is not included in the bitstream data, decode a reconstructed image block to be encoded; and save the reconstructed image block into the cache list.

[0099] In a possible implementation, the block to be encoded belongs to a first frame of image, and the at least two frames of images include the first frame of image.

[0100] It should be noted that the implementation of each unit may also correspond to the corresponding description of the Figure 3 illustrated embodiment.

[0101] Please refer to Figure 6 , Figure 6 FIG. is a schematic structural diagram of a composition processing apparatus 600 provided by an embodiment of the present application. The composition processing apparatus 600 may be a node or a device in a node, such as a chip or an integrated circuit. The composition processing apparatus 60 may include at least one memory 601 and at least one processor 602. Optionally, a bus 603 may also be included. Further optionally, a communication interface 604 may also be included, where the memory 601, the processor 602, and the communication interface 604 are connected through the bus 603.

[0102] Among them, the memory 601 is used to provide a storage space, and data such as an operating system and a computer program may be stored in the storage space. The memory 601 may be one or a combination of a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM), or a compact disc read-only memory (CD-ROM), etc.

[0103] The processor 602 is a module that performs arithmetic operations and / or logical operations, and can specifically be one or a combination of multiple processing modules such as a central processing unit (CPU), a graphics processing unit (GPU), a microprocessor unit (MPU), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA), and a complex programmable logic device (CPLD).

[0104] The communication interface 604 is used to receive data sent from the outside and / or send data to the outside, and can be a wired link interface including, for example, an Ethernet cable, or a wireless link (Wi-Fi, Bluetooth, general wireless transmission, etc.) interface. Optionally, the communication interface 604 can also include a transmitter (such as a radio frequency transmitter, an antenna, etc.) coupled to the interface, or a receiver, etc.

[0105] In some possible implementation manners, the processor 602 in the composition processing device 600 is used to read the computer program stored in the memory 601 and execute the foregoing screen encoding method, for example Figure 2 the video encoding method described in the embodiments. Specifically, it is used to execute:

[0106] Obtain a block to be encoded; select a prediction block of the block to be encoded from a cache list; where the cache list includes image blocks in at least two frames of images; output bitstream data, where the bitstream data includes the residual between the prediction block and the block to be encoded, and the identifier of the prediction block.

[0107] In one possible implementation manner, the processor 602 is specifically used to: perform a matching determination on the image blocks in the cache list and the block to be encoded to obtain a determination result.

[0108] In one possible implementation manner, the processor 602 is specifically used to: select a prediction block of the block to be encoded from the cache list when the determination result is that they are similar in matching.

[0109] In one possible implementation manner, the processor 602 is specifically used to: obtain a reconstructed image block of the block to be encoded and save the reconstructed image block to the cache list when the determination result is that they are not similar in matching.

[0110] In a possible implementation, the block to be encoded belongs to the first frame image, and the at least two frames of images include the first frame image.

[0111] In some possible implementations, the processor 602 in the composition processing device 600 is configured to read the computer program stored in the memory 601 and execute the foregoing screen decoding method, such as Figure 3 the video decoding method described in the embodiments. Specifically, it is configured to execute:

[0112] Receive bitstream data, where the bitstream data includes an identifier of a prediction block of the block to be decoded and a residual between the block to be decoded and the prediction block;

[0113] Select the prediction block of the block to be decoded from the cache list according to the identifier of the prediction block; wherein, the cache list includes image blocks in at least two frames of images;

[0114] Determine the block to be decoded according to the prediction block and the residual.

[0115] In a possible implementation, the processor 602 is specifically configured to: if the identifier of the prediction block of the encoded block is not included in the bitstream data, decode to obtain the reconstructed image block to be encoded; save the reconstructed image block to the cache list.

[0116] In a possible implementation, the block to be encoded belongs to the first frame image, and the at least two frames of images include the first frame image.

[0117] Please refer to Figure 7 , Figure 7 which is a schematic structural diagram of a board 700 provided by an embodiment of the present application. As can be seen from Figure 3 the board includes a storage device 704 for storing data, which includes one or more storage units 710. The storage device can be connected and data-transferred to the control device 708 and the chip 702 described above through, for example, a bus. Further, the board further includes an external interface device 706, which is configured to perform data relaying or switching functions between a chip (or a chip in a chip package structure) and an external device 76 (such as a server or a computer, etc.). For example, the data to be processed can be transmitted from the external device to the chip through the external interface device. Again, for example, the calculation result of the chip can be transmitted back to the external device via the external interface device. According to different application scenarios, the external interface device can have different interface forms. For example, it can adopt a standard PCIE interface, etc.

[0118] The chip 702 can be a system-on-chip (SoC) and integrates one or more such as Figure 6The combined processing device shown in. The chip can be connected to other related components through an external interface device (such as Figure 7 the external interface device 706 shown in). The related components can be, for example, a camera, a display, a mouse, a keyboard, a network card, or a wifi interface. In some application scenarios, other processing units (such as video codecs) and / or interface modules (such as DRAM interfaces) can be integrated on the chip.

[0119] In one or more embodiments, the control device in the disclosed board can be configured to regulate the state of the chip. For this purpose, in one application scenario, the control device can include a microcontroller unit (MCU) for regulating the operating state of the chip.

[0120] According to the above combination Figure 6 and Figure 7 description, those skilled in the art can understand that the present disclosure also discloses an electronic device or apparatus, which can include one or more of the above boards, one or more of the above chips, and / or one or more of the above combined processing devices.

[0121] According to different application scenarios, the electronic devices or apparatuses disclosed in this disclosure may include servers, cloud servers, server clusters, data processing devices, robots, computers, printers, scanners, tablets, smart terminals, PC devices, Internet of Things (IoT) terminals, mobile terminals, mobile phones, dash cams, navigators, sensors, cameras, video cameras, projectors, watches, earphones, mobile storage devices, wearable devices, vision terminals, autonomous driving terminals, transportation means, household appliances, and / or medical devices. The transportation means include airplanes, ships, and / or vehicles; the household appliances include televisions, air conditioners, microwave ovens, refrigerators, rice cookers, humidifiers, washing machines, electric lights, gas stoves, and range hoods; the medical devices include nuclear magnetic resonance (NMR) spectrometers, B-ultrasound devices, and / or electrocardiographs. The electronic devices or apparatuses disclosed in this disclosure may also be applied to fields such as the Internet, IoT, data centers, energy, transportation, public management, manufacturing, education, power grids, telecommunications, finance, retail, construction sites, and healthcare. Further, the electronic devices or apparatuses disclosed in this disclosure may also be used in application scenarios related to artificial intelligence, big data, and / or cloud computing, such as the cloud, edge, and terminal. In one or more embodiments, the electronic devices or apparatuses with high computing power according to the solution of this disclosure may be applied to cloud devices (such as cloud servers), while the electronic devices or apparatuses with low power consumption may be applied to terminal devices and / or edge devices (such as smartphones or cameras). In one or more embodiments, the hardware information of cloud devices is compatible with the hardware information of terminal devices and / or edge devices, so that appropriate hardware resources can be matched from the hardware resources of cloud devices according to the hardware information of terminal devices and / or edge devices to simulate the hardware resources of terminal devices and / or edge devices, so as to complete the unified management, scheduling, and collaborative work of end-cloud integration or cloud-edge-end integration.

[0122] It should be noted that, for the purpose of simplicity, some methods and their embodiments in this disclosure are expressed as a series of actions and their combinations. However, those skilled in the art can understand that the solution of this disclosure is not limited by the order of the described actions. Therefore, based on the disclosure or teachings of this disclosure, those skilled in the art can understand that some of the steps can be executed in other orders or simultaneously. Further, those skilled in the art can understand that the embodiments described in this disclosure can be regarded as optional embodiments, that is, the actions or modules involved are not necessarily required for the implementation of certain solutions of this disclosure. In addition, according to the differences in the solutions, the descriptions of some embodiments in this disclosure also have different focuses. In view of this, those skilled in the art can understand that the parts not detailed in a certain embodiment of this disclosure can also refer to the relevant descriptions of other embodiments.

[0123] In terms of specific implementation, based on the disclosure and teachings of the present disclosure, those skilled in the art can understand that several embodiments disclosed in the present disclosure can also be implemented in other ways not disclosed herein. For example, for each unit in the foregoing embodiments of the electronic device or apparatus, it is divided herein based on consideration of logical functions, and there may be other division methods in actual implementation. For another example, multiple units or components can be combined or integrated into another system, or some features or functions of the units or components can be selectively disabled. Regarding the connection relationship between different units or components, the connections discussed in conjunction with the accompanying drawings above can be direct or indirect couplings between the units or components. In some scenarios, the foregoing direct or indirect couplings involve communication connections using interfaces, where the communication interface can support signal transmission in electrical, optical, acoustic, magnetic, or other forms.

[0124] In the present disclosure, the units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units. The foregoing components or units may be located at the same position or distributed across multiple network units. Additionally, according to actual needs, some or all of the units can be selected to achieve the objectives of the solutions described in the embodiments of the present disclosure. Further, in some scenarios, multiple units in the embodiments of the present disclosure can be integrated into one unit or each unit physically exists separately.

[0125] In some implementation scenarios, the above-mentioned integrated units can be implemented in the form of software program modules. If implemented in the form of software program modules and sold or used as an independent product, the integrated units can be stored in a computer-readable memory. Based on this, when the solution of the present disclosure is embodied in the form of a software product (such as a computer-readable storage medium), the software product can be stored in the memory, which may include several instructions for causing a computer device (such as a personal computer, a server, or a network device, etc.) to execute some or all of the steps of the method described in the embodiments of the present disclosure. The foregoing memory may include, but is not limited to, various media that can store program codes, such as USB flash drives, flash memory drives, read-only memory (ROM), random access memory (RAM), mobile hard disks, magnetic disks, or optical discs.

[0126] In some other implementation scenarios, the above integrated units can also be implemented in the form of hardware, i.e., a specific hardware circuit, which can include digital circuits and / or analog circuits, etc. The physical implementation of the hardware structure of the circuit can include, but is not limited to, physical devices, and the physical devices can include, but are not limited to, devices such as transistors or memristors. In view of this, various devices described herein (such as computing devices or other processing devices) can be implemented by appropriate hardware processors, such as CPUs, GPUs, FPGAs, DSPs, and ASICs, etc. Further, the aforementioned storage unit or storage device can be any appropriate storage medium (including magnetic storage media or magneto-optical storage media, etc.), which can be, for example, Resistive Random Access Memory (RRAM), Dynamic Random Access Memory (DRAM), Static Random Access Memory (SRAM), Enhanced Dynamic Random Access Memory (EDRAM), High Bandwidth Memory (HBM), Hybrid Memory Cube (HMC), ROM, and RAM, etc.

[0127] Although multiple embodiments of the present disclosure have been shown and described herein, it will be apparent to those skilled in the art that such embodiments are provided by way of example only. Many changes, variations, and alternative approaches can be contemplated by those skilled in the art without departing from the spirit and scope of the present disclosure. It should be understood that various alternatives to the embodiments of the present disclosure described herein can be employed in practicing the present disclosure. The appended claims are intended to define the scope of protection of the present disclosure and thus cover equivalents or alternatives within the scope of these claims.

Claims

1. A screen encoding method, characterized in that, comprising: obtaining a block to be encoded; matching and judging the image blocks in the first cache list with the block to be encoded to obtain a judgment result; wherein, the first cache list includes image blocks in at least two frames of images; the image value of the image block is the reconstruction value of the encoded image block; when the judgment result is that they are similar in matching, selecting a prediction block of the block to be encoded from the first cache list; outputting bitstream data, wherein the bitstream data includes the residual between the prediction block and the block to be encoded, and the identifier of the prediction block.

2. The method according to claim 1, characterized in that, when the judgment result is that they are not similar in matching, obtaining a reconstructed image block of the block to be encoded, and saving the reconstructed image block into the first cache list.

3. The method according to claim 1 or 2, characterized in that, the block to be encoded belongs to the first frame of image, and the at least two frames of images include the first frame of image.

4. A screen decoding method, characterized in that, comprising: receiving bitstream data, the bitstream data including the identifier of the prediction block of the block to be decoded, and the residual between the block to be decoded and the prediction block; selecting the prediction block of the block to be decoded from the second cache list according to the identifier of the prediction block; wherein, the second cache list includes image blocks in at least two frames of images; determining the block to be decoded according to the prediction block and the residual.

5. The method according to claim 4, characterized in that, after receiving the bitstream data and before selecting the prediction block of the block to be decoded from the cache list according to the identifier of the prediction block, further comprising: if the bitstream data does not include the identifier of the prediction block of the encoded block, decoding to obtain the reconstructed image block to be encoded; saving the reconstructed image block into the second cache list.

6. The method according to claim 4 or 5, characterized in that, the block to be encoded belongs to the first frame of image, and the at least two frames of images include the first frame of image.

7. A video encoder, characterized in that, comprising: an obtaining unit, configured to obtain a block to be encoded; matching and judging the image blocks in the first cache list with the block to be encoded to obtain a judgment result; wherein, the first cache list includes image blocks in at least two frames of images; the image value of the image block is the reconstruction value of the encoded image block; a selecting unit, configured to select a prediction block of the block to be encoded from the first cache list when the judgment result is that they are similar in matching; an output unit, configured to output bitstream data, wherein the bitstream data includes the residual between the prediction block and the block to be encoded, and the identifier of the prediction block.

8. A video decoder, characterized in that, comprising: a receiving unit, configured to receive bitstream data, the bitstream data including the identifier of the prediction block of the block to be decoded, and the residual between the block to be decoded and the prediction block; a selecting unit, configured to select the prediction block of the block to be decoded from the second cache list according to the identifier of the prediction block; wherein, the second cache list includes image blocks in at least two frames of images; A determination unit for determining the block to be decoded according to the predicted block and the residual.

9. A computer-readable storage medium, characterized in that, a computer program is stored in the computer-readable storage medium, and when the computer program runs on one or more processors, the method according to any one of claims 1-6 is executed.

10. A chip system, characterized in that, the chip system includes at least one processor, a memory and an interface circuit, a computer program is stored in the memory, and when the computer program runs on one or more processors, the method according to any one of claims 1-6 is executed.

Citation Information

Patent Citations

  • Video encoding method and device, video decoding method and device and computer storage medium

    CN111435989A

  • Candidate motion vector list acquisition method and device and codec

    CN111953997A