Pixel value prediction method and related equipment

By utilizing the fusion of a plurality of first pixel values and second pixel values in the video or image encoding process, the problem of insufficient pixel value prediction accuracy in the prior art is solved, and the prediction effect and performance of the encoding and decoding are improved.

CN120302047APending Publication Date: 2025-07-11VIVO MOBILE COMM CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410037037.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-01-09
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

During the existing video or image encoding process, the accuracy of the pixel value prediction method is insufficient, resulting in poor encoding and decoding performance.

Method used

By determining a plurality of first pixel values in the current processing unit and fusing in combination with the second pixel values in at least one second unit, the number of pixel values in the first unit is ensured to be greater than the number in the second unit, thereby improving prediction accuracy.

Benefits of technology

Improve the prediction accuracy of pixel values, thereby improving the prediction effect and performance of codecs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120302047A_ABST
    Figure CN120302047A_ABST
Patent Text Reader

Abstract

The invention discloses a pixel value prediction method and related equipment, and relates to the technical field of coding and decoding, and the method comprises the steps: determining a plurality of first pixel values in a first unit corresponding to a current processing unit based on a first position where a current processing pixel point in the current processing unit is located; the first pixel value is a pixel reconstruction value or a pixel prediction value; determining at least one second pixel value in at least one second unit corresponding to the current processing unit based on the first position; the second pixel value is a pixel reconstruction value or a pixel prediction value, and the number of the plurality of first pixel values is greater than the number of the second pixel values determined in any second unit in the at least one second unit; and fusing the plurality of first pixel values and the at least one second pixel value to obtain a predicted pixel value of the current processing pixel point. According to the method, the prediction accuracy of the pixel value can be improved, so that the prediction effect and the coding and decoding performance can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of encoding and decoding, and more particularly, to a pixel value prediction method and related devices. Background Art

[0002] During the encoding process of a video or an image, the image is divided into multiple encoding blocks. The encoder can obtain a prediction block of the current encoding block using various prediction modes, and then take the difference between the original block and the prediction block as the residual block, and perform transformation, quantization, and entropy encoding on the residual block to obtain a bitstream. Correspondingly, during the decoding process of a video or an image, the decoder can obtain the residual block of the current decoding block and the corresponding prediction mode by decoding the bitstream, then obtain the prediction block of the current decoding block using the corresponding prediction mode, and take the sum of the residual block and the prediction block as the reconstructed block, thereby realizing the reconstruction of the image.

[0003] In addition, the encoder or decoder can also obtain the final prediction block by fusing multiple prediction blocks.

[0004] However, with the development of technology, there is still a need to pursue better prediction methods to improve the prediction accuracy, and thus improve the prediction effect and encoding / decoding performance. Summary of the Invention

[0005] Embodiments of this application provide a pixel value prediction method and related devices, which can improve the prediction accuracy of pixel values, and thus can improve the prediction effect and encoding / decoding performance.

[0006] In a first aspect, a pixel value prediction method is provided, which is executed by a decoding end or an encoding end. The method includes:

[0007] Based on a first position where a current processing pixel point is located in a current processing unit, determine a plurality of first pixel values in a first unit corresponding to the current processing unit; the first pixel values are pixel reconstruction values or pixel prediction values;

[0008] Based on the first position, determine at least one second pixel value in at least one second unit corresponding to the current processing unit; the second pixel values are pixel reconstruction values or pixel prediction values, and the number of the plurality of first pixel values is greater than the number of second pixel values determined in any one of the at least one second unit;

[0009] Fuse the plurality of first pixel values and the at least one second pixel value to obtain a predicted pixel value of the current processing pixel point.

[0010] In a second aspect, a pixel value prediction device is provided, including:

[0011] A determination unit, configured to:

[0012] Based on the first position where the currently processed pixel point is located in the current processing unit, determine a plurality of first pixel values in the first unit corresponding to the current processing unit; the first pixel value is a pixel reconstruction value or a pixel prediction value.

[0013] Based on the first position, determine at least one second pixel value in at least one second unit corresponding to the current processing unit; the second pixel value is a pixel reconstruction value or a pixel prediction value, and the number of the plurality of first pixel values is greater than the number of second pixel values determined in any one of the at least one second unit.

[0014] A fusion unit is configured to fuse the plurality of first pixel values and the at least one second pixel value to obtain a predicted pixel value of the currently processed pixel point.

[0015] In a third aspect, an electronic device is provided. The terminal includes a processor and a memory. The memory stores a program or instruction that can be run on the processor. When the program or instruction is executed by the processor, the steps of the method described in the first aspect are implemented.

[0016] In a fourth aspect, an electronic device is provided, including a processor and a communication interface. Among them, the processor is configured to:

[0017] Based on the first position where the currently processed pixel point is located in the current processing unit, determine a plurality of first pixel values in the first unit corresponding to the current processing unit; the first pixel value is a pixel reconstruction value or a pixel prediction value.

[0018] Based on the first position, determine at least one second pixel value in at least one second unit corresponding to the current processing unit; the second pixel value is a pixel reconstruction value or a pixel prediction value, and the number of the plurality of first pixel values is greater than the number of second pixel values determined in any one of the at least one second unit.

[0019] Fuse the plurality of first pixel values and the at least one second pixel value to obtain a predicted pixel value of the currently processed pixel point.

[0020] In a fifth aspect, an electronic device is provided, including: a memory configured to store video data, and a processing circuit configured to implement the steps of the method described in the first aspect.

[0021] In a sixth aspect, a readable storage medium is provided. The readable storage medium stores a program or instruction. When the program or instruction is executed by a processor, the steps of the method described in the first aspect are implemented.

[0022] In a seventh aspect, there is provided an encoding and decoding system, including: an encoding end device and a decoding end device, both the encoding end device and the decoding end device can be used to execute the steps of the method described in the first aspect.

[0023] In an eighth aspect, there is provided a chip, the chip includes a processor and a communication interface, the communication interface is coupled to the processor, and the processor is used to run programs or instructions to implement the steps of the method described in the first aspect.

[0024] In a ninth aspect, there is provided a computer program / program product, the computer program / program product is stored in a storage medium, and the program / program product is executed by at least one processor to implement the steps of the method described in the first aspect.

[0025] In the embodiments of the present application, the pixel value prediction method includes: based on the first position where the current processing pixel point is located in the current processing unit, determining a plurality of first pixel values in the first unit corresponding to the current processing unit; the first pixel value is a pixel reconstruction value or a pixel prediction value; based on the first position, determining at least one second pixel value in at least one second unit corresponding to the current processing unit; the second pixel value is a pixel reconstruction value or a pixel prediction value, and the number of the plurality of first pixel values is greater than the number of second pixel values determined in any one of the at least one second unit; fusing the plurality of first pixel values and the at least one second pixel value to obtain a predicted pixel value of the current processing pixel point. Since when fusing the plurality of first pixel values and the at least one second pixel value, the number of the plurality of first pixel values is greater than the number of second pixel values determined in any one of the second units, equivalently, the predicted pixel value of the current processing pixel point can be obtained by increasing the number of first pixel values in the first unit, that is, the prediction accuracy of the pixel value can be improved, and further the prediction effect and the encoding and decoding performance can be improved. Description of the Drawings

[0026] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments of the present application. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0027] Figure 1 is a schematic diagram of the encoding and decoding system provided by the embodiments of the present application.

[0028] Figure 2 is a schematic structural diagram of the encoder provided by the embodiments of the present application.

[0029] Figure 3 It is a schematic structural diagram of a decoder provided by an embodiment of the present application.

[0030] Figure 4 It is an example of an intra mode used in VVC provided by an embodiment of the present application.

[0031] Figure 5 It is an example of the position of reconstructed pixel values used when determining model parameters provided by an embodiment of the present application.

[0032] Figure 6 It is an example of the principle of IntraTMP mode provided by an embodiment of the present application.

[0033] Figure 7 It is a schematic flowchart of a pixel value prediction method provided by an embodiment of the present application.

[0034] Figure 8 It is an example of the positions of multiple first pixel values provided by an embodiment of the present application.

[0035] Figure 9 It is a schematic block diagram of a pixel value prediction device provided by an embodiment of the present application.

[0036] Figure 10 It is a schematic block diagram of an electronic device provided by an embodiment of the present application.

[0037] Figure 11 It is a schematic hardware structure diagram of another electronic device provided by an embodiment of the present application. Detailed implementation manners

[0038] Next, the technical solutions in the embodiments of the present application will be clearly described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, rather than all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art belong to the scope of protection of the present application.

[0039] The terms "first", "second", etc. in the present application are used to distinguish similar objects, rather than to describe a specific order or sequence. It should be understood that such terms can be interchanged under appropriate circumstances so that the embodiments of the present application can be implemented in an order other than those illustrated or described herein, and the objects distinguished by "first" and "second" are usually of the same category, and the number of objects is not limited. For example, the first object can be one or multiple. In addition, "or" in the present application means at least one of the connected objects. For example, "A or B" covers three scenarios, namely, Scenario 1: including A and not including B; Scenario 2: including B and not including A; Scenario 3: including both A and B. The character " / " generally indicates an "or" relationship between the associated objects before and after.

[0040] Figure 1 It is a schematic diagram of the codec system 10 provided by an embodiment of the present application. The technical solution of the embodiment of the present application relates to encoding and decoding (CODEC) (including encoding or decoding) video data. Among them, the video data includes raw unencoded video, encoded video, decoded (for example, reconstructed) video, or syntax elements, etc.

[0041] As Figure 1 shown, the codec system 10 includes a source device 100, and the source device 100 provides encoded video data to be decoded and displayed by the destination device 110. Specifically, the source device 100 provides video data to the destination device 110 via the communication medium 120. The source device 100 and the destination device 110 may include any one or more of a desktop computer, a notebook (i.e., laptop) computer, a tablet computer, a set-top box, a mobile phone, a wearable device (such as a smart watch or a wearable camera), a television, a camera, a display device, a vehicle-mounted device, a virtual reality (VR) device, an augmented reality (AR) device, a mixed reality (MR) device, a digital media player, a video game console, a video conferencing device, a video streaming device, a broadcast receiver device, a broadcast transmitter device, a spacecraft, an aircraft, a robot, a satellite, etc.

[0042] In Figure 1 the example of, the source device 100 includes a data source 101, a memory 102, an encoder 200, and an output interface 104. The destination device 110 includes an input interface 111, a decoder 300, a memory 113, and a display device 114. The source device 100 represents an example of a video encoding device, and the destination device 110 represents an example of a video decoding device. In other examples, the source device 100 and the destination device 110 may not include Figure 1 some components in Figure 1 or may also include other components outside

[0043] Although Figure 1 the source device 100 and the destination device 110 are shown as separate devices, in some examples, the two may also be integrated into one device. In such embodiments, the functions corresponding to the source device 100 and the functions corresponding to the destination device 110 may be implemented using the same hardware or software, or using separate hardware or software, or any combination thereof.

[0044] In some examples, the source device 100 and the destination device 110 can perform one-way video transmission or two-way video transmission. If it is two-way video transmission, the source device 100 and the destination device 110 can operate in a substantially symmetric manner, that is, each of the source device 100 and the destination device 110 includes an encoder and a decoder.

[0045] The data source 101 represents the source of video data (i.e., the original, unencoded video data) and provides consecutive pictures containing the video data to the encoder 200, and the encoder 200 encodes the data of the pictures. The data source 101 of the source device 100 can include a video capture device (such as a video camera), a video archive containing previously captured original video, or a video feed interface for receiving video from a video content provider. Alternatively, the data source 101 can generate computer graphics-based data as the source video, or combine live video, archived video, and computer-generated video. In these cases, the encoder 200 encodes the captured, pre-captured, or computer-generated video data. The encoder 200 can rearrange the pictures from the received order (sometimes referred to as the "display order") into the encoding order. The encoder 200 can generate a bitstream including the encoded video data. The source device 100 can then output the encoded video data to the communication medium 120 via the output interface 104 for reception or retrieval by, for example, the input interface 111 of the destination device 110.

[0046] The memory 102 of the source device 100 and the memory 113 of the destination device 110 represent general-purpose memories. In some examples, the memory 102 can store the original video data from the data source 101, and the memory 113 can store the decoded video data from the decoder 300. Additionally or alternatively, the memories 102, 113 can store software instructions that can be executed by, for example, the encoder 200 and the decoder 300. Although the memory 102 and the memory 113 are shown separately from the encoder 200 and the decoder 300 in this example, it should be understood that the encoder 200 and the decoder 300 can also include internal memories for functionally similar or equivalent purposes. If the encoder 200 and the decoder 300 are deployed on the same hardware device, the memory 102 and the memory 113 can be the same memory. Furthermore, the memories 102, 113 can store, for example, the encoded video data output from the encoder 200 and input to the decoder 300. In some examples, portions of the memories 102, 113 can be allocated as one or more video buffers, for example, for storing the original, decoded, or encoded video data.

[0047] In some examples, the source device 100 may output the encoded data from the output interface 104 to the memory 113. Similarly, the destination device 110 may access the encoded data from the memory 113 via the input interface 111. The memory 113 or the memory 102 may include any of a variety of distributed or local access data storage media, such as hard drives, Blu-ray discs, Digital Versatile Discs (DVDs), Compact Disc Read-Only Memories (CD-ROMs), flash memories, volatile or non-volatile memories, or any other suitable digital storage media for storing encoded video data.

[0048] The output interface 104 may include any type of medium or device capable of sending the encoded video data from the source device 100 to the destination device 110. For example, the output interface 104 may include a transmitter or transceiver, such as an antenna, configured to send the encoded video data from the source device 100 directly in real time to the destination device 110. The encoded video data may be modulated according to the communication standards of a wireless communication protocol and sent to the destination device 110.

[0049] The communication medium 120 may include transient media, such as wireless broadcasts or wired network transmissions. For example, the communication medium 120 may include the radio frequency (RF) spectrum or one or more physical transmission lines (e.g., cables). The communication medium 120 may form part of a packet-based network, such as a local area network, a wide area network, or a global network such as the Internet. The communication medium 120 may also take the form of a storage medium (e.g., a non-transitory storage medium), such as a hard drive, a flash drive, a compact disc, a digital video disc, a Blu-ray disc, a volatile or non-volatile memory, or any other suitable digital storage medium for storing encoded video data.

[0050] In some embodiments, the communication medium 120 may include a router, a switch, a base station, or any other device that can be used to facilitate communication from the source device 100 to the destination device 110. For example, a server (not shown) may receive the encoded video from the source device 100 and provide the encoded video data to the destination device 110, e.g., via network transmission to the destination device 110. The server may include, for example, a web server (for a website), a server configured to provide file transfer protocol services such as File Transfer Protocol (FTP) or File Delivery Over Unidirectional Transport (FLUTE) protocol, a content delivery network (CDN) device, a Hypertext Transfer Protocol (HTTP) server, a Multimedia Broadcast Multicast Services (MBMS) or evolved Multimedia Broadcast Multicast Service (eMBMS) server, or a Network-attached storage (NAS) device, etc. The server may implement one or more HTTP streaming protocols such as the MPEG Media Transport (MMT) protocol, the Dynamic Adaptive Streaming over HTTP (DASH) protocol, the HTTP Live Streaming (HLS) protocol, or the Real Time Streaming Protocol (RTSP), etc.

[0051] The destination device 110 may access the encoded video data from the server, e.g., via a wireless channel (e.g., Wi-Fi connection) or a wired connection (e.g., Digital subscriber line (DSL), cable modem, etc.) for accessing the encoded video data stored on the server.

[0052] Output interface 104 and input interface 111 may represent a wireless transmitter / receiver, a modem, a wired networking component (e.g., an Ethernet card), a wireless communication component operating according to the IEEE 802.11 standard or IEEE 802.15 standard (e.g., ZigBeeTM), the Bluetooth standard, etc., or other physical components. In an example where output interface 104 and input interface 111 include wireless components, output interface 104 and input interface 111 may be configured to transmit data, such as encoded video data, according to WIFI, Ethernet, a cellular network (such as 4G, LTE (Long Term Evolution), advanced LTE, 5G, 6G, etc.).

[0053] The technology provided by the embodiments of this application can be applied to support video encoding and decoding in one or more multimedia applications such as the following: video conferencing, over-the-air television broadcasting, cable television transmission, satellite television transmission, Internet streaming video transmission, digital video encoded onto a data storage medium, decoding of digital video stored on a data storage medium, or other applications.

[0054] The input interface 111 of the destination device 110 receives the encoded video bitstream from the communication medium 120. The encoded video bitstream may include syntax elements and encoded data units (e.g., sequences, groups of pictures, pictures, slices, blocks, etc.), where the syntax elements are used to decode the encoded data units to obtain decoded video data. The display device 114 displays the decoded video data to the user. The display device 114 may include a cathode ray tube (CRT), a liquid-crystal display (LCD), a plasma display, an organic light-emitting diode (OLED) display, or other types of display devices.

[0055] The encoder 200 and the decoder 300 may be implemented as one or more of various processing circuits, which may include a microprocessor, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA), discrete logic, hardware, or any combination thereof. When the technology is implemented in whole or in part in software, the device may store instructions for the software in a suitable non-transitory computer-readable storage medium and use one or more processors to execute the instructions in hardware to perform the technology provided by the embodiments of this application.

[0056] The encoder 200 and the decoder 300 can perform processing based on the following video coding and decoding standards: H.263, H.264, H.265 (also known as High Efficiency Video Coding, HEVC), H.266 (also known as Versatile Video Coding, VVC), Moving Picture Experts Group 2 (MPEG-2), MPEG-4, VP8, VP9, Alliance for Open Media Video 1 (AV1), Audio Video Coding Standard 1 (AVS1), AVS2, AVS3, or the next-generation video standard protocol. The embodiments of the present application do not make specific limitations.

[0057] Generally, the encoder 200 and the decoder 300 can perform block-based coding and decoding of pictures. The term "block" generally refers to a structure including data to be processed (e.g., encoded, decoded, or otherwise used during the encoding or decoding process). For example, a block can include a two-dimensional matrix of samples of luminance or chrominance data. For example, the encoder 200 and the decoder 300 can perform coding and decoding on video data represented in the YUV format.

[0058] See Figure 2 , which is a schematic structural diagram of the encoder 200 provided by the embodiments of the present application. The encoder 200 can be the Figure 1 encoder 200 in Figure 2 . In the example of

[0059] The memory 201 can store video data to be encoded. For example, the encoder 200 can receive and store video data from the data source 104 shown in Figure 1 . In some examples, the memory 201 can be on the same chip as other components of the encoder 200 (as shown in Figure 2 ), or can be independent of the chip where these components are located.

[0060] The coding parameter determination unit 210 includes a mode selection unit 211, an inter-frame prediction unit 212, and an intra-frame prediction unit 213. The inter-frame prediction unit 212 is used to obtain a first prediction block of the current block by using the inter-frame prediction mode. The intra-frame prediction unit 213 is used to obtain a second prediction block of the current block by using the intra-frame prediction mode. The mode selection unit 211 is used to obtain a target prediction block based on the first prediction block and the second prediction block, and determine the final prediction mode. In addition, the coding parameter determination unit 210 may further include other functional units, such as a functional unit for determining the partitioning method of a coding unit (CU), a functional unit for determining the transform type of the residual data of the CU, or a functional unit for determining the quantization parameter of the residual data of the CU, etc.

[0061] For the convenience of description and understanding, in the embodiments of the present application, the CU to be processed in the current image is referred to as the current CU, and the image block to be processed in the current CU is referred to as the current block or the image block to be processed. For example, in encoding, it refers to the block currently being encoded; in decoding, it refers to the block currently being decoded.

[0062] The inter-frame prediction unit 212 may include a motion estimation unit and a motion compensation unit. For the inter-frame prediction of the current block, the motion estimation unit may perform a motion search to identify one or more matching reference blocks in one or more reference pictures (for example, one or more previously encoded and decoded pictures stored in the DPB 209).

[0063] The motion estimation unit may form one or more motion vectors (MVs) of the position of the reference block in the reference picture relative to the position of the current block in the current picture. The motion compensation unit may obtain a predicted value with the accuracy indicated by the motion vector through interpolation.

[0064] The coding parameter determination unit 210 may provide the target prediction block to the residual generation unit 202. The residual generation unit 202 receives the original unencoded video data of the current block from the memory 201, and calculates the residual between the current block and the target prediction block to obtain a residual block. In some examples, the function of the residual generation unit 202 may be implemented by using one or more subtractor circuits that perform binary subtraction.

[0065] As an example, the coding parameter determination unit 210 may provide syntax elements representing coding parameters to the entropy coding unit 220 for encoding. The coding parameters include one or more of the partitioning method of the CU, the final prediction mode, the transform type of the residual data of the CU, or the quantization parameter of the residual data of the CU, etc.

[0066] The transform processing unit 203 performs a transform on the residual block output by the residual generation unit 202 to obtain a transform coefficient block. This transform may include a Discrete Cosine Transform (DCT), an integer transform, a directional transform, a Karhunen-Loeve transform, etc. In some examples, the encoder 200 may not include the transform processing unit 203.

[0067] The quantization unit 204 may quantize the transform coefficients in the transform coefficient block according to the quantization parameter (QP) value associated with the current block to generate a quantized transform coefficient block.

[0068] The inverse quantization unit 205 and the inverse transform processing unit 206 may perform inverse quantization and inverse transform on the transform coefficient block respectively to obtain a reconstructed residual block. The reconstruction unit 207 may generate a reconstructed block corresponding to the current block based on the reconstructed residual block and the target prediction block generated by the coding parameter determination unit 210.

[0069] The filter unit 208 may perform one or more filter operations on the reconstructed block. For example, the filter unit 208 may be a deblocking filter (DBF), an adaptive loop filter (ALF), a sample adaptive offset (SAO) filter, etc. In some examples, the encoder 200 may not include the filter unit 208.

[0070] The encoder 200 stores the reconstructed picture obtained from the reconstructed block in the DPB 209. For example, in an example where the operation of the filter unit 208 is not required, the reconstruction unit 207 may store the reconstructed block in the DPB 209. In an example where the operation of the filter unit 208 is required, the filter unit 208 may store the filtered reconstructed block in the DPB 209. The inter prediction unit 212 obtains the reconstructed picture from the DPB 209 to perform inter prediction on the blocks of the subsequent picture to be encoded. In some examples, the DPB 209 may be replaced by other types of memories.

[0071] The entropy coding unit 220 may perform entropy coding on the syntax elements of other components in the encoder 200 and output the encoded video data. For example, the entropy coding unit 220 may perform entropy coding on the quantized transform coefficient block from the quantization unit 204. As another example, the entropy coding unit 220 may perform entropy coding on the syntax elements (such as motion information for inter prediction or intra mode information for intra prediction) from the coding parameter determination unit 210.

[0072] It can be understood that Figure 2 The composition of the encoder 200 described above is only illustrative and does not constitute a limitation on the embodiments of the present application.

[0073] Figure 3 FIG. 6 is a schematic structural diagram of a decoder 300 provided by an embodiment of the present application. The decoder 300 may be Figure 1 the decoder 300 described above. In Figure 3 In the example of FIG. 6, the decoder 300 includes a coded picture buffer (CPB) 301, an entropy decoding unit 302, a prediction processing unit 310, an inverse quantization unit 303, an inverse transform processing unit 304, a reconstruction unit 305, a filter unit 306, and a DPB 307.

[0074] The entropy decoding unit 302 may receive the encoded video data from the CPB 301 and perform entropy decoding on the video data to obtain syntax elements, where the syntax elements indicate coding parameters, and the coding parameters include one or more of the partitioning method of the CU, the final prediction mode, the transform type of the residual data of the CU, or the quantization parameter of the residual data of the CU, etc.

[0075] In the case where the syntax element includes the final prediction mode, the prediction processing unit 310 obtains the final prediction mode. If the final prediction mode is an inter prediction mode, the prediction block of the current CU may be obtained through the inter prediction unit 311 of the prediction processing unit 310; if the final prediction mode is an intra prediction mode, the prediction block of the current CU may be obtained through the intra prediction unit 312 of the prediction processing unit 310. In some examples, the prediction processing unit 310 may further include a unit for performing a prediction function according to other prediction modes.

[0076] The CPB 301 may obtain the encoded video data from the communication medium 120 as shown in Figure 1 FIG. 6 and store it. The DPB 307 is used to store the decoded pictures. Optionally, the CPB 301 and the DPB 307 may also be replaced with other types of memories, and the present application does not make specific limitations. In some examples, the CPB 301 may be on the same chip as other components of the decoder 300 (as shown in the figure), or may be independent of the chip where these components are located.

[0077] The decoder 300 can perform reconstruction operations on each block separately. The entropy decoding unit 302 can perform entropy decoding on the syntax elements of the quantized transform coefficients and the transform information (such as QP or transform mode indication) to obtain the quantized transform coefficients. The inverse quantization unit 303 performs inverse quantization on the quantized transform coefficients to obtain a transform coefficient block including transform coefficients. The inverse transform processing unit 304 performs an inverse transform on the transform coefficient block to generate a residual block corresponding to the current block, and this inverse transform is the reverse operation of the above-mentioned transform.

[0078] The reconstruction unit 305 can reconstruct the current block according to the prediction block and the residual block. For example, the reconstruction unit 305 can add the samples of the residual block to the corresponding samples of the prediction block to reconstruct the current block.

[0079] The filter unit 306 can perform one or more filter operations on the reconstructed block. For example, the type of the filter unit 306 can refer to the type of the filter unit 208, which will not be elaborated here. In some examples, the operation of the filter unit 306 can be skipped.

[0080] The decoder 300 can store the reconstructed picture obtained from the reconstructed block in the DPB 307. For example, in an example where the operation of the filter unit 306 is not performed, the reconstruction unit 305 can store the reconstructed block in the DPB 307. In an example where the operation of the filter unit 306 is performed, the filter unit 306 can store the filtered reconstructed block in the DPB 307. The decoder 300 can output the decoded picture (such as decoded video) from the DPB 307 for subsequent presentation to a display device (such as Figure 1 the display device 114).

[0081] To facilitate a better understanding of the embodiments of the present application, the related technologies of the present application are described.

[0082] In video coding, a frame of image is divided into many macroblocks, and a prediction block is obtained by using intra-frame prediction or inter-frame prediction. The difference between the original block and the prediction block is the residual block, and then the residual block is transformed, quantized, and entropy encoded.

[0083] The image types in video standards usually include I pictures, P pictures, and B pictures. An I picture can be decoded independently without referring to other pictures. A P picture uses multiple pictures (past) that are located before the current picture in the display order as reference pictures. A B picture can use multiple pictures (past) that are located before the current picture in the display order and multiple pictures (future) that are located after the current picture in the display order as reference pictures.

[0084] (1) Intra-frame prediction.

[0085] Intra prediction has many prediction modes to handle various types of textures in an image, including direct current (DC), planar, and some angular prediction modes. Using the reconstructed pixels in the upper row and the left column of the current block as reference pixels, the predicted value of the current predicted block is obtained through the specified prediction mode, achieving the purpose of removing spatial redundancy.

[0086] There are 65 angular prediction modes in VVC.

[0087] Figure 4 It is an example of an intra mode used in VVC provided by an embodiment of the present application.

[0088] Such as Figure 4 shown, the intra modes used in VVC include direct current (DC), planar, and 65 angular modes, that is, a total of 67 prediction modes. Among them, the planar mode is usually used to process some blocks with gradually changing textures, the DC mode is usually used to process some flat areas, and the angular prediction mode is usually used to process blocks with obvious angular textures. Of course, in addition to the above 67 modes, VVC also provides wide-angle modes for some rectangular blocks with a large difference between length and width. For example, Figure 4 the modes indicated by the dotted lines in

[0089] Each angular mode is equivalent to making an offset in the horizontal or vertical direction.

[0090] The corresponding offset values (intraPredAngle) for different angular modes (predModeIntra) can be seen in Table 1.

[0091] Table 1

[0092]

[0093]

[0094] (2) Inter-component model prediction mode.

[0095] To reduce the redundancy between components (i.e., the luminance component and the chrominance component), the inter-component model relationship is established through the luminance reconstructed pixels and the chrominance reconstructed pixels, and the chrominance prediction pixels of the current coding unit are obtained using the encoded luminance reconstructed pixels and the model in the same coding unit.

[0096] Inter-component linear single model prediction mode:

[0097] Based on the encoded luminance reconstructed pixels in the same coding unit, the chrominance prediction pixels are obtained through the following linear model.

[0098] pred C(i,j) = α·rec L ′(t,j) + β;

[0099] where the parameters α and β are obtained by using the minimum linear mean square error for the adjacent luminance reconstructed pixel values and chrominance reconstructed pixel values. The positions of the reconstructed pixel values used to obtain the parameters can be above, to the left, or both above and to the left. For example, it can be Figure 5 as shown above and to the left, and the finally selected position will be written into the bitstream.

[0100] Inter-component linear multi-model prediction mode:

[0101] In the inter-component linear multi-model, the luminance reconstructed pixels encoded in the same coding unit are thresholded into two categories by the average value of the luminance pixels. Each category obtains its respective model parameters by using the minimum linear mean square error.

[0102] (3) Intra template matching prediction (IntraTMP).

[0103] Figure 6 is an example of the principle of the IntraTMP mode provided by this application.

[0104] As Figure 6 shown, IntraTMP searches for the block (matching block) that best matches the template (current template) of the current block within a predetermined range ( Figure 6 R1 to R4 in) in the reconstructed area of the current frame as the predicted value of the current block. The template described here is composed of the reconstructed pixel values of the decoded (or encoded) pixels adjacent to the current block, and can include: one or more rows above, and / or one or more columns to the left.

[0105] (4) Prediction block fusion.

[0106] Multiple prediction modes or multiple reference rows are used for intra-frame prediction to obtain multiple prediction blocks, and the multiple prediction blocks are fused using methods such as weighted average, linear fitting, and non-linear fitting to obtain the final prediction block.

[0107] Fusion mode 1:

[0108] Fuse the intra-frame prediction mode and the inter-component linear model mode to obtain the predicted pixels of the chrominance component.

[0109] pred = (w0 * pred0 + w1 * pred1 + (1 << (shift1))) >> shift;

[0110] Among them, pred0 is calculated using the traditional intra prediction mode, and pred1 is calculated using the inter-component linear model mode. According to the adjacent coding mode information, w0 and w1 can be {1, 3}, {3, 1}, or {2, 2}.

[0111] Fusion mode 2:

[0112] Use the luminance reconstructed pixels, chrominance prediction pixels, and derived model parameters corresponding to the chrominance block to obtain the chrominance prediction pixels of the coding block.

[0113] pred C (i, j) = α0·rec′L(i, j) + α1·pred′ C (i, j) + α2·midValue;

[0114] Where rec′ L (i, j) are the encoded luminance reconstructed pixels in the same coding unit, and pred′ C (i, j) are the chrominance prediction pixels obtained by the chrominance intra prediction mode (non-inter-component model prediction) of the coding unit. midValue is determined by the video bit depth (optional, 10-bit depth, midValue = 512). The model parameters α0, α1, and α2 are obtained by fitting the luminance reconstructed pixels and chrominance prediction pixels of the template.

[0115] IntraTMP prediction fusion mode:

[0116] Select the top N matching blocks with the smallest template cost, and calculate the template cost using the sum of absolute differences (SAD). Use the N matching blocks (refBlock n ) as N prediction values and perform fusion according to the following formula:

[0117]

[0118] midValue is determined by the video bit depth (optional, 10-bit depth, midValue = 512). The fusion weights (w n ) of the prediction values can be calculated based on the template cost or obtained by template fitting.

[0119] It should be noted that in the usual fusion methods, the quantitative relationships between different input items are all one-to-one. For example: in the formula of chrominance fusion mode 2, for the chrominance prediction value of the pixel at the (i, j) position, the downsampled luminance reconstructed pixel rec′ L (i, j) and the chrominance prediction pixel pred′ C(i,j) is obtained by fusion; in the formula for predicting the fusion mode of IntraTMP, the predicted value of the pixel at the (i,j) position uses the corresponding positions of N matching blocks in the refBlock n (i,j) is obtained by fusion.

[0120] It can be seen from the analysis that since the fusion process of pixel values does not consider the influence of the different importance of input items on the final predicted value, the fusion process cannot handle the rich and diverse textures in real video content well. In view of this, the pixel value prediction method proposed in this application can fuse different numbers of pixel values for different input items, which can perform better fusion prediction and improve the prediction accuracy.

[0121] The following introduces the pixel value prediction method provided by the embodiments of this application with reference to the accompanying drawings. The pixel value prediction method provided by the embodiments of this application can be executed by the encoding end, for example Figure 1 or Figure 2 the encoder 200 shown. Alternatively, the pixel value prediction provided by the embodiments of this application can be executed by the decoding end, for example Figure 1 or Figure 3 the decoder 300 described. Among them, the encoding end and the decoding end can be implemented by software, hardware, or a combination thereof. When implemented by hardware, the encoding end can be referred to as an encoding end device or a video encoding device, and the decoding end can be referred to as a decoding end device or a video decoding device.

[0122] Figure 7 is a schematic flowchart of the pixel value prediction method 400 according to the embodiments of this application.

[0123] As Figure 7 shown, the pixel value prediction method 400 may include at least some of the following contents:

[0124] S410, based on the first position where the current processing pixel point is located in the current processing unit, determine a plurality of first pixel values in the first unit corresponding to the current processing unit; the first pixel value is a pixel reconstruction value or a pixel prediction value.

[0125] Exemplarily, the current processing unit may be the current decoding unit or the current encoding unit.

[0126] Exemplarily, the first unit may be a prediction block, a matching block, or a reconstruction block corresponding to the current processing unit. For example, the first unit may be a prediction block obtained by predicting the current processing unit using a first prediction mode. Again, the first unit may be the best matching block that is the most template-matched to the current processing unit obtained by IntraTMP. Again, the first unit may be a reconstruction block of the luminance component of the current processing unit.

[0127] S420. Based on the first position, in at least one second unit corresponding to the current processing unit, determine at least one second pixel value; the second pixel value is a pixel reconstruction value or a pixel prediction value, and the number of the plurality of first pixel values is greater than the number of second pixel values determined in any one of the at least one second unit.

[0128] Exemplarily, the second unit may be a prediction block, a matching block, or a reconstruction block corresponding to the current processing unit. For example, the at least one second unit may be at least one prediction block obtained by predicting the current processing unit using at least one second prediction mode other than the first prediction mode. For another example, the at least one second unit may be at least one matching block obtained by using IntraTMP other than the best matching block that is most matched with the template of the current processing unit. For another example, the second unit may be a prediction block of the chrominance component of the current processing unit.

[0129] S430. Fuse the plurality of first pixel values and the at least one second pixel value to obtain a predicted pixel value of the current processed pixel point.

[0130] Exemplarily, the predicted pixel value may be a pixel value obtained by performing a fusion calculation on the plurality of first pixel values and the at least one second pixel value.

[0131] In this embodiment, when fusing the plurality of first pixel values and the at least one second pixel value, since the number of the plurality of first pixel values is greater than the number of second pixel values determined in any one of the second units, therefore, the first pixel values in the first unit and the second pixel values in the second unit can be utilized more flexibly to obtain the predicted pixel value of the current processed pixel point, which can improve the prediction accuracy of the pixel value, and further improve the prediction effect and the codec performance.

[0132] In some embodiments, S410 includes:

[0133] Determine the plurality of first pixel values from the pixel value at the first position in the first unit and the pixel values at adjacent positions to the first position in the first unit.

[0134] Exemplarily, the pixel value at the first position in the first unit may also be referred to as the pixel value at the position of the first position in the first unit.

[0135] Exemplarily, the pixel values at adjacent positions to the first position in the first unit may also be referred to as the pixel values at positions adjacent to the first position in the first unit.

[0136] Exemplarily, when determining the plurality of first pixel values among the pixel value at the first position in the first unit and the pixel values at adjacent positions to the first position in the first unit, some or all of the pixel values among the pixel value at the first position in the first unit and the pixel values at adjacent positions to the first position in the first unit may be determined as the plurality of first pixel values. The number of the plurality of first pixel values may be a value predefined by the protocol.

[0137] In this embodiment, since both the pixel value at the first position in the first unit and the pixel values at adjacent positions to the first position in the first unit have a relatively high correlation with the pixel value at the first position in the current processing unit, therefore, among the pixel value at the first position in the first unit and the pixel values at adjacent positions to the first position in the first unit, the plurality of first pixel values are determined, and then the plurality of first pixel values and the at least one second pixel value are fused to obtain the predicted pixel value of the current processing pixel point, which can improve the prediction accuracy of the pixel value, and further improve the prediction effect and the coding and decoding performance.

[0138] In some embodiments, the adjacent positions include positions adjacent to the first position in at least one direction.

[0139] Exemplarily, the at least one direction includes at least one of the following: up, down, left, right, upper right, lower right, upper left, lower left.

[0140] In some embodiments, S410 includes:

[0141] Among the pixel values in the first region where the first position is located in the first unit, the plurality of first pixel values are determined.

[0142] Exemplarily, the length or width of the first region is a value predefined by the protocol. For example, the number of pixel values of the length or width of the first region is the same as the value predefined by the protocol.

[0143] Exemplarily, when determining the plurality of first pixel values among the pixel values in the first region where the first position is located in the first unit, some or all of the pixel values in the first region where the first position is located in the first unit may be determined as the plurality of first pixel values. The number of the plurality of first pixel values may be a value predefined by the protocol.

[0144] In this embodiment, when fusing the multiple first pixel values and the at least one second pixel value, the number of the multiple first pixel values is greater than the number of the second pixel values determined in any one of the second units. Equivalently, the predicted pixel value of the current processed pixel point can be obtained by increasing the number of the first pixel values in the first unit, that is, the prediction accuracy of the pixel value can be improved, and further the prediction effect and the coding and decoding performance can be improved.

[0145] In some embodiments, S420 includes:

[0146] Determine the pixel value at the first position in any one of the second units as the second pixel value among the at least one second pixel value.

[0147] Exemplarily, the pixel value at the first position in any one of the second units can be understood as the pixel value at the same position as the first position in any one of the units.

[0148] Of course, in other alternative embodiments, one or more second pixel values can also be determined in any one of the second units, as long as the number of the multiple first pixel values is greater than the number of the second pixel values in any one of the second units. Even more, as long as the number of the multiple first pixel values is greater than the number of the second pixel values in any one of the second units, the number of the second pixel values determined in different units in the at least one second unit can be the same or different, and the present application does not make specific limitations thereto.

[0149] In some embodiments, S430 includes:

[0150] Obtain the model parameters of the first model; the first model is used to fuse the multiple first pixel values and the at least one second pixel value;

[0151] Based on the model parameters, fuse the multiple first pixel values and the at least one second pixel value to obtain the predicted pixel value.

[0152] Exemplarily, the first model can be any machine learning model or AI model.

[0153] Exemplarily, the first model can be a model that performs fusion using fusion methods such as weighted average, linear fitting, and non - linear fitting.

[0154] In some embodiments, the current processing unit includes a chrominance component unit, the first unit includes a luminance component reconstruction unit corresponding to the chrominance component unit, the luminance component reconstruction unit is a reconstructed unit with or without downsampling, and the at least one second unit includes a chrominance component prediction unit obtained by predicting the chrominance component unit using the prediction mode of the chrominance component;

[0155] Among them, obtaining the model parameters of the first model includes:

[0156] Fitting the reconstructed pixel values of the template of the luminance component unit corresponding to the chrominance component unit and the predicted pixel values of the template of the chrominance component unit to obtain the model parameters.

[0157] Exemplarily, the template of the luminance component unit includes the reconstructed pixel values of the adjacent decoded (or encoded) luminance component units, for example, it may include: the upper row / rows, and / or the left column / columns. Similarly, the template of the chrominance component unit includes the reconstructed pixel values of the adjacent decoded (or encoded) chrominance component units, for example, it may include: the upper row / rows, and / or the left column / columns.

[0158] Exemplarily, when fitting the reconstructed pixel values of the template of the luminance component unit and the predicted pixel values of the template of the chrominance component unit to obtain the model parameters, multiple reconstructed pixel values (the number is equal to the number of the multiple first pixel values) in the template of the luminance component unit, and one or more predicted pixel values (the number is equal to the number of the second pixel values determined in any one of the second units) in the template of the chrominance component unit and with a number less than the number of the reconstructed pixel values in the template of the used luminance component unit can be fitted to obtain the model parameters.

[0159] In this embodiment, since the luminance component reconstruction unit usually contains richer texture information than the chrominance component prediction unit, therefore, the first unit includes the luminance component reconstruction unit, and the at least one second unit includes the chrominance component prediction unit; equivalently, when fusing the reconstructed pixel values in the luminance component reconstruction unit and the predicted pixel values in the chrominance component prediction unit, the number of the reconstructed pixel values used in the luminance component reconstruction unit is increased, and thus the chrominance prediction pixel values can be predicted more accurately using richer luminance texture information, which can improve the coding performance of the coded chrominance component.

[0160] In some embodiments, when the size of the chrominance component unit is the same as the size of the luminance component unit, the luminance component reconstruction unit is a reconstruction unit without downsampling; or, when the size of the chrominance component unit is different from the size of the luminance component unit, the luminance component reconstruction unit is a reconstruction unit with or without downsampling.

[0161] Exemplarily, for videos and images in the 4:4:4 format, the luminance component unit and the chrominance component unit have the same size. In this case, the luminance component reconstruction unit is a reconstruction unit without downsampling.

[0162] For videos and images in other formats, such as 4:2:0, the chrominance component unit is smaller in size than the luminance component unit. In this case, the luminance component reconstruction unit is a reconstruction unit with or without downsampling. At this time, if the luminance reconstruction value without downsampling is used, the multiple first pixel values may include pixel values determined from any at least 1 luminance reconstruction pixel value (L0 to L5) adjacent to the first position (C) as shown in Figure 8 the luminance reconstruction pixel values shown.

[0163] In some embodiments, the first unit includes the first matching block with the smallest template error among the multiple matching blocks corresponding to the current processing unit, and the at least one second unit includes the second matching blocks other than the first matching block among the multiple matching blocks;

[0164] Among them, obtaining the model parameters of the first model includes:

[0165] Fitting the reconstructed pixel values of the templates of the multiple matching blocks to obtain the model parameters.

[0166] Exemplarily, the template of any one of the multiple matching blocks includes the adjacent decoded (or encoded) reconstructed pixel values of the any one matching block. For example, it may include: the upper one / several rows, and / or the left one / several columns.

[0167] Exemplarily, when fitting the reconstructed pixel values of the templates of the multiple matching blocks to obtain the model parameters, multiple reconstructed pixel values (the number is equal to the number of the multiple first pixel values) of the template in the first matching block, and one or more reconstructed pixel values (the number is equal to the number of the second pixel values determined in any one of the second units) in the template of the second matching block and with a number less than the number of the reconstructed pixel values in the template of the used first matching block can be fitted to obtain the model parameters.

[0168] In this embodiment, since the matching block with the smallest template error (i.e., the first matching block) has the highest similarity to the current coding block, the first unit includes the first matching block, and the at least one second unit includes the second matching blocks other than the first matching block among the multiple matching blocks. Equivalently, when fusing the reconstructed pixel values in the multiple matching blocks, by increasing the number of reconstructed pixel values used in the first matching block, a more accurate prediction value can be obtained by utilizing the high correlation between the first matching block and the current processing unit, and the coding performance of coding the current processing unit can be improved.

[0169] In some embodiments, the model parameters include the weight of each first pixel value among the multiple first pixel values and the weight of each second pixel value among the at least one second pixel value.

[0170] Of course, in other alternative embodiments, the model parameters may further include a bias parameter and the weight of the bias parameter. For example, the bias parameter may be the midValue mentioned above, that is, the bias parameter may be determined by the video bit depth. For example, for a 10-bit depth, midValue = 512. Even, the model parameters may also be non-linear parameters, which are not specifically limited in this application.

[0171] The solution of this application will be described below with reference to specific embodiments.

[0172] Embodiment 1:

[0173] Step 101: Obtain the template of the chrominance block and the corresponding template of the luminance block of the current coding unit, and obtain the predicted pixel values of the template of the chrominance block and the reconstructed pixel values of the template of the luminance block, and downsample the reconstructed pixel values of the template of the luminance block to a downsampled template with the same size as the template of the chrominance block.

[0174] Step 102:

[0175] Fit the model parameters with the predicted pixel values of the template of the chrominance block and the reconstructed pixel values of the downsampled template.

[0176] Step 103:

[0177] Obtain the intra prediction mode (non-component-inter model prediction) predicted pixel values of the chrominance block of the current coding unit.

[0178] Step 104:

[0179] Use the reconstructed pixel values of the downsampled luminance reconstruction block corresponding to the current coding unit, which have been downsampled to the same size as the chrominance block, the predicted pixel values of the chrominance block, and the model parameters to obtain the chrominance predicted pixel values of the current coding unit according to the following formula.

[0180] pred C (ij) = α0·rec′ L (i,j) + α1·rec′ L (i - 1,j) + α2·rec′ L (i + 1,j) + α3·rec′(i, j - 1 + α4·recL′ij + 1 + α5·predC′i,j + α6·midValue;

[0181] where rec′ L is the downsampled luminance reconstruction block after downsampling, and rec′ L (i,j) is the reconstructed luminance reconstruction pixel value corresponding to the current chrominance position (i,j), and rec′ L (i - 1, j) is the luminance reconstruction pixel value to the left of the current chrominance position (i,j), and rec′ L (i + 1, j) is the luminance reconstruction pixel value to the right of the current chrominance position (i,j), and rec′ L (u, j - 1) is the luminance reconstruction pixel value above the current chrominance position (i,j), and rec′ L (i, j + 1) is the luminance reconstruction pixel value below the current chrominance position (i,j), and pred′ C (i,j) is the chrominance prediction pixel value at the current chrominance position (i,j) in the chrominance prediction block obtained by the current coding unit using the intra - chrominance prediction mode (non - inter - component model prediction). midValue is determined by the video bit depth (optional, for 10 - bit depth, midValue = 512). The model parameters α0, α1, α2, α3, α4, α5, and α6 are obtained by fitting with the luminance reconstruction pixels and chrominance prediction pixels of the template.

[0182] It should be noted that in this embodiment, the luminance reconstruction pixel values at the current chrominance position (i,j) and its four upper, lower, left, and right positions are used. It can also be the luminance reconstruction pixel values at other adjacent (or non - adjacent) positions, which are not limited here.

[0183] In this embodiment, since the luminance block usually contains richer texture information than the chrominance block, using the current chrominance position and the luminance pixel reconstruction values in its neighborhood, which are more in number than the chrominance prediction pixel values, as the input for fusion can use the richer luminance texture information to more accurately predict the chrominance prediction pixel value at the current chrominance position (i,j), and can improve the coding performance of the coded chrominance component.

[0184] Embodiment 2:

[0185] Step 201:

[0186] Obtain a candidate list, where the candidate list contains N matching blocks arranged in ascending order of template error.

[0187] Step 202:

[0188] Obtain the templates of the current coding unit and the templates of N matching blocks, and fit the model parameters with the reconstructed pixel values of the templates.

[0189] Step 203:

[0190] Obtain the reconstructed pixel values of N matching blocks, and obtain the predicted pixel values of the current coding unit according to the following formula with the model parameters.

[0191] pred(i, j) = w0 * refBlock0(i, j) + w1 * refBlock0(i - 1, j) + w2 * refBlock0(i + 1, j) + w3 * refBlock0(i, j - 1) + w4 * refBlock0(i, j + 1) + ∑n = 1N - 1wn + 4 * refBlockn(i, j) + wN + 4 * midValue;

[0192] Where, refBlock(i, j) is the reconstructed pixel value of the matching block in the candidate list at the current position (i, j), refBlock0 is the matching block Block0 with the smallest template error, refBlock0(i, j) is the reconstructed pixel value at the current position (i, j) in Block0, refBlock0(i - 1, j) is the reconstructed pixel value to the left of the current position (i, j), refBlock0(i + 1, j) is the reconstructed pixel value to the right of the current position (i, j), refBlock0(i, j - 1) is the reconstructed pixel value above the current position (i, j), refBlock0(i, j + 1) is the reconstructed pixel value below the current position (i, j), midValue is determined by the video bit depth (optional, 10-bit depth, midValue = 512), and the model parameters w0 to w N+4 Are obtained by fitting with the reconstructed pixel values of N templates.

[0193] It should be noted that in this embodiment, the reconstructed pixel values of the current position (i, j) and its 4 positions above, below, left, and right can also be the reconstructed pixel values of other adjacent (or non-adjacent) positions, which are not limited here.

[0194] In this embodiment, since the matching block with the smallest template error (the best matching block) has the highest similarity with the current coding block, the reconstructed pixel values of the current position and its neighborhood in the best matching block, which are more in number than those of other matching blocks, are used as the input for fusion. More accurate prediction values can be obtained by using its high correlation, improving the coding performance of the IntraTMP mode.

[0195] The pixel value prediction method provided by the embodiments of this application may be executed by a pixel value prediction device. In the embodiments of this application, taking the pixel value prediction device executing the pixel value prediction method as an example, the pixel value prediction device provided by the embodiments of this application is described.

[0196] Figure 9 FIG. 4 shows a schematic block diagram of a pixel value prediction device 500 according to an embodiment of the present application.

[0197] As Figure 9 shown, the pixel value prediction device 500 includes:

[0198] A determination unit 510, configured to:

[0199] Based on the first position where the current processing pixel point is located in the current processing unit, determine a plurality of first pixel values in the first unit corresponding to the current processing unit; the first pixel value is a pixel reconstruction value or a pixel prediction value;

[0200] Based on the first position, determine at least one second pixel value in at least one second unit corresponding to the current processing unit; the second pixel value is a pixel reconstruction value or a pixel prediction value, and the number of the plurality of first pixel values is greater than the number of second pixel values determined in any one of the at least one second unit;

[0201] A fusion unit 520, configured to fuse the plurality of first pixel values and the at least one second pixel value to obtain a predicted pixel value of the current processing pixel point.

[0202] In some embodiments, the determination unit 510 is specifically configured to:

[0203] Determine the plurality of first pixel values from the pixel value at the first position in the first unit and the pixel values at adjacent positions to the first position in the first unit.

[0204] In some embodiments, the adjacent positions include positions adjacent to the first position in at least one direction.

[0205] In some embodiments, the determination unit 510 is specifically configured to:

[0206] Determine the plurality of first pixel values from the pixel values in the first region where the first position is located in the first unit.

[0207] In some embodiments, the determination unit 510 is specifically configured to:

[0208] Determine the pixel value at the first position in any one of the second units as the second pixel value among the at least one second pixel value.

[0209] In some embodiments, the fusion unit 520 is specifically configured to:

[0210] Obtain the model parameters of the first model; the first model is used to fuse the multiple first pixel values and the at least one second pixel value;

[0211] Based on the model parameters, fuse the multiple first pixel values and the at least one second pixel value to obtain the predicted pixel value.

[0212] In some embodiments, the current processing unit includes a chrominance component unit, the first unit includes a luminance component reconstruction unit corresponding to the chrominance component unit, the luminance component reconstruction unit is a reconstruction unit with or without downsampling, and the at least one second unit includes a chrominance component prediction unit obtained by predicting the chrominance component unit using the prediction mode of the chrominance component;

[0213] Wherein, the fusion unit 520 is specifically configured to:

[0214] Fit the reconstructed pixel values of the template of the luminance component unit corresponding to the chrominance component unit and the predicted pixel values of the template of the chrominance component unit to obtain the model parameters.

[0215] In some embodiments, when the size of the chrominance component unit is the same as the size of the luminance component unit, the luminance component reconstruction unit is a reconstruction unit without downsampling; or, when the size of the chrominance component unit is different from the size of the luminance component unit, the luminance component reconstruction unit is a reconstruction unit with or without downsampling.

[0216] In some embodiments, the first unit includes the first matching block with the smallest template error among the multiple matching blocks corresponding to the current processing unit, and the at least one second unit includes the second matching blocks other than the first matching block among the multiple matching blocks;

[0217] Wherein, the fusion unit 520 is specifically configured to:

[0218] Fit the reconstructed pixel values of the templates of the multiple matching blocks to obtain the model parameters.

[0219] In some embodiments, the model parameters include the weight of each of the multiple first pixel values and the weight of each of the at least one second pixel value.

[0220] It should be understood that the pixel value prediction device 500 provided in the embodiments of the present application may correspond to the execution subject in the method embodiments of the present application, and the above (or other) operations or functions of each unit in the pixel value prediction device 500 respectively implement Figure 7 the corresponding processes of the method embodiments, and for the sake of brevity, they will not be elaborated here.

[0221] In the embodiments of the present application, when fusing the multiple first pixel values and the at least one second pixel value, the number of the multiple first pixel values is greater than the number of the second pixel values determined in any one of the second units. Equivalently, by increasing the number of the first pixel values in the first unit, the predicted pixel value of the currently processed pixel point can be obtained, that is, the prediction accuracy of the pixel value can be improved, and further the prediction effect and the coding and decoding performance can be improved.

[0222] The pixel value prediction device provided in the embodiments of the present application can implement Figure 7 the respective processes implemented by the method embodiments, and achieve the same technical effects. For the sake of avoiding repetition, they will not be elaborated here.

[0223] The embodiments of the present application further provide an electronic device 600, as Figure 10 shown, including a processor 601 and a memory 602. A program or instruction that can run on the processor 601 is stored on the memory 602. For example, when the electronic device 600 is an encoding end device, when the program or instruction is executed by the processor 601, it implements the respective steps of the above pixel value prediction method embodiments, and can achieve the same technical effects. When the electronic device 600 is a decoding end device, when the program or instruction is executed by the processor 601, it implements the respective steps of the above pixel value prediction method embodiments, and can achieve the same technical effects. For the sake of avoiding repetition, they will not be elaborated here. Optionally, the memory 602 may be Figure 1 the memory 102 or the memory 113 in the embodiments shown, and the processor 601 may implement Figures 1-3 the functions of the encoder 200 or the decoder 300 in the embodiments shown.

[0224] The embodiments of the present application further provide an electronic device, including: a memory configured to store video data; and a processing circuit configured to implement the respective steps of the above pixel value prediction method embodiments. Optionally, the memory may be Figure 1 the memory 102 or the memory 113 in the embodiments shown, and the processing circuit may implement Figures 1-3 the functions of the encoder 200 or the decoder 300 in the embodiments shown.

[0225] An embodiment of the present application further provides an electronic device, including a processor and a communication interface. The communication interface is coupled to the processor, and the processor is configured to run programs or instructions to implement the steps in the above-mentioned embodiment of the pixel value prediction method. This device embodiment corresponds to the above-mentioned method embodiment, and all implementation processes and implementation manners of the above-mentioned method embodiment can be applied to this terminal embodiment and can achieve the same technical effects.

[0226] The above-mentioned electronic device can be a terminal or other devices other than terminals, such as a server, a Network Attached Storage (NAS), etc.

[0227] Among them, the terminal can be a mobile phone, a tablet personal computer, a laptop computer, a notebook computer, a personal digital assistant (PDA), a handheld computer, a netbook, an ultra-mobile personal computer (UMPC), a mobile Internet device (MID), an augmented reality (AR), a virtual reality (VR) device, a mixed reality (MR) device, a robot, a wearable device, a flight vehicle, a vehicle user equipment (VUE), a shipborne device, a pedestrian user equipment (PUE), a smart home (home appliances with wireless communication functions, such as refrigerators, TVs, washing machines, or furniture, etc.), a game console, a personal computer (PC), a teller machine, or a self-service machine, etc. Wearable devices include: smart watches, smart bracelets, smart earphones, smart glasses, smart jewelry (smart bracelets, smart bracelets, smart rings, smart necklaces, smart anklets, smart ankle chains, etc.), smart wristbands, smart clothing, etc. Among them, the vehicle user equipment can also be referred to as a vehicle terminal, a vehicle controller, a vehicle module, a vehicle component, a vehicle chip, or a vehicle unit, etc. It should be noted that the specific type of the terminal is not limited in the embodiment of the present application.

[0228] The server can be an independent physical server, a server cluster or a distributed system composed of multiple physical servers, or a cloud server. The cloud server can provide cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, Content Delivery Networks (CDNs), or cloud computing services based on big data and artificial intelligence platforms, etc.

[0229] Exemplarily, the above-mentioned electronic device can include but is not limited to Figure 1 the types of the source device 100 or the destination device 110 shown.

[0230] Taking the electronic device as a terminal as an example, Figure 11 It is a schematic diagram of the hardware structure of a terminal for implementing an embodiment of the present application.

[0231] The terminal 700 includes but is not limited to at least some components such as a radio frequency unit 701, a network module 702, an audio output unit 703, an input unit 704, a sensor 705, a display unit 706, a user input unit 707, an interface unit 708, a memory 709, and a processor 710.

[0232] Those skilled in the art can understand that the terminal 700 may further include a power source (such as a battery) for supplying power to each component. The power source can be logically connected to the processor 710 through a power management system, so as to realize functions such as management of charging, discharging, and power consumption management through the power management system. Figure 11 The terminal structure shown in does not constitute a limitation on the terminal. The terminal may include more or fewer components than shown in the figure, or combine certain components, or have different component arrangements, which will not be elaborated here.

[0233] It should be understood that in the embodiments of the present application, the input unit 704 may include a Graphics Processing Unit (GPU) 7041 and a microphone 7042. The graphics processor 7041 processes the image data of static pictures or videos obtained by an image acquisition device (such as a camera) in the video acquisition mode or the image acquisition mode, or may process the obtained point cloud data. The display unit 706 may include a display panel 7061, and the display panel 7061 may be configured in the form of, for example, a liquid crystal display, an organic light emitting diode, etc. The user input unit 707 includes at least one of a touch panel 7071 and other input devices 7072. The touch panel 7071 is also referred to as a touch screen. The touch panel 7071 may include two parts: a touch detection device and a touch controller. The other input devices 7072 may include, but are not limited to, a physical keyboard, function keys (such as volume control keys, power on / off keys, etc.), a trackball, a mouse, and a joystick, which will not be elaborated here.

[0234] In the embodiments of the present application, after receiving downlink data from a network-side device, the radio frequency unit 701 may transmit it to the processor 710 for processing; in addition, the radio frequency unit 701 may send uplink data to the network-side device. Generally, the radio frequency unit 701 includes, but is not limited to, an antenna, an amplifier, a transceiver, a coupler, a low noise amplifier, a duplexer, etc.

[0235] The memory 709 can be used to store software programs or instructions as well as various data. The memory 709 may mainly include a first storage area for storing programs or instructions and a second storage area for storing data. Among them, the first storage area may store an operating system, application programs or instructions required for at least one function (such as a sound playback function, an image playback function, etc.). In addition, the memory 709 may include volatile memory or non-volatile memory. Among them, the non-volatile memory may be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory may be a random access memory (RAM), a static random access memory (SRAM), a dynamic random access memory (DRAM), a synchronous dynamic random access memory (SDRAM), a double data rate synchronous dynamic random access memory (DDR SDRAM), an enhanced synchronous dynamic random access memory (ESDRAM), a synchlink dynamic random access memory (SLDRAM), and a direct rambus random access memory (DRRAM). The memory 709 in the embodiments of the present application includes but is not limited to these and any other suitable types of memory.

[0236] The processor 710 may include one or more processing units; optionally, the processor 710 integrates an application processor and a modem processor. Among them, the application processor mainly processes operations related to the operating system, user interface, and application programs, etc., and the modem processor mainly processes wireless communication signals, such as a baseband processor. It can be understood that the above modem processor may not be integrated into the processor 710 either.

[0237] Among them, the processor 710 is used for:

[0238] Based on the first position where the current processing pixel point is located in the current processing unit, determine a plurality of first pixel values in the first unit corresponding to the current processing unit; the first pixel value is a pixel reconstruction value or a pixel prediction value;

[0239] Based on the first position, determine at least one second pixel value in at least one second unit corresponding to the current processing unit; the second pixel value is a pixel reconstruction value or a pixel prediction value, and the number of the multiple first pixel values is greater than the number of second pixel values determined in any one of the at least one second unit;

[0240] Fuse the multiple first pixel values and the at least one second pixel value to obtain a predicted pixel value of the current processed pixel point

[0241] In the embodiments of the present application, when fusing the multiple first pixel values and the at least one second pixel value, since the number of the multiple first pixel values is greater than the number of second pixel values determined in any one of the at least one second unit, equivalently, the predicted pixel value of the current processed pixel point can be obtained by increasing the number of first pixel values in the first unit, that is, the prediction accuracy of the pixel value can be improved, and further the prediction effect and the coding and decoding performance can be improved.

[0242] It can be understood that the implementation processes of the various implementation manners mentioned in this embodiment can refer to the relevant descriptions of the method embodiments and achieve the same or corresponding technical effects. To avoid repetition, they will not be elaborated here.

[0243] The embodiments of the present application further provide a readable storage medium, on which a program or instruction is stored. When the program or instruction is executed by a processor, it implements each process of the above-mentioned pixel value prediction method embodiment and can achieve the same technical effect. To avoid repetition, it will not be elaborated here.

[0244] Wherein, the processor is the processor in the terminal described in the above embodiment. The readable storage medium includes computer-readable storage media, such as ROM, RAM, magnetic disks or optical discs, etc. In some examples, the readable storage medium may be a non-transitory readable storage medium.

[0245] The embodiments of the present application further provide a chip, which includes a processor and a communication interface. The communication interface is coupled to the processor. The processor is used to run a program or instruction to implement each process of the above-mentioned pixel value prediction method embodiment and can achieve the same technical effect. To avoid repetition, it will not be elaborated here.

[0246] It should be understood that the chip mentioned in the embodiments of the present application may include a system-on-chip (also referred to as a system chip, a chip system or a system-on-chip), or may include an independent display chip, etc.

[0247] Another embodiment of the present application further provides a computer program / program product. The computer program / program product is stored in a storage medium and is executed by at least one processor to implement each process of the above-mentioned pixel value prediction method embodiment, and can achieve the same technical effect. To avoid repetition, it will not be elaborated here.

[0248] An embodiment of the present application further provides an encoding and decoding system, including: an encoding end device and a decoding end device. The encoding end device can be used to execute the steps of the pixel value prediction method as described above, and the decoding end device can be used to execute the steps of the pixel value prediction method as described above.

[0249] It should be noted that in this article, the term "including", "comprising" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not explicitly listed, or further includes elements inherent to such process, method, article or device. Without more limitations, an element defined by the statement "including one..." does not exclude the existence of additional identical elements in the process, method, article or device including that element. In addition, it should be pointed out that the scope of the methods and devices in the embodiments of the present application is not limited to performing functions in the order shown or discussed, and may also include performing functions in a substantially simultaneous manner or in a reverse order according to the functions involved. For example, the described methods may be performed in an order different from that described, and various steps may be added, omitted or combined. In addition, the features described with reference to certain examples may be combined in other examples.

[0250] Through the description of the above embodiments, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be implemented by means of a computer software product plus a necessary general hardware platform, and of course, can also be implemented by hardware. The computer software product is stored in a storage medium (such as ROM, RAM, magnetic disk, optical disk, etc.) and includes several instructions for enabling a terminal or a network-side device to execute the methods described in various embodiments of the present application.

[0251] The embodiments of the present application have been described above in conjunction with the accompanying drawings. However, the present application is not limited to the above specific embodiments. The above specific embodiments are merely illustrative and not restrictive. Under the inspiration of the present application, those of ordinary skill in the art can also make many forms of embodiments without departing from the purpose of the present application and the scope protected by the claims. These embodiments are all within the protection scope of the present application.

Claims

1. A pixel value prediction method, characterized in that Performed by a decoding end or an encoding end, the method includes: Based on a first position where a currently processed pixel point is located in a currently processed unit, determining a plurality of first pixel values in a first unit corresponding to the currently processed unit; the first pixel values being pixel reconstruction values or pixel prediction values; Based on the first position, determining at least one second pixel value in at least one second unit corresponding to the currently processed unit; the second pixel values being pixel reconstruction values or pixel prediction values, and the number of the plurality of first pixel values being greater than the number of second pixel values determined in any one of the at least one second unit; Fusing the plurality of first pixel values and the at least one second pixel value to obtain a predicted pixel value of the currently processed pixel point.

2. The method according to claim 1, wherein The determining a plurality of first pixel values in a first unit corresponding to the currently processed unit based on a first position where a currently processed pixel point is located in the currently processed unit includes: Determining the plurality of first pixel values from among pixel values at the first position in the first unit and pixel values at adjacent positions to the first position in the first unit.

3. The method according to claim 2, characterized in that, The adjacent positions include positions adjacent to the first position in at least one direction.

4. The method according to claim 1, wherein The determining a plurality of first pixel values in a first unit corresponding to the currently processed unit based on a first position where a currently processed pixel point is located in the currently processed unit includes: Determining the plurality of first pixel values from among pixel values in a first region where the first position is located in the first unit.

5. The method according to any one of claims 1 to 4, characterized in that, The determining at least one second pixel value in at least one second unit corresponding to the currently processed unit based on the first position includes: Determining the pixel value at the first position in any one of the second units as the second pixel value among the at least one second pixel value.

6. The method according to any one of claims 1 to 5, characterized in that, The fusing the plurality of first pixel values and the at least one second pixel value to obtain a predicted pixel value of the currently processed pixel point includes: Obtaining model parameters of a first model; the first model being used to fuse the plurality of first pixel values and the at least one second pixel value; Based on the model parameters, fusing the plurality of first pixel values and the at least one second pixel value to obtain the predicted pixel value.

7. The method according to claim 6, wherein The currently processed unit includes a chrominance component unit, the first unit includes a luminance component reconstruction unit corresponding to the chrominance component unit, the luminance component reconstruction unit being a reconstructed unit with or without downsampling, and the at least one second unit includes a chrominance component prediction unit obtained by predicting the chrominance component unit using a prediction mode of the chrominance component; Wherein, the obtaining model parameters of the first model includes: Fitting a reconstructed pixel value of a template of a luminance component unit corresponding to the chrominance component unit and a predicted pixel value of a template of the chrominance component unit to obtain the model parameters.

8. The method according to claim 7, characterized in that, When the size of the chrominance component unit is the same as the size of the luminance component unit, the luminance component reconstruction unit is a reconstruction unit without downsampling; or, when the size of the chrominance component unit is different from the size of the luminance component unit, the luminance component reconstruction unit is a reconstruction unit with or without downsampling.

9. The method according to claim 6, characterized in that, The first unit includes a first matching block with the smallest template error among a plurality of matching blocks corresponding to the current processing unit, and the at least one second unit includes second matching blocks among the plurality of matching blocks other than the first matching block; Wherein, obtaining the model parameters of the first model includes: Fitting the reconstructed pixel values of the templates of the plurality of matching blocks to obtain the model parameters.

10. The method according to any one of claims 7 to 9, characterized in that The model parameters include the weights of each of the plurality of first pixel values and the weights of each of the at least one second pixel value.

11. A pixel value prediction device, characterized in that, Includes: A determination unit for: Based on a first position where a current processing pixel point is located in a current processing unit, determining a plurality of first pixel values in a first unit corresponding to the current processing unit; the first pixel values are pixel reconstruction values or pixel prediction values; Based on the first position, determining at least one second pixel value in at least one second unit corresponding to the current processing unit; the second pixel values are pixel reconstruction values or pixel prediction values, and the number of the plurality of first pixel values is greater than the number of second pixel values determined in any one of the at least one second unit; A fusion unit for fusing the plurality of first pixel values and the at least one second pixel value to obtain a predicted pixel value of the current processing pixel point.

12. The device according to claim 11, wherein, The determination unit is specifically configured to: Determine the plurality of first pixel values among the pixel value at the first position in the first unit and the pixel values at adjacent positions to the first position in the first unit.

13. The device according to claim 12, wherein The adjacent positions include positions adjacent to the first position in at least one direction.

14. The device according to claim 11, wherein The determination unit is specifically configured to: Determine the plurality of first pixel values among the pixel values in a first region where the first position is located in the first unit.

15. The device according to any one of claims 11 to 14, characterized in that, The determination unit is specifically configured to: Determine the pixel value at the first position in any one of the second units as the second pixel value among the at least one second pixel value.

16. The device according to any one of claims 11 to 15, characterized in that The fusion unit is specifically configured to: Obtain the model parameters of a first model; the first model is used to fuse the plurality of first pixel values and the at least one second pixel value; Based on the model parameters, fuse the plurality of first pixel values and the at least one second pixel value to obtain the predicted pixel value.

17. The device according to claim 16, characterized in that, The current processing unit includes a chrominance component unit, the first unit includes a luminance component reconstruction unit corresponding to the chrominance component unit, the luminance component reconstruction unit is a reconstruction unit with or without downsampling, and the at least one second unit includes a chrominance component prediction unit obtained by predicting the chrominance component unit using a prediction mode of the chrominance component; Wherein, the fusion unit is specifically configured to: Fitting the reconstructed pixel values of the template of the luminance component unit corresponding to the chrominance component unit and the predicted pixel values of the template of the chrominance component unit to obtain the model parameters.

18. The device according to claim 17, characterized in that, When the size of the chrominance component unit is the same as the size of the luminance component unit, the luminance component reconstruction unit is a reconstruction unit without downsampling; or, when the size of the chrominance component unit is different from the size of the luminance component unit, the luminance component reconstruction unit is a reconstruction unit with or without downsampling.

19. The device according to claim 16, characterized in that, The first unit includes the first matching block with the smallest template error among multiple matching blocks corresponding to the current processing unit, and the at least one second unit includes second matching blocks among the multiple matching blocks other than the first matching block; Wherein, the fusion unit is specifically configured to: Fit the reconstructed pixel values of the templates of the multiple matching blocks to obtain the model parameters.

20. The device according to any one of claims 17 to 19, characterized in that, The model parameters include the weights of each of the multiple first pixel values and the weights of each of the at least one second pixel value.

21. An electronic device, characterized in that, It includes a processor and a memory, and the memory stores programs or instructions that can run on the processor. When the programs or instructions are executed by the processor, the steps of the pixel value prediction method according to any one of claims 1 to 10 are implemented.

22. A readable storage medium, characterized in that, Programs or instructions are stored on the readable storage medium. When the programs or instructions are executed by the processor, the steps of the pixel value prediction method according to any one of claims 1 to 10 are implemented.

23. A chip, characterized in that, The chip includes a processor and a communication interface. The communication interface is coupled to the processor, and the processor is used to run programs or instructions to implement the steps of the pixel value prediction method according to any one of claims 1 to 10.