Video encoding method and apparatus, video decoding method and apparatus, and medium and electronic device
By fitting the mapping relationship between color components in video encoding and decoding for prediction, the problems of low efficiency and poor accuracy of traditional encoding and decoding are solved, and more efficient and accurate video encoding is achieved.
Patent Information
- Application Number
- PCT/CN2024/131391
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-09
- Filing Date
- 2024-11-11
- Publication Date
- 2025-06-12
AI Technical Summary
Traditional audio and video codec solutions have problems such as low encoding and codec efficiency and poor accuracy.
By obtaining the reconstruction value of the first color component of the current block, and fitting the mapping relationship between the first color component and the second color component according to at least one image area in the current block and the reference area, the mapping process is performed to obtain the predicted value of the second color component.
The encoding efficiency and accuracy of videos are improved, and the problem of independent prediction of different color components in traditional encoding schemes is overcome.
Smart Images

Figure CN2024131391_12062025_PF_FP_ABST
Abstract
Description
Video encoding and decoding method, device, medium and electronic equipment
[0001] This application claims priority to the Chinese patent application filed with the China Patent Office on December 9, 2023, with application number 202311692581.5 and invention name “Video Coding and Decoding Method, Device, Medium and Electronic Device”, the entire contents of which are incorporated by reference into this application. Technical Field
[0002] The present application belongs to the field of audio and video technology, and specifically relates to a video encoding method, a video decoding method, a video encoding device, a video decoding device, a computer-readable medium, an electronic device, and a computer program product. Background Art
[0003] To accommodate large-scale data transmission of audio and video, the original audio and video data must typically be encoded at the data transmitter to form a compressed data stream. After transmitting the data stream to the data receiver, the stream is decoded to restore the predicted and reconstructed audio and video data. Traditional audio and video codec solutions suffer from low coding and decoding efficiency and poor accuracy.
[0004] Summary of the Invention
[0005] The present application provides a video encoding method, a video decoding method, a video encoding device, a video decoding device, a computer-readable medium, an electronic device, and a computer program product, the purpose of which is to improve encoding efficiency and encoding accuracy.
[0006] According to one aspect of an embodiment of the present application, a video decoding method is provided, the method comprising:
[0007] Obtaining a reconstructed value of a first color component of a current block, where the current block is an image block to be decoded in a current video frame;
[0008] fitting a mapping relationship between a first color component of the current block and a second color component of the current block according to the current block and at least one image region in a reference region, wherein the reference region includes one or more decoded image regions that are the nearest neighbor or the next nearest neighbor to the current block;
[0009] Mapping processing is performed on the reconstructed value of the first color component of the current block according to the mapping relationship to obtain a predicted value of the second color component of the current block.
[0010] According to one aspect of an embodiment of the present application, a video encoding method is provided, the method comprising:
[0011] Obtaining a reconstructed value of a first color component of a current block, where the current block is an image block to be encoded in a current video frame;
[0012] Fitting a mapping relationship between a first color component of the current block and a second color component of the current block according to the current block and at least one image region in a reference region, wherein the reference region includes one or more coded and reconstructed image regions that are the nearest or next-nearest neighbors of the current block;
[0013] Mapping processing is performed on the reconstructed value of the first color component of the current block according to the mapping relationship to obtain a predicted value of the second color component of the current block.
[0014] According to one aspect of an embodiment of the present application, a video decoding apparatus is provided, the apparatus comprising:
[0015] A first acquisition module is configured to acquire a reconstructed value of a first color component of a current block, where the current block is an image block to be decoded in a current video frame;
[0016] a first fitting module configured to fit a mapping relationship between a first color component of the current block and a second color component of the current block according to the current block and at least one image region in a reference region, wherein the reference region includes one or more decoded image regions that are the nearest neighbor or the next nearest neighbor to the current block;
[0017] The first mapping module is configured to perform mapping processing on the reconstructed value of the first color component of the current block according to the mapping relationship to obtain a predicted value of the second color component of the current block.
[0018] According to one aspect of an embodiment of the present application, a video encoding apparatus is provided, the apparatus comprising:
[0019] a second acquisition module configured to acquire a reconstructed value of a first color component of a current block, where the current block is an image block to be encoded in a current video frame;
[0020] a second fitting module configured to fit a mapping relationship between a first color component of the current block and a second color component of the current block according to the current block and at least one image region in a reference region, wherein the reference region includes one or more encoded and reconstructed image regions that are the nearest or next nearest neighbor to the current block;
[0021] The second mapping module is configured to perform mapping processing on the reconstructed value of the first color component of the current block according to the mapping relationship to obtain a predicted value of the second color component of the current block.
[0022] According to one aspect of an embodiment of the present application, a computer-readable medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the video encoding and decoding method in the above technical solution is implemented.
[0023] According to one aspect of an embodiment of the present application, an electronic device is provided, comprising: a processor; and a memory for storing executable instructions of the processor; wherein the processor is configured to execute the executable instructions to implement the video encoding and decoding method in the above technical solution.
[0024] According to one aspect of an embodiment of the present application, a computer program product is provided, including a computer program, which implements the video encoding and decoding method in the above technical solution when executed by a processor.
[0025] In the technical solution provided in the embodiments of the present application, a reconstructed value of the first color component of the current block is obtained, and then a mapping relationship between the first color component and the second color component is fitted based on the current block and at least one image region in the reference region. The reconstructed value of the first color component of the current block is mapped according to the mapping relationship to obtain a predicted value of the second color component of the current block. The embodiments of the present application use color component mapping prediction to predict the second color component, overcoming the problem of poor prediction accuracy caused by independent prediction of different color components in traditional encoding schemes, thereby improving video encoding efficiency and accuracy.
[0026] It should be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the present application. BRIEF DESCRIPTION OF THE DRAWINGS
[0027] FIG1 schematically shows a diagram of an exemplary system architecture to which the technical solutions of the embodiments of the present application can be applied.
[0028] FIG2 schematically shows how a video encoding device and a video decoding device are placed in a streaming environment.
[0029] FIG3 schematically shows a basic flow chart of a video encoder, in which intra-frame prediction is taken as an example for explanation.
[0030] FIG4 shows a flowchart of the steps of a video decoding method in one embodiment of the present application.
[0031] FIG5 is a schematic diagram showing a process of sampling color components of a current block in one embodiment of the present application.
[0032] FIG6 shows a flowchart of fitting the mapping relationship between color components in one embodiment of the present application.
[0033] FIG7 is a schematic diagram showing the relationship between a designated pixel point and neighboring pixel points in one embodiment of the present application.
[0034] FIG8 shows a schematic diagram of the distribution of reference areas corresponding to the current block in one embodiment of the present application.
[0035] FIG9 shows a schematic diagram of a region template in which some sub-regions are combined into a reference region in one embodiment of the present application.
[0036] FIG10 is a schematic diagram showing a region template of a reference region selected for a current block in one embodiment of the present application.
[0037] FIG. 11 is a schematic diagram showing a method of selecting a sampling position based on the position coordinates of a pixel point in one embodiment of the present application.
[0038] FIG12 is a schematic diagram showing a method of selecting sampling positions based on a round-trip scanning method in one embodiment of the present application.
[0039] FIG13 shows a schematic diagram of selecting sampling positions based on a ZigZag scanning method in one embodiment of the present application.
[0040] FIG14 shows a schematic diagram of a sampling window for downsampling the first color component in one embodiment of the present application.
[0041] FIG15 shows a flowchart of the steps of a video encoding method in one embodiment of the present application.
[0042] FIG16 schematically shows a structural block diagram of a video decoding device provided in an embodiment of the present application.
[0043] FIG17 schematically shows a structural block diagram of a video encoding device provided in an embodiment of the present application.
[0044] FIG18 schematically shows a block diagram of a computer system structure of an electronic device suitable for implementing an embodiment of the present application. DETAILED DESCRIPTION
[0045] Example embodiments will now be described more fully with reference to the accompanying drawings. However, example embodiments can be implemented in many forms and should not be construed as limited to the examples set forth herein; rather, these embodiments are provided so that this application will be thorough and complete and will fully convey the concepts of the example embodiments to those skilled in the art.
[0046] The relevant technical terms involved in the embodiments of this application are explained as follows.
[0047] Video coding generally refers to the processing of a sequence of pictures to form a video or video sequence. In the field of video coding, the terms "picture", "frame" or "image" can be used as synonyms. The video coding used in the embodiments of the present application represents video encoding or video decoding. Video encoding is performed on the source side and generally includes processing (for example, by compression) the original video picture to reduce the amount of data required to represent the video picture, thereby more efficiently storing and / or transmitting. Video decoding is performed on the destination side and generally includes inverse processing relative to the encoder to reconstruct the video picture. The "encoding" of the video frames involved in the embodiments should be understood as "encoding" or "decoding" involving a sequence of video images. The combination of the encoding part and the decoding part is also called codec (encoding and decoding).
[0048] Each picture in a video image sequence is typically divided into a set of non-overlapping blocks, which are typically encoded at the block level. In other words, the encoder typically processes, or encodes, the video at the block level (also called an image block or video block). For example, a prediction block is generated through spatial (intra-picture) prediction and temporal (inter-picture) prediction, the prediction block is subtracted from the current block (the block currently being processed or to be processed) to obtain a residual block, the residual block is transformed in the transform domain, and the residual block is quantized to reduce the amount of data to be transmitted (compressed). The decoder applies the inverse of the encoder's processing to the coded or compressed block to reconstruct the current block for representation. Furthermore, the encoder replicates the decoder processing loop, so that the encoder and decoder generate the same prediction (e.g., intra-frame prediction and inter-frame prediction) and / or reconstruction for processing, or encoding, subsequent blocks.
[0049] The term "block" refers to a portion of a picture or frame. In the embodiments of the present application, the current block refers to the block currently being processed. For example, in encoding, it refers to the block currently being encoded; in decoding, it refers to the block currently being decoded.
[0050] FIG1 schematically shows a diagram of an exemplary system architecture to which the technical solutions of the embodiments of the present application can be applied.
[0051] 1 , a system architecture 100 includes a plurality of terminal devices that can communicate with each other via, for example, a network 150. For example, the system architecture 100 may include a first terminal device 110 and a second terminal device 120 interconnected via the network 150. In the embodiment of FIG1 , the first terminal device 110 and the second terminal device 120 perform unidirectional data transmission.
[0052] For example, the first terminal device 110 can encode video data (such as a video picture stream captured by the terminal device 110) for transmission to the second terminal device 120 via the network 150. The encoded video data is transmitted in the form of one or more encoded video streams. The second terminal device 120 can receive the encoded video data from the network 150, decode the encoded video data to restore the video data, and display the video picture based on the restored video data.
[0053] In one embodiment of the present application, the system architecture 100 may include a third terminal device 130 and a fourth terminal device 140 for performing bidirectional transmission of encoded video data, such as during a video conference. For bidirectional data transmission, each of the third terminal device 130 and the fourth terminal device 140 may encode video data (e.g., a video image stream captured by the terminal device) for transmission to the other of the third terminal device 130 and the fourth terminal device 140 via a network 150. Each of the third terminal device 130 and the fourth terminal device 140 may also receive the encoded video data transmitted by the other of the third terminal device 130 and the fourth terminal device 140, decode the encoded video data to recover the video data, and display the video image on an accessible display device based on the recovered video data.
[0054] In the embodiment of FIG1 , the first terminal device 110, the second terminal device 120, the third terminal device 130, and the fourth terminal device 140 may be servers, personal computers, and smartphones, but the principles disclosed herein are not limited thereto. The embodiments disclosed herein are applicable to laptop computers, tablet computers, media players, and / or dedicated video conferencing equipment. Network 150 represents any number of networks, including, for example, wired and / or wireless communication networks, for transmitting encoded video data between the first terminal device 110, the second terminal device 120, the third terminal device 130, and the fourth terminal device 140. The communication network 150 may exchange data using circuit-switched and / or packet-switched channels. The network may include a telecommunications network, a local area network, a wide area network, and / or the Internet. For the purposes of this application, unless otherwise explained below, the architecture and topology of network 150 may be immaterial to the operations disclosed herein.
[0055] In one embodiment of the present application, FIG2 schematically illustrates the placement of a video encoding device and a video decoding device in a streaming environment. The subject matter disclosed herein is equally applicable to other video-enabled applications, including, for example, video conferencing, digital television, and storing compressed video on digital media such as CDs, DVDs, and memory sticks.
[0056] The streaming system may include an acquisition subsystem 213, which may include a video source 201, such as a digital camera, that creates an uncompressed video picture stream 202. In one embodiment, the video picture stream 202 includes samples captured by the digital camera. The video picture stream 202 is depicted as a thicker line to emphasize the higher data volume of the video picture stream compared to the encoded video data 204 (or the encoded video stream 204). The video picture stream 202 may be processed by an electronic device 220, which includes a video encoding device 203 coupled to the video source 201. The video encoding device 203 may include hardware, software, or a combination of hardware and software to implement or embody various aspects of the disclosed subject matter, as described in greater detail below. The encoded video data 204 (or the encoded video stream 204) is depicted as a thinner line to emphasize the lower data volume of the encoded video data 204 (or the encoded video stream 204), which may be stored on a streaming server 205 for future use. One or more streaming client subsystems, such as client subsystem 206 and client subsystem 208 in FIG2 , can access streaming server 205 to retrieve copies 207 and 209 of encoded video data 204. Client subsystem 206 can include, for example, a video decoding device 210 in electronic device 230. Video decoding device 210 decodes the incoming copy 207 of the encoded video data and generates an output video picture stream 211 that can be presented on a display 212 (e.g., a screen) or another presentation device. In some streaming systems, the encoded video data 204, video data 207, and video data 209 (e.g., video code streams) can be encoded according to certain video encoding / compression standards.
[0057] It should be noted that the electronic device 220 and the electronic device 230 may include other components not shown in the figure. For example, the electronic device 220 may include a video decoding device, and the electronic device 230 may also include a video encoding device.
[0058] In one embodiment of the present application, taking the international video coding standards HEVC (High Efficiency Video Coding, H.265), VVC (Versatile Video Coding, H.266), and China's national video coding standard AVS (Audio Video Coding Standard) as examples, after a video frame image is input, the video frame image will be divided into several non-overlapping processing units according to a block size, and each processing unit will perform similar compression operations. This processing unit is called CTU (Coding Tree Unit), or LCU (Largest Coding Unit). The CTU can be further divided into more refined parts to obtain one or more basic coding units CU. CU is the most basic element in a coding link.
[0059] FIG3 schematically shows a basic flow chart of a video encoder, in which intra-frame prediction is taken as an example for explanation.
[0060] Among them, the original image signal s k [x,y] and predicted image signal Perform difference operation to obtain the residual signal u k [x,y], residual signal u k [x,y] is transformed and quantized to obtain the quantized coefficients. The quantized coefficients are entropy coded to obtain the encoded bit stream, and the reconstructed residual signal u' is obtained by inverse quantization and inverse transformation. k [x,y], predicted image signal and the reconstructed residual signal u' k [x,y] superposition generates image signal Image signal On the one hand, it is input to the intra-frame mode decision module and the intra-frame prediction module for intra-frame prediction processing, and on the other hand, the reconstructed image signal s' is output through loop filtering. k [x,y], reconstructed image signal s' k [x,y] can be used as the reference image for the next frame for motion estimation and motion compensation prediction. Then based on the result of motion compensation prediction s' r [x+m x ,y+m y ] and intra prediction results Get the predicted image signal of the next frame And continue to repeat the above process until the encoding is completed.
[0061] The encoding operations for each CU involved in the above video encoding process are described in detail as follows.
[0062] Predictive Coding: Predictive coding includes methods such as intra-frame prediction and inter-frame prediction. The original video signal is predicted by a selected reconstructed video signal to obtain a residual video signal. The encoder needs to decide which predictive coding mode to use for the current CU and inform the decoder. Intra-frame prediction refers to the predicted signal coming from an already coded and reconstructed area within the same image; inter-frame prediction refers to the predicted signal coming from a previously coded image different from the current image (called a reference image).
[0063] Transform & Quantization: After the residual video signal undergoes transformations such as the Discrete Fourier Transform (DFT) and Discrete Cosine Transform (DCT), the signal is converted to the transform domain, where the coefficients are known as transform coefficients. The transform coefficients are then subjected to a lossy quantization operation, which loses some information, making the quantized signal more suitable for compression. In some video coding standards, more than one transform scheme may be available, so the encoder must select one for the current CU and inform the decoder. The level of quantization is typically determined by the quantization parameter (QP). A larger QP value means that coefficients with a larger value range will be quantized to the same output, which generally results in greater distortion and a lower bitrate. Conversely, a smaller QP value means that coefficients with a smaller value range will be quantized to the same output, which generally results in less distortion and a higher bitrate.
[0064] Entropy Coding or Statistical Coding: The quantized transform domain signal is statistically compressed and encoded based on the frequency of occurrence of each value, and finally a binary (0 or 1) compressed code stream is output. At the same time, the encoding generates other information, such as the selected coding mode, motion vector data, etc., which also need to be entropy coded to reduce the bit rate. Statistical coding is a lossless coding method that can effectively reduce the bit rate required to express the same signal. Common statistical coding methods include variable length coding (VLC) or context-based binary arithmetic coding (CABAC).
[0065] The context-based adaptive binary arithmetic coding (CABAC) process consists of three main steps: binarization, context modeling, and binary arithmetic coding. After binarization, the input syntax elements can be encoded using both the normal coding mode and the bypass coding mode. In the bypass coding mode, instead of assigning a specific probability model to each binary bit, the input binary bit values are directly encoded using a simple bypass encoder, speeding up both encoding and decoding. Generally, different syntax elements are not completely independent, and even the same syntax elements have some memory. Therefore, based on conditional entropy theory, conditional coding using other coded syntax elements can further improve coding performance compared to independent or memoryless coding. This coded symbol information used as a condition is called context. In the normal coding mode, the binary bits of the syntax elements are sequentially fed into the context modeler. The encoder assigns an appropriate probability model to each input binary bit based on the values of previously coded syntax elements or binary bits. This process is known as context modeling. The context model corresponding to the syntax element can be located using ctxIdxInc (context index increment) and ctxIdxStart (context index Start). After the bin value and the assigned probability model are fed into the binary arithmetic encoder for encoding, the context model needs to be updated based on the bin value, which is the adaptive process in encoding.
[0066] Loop Filtering: The changed and quantized signal will be reconstructed through inverse quantization, inverse transformation and prediction compensation operations to obtain a reconstructed image. Compared with the original image, due to the influence of quantization, some information of the reconstructed image is different from the original image, that is, the reconstructed image will produce distortion. Therefore, the reconstructed image can be filtered, such as deblocking filter (DB), SAO (Sample Adaptive Offset) or ALF (Adaptive Loop Filter) and other filters, which can effectively reduce the degree of distortion caused by quantization. Since these filtered reconstructed images will be used as a reference for subsequent encoded images to predict future image signals, the above filtering operation is also called loop filtering, that is, filtering operation within the encoding loop.
[0067] Based on the above encoding process, at the decoding end, after obtaining the compressed code stream (i.e., bitstream), entropy decoding is performed on each CU to obtain various mode information and quantization coefficients. The quantization coefficients are then dequantized and inversely transformed to obtain a residual signal. Furthermore, based on the known coding mode information, the prediction signal corresponding to the CU can be obtained. The residual signal is then added to the prediction signal to obtain a reconstructed signal. The reconstructed signal then undergoes loop filtering and other operations to produce the final output signal.
[0068] The following describes in detail the technical solutions provided by the present application, including the video encoding method, video decoding method, video encoding device, video decoding device, computer-readable medium, electronic device, and computer program product, in combination with specific implementation methods.
[0069] Figure 4 shows a flowchart of the steps of a video decoding method in an embodiment of the present application. The video decoding method can be executed by a terminal device or a server that receives the encoded data. The embodiment of the present application uses the video decoding method executed by a terminal device as an example to illustrate. The terminal device can be, for example, the video decoding device 210 shown in Figure 2.
[0070] As shown in FIG4 , the video decoding method in the embodiment of the present application includes the following steps S410 to S430 .
[0071] S410: Obtain a reconstructed value of a first color component of a current block, where the current block is an image block to be decoded in a current video frame.
[0072] S420: Fitting a mapping relationship between the first color component and the second color component according to the current block and at least one image region in a reference region, where the reference region includes one or more decoded image regions that are nearest or next nearest neighbors to the current block.
[0073] S430: Mapping the reconstructed value of the first color component of the current block according to the mapping relationship to obtain a predicted value of the second color component of the current block.
[0074] In the video decoding method provided in an embodiment of the present application, the reconstructed value of the first color component of the current block is obtained, and then the mapping relationship between the first color component and the second color component is fitted according to the current block and at least one image area in the reference area, so that the reconstructed value of the first color component of the current block is mapped according to the mapping relationship to obtain the predicted value of the second color component of the current block.
[0075] The predicted pixels of the inter-frame prediction or intra-frame block copy mode are obtained by searching in the reconstructed reference frame or the reconstructed area of the current frame or searching in its sub-pixel interpolated image. For the chrominance component, it is generally assumed that it has a similar motion vector or block vector as the luminance component, so the chrominance motion vector or block vector is simply derived based on the luminance motion vector or block vector. Since this matching search process not only needs to take into account the matching degree of the image, but also the encoding cost of the motion vector or block vector, there is still room for improvement in the prediction accuracy. The embodiment of the present application predicts the second color component by using color component mapping prediction, which overcomes the problem of poor prediction accuracy caused by independent prediction of different color components in traditional encoding schemes, thereby improving the encoding efficiency and accuracy of the video.
[0076] It should be understood that the fitting referred to in this application refers to establishing a mathematical model or function by analyzing the color components of the current block and at least one image region in the reference region. This model can describe the mapping relationship between the first color component of the current block and the second color component of the current block. Thus, the reconstructed value of the first color component of the current block can be mapped using this mapping relationship to achieve prediction of the second color component of the current block.
[0077] The following is a detailed description of each method step in the embodiment of the present application in conjunction with specific implementation methods.
[0078] In step S410 , a reconstructed value of a first color component of a current block is obtained, where the current block is an image block to be decoded in a current video frame.
[0079] Segmenting the current video frame may yield multiple image regions, each of which, or a combination of multiple image regions, may be considered an image block. The current block is an image block to be decoded in the current video frame. The first color component may be any one of a luminance component Y, a chrominance component U, and a chrominance component V.
[0080] In one embodiment of the present application, color component sampling of the current block in the color space may form a luminance block and a chrominance block.
[0081] Color video images contain not only a luma component (Y) but also chroma components (U, V). Such images are also called YUV images. When encoding YUV images, in addition to encoding the luma component, the chroma components also need to be encoded. Because the human eye is more sensitive to brightness and less so to color, to save storage space and improve coding efficiency, the luma component is sampled at full resolution, while the chroma components do not need to be sampled at full resolution. Depending on the sampling method for the luma and chroma components in color video, video sequences typically have various formats: 4:4:4 YUV images, 4:2:2 YUV images, 4:2:0 YUV images, and so on.
[0082] Figure 5 shows a schematic diagram of the process of sampling the color components of the current block in one embodiment of the present application. As shown in Figure 5, the current block 501 is an image block of size 16×16. When sampling the color components of the current block 501, different resolution formats can be selected for sampling.
[0083] The 4:4:4 format indicates that the chroma components are not downsampled. The 4:4:4 format is the format with the highest chroma component resolution. When sampling in the 4:4:4 format, one Y component corresponds to a set of UV components. The data in four adjacent pixels contains 4 Y, 4 U, and 4 V components.
[0084] The 4:2:2 format means that the chroma components are horizontally downsampled 2:1 relative to the luma components, with no vertical downsampling. For every two U or V samples, each row contains four Y samples. When sampling in the 4:2:2 format, every two Y components share a set of UV components, and the data in four adjacent pixels contains 4 Y, 2 U, and 2 V components.
[0085] The 4:2:0 format means that the chroma components are horizontally downsampled 2:1 relative to the luma components, and vertically downsampled 2:1. The 4:2:0 format is the format with the lowest resolution for the chroma components and is also the most common format. In the 4:2:0 format, the chroma samples are only half of the luma samples in each row (i.e. horizontally), and only half of the luma samples in each column (i.e. vertically). When sampling in the 4:2:0 format, the U component and the Y component appear alternately. For example, if the three YUV components appear in a ratio of 4:2:0 in the first row, then in the second row they appear in a ratio of 4:0:2; there are 4 Y, 1 U, and 1 V in 4 adjacent pixels.
[0086] In the case of a video image in 4:2:0 format, if the luminance component of an image block is a 2M×2N image block, then the chrominance component of the image block is an M×N image block. For example, if the resolution of the image block is 720*480, then the resolution of the luminance component of the image block is 720*480, and the resolution of the chrominance component of the image block is 360*240.
[0087] In the embodiment of the present application, the 4:2:0 format is taken as an example. After sampling the current coding unit 501, a luminance block 502 and a chrominance block 503 can be obtained. The luminance block 502 is a 16×16 image block, and the corresponding chrominance block 503 is an 8×8 image block.
[0088] The prediction coding mode of the current block can include inter-frame prediction, intra-frame prediction, or other prediction methods. In video coding, the main redundant information is temporal redundancy, followed by spatial redundancy. Video coding eliminates temporal redundancy through inter-frame prediction and spatial redundancy through intra-frame prediction.
[0089] In one embodiment of the present application, obtaining the reconstructed value of the first color component of the current block may further include: obtaining a prediction coding mode of the current block; when the prediction coding mode is inter-frame prediction, performing inter-frame prediction based on a first reference block corresponding to the current block to obtain a reconstructed value of the first color component of the current block, where the first reference block is a decoded image block in a reference video frame corresponding to the current video frame; when the prediction coding mode is intra-frame block copying, performing intra-frame block copying prediction based on a second reference block corresponding to the current block to obtain a reconstructed value of the first color component of the current block, where the second reference block is a decoded image block in the current video frame.
[0090] Inter-frame prediction is a method for predicting and encoding the current frame by leveraging the correlation between previous and next frames in a video sequence. In a video sequence, adjacent frames typically have a high degree of pixel similarity, so information from the previous or next frame can be used to predict the pixel values of the current frame.
[0091] Inter-frame prediction can be achieved through processes such as motion estimation, motion compensation, and residual coding. Motion estimation is to find the best motion vector to describe the motion relationship between the two frames by comparing the difference between the current frame and the previous or next frame. Common motion estimation algorithms include full search algorithm, block matching algorithm, etc. Motion compensation is to use the found motion vector to correct the previous or next frame to achieve pixel prediction of the current frame. Motion compensation can be achieved by applying the motion vector to the reference frame for pixel replication or interpolation. Residual coding is to obtain residual data by encoding the difference between the pixel values of the current frame and the pixel values of the predicted frame. The residual data represents the pixel information in the current frame that cannot be obtained through motion prediction and requires additional encoding and transmission.
[0092] Intra Block Copy (IBC) can be considered a special inter-frame prediction mode. Its implementation principles are nearly identical to those used for motion compensation in inter-frame prediction models. The difference is that, whereas inter-frame prediction selects reference blocks for motion compensation from a reference frame different from the current one, IBC selects reference blocks within the current frame. In IBC, the motion vector represents the relative displacement from the current block's position to the reference block's position within the current frame.
[0093] In step S420, a mapping relationship between the first color component and the second color component is fitted according to the current block and at least one image region in a reference region, where the reference region includes one or more decoded image regions that are the nearest or next nearest neighbor to the current block.
[0094] The second color component is another color component different from the first color component in the color space. For example, when the first color component is the brightness component Y, the second color component may be the chrominance component U or the chrominance component V.
[0095] Taking the YUV color space as an example, in the embodiment of the present application, for the current block with the YUV image format, multiple mapping relationships as shown below can be fitted between multiple color components.
[0096] (1) Predict the chrominance component U based on the luminance component Y, and obtain the mapping relationship U=f(Y).
[0097] (2) Predict the chrominance component V based on the luminance component Y, and obtain the mapping relationship V=f(Y).
[0098] (3) Predict the chrominance component V based on the chrominance component U, and obtain the mapping relationship V = f(U)
[0099] (4) The chrominance component V is predicted based on the luminance component Y and the chrominance component U, and a mapping relationship V = f(Y, U) is obtained.
[0100] (5) The chrominance component U is predicted based on the luminance component Y and the chrominance component V, and a mapping relationship U = f(Y, V) is obtained.
[0101] FIG6 shows a flowchart for fitting a mapping relationship between color components in one embodiment of the present application. As shown in FIG6 , based on the above embodiment, fitting a mapping relationship between a first color component and a second color component based on the current block and at least one image region in the reference region may further include the following steps S610 to S630.
[0102] S610: Obtain an initial prediction model corresponding to the current block, where the input item of the initial prediction model includes the first color component of the specified pixel point, the output item of the initial prediction model is the second color component of the specified pixel point, and the initial prediction model includes at least one of multiple candidate models.
[0103] In one embodiment of the present application, the input item of the initial prediction model also includes the first color component of one or more neighborhood pixels, where the neighborhood pixels are the nearest or next nearest neighbors of the designated pixel.
[0104] FIG7 is a schematic diagram showing the relationship between a designated pixel point and neighboring pixel points in one embodiment of the present application.
[0105] As shown in FIG7 , C represents a designated pixel point, and the designated pixel points in the first color component a and the second color component b may be pixel sample points at associated positions or the same position.
[0106] The neighborhood pixels of the designated pixel point C may include multiple nearest neighboring pixel points, such as the multiple pixel points N, S, W, and E located above, below, to the left, and to the right of the designated pixel point C as shown in FIG7 .
[0107] The neighborhood pixels of the designated pixel point C may include multiple next-nearest neighboring pixels, such as the multiple pixels NW, NE, SW and SE located at the upper left, upper right, lower left and lower right positions of the designated pixel point C as shown in FIG7 .
[0108] Each such connection between the first color component a and the second color component b is called a cross-component matching pair, generating a mapping equation. Multiple such mapping equations can be used to solve the model's weighted parameters. The input location of the first color component a in the figure is only an example; a wider range of pixel sample points can also be selected.
[0109] In one embodiment of the present application, the initial prediction model includes one or more combination items with independent weighting parameters, and the combination item takes the first color components of at least two pixels as input items.
[0110] The initial prediction model in the embodiment of the present application can be composed of at least one monomial, wherein each monomial can have an independent weighting parameter. When a monomial has at least two first color components of pixels as input items, the monomial is called a combination item.
[0111] In one embodiment of the present application, the initial prediction model may include at least two combinations of different orders, where the order is the highest power of the input items in the combination. The power operation can introduce nonlinear factors to improve the model's fit to the color component mapping relationship.
[0112] In one embodiment of the present application, the initial prediction model can be obtained by performing weighted operations on one or more of the following monomials: x, mx±ny, xy, x k , (mx±ny) k , (mx±ny)(pz±qf), (mx±ny)x, B.
[0113] Where m, n, p, and q represent weighting coefficients for fixed weighting of the input items; x, y, z, and f represent the first color components of a specified pixel or neighboring pixels; B represents a constant bias term; and k represents the order of the power operation on the input item, where k is an integer greater than 1.
[0114] In one embodiment of the present application, the initial prediction model may include one or more of the following multiple candidate models.
[0115] (1)C b =p0C+p1N+p2S+p3W+p4E+p5C 2 +p6B.
[0116] (2)C b =p0C+p1N+p2S+p3W+p4E+p5B.
[0117] (3)C b =p0C+p1B.
[0118] (12)C b =p0C+p1S+p2W+p3E+p4SW+p5SE+p6C 2 +p7B.
[0119] S620: Selecting training sample pairs in at least one image area of the current block and the reference area, the training sample pairs including one or more of the following sample pairs: a predicted value of the first color component and a predicted value of the second color component of a sample pixel point located in the current block, and a reconstructed value of the first color component and a reconstructed value of the second color component of a sample pixel point located in the reference area.
[0120] In one embodiment of the present application, a mapping relationship between the first color component and the second color component can be fitted based on the predicted values of the color components of the current block itself. On this basis, the training sample pairs used for fitting the mapping relationship can include the predicted values of the first color component and the predicted values of the second color component of sample pixels located in the current block.
[0121] In one embodiment of the present application, the predicted value of the first color component and the predicted value of the second color component of the sample pixel point located in the current block are obtained according to the following method: obtaining the prediction coding mode of the current block; when the prediction coding mode is inter-frame prediction, performing inter-frame prediction based on a first reference block corresponding to the current block to obtain the predicted value of the first color component and the predicted value of the second color component of the sample pixel point in the current block, where the first reference block is a decoded image block in a reference video frame corresponding to the current video frame; when the prediction coding mode is intra-frame block copying, performing intra-frame block copying prediction based on a second reference block corresponding to the current block to obtain the predicted value of the first color component and the predicted value of the second color component of the sample pixel point in the current block, where the second reference block is a decoded image block in the current video frame.
[0122] In one embodiment of the present application, a mapping relationship between the first color component and the second color component can be fitted based on the reconstructed values of the color components of the reference area. On this basis, the training sample pairs used for fitting the mapping relationship can include the reconstructed values of the first color component and the reconstructed values of the second color component of the sample pixels located in the reference area.
[0123] In one embodiment of the present application, a mapping relationship between the first color component and the second color component can be fitted based on both the predicted values of the color components of the current block and the reconstructed values of the color components of the reference area. On this basis, the training sample pairs used for fitting the mapping relationship can include the predicted values of the first color component and the predicted values of the second color component of sample pixels located in the current block, and the reconstructed values of the first color component and the reconstructed values of the second color component of sample pixels located in the reference area.
[0124] In one embodiment of the present application, the reference area corresponding to the current block is composed of one or more nearest neighbor areas or second-nearest neighbor areas, the nearest neighbor area includes an image area with a specified image size located above or to the left of the current block, and the second-nearest neighbor area includes an image area with a specified image size located above, below, or above the current block.
[0125] FIG8 shows a schematic diagram of the distribution of reference areas corresponding to the current block in one embodiment of the present application.
[0126] As shown in FIG8 , the reference region corresponding to the current block 801 may be composed of multiple sub-regions 802, wherein each sub-region 802 may be the nearest neighbor region or the next nearest neighbor region of the current block. The nearest neighbor region may, for example, include the image region B located above the current block 801 or the image region D located to the left of the current block 801, and the next nearest neighbor region may, for example, include the image region A located to the upper left of the current block 801, the image region E located to the lower left of the current block 801, or the image region C located to the upper right of the current block 801.
[0127] One or more of the image areas AE may be combined to form a reference area of the current block 801 .
[0128] In one embodiment of the present application, the image sizes of the sub-regions constituting the reference region are specified as follows.
[0129] The nearest neighbor region B above the current block 801 has the same image size as the current block 801 in the horizontal direction, and the nearest neighbor region B above the current block 801 has a specified image size in the vertical direction.
[0130] The nearest neighbor region D on the left side of the current block 801 has the same image size as the current block 801 in the vertical direction, and the nearest neighbor region D on the left side of the current block 801 has a specified image size in the horizontal direction.
[0131] The next neighboring region C above and to the right of the current block 801 has the same image size as the current block 801 in the horizontal direction, and has a specified image size in the vertical direction.
[0132] The next neighboring region E below the left of the current block 801 has the same image size as the current block 801 in the vertical direction, and the next neighboring region below the left of the current block 801 has a specified image size in the horizontal direction.
[0133] The next neighboring area A at the upper left of the current block 801 has a specified image size in both the horizontal and vertical directions.
[0134] The designated image size may be a preset value greater than or equal to one, for example, the designated image size may be set to 6. When the designated image size is greater than one, color component prediction may be performed using multiple layers of neighborhood pixels, thereby improving the accuracy of color component prediction.
[0135] In one embodiment of the present application, the sub-regions of the reference region constituting the current block may have the same specified image size or different specified image sizes. For example, the vertical size of image region C may be the same as or different from the horizontal size of image region E.
[0136] In one embodiment of the present application, pixels that have been decoded and reconstructed or are allowed to be available can be selected from the sub-region as reference pixels for color component prediction of the current block. For example, when some pixels in image region C have been decoded and reconstructed, while other pixels have not been decoded and reconstructed, pixels that have been decoded and reconstructed can be selected from image region C as reference pixels for color component prediction of the current block.
[0137] In one embodiment of the present application, when the decoded and reconstructed pixels in a sub-region do not meet the aforementioned size requirements, the sub-neighboring region can be configured as unavailable. For example, when the lower right corner of image region C is not reconstructed or exceeds the image boundary, image region C can be configured as unavailable; for another example, when the lower right corner of image region E is not reconstructed or exceeds the image boundary, image region E can be configured as unavailable.
[0138] In one embodiment of the present application, all the sub-regions shown in FIG. 8 may be combined to form a reference region of the current block, or a portion of the sub-regions may be combined to form a reference region of the current block.
[0139] Figure 9 shows a schematic diagram of a region template that combines some subregions into a reference region in one embodiment of the present application. As shown in Figure 9, based on the combination of different subregions, ten exemplary candidate region templates can be formed. In this embodiment of the present application, one or more of these candidate region templates can be specified for the current block.
[0140] Figure 10 shows a schematic diagram of a region template of a reference region selected for a current block in one embodiment of the present application. As shown in Figure 10, in this embodiment of the present application, the reference region corresponding to the current block includes at least one of a full region combination, a left region combination, or an upper region combination.
[0141] The full region combination includes the nearest neighbor regions located on the left and above the current block and the next nearest neighbor regions located on the upper left, lower left and upper right of the current block.
[0142] The left region combination includes the nearest neighbor region located on the left side of the current block and the next nearest neighbor region located on the lower left side of the current block.
[0143] The upper region combination includes the nearest neighbor region located above the current block and the next nearest neighbor region located above and to the right of the current block.
[0144] In one embodiment of the present application, an indicator field can be used in the video bitstream to identify the region template of the reference region used by the current block. For example, when the indicator field value is 1, it indicates that the reference region selected by the current block is the full region combination shown in Figure 10; when the indicator field value is 01, it indicates that the reference region selected by the current block is the left region combination shown in Figure 10; when the indicator field value is 00, it indicates that the reference region selected by the current block is the upper region combination shown in Figure 10.
[0145] In one embodiment of the present application, after determining the reference area used when predicting the color components of the current block, availability information of one or more sub-areas constituting the reference area can be obtained, and then the area range of the reference area can be adjusted according to the availability information of the one or more sub-areas.
[0146] In one embodiment of the present application, adjusting the area range of the reference area based on the availability information of the one or more sub-areas may further include: removing sub-areas in an unavailable state from the reference area; and configuring the reference area to an unavailable state when all sub-areas in the reference area are in an unavailable state.
[0147] For example, the value of the indication field corresponding to the current block obtained by parsing the video code stream is 1, indicating that the reference area used for color component prediction of the current block is the full area combination including the five sub-areas AE shown in FIG. 10 .
[0148] When performing coding prediction on the current block, availability information of each sub-region in the reference region may be obtained, and the region range of the reference region may be adjusted according to the availability information.
[0149] For example, sub-regions A, B, and C are image regions of other image blocks located above the current block. If the image blocks where sub-regions A, B, and C are located have not yet completed encoding and reconstruction, sub-regions A, B, and C are in an unavailable state. At this time, the area range of the reference area can be adjusted from A+B+C+D+E to D+E.
[0150] For another example, sub-regions D and E are image regions of other image blocks located to the left of the current block. If the image blocks containing sub-regions D and E have not yet been coded and reconstructed, sub-regions D and E are also unavailable. In this case, all five sub-regions A to E are unavailable, so the entire reference region based on this full region combination can be configured as unavailable.
[0151] In one embodiment of the present application, the sample pixels used to train the prediction model may be all or a portion of the pixels selected from at least one image region between the current block and the reference region. When a portion of the pixels is selected to train the prediction model, pixel sampling may be performed within at least one image region between the current block and the reference region according to a preset sampling rule.
[0152] In one embodiment of the present application, selecting training sample pairs in at least one image area of the current block and the reference area may further include: obtaining a pixel sampling mode of the current block, where the pixel sampling mode includes full pixel sampling or partial pixel sampling; when the pixel sampling mode of the current block is full pixel sampling, selecting all pixels in at least one image area of the current block and the reference area as sample pixels, and obtaining a training sample pair consisting of a first color component and a second color component of the sample pixels; when the pixel sampling mode of the current block is partial pixel sampling, selecting a portion of pixels with specified sampling positions in at least one image area of the current block and the reference area as sample pixels, and obtaining a training sample pair consisting of the first color component and the second color component of the sample pixels.
[0153] In one embodiment of the present application, the designated sampling position includes at least one of the following three sampling positions.
[0154] The first sampling position: a designated sampling position where the position coordinates of the pixel point meet the preset coordinate value conditions.
[0155] In one embodiment of the present application, the coordinate value condition includes: at least one of the horizontal position coordinate or the vertical position coordinate of the pixel point is an even number; or, at least one of the horizontal position coordinate or the vertical position coordinate of the pixel point is an odd number.
[0156] FIG. 11 is a schematic diagram showing a method of selecting a sampling position based on the position coordinates of a pixel point in one embodiment of the present application.
[0157] As shown in Figure 11, in the reference area, a coordinate system is established with the upper left corner as the coordinate origin (0, 0), the horizontal position coordinate x represents the sequential position of the pixel points arranged from left to right in the horizontal direction, and the vertical position coordinate y represents the sequential position of the pixel points arranged from top to bottom in the horizontal direction.
[0158] According to the preset coordinate value conditions, even positions, odd positions or all positions can be selected as designated sampling positions in the horizontal direction, and even positions, odd positions or all positions can be selected as designated sampling positions in the vertical direction.
[0159] For example, in the embodiment shown in FIG11 , pixel positions that are even positions in the horizontal direction and all positions in the vertical direction are selected as designated sampling positions, that is, pixel positions in the shaded portion of the figure are selected as designated sampling positions.
[0160] The second sampling position: a specified sampling position selected along the preset pixel scanning direction.
[0161] In one embodiment of the present application, pixels in the reference area can be scanned using any scanning method, such as a round-trip scan or a ZigZag scan, so that designated sampling positions are selected along a preset pixel scanning direction in accordance with a sequence and a preset sampling rule. The preset sampling rule may be, for example, interval sampling, where a designated sampling position is selected after one or more scanned pixels have been scanned.
[0162] FIG12 is a schematic diagram showing a method of selecting sampling positions based on a round-trip scanning method in one embodiment of the present application.
[0163] As shown in Figure 12, a row of pixels is scanned from left to right in the reference area. After reaching the boundary, the next row of pixels is scanned from right to left. During the pixel scanning process, a designated sampling position is selected for every other pixel along the scanning direction. For example, the arrow in the figure indicates the scanning direction, and the shaded pixel positions are the selected designated sampling positions.
[0164] FIG13 shows a schematic diagram of selecting sampling positions based on a ZigZag scanning method in one embodiment of the present application.
[0165] As shown in Figure 13, pixels are scanned from the upper left corner to the lower right corner of the reference area using a ZigZag scanning method. During the pixel scanning process, a designated sampling position is selected for every other pixel along the scanning direction. For example, the arrow in the figure indicates the scanning direction, and the shaded pixel positions are the selected designated sampling positions.
[0166] The third sampling position: a designated sampling position where the reconstructed value of the first color component falls within a preset value range.
[0167] Taking the first color component as the brightness component as an example, the embodiment of the present application can select the position of the pixel point whose brightness value is greater than or less than a certain threshold as the designated sampling position.
[0168] For example, a prediction model can be fitted by selecting sample pixel points in the reference area whose brightness values are greater than a set threshold, and another prediction model can be fitted by selecting sample pixel points whose brightness values are less than or equal to the set threshold; when generating a predicted image, the corresponding first prediction model can be used for pixel points whose brightness values are greater than the set threshold, and the corresponding second prediction model can be used for pixel points whose brightness values are less than or equal to the set threshold.
[0169] S630: Parameter fitting is performed on the initial prediction model according to the training samples to obtain a target prediction model.
[0170] In one embodiment of the present application, the prediction model can fit the model parameters through online training or offline training.
[0171] Among them, online training refers to training and fitting during the encoding and decoding process of the current video frame, wherein the model parameters of each current block can be calculated online during encoding and decoding.
[0172] Offline training involves fitting and calculating model parameters offline, outside of the encoding and decoding process. During offline training, model parameters are trained offline based on a pre-collected sample dataset. The trained target prediction model is then used directly during encoding and decoding, without further model parameter calculations during the encoding and decoding of the current video frame. Compared to online training, offline training reduces model prediction accuracy, but offers faster encoding and decoding speeds. Therefore, it is suitable for applications where encoding and decoding quality is less critical but speed is more important.
[0173] In one embodiment of the present application, during the training process of the prediction model, multiple sample pixels can be sampled in at least one image area in the current block and the reference area, and then the initial prediction model is parameter-fitted according to the first color component and the second color component of the sample pixel to obtain the target prediction model.
[0174] For example, in the embodiment of the present application, the first color component and the second color component of the sample pixel point can be input into the initial prediction model to establish multiple equations, namely, Ax=b.
[0175] Wherein, A is a matrix with M rows and N columns, M and N represent the number of equations established according to the initial prediction model and the number of model parameters in the initial prediction model, respectively, and x represents a parameter vector composed of the model parameters in the initial prediction model as elements.
[0176] Solving the parameter vector x in the above equation can obtain all model parameters p i value.
[0177] In one embodiment of the present application, the LDL decomposition method or the Gaussian elimination method can be selected to solve the above equation. Taking the LDL decomposition method as an example, the solution steps may include:
[0178] (1) Transform the equation into A T Ax=A T b;
[0179] (2) To A T A is decomposed to obtain LDL T x=A T b;
[0180] (3) Solve LY = A T bGet the matrix Y;
[0181] (4) Solving DL Tx=Y to get the parameter vector x, thereby obtaining the model parameter p i .
[0182] In one embodiment of the present application, when the luminance component Y is selected as the first color component and the chrominance component U and the chrominance component V are selected as the second color component, that is, the chrominance component U and the chrominance component V are predicted based on the luminance component Y, since both prediction processes require A T A is decomposed, and the two chrominance component prediction processes can be combined. T The decomposition process of A can reduce the computational complexity and improve the coding efficiency.
[0183] In step S430 , mapping processing is performed on the reconstructed value of the first color component of the current block according to the mapping relationship to obtain a predicted value of the second color component of the current block.
[0184] Based on the above embodiments, it can be seen that based on the fitted target prediction model, the reconstructed value of the first color component of the current block can be input into the target prediction model as an input item to obtain the predicted value of the second color component output by the target prediction model.
[0185] In one embodiment of the present application, color component prediction may be performed on a portion of image blocks in the current video frame. Based on this, it may be explicitly or implicitly identified which image blocks perform a color component prediction method for predicting a second color component based on a first color component.
[0186] In one embodiment of the present application, mapping the reconstructed value of the first color component of the current block according to the mapping relationship to obtain the predicted value of the second color component of the current block may further include: obtaining a color component prediction condition of the current block, the color component prediction condition is used to indicate whether to predict another color component based on a color component of the current block; when the current block meets the color component prediction condition, mapping the reconstructed value of the first color component of the current block according to the mapping relationship to obtain the predicted value of the second color component of the current block.
[0187] In one embodiment of the present application, the color component prediction condition includes at least one of the following four conditions:
[0188] The first color component prediction condition is that the index corresponding to the current block has a specified index value.
[0189] For example, for a current block, an index can be decoded in the sequence header / image header / slice header / maximum coding block, which is used to indicate whether all image blocks in the current sequence / image / slice / maximum coding block execute the color component prediction method of predicting the second color component based on the first color component.
[0190] The first color component prediction condition: the reconstruction residual of the first color component of the pixel point in the current block falls within a preset value range.
[0191] For example, within a current block, the color component prediction method for predicting the second color component based on the first color component is performed for the pixel only when the reconstruction residual of the first color component of the pixel is greater than zero or less than zero. When the reconstruction residual of the first color component of the pixel is zero, the color component prediction step for predicting the second color component based on the first color component of the pixel can be skipped.
[0192] The third color component prediction condition: the image features of the current block itself meet the preset feature conditions.
[0193] In one embodiment of the present application, the image features of the current block itself meet preset feature conditions, which may include: the image size of the current block falls within a specified size range; or the position of the current block falls within a specified area range.
[0194] In an optional implementation, whether the current block executes the color component prediction method of predicting the second color component based on the first color component may be indicated based on whether the image size of the current block meets a preset range limit.
[0195] In an optional embodiment, whether to execute the color component prediction method for predicting the second color component based on the first color component for the current block may be indicated based on whether the position of the current block meets a preset range limit. For example, if the current block is located in the upper left corner of the current video frame, the color component prediction method for predicting the second color component based on the first color component is not executed.
[0196] The fourth color component prediction condition: the image features of the reference area corresponding to the current block meet the preset feature conditions.
[0197] In one embodiment of the present application, the image features of the reference area corresponding to the current block meet preset feature conditions, including: the area of the reference area corresponding to the current block is greater than a specified area threshold; or, the number of specified sampling positions in the reference area corresponding to the current block is greater than the number of model parameters of the prediction model; the prediction model is used to indicate the mapping relationship between the first color component and the second color component.
[0198] In an optional implementation, the color component prediction method of predicting the second color component based on the first color component is performed on the current block only when the area of the reference region is greater than a specified area threshold.
[0199] In an optional embodiment, the color component prediction method of predicting the second color component based on the first color component is performed on the current block only when the number of specified sampling positions in the reference area (i.e., the number of matching pairs of the first color component and the second color component) is greater than the number of model parameters of the prediction model.
[0200] In one embodiment of the present application, before mapping the reconstructed value of the first color component of the current block according to the mapping relationship, the first color component may be downsampled to obtain a first color component that can form a matching pair with the second color component.
[0201] Taking the YUV image block shown in Figure 5 as an example, when the first color component is luma Y and the second color component is chroma U or V, the luma and chroma blocks obtained by sampling the color components using a 420 sampling format have different block sizes. The resolution of the luma component is twice that of the chroma components in both the horizontal and numerical directions. In this case, to form a matching pair of luma and chroma components, the luma component can be downsampled.
[0202] In one embodiment of the present application, a method for downsampling a first color component may include: determining a sampling window with a specified window size based on a pixel position of a second color component of a current block; downsampling a reconstructed value of the first color component of the current block in the sampling window to obtain a reconstructed value of the first color component that matches the pixel position.
[0203] In one embodiment of the present application, the sampling window determined according to the pixel position of the second color component is used to cover the four nearest first color components and the two next-nearest first color components on the left.
[0204] FIG14 shows a schematic diagram of a sampling window for downsampling the first color component in one embodiment of the present application.
[0205] As shown in FIG14 , the blocks distributed in an array represent a first color component 1401, and the five-pointed stars distributed within the array represent a second color component 1402. In the embodiment of the present application, a sampling window 1403 having a specified window size corresponding to the pixel location of the second color component 1402 can be determined. For example, the sampling window 1403 in the embodiment of the present application can cover the four first color components that are the closest neighbors to the second color component and the two first color components that are the next closest neighbors to the left.
[0206] In one embodiment of the present application, downsampling the reconstructed value of the first color component of the current block in a sampling window to obtain a reconstructed value of the first color component that matches the pixel position may further include: obtaining the positional relationship between the reconstructed value of the first color component of the current block and the pixel position in the sampling window; performing a weighted operation on the reconstructed value of the first color component according to the positional relationship to obtain a reconstructed value of the first color component that matches the pixel position.
[0207] For example, the embodiment of the present application can downsample the first color component according to the following formula.
[0208] L′(x,y)=(L(x-1,y)+2L(x,y)+L(x+1,y)+L(x-1,y+1)+2L(x,y+1)+L(x+1,y+1)) / 8
[0209] Wherein, L and L′ represent the original first color component and the downsampled first color component, respectively.
[0210] x and y represent the horizontal and vertical position coordinates of the first color component, respectively. Taking the sampling window shown in Figure 14 as an example, the two first color components nearest to the left of the second color component, namely (x, y) and (x, y+1), can be assigned a weight coefficient of 1 / 4; the other four first color components, namely (x-1, y), (x-1, y+1), (x+1, y), and (x+1, y+1), can be assigned a weight coefficient of 1 / 8.
[0211] In one embodiment of the present application, when the two first color components of the next nearest neighbor on the left are not available, that is, (x-1, y) or (x-1, y+1) are not available, the two first color components of the nearest neighbor on the left can be used instead, that is, (x, y) or (x, y+1) can be used instead.
[0212] In one embodiment of the present application, when downsampling the first color component of the current block is selected, the first color component serving as a model training sample can be synchronously downsampled within at least one image region in the current block and the reference region. That is, before fitting the mapping relationship between the first color component and the second color component based on the at least one image region in the current block and the reference region, a sampling window with a specified window size is determined based on the pixel positions of the second color component within the at least one image region in the current block and the reference region; the first color component is downsampled within the sampling window to obtain a first color component that matches the pixel positions.
[0213] In one embodiment of the present application, the first color component may be downsampled, or downsampling may not be performed, and the first color component having the same or corresponding position as the second color component may be selected for color component prediction. Taking FIG. 14 as an example, the first color component 1401 having a position coordinate (x, y) corresponding to the second color component 1402 may be selected for color component prediction.
[0214] In one embodiment of the present application, when the sampling format of the first color component and the second color component is the same, before the color component is predicted by the prediction model, the first color component can be filtered to reduce redundant information or interference information in the first color component, thereby further improving the prediction accuracy of the second color component, that is, improving the video encoding accuracy of the current block.
[0215] On this basis, when fitting the prediction model, multiple filters can be selected at the same time to filter the sample data of the first color component, and then different first color component samples can be selected to train the fitting prediction model to improve the fitting training effect of the prediction model.
[0216] In one embodiment of the present application, after downsampling or filtering the first color component, the pixel position of the first color component can be boundary-extended as needed to obtain neighboring pixel points for the designated pixel point.
[0217] It should be noted that the above embodiments introduce a variety of different color component prediction schemes. For example, the reference area includes a variety of different candidate area templates, and the prediction model also includes a variety of different candidate prediction models. By combining a variety of different candidate area templates with a variety of different candidate prediction models, a variety of different color component prediction schemes can be obtained.
[0218] For different image blocks, you can choose to use the same color component prediction scheme or different color component prediction schemes.
[0219] In one embodiment of the present application, for example, different color component prediction schemes (candidate region templates, candidate prediction models) may be selected for image blocks with different block sizes.
[0220] In one embodiment of the present application, the color component prediction scheme used by the current block can be indicated in the data code stream through an explicit index or an implicit index. For example, the candidate area template and / or candidate prediction model used by the current block can be indicated by one index or a combination of multiple indexes.
[0221] In one embodiment of the present application, a plurality of different color component prediction schemes may be used for a current block, and then the prediction results of the plurality of schemes may be weighted to obtain a final color component prediction result.
[0222] In one embodiment of the present application, the final predicted image can be obtained by weighting the color component prediction results provided by the above embodiments and the prediction results of the common inter-frame mode or the intra-frame block copy mode.
[0223] In one embodiment of the present application, inter-frame prediction or intra-frame block copy prediction is performed on the current block to obtain an initial prediction value of the second color component of the current block; after the reconstructed value of the first color component of the current block is mapped according to the mapping relationship to obtain the prediction value of the second color component of the current block, a weighted operation is performed on the initial prediction value and the prediction value of the second color component according to a preset weight coefficient to obtain an updated prediction value of the second color component.
[0224] For example, in an embodiment of the present application, the final prediction result pred of the second color component of the current block may be determined according to the following formula.
[0225] pred=w*pred1+(1-w)*pred2
[0226] Among them, pred1 represents the predicted value of the second color component obtained based on the first color component according to the scheme provided in any of the above embodiments, pred2 represents the initial predicted value of the second color component obtained according to the ordinary inter-frame prediction mode or the intra-frame block copy mode, and w represents the weight coefficient for weighted operation of the two prediction results.
[0227] Figure 15 shows a step flow chart of a video encoding method in an embodiment of the present application. The video encoding method can be executed by a terminal device or a server that sends encoded data. The embodiment of the present application uses the video encoding method executed by a terminal device as an example to illustrate. The terminal device can be, for example, the video encoding device 203 shown in Figure 2.
[0228] As shown in FIG. 15 , the video encoding method in the embodiment of the present application includes the following steps S1510 to S1530 .
[0229] S1510: Obtain a reconstructed value of a first color component of a current block, where the current block is an image block to be encoded in a current video frame.
[0230] S1520: Fitting a mapping relationship between the first color component and the second color component according to the current block and at least one image region in a reference region, where the reference region includes one or more encoded and reconstructed image regions that are the nearest or next nearest neighbor to the current block.
[0231] S1530: Mapping the reconstructed value of the first color component of the current block according to the mapping relationship to obtain a predicted value of the second color component of the current block.
[0232] The various method steps of the video encoding method in the embodiment of the present application correspond one-to-one to the video decoding method in the above embodiment, and will not be repeated here.
[0233] The following describes the implementation process of some embodiments of the technical solution of the present application in multiple application scenarios.
[0234] In the first application scenario, the inter-frame prediction coding mode or the intra-frame block copy coding mode is used to predict the luminance component Y of the current block, and the chrominance component U and the chrominance component V are predicted using the luminance component Y. The prediction model used is as follows.
[0235] In this application scenario, the reference area corresponding to the current block is used to train the prediction model, that is, the pixels in the reference area are selected as sample pixels for training the prediction model. Accordingly, the training sample includes the luminance component reconstruction value and the chrominance component reconstruction value of the sample pixel in the reference area.
[0236] The reference area may be selected from a plurality of image areas that are the nearest neighbor and the next nearest neighbor to the current block as shown in FIG8 . That is, the two nearest neighbor areas located above and to the left of the current block and the three next nearest neighbor areas located above, below, and to the right of the current block are selected.
[0237] After determining the prediction model and the image used to select sample pixels, the luminance components in the current block and the reference area can be downsampled, for example, using the sampling window shown in Figure 14. The downsampling process is as follows.
[0238] L′(x,y)=(L(x-1,y)+2L(x,y)+L(x+1,y)+L(x-1,y+1)+2L(x,y+1)+L(x+1,y+1)) / 8
[0239] After downsampling the luma component reconstructed values, luma component reconstructed values whose positional coordinates in both the horizontal and vertical directions are not odd numbers are selected as sample pixels in the reference area to construct sample pairs of luma component reconstructed values and chroma component reconstructed values. In other optional implementations, luma component reconstructed values whose positional coordinates in at least one of the horizontal and vertical directions are even numbers may also be selected as sample pixels in the reference area.
[0240] Input the sample pairs into the prediction model, establish the equation and solve the various model parameters p in the prediction model. i, respectively obtain a first prediction model for predicting the chrominance component U according to the luminance component Y and a second prediction model for predicting the chrominance component V according to the luminance component Y.
[0241] After downsampling the luminance reconstruction value in the current block, the values are input into the first prediction model and the second prediction model that have been fitted and trained, respectively, to obtain prediction results of the chrominance component U and the chrominance component V.
[0242] In the second application scenario, the inter-frame prediction coding mode or the intra-frame block copy coding mode is used to predict the luminance component Y, chrominance component U, and chrominance component V of the current block, and the luminance component Y is used to predict the chrominance component U and chrominance component V respectively. The prediction model used is as follows.
[0243] C b =p0C+p1S+p2W+p3E+p4SW+p5SE+p6C 2 +p7B
[0244] In this application scenario, the current block is used as the image region for training the prediction model. Specifically, pixels in the current block are selected as sample pixels for training the prediction model. Accordingly, the training sample pairs include the predicted values for the luminance and chrominance components of the sample pixels in the current block.
[0245] After the current block is determined as the image region for training the prediction model, the brightness component prediction value in the current block can be downsampled, for example, using the sampling window shown in Figure 14. The downsampling process is as follows.
[0246] L′(x,y)=(L(x-1,y)+2L(x,y)+L(x+1,y)+L(x-1,y+1)+2L(x,y+1)+L(x+1,y+1)) / 8
[0247] After completing the downsampling of the luminance component, the luminance components whose position coordinates are not odd in the horizontal and vertical directions are selected as sample pixels in the current block to construct sample pairs of luminance component prediction values and chrominance component prediction values.
[0248] Input the sample pairs into the prediction model, establish the equation and solve the various model parameters p in the prediction model. i , respectively obtain a first prediction model for predicting the chrominance component U according to the luminance component Y and a second prediction model for predicting the chrominance component V according to the luminance component Y.
[0249] After downsampling the luminance reconstruction value in the current block, the values are input into the first prediction model and the second prediction model that have been fitted and trained, respectively, to obtain prediction results of the chrominance component U and the chrominance component V.
[0250] In the third application scenario, the inter-frame prediction coding mode or the intra-frame block copy coding mode is used to perform coding prediction on the luminance component Y, the chrominance component U, and the chrominance component V, and the luminance component Y is used to predict the chrominance component U and the chrominance component V. The prediction model used is as follows.
[0251] C b =p0C+p1S+p2W+p3E+p4SW+p5SE+p6C 2 +p7B
[0252] In this application scenario, the current block is used as the image region for training the prediction model. Specifically, pixels in the current block are selected as sample pixels for training the prediction model. Accordingly, the training sample pairs include the predicted values for the luminance and chrominance components of the sample pixels in the current block.
[0253] In this application scenario, the luminance component in the current block is not downsampled, but the model is directly fitted. That is, all luminance components in the current block are used as sample pixels to construct sample pairs of luminance component prediction values and chrominance component prediction values.
[0254] Input the sample pairs into the prediction model, establish the equation and solve the various model parameters p in the prediction model. i , respectively obtain a first prediction model for predicting the chrominance component U according to the luminance component Y and a second prediction model for predicting the chrominance component V according to the luminance component Y.
[0255] The luminance reconstruction value in the current block is input into the first prediction model and the second prediction model that have been fitted and trained, respectively, to obtain the prediction results of the chrominance component U and the chrominance component V.
[0256] The prediction results of the chrominance component U and the chrominance component V are weighted with the conventional prediction results according to the following formula to obtain the final prediction results pred of the chrominance component U and the chrominance component V. The conventional prediction results may include, for example, the prediction results of the inter-frame prediction or the prediction results of the intra-frame block copy.
[0257] pred=w*pred1+(1-w)*pred2
[0258] Among them, pred1 represents the predicted value of the chrominance component U or the chrominance component V obtained by predicting the luminance component Y, pred2 represents the predicted value of the chrominance component U or the chrominance component V obtained by predicting the ordinary inter-frame prediction mode or the intra-frame block copy mode, and w represents the weight coefficient for weighted operation of the two prediction results, which can be, for example, 0.75.
[0259] Based on the introduction of the above embodiments and application scenarios, it can be seen that the embodiment of the present application proposes a color component prediction method. By utilizing the similarity of the mapping relationship between the components of the adjacent regions or the predicted image and the reconstructed image, a cross-component prediction model is generated based on the predicted image of the adjacent reconstructed pixels or the current block, and then the cross-component prediction model is used to generate the target predicted image of the current coding block. Compared with conventional coding prediction schemes, the embodiment of the present application can further utilize the reconstructed pixels of the current block to improve the prediction accuracy of the color components, thereby improving coding efficiency.
[0260] It should be noted that although the steps of the method of the present application are described in a specific order in the drawings, this does not require or imply that the steps must be performed in this specific order, or that all steps must be performed to achieve the desired results. Additionally or alternatively, some steps may be omitted, multiple steps may be combined into one step, and / or one step may be decomposed into multiple steps.
[0261] The following describes an apparatus embodiment of the present application, which can be used to execute the video encoding and decoding method in the above-mentioned embodiment of the present application.
[0262] FIG16 schematically shows a block diagram of a video decoding apparatus according to an embodiment of the present application. As shown in FIG16 , the video decoding apparatus 1600 includes:
[0263] A first acquisition module 1610 is configured to obtain a reconstructed value of a first color component of a current block, where the current block is an image block to be decoded in a current video frame;
[0264] A first fitting module 1620 is configured to fit a mapping relationship between the first color component and the second color component according to the current block and at least one image region in a reference region, where the reference region includes one or more decoded image regions that are the nearest neighbor or the next nearest neighbor to the current block;
[0265] The first mapping module 1630 is configured to perform mapping processing on the reconstructed value of the first color component of the current block according to the mapping relationship to obtain a predicted value of the second color component of the current block.
[0266] In one embodiment of the present application, based on the above embodiment, the first acquisition module 1610 can be further configured to: obtain a prediction coding mode of the current block; when the prediction coding mode is inter-frame prediction, perform inter-frame prediction based on a first reference block corresponding to the current block to obtain a reconstructed value of the first color component of the current block, and the first reference block is a decoded image block in a reference video frame corresponding to the current video frame; when the prediction coding mode is intra-frame block copy, perform intra-frame block copy prediction based on a second reference block corresponding to the current block to obtain a reconstructed value of the first color component of the current block, and the second reference block is a decoded image block in the current video frame.
[0267] In one embodiment of the present application, based on the above embodiment, the video decoding device 1600 further includes:
[0268] The weighting module is configured to perform inter-frame prediction or intra-frame block copy prediction on the current block to obtain an initial prediction value of the second color component of the current block; and perform a weighted operation on the initial prediction value and the prediction value of the second color component according to a preset weight coefficient to obtain an updated prediction value of the second color component.
[0269] In one embodiment of the present application, based on the above embodiment, the first fitting module 1620 may further include:
[0270] a model acquisition module configured to acquire an initial prediction model corresponding to the current block, wherein an input item of the initial prediction model includes a first color component of a specified pixel point, an output item of the initial prediction model is a second color component of the specified pixel point, and the initial prediction model includes at least one of a plurality of candidate models;
[0271] a sample pair selection module configured to select a training sample pair in at least one image area between the current block and the reference area, the training sample pair comprising one or more of the following sample pairs: a predicted value of the first color component and a predicted value of the second color component of a sample pixel located in the current block, and a reconstructed value of the first color component and a reconstructed value of the second color component of a sample pixel located in the reference area;
[0272] The model fitting module is configured to perform parameter fitting on the initial prediction model according to the training sample to obtain a target prediction model.
[0273] In one embodiment of the present application, based on the above embodiment, the predicted value of the first color component and the predicted value of the second color component of the sample pixel point located in the current block are obtained according to the following method: obtaining a prediction coding mode of the current block; when the prediction coding mode is inter-frame prediction, performing inter-frame prediction based on a first reference block corresponding to the current block to obtain the predicted value of the first color component and the predicted value of the second color component of the sample pixel point in the current block, where the first reference block is a decoded image block in a reference video frame corresponding to the current video frame; when the prediction coding mode is intra-frame block copying, performing intra-frame block copying prediction based on a second reference block corresponding to the current block to obtain the predicted value of the first color component and the predicted value of the second color component of the sample pixel point in the current block, where the second reference block is a decoded image block in the current video frame.
[0274] In one embodiment of the present application, based on the above embodiment, the input item of the initial prediction model also includes the first color component of one or more neighborhood pixels, and the neighborhood pixels are the pixels that are the nearest neighbors or second-nearest neighbors of the specified pixel.
[0275] In one embodiment of the present application, based on the above embodiment, the initial prediction model includes one or more combination items with independent weighting parameters, and the combination item takes the first color component of one or more pixel points as an input item.
[0276] In one embodiment of the present application, based on the above embodiment, the initial prediction model includes at least two combination terms with different orders, and the order is the highest power of the input terms in the combination terms.
[0277] In one embodiment of the present application, based on the above embodiment, the reference area is composed of one or more nearest neighbor areas or second-nearest neighbor areas, the nearest neighbor area includes an image area with a specified image size located above or to the left of the current block, and the second-nearest neighbor area includes an image area with a specified image size located above, below, or above the current block.
[0278] In one embodiment of the present application, based on the above embodiment, the next-nearest neighboring region to the upper right of the current block has the same image size as the current block in the horizontal direction, and the next-nearest neighboring region to the upper right of the current block has a specified image size in the vertical direction, and the specified image size is greater than or equal to one;
[0279] The next neighboring area below the left of the current block has the same image size as the current block in the vertical direction, and the next neighboring area above the right of the current block has the specified image size in the horizontal direction;
[0280] The next neighboring area above and to the left of the current block has the specified image size in both the horizontal direction and the vertical direction.
[0281] In one embodiment of the present application, based on the above embodiment, the reference area includes at least one of a full area combination, a left area combination, and an upper area combination;
[0282] The full region combination includes the nearest neighbor regions located on the left and above the current block and the next nearest neighbor regions located on the upper left, lower left, and upper right of the current block;
[0283] The left region combination includes a nearest neighbor region located on the left side of the current block and a next nearest neighbor region located on the lower left side of the current block;
[0284] The upper region combination includes a nearest neighbor region located above the current block and a next nearest neighbor region located to the upper right of the current block.
[0285] In one embodiment of the present application, based on the above embodiment, the sample pair selection module can be further configured to: obtain a pixel sampling mode of the current block, where the pixel sampling mode includes full pixel sampling or partial pixel sampling; when the pixel sampling mode of the current block is full pixel sampling, all pixels are selected as sample pixels in at least one image area between the current block and the reference area, and a training sample pair consisting of a first color component and a second color component of the sample pixels is obtained; when the pixel sampling mode of the current block is partial pixel sampling, a portion of pixels with a specified sampling position are selected as sample pixels in at least one image area between the current block and the reference area, and a training sample pair consisting of the first color component and the second color component of the sample pixels is obtained.
[0286] In one embodiment of the present application, based on the above embodiment, the designated sampling position includes at least one of the following sampling positions:
[0287] The pixel point's position coordinates meet the specified sampling position of the preset coordinate value conditions;
[0288] A designated sampling position is selected along a preset pixel scanning direction;
[0289] The reconstructed value of the first color component falls within a specified sampling position within a preset value range.
[0290] In one embodiment of the present application, based on the above embodiment, the coordinate value conditions include:
[0291] At least one of the horizontal position coordinate and the vertical position coordinate of the pixel point is an even number;
[0292] Alternatively, at least one of the horizontal position coordinate and the vertical position coordinate of the pixel point is an odd number.
[0293] In one embodiment of the present application, based on the above embodiment, the first mapping module 1630 may further include:
[0294] a prediction condition acquisition module configured to acquire a color component prediction condition of the current block, wherein the color component prediction condition is used to indicate whether to predict another color component based on a color component of the current block;
[0295] The reconstruction value mapping module is configured to, when the current block meets the color component prediction condition, map the reconstruction value of the first color component of the current block according to the mapping relationship to obtain the prediction value of the second color component of the current block.
[0296] In one embodiment of the present application, based on the above embodiment, the color component prediction condition includes at least one of the following conditions:
[0297] An index corresponding to the current block has a specified index value;
[0298] The reconstructed residual of the first color component of the pixel point in the current block falls within a preset value range;
[0299] The image features of the current block itself meet the preset feature conditions;
[0300] The image features of the reference area meet preset feature conditions.
[0301] In one embodiment of the present application, based on the above embodiments, the reconstruction residual of the first color component of the pixel point in the current block falls within a preset numerical range, including: the reconstruction residual of the first color component of the pixel point in the current block is greater than zero or less than zero.
[0302] In one embodiment of the present application, based on the above embodiment, the image features of the current block itself meet preset feature conditions, including: the image size of the current block falls within a specified size range; or the position of the current block falls within a specified area range.
[0303] In one embodiment of the present application, based on the above embodiments, the image features of the reference area meet preset feature conditions, including: the area of the reference area is greater than a specified area threshold; or, the number of specified sampling positions in the reference area is greater than the number of model parameters of the prediction model; the prediction model is used to indicate the mapping relationship between the first color component and the second color component.
[0304] In one embodiment of the present application, based on the above embodiment, the video decoding device 1600 may further include:
[0305] The first downsampling module is configured to determine a sampling window with a specified window size based on the pixel position of the second color component of the current block; downsample the reconstructed value of the first color component of the current block in the sampling window to obtain a reconstructed value of the first color component that matches the pixel position.
[0306] In one embodiment of the present application, based on the above embodiment, the first downsampling module is further configured to: obtain the positional relationship between the reconstructed value of the first color component of the current block and the pixel position in the sampling window; perform a weighted operation on the reconstructed value of the first color component according to the positional relationship to obtain a reconstructed value of the first color component that matches the pixel position.
[0307] In one embodiment of the present application, based on the above embodiment, the video decoding device 1600 may further include:
[0308] The second downsampling module is configured to determine a sampling window with a specified window size according to the pixel position of the second color component in at least one image area between the current block and the reference area; and downsample the first color component in the sampling window to obtain the first color component that matches the pixel position.
[0309] FIG17 schematically shows a block diagram of a video encoding apparatus according to an embodiment of the present application. As shown in FIG17 , the video encoding apparatus 1700 includes:
[0310] A second acquisition module 1710 is configured to obtain a reconstructed value of a first color component of a current block, where the current block is an image block to be encoded in a current video frame;
[0311] A second fitting module 1720 is configured to fit a mapping relationship between the first color component and the second color component according to the current block and at least one image region in a reference region, where the reference region includes one or more encoded and reconstructed image regions that are the nearest or next nearest neighbor to the current block;
[0312] The second mapping module 1730 is configured to perform mapping processing on the reconstructed value of the first color component of the current block according to the mapping relationship to obtain a predicted value of the second color component of the current block.
[0313] In the embodiment of the present application, the specific implementation method of each module in the video encoding device 1700 corresponds to the various modules in the video decoding device 1600. Please refer to the relevant description in the above embodiment and will not be repeated here.
[0314] The specific details of the video decoding device and the video encoding device provided in each embodiment of the present application have been described in detail in the corresponding method embodiments and will not be repeated here.
[0315] FIG18 schematically shows a block diagram of a computer system structure of an electronic device for implementing an embodiment of the present application.
[0316] It should be noted that the computer system 1800 of the electronic device shown in FIG18 is merely an example and should not impose any limitations on the functions and scope of use of the embodiments of the present application.
[0317] As shown in Figure 18, the computer system 1800 includes a central processing unit (CPU) 1801, which can perform various appropriate actions and processes according to the program stored in the read-only memory 1802 (ROM) or the program loaded from the storage part 1808 into the random access memory 1803 (RAM). Various programs and data required for system operation are also stored in the random access memory 1803. The CPU 1801, the read-only memory 1802, and the random access memory 1803 are connected to each other via a bus 1804. An input / output interface 1805 (i.e., an I / O interface) is also connected to the bus 1804.
[0318] The following components are connected to the input / output interface 1805: an input section 1806 including a keyboard, a mouse, and the like; an output section 1807 including devices such as a cathode ray tube (CRT), a liquid crystal display (LCD), and speakers; a storage section 1808 including a hard disk; and a communication section 1809 including a network interface card such as a local area network card or a modem. The communication section 1809 performs communication processing via a network such as the Internet. A drive 1810 is also connected to the input / output interface 1805 as needed. Removable media 1811, such as a magnetic disk, an optical disk, a magneto-optical disk, or a semiconductor memory, is installed in the drive 1810 as needed, so that computer programs read from the removable media can be installed in the storage section 1808 as needed.
[0319] In particular, according to an embodiment of the present application, the processes described in the various method flow charts can be implemented as computer software programs. For example, an embodiment of the present application includes a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for executing the methods shown in the flow charts. In such an embodiment, the computer program can be downloaded and installed from a network via the communication section 1809 and / or installed from a removable medium 1811. When the computer program is executed by the central processing unit 1801, the various functions defined in the system of the present application are performed.
[0320] It should be noted that the computer-readable medium shown in the embodiments of the present application may be a computer-readable signal medium or a computer-readable storage medium or any combination of the two. The computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or component, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM), a flash memory, an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present application, a computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in combination with an instruction execution system, device or device. In the present application, a computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, which carries a computer-readable program code. Such propagated data signals may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium that can transmit, propagate, or transport a program for use by or in connection with an instruction execution system, apparatus, or device. Program code embodied on a computer-readable medium may be transmitted using any suitable medium, including but not limited to wireless, wired, or any suitable combination thereof.
[0321] The flowcharts and block diagrams in the accompanying drawings illustrate the possible implementation architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present application. In this regard, each box in the flowchart or block diagram can represent a module, program segment, or a part of code, and the above-mentioned module, program segment, or a part of code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in an order different from that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram or flowchart, and the combination of the boxes in the block diagram or flowchart, can be implemented with a dedicated hardware-based system that performs the specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.
[0322] It should be noted that, although several modules or units of the device for action execution are mentioned in the above detailed description, this division is not mandatory. In fact, according to the embodiment of the application, the features and functions of two or more modules or units described above can be concretized in one module or unit. On the contrary, the features and functions of one module or unit described above can be further divided into multiple modules or units to be concretized.
[0323] Through the description of the above embodiments, it is easy for those skilled in the art to understand that the example embodiments described herein can be implemented by software or by combining software with necessary hardware. Therefore, the technical solution according to the embodiments of the present application can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (which can be a CD-ROM, a USB flash drive, a mobile hard disk, etc.) or on a network, and includes several instructions to enable a computing device (which can be a personal computer, a server, a touch terminal, or a network device, etc.) to execute the method according to the embodiments of the present application.
[0324] Those skilled in the art will readily appreciate other embodiments of the present invention after considering the specification and practicing the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of the present invention that follow the general principles of this application and include common knowledge or customary techniques in the art that are not disclosed herein.
[0325] It should be understood that the present application is not limited to the exact structures described above and shown in the drawings, and that various modifications and changes may be made without departing from the scope thereof. The scope of the present application is limited only by the appended claims.
Claims
1. A video decoding method, characterized in that: include: Obtaining a reconstructed value of a first color component of a current block, where the current block is an image block to be decoded in a current video frame; Fitting a mapping relationship between a first color component of the current block and a second color component of the current block according to the current block and at least one image region in a reference region, wherein the reference region includes one or more decoded image regions that are the nearest or second nearest neighbor to the current block; The reconstructed value of the first color component of the current block is mapped according to the mapping relationship to obtain a predicted value of the second color component of the current block.
2. The video decoding method according to claim 1, characterized in that: The method further comprises: Performing inter-frame prediction or intra-frame block copy prediction on the current block to obtain an initial prediction value of the second color component of the current block; A weighted operation is performed on the initial prediction value and the prediction value of the second color component according to a preset weight coefficient to obtain an updated prediction value of the second color component.
3. The video decoding method according to claim 1 or 2, characterized in that: The obtaining of the reconstruction value of the first color component of the current block includes: Obtaining a prediction coding mode of the current block; When the predictive coding mode is inter-frame prediction, inter-frame prediction is performed according to a first reference block corresponding to the current block to obtain a reconstructed value of a first color component of the current block, where the first reference block is a decoded image block in a reference video frame corresponding to the current video frame; When the prediction coding mode is intra block copy, intra block copy prediction is performed according to a second reference block corresponding to the current block to obtain a reconstructed value of the first color component of the current block, and the second reference block is a decoded image block in the current video frame.
4. The video decoding method according to any one of claims 1 to 3, characterized in that: The step of fitting a mapping relationship between a first color component of the current block and a second color component of the current block according to the current block and at least one image region in the reference region comprises: Acquire an initial prediction model corresponding to the current block, wherein an input item of the initial prediction model includes a first color component of a specified pixel point, an output item of the initial prediction model is a second color component of the specified pixel point, and the initial prediction model includes at least one of a plurality of candidate models; Selecting a training sample pair in at least one image area of the current block and the reference area, the training sample pair comprising one or more of the following sample pairs: a predicted value of a first color component and a predicted value of a second color component of a sample pixel point located in the current block, and a reconstructed value of the first color component and a reconstructed value of the second color component of a sample pixel point located in the reference area; The target prediction model is obtained by performing parameter fitting on the initial prediction model according to the training samples.
5. The video decoding method according to claim 4, characterized in that: The predicted value of the first color component and the predicted value of the second color component of the sample pixel point located in the current block are obtained according to the following method: Obtaining a prediction coding mode of the current block; When the prediction coding mode is inter-frame prediction, inter-frame prediction is performed according to a first reference block corresponding to the current block to obtain a prediction value of a first color component and a prediction value of a second color component of a sample pixel point in the current block, wherein the first reference block is a decoded image block in a reference video frame corresponding to the current video frame; When the prediction coding mode is intra block copy, intra block copy prediction is performed according to a second reference block corresponding to the current block to obtain a prediction value of a first color component and a prediction value of a second color component of a sample pixel point in the current block, wherein the second reference block is a decoded image block in the current video frame.
6. The video decoding method according to claim 4 or 5, characterized in that: Before acquiring the reconstructed value of the first color component and the reconstructed value of the second color component of the pixel point from the reference area, the method further includes: Acquire availability information of one or more sub-areas constituting the reference area; The area range of the reference area is adjusted according to the availability information of the one or more sub-areas.
7. The video decoding method according to claim 6, characterized in that: The adjusting the area range of the reference area according to the availability information of the one or more sub-areas includes: removing the sub-area in an unusable state from the reference area; When all sub-areas in the reference area are in an unavailable state, the reference area is configured to be in an unavailable state.
8. The video decoding method according to any one of claims 4 to 7, characterized in that: The selecting a training sample pair in at least one image area between the current block and the reference area comprises: Acquire a pixel sampling mode of the current block, where the pixel sampling mode includes full pixel sampling or partial pixel sampling; When the pixel sampling mode of the current block is full pixel sampling, all pixels are selected in at least one image area of the current block and the reference area as sample pixels to obtain a training sample pair consisting of a first color component and a second color component of the sample pixels; When the pixel sampling mode of the current block is partial pixel sampling, partial pixel points with specified sampling positions are selected as sample pixel points in at least one image area in the current block and the reference area to obtain a training sample pair consisting of a first color component and a second color component of the sample pixel points.
9. The video decoding method according to claim 8, characterized in that: The designated sampling position includes at least one of the following sampling positions: The position coordinates of the pixel point meet the specified sampling position of the preset coordinate value conditions; A designated sampling position is selected along a preset pixel point scanning direction; The first color component reconstruction value falls within the specified sampling position within a preset value range.
10. The video decoding method according to claim 9, characterized in that: The coordinate numerical conditions include: At least one of the horizontal position coordinate and the vertical position coordinate of the pixel point is an even number; Alternatively, at least one of the horizontal position coordinate and the vertical position coordinate of the pixel point is an odd number.
11. The video decoding method according to any one of claims 4 to 10, characterized in that: The input item of the initial prediction model also includes the first color component of one or more neighborhood pixels, where the neighborhood pixels are the pixels that are the nearest neighbors or the second nearest neighbors of the designated pixel.
12. The video decoding method according to any one of claims 4 to 11, characterized in that: The initial prediction model includes one or more combination items with independent weighting parameters, and the combination item takes the first color component of one or more pixel points as an input item.
13. The video decoding method according to claim 12, characterized in that: The initial prediction model includes at least two combination terms with different orders, where the order is the highest power of input terms in the combination terms.
14. The video decoding method according to any one of claims 1 to 13, characterized in that: The reference area includes one or more nearest neighbor areas or one or more second-nearest neighbor areas, the nearest neighbor area includes an image area with a specified image size located above or to the left of the current block, and the second-nearest neighbor area includes an image area with a specified image size located above the left, below the left, or above the right of the current block.
15. The video decoding method according to claim 14, characterized in that: The next neighboring area to the upper right of the current block has the same image size as the current block in the horizontal direction, and the next neighboring area to the upper right of the current block has a specified image size in the vertical direction, and the specified image size is greater than or equal to one; The next nearest neighboring area at the lower left of the current block has the same image size as the current block in the vertical direction, and the next nearest neighboring area at the upper right of the current block has the specified image size in the horizontal direction; The next neighboring area at the upper left of the current block has the specified image size both in the horizontal direction and in the vertical direction.
16. The video decoding method according to any one of claims 1 to 15, characterized in that: The reference area includes at least one of a full area combination, a left area combination and an upper area combination; The full region combination includes the nearest neighbor regions located on the left and above the current block and the next nearest neighbor regions located on the upper left, lower left and upper right of the current block; The left region combination includes a nearest neighbor region located on the left side of the current block and a second nearest neighbor region located on the lower left side of the current block; The upper region combination includes a nearest neighbor region located above the current block and a next nearest neighbor region located to the upper right of the current block.
17. The video decoding method according to any one of claims 1 to 16, characterized in that: The mapping process is performed on the reconstructed value of the first color component of the current block according to the mapping relationship to obtain the predicted value of the second color component of the current block, including: Acquire a color component prediction condition of the current block, where the color component prediction condition is used to indicate whether to predict another color component based on a color component of the current block; When the current block meets the color component prediction condition, the reconstructed value of the first color component of the current block is mapped according to the mapping relationship to obtain the predicted value of the second color component of the current block.
18. The video decoding method according to claim 17, characterized in that: The color component prediction condition includes at least one of the following conditions: An index corresponding to the current block has a specified index value; The reconstruction residual of the first color component of the pixel point in the current block falls within a preset value range; The image feature of the current block itself meets the preset feature condition; The image features of the reference area meet preset feature conditions.
19. The video decoding method according to claim 18, characterized in that: The reconstructed residual of the first color component of the pixel point in the current block falls within a preset value range, including: The reconstruction residual of the first color component of the pixel point in the current block is greater than zero or less than zero.
20. The video decoding method according to claim 18 or 19, characterized in that: The image features of the current block itself meet the preset feature conditions, including: The image size of the current block falls within a specified size range; Alternatively, the position of the current block falls within a specified area.
21. The video decoding method according to any one of claims 18 to 20, characterized in that: The image features of the reference area meet preset feature conditions, including: The area of the reference region is greater than a specified area threshold; Alternatively, the number of designated sampling positions in the reference area is greater than the number of model parameters of a prediction model; the prediction model is used to indicate a mapping relationship between a first color component and a second color component.
22. The video decoding method according to any one of claims 1 to 21, characterized in that: Before mapping the reconstructed value of the first color component of the current block according to the mapping relationship, the method further includes: Determining a sampling window with a specified window size according to pixel positions of the second color component of the current block; The reconstructed value of the first color component of the current block is downsampled in the sampling window to obtain a reconstructed value of the first color component that matches the pixel point position.
23. The video decoding method according to claim 22, characterized in that: The step of downsampling the reconstructed value of the first color component of the current block in the sampling window to obtain the reconstructed value of the first color component matching the pixel point position includes: Acquire a positional relationship between a reconstructed value of a first color component of the current block and a position of the pixel point in the sampling window; A weighted operation is performed on the reconstructed value of the first color component according to the positional relationship to obtain a reconstructed value of the first color component that matches the position of the pixel point.
24. The video decoding method according to claim 22 or 23, characterized in that: Before fitting the mapping relationship between the first color component of the current block and the second color component of the current block according to the current block and at least one image area in the reference area, the method further includes: Determining a sampling window with a specified window size in at least one image area between the current block and the reference area according to pixel positions of the second color component; The first color component is down-sampled in the sampling window to obtain a first color component that matches the pixel point position.
25. A video encoding method, characterized in that: include: Obtaining a reconstructed value of a first color component of a current block, where the current block is an image block to be encoded in a current video frame; Fitting a mapping relationship between a first color component of the current block and a second color component of the current block according to the current block and at least one image region in a reference region, wherein the reference region includes one or more encoded and reconstructed image regions that are the nearest or next nearest neighbor to the current block; The reconstructed value of the first color component of the current block is mapped according to the mapping relationship to obtain a predicted value of the second color component of the current block.
26. A video decoding device, characterized in that: include: A first acquisition module is configured to acquire a reconstructed value of a first color component of a current block, where the current block is an image block to be decoded in a current video frame; A first fitting module is configured to fit a mapping relationship between a first color component of the current block and a second color component of the current block according to the current block and at least one image region in a reference region, wherein the reference region includes one or more decoded image regions that are the nearest or second nearest neighbor to the current block; The first mapping module is configured to perform mapping processing on the reconstructed value of the first color component of the current block according to the mapping relationship to obtain the predicted value of the second color component of the current block.
27. A video encoding device, characterized in that: include: A second acquisition module is configured to acquire a reconstructed value of a first color component of a current block, where the current block is an image block to be encoded in a current video frame; A second fitting module is configured to fit a mapping relationship between a first color component of the current block and a second color component of the current block according to the current block and at least one image region in a reference region, wherein the reference region includes one or more encoded and reconstructed image regions that are the nearest or next nearest neighbor to the current block; The second mapping module is configured to perform mapping processing on the reconstructed value of the first color component of the current block according to the mapping relationship to obtain a predicted value of the second color component of the current block.
28. A computer readable medium, characterized in that The computer readable medium stores a computer program, and when the computer program is executed by a processor, the method according to any one of claims 1 to 25 is implemented.
29. An electronic device, characterized in that: include: processor; as well as A memory, configured to store executable instructions of the processor; The processor is configured to execute the executable instructions to implement the method according to any one of claims 1 to 25.
30. A method for processing a code stream, characterized in that: A video code stream is stored on a non-transitory computer-readable medium, wherein the video code stream is decoded based on the video decoding method according to any one of claims 1 to 24, or generated according to the video encoding method according to claim 25.
Citation Information
Patent Citations
Image component prediction method, encoder, decoder and storage medium
CN113676732A
Chroma prediction method and device, coding equipment, decoding equipment and storage medium
CN115118990A
Video coding method and device using mixed cross-component prediction
WO2023224280A1