Video coding and decoding method and device, computer readable medium and electronic equipment

By processing skip patterns based on pixel regions in video frames, the system accurately identifies and generates reconstructed data, solving the problems of complex block partitioning and high-bit description in existing technologies, and improving encoding and decoding efficiency and flexibility.

CN121967700APending Publication Date: 2026-05-01TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
TENCENT TECHNOLOGY (SHENZHEN) CO LTD
Filing Date
2024-10-31
Publication Date
2026-05-01

AI Technical Summary

Technical Problem

In existing technologies, video encoders require complex block division and a large amount of bit information to describe the outline of skipped areas when processing static regions, resulting in low encoding and decoding efficiency.

Method used

By determining the skipped region in the video frame based on a preset pixel region size, and generating reconstructed data according to the positional relationship between the skipped region and the block to be processed, the number of bits required to describe the outline of the skipped region is reduced.

Benefits of technology

It improves encoding and decoding efficiency, reduces encoding and decoding complexity, adapts to static regions of arbitrary shapes, and enhances encoding flexibility.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121967700A_ABST
    Figure CN121967700A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a video encoding and decoding method and device, a computer readable medium and electronic equipment. The video decoding method comprises the following steps: determining a plurality of pixel regions in a current frame based on a preset pixel region size, wherein each pixel region comprises at least one pixel; determining a skipped area in the current frame according to whether each pixel area is processed in a skipped mode or not; and generating reconstruction data corresponding to the current block according to the position relationship between the current block to be decoded and the skipping region. According to the technical scheme of the embodiment of the invention, the coding and decoding efficiency can be improved, and the coding and decoding complexity can be reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the fields of computer and communication technology, and more specifically, to a video encoding / decoding method, apparatus, computer-readable medium, and electronic device. Background Technology

[0002] During video encoding and decoding, some regions may remain unchanged between consecutive video frames. To optimize the encoding efficiency of such videos, a skip mode has been proposed. This mode allows the encoder to skip encoding a region when it detects that the content of that region has not changed between consecutive video frames, and directly reuse the data from the previous frame.

[0003] However, in order to accurately describe the outlines of these regions (referred to as skip regions for ease of description), the encoder needs to perform complex block division of video frames and use a lot of bit information to describe the location of skip regions, which seriously reduces the efficiency of video encoding and decoding. Summary of the Invention

[0004] The embodiments of this application provide a video encoding / decoding method, apparatus, computer-readable medium, and electronic device, which can improve encoding / decoding efficiency and reduce encoding / decoding complexity.

[0005] Other features and advantages of this application will become apparent from the following detailed description, or may be learned in part by practice of this application.

[0006] In a first aspect, embodiments of this application provide a video decoding method, comprising: determining multiple pixel regions in a current frame based on a preset pixel region size, wherein each pixel region contains at least one pixel; determining a skip region in the current frame based on whether each pixel region is processed using a skip mode; and generating reconstructed data corresponding to the current block based on the positional relationship between the current block to be decoded and the skip region.

[0007] Secondly, embodiments of this application provide a video encoding method, comprising: determining multiple pixel regions in a current frame based on a preset pixel region size, each pixel region containing at least one pixel; determining a skip region in the current frame based on whether each pixel region is processed using a skip mode; and encoding the current block according to the positional relationship between the current block to be encoded and the skip region.

[0008] Thirdly, embodiments of this application provide a video decoding apparatus, comprising: a determining unit configured to determine multiple pixel regions in a current frame based on a preset pixel region size, wherein each pixel region contains at least one pixel; a processing unit configured to determine a skipped region in the current frame based on whether each pixel region is processed using a skipped mode; and a generating unit configured to generate reconstructed data corresponding to the current block based on the positional relationship between the current block to be decoded and the skipped region.

[0009] Fourthly, embodiments of this application provide a video encoding apparatus, comprising: a determining unit configured to determine multiple pixel regions in a current frame based on a preset pixel region size, wherein each pixel region contains at least one pixel; a processing unit configured to determine a skipped region in the current frame based on whether each pixel region is processed using a skipped mode; and an encoding unit configured to encode the current block based on the positional relationship between the current block to be encoded and the skipped region.

[0010] Fifthly, embodiments of this application provide a computer-readable medium having a computer program stored thereon, which, when executed by a processor, implements the video decoding method or video encoding method as described in the above embodiments.

[0011] Sixthly, embodiments of this application provide an electronic device, including: one or more processors; and a storage device for storing one or more computer programs, which, when executed by the one or more processors, cause the electronic device to implement the video decoding method or video encoding method as described in the above embodiments.

[0012] In a seventh aspect, embodiments of this application provide a computer program product comprising a computer program stored in a computer-readable storage medium. A processor of an electronic device reads from and executes the computer program from the computer-readable storage medium, causing the electronic device to perform the video decoding or video encoding methods provided in the various alternative embodiments described above.

[0013] In some embodiments of this application, multiple pixel regions can be determined in the current frame based on a preset pixel region size. Each pixel region contains at least one pixel. Then, based on whether each pixel region is processed using a skip mode, skipped regions in the current frame are determined. Furthermore, based on the positional relationship between the current block to be decoded and the skipped regions, reconstructed data corresponding to the current block is generated. It is evident that the technical solutions of this application can accurately and efficiently identify skipped regions in the current frame, reducing the number of bits required to describe the outline of the skipped regions, thereby significantly improving encoding and decoding efficiency. Simultaneously, since the technical solutions of this application are based on pixel region processing, they can adapt to static regions of arbitrary shapes, reducing encoding and decoding complexity while improving encoding flexibility.

[0014] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and do not limit this application. Attached Figure Description

[0015] Figure 1 A schematic diagram of an exemplary system architecture to which the technical solutions of the embodiments of this application can be applied is shown;

[0016] Figure 2 This diagram illustrates the placement of the video encoding and decoding devices in a streaming system.

[0017] Figure 3 A basic flowchart of a video encoder is shown;

[0018] Figure 4 A schematic diagram of an inter-frame prediction process is shown;

[0019] Figure 5 A schematic diagram of an inter-frame prediction process is shown;

[0020] Figure 6 A schematic diagram of the shape of a skipped area is shown;

[0021] Figure 7 The diagram shows the shapes of some skipped areas;

[0022] Figure 8 A flowchart of a video decoding method according to an embodiment of this application is shown;

[0023] Figure 9 A flowchart of a video encoding method according to an embodiment of this application is shown;

[0024] Figure 10 A block diagram of a video decoding apparatus according to an embodiment of this application is shown;

[0025] Figure 11 A block diagram of a video encoding apparatus according to an embodiment of this application is shown;

[0026] Figure 12 A schematic diagram of the structure of a computer system suitable for implementing the electronic device of the present application is shown. Detailed Implementation

[0027] Exemplary embodiments will now be described in a more comprehensive manner with reference to the accompanying drawings. However, the exemplary embodiments can be implemented in various forms and should not be construed as limited to these examples; rather, these embodiments are provided so that this application will be more comprehensive and complete, and will fully convey the concept of the exemplary embodiments to those skilled in the art.

[0028] Furthermore, the features, structures, or characteristics described in this application can be combined in any suitable manner in one or more embodiments. Numerous specific details are provided in the following description to provide a full understanding of the embodiments of this application. However, those skilled in the art will recognize that when implementing the technical solutions of this application, not all the detailed features in the embodiments may be used, one or more specific details may be omitted, or other methods, elements, devices, steps, etc., may be employed.

[0029] In this application embodiment, the terms "module" or "unit" refer to a computer program or part of a computer program that has a predetermined function and works with other related parts to achieve a predetermined goal, and can be implemented wholly or partially using software, hardware (such as processing circuitry or memory), or a combination thereof. Similarly, a processor (or multiple processors or memory) can be used to implement one or more modules or units. Furthermore, each module or unit can be part of an overall module or unit that includes the functionality of that module or unit.

[0030] The block diagrams shown in the accompanying drawings are merely functional entities and do not necessarily correspond to physically independent entities. That is, these functional entities can be implemented in software, in one or more hardware modules or integrated circuits, or in different network and / or processor devices and / or microcontroller devices.

[0031] The flowcharts shown in the accompanying drawings are merely illustrative and do not necessarily include all content and operations / steps, nor do they necessarily have to be performed in the described order. For example, some operations / steps can be broken down, while others can be combined or partially combined; therefore, the actual execution order may change depending on the specific circumstances.

[0032] It should be noted that "multiple" in this article refers to two or more. "And / or" describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A alone, A and B simultaneously, or B alone. The character " / " generally indicates that the preceding and following related objects have an "or" relationship.

[0033] Figure 1 A schematic diagram of an exemplary system architecture to which the technical solutions of the embodiments of this application can be applied is shown.

[0034] like Figure 1 As shown, system architecture 100 includes multiple terminal devices that can communicate with each other via, for example, a network 150. For instance, system architecture 100 may include a first terminal device 110 and a second terminal device 120 interconnected via network 150. Figure 1 In one embodiment, the first terminal device 110 and the second terminal device 120 perform unidirectional data transmission.

[0035] For example, the first terminal device 110 can encode video data (e.g., a video image stream captured by the terminal device 110) to transmit it to the second terminal device 120 via the network 150. The encoded video data is transmitted in the form of one or more encoded video streams. The second terminal device 120 can receive the encoded video data from the network 150, decode the encoded video data to recover the video data, and display video images based on the recovered video data.

[0036] In one embodiment of this application, system architecture 100 may include a third terminal device 130 and a fourth terminal device 140 that perform bidirectional transmission of encoded video data, such as during a video conference. For bidirectional data transmission, each of the third terminal device 130 and the fourth terminal device 140 may encode video data (e.g., a video image stream captured by the terminal device) for transmission over network 150 to the other terminal device. Each of the third terminal device 130 and the fourth terminal device 140 may also receive encoded video data transmitted by the other terminal device, decode the encoded video data to recover the video data, and display the video images on an accessible display device based on the recovered video data.

[0037] exist Figure 1 In the embodiments shown, the first terminal device 110, the second terminal device 120, the third terminal device 130 and the fourth terminal device 140 may be servers or terminals, but the principles disclosed in this application are not limited to these.

[0038] Servers can be standalone physical servers, server clusters or distributed systems composed of multiple physical servers, or cloud servers providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms. Terminals can be smartphones, tablets, laptops, desktop computers, smart speakers, smart voice interaction devices, smartwatches, smart home appliances, in-vehicle terminals, aircraft, etc., but are not limited to these.

[0039] Figure 1 The network 150 shown represents any number of networks, including, for example, wired and / or wireless communication networks, that transmit encoded video data between the first terminal device 110, the second terminal device 120, the third terminal device 130, and the fourth terminal device 140. The communication network 150 may exchange data in circuit-switched and / or packet-switched channels. This network may include telecommunications networks, local area networks (LANs), wide area networks (WANs), and / or the Internet. For the purposes of this application, unless explained below, the architecture and topology of network 150 may be irrelevant to the operation of the disclosure herein.

[0040] In one embodiment of this application, Figure 2 The illustration shows the placement of video encoding and decoding devices in a streaming environment. The subject matter disclosed in this application is equally applicable to other video-enabled applications, including, for example, video conferencing, digital television (TV), and storing compressed video on digital media including CDs, DVDs, memory sticks, etc.

[0041] The streaming system may include an acquisition subsystem 213, which may include a video source 201 such as a digital camera, which creates an uncompressed video image stream 202. In an embodiment, the video image stream 202 includes samples captured by a digital camera. The video image stream 202 is depicted as a thick line to emphasize the high data volume of the video image stream compared to encoded video data 204 (or encoded video bitstream 204). The video image stream 202 may be processed by an electronic device 220, which includes a video encoding device 203 coupled to the video source 201. The video encoding device 203 may include hardware, software, or a combination of hardware and software to implement or enforce aspects of the disclosed subject matter as described in more detail below. The encoded video data 204 (or encoded video bitstream 204) is depicted as a thin line to emphasize the lower data volume of the encoded video data 204 (or encoded video bitstream 204), which may be stored on a streaming server 205 for future use. One or more streaming client subsystems, such as Figure 2 Client subsystems 206 and 208 can access streaming server 205 to retrieve copies 207 and 209 of encoded video data 204. Client subsystem 206 may include, for example, a video decoding device 210 in electronic device 230. Video decoding device 210 decodes the incoming copy 207 of the encoded video data and produces an output video picture stream 211 that can be displayed on display 212 (e.g., a screen) or another presentation device. In some streaming systems, the encoded video data 204, video data 207, and video data 209 (e.g., video stream) may be encoded according to certain video encoding / compression standards.

[0042] It should be noted that electronic devices 220 and 230 may include other components not shown in the figures. For example, electronic device 220 may include a video decoding device, and electronic device 230 may also include a video encoding device.

[0043] In one embodiment of this application, taking High Efficiency Video Coding (HEVC) and Versatile Video Coding (VVC) from international video coding standards, as well as the Chinese national video coding standard AVS, as examples, after an input video frame image, the video frame image is divided into several non-overlapping processing units according to a block size. Each processing unit performs a similar compression operation. This processing unit is called a Coding Tree Unit (CTU) or Largest Coding Unit (LCU). The CTU can be further subdivided into more refined units to obtain one or more basic Coding Units (CUs). The CU is the most basic element in a coding process.

[0044] In another embodiment, this processing unit can also be called a tile, which is a rectangular area of ​​a multimedia data frame that can be independently decoded and encoded. In the Alliance for Open Media Video 1 (AV1) standard, the tile can be further subdivided into one or more superblocks (SBs). The SB is the starting point for block partitioning and can be further divided into multiple subblocks. The superblocks are then further subdivided into one or more blocks. Each block is the most basic element in a coding process. Optionally, an SB can contain several blocks (Bs).

[0045] The above method of dividing video frame images can be called a block partition structure. The following introduces some concepts in the encoding process:

[0046] Predictive coding includes intra-frame prediction and inter-frame prediction. The original video signal is predicted from a selected reconstructed video signal to obtain a residual video signal. The encoder needs to decide which predictive coding mode to choose for the current coding unit (or coding block) and inform the decoder. Intra-frame prediction refers to the predicted signal coming from a region within the same image that has already been encoded and reconstructed; inter-frame prediction refers to the predicted signal coming from another encoded image (called a reference image) that is different from the current image.

[0047] Transform and Quantization: After the residual video signal undergoes transformation operations such as Discrete Fourier Transform (DFT) and Discrete Cosine Transform (DCT), the signal is transformed into the transform domain, and these are called transform coefficients. The transform coefficients are then subjected to lossy quantization, losing some information to make the quantized signal more suitable for compression. In some video coding standards, there may be more than one transform method to choose from. Therefore, the encoder needs to select one of the transform methods for the current coding unit (or coding block) and inform the decoder. The fineness of quantization is usually determined by the quantization parameter (QP). A larger QP value means that coefficients with a wider range of values ​​will be quantized into the same output, which usually leads to greater distortion and a lower bit rate. Conversely, a smaller QP value means that coefficients with a smaller range of values ​​will be quantized into the same output, which usually leads to less distortion and a higher bit rate.

[0048] Entropy coding, or statistical coding, involves statistically compressing the quantized transform-domain signal based on the frequency of each value, ultimately outputting a binary (0 or 1) compressed bitstream. Simultaneously, other information generated during encoding, such as the selected coding mode and motion vector data, also requires entropy coding to reduce the bit rate. Statistical coding is a lossless coding method that effectively reduces the bit rate required to represent the same signal. Common statistical coding methods include Variable Length Coding (VLC) and Content-Adaptive Binary Arithmetic Coding (CABAC).

[0049] Context-Based Binary Arithmetic Coding (CABAC) primarily involves three steps: binarization, context modeling, and binary arithmetic coding. After binarizing the input syntax elements, the binary data can be encoded using either a regular coding mode or a bypass coding mode. The bypass coding mode eliminates the need to assign a specific probability model to each binary bit; the input binary bit bin value is directly encoded using a simple bypass encoder, thus accelerating the overall encoding and decoding speed. Generally, different syntax elements are not completely independent, and even identical syntax elements possess a certain degree of memory. Therefore, according to conditional entropy theory, using other encoded syntax elements for conditional coding can further improve coding performance compared to independent coding or memoryless coding. This encoded symbol information used as conditions is called the context. In the regular coding mode, the binary bits of the syntax elements sequentially enter the context modeler. The encoder assigns an appropriate probability model to each input binary bit based on the values ​​of previously encoded syntax elements or binary bits; this process is called context modeling. The context model corresponding to a grammatical element can be located using the context index increment (ctxIdxInc) and the context index start (ctxIdxStart). After the bin value and the assigned probability model are fed into the binary arithmetic encoder for encoding, the context model needs to be updated based on the bin value, which is the adaptive process in encoding.

[0050] Loop Filtering: The transformed and quantized signal undergoes inverse quantization, inverse transform, and prediction compensation to obtain a reconstructed image. Due to the effects of quantization, the reconstructed image differs from the original image in some aspects, resulting in distortion. Therefore, filtering operations can be performed on the reconstructed image, such as deblocking filters (DB), sample adaptive offset (SAO), or adaptive loop filters (ALF), to effectively reduce the distortion caused by quantization. Since these filtered reconstructed images will serve as a reference for subsequent coded images to predict future image signals, the aforementioned filtering operations are also called loop filtering, i.e., filtering operations within the coding loop.

[0051] In one embodiment of this application, Figure 3 A basic flowchart of a video encoder is shown, illustrating the process using intra-frame prediction as an example. The original image signal s...k [x,y] and the predicted image signal Perform the difference operation to obtain the residual signal u. k [x,y], residual signal u k After transformation and quantization, [x,y] is obtained as quantization coefficients. These coefficients are then used to obtain the encoded bitstream through entropy encoding, and to obtain the reconstructed residual signal u' through inverse quantization and inverse transform. k [x,y], predict image signal With the reconstructed residual signal u' k [x,y] superimposed to generate image signals Image signal On one hand, the signal is input to the intra-frame mode decision module and the intra-frame prediction module for intra-frame prediction processing; on the other hand, the reconstructed image signal s' is output through loop filtering. k [x,y], reconstruct the image signal s' k [x,y] can be used as a reference image for the next frame for motion estimation and motion compensation prediction. Then, based on the result s' of the motion compensation prediction... r [x+m x ,y+m y ] and intra-frame prediction results Obtain the predicted image signal for the next frame. And continue repeating the above process until the coding is complete.

[0052] Based on the above encoding process, at the decoding end, for each encoding unit (or encoding block), after acquiring the compressed bitstream (i.e., bitstream), entropy decoding is performed to obtain various mode information and quantization coefficients. Then, the quantization coefficients undergo inverse quantization and inverse transform processing to obtain the residual signal. On the other hand, based on the known encoding mode information, the prediction signal corresponding to the encoding unit (or encoding block) can be obtained. Then, the residual signal and the prediction signal are added together to obtain the reconstructed signal. The reconstructed signal then undergoes loop filtering and other operations to generate the final output signal.

[0053] In the field of coding technology, inter-frame prediction is a commonly used predictive coding technique, such as... Figure 4 As shown, inter-frame prediction utilizes the temporal correlation of video, using pixels from neighboring encoded images to predict pixels in the current image, effectively removing temporal redundancy and saving bits of encoded residual data. Here, P represents the current frame, Pr represents the reference frame, B represents the current coded block, and Br represents the reference block of B. The coordinates of B' in the reference frame are the same as the coordinates of B in the current frame, and the coordinates of Br are (x...). r ,y rThe coordinates of B' are (x, y). The displacement between the current coded block and its reference block is called the motion vector (MV), where MV = (x, y). r -x,y r -y). In other words, inter-frame prediction refers to the process of searching for a reference block in neighboring encoded images (i.e., reference frames) based on the current block to be encoded in the current frame, with the aim of removing temporal redundancy in the video signal. For example... Figure 5 As shown, the current block to be encoded in the current frame is searched within a certain range in the reference frame (i.e., the search area formed by the search box) according to the block matching criteria to obtain the best matching block. Optionally, commonly used block matching criteria in video coding include: minimum mean square error (MSE), sum of absolute differences (SAD), and other matching criteria.

[0054] As described above, the predicted pixels in the inter-frame prediction mode are obtained by searching within the reconstructed reference frame or in its sub-pixel interpolated image. For the chroma component, it is generally assumed that it has a similar motion vector or block vector to the luma component, so the motion vector or block vector of the chroma component is simply derived from the motion vector or block vector of the luma component. Since this matching search process considers not only the matching degree of the image but also the encoding cost of the motion vector, it is not necessarily optimal in terms of matching degree alone; that is, there is still room for improvement in prediction accuracy. If the motion vector of a coded block can be completely derived from the predefined predicted motion vector (e.g., motion vector prediction of spatially adjacent blocks), without needing to additionally identify the residual value of the motion vector prediction, and the residual after prediction of the coded block is also all 0, then the skip mode can be used to encode the coded block. Specifically, for the current block using skip mode with a motion vector of 0, reconstruction can be completed by copying pixels at the same position in the reference frame.

[0055] In practical applications, when the background or subtitles of a video content remain largely unchanged between consecutive video frames, the skip mode can be used. Optionally, the area within the video content processed using skip mode can be called the skip region. The skip region within a video frame can be of any shape, and its shape may change across different video frames. For example... Figure 6As shown, the skipped area in the current frame can be rectangular in shape and includes parts of pixel regions C31, C32, and C33 in the current frame. During reconstruction, pixels at the same position in the reference frame can be copied as reconstructed pixels. Figure 7 As shown, the shape of the skipped area can be arbitrary, such as the rectangle shown in shape 1 and shape 2, or the trapezoid shown in shape 3, etc.

[0056] In related technologies, to accurately describe the outline of skipped regions, encoders need to perform complex block division on video frames. This typically involves multiple, progressively finer block division processes until the boundaries of skipped regions can be delineated relatively accurately. This fine division not only increases coding complexity but also necessitates the use of a large amount of bit information in the video bitstream to describe the position and shape of these divided blocks. Furthermore, for each block obtained through division, the encoder requires additional bits to identify whether the block uses skipped mode. This additional identification information further increases the overhead of the encoded data, thus offsetting to some extent the coding efficiency improvement brought by skipped mode.

[0057] Based on the aforementioned technical problems, the technical solution of this application proposes a novel video encoding and decoding scheme that can accurately and efficiently identify skipped regions in the current frame, reducing the number of bits required to describe the outline of the skipped regions, thereby significantly improving encoding and decoding efficiency. Furthermore, since the technical solution of this application is based on pixel regions, it can adapt to static regions of arbitrary shapes, reducing encoding and decoding complexity while improving encoding flexibility.

[0058] The implementation details of the technical solutions in the embodiments of this application are described in detail below:

[0059] Figure 8 A flowchart of a video decoding method according to an embodiment of this application is shown. This video decoding method can be executed by a device with computing processing capabilities, such as a terminal device or a server. (Refer to...) Figure 8 As shown, this video decoding method includes at least S810 to S840, which are described in detail below:

[0060] In S810, multiple pixel regions are determined in the current frame based on a preset pixel region size, and each pixel region contains at least one pixel.

[0061] In some optional embodiments, the preset pixel area size can be pre-negotiated between the video encoding device and the video decoding device, such as 1 pixel size, 2×2 pixel size, 3×3 pixel size, etc.

[0062] In some optional embodiments, the preset pixel region size can be determined by the video encoding device and indicated to the video decoding device by adding a flag bit to the video bitstream. Optionally, the flag bit may include one or more of the following flag bits: flag bits contained in the sequence header, flag bits contained in the image header, flag bits contained in the strip header, flag bits contained in the coding tree unit (CTU) header, and flag bits contained in the coding block. In this case, the video decoding device can decode the video bitstream to obtain the flag bit used to indicate the pixel region size.

[0063] It should be noted that the flags contained in the sequence header can indicate the size of the pixel region set for the entire image sequence; the flags contained in the image header can indicate the size of the pixel region set for the entire image; the flags contained in the strip header can indicate the size of the pixel region set for the entire strip; the flags contained in the coding tree unit (CTU) header can indicate the size of the pixel region set for the entire CTU; and the flags contained in the coding block can indicate the size of the pixel region set for the current coding block.

[0064] In S820, the skipped region in the current frame is determined based on whether each pixel region is processed using the skip mode.

[0065] In some optional embodiments, whether a pixel region should be processed in skip mode can be determined based on the pixels it contains and whether each pixel in the current frame is processed in skip mode. For example, a pixel region in which all its pixels are processed in skip mode can be defined as a pixel region processed in skip mode. That is, if all pixels in a pixel region are processed in skip mode, then it can be determined that the pixel region is processed in skip mode. Optionally, a pixel region can also be defined as a pixel region in which most pixels are processed in skip mode. For example, a proportional threshold can be set; if more than the proportional threshold of pixels are processed in skip mode, then it can be determined that the pixel region is processed in skip mode.

[0066] In some alternative embodiments, when determining whether each pixel should be processed in skip mode, the video decoding device can directly decode the corresponding flag bit from the video bitstream to determine whether each pixel should be processed in skip mode based on the flag bit. In this case, the video encoding device needs to indicate whether each pixel should be processed in skip mode through the flag bit in the video bitstream.

[0067] In some optional embodiments, the video decoding device can also determine whether each pixel should be processed using the skip mode by deduction. For example, if the pixel difference at the target pixel location is less than or equal to a pixel threshold for a preset number of consecutive image frames adjacent to and preceding the current frame, then the pixel at the target pixel location can be determined to be processed using the skip mode. In this embodiment, if the pixel difference at a certain pixel location for N consecutive frames (N≥2) preceding the current frame is less than or equal to a set pixel threshold (e.g., 0 or other values), then the pixel at that pixel location can be determined to be processed using the skip mode.

[0068] For example, if the motion vector at the target pixel location is 0 for a predetermined number of consecutive image frames adjacent to and preceding the current frame, and the residual value at the target pixel location is less than or equal to a residual threshold, then the pixel at the target pixel location is determined to be processed using a skip mode. In this embodiment, if the motion vector at a certain pixel location is 0 for N consecutive frames (N≥2) preceding the current frame, and the residual value at that pixel location is less than or equal to a set residual threshold (e.g., 0 or another value), then the pixel at that pixel location can be determined to be processed using a skip mode.

[0069] Continue to refer to Figure 8 As shown, in S830, the reconstruction data corresponding to the current block is generated based on the positional relationship between the current block to be decoded and the skipped area.

[0070] In some optional embodiments, the video contains a sequence of video image frames, which includes a series of images. Each image can be further divided into slices, and each slice can be further divided into a series of LCUs (or CTUs). Each LCU contains several CUs. Video image frames are encoded in blocks. In some newer video coding standards, such as H.264, there are macroblocks (MBs), which can be further divided into multiple prediction blocks for predictive coding. In the HEVC standard, basic concepts such as coding units (CUs), prediction units (PUs), and transform units (TUs) are used to functionally divide various block units, and a novel tree-based structure is used for description. For example, a CU can be divided into smaller CUs according to a quadtree, and these smaller CUs can be further divided to form a quadtree structure. In the embodiments of this application, the current block can be a CU, or a block smaller than a CU, such as a smaller block obtained by dividing a CU.

[0071] In some optional embodiments, if the current block is within a skipped region, the pixel value at the corresponding position of the current block in the previous frame or the reference frame can be used as the reconstructed data for the current block. This embodiment allows the reconstructed data for a current block within a skipped region to be directly obtained by copying the pixel value at the corresponding position (e.g., the same position) in the previous or reference frame. This not only saves the number of bits used to transmit the current block data during encoding but also reduces the computational burden on the encoder and decoder, improving overall encoding efficiency and decoding speed.

[0072] In some optional embodiments, if some pixels in the current block are within the skipped region, then after decoding the current block to obtain the corresponding decoded data, the pixel values ​​at the corresponding positions in the decoded data from the previous frame or the reference frame of the current frame can be used to replace the data at the corresponding positions to obtain the reconstructed data for the current block. This embodiment's technical solution allows for both preserving the decoding results of non-skipped pixels in the current block and optimizing the skipped pixels, thereby reducing the amount of data required for encoding while maintaining high image quality, thus improving overall encoding efficiency.

[0073] In some optional embodiments, the video encoding device may add a flag bit to the video bitstream to indicate whether the current frame uses a decoding scheme based on skipped regions; that is, a flag bit is added to the video bitstream to indicate whether the current frame uses the technical solution of the aforementioned embodiments. In this case, the video decoding device can determine whether the current frame uses the technical solution of the aforementioned embodiments by decoding the flag bit in the video bitstream.

[0074] In some optional embodiments, the video encoding device may add flag bits to the video bitstream to indicate whether each image block in the current frame adopts a decoding scheme based on skipped regions. That is, flag bits are added to the video bitstream to indicate whether each image block in the current frame adopts the technical solution of the aforementioned embodiments. In this case, the video decoding device can determine whether each image block in the current frame adopts the technical solution of the aforementioned embodiments by decoding the flag bits in the video bitstream. For example, some image blocks in the current frame may adopt the technical solution of the aforementioned embodiments, while other image blocks may not.

[0075] In some optional embodiments, the video encoding device may add flag bits to the video bitstream to indicate whether each stripe in the current frame uses a decoding scheme based on skipped regions. That is, flag bits are added to the video bitstream to indicate whether each stripe in the current frame uses the technical solution of the aforementioned embodiments. In this case, the video decoding device can determine whether each stripe in the current frame uses the technical solution of the aforementioned embodiments by decoding the flag bits in the video bitstream. For example, some stripes in the current frame may use the technical solution of the aforementioned embodiments, while other stripes may not.

[0076] In some alternative embodiments, the video encoding device may also employ multiple flags to indicate whether a skip-region-based decoding scheme is used. For example, multiple of the following flags may be added: a first flag indicating whether the current frame uses skip-region-based decoding, a second flag indicating whether each image block in the current frame uses skip-region-based decoding, and a third flag indicating whether each strip in the current frame uses skip-region-based decoding.

[0077] Optionally, if the first flag bit and the second flag bit mentioned above are used to indicate whether to use a scheme for decoding based on skipped regions, then only if the first flag bit indicates that the current frame uses a scheme for decoding based on skipped regions and the second flag bit indicates that a specified image block in the current frame uses a scheme for decoding based on skipped regions, then it means that the specified image block needs to use a scheme for decoding based on skipped regions.

[0078] Figure 8 This description focuses on the technical solutions of the embodiments of this application from the perspective of video decoding. The following is a combination of... Figure 9 The technical solutions of the embodiments of this application will be described again from the perspective of video encoding.

[0079] Figure 9 A flowchart of a video encoding method according to an embodiment of this application is shown. This video encoding method can be executed by a device with computing processing capabilities, such as a terminal device or a server. (Refer to...) Figure 9 As shown, this video encoding method includes at least S910 to S930, which are detailed below:

[0080] In S910, multiple pixel regions are determined in the current frame based on a preset pixel region size, and each pixel region contains at least one pixel.

[0081] Optionally, the specific implementation details of S910 can be referred to the implementation details of S810 in the aforementioned embodiments, and will not be repeated here.

[0082] In S920, the skipped region in the current frame is determined based on whether each pixel region is processed using the skip mode.

[0083] Optionally, the specific implementation details of S920 can refer to the implementation details of S820 in the aforementioned embodiments. It is worth noting that when determining whether each pixel should be processed using the skip mode, the video encoding device can also compare the current frame with the previous M frames (M≥1) to determine this.

[0084] For example, if the pixel difference at the target pixel location is less than or equal to the pixel threshold for a predetermined number of consecutive image frames, including the current frame, then the pixel at the target pixel location can be processed using the skip mode.

[0085] For example, if the motion vector at the target pixel location is 0 for a consecutive preset number of image frames, including the current frame, and the residual value at the target pixel location is less than or equal to the residual threshold, then the pixel at the target pixel location is determined to be processed using the skip mode.

[0086] In S930, the current block is encoded based on the positional relationship between the current block to be encoded and the skipped region.

[0087] In some alternative embodiments, if the current block is within the skip region, the video encoding device can skip the encoding process of the current block, such as skipping the transformation and quantization process of the current block. Correspondingly, the video decoding device can skip the inverse quantization and inverse transformation processing of the current block and directly copy the pixel value at the corresponding position (such as the same position) in the previous frame image or reference frame image as the reconstruction data.

[0088] In some optional embodiments, if some pixels in the current block are located within a skipped region, the video encoding device can exclude pixel distortion within the skipped region when determining the encoding mode and parameters of the current block. This embodiment's technical solution, by excluding pixel distortion within the skipped region, allows the encoding process to focus on optimizing pixels in non-skipped regions that have a greater impact on video quality. This effectively reduces unnecessary computational overhead and bitrate consumption while ensuring video quality, thus improving video encoding and decoding efficiency.

[0089] In some optional embodiments, if some pixels in the current block are located within a skipped region, the video encoding device can fill in some pixels using pixels adjacent to the skipped region in the current block, and determine the encoding mode and encoding parameters of the current block based on the filled pixels. This embodiment's technical solution, by utilizing pixels adjacent to the skipped region for filling, not only maintains the continuity and smoothness of the image at the edges of the skipped region, but also effectively avoids the adverse effects of missing pixel information caused by the skipped region on the selection of encoding mode and parameters.

[0090] It should be noted that the processing of video encoding equipment is similar to other processing of video decoding equipment. For details, please refer to the aforementioned processing of video decoding equipment, which will not be repeated here.

[0091] In summary, the technical solution of this application embodiment can accurately and efficiently identify skipped regions in the current frame, reducing the number of bits required to describe the outline of the skipped region, thereby significantly improving encoding and decoding efficiency. Furthermore, since the technical solution of this application embodiment is based on pixel region processing, it can adapt to static regions of arbitrary shapes, reducing encoding and decoding complexity while improving encoding flexibility.

[0092] In a specific application scenario of this application, when determining the shape of the skip region for the current image frame to be processed, the size of the pixel block (i.e., the pixel region in the aforementioned embodiments) used to determine the skip region can be defined first. For example, the size of the pixel block can be 1 pixel, 2×2 pixels, etc. For the current frame, a binary map can be constructed in units of pixel blocks, describing whether each pixel block is processed using the skip mode. The pixel blocks processed using the skip mode constitute the skip region. Optionally, the size of the pixel block can be defined or updated in the image header, sequence header, or other high-level syntax.

[0093] In some alternative embodiments, possible methods for deriving the value of each pixel block in a binary map include:

[0094] Method 1: Calculate the difference between pixel pairs at the same position in consecutive frames. If the difference at a certain position is 0 (or less than a certain threshold) for M consecutive frames (M≥2), then that position is considered a skipped pixel (i.e., a pixel processed using the skipped mode). If all pixels within a pixel block are processed using the skipped mode, then that pixel block can be defined as a skipped pixel block. On the binary map, skipped pixel blocks are marked as 1, and other pixel blocks are marked as 0.

[0095] It should be noted that, for video encoding devices, when implementing method 1, consecutive M frames can include the current frame; for video decoding devices, when implementing method 1, consecutive M frames can be M frames adjacent to the current frame and located before the current frame.

[0096] Method 2: For a given pixel block, if the pixel motion vector at the corresponding position is 0 for M consecutive frames (M≥2), and the residual pixel value at the corresponding position is 0 (or less than a certain threshold), then that position is considered a skipped pixel. If all pixels within a pixel block are processed using the skipped mode, then that pixel block can be defined as a skipped pixel block. On the binary map, skipped pixel blocks are marked as 1, and other pixel blocks are marked as 0.

[0097] It should be noted that, for video encoding devices, when implementing method 2, consecutive M frames can include the current frame; for video decoding devices, when implementing method 2, consecutive M frames can be M frames that are adjacent to the current frame and located before the current frame.

[0098] In some alternative embodiments, switches can be set at the frame level or the region level to determine whether to use the skip region method. For example, if a frame-level switch is used, all skipped pixel blocks within the current frame do not need to be encoded or decoded; the pixel values ​​at the corresponding positions in the previous frame or reference frame are directly copied.

[0099] For example, if using a region-level switch, the current frame could be divided into fixed sub-images or strip sizes, for instance. Figure 6 The current frame is divided into 12 sub-images, and each sub-image can be individually configured with a switch to determine whether to use the skip region method. Figure 6 In the example shown, sub-images C31, C32, and C33 can be considered for using the skip region method. For each sub-image that chooses to use the skip region method, the skipped pixel blocks within it do not need to be encoded or decoded; the pixel values ​​at the corresponding positions in the previous frame or reference frame are directly copied.

[0100] In some alternative embodiments, for a video encoding device, if the current block is entirely within the skip area, the encoding process of the current block can be skipped (e.g., skipping the transform and quantization process), and the transmission of syntax elements related to the current block in the bitstream can be ignored.

[0101] Optionally, in an exemplary scheme, if some pixels in the current block are within the skip region, the encoder can exclude pixels in the current block that belong to the skip region by taking into account pixel distortion when evaluating various encoding parameters and mode selections of the current block.

[0102] In another exemplary scheme, if some pixels in the current block are within the skip region, the encoder can extend the pixels in the current block that are adjacent to the skip region and belong to the boundary of the non-skip region to the position of the skip pixels in the current block by copying them when evaluating various encoding parameters and mode selections of the current block, thereby maintaining the pixel continuity during the non-skip region encoding process.

[0103] In some optional embodiments, for the video decoding device, if the current block is entirely within the skip area, the decoding process of the current block can be skipped, and the pixel values ​​at the corresponding positions in the previous frame or reference frame can be directly used as the reconstruction data of the current block.

[0104] Optionally, if some pixels in the current block are within the skipped area, the decoding block can be decoded normally. After decoding is complete, the pixels in the skipped area of ​​the decoding block can be replaced by pixels at the corresponding positions in the previous frame or reference frame.

[0105] The technical solutions in the above embodiments can reduce the number of bits required for block partitioning and identifying skip patterns while maintaining accurate description of the skip region outline, effectively improving the efficiency of predictive coding.

[0106] The following describes an apparatus embodiment of this application, which can be used to perform the methods described in the above embodiments of this application. For details not disclosed in the apparatus embodiments of this application, please refer to the method embodiments described above.

[0107] Figure 10 A block diagram of a video decoding apparatus according to an embodiment of the present application is shown. The video decoding apparatus can be installed in a device with computing processing capabilities, such as a terminal device or a server.

[0108] Reference Figure 10 As shown, a video decoding apparatus 1000 according to an embodiment of this application includes: a determining unit 1002, a processing unit 1004, and a generating unit 1006.

[0109] The determining unit 1002 is configured to determine multiple pixel regions in the current frame based on a preset pixel region size, and each pixel region contains at least one pixel; the processing unit 1004 is configured to determine the skipped region in the current frame based on whether each pixel region is processed using a skipped mode; and the generating unit 1006 is configured to generate the reconstructed data corresponding to the current block based on the positional relationship between the current block to be decoded and the skipped region.

[0110] In some embodiments of this application, based on the foregoing scheme, the processing unit 1004 is further configured to: determine whether each pixel region is processed in a skip mode based on the pixels contained in each pixel region and whether each pixel in the current frame is processed in a skip mode.

[0111] In some embodiments of this application, based on the foregoing scheme, the processing unit 1004 is further configured to: if the pixel difference at the target pixel position is less than or equal to a pixel threshold of a preset number of consecutive image frames adjacent to the current frame and located before the current frame, then determine that the pixel at the target pixel position is processed in a skip mode.

[0112] In some embodiments of this application, based on the foregoing scheme, the processing unit 1004 is further configured to: if the motion vector of a preset number of consecutive image frames adjacent to the current frame and located before the current frame is 0 at the target pixel position, and the residual value at the target pixel position is less than or equal to the residual threshold, then determine that the pixel at the target pixel position is processed using a skip mode.

[0113] In some embodiments of this application, based on the foregoing scheme, the processing unit 1004 is further configured to: determine whether each pixel is processed using a skip mode based on the flag bit obtained from the video bitstream.

[0114] In some embodiments of this application, based on the foregoing scheme, the processing unit 1004 is configured to: determine the pixel region in which all included pixels are processed using the skip mode as the pixel region to be processed using the skip mode.

[0115] In some embodiments of this application, based on the foregoing scheme, the generation unit 1006 is configured to: if the current block is within the skipped area, then the pixel value at the position corresponding to the current block in the previous frame image of the current frame or the reference frame image of the current frame is used as the reconstruction data corresponding to the current block.

[0116] In some embodiments of this application, based on the foregoing scheme, the generation unit 1006 is configured as follows: if some pixels in the current block are within the skipped area, after decoding the current block to obtain the decoded data corresponding to the current block, the pixel values ​​at the positions corresponding to the some pixels in the previous frame image of the current frame or the reference frame image of the current frame are used to replace the data at the corresponding positions in the decoded data to obtain the reconstructed data corresponding to the current block.

[0117] In some embodiments of this application, based on the foregoing scheme, the determining unit 1002 is further configured to: decode the video bitstream to obtain a flag bit used to indicate the size of the pixel region; wherein the flag bit includes one or more of the following flag bits: flag bits contained in the sequence header, flag bits contained in the image header, flag bits contained in the strip header, flag bits contained in the coding tree unit (CTU) header, and flag bits contained in the coding block.

[0118] In some embodiments of this application, based on the foregoing scheme, the determining unit 1002 is further configured to perform at least one of the following steps:

[0119] The video stream is decoded to obtain a flag indicating whether the current frame is decoded based on the skipped region;

[0120] The video stream is decoded to obtain a flag indicating whether each image block in the current frame is decoded based on the skipped region;

[0121] The video stream is decoded to obtain a flag indicating whether each strip in the current frame is decoded based on the skipped region.

[0122] Figure 11 A block diagram of a video encoding apparatus according to an embodiment of the present application is shown. The video encoding apparatus can be installed in a device with computing processing capabilities, such as a terminal device or a server.

[0123] Reference Figure 11 As shown, a video encoding apparatus 1100 according to an embodiment of this application includes: a determining unit 1102, a processing unit 1104, and an encoding unit 1106.

[0124] The determining unit 1102 is configured to determine multiple pixel regions in the current frame based on a preset pixel region size, and each pixel region contains at least one pixel; the processing unit 1104 is configured to determine the skipped region in the current frame based on whether each pixel region is processed using a skipped mode; and the encoding unit 1106 is configured to encode the current block based on the positional relationship between the current block to be encoded and the skipped region.

[0125] In some embodiments of this application, based on the foregoing scheme, the encoding unit 1106 is configured to: skip the encoding process of the current block if the current block is within the skipped area.

[0126] In some embodiments of this application, based on the foregoing scheme, the encoding unit 1106 is configured to: if some pixels in the current block are within the skipped area, then when determining the encoding mode and encoding parameters of the current block, exclude pixel distortion in the current block that is within the skipped area.

[0127] In some embodiments of this application, based on the foregoing scheme, the encoding unit 1106 is configured to: if some pixels in the current block are within the skipped area, then use pixels in the current block adjacent to the skipped area to fill the some pixels, and determine the encoding mode and encoding parameters of the current block through the filled pixels.

[0128] Figure 12 A schematic diagram of a computer system suitable for implementing an electronic device according to the embodiments of this application is shown. The electronic device may be a video encoding device or a video decoding device as described in the foregoing embodiments.

[0129] It should be noted that, Figure 12 The computer system 1200 of the electronic device shown is merely an example and should not impose any limitation on the functionality and scope of use of the embodiments of this application.

[0130] like Figure 12 As shown, the computer system 1200 may include a Central Processing Unit (CPU) 1201, which can perform various appropriate actions and processes based on a program stored in Read-Only Memory (ROM) 1202 or a program loaded from storage portion 1208 into Random Access Memory (RAM) 1203, such as performing the methods described in the above embodiments. Various programs and data required for system operation are also stored in RAM 1203. The CPU 1201, ROM 1202, and RAM 1203 are interconnected via bus 1204. An input / output (I / O) interface 1205 is also connected to bus 1204.

[0131] The following components can be connected to I / O interface 1205: an input section 1206 including a keyboard, mouse, etc.; an output section 1207 including a cathode ray tube (CRT), liquid crystal display (LCD), etc., and speakers, etc.; a storage section 1208 including a hard disk, etc.; and a communication section 1209 including a network interface card such as a LAN (Local Area Network) card, modem, etc. The communication section 1209 performs communication processing via a network such as the Internet. A drive 1210 is also connected to I / O interface 1205 as needed. Removable media 1211, such as a disk, optical disk, magneto-optical disk, semiconductor memory, etc., are installed on drive 1210 as needed so that computer programs read from them can be installed into storage section 1208 as needed.

[0132] Specifically, according to embodiments of this application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this application include a computer program product comprising a computer program carried on a computer-readable medium for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via communication section 1209, and / or installed from removable medium 1211. When the computer program is executed by central processing unit (CPU) 1201, it performs various functions defined in the system of this application.

[0133] It should be noted that the computer-readable medium shown in the embodiments of this application can be a computer-readable signal medium or a computer-readable storage medium, or any combination of the two. A computer-readable storage medium can be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM), flash memory, optical fiber, portable compact disc read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this application, a computer-readable storage medium can be any tangible medium containing or storing a computer program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In this application, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, carrying a computer-readable computer program. The transmitted data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. The computer-readable signal medium can also be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The computer program contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to wireless, wired, etc., or any suitable combination thereof.

[0134] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this application. Each block in a flowchart or block diagram may represent a module, segment, or portion of code, which contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in a block diagram or flowchart, and combinations of blocks in a block diagram or flowchart, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and a computer program.

[0135] The units described in the embodiments of this application can be implemented in software or hardware, and the described units can also be located in a processor. The names of these units do not necessarily limit the specific unit itself.

[0136] In another aspect, this application also provides a computer-readable medium, which may be included in the electronic device described in the above embodiments; or it may exist independently and not assembled into the electronic device. The computer-readable medium carries one or more computer programs, which, when executed by the electronic device, cause the electronic device to perform the methods described in the above embodiments.

[0137] It should be noted that although several modules or units for the device used to perform actions have been mentioned in the detailed description above, this division is not mandatory. In fact, according to the embodiments of this application, the features and functions of two or more modules or units described above can be embodied in one module or unit. Conversely, the features and functions of one module or unit described above can be further divided and embodied by multiple modules or units.

[0138] From the above description of the embodiments, those skilled in the art will readily understand that the exemplary embodiments described herein can be implemented by software or by combining software with necessary hardware. Therefore, the technical solutions according to the embodiments of this application can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (such as a CD-ROM, USB flash drive, external hard drive, etc.) or on a network, and includes several instructions to cause an electronic device to execute the method according to the embodiments of this application.

[0139] For example, an electronic device can be a video decoding device, then the video decoding device can perform... Figure 8 The video decoding method shown; for example, an electronic device can be a video encoding device, then the video encoding device can perform... Figure 9 The video encoding method shown.

[0140] Other embodiments of this application will readily occur to those skilled in the art upon consideration of the specification and practice of the embodiments disclosed herein. This application is intended to cover any variations, uses, or adaptations of this application that follow the general principles of this application and include common knowledge or customary techniques in the art not disclosed herein.

[0141] It should be understood that this application is not limited to the precise structure described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of this application is limited only by the appended claims.

Claims

1. A video decoding method, characterized in that, include: Multiple pixel regions are determined in the current frame based on a preset pixel region size, and each pixel region contains at least one pixel. The skipped regions in the current frame are determined based on whether each pixel region is processed using the skip mode. Based on the positional relationship between the current block to be decoded and the skipped region, the reconstruction data corresponding to the current block is generated.

2. The video decoding method according to claim 1, characterized in that, The video decoding method further includes: Based on the pixels contained in each pixel region and whether each pixel in the current frame is processed using the skip mode, it is determined whether each pixel region is processed using the skip mode.

3. The video decoding method according to claim 2, characterized in that, The video decoding method further includes: If the pixel difference at the target pixel location is less than or equal to a pixel threshold among a preset number of consecutive image frames adjacent to and preceding the current frame, then the pixel at the target pixel location is determined to be processed using a skip mode.

4. The video decoding method according to claim 2, characterized in that, The video decoding method further includes: If the motion vector of a preset number of consecutive image frames adjacent to and preceding the current frame is 0 at the target pixel location, and the residual value at the target pixel location is less than or equal to the residual threshold, then the pixel at the target pixel location is determined to be processed using the skip mode.

5. The video decoding method according to claim 2, characterized in that, The video decoding method further includes: Based on the flag bits obtained from decoding the video stream, it is determined whether each pixel should be processed using the skip mode.

6. The video decoding method according to claim 2, characterized in that, Determining whether to process each pixel region using skip mode based on the pixels contained in each pixel region and whether each pixel in the current frame is processed using skip mode includes: The pixel region in which all contained pixels are processed using the skip mode is defined as the pixel region to be processed using the skip mode.

7. The video decoding method according to claim 1, characterized in that, Based on the positional relationship between the current block to be decoded and the skipped region, reconstructed data corresponding to the current block is generated, including: If the current block is within the skipped area, the pixel value at the position corresponding to the current block in the previous frame image or the reference frame image of the current frame is used as the reconstruction data corresponding to the current block.

8. The video decoding method according to claim 1, characterized in that, Based on the positional relationship between the current block to be decoded and the skipped region, reconstructed data corresponding to the current block is generated, including: If some pixels in the current block are within the skipped area, after decoding the current block to obtain the decoded data corresponding to the current block, the pixel values ​​at the positions corresponding to the some pixels in the previous frame image of the current frame or the reference frame image of the current frame are used to replace the data at the corresponding positions in the decoded data to obtain the reconstructed data corresponding to the current block.

9. The video decoding method according to claim 1, characterized in that, The video decoding method further includes: The video stream is decoded to obtain a flag bit indicating the size of the pixel region; wherein the flag bit includes one or more of the following flag bits: flag bits contained in the sequence header, flag bits contained in the image header, flag bits contained in the strip header, flag bits contained in the coding tree unit (CTU) header, and flag bits contained in the coding block.

10. The video decoding method according to any one of claims 1 to 9, characterized in that, The video decoding method further includes at least one of the following steps: The video stream is decoded to obtain a flag indicating whether the current frame is decoded based on the skipped region; The video stream is decoded to obtain a flag indicating whether each image block in the current frame is decoded based on the skipped region; The video stream is decoded to obtain a flag indicating whether each strip in the current frame is decoded based on the skipped region.

11. A video encoding method, characterized in that, include: Multiple pixel regions are determined in the current frame based on a preset pixel region size, and each pixel region contains at least one pixel. The skipped regions in the current frame are determined based on whether each pixel region is processed using the skip mode. The current block is encoded based on the positional relationship between the current block to be encoded and the skipped region.

12. The video encoding method according to claim 11, characterized in that, The current block is encoded based on its positional relationship with the skipped region, including: If the current block is within the skipped region, the encoding process for the current block is skipped.

13. The video encoding method according to claim 11, characterized in that, The current block is encoded based on its positional relationship with the skipped region, including: If some pixels in the current block are within the skipped region, then when determining the encoding mode and encoding parameters of the current block, pixel distortion within the skipped region in the current block is excluded.

14. The video encoding method according to claim 11, characterized in that, The current block is encoded based on its positional relationship with the skipped region, including: If some pixels in the current block are within the skipped area, then the pixels in the current block that are adjacent to the skipped area are used to fill the partial pixels, and the encoding mode and encoding parameters of the current block are determined by the pixels after filling.

15. A video decoding device, characterized in that, include: The determining unit is configured to determine multiple pixel regions in the current frame based on a preset pixel region size, and each pixel region contains at least one pixel. The processing unit is configured to determine the skipped region in the current frame based on whether each pixel region is processed in a skipped mode. The generation unit is configured to generate reconstruction data corresponding to the current block based on the positional relationship between the current block to be decoded and the skipped region.

16. A video encoding apparatus, characterized in that, include: The determining unit is configured to determine multiple pixel regions in the current frame based on a preset pixel region size, and each pixel region contains at least one pixel. The processing unit is configured to determine the skipped region in the current frame based on whether each pixel region is processed in a skipped mode. The encoding unit is configured to encode the current block according to the positional relationship between the current block to be encoded and the skipped region.

17. A computer-readable medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the video decoding method of any one of claims 1 to 10, or the video encoding method of any one of claims 11 to 14.

18. An electronic device, characterized in that, include: One or more processors; A memory for storing one or more computer programs that, when executed by one or more processors, cause the electronic device to implement the video decoding method of any one of claims 1 to 10, or the video encoding method of any one of claims 11 to 14.

19. A computer program product, characterized in that, The computer program product includes a computer program stored in a computer-readable storage medium. The processor of the electronic device reads from the computer-readable storage medium and executes the computer program, causing the electronic device to perform the video decoding method of any one of claims 1 to 10, or to implement the video encoding method of any one of claims 11 to 14.