Video data processing method and device, display device, and storage medium

The video data processing method optimizes decoding efficiency by deriving motion vectors based on displayed regions, addressing resource waste in flexible display screens and enhancing user experience.

JP2025540719APending Publication Date: 2025-12-16BOE TECHNOLOGY GROUP CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025530430
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-11-25
Filing Date
2023-11-22
Publication Date
2025-12-16

AI Technical Summary

Technical Problem

Video coding methods for flexible display screens, such as foldable devices, result in inefficient use of decoding resources due to the need to decode non-displayed portions of the screen, leading to resource waste and reduced user experience.

Method used

A video data processing method that determines a first inter-prediction mode to derive a motion vector based on a base region corresponding to the displayed area, allowing partial decoding of the video bitstream, thereby optimizing resource usage.

Benefits of technology

Reduces decoding resource consumption for non-displayed areas, enhancing video coding efficiency and improving user experience on flexible display devices.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025540719000001_ABST
    Figure 2025540719000001_ABST
Patent Text Reader

Abstract

A video data processing method, a video data processing device, a display device, and a storage medium, the method including: determining (S101) to code a current video block of a video using a first inter-prediction mode; and performing (S102) a conversion between the current video block and a bitstream of the video based on the determination, wherein, in the first inter-prediction mode, deriving a motion vector for the current video block based on a base region of the video corresponding to a first display mode. The method allows partial decoding of the bitstream based on the region that actually needs to be displayed, thereby reducing decoding resource consumption for the portion not displayed and improving coding efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] SUMMARY Embodiments of the present disclosure relate to a video data processing method, a video data processing device, a display device, and a computer-readable storage medium. [Background technology]

[0002] Digital video capabilities can be incorporated into a variety of devices, including digital televisions, digital live streaming systems, wireless broadcasting systems, portable or desktop computers, tablet computers, e-readers, digital cameras, digital recording devices, digital media players, video game devices, video game consoles, smartphones, video teleconferencing devices, and video streaming devices. Digital video devices can implement video coding technologies, such as those described in standards defined by MPEG-2, MPEG-4, ITU-TH.263, ITU-TH.264 / MPEG-4, Part 10, Advanced Video Coding (AVC), High Efficiency Video Coding (HEVC), ITU-TH.265 / High Efficiency Video Coding, and extensions to such standards. By implementing these video coding technologies, video devices can more efficiently transmit, receive, encode, decode, and / or store digital video information. Summary of the Invention

[0003] At least one embodiment of the present disclosure provides a video data processing method, the video data processing method including: determining, for a current video block of a video, to code using a first inter-prediction mode; and, based on the determination, performing a conversion between the current video block and a bitstream of the video. In the first inter-prediction mode, deriving a motion vector for the current video block based on a base region corresponding to a first display mode of the video.

[0004] For example, in a method according to at least one embodiment of the present disclosure, for a current video frame of the video, a first display area along the expansion direction from a video expansion start position defined by the first display mode is set as the base area.

[0005] For example, in a method according to at least one embodiment of the present disclosure, in response to the current video block being located within the first display area, the motion vector is within a prediction range of a first motion vector.

[0006] For example, in a method according to at least one embodiment of the present disclosure, the prediction range of the first motion vector is determined based on the position of the current video block, the prediction accuracy of the motion vector, and the boundary of the first display area.

[0007] For example, in a method according to at least one embodiment of the present disclosure, the current video frame includes the first display area and at least one display sub-area laid out adjacently in order along the expansion direction from left to right.

[0008] In response to the current video block being located in a first display sub-area to the right of the first display area in the current video frame, the motion vector is within a prediction range of a second motion vector.

[0009] For example, in a method according to at least one embodiment of the present disclosure, the prediction range of the second motion vector is determined based on the position of the current video block, the prediction accuracy of the motion vector, the boundary of the first display area, and the width of the first display sub-area.

[0010] For example, in a method according to at least one embodiment of the present disclosure, in response to the current video block being located within the kth first display sub-area to the right of the first display area and k=1, the prediction range of the second motion vector is equal to the prediction range of the first motion vector.

[0011] For example, in a method according to at least one embodiment of the present disclosure, in response to the current video block being located within a kth first display sub-area on the right side of the first display area and k being an integer greater than 1, a first right boundary of a prediction range of the first motion vector is different from a second right boundary of a prediction range of the second motion vector.

[0012] For example, in a method according to at least one embodiment of the present disclosure, the current video frame includes the first display area and at least one display sub-area laid out adjacently in order from top to bottom along the expansion direction.

[0013] In response to the current video block being located in a second display sub-area below the first display area, the motion vector is within a prediction range of a third motion vector.

[0014] For example, in a method according to at least one embodiment of the present disclosure, the prediction range of the third motion vector is determined based on the position of the current video block, the prediction accuracy of the motion vector, the boundary of the first display area, and the height of the second display sub-area.

[0015] For example, in a method according to at least one embodiment of the present disclosure, in response to the current video block being located within the mth second display sub-area below the first display area and m=1, the prediction range of the third motion vector is equal to the prediction range of the first motion vector.

[0016] For example, in a method according to at least one embodiment of the present disclosure, in response to the current video block being located within an m-th second display sub-area below the first display area and m being an integer greater than 1, a first lower boundary of a prediction range of the first motion vector is different from a third lower boundary of a prediction range of the third motion vector.

[0017] For example, in a method according to at least one embodiment of the present disclosure, in response to the current video block being located outside the base region, a predicted value of a temporal domain candidate motion vector in a prediction candidate list of motion vectors of the current video block is calculated based on a predicted value of a spatial domain candidate motion vector.

[0018] For example, in a method according to at least one embodiment of the present disclosure, in response to the current video block being located within the base region, all reference pixels used by the current video block are within the base region.

[0019] For example, in a method according to at least one embodiment of the present disclosure, the first inter prediction mode includes a merge prediction mode, an advanced motion vector prediction (AMVP) mode, a merge mode with motion vector difference, a bidirectional weighted prediction mode, or an affine prediction mode.

[0020] At least one embodiment of the present disclosure further provides a video data processing method, the video data processing method including receiving a bitstream of video, determining that a current video block of the video is coded using a first inter-prediction mode, and decoding the bitstream based on the determination, wherein in the first inter-prediction mode, a motion vector for the current video block is derived based on a base region corresponding to a first display mode of the video.

[0021] For example, in a method according to at least one embodiment of the present disclosure, the step of decoding the bitstream includes a step of determining a decoding target region of a current video frame of the video, the decoding target region including at least a first display region corresponding to the base region.

[0022] For example, in a method according to at least one embodiment of the present disclosure, the step of determining the area to be decoded includes a step of determining the area to be decoded based on at least one of the number of pixels and number of coding units to be displayed in the current video frame, the number of coding units in the first display area, the number of displayed pixels in the previous video frame, and the number of coding units in the decoded area of ​​the previous video frame.

[0023] For example, in a method according to at least one embodiment of the present disclosure, the step of determining the area to be decoded includes a step of determining that the area to be decoded includes the decoded area of ​​the previous video frame and one new display sub-area in response to a number of coding units to be displayed of the current video frame being greater than a number of coding units of the decoded area of ​​the previous video frame, or in response to a number of coding units to be displayed of the current video frame being equal to a number of coding units of the decoded area of ​​the previous video frame and a number of pixels to be displayed of the current video frame being greater than a number of displayed pixels of the previous video frame.

[0024] For example, in a method according to at least one embodiment of the present disclosure, the step of determining the area to be decoded includes a step of determining that the area to be decoded of the current video frame includes the area to be displayed of the current video frame in response to the number of coding units to be displayed of the current video frame being greater than the number of coding units of the first display area and the number of coding units to be displayed of the current video frame being less than the number of coding units of the decoded area of ​​the previous video frame, or in response to the number of coding units to be displayed of the current video frame being greater than the number of coding units of the first display area, the number of coding units to be displayed of the current video frame being equal to the number of coding units of the decoded area of ​​the previous video frame, and the number of pixels to be displayed of the current video frame being less than or equal to the number of displayed pixels of the previous video frame.

[0025] At least one embodiment of the present disclosure further provides a video data processing apparatus, comprising: a determining module; and an executing module. The determining module is configured to determine to code a current video block of the video using a first inter-prediction mode. The executing module is configured to perform conversion between the current video block and a bitstream of the video based on the determination. In the first inter-prediction mode, a motion vector for the current video block is derived based on a base region of the video corresponding to a first display mode.

[0026] At least one embodiment of the present disclosure further provides a display device, including a video data processing device and a sliding scrolling screen, wherein the video data processing device is configured to decode a received bitstream according to the method of any one of claims 1 to 20, and transmit decoded pixel values ​​to the sliding scrolling screen for display.

[0027] For example, in a display device according to at least one embodiment of the present disclosure, in response to the sliding scroll screen including a display area and a non-display area in operation, the video data processing device decodes the bitstream based on the size of the display area at the current time point and the previous frame time point.

[0028] For example, a display device according to at least one embodiment of the present disclosure further includes a curl state determination device configured to detect a size of a display area of ​​the sliding scroll screen and send the size of the display area to the video data processing device, so that the video data processing device decodes the bitstream based on the sizes of the display area at a current time point and a previous frame time point.

[0029] For example, at least one embodiment of the present disclosure further provides a video data processing apparatus, comprising: a processor; and a memory including one or more computer program modules stored in the memory and configured to be executed by the processor, the one or more computer program modules including instructions for performing a video data processing method according to any of the previous embodiments.

[0030] For example, at least one embodiment of the present disclosure further provides a computer-readable storage medium having stored thereon computer instructions that, when executed by a processor, implement the steps of the video data processing method according to any of the above embodiments.

[0031] In order to more clearly describe the technical solutions of the embodiments of the present disclosure, the drawings of the embodiments will be briefly described below. Obviously, the drawings described below are only related to some embodiments of the present disclosure, and are not intended to limit the present disclosure. [Brief explanation of the drawings]

[0032] [Figure 1] FIG. 1 is a schematic diagram of a structure of a sliding scroll screen in accordance with at least one embodiment of the present disclosure. [Figure 2] FIG. 1 is a block diagram of an example video coding system in accordance with at least one embodiment of the present disclosure. [Figure 3] FIG. 1 is a block diagram of an example video encoder in accordance with at least one embodiment of the present disclosure. [Figure 4] FIG. 2 is a block diagram of an example video decoder in accordance with at least one embodiment of the present disclosure. [Figure 5] FIG. 1 is a schematic diagram of a full intra-configuration coding structure in accordance with at least one embodiment of the present disclosure. [Figure 6] FIG. 1 is a schematic diagram of an encoding structure for a low-latency configuration in accordance with at least one embodiment of the present disclosure. [Figure 7A]FIG. 1 is a schematic diagram of inter-prediction coding in accordance with at least one embodiment of the present disclosure. [Figure 7B] 1 is a schematic flowchart of an inter-prediction technique in accordance with at least one embodiment of the present disclosure. [Figure 8A] FIG. 1 is a schematic diagram of affine motion compensation in accordance with at least one embodiment of the present disclosure. [Figure 8B] FIG. 10 is a schematic diagram of another affine motion compensation scheme in accordance with at least one embodiment of the present disclosure. [Figure 9] 1 is a schematic diagram of a video data processing method in accordance with at least one embodiment of the present disclosure. [Figure 10] FIG. 1 is a schematic diagram of video coding for a sliding scrolling screen in accordance with at least one embodiment of the present disclosure. [Figure 11] FIG. 1 is a schematic diagram of a current video frame partitioning scheme in accordance with at least one embodiment of the present disclosure. [Figure 12A] 1 is a schematic diagram of a rotation axis direction of a sliding scroll screen in accordance with at least one embodiment of the present disclosure. [Figure 12B] 10 is a schematic diagram of a rotation axis direction of another sliding scroll screen in accordance with at least one embodiment of the present disclosure. FIG. [Figure 13] FIG. 1B is an encoding schematic diagram in which a current video block is located within a first display area in accordance with at least one embodiment of the present disclosure. [Figure 14] FIG. 1B is an encoding schematic diagram in which a current video block is located at the boundary of a first display area in accordance with at least one embodiment of the present disclosure. [Figure 15] 1 is a schematic diagram of another video data processing method in accordance with at least one embodiment of the present disclosure. [Figure 16] 1 is a schematic block diagram of a video coding system in a low-delay configuration in accordance with at least one embodiment of the present disclosure. [Figure 17] 1 is a schematic flowchart of a method for processing video data in a low-latency configuration in accordance with at least one embodiment of the present disclosure. [Figure 18]1 is a schematic flow chart of a video data processing apparatus in accordance with at least one embodiment of the present disclosure. [Figure 19] 1 is a schematic block diagram of another video data processing apparatus in accordance with at least one embodiment of the present disclosure. [Figure 20] 1 is a schematic block diagram of yet another video data processing apparatus in accordance with at least one embodiment of the present disclosure. [Figure 21] FIG. 1 is a schematic block diagram of a non-transitory readable storage medium in accordance with at least one embodiment of the present disclosure. DETAILED DESCRIPTION OF THE INVENTION

[0033] In order to make the objectives, technical solutions and advantages of the embodiments of the present disclosure clearer, the technical solutions of the embodiments of the present disclosure will be clearly and completely described below with reference to the drawings. Obviously, the described embodiments are only a part of the embodiments of the present disclosure, but not all of the embodiments. Based on the described embodiments of the present disclosure, any other embodiments obtained by those skilled in the art without any creative work will also fall within the scope of protection of the present application.

[0034] Flowcharts are used in this disclosure to describe operations performed by systems according to embodiments of the present application. It should be understood that the operations described above or below do not necessarily have to be performed in exact order. Rather, various steps may be processed in reverse order or simultaneously, as appropriate. At the same time, other operations may be added to these processes, or one or more steps of the operations may be removed from these processes.

[0035] Unless otherwise defined, technical or scientific terms used in this disclosure should have the ordinary meaning understood by one of ordinary skill in the art to which this disclosure belongs. The terms "first," "second," and similar terms used in this disclosure do not denote any order, quantity, or importance, but merely serve to distinguish different components. Similarly, similar terms such as "one," "an," or "the" do not denote a limitation of quantity, but rather indicate the presence of at least one. Similar terms such as "comprises" or "has" mean that the device or component described before the term covers the device or component listed thereafter and its equivalents, and do not exclude other devices or components. Similar terms such as "connected" or "coupled" are not limited to physical or mechanical connections, but may also include electrical connections, whether direct or indirect. Terms such as "upper," "lower," "left," "right," and the like only refer to relative positions; if the absolute positions of the described objects change, the relative positions may change accordingly.

[0036] Due to the increasing demand for high-definition video, video coding methods and techniques are widespread in modern technology. Video codecs typically include electronic circuits or software that compress or decompress digital video and are continually being improved to provide higher coding efficiency. Video codecs convert uncompressed video into a compressed format and vice versa. There is a complex relationship between video quality, the amount of data required to represent the video (determined by the bit rate), the complexity of the encoding and decoding algorithms, sensitivity to data loss and errors, ease of editing, random access, and end-to-end delay. Compression formats typically conform to standard video compression specifications, such as the High Efficiency Video Coding (HEVC) standard (also known as H.265), the yet-to-be-finalized Versatile Video Coding (VVC) standard (also known as H.266), or other current and / or future video coding standards.

[0037] As can be appreciated, embodiments of the techniques described herein can be applied to conventional video coding standards (e.g., AVC, HEVC, and VVC) and future standards to improve compression performance. Descriptions of coding operations herein may refer to conventional video coding standards, and it can be appreciated that the methods provided in this disclosure are not limited to the described video coding standards.

[0038] With the emergence of terminal products such as foldable screen mobile phones and foldable screen tablets, research into flexible display screens, such as flexible sliding scroll screens, has attracted increasing attention. FIG. 1 is a schematic diagram of the structure of a sliding scroll screen according to at least one embodiment of the present disclosure. As shown in FIG. 1, the sliding scroll screen typically includes a fully expanded state and a partially expanded state. The unscrolled portion of the sliding scroll screen is considered a display area, and the scrolled portion is considered a non-display area. It should be noted that in the embodiments of the present disclosure, the sliding scroll screen may be any type of display screen with a variable display area, including but not limited to the sliding scroll screen structure shown in FIG. 1. Typically, during the actual use of a sliding scroll screen, such as the scrolling process, the video screen of the scrolled portion of the sliding scroll screen does not need to be displayed, but the screen of that portion is still decoded, resulting in a waste of decoding resources.

[0039] In order to solve at least the above technical problem, at least one embodiment of the present disclosure provides a video data processing method, the method including: determining to code a current video block of a video using a first inter-prediction mode; and performing conversion between the current video block and a bitstream of the video based on the determination. In the first inter-prediction mode, deriving a motion vector for the current video block based on a base region corresponding to a first display mode of the video.

[0040] Correspondingly, at least one embodiment of the present disclosure further provides a video data processing device, a display device, and a computer-readable storage medium corresponding to the above video data processing method.

[0041] A video data processing method according to at least one embodiment of the present disclosure derives a motion vector of a current video block based on a base region corresponding to a first display mode in a video, thereby enabling the video bitstream to be partially decoded based on the display region actually displayed, thereby reducing the consumption of decoding resources for the non-displayed portion, effectively improving the efficiency of video coding, and further improving the user's product usage experience.

[0042] It should be noted that in embodiments of the present disclosure, the meanings of terms describing the positions of neighboring blocks or reference pixels relative to a current video block, such as "above," "below," "left," and "right," are consistent with those defined in video coding standards (e.g., AVC, HEVC, and VVC). For example, in some examples, "left" and "right" refer to opposite sides in the horizontal direction, and "above" and "below" refer to opposite sides in the vertical direction.

[0043] Hereinafter, a number of examples or embodiments and the layout design method according to the present disclosure will be described without limitation by the examples. As will be explained below, different features in these specific examples or embodiments can be combined with each other when not mutually contradictory, to obtain new examples or embodiments, and all of these new examples or embodiments fall within the scope of protection of the present disclosure.

[0044] At least one embodiment of the present disclosure provides a coding system. As can be appreciated, the present disclosure can be implemented using a codec with the same structure for both the encoding and decoding sides.

[0045] 2 is a block diagram illustrating an example video coding system 1000 operable in accordance with some embodiments of the present disclosure. The techniques of this disclosure generally relate to coding (encoding and / or decoding) video data. Generally, video data includes any data for processing video, and thus may include unencoded original video, encoded video, decoded (e.g., reconstructed) video, and video metadata such as grammar data. A video may include one or more pictures, or referred to as a picture sequence.

[0046] 2 , in this example, system 1000 includes a source device 102, which is used to provide encoded video data to be decoded by a destination device 116 for display, and the encoded video data is transmitted to the decoding side by forming a bitstream, which may also be referred to as a bitstream. Specifically, source device 102 provides the encoded video data to destination device 116 via computer-readable medium 110. Source device 102 and destination device 116 may be embodied as a variety of devices, such as a desktop computer, a notebook (i.e., portable) computer, a tablet computer, a mobile device, a set-top box, a smartphone, a handheld phone, a television, a camera, a display device, a digital media player, a video game console, a video streaming device, etc. In some cases, source device 102 and destination device 116 may be configured for wireless communication and may therefore be referred to as wireless communication devices.

[0047] In the example of FIG. 2 , source device 102 includes video source 104, memory 106, video encoder 200, and output interface 108. Destination device 116 includes input interface 122, video decoder 300, memory 120, and display device 118. According to some embodiments of the present disclosure, video encoder 200 of source device 102 and video decoder 300 of destination device 116 may be configured to be used to implement encoding and decoding methods according to some embodiments of the present disclosure. Thus, source device 102 represents an example of a video coding device, and destination device 116 represents an example of a video decoding device. In other examples, source device 102 and destination device 116 may include other assemblies or configurations. For example, source device 102 may receive video data from an external video source, such as an external camera. Similarly, destination device 116 may not incorporate integrated display device 118 but may be connected to an external display device.

[0048] The system 1000 shown in FIG. 2 is merely an example. In general, any digital video encoding and / or decoding device can perform the encoding and decoding methods according to some embodiments of this disclosure. Source device 102 and destination device 116 are merely examples of such coding devices, with source device 102 generating and transmitting a bitstream to destination device 116. This disclosure refers to "coding" devices as devices that perform data coding (encoding and / or decoding). Accordingly, video encoder 200 and video decoder 300 each represent examples of coding devices.

[0049] In some examples, devices 102, 116 are operated in a substantially symmetrical manner, such that devices 102, 116 both include video encoding and decoding assemblies, i.e., devices 102, 116 are both capable of implementing video encoding and decoding processes. Thus, system 1000 can support unidirectional or bidirectional video transmission between video devices 102, 116 and can be used for video streaming, video playback, video broadcasting, or video telephony.

[0050] Generally, the video source 104 represents a video data source (i.e., unencoded original video data) and provides a continuous series of pictures (also called "frames") of video data to the video encoder 200, which encodes the picture data. The video source 104 of the source device 102 may include a video capture device, such as a video camera, a video archive containing previously captured original video, and / or a video feed interface for receiving video from a video content provider. As another alternative solution, the video source 104 may generate computer-generated data as a combination of source video or live video, archival video, and computer-generated video. In various cases, the video encoder 200 encodes captured, pre-captured, or computer-generated video data. The video encoder 200 may reorder the pictures from the order in which they are received (sometimes called "display order") to a coding order for encoding. The video encoder 200 may generate a bitstream containing the encoded video data. Source device 102 then outputs the generated bitstream via output interface 108 to computer-readable medium 110, where it can be received and / or retrieved, such as by input interface 122 of destination device 116, for example.

[0051] Memory 106 of source device 102 and memory 120 of destination device 116 represent general-purpose memory. In some examples, memory 106 and memory 120 may store original video data, such as original video data from video source 104 and decoded video data from video decoder 300. Optionally, memory 106 and memory 120 may also store software instructions executed by video encoder 200, video decoder 300, etc., respectively. While video encoder 200 and video decoder 300 are presented separately in this example, it should be understood that video encoder 200 and video decoder 300 may include internal memory to achieve functionally similar or equivalent purposes. Memory 106 and memory 120 may also store encoded video data output from video encoder 200 and input to video decoder 300, etc. In some examples, portions of memory 106 and memory 120 may be allocated as one or more video buffers, for example, to store decoded original video data and / or encoded original video data.

[0052] The computer-readable medium 110 may represent any type of medium or device capable of transmitting encoded video data from the source device 102 to the destination device 116. In some examples, the computer-readable medium 110 may represent a communications medium, allowing the source device 102 to transmit a bitstream directly and in real time to the destination device 116, such as via a radio frequency network or a computer network. Based on a communications standard, such as a wireless communications protocol, the output interface 108 may modulate a transmission signal containing the encoded video data, and the input interface 122 may modulate a received transmission signal. The communications medium may include wireless or wired communications media, or may include both, such as a radio frequency (RF) spectrum or one or more physical transmission lines. The communications medium may form part of a packet-based network, such as a local area network, a wide area network, or a global network such as the Internet. The communications medium may include a router, a switch, a base station, or any other device capable of facilitating communication from the source device 102 to the destination device 116.

[0053] In some examples, source device 102 may output the encoded data from output interface 108 to storage device 112. Similarly, destination device 116 may access the encoded data from storage device 112 via input interface 122. Storage device 112 may include a variety of distributed or locally accessed data storage media, such as a hard disk drive, a Blu-ray disc, a digital video disc (DVD), a read-only optical disc drive (CD-ROM), flash memory, volatile or non-volatile memory, or any other suitable digital storage medium for storing encoded video data.

[0054] In some examples, source device 102 can output the encoded data to file server 114 or another intermediate storage device capable of storing the encoded video generated by source device 102. Destination device 116 can access the stored video data from file server 114 online or via download. File server 114 may be any type of server device capable of storing the encoded data and transmitting the encoded data to destination device 116. File server 114 may represent a network server (e.g., for a website), a file transfer protocol (FTP) server, a content delivery network device, or a network-attached storage (NAS) device. Destination device 116 can access the encoded data from file server 114 via any standard data connection, including an Internet connection. This may include a wireless channel, such as that applied to a Wi-Fi connection, to access the encoded video data stored on file server 114, a wired connection, such as a digital subscriber line (DSL) or cable modem, or a combination of a wireless channel and a wired connection. File server 114 and input interface 122 may be configured to operate based on a streaming transmission protocol, a download transmission protocol, or a combination thereof.

[0055] Output interface 108 and input interface 122 may represent wired networking assemblies such as wireless transmitters / receivers, modems, Ethernet cards, wireless communication assemblies operating in accordance with any one of the various IEEE 802.11 standards, or other physical assemblies. In examples in which output interface 108 and input interface 122 comprise wireless assemblies, output interface 108 and input interface 122 may be configured to transfer data, such as data encoded in accordance with Fourth Generation Mobile Communications (4G), 4G Long Term Evolution (4G-LTE), Advanced LTE (LTE Advanced), Fifth Generation Mobile Communications (5G), or other cellular communication standards. In some examples in which output interface 108 comprises a wireless transmitter, output interface 108 and input interface 122 may be configured to transfer data, such as data encoded in accordance with other wireless standards, such as the IEEE 802.11 specification, the IEEE 802.15 specification (e.g., ZigBee), the Bluetooth standard, etc. In some examples, source device 102 and / or destination device 116 may include corresponding system-on-chip (SoC) devices. For example, source device 102 may include SoC devices to perform the functionality of video encoder 200 and / or output interface 108, and destination device 116 may include SoC devices to perform the functionality of, for example, video decoder 300 and / or input interface 122.

[0056] The techniques of this disclosure can be applied to video encoding to support multiple multimedia applications, such as over-the-air television broadcasting, cable television transmission, satellite television transmission, Internet streaming video transmission such as HTTP-based dynamic adaptive streams, digital video encoded on a data storage medium, decoding of digital video stored on a data storage medium, or other applications.

[0057] Input interface 122 of destination device 116 receives a bitstream from computer-readable medium 110 (e.g., storage device 112, file server 114, etc.). The bitstream may include signaling information defined by video encoder 200 that is also used by video decoder 300, such as grammar elements with values ​​that describe the nature and / or processing of video blocks or other coding units (e.g., stripes, pictures, groups of pictures, sequences, etc.).

[0058] Display device 118 displays decoded pictures of the decoded video data to a user. Display device 118 may be any type of display device, such as a cathode ray tube (CRT)-based device, a liquid crystal display (LCD), a plasma display, an organic light emitting diode (OLED) display, or other type of display device.

[0059] 2, in some examples, video encoder 200 and video decoder 300 may be integrated with an audio encoder and / or audio decoder, respectively, and may include an appropriate multiplexing-demultiplexing (MUX-DEMUX) unit or other hardware and / or software to process multiplexed streams containing both audio and video in a common data stream. Where applicable, the MUX-DEMUX unit may conform to the ITU H.223 multiplexer protocol or other protocols, such as the User Datagram Protocol (UDP).

[0060] Both the video encoder 200 and the video decoder 300 may be implemented as any suitable codec circuit, such as a microprocessor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), a discrete logic device, software, hardware, firmware, or any combination thereof. When the techniques are implemented partially as software, a device can store instructions for the software on a suitable non-transitory computer-readable medium and use one or more processors to execute the instructions in hardware to perform the techniques of this disclosure. Both the video encoder 200 and the video decoder 300 may be included in one or more encoders or decoders, or either the encoder or decoder may be integrated as part of a combined encoder / decoder (CODEC) in a corresponding device. Devices including the video encoder 200 and / or the video decoder 300 may be integrated circuits, microprocessors, and / or wireless communication devices such as cellular phones.

[0061] Video encoder 200 and video decoder 300 may operate based on a video coding standard, such as ITU-T H.265 (also known as High Efficiency Video Coding (HEVC)), or based on HEVC extensions such as multiview and / or scalable video coding extensions. Optionally, video encoder 200 and video decoder 300 may operate based on other proprietary or industry standards, such as the currently developed Joint Search and Test Model (JEM) or Versatile Video Coding (VVC) standards. The techniques described herein are not limited to any particular coding standard.

[0062] Generally, the video encoder 200 and the video decoder 300 can code video data represented in YUV (e.g., Y, Cb, Cr) format. That is, rather than coding red, green, and blue (RGB) data of a picture sample point, the video encoder 200 and the video decoder 300 can code luma and chroma components, where the chroma components may include red and blue hues. In some examples, the video encoder 200 converts received RGB-formatted data to YUV format before encoding, and the video decoder 300 converts the YUV format to RGB format. Optionally, a pre-processing unit and a post-processing unit (not shown) can perform these conversions.

[0063] In general, video encoder 200 and video decoder 300 may perform a coding process based on blocks of a picture. The terms "block" or "video block" generally refer to a structure that contains data to be processed (e.g., encoded, decoded, or otherwise used in the encoding and / or decoding process). For example, a block may contain a two-dimensional matrix of luma and / or chroma data sample points. In general, the coding process may be performed by first dividing the picture into multiple blocks, and the block in the picture that is being coded may be referred to as the "current block" or "current video block."

[0064] Additionally, embodiments of the present disclosure may relate to processes that include encoding or decoding picture data by coding a picture. Similarly, the present disclosure may relate to processes that include encoding or decoding block data by coding blocks of a picture, such as predictive and / or residual coding. A bitstream obtained by an encoding process typically includes a series of values ​​for grammar elements, which represent a coding strategy (e.g., a coding mode) and information on dividing a picture into blocks. Therefore, encoding a picture or a block can generally be understood as encoding the values ​​of the grammar elements that form the picture or block.

[0065] HEVC defines various blocks, including coding units (CUs), prediction units (PUs), and transform units (TUs). Based on HEVC, a video encoder (e.g., video encoder 200) divides coding tree units (CTUs) into CUs based on a quadtree structure. That is, the video encoder divides CTUs and CUs into four equal, non-overlapping blocks, and each node of the quadtree has zero or four subnodes. A node without subnodes may be called a "leaf node," and a CU of such a leaf node may include one or more PUs and / or one or more TUs. The video encoder can further divide PUs and TUs. For example, in HEVC, a residual quadtree (RQT) represents the division into TUs. In HEVC, a PU represents inter-predicted data, and a TU represents residual data. A CU for intra prediction includes intra-prediction information, such as an intra-mode indication.

[0066] In VVC, a quadtree with a nested multitype tree (partitioned using binary and ternary trees) replaces the concept of multiple partition unit types; that is, it eliminates the separation of CU, PU, ​​and TU concepts unless a CU whose size is too large for the maximum transform length is required and supports greater flexibility in CU partition shape. In the coding tree structure, a CU may have a square or rectangular shape. First, a CTU is divided by the quadtree structure. Then, the quadtree leaf nodes can be further divided by the multitype tree structure. The leaf nodes of the multitype tree are called coding units (CUs), and unless the CU is too large for the maximum transform length, this division is used for prediction and transform processing and does not require any further division. This means that in many cases, CUs, PUs, and TUs have the same block size in a quadtree with a nested multitype tree coding block structure.

[0067] Video encoder 200 and video decoder 300 may be configured to use quadtree partitioning according to HEVC, JEM-based quadtree and binary tree (QTBT) partitioning, or other partitioning structures. It should be understood that the techniques of this disclosure may also be applied to video encoders configured to use quadtree partitioning or other partitioning types. Video encoder 200 encodes video data of a CU to represent prediction information and / or residual information and other information. The prediction information indicates how to predict the CU to form a prediction block of the CU. The residual information typically represents sample-point-by-sample-point differences between sample points of the CU before encoding and sample points of the prediction block.

[0068] The video encoder 200 may further generate grammar data for the video decoder 300, such as block-based grammar data, picture-based grammar data, and sequence-based grammar data, for example, in picture headers, block headers, stripe headers, etc., or other grammar data to generate, for example, a sequence parameter set (SPS), a picture parameter set (PPS), or a video parameter set (VPS). The video decoder 300 may also decode such grammar data to determine how to decode the corresponding video data. For example, the grammar data may include various grammar elements, flags, parameters, etc., used to represent video coding information.

[0069] In this manner, video encoder 200 can generate a bitstream that includes encoded video data, such as grammar elements describing the division into picture blocks (e.g., CUs) and prediction and / or residual information for the blocks. Finally, video decoder 300 can receive the bitstream and decode the encoded video data.

[0070] Generally, the video decoder 300 decodes encoded video data in a bitstream by performing the inverse process of that performed by the video encoder 200. For example, the video decoder 300 may decode values ​​of grammar elements in the bitstream in a manner substantially similar to that of the video encoder 200. The grammar elements may define picture CTUs based on partition information and define CUs of the CTUs by partitioning each CTU based on a partition structure corresponding to a QTBT structure, etc. The grammar elements may further define prediction information and residual information of a block (e.g., a CU) of video data. The residual information may be represented, for example, by quantized transform coefficients. The video decoder 300 may reconstruct a residual block of the block by performing inverse quantization and inverse transform on the quantized transform coefficients of the block. The video decoder 300 forms a prediction block of the block using a prediction mode (intra- or inter-prediction) and related prediction information (e.g., motion information for inter-prediction) signaled in the bitstream. The video decoder 300 can then reconstruct the original block by combining the prediction block and the residual block (on a sample point by sample point basis). The video decoder 300 can also perform additional processing, such as performing a deblocking process to reduce visual artifacts along block boundaries.

[0071] Figure 3 is a block diagram illustrating an exemplary video encoder according to some embodiments of the present disclosure, and correspondingly, Figure 4 is a block diagram illustrating an exemplary video decoder according to some embodiments of the present disclosure, for example, the encoder shown in Figure 3 may be implemented as video encoder 200 in Figure 2, and the decoder shown in Figure 4 may be implemented as video decoder 300 in Figure 2. Below, a codec according to some embodiments of the present disclosure will be described in detail with reference to Figures 3 and 4.

[0072] 3 and 4 are provided for purposes of interpretation and should not be considered as limiting the techniques broadly illustrated and described in this disclosure. For purposes of interpretation, this disclosure describes video encoder 200 and video decoder 300 in the context of evolving video coding standards (e.g., the HEVC video coding standard or the H.266 video coding standard), but the techniques of this disclosure are not limited to these video coding standards.

[0073] The illustration of each unit (also called a module) in FIG. 3 helps to understand the operations performed by video encoder 200. These units may be implemented as fixed-function circuits, programmable circuits, or a combination of both. A fixed-function circuit refers to a circuit that provides a specific function and is preconfigured for the operations it can perform. A programmable circuit refers to a circuit that can be programmed to perform multiple tasks and provides flexibility in the operations it can perform. For example, a programmable circuit can execute software or firmware that causes the programmable circuit to operate in a manner defined by the software or firmware instructions. Although a fixed-function circuit can execute software instructions (e.g., receive parameters or output parameters), the type of operation performed by the fixed-function circuit is typically fixed. In some examples, one or more units may be different circuit blocks (fixed-function circuit blocks or programmable circuit blocks), and in some examples, one or more units may be integrated circuits.

[0074] The video encoder 200 shown in Figure 3 may include an arithmetic logic unit (ALU), a basic functional unit (EFU), digital circuitry, analog circuitry, and / or a programmable core formed by programmable circuitry. In examples where the operations of the video encoder 200 are performed using software executed by programmable circuitry, the memory 106 (Figure 2) may store target code of the software received and executed by the video encoder 200, or other memory (not shown) within the video encoder 200 may be used to store such instructions.

[0075] In the example of FIG. 3, the video encoder 200 can receive input video, for example, from a video data memory or can receive input video directly from a video acquisition device. The video data memory can store video data to be encoded by the video encoder 200 assembly. The video encoder 200 can receive video data stored in the video data memory from, for example, a video source 104 (see FIG. 2). The decoding cache memory can be used as a reference picture memory for storing reference video data, which the video encoder 200 can use when predicting subsequent video data. The video data memory and the decoding cache memory can be formed by multiple memory devices, such as dynamic random access memory (DRAM), including synchronous dynamic random access memory (SDRAM), magnetoresistive random access memory (MRAM), and resistive random access memory (RRAM), or other types of memory devices. The video data memory and the decoding cache memory can be provided by the same storage device or different storage devices. In various examples, the video data memory can be located on the same chip as other assemblies of the video encoder 200, as shown in FIG. 3, or can be located on a different chip from other assemblies.

[0076] In this disclosure, references to video data memory should not be construed as being limited to memory internal to video encoder 200 (unless specifically described as such) or limited to memory external to video encoder 200 (unless specifically described as such). More precisely, references to video data memory should be understood as a reference memory that stores video data for encoding (e.g., video data of a current block to be encoded) received by video encoder 200. Additionally, memory 106 in FIG. 2 may also provide temporary storage for the output of each unit in video encoder 200.

[0077] The mode selection unit typically tests combinations of coding parameters and the rate-distortion values ​​obtained by these combinations by matching multiple coding channels. The coding parameters may include the division of CTUs into CUs, prediction modes of CUs, transformation types of CU residual data, quantization parameters of CU residual data, etc. The mode selection unit can finally select a coding parameter combination that has a better rate-distortion value than other tested combinations.

[0078] Video encoder 200 may divide a picture retrieved from video memory into a series of CTUs and encapsulate one or more CTUs within a stripe. The mode selection unit may divide the CTUs of a picture based on a tree structure (e.g., the QTBT structure described above or the quadtree structure of HEVC). As described above, video encoder 200 may form one or more CUs by dividing the CTUs based on the tree structure. Such a CU may generally be referred to as a "block" or a "video block."

[0079] In general, the mode selection unit further controls its assembly (e.g., a motion estimation unit, a motion compensation unit, and an intra prediction unit) to generate a prediction block for a current block (e.g., a current CU or an overlapping portion of a PU and a TU in HEVC). For inter prediction of the current block, the motion estimation unit can identify one or more closely matching reference blocks in one or more reference pictures (e.g., one or more decoded pictures stored in a decoding cache memory) by performing a motion search. Specifically, the motion estimation unit can calculate a value representing the degree of similarity between a potential reference block and the current block based on, for example, a sum of absolute differences (SAD), a sum of squared differences (SSD), a mean absolute difference (MAD), a mean squared difference (MSD), etc., and the motion estimation unit can usually perform these calculations using the sample-point differences between the current block and the reference block under consideration. The motion estimation unit can identify the reference block with the lowest value resulting from these calculations, thereby indicating the reference block that most closely matches the current block.

[0080] The motion estimation unit may generate one or more motion vectors (MVs), which define the location of a reference block in a reference picture relative to the location of a current block in the current picture. The motion estimation unit may then provide the motion vectors to a motion compensation unit. For example, for unidirectional inter prediction, the motion estimation unit may provide a single motion vector, and for bidirectional inter prediction, the motion estimation unit may provide two motion vectors. The motion compensation unit may then use the motion vectors to generate a predictive block. For example, the motion compensation unit may use the motion vectors to retrieve data for a reference block. As another example, if the motion vector has fractional sample point precision, the motion compensation unit may interpolate the predictive block based on one or more interpolation filters. Also, for bidirectional inter prediction, the motion compensation unit may retrieve data for two reference blocks identified by corresponding motion vectors and combine the retrieved data by averaging or weighted averaging per sample point, etc.

[0081] As another example, for intra prediction, the intra prediction unit may generate a prediction block from sample points neighboring the current block. For example, for a directional mode, the intra prediction unit may generate a prediction block by mathematically combining the values ​​of neighboring sample points and filling these calculated values ​​along a direction defined in the current block. As another example, for a DC mode, the intra prediction unit may calculate average values ​​of sample points neighboring the current block and generate a prediction block that includes the obtained average values ​​of each sample point of the prediction block.

[0082] For other video coding techniques, such as intra block copy mode coding, affine mode coding, and linear model (LM) mode coding, for example, the mode select unit may generate a predictive block for the current block being coded via a corresponding unit associated with the coding technique. In some examples, such as palette mode coding, the mode select unit may not generate a predictive block, but instead generate grammar elements that indicate how to reconstruct the block based on a selected palette. In such modes, the mode select unit may provide these grammar elements to the entropy coding unit for coding.

[0083] As described above, the residual unit receives the current block and the corresponding predicted block. The residual unit then generates a residual block for the current block. To generate the residual block, the residual unit calculates the sample-point-by-sample difference between the predicted block and the current block.

[0084] A transform unit (shown as "Transform & Sample & Quantize" in FIG. 3) generates blocks of transform coefficients (e.g., referred to as "transform coefficient blocks") by applying one or more transforms to a residual block. A transform unit may form a transform coefficient block by applying various transforms to a residual block. For example, a transform unit may apply a discrete cosine transform (DCT), a directional transform, a Karhunen-Loeve transform (KLT), or a conceptually similar transform to a residual block. In some examples, a transform unit may perform multiple transforms on a residual block, e.g., a linear transform and a secondary transform, e.g., a rotation transform. In some examples, a transform unit may not apply a transform to a residual block.

[0085] Subsequently, the transform unit may quantize the transform coefficients in the transform coefficient block to generate a quantized transform coefficient block. The transform unit may quantize the transform coefficients of the transform coefficient block based on a quantization parameter (QP) value associated with the current block. The video encoder 200 may adjust the degree of quantization applied to the coefficient block associated with the current block by adjusting the QP value associated with the CU (e.g., via a mode selection unit). Quantization may cause information loss, and thus the precision of the quantized transform coefficients may be lower than the precision of the original transform coefficients.

[0086] The encoder 200 may also include a coding control unit, which can generate control information for operations during the encoding process. Subsequently, the inverse quantization and inverse transform unit ("Inverse Quantization & Inverse Transform" shown in FIG. 3) can obtain a reconstructed residual block from the transform coefficient block by applying inverse quantization and inverse transform, respectively, to the quantized transform coefficient block. The reconstruction unit can generate a reconstructed block (which may have some distortion) corresponding to the current block based on the reconstructed residual block and the prediction block generated by the mode selection unit. For example, the reconstruction unit can generate the reconstructed block by adding sample points of the reconstructed residual block to corresponding sample points of the prediction block generated by the mode selection unit.

[0087] The reconstruction block may perform one or more filtering operations by a filtering process, such as the loop filtering unit shown in FIG. 3. For example, the filtering process may include a deblocking operation to reduce blocking artifacts along CU edges. In some examples, operations in the filtering process may be skipped.

[0088] Subsequently, the video encoder 200 can store the reconstructed blocks in a decoding cache memory, e.g., by loop filtering. In examples that skip filtering, the reconstruction unit can store the reconstructed blocks in the decoding cache memory. In examples that require filtering, the filtered reconstructed blocks can be stored in the decoding cache memory. The motion estimation unit and motion compensation unit can perform inter prediction on blocks of a subsequently coded picture by retrieving reference pictures formed by the reconstructed (and possibly filtered) blocks from the decoding cache memory. Also, the intra prediction unit can perform intra prediction on other blocks in the current picture using the reconstructed blocks in the decoding cache memory of the current picture.

[0089] The operations described above are block-wise. The description should be understood as operations for luma-coding blocks and / or chroma-coding blocks. As mentioned above, in some examples, the luma-coding blocks and chroma-coding blocks are the luma and chroma components of a CU. In some examples, the luma-coding blocks and chroma-coding blocks are the luma and chroma components of a PU.

[0090] In general, the entropy coding unit can entropy code grammar elements received from other functional assemblies of the video encoder 200. For example, the entropy coding unit can entropy code quantized transform coefficient blocks from the transform unit. For example, the entropy coding unit can entropy code prediction grammar elements (e.g., motion information for inter prediction or intra mode information for intra prediction) from the mode selection unit to generate entropy-coded data. For example, the entropy coding unit can perform a context-adaptive variable-length coding (CAVLC) operation, a context-adaptive binary arithmetic coding (CABAC) operation, a variable-length coding operation, a grammar-based context-adaptive binary arithmetic coding (SBAC) operation, a probability interval partitioning entropy (PIPE) coding operation, an exponential-Golomb coding operation, or other types of entropy coding operations on the data. In some examples, the entropy coding unit can operate in a bypass mode in which the grammar elements are not entropy coded. The video encoder 200 can output a bitstream including the entropy-coded grammar elements necessary to reconstruct blocks of a stripe or picture.

[0091] 4 is a block diagram illustrating an exemplary video decoder according to some embodiments of the present disclosure; for example, the decoder illustrated in FIG. 4 may be video decoder 300 in FIG. 2. As can be understood, providing FIG. 4 is for purposes of interpretation and not for purposes of limiting the techniques broadly illustrated and described in this disclosure. For purposes of interpretation, video decoder 300 is described based on HEVC technology. However, the techniques of the present disclosure can be performed by video decoding devices configured for other video coding standards.

[0092] As can be understood, in practical applications, the basic structure of the video decoder 300 may be similar to that of the video encoder shown in FIG. 3, whereby both the encoder 200 and the decoder 300 include video encoding and decoding assemblies, i.e., both the encoder 200 and the decoder 300 can implement video encoding and decoding processes. In such a case, the encoder 200 and the decoder 300 may be collectively referred to as a codec. Therefore, a system consisting of the encoder 200 and the decoder 300 can support unidirectional or bidirectional video transmission between devices, and can be used for, for example, video streaming, video playback, video broadcasting, or video telephony. As can be understood, the video decoder 300 may include more, fewer, or different functional assemblies than those shown in FIG. 4. For ease of understanding, FIG. 4 illustrates assemblies related to the decoding and conversion process according to some embodiments of the present disclosure.

[0093] In the example of FIG. 4, the video decoder 300 includes a memory, an entropy decoding unit, a prediction processing unit, an inverse quantization and inverse transform unit (shown as the "inverse quantization & inverse transform unit" in FIG. 4), a reconstruction unit, a filter unit, a decoding cache memory, and a bit-depth inverse transform unit. The prediction processing unit may include a motion compensation unit and an intra prediction unit. The prediction processing unit may further include, for example, an addition unit, thereby enabling prediction based on other prediction modes. By way of example, the prediction processing unit may include a palette unit, an intra block copy unit (which may form part of the motion compensation unit), an affine unit, a linear model (LM) unit, etc. In other examples, the video decoder 300 may include more, fewer, or different functional assemblies.

[0094] As shown in FIG. 4, decoder 300 may first receive a bitstream containing encoded video data. For example, the memory in FIG. 4 may be called a coding picture buffer (CPB), which is used to store the bitstream containing the encoded video data and wait for the bitstream to be decoded by the video decoder 300 assembly. The video data stored in the CPB may be obtained, for example, from computer-readable medium 110 (FIG. 2). The CPB may also store temporary data output from each unit of video decoder 300. The decoding cache memory typically stores decoded pictures, which video decoder 300 can output and / or use as reference video data when decoding subsequent data or pictures in the bitstream. The CPB memory and the decoding cache memory may be formed by multiple memory devices, such as dynamic random access memories (DRAMs), including synchronous dynamic random access memories (SDRAMs), magnetoresistive random access memories (MRAMs), and resistive random access memories (RRAMs), or other types of memory devices. The CPB memory and the decoding cache memory may be provided by the same storage device or different storage devices. In various examples, the CPB memory may be located on the same chip as the other assemblies of video decoder 300, or may not be located on the same chip as the other assemblies, as shown.

[0095] The various units shown in FIG. 4 are presented to aid in understanding the operations performed by video decoder 300. These units may be implemented as fixed-function circuits, programmable circuits, or a combination of both. Similar to FIG. 3, a fixed-function circuit refers to a circuit that provides a specific function and is preconfigured for the operations it can perform. A programmable circuit refers to a circuit that can be programmed to perform multiple tasks and provides flexibility in the operations it can perform. For example, a programmable circuit can execute software or firmware that causes the programmable circuit to operate in a manner defined by the software or firmware instructions. While a fixed-function circuit can execute software instructions (e.g., receive parameters or output parameters), the type of operation performed by the fixed-function circuit is typically fixed. In some examples, one or more units may be different circuit blocks (either fixed-function circuit blocks or programmable circuit blocks), and in some examples, one or more units may be integrated circuits.

[0096] Video decoder 300 may include a programmable core formed by ALUs, EFUs, digital circuits, analog circuits, and / or programmable circuits. In examples where the operation of video decoder 300 is performed by software executed in programmable circuits, on-chip or off-chip memory may store software instructions (e.g., target code) received and executed by video decoder 300.

[0097] Subsequently, the entropy decoding unit can perform entropy decoding on the received bitstream to parse the coded information corresponding to the picture therefrom.

[0098] Subsequently, the decoder 300 can perform a decoding conversion process based on the analyzed coding information to generate display video data. According to some embodiments of the present disclosure, the operations that can be performed by the decoder 300 located on the decoding side can refer to the decoding conversion process shown in Figure 4, which can be understood as including a general decoding process to generate a display picture for display by a display device.

[0099] In the decoder 300 shown in Figure 4, the entropy decoding unit can receive a bitstream containing encoded video, for example from the memory 120, and perform entropy decoding on it to recreate the grammar elements. The inverse quantization and inverse transform unit ("Inverse Quantization & Inverse Transform" shown in Figure 4), the reconstruction unit, and the filter unit can generate decoded video based on the grammar elements extracted from the bitstream, for example, generating decoded pictures.

[0100] Generally, video decoder 300 reconstructs a picture block by block. Video decoder 400 may perform a reconstruction operation on each block independently, and the block currently being reconstructed (i.e., decoded) may be referred to as the "current block."

[0101] Specifically, the entropy decoding unit may perform entropy decoding on transformation information, such as grammar elements and quantization parameters (QPs) and / or mode conversion instructions, that define the quantized transform coefficients of the quantized transform coefficient block. The inverse quantization and inverse transform unit may use the QP associated with the quantized transform coefficient block to determine the degree of quantization and may also determine the degree of inverse quantization to be applied. For example, the inverse quantization and inverse transform unit may inverse quantize the quantized transform coefficients by performing a bitwise left-shift operation. The inverse quantization and inverse transform unit may thereby form a transform coefficient block including the transform coefficients. After forming the transform coefficient block, the inverse quantization and inverse transform unit may generate a residual block associated with the current block by applying one or more inverse transforms to the transform coefficient block. For example, the inverse quantization and inverse transform unit may apply an inverse DCT, an inverse integer transform, an inverse Karhunen-Loeve transform (KLT), an inverse rotational transform, an inverse transform, or other inverse transform to the coefficient block.

[0102] The prediction processing unit may also generate a prediction block based on a grammar element of the prediction information that is entropy decoded by the entropy decoding unit. For example, if the grammar element of the prediction information indicates that the current block is inter-predicted, the motion compensation unit may generate a prediction block. In this case, the grammar element of the prediction information may indicate a reference picture in the decoding cache memory (to retrieve the reference block from this reference picture) and may indicate identifying a motion vector of the reference block in the reference picture relative to the position of the current block in the current picture. The motion compensation unit may generally perform an inter-prediction process in a manner essentially similar to that described for the motion compensation unit in FIG. 3.

[0103] As another example, if the grammar element for prediction information indicates that intra prediction is to be performed on the current block, the intra prediction unit may generate a predicted block based on the intra prediction mode indicated by the grammar element for prediction information. Similarly, the intra prediction unit may generally perform an intra prediction process in a manner essentially similar to that described for the intra prediction unit in Figure 3. The intra prediction unit may retrieve data of sample points adjacent to the current block from a decoding cache memory.

[0104] The reconstruction unit may use the prediction block and the residual block to reconstruct the current block, for example, by adding sample points of the residual block to corresponding sample points of the prediction block.

[0105] The filter unit may then perform one or more filter operations on the reconstructed block. For example, the filter unit may perform a deblocking operation to reduce blocking artifacts along the edges of the reconstructed block. As can be appreciated, the filtering operation need not be performed in all instances, i.e., the filtering operation may be skipped in some cases.

[0106] The video decoder 300 may store the reconstructed blocks in a decoding cache memory. As described above, the decoding cache memory may provide reference information to, for example, a motion compensation or motion estimation unit, such as sample points of the current picture for intra-prediction and sample points of previously decoded pictures for subsequent motion compensation. The video decoder 300 may also output the decoded pictures from the decoding cache memory for subsequent presentation to a display device (e.g., display device 118 of FIG. 2).

[0107] FIG. 5 is a schematic diagram of an All Intra (AI) coding structure in accordance with at least one embodiment of the present disclosure.

[0108] For example, as shown in Figure 5, in a full intra AI configuration, all frames in a video are coded according to the I-frame during the coding process, i.e., the coding process is completely independent and has no dependency on other frames. At the same time, the quantization parameter (QP) during the coding process does not vary depending on the coding position, and is always equal to the QP value (QPI) of the first frame. As shown in Figure 5, in a full intra AI configuration, the playback order and coding order of all frames in a video are the same, i.e., the playback order count (POC) and coding order count (EOC) of video frames are the same.

[0109] FIG. 6 is a schematic diagram of a coding structure for a low-delay (LD) configuration in accordance with at least one embodiment of the present disclosure.

[0110] In practical applications, low-delay LD configurations are typically applied to real-time communication environments with low-delay needs. As shown in Figure 6, in LD configurations, all P frames or B frames use generalized P / B frame prediction, and the EOC of all frames still matches the POC. For low-delay configurations, a "1+x" solution is proposed, where "1" is one nearest neighbor reference frame and "x" is x high-quality reference frames.

[0111] 7A is a schematic diagram of inter-prediction coding in accordance with at least one embodiment of the present disclosure, and FIG. 7B is a schematic flow chart of an inter-prediction technique in accordance with at least one embodiment of the present disclosure.

[0112] The main concept of video predictive coding is to eliminate correlation between pixels through prediction. Based on the location of reference pixels, video predictive coding techniques are mainly divided into two types: intra prediction (1), which generates a predicted value using coded pixels in the current image (current video frame), and inter prediction (2), which generates a predicted value using reconstructed pixels from a coded image preceding the current image (current video frame). Inter predictive coding refers to effectively eliminating redundancy in the video temporal domain by utilizing correlation in the video time domain and predicting pixels in the current image using pixels from adjacent coded images. As shown in Figures 7A and 7B, the inter predictive coding algorithm in the HEVC / H265 standard uses a coded image as a reference image for the current image to obtain motion information in the reference image for each block in the current image. The motion information is usually represented by a motion vector and a reference frame index. The reference image may be forward, backward, or bidirectional. When using inter coding techniques to obtain information for the current block, motion information may be directly inherited from adjacent blocks, or motion estimation may be used to search for matching blocks in the reference image to obtain corresponding motion information. Then, a motion compensation process is performed to obtain a prediction value for the current block.

[0113] Each pixel block of the current image finds one best matching block in the previous coded image, a process called motion estimation. The image for prediction is called the reference image, the displacement from the reference block to the current block (i.e., the current pixel block) is called the motion vector, and the difference between the current block and the reference block is called the prediction residual.

[0114] The video coding standard defines three types of images: I frame images, P frame images, and B frame images. I frame images can only use intra-coding, while P frame images and B frame images can use inter-prediction coding. The prediction method for P frame images is to predict the current image using the previous frame image, and this method is called "forward prediction." That is, a matching block (reference block) for the current block is found in the forward reference image. B frame images can use three prediction methods: forward prediction, backward prediction, and bidirectional prediction.

[0115] Inter prediction coding relies on inter correlation and includes processes such as motion estimation, motion compensation, etc. For example, in some examples, the main process of inter prediction includes steps 1 to 7.

[0116] Step 1: Create a candidate list of motion vectors (MVs), perform Lagrangian rate-distortion optimization (RDO) calculations, and select the MV with the smallest distortion as the initial MV.

[0117] Step 2: The point with the smallest matching error in step 1 is found as the starting point for the next search.

[0118] Step 3: Starting with a step size of 1 and incrementing by an exponent of 2, we perform an 8-point diamond search, and can set the maximum number of searches in that step (one traverse with a step size counts as one).

[0119] Step 4: If the optimal step size obtained by the search in Step 3 is 1, then a two-point diamond search needs to be performed using the optimal point as the starting point, because two of the eight adjacent points of this optimal point have not been searched in the previous eight-point search.

[0120] Step 5: If the optimal step size found in step 3 is greater than a certain threshold (iRaster), a raster scan is performed with a step size of iRaster, starting from the point found in step 2 (i.e., traversing all points within the motion search range).

[0121] Step 6: After going through the previous steps 1-5, steps 3 and 4 are repeated again, starting from the optimal point obtained.

[0122] Step 7: The MV corresponding to the best matching point is saved as the final MV and the sum of absolute errors (SAD).

[0123] In the video coding standard H.265 / HEVC, inter-prediction techniques mainly include Merge and Advanced Motion Vector Prediction (AMVP) techniques, which use temporal and spatial motion video prediction concepts. The core concept of both these two technologies is to create a predicted MV list for one candidate motion vector and select the MV with the best performance as the predicted MV for the current coding block. For example, in Merge mode, one MV candidate list is created for the current prediction unit (PU), and there are five candidate MVs (and their corresponding reference images) in the list. These five candidate MVs are traversed and the rate-distortion cost is calculated, and the one with the lowest rate-distortion cost is finally selected as the optimal MV for the Merge mode. If the encoding / decoding sides create the candidate list in the same manner, the encoder only needs to transmit the index of the optimal MV in the candidate list, thereby significantly reducing the number of coding bits for motion information.

[0124] The video coding standard VVC uses the same motion vector prediction technology as HEVC, but with some optimizations, such as extending the length of the merge motion vector candidate list and modifying the candidate list construction process, and also adding some new prediction techniques, such as affine transformation technology and adaptive motion vector accuracy technology.

[0125] 8A and 8B show schematic diagrams of affine motion compensation according to at least one embodiment of the present disclosure. Figure 8A shows a two-control-point affine transformation, i.e., the affine motion vector of the current block is generated by two control points (four parameters). Figure 8B shows a three-control-point affine transformation, i.e., the affine motion vector of the current block is generated by three control points (six parameters).

[0126] For a four-parameter affine motion model, the calculation method for the motion vector of a sub-block whose center pixel is (x,y) is as follows.

[0127]

number

[0128] The reference point vectors a and b are (a h ,a v ), (b h ,b v ), and the horizontal prediction vector of the four-parameter affine motion model for the two points of the current center pixel is (MV h ,MV v ) and can be expressed in terms of a, b vectors as follows (similarly, the vertical prediction of the two-point four-parameter affine motion model of the central pixel can be found), where w and h denote the width and height of the current block, respectively:

[0129] When three control points a, b, and c are used, that is, when a six-parameter affine motion model is used, the motion vector of a sub-block whose center pixel is (x, y) is calculated as follows.

[0130]

number

[0131] The 6-parameter affine motion model has one more reference point c than the 4-parameter affine motion model, and the motion vector (c h ,c v )

[0132] FIG. 9 is a schematic diagram of a video data processing method in accordance with at least one embodiment of the present disclosure.

[0133] For example, at least one embodiment of the present disclosure provides a video data processing method 10. The video data processing method 10 can be applied to various application scenarios related to video coding, such as terminals such as mobile phones and computers, and can also be applied to video websites / video platforms, etc., although the embodiments of the present disclosure are not specifically limited thereto. For example, as shown in FIG. 9, the video data processing method 10 includes the following operations S101 to S102.

[0134] Step S101: Determine to use a first inter-prediction mode for coding a current video block of a video.

[0135] Step S102: Based on the determination, perform conversion between the current video block and the video bitstream: in a first inter-prediction mode, derive a motion vector for the current video block based on a base region corresponding to a first display mode in the video.

[0136] For example, in at least one embodiment of the present disclosure, the video may be a filmed image, a video downloaded from a network, or a locally stored video, and may be an LDR video, an SDR video, etc., and the embodiments of the present disclosure are not limited thereto.

[0137] For example, in at least one embodiment of the present disclosure, a first display mode of a video may define that the video screen gradually increases from a certain expansion starting position along a certain expansion direction (e.g., horizontally or vertically). For example, in some examples, one or more grammar elements associated with the first display mode exist in the video bitstream. For example, in some examples, the video screen expands from left to right (i.e., horizontally), and further, in some examples, the video screen expands from top to bottom (i.e., vertically), or vice versa. For example, in one example, the first display mode of a video defines information such as the size and position of a base region. It should be noted that the "first display mode" is not limited to any one or several specific display modes, nor is it limited to a specific order.

[0138] For example, in at least one embodiment of the present disclosure, any inter-prediction mode can be selected for the current video block of a video. For example, in the H.265 / HEVC standard, the current video block can be inter-predictively encoded by selecting Merge mode or AMVP mode. For example, for Merge mode, a motion vector MV candidate list is constructed for the current block, and the candidate list includes five candidate MVs. These five candidate MVs typically include two types: spatial and temporal. The spatial domain provides a maximum of four candidate MVs, and the temporal domain provides a maximum of one candidate MV. If the number of candidate MVs in the current MV candidate list is five or less, it must be replenished using a zero vector (0,0) to reach the specified number. Similar to Merge mode, the MV candidate list constructed in AMVP mode also includes two types: spatial and temporal; the difference is that the length of the AMVP list is only two.

[0139] The H.266 / VVC standard extends the size of the merge mode candidate list to include up to six candidate MVs. The VVC standard also introduces new inter prediction techniques, such as affine prediction mode, combined intra-inter prediction (CIIP) mode, geometric partitioning prediction mode (TPM), bidirectional optical flow (BIO) method, bidirectional weighted prediction (BCW), merge mode with motion vector difference, etc. In embodiments of the present disclosure, for a current video block, the available inter prediction modes may include any one of the above inter prediction modes, and embodiments of the present disclosure are not limited thereto.

[0140] For example, in at least one embodiment of the present disclosure, rate-distortion optimization (RDO) is used to select an optimal inter-prediction mode and motion vector to obtain better compression performance while maintaining image quality. For example, in an embodiment of the present disclosure, a "first inter-prediction mode" is used to indicate the inter-prediction mode for a current video block. It should be noted that the "first inter-prediction mode" is not limited to any particular inter-prediction mode or a particular order.

[0141] For example, in at least one embodiment of the present disclosure, the first inter prediction mode may be any of the above-mentioned available inter prediction modes, such as merge mode, advanced motion vector prediction (AMVP) mode, merge mode with motion vector difference, bidirectional weighted prediction mode, or affine prediction mode, and the embodiments of the present disclosure are not limited thereto.

[0142] For example, in at least one embodiment of the present disclosure, for step S102, converting between the current video block and the bitstream may include encoding the current video block into the bitstream or decoding the current video block from the bitstream. For example, the conversion process may include an encoding process or a decoding process, and embodiments of the present disclosure are not limited thereto.

[0143] For example, in at least one embodiment of the present disclosure, in a first inter-prediction mode, a motion vector for a current video block is derived based on a base region corresponding to a first display mode in the video. For example, in some examples, the motion vector for the current video block is related to the position and / or size of the base region. For example, in some examples, the reference pixels used for the current video block are restricted to within a specific region. In this way, by restricting the use of reference pixels to only within a region effective for the inter-prediction mode of the current video block, the video bitstream can be partially decoded based on the display region actually displayed, thereby reducing the consumption of decoding resources for non-displayed portions, effectively improving video coding efficiency, and further improving the user's product usage experience.

[0144] For example, in embodiments of the present disclosure, the base region refers to a region that is displayed throughout the video and can be determined by the first display mode of the video. For example, in embodiments of the present disclosure, the base region (also referred to herein as the first display region) is a fixed region having a predetermined length along the video expansion direction from the video expansion start position defined by the first display mode. For example, the position of the base region may change depending on the first display mode of the video. For example, in the example shown in FIG. 1, the first display mode of the video defines the video screen to gradually increase from the leftmost dashed line along the left-to-right direction. In such a case, the fixed region is a region that starts from the leftmost side of the video screen and extends a predetermined length from left to right. For example, in some examples, the first display mode of the video defines the video screen to start from the top and gradually increase along the top-to-bottom direction. In such a case, the fixed region is a region that starts from the top of the video screen and extends a predetermined length from top to bottom. For example, in some examples, the ratio of the length of the base region to the entire display area is 1:1. Furthermore, for example, in some examples, the length ratio between the base area and the entire display area is 1:2. The embodiments of the present disclosure do not limit the position / size of the base area, and can be set according to actual circumstances.

[0145] For example, in at least one embodiment of the present disclosure, one or more grammar elements associated with the first display mode can define the application of the first display mode, the video expansion direction, the size of the base area, etc., and the embodiments of the present disclosure are not limited thereto and can be set according to actual circumstances.

[0146] FIG. 10 is a schematic diagram of video coding for a sliding scrolling screen in accordance with at least one embodiment of the present disclosure.

[0147] For example, as shown in FIG. 10, in at least one embodiment of the present disclosure, when the video data processing method 10 is applied to a display device with a sliding scrolling screen, video content corresponding to the scrolled area may not be displayed during the scrolling process of the sliding scrolling screen, and therefore the bitstream corresponding to the scrolled area may not be decoded. In FIG. 10, the shaded area represents the area not participating in decoding. As the sliding scrolling screen gradually scrolls from the fully expanded state, the size of the area not participating in decoding gradually increases. Therefore, a technical effect of partially decoding the received bitstream based on the actual display area of ​​the sliding scrolling screen can be achieved, thereby reducing resource consumption during the decoding process, thereby improving the durability of products with sliding scrolling screens and improving the user experience.

[0148] For example, in at least one embodiment of the present disclosure, as shown in FIG. 10, during the scrolling process of a sliding scroll screen, partial decoding of a P frame image is required to satisfy the requirement that the MV prediction range of the current video frame to be displayed is a subset of the display area of ​​the reference frame. Tile coding, designed in conventional coding standards, can divide the screen into regions, but improves parallel coding capabilities by encoding only the coding tree units (CTUs) included in a Tile in scan order. Because Tile only limits the range of intra prediction, i.e., intra prediction modes do not utilize pixel information beyond the Tile range. However, for inter-coding and loop filtering modules, there is a possibility of crossing the Tile boundary. Therefore, Tile requires decoding the entire frame image of the reference frame, and independent decoding within the limited range is not possible. Therefore, conventional Tile coding cannot meet the coding needs of products such as sliding scroll screens. In response to the above technical issues, at least one embodiment of the present disclosure provides a progressive coding structure.

[0149] FIG. 11 is a schematic diagram of a current video frame partitioning scheme in accordance with at least one embodiment of the present disclosure.

[0150] For example, in at least one embodiment of the present disclosure, the current video frame includes a first display region and at least one display sub-region laid out adjacently in sequence along the deployment direction (left to right), as shown in Figure 11. For example, in the example shown in Figure 11, the first display region is represented as a basic_tile region on the left, and at least one display sub-region is represented as at least one enhanced_tile region on the right.

[0151] For example, in some examples, the first display area may be a fixed display area, such as a fixed display area of ​​a general display screen. For example, in a sliding scroll screen application scenario, as shown in FIG. 10, the first display area is a display area that is displayed throughout the scrolling process of the sliding scroll screen. For example, in some examples, the first display area is half the area of ​​the video screen, for example, occupying half the number of CTUs of the video screen. Further, for example, in some examples, the first display area is one-third the area of ​​the video screen, for example, occupying one-third the number of CTUs of the video screen. Further, for example, in some examples, the first display area is the entire area of ​​the video screen. It should be noted that the embodiments of the present disclosure are not specifically limited thereto and can be configured according to actual needs. It should be further noted that in the embodiments of the present disclosure, the term "first display area" is used to refer to a fixedly displayed display area (i.e., a base area), and is not limited to any specific display area or to a specific order.

[0152] For example, in at least one embodiment of the present disclosure, other display areas in the current video frame than the first display area can be evenly divided into at least one display sub-area, such as at least one enhanced_tile as shown in FIG. 11. For example, the number of display sub-areas varies depending on the scrolling state of the sliding scroll screen. It is worth noting that in the embodiment of the present disclosure, each of the at least one display sub-area has the same size, and the width or height of each display sub-area is greater than one CTU.

[0153] FIG. 12A is a schematic diagram of a rotation axis direction of a sliding scroll screen in accordance with at least one embodiment of the present disclosure, and FIG. 12B is a schematic diagram of a rotation axis direction of another sliding scroll screen in accordance with at least one embodiment of the present disclosure.

[0154] For example, in at least one embodiment of the present disclosure, the video encoding direction is typically from left to right in the horizontal direction and then from top to bottom in the vertical direction. In the example shown in FIG. 12A, the rotation axis direction of the sliding scroll screen is considered to be perpendicular to the video encoding direction. For example, in at least one embodiment of the present disclosure, when the sliding scroll screen scrolls unfolds in the horizontal direction, the rotation axis moves from left to right, and the display target area becomes larger and larger. In the example shown in FIG. 12B, the rotation axis direction of the sliding scroll screen is considered to be parallel to the video encoding direction. For example, in at least one embodiment of the present disclosure, when the sliding scroll screen scrolls unfolds in the vertical direction, the rotation axis moves from top to bottom, and the display target area becomes larger and larger.

[0155] For example, in the example shown in Figure 12A, the current video frame may include a first display region (basic_tile) on the leftmost side and at least one display sub-region (enhanced_tile) laid out adjacent to each other in the expansion direction (left to right). For example, in the example shown in Figure 12B, the current video frame may include a first display region (basic_tile) on the topmost side and at least one display sub-region (enhanced_tile) below the first display region laid out adjacent to each other in the expansion direction (top to bottom).

[0156] It is worth noting that in the embodiments of the present disclosure, the "first display sub-area" refers to the display sub-area (enhanced_tile) on the right side of the base area / first display area (basic_tile), and the "second display sub-area" refers to the display sub-area (enhanced_tile) below the base area / first display area (basic_tile). The "first display sub-area" and the "second display sub-area" are not limited to any particular one or several display sub-areas, nor are they limited to a particular order.

[0157] For example, in at least one embodiment of the present disclosure, in response to the current video block being located within a first display area of ​​the current video frame, the motion vector of the current video block is within a prediction range of the first motion vector.

[0158] For example, in at least one embodiment of the present disclosure, an independent encoding scheme is used for the base region (basic_tile), as shown in Figure 11. For example, for each video frame of a video, the base region basic_tile needs to be decoded. By defining one effective MV prediction range (i.e., a first MV prediction range) for a current video block located within the basic_tile region, all MVs associated with the current video block can be within the MV prediction range.

[0159] FIG. 13 illustrates a schematic diagram of encoding a video block located within a first display area in accordance with at least one embodiment of the present disclosure.

[0160] For example, in at least one embodiment of the present disclosure, the prediction range of the first motion vector is determined based on the position of the current video block, the prediction accuracy of the motion vector, and the boundary of the first display area.

[0161] For example, in some examples, as shown in FIG. 13, the four boundaries of the first display area basic_tile are respectively a left boundary

number

number

number

number

[0162]

number

[0163] In the above formulas (1) to (4),

number

number

number

number

[0164] For example, in at least one embodiment of the present disclosure, the first MV prediction range defined by Equations (1)-(4) is the limiting range of the initial MV (corresponding to the search starting point) of the current video block, and is also the limiting range of the final MV (corresponding to the best matching point) of the current video block, thus ensuring that the final MV of the current video block is within the corresponding MV prediction range (first MV prediction range).

[0165] For example, in at least one embodiment of the present disclosure, in response to the current video block being located within a base region, all reference pixels used by the current video block are within the base region, which may similarly be defined by equations (1)-(4) above.

[0166] For example, in at least one embodiment of the present disclosure, for a current video block located at position (x, y) and within basic_tile, all associated motion vectors should satisfy the following equations (5) and (6):

[0167]

number

[0168] In equations (5) and (6),

number

number

[0169] For example, in at least one embodiment of the present disclosure, for an MV candidate list of a current video block, it can be determined whether each MV in the candidate list is within the corresponding MV prediction range. For example, in some examples, if it is determined that any MV is not within the MV prediction range, the MV is removed from the candidate list and the MV is not selected as the initial MV, thereby ensuring that the search starting point for the MV is within the corresponding MV prediction range.

[0170] FIG. 14 is a schematic diagram illustrating a current video block located at a boundary of a first display area in accordance with at least one embodiment of the present disclosure.

[0171] For example, in at least one embodiment of the present disclosure, in the process of selecting an initial MV and determining a search starting point, regardless of whether inter prediction is performed using the Merge method or the AMVP method, motion information of reference blocks at corresponding positions in adjacent coding frames in the time domain needs to be used when creating a temporal domain candidate list. As shown in Figure 14, in building the temporal domain candidate list, motion information of reference block H is usually used, and if reference block H is unavailable, it is replaced with reference block C.

[0172] 14, when the current video block is at the right boundary of the first display area basic_tile, the motion information of the reference block located at position H cannot be obtained (because it has not been decoded). Therefore, in the encoding process, a very large error value may be assigned to the motion information of reference block H, which prevents the motion information from becoming the optimal candidate MV, that is, the motion information of reference block C is selected as the temporal domain candidate list.

[0173] For example, in at least one embodiment of the present disclosure, a current video frame includes a first display region and at least one display sub-region laid out adjacently in order along a left-to-right expansion direction, and in response to the current video block being located within the first display sub-region to the right of the first display region in the current video frame, a motion vector of the current video block is constrained within a prediction range of a second motion vector.

[0174] For example, in at least one embodiment of the present disclosure, the prediction range of the second motion vector is determined based on the position of the current video block, the prediction accuracy of the motion vector, the boundary of the first display region, and the width of the first display sub-region.

[0175] For example, in at least one embodiment of the present disclosure, as shown in Figure 12A, the current video frame includes a first display region basic_tile on the left and at least one display sub-region enhanced_tile on the right, where each display sub-region enhanced_tile has a width equal to or greater than the width of one CTU.

[0176] For example, in some examples, the left boundary of the MV prediction range (second MV prediction range) of the display sub-region enhanced_tile is the left boundary of the basic_tile of the current video frame, and the right boundary of the second MV prediction range is the right boundary of the enhanced_tile adjacent to the enhanced_tile to the left of the enhanced_tile in which the current video block is located. For example, if there is no other enhanced_tile to the left of the enhanced_tile in which the current video block is located, the right boundary of the second MV prediction range is the right boundary of the basic_tile.

[0177] For example, in at least one embodiment of the present disclosure, the MV prediction range of the current video block located in the first display sub-area enhanced_tile to the right of the first display area basic_tile, i.e., the second MV prediction range, is determined by the following equations (7) to (12).

[0178]

number

[0179] In equations (7) to (12), x and y represent the position of the current video block (e.g., PU), and the four boundaries of the first display area basic_tile are respectively:

number

number

number

number

number

number

number

number

[0180] Similar to equations (5) and (6), for a current video block located within the kth enhanced_tile in the left-to-right direction to the right of the first display area basic_tile, all MVs associated with the current video block should satisfy the constraints of equations (11) and (12).

[0181]

number

[0182] In equations (10) and (11),

number

number

[0183] For example, in at least one embodiment of the present disclosure, in response to the current video block being located within the kth display sub-area to the right of the first display area and k=1, the second MV prediction range is equal to the first MV prediction range.

[0184] For example, in at least one embodiment of the present disclosure, in response to the current video block being located within the kth display sub-area to the right of the first display area and k being an integer greater than 1, the first right boundary of the first MV prediction range is different from the second right boundary of the second MV prediction range.

[0185] It should be noted that in the embodiments of the present disclosure, the "first right boundary" refers to the right boundary of the first MV prediction range, and the "second right boundary" refers to the right boundary of the second MV prediction range. The "first right boundary" and the "second right boundary" are not limited to any particular boundary or boundaries, nor are they limited to a particular order.

[0186] For example, as can be seen from the above equations (8) and (2), when k=1, the right boundary of the second MV prediction range

number

number

number

number

[0187] For example, in at least one embodiment of the present disclosure, a current video frame includes a first display region and at least one display sub-region laid out adjacently in order along a top-to-bottom expansion direction, and in response to the current video block being located within a second display sub-region below the first display region, a motion vector of the current video block is constrained within a prediction range of a third motion vector.

[0188] For example, in at least one embodiment of the present disclosure, the prediction range of the third motion vector is determined based on the position of the current video block, the prediction accuracy of the motion vector, the boundary of the first display region, and the height of the second display sub-region.

[0189] For example, in at least one embodiment of the present disclosure, as shown in Figure 12B, the current video frame includes a first display region basic_tile above and at least one display sub-region enhanced_tile below, where each display sub-region enhanced_tile has a height equal to or greater than the height of one CTU.

[0190] For example, in at least one embodiment of the present disclosure, the MV prediction range of the current video block located in the second display sub-area enhanced_tile below the first display area basic_tile, i.e., the third MV prediction range, is determined by the following equations (13) to (16):

[0191]

number

[0192] For example, the MV prediction range of the current video block located in the m-th enhanced_tile in the top-to-bottom direction below the basic_tile, i.e., the third MV prediction range, is the left boundary defined by Equations (13)-(16).

number

number

number

number

number

number

number

number

[0193] Similar to equations (11) and (12), when the current video block is located in the mth enhanced_tile below the first display area basic_tile, all MVs associated with the current video block should satisfy the constraints of equations (17) and (18).

[0194]

number

[0195] In equations (17) and (18),

number

number

[0196] For example, in at least one embodiment of the present disclosure, in response to the current video block being located within the mth second display sub-area below the first display area and m=1, the prediction range of the third motion vector is equal to the prediction range of the first motion vector.

[0197] For example, in at least one embodiment of the present disclosure, in response to the current video block being located within the mth second display sub-area below the first display area and m being an integer greater than 1, the first lower boundary of the prediction range of the first motion vector is different from the third lower boundary of the prediction range of the third motion vector.

[0198] It should be noted that in the embodiments of the present disclosure, the "first lower boundary" is used to indicate the lower boundary of the first MV prediction range, and the "third lower boundary" is used to indicate the lower boundary of the third MV prediction range. The "first lower boundary" and the "third lower boundary" are not limited to any particular boundary or boundaries, nor are they limited to a particular order.

[0199] For example, as can be seen based on the above equations (4) and (16), when m=1, the lower boundary of the third MV prediction range

number

number

number

number

[0200] For example, in at least one embodiment of the present disclosure, in response to the current video block being located outside the first display region or the base region of the current video frame, a predicted value of a temporal domain candidate motion vector in the motion vector prediction candidate list is calculated using a predicted value of a spatial domain candidate motion vector.

[0201] For example, in some examples, when the current video block is located outside the first display area / base area (e.g., located to the right or below the first display area), the motion information of reference block H and reference block C cannot be obtained in the process of constructing the temporal domain candidate list, regardless of whether Merge or AMVP mode is used. For example, in some examples, the motion vector proportional scaling MV of the first reference block is selected according to the order of the spatial domain candidate list and added to the temporal domain candidate list, as shown in the following equation (19):

[0202]

number

[0203]

number

[0204] It is important to note that the embodiments of the present disclosure do not specifically limit which predicted value of the spatial domain candidate motion vector is used to replace the predicted value of the temporal domain candidate motion vector, and can be set according to actual needs.

[0205] For example, in at least one embodiment of the present disclosure, performing the conversion between the current video block of the current video frame and the video bitstream may include a decoding process. For example, in some examples, the entire received bitstream is decoded for display. Furthermore, for example, in some examples, when a display terminal performs only partial display, for example, when the display terminal has a sliding scroll screen as shown in FIG. 1, only a partial decoding of the received bitstream is required, thereby reducing the use of decoding resources and improving video coding efficiency.

[0206] FIG. 15 is a schematic diagram of another method for processing video data in accordance with at least one embodiment of the present disclosure.

[0207] For example, at least one embodiment of the present disclosure provides another video data processing method 30. The video data processing method 30 can be applied to various application scenarios related to video decoding (i.e., applied to the decoding side). For example, as shown in FIG. 15, the video data processing method 30 includes the following operations S301 to S303.

[0208] Step S301: Receive a video bitstream.

[0209] Step S302: Determine that a current video block of a video is coded using a first inter-prediction mode.

[0210] Step S303: Based on the determination, decode the bitstream, and derive, in a first inter prediction mode, a motion vector of the current video block based on a base region corresponding to a first display mode in the video.

[0211] For example, in at least one embodiment of the present disclosure, a decoding side can determine whether the video applies a first display mode and the corresponding video expansion direction based on a received video bitstream. For example, in some examples, when the received bitstream includes a grammar element "enhanced_tile_enabled_hor" (or the value of the grammar element is 1), the current video applies the first display mode and the expansion direction is horizontal (e.g., from left to right). Furthermore, in some other examples, when the received bitstream includes a grammar element "enhanced_tile_enabled_ver" (or the value of the grammar element is 1), the current video applies the first display mode and the expansion direction is vertical (e.g., from top to bottom).

[0212] It should be noted that in the embodiments of the present disclosure, the application of the first display mode during the decoding process is not only determined based on the relevant grammar elements in the received bitstream, but also takes into consideration the actual situation of the display terminal. For example, in some cases, when the video display mode of the display terminal does not match the first display mode indicated in the bitstream, the first display mode is not applied. For example, when the relevant grammar elements in the bitstream indicate that the current video is to be applied in the first display mode and that the expansion direction is horizontal, and at the same time, when the video display mode of the display terminal is to be vertically scrolled, it is determined that the first display mode is not applied to the current video. Furthermore, when the relevant grammar elements in the bitstream indicate that the current video is to be applied in the first display mode and that the expansion direction is horizontal, and at the same time, when the video display mode of the display terminal is to be a general display (e.g., full-screen display) and no scroll expansion is required, it is determined that the first display mode is not applied to the current video. The embodiments of the present disclosure are not specifically limited to this, and can be configured according to actual situations.

[0213] For example, in at least one embodiment of the present disclosure, for step S303, decoding the bitstream includes determining a target region for decoding of the current video frame, and decoding the bitstream based on the target region for decoding, where the target region for decoding includes at least a first display region corresponding to the base region.

[0214] For example, in at least one embodiment of the present disclosure, as shown in Figure 11, encoding a current video frame mainly includes encoding a base region / first display region basic_tile and encoding at least one display sub-region enhanced_tile that are adjacently laid out in left-to-right order. For example, the basic_tile region of each video frame in a video needs to be decoded, and the number of enhanced_tile regions to be decoded can be determined based on the size of the display region and the number of enhanced_tile regions to be decoded in the previous frame. For example, in some examples, the current video frame can be decoded by at most one enhanced_tile more than the corresponding reference frame.

[0215] For example, in at least one embodiment of the present disclosure, only the display area to be displayed needs to be decoded, and therefore, the decoding target area needs to be limited in the decoding process. For example, in the embodiment of the present disclosure, the decoding target area is limited by the boundary of the coding unit, or the coding unit is used as the unit of the decoding target area. It is necessary to explain that in the embodiment of the present disclosure, the coding unit is a coding tree unit (CTU) as an example.

[0216] For example, in at least one embodiment of the present disclosure, the number of pixels to be displayed in the current video frame (

number

number

number

number

number

number

[0217] For example, in at least one embodiment of the present disclosure, the number of CTUs to be displayed in the current video frame (

number

number

number

number

number

number

number

number

number

[0218] For example, in at least one embodiment of the present disclosure, the current video frame decodes at most one display sub-region enhanced_tile more than the previous video frame, ie, one new display sub-region enhanced_tile more.

[0219] For example, in at least one embodiment of the present disclosure, the number of CTUs to be displayed in the current video frame (

number

number

number

number

number

number

number

number

number

number

number

number

[0220] For example, in at least one embodiment of the present disclosure, the decoding target region (

number

[0221]

number

[0222] for example,

number

number

number

number

[0223] For example, in at least one embodiment of the present disclosure, the decoding target area of ​​the current video frame includes only a fixedly displayed display area (e.g., a first display area), i.e.

number

number

number

number

number

number

number

number

[0224] For example, in at least one embodiment of the present disclosure, CTUs in areas that do not need to be displayed can be directly filled with 0 pixels without being decoded, thereby improving video coding efficiency, simplifying the coding process, and saving product energy.

[0225] It is worth noting that in the embodiments of the present disclosure, the CTUs in the areas that do not need to be displayed can be filled with other pixels, and are not necessarily 0 pixels, but can be set according to actual needs.

[0226] For example, in at least one embodiment of the present disclosure, when the display terminal is in a state where the entire screen is scrolled, that is, when there is no area that needs to be displayed, the decoding target area (

number

number

[0227] For example, in at least one embodiment of the present disclosure, a general description of the video coding system in Figure 16 may refer to the related descriptions of Figures 2-4, and a detailed description thereof will be omitted here. In an embodiment of the present disclosure, for an encoding process in inter prediction mode, the range of a motion vector associated with a current video block is restricted to avoid the use of invalid reference pixel information. For a decoding process, a bitstream can be partially decoded based on an area that is actually displayed, thereby reducing the consumption of decoding resources for non-displayed portions and improving coding efficiency.

[0228] FIG. 17 is a simplified flowchart of a method for processing video data in an LDP configuration in accordance with at least one embodiment of the present disclosure.

[0229] For example, in at least one embodiment of the present disclosure, a video data processing method is provided, as shown in Figure 17. The video data processing method includes steps S201-S206.

[0230] Step S201: At the encoding side, divide the current video frame into basic_tile and enhanced_tile. For example, in the example shown in FIG. 12A, the encoder divides a frame image into two tiles, left and right, with the left being the basic_tile and the right being at least one enhanced_tile. For example, in some examples, allocation is performed in a 1:1 ratio according to the number of CTUs in a row. For example, in other examples, allocation is performed in a 1:2 ratio according to the number of CTUs in a row. Furthermore, for example, when the number of CTUs in a row is odd, the total number of CTUs in the multiple enhanced_tiles in each row is one more than the number of CTUs in a row of basic_tile. The embodiments of the present disclosure do not limit the specific division scheme, and it can be set according to actual needs.

[0231] Step S202: For the left basic_tile, use an independent tile coding method that removes right-side coupling, and for the right enhanced_tile, use a method that depends on the direction of the left basic_tile and restrict the coding method and range of MV. For example, this step involves modifying the initial MV selection process and motion search algorithm in the inter-encoded AMVP and Merge process, so that the bitstream can meet the needs of scroll development and scrolling during decoding.

[0232] Step S203: The number of pixels not yet scrolled on the slide scroll screen at the current time

number

number

[0233] Step S204: The decoding side receives a video bitstream (not limited to H.264 / H.265 / H.266).

[0234] Step S205: Number of CTUs in the decoding area of ​​the previous frame

number

[0235] Step S206: The number of unscrolled pixels on the slide scroll screen at the current time and the previous frame

number

number

number

[0236] Step S207: The area to be decoded is decoded, and the undecoded area is filled.

[0237] Step S208: The content of the decoding target area is transmitted to the display terminal and displayed.

[0238] It is necessary to explain that the specific operations of steps S201 to S206 shown in FIG. 17 have been described in detail above, and therefore will not be described again here.

[0239] Therefore, a video data processing method according to at least one embodiment of the present disclosure can partially decode a video bitstream based on the region to be displayed, thereby reducing the consumption of decoding resources for the non-displayed portion and improving coding efficiency.

[0240] It should be noted that in the embodiments of the present disclosure, the order in which the steps of the video data processing method 10 are performed is not limited, and although the steps are described above in a specific order, this is not a limitation of the embodiments of the present disclosure. The steps of the video data processing method 10 can be performed serially or in parallel, which can be determined according to actual needs. For example, the video data processing method 10 may include more or fewer steps, and the embodiments of the present disclosure are not limited thereto.

[0241] FIG. 18 is a schematic block diagram of a video data processing device in accordance with at least one embodiment of the present disclosure.

[0242] For example, at least one embodiment of the present disclosure provides a video data processing device 40, as shown in FIG. 18 . The video data processing device 40 includes a determination module 401 and an execution module 402. The determination module 401 is configured to determine to code a current video block of a video using a first inter-prediction mode. For example, the determination module 401 may implement step S101, and specific implementation methods thereof may refer to the related description of step S101, and detailed descriptions thereof will be omitted here. The execution module 402 is configured to perform conversion between the current video block and the video bitstream based on the determination, and in the first inter-prediction mode, derive a motion vector for the current video block based on a base region corresponding to a first display mode in the video. For example, the execution module 402 may implement step S102, and specific implementation methods thereof may refer to the related description of step S102, and detailed descriptions thereof will be omitted here.

[0243] It should be noted that the determination module 401 and the execution module 402 can be implemented by software, hardware, firmware, or any combination thereof, for example, as a determination circuit 401 and an execution circuit 402, respectively, and the embodiments of the present disclosure do not limit these specific embodiments.

[0244] It should be understood that the video data processing device 40 according to at least one embodiment of the present disclosure can achieve technical effects similar to those of the above-mentioned video data processing method 10. For example, in the video data processing device 40 according to at least one embodiment of the present disclosure, the above-mentioned method can partially decode a bitstream based on the area that actually needs to be displayed, thereby reducing the consumption of decoding resources for the non-displayed portion and improving coding efficiency.

[0245] However, in the embodiments of the present disclosure, the video data processing device 40 may include more or fewer circuits or units, and the connection relationships between each circuit or unit are not limited and can be determined according to actual needs. The specific configuration of each circuit is not limited and may be configured by analog devices, digital chips, or other applicable methods based on circuit principles.

[0246] For example, at least one embodiment of the present disclosure further provides a display device including a video data processing device and a sliding scrolling screen. The video data processing device is configured to decode a received bitstream based on the method according to the at least one embodiment and transmit the decoded pixel values ​​to the sliding scrolling screen for display. For example, in some examples, when the sliding scrolling screen is fully unfolded and there is no non-display area, the video data processing device fully decodes the received bitstream. Furthermore, in some other examples, for example, when the sliding scrolling screen includes a scrolling portion and an unfolded portion, i.e., there is a display area and a non-display area, as shown in FIG. 1, the video data processing device partially decodes the received bitstream.

[0247] For example, in at least one embodiment of the present disclosure, in response to the sliding scrolling screen including a display area and a non-display area in operation, the video data processing device decodes the bitstream based on the size of the display area at the current time point and the previous frame time point. For example, as shown in FIG. 1, when the sliding scrolling screen is in a partially expanded state, the video data processing device only needs to decode content corresponding to the display area. For example, the video data processing device may determine the decoding target area based on the size of the display area at the current time point and the previous frame time point. For example, the video data processing device may determine the number of pixels to be displayed in the video frames at the current time point and the previous frame time point based on the size of the display area at the current time point and the previous frame time point.

number

number

[0248] For example, in at least one embodiment of the present disclosure, the display device further includes a curl state determination device in addition to the video data processing device and the sliding scroll screen. For example, the curl state determination device is configured to detect the size of the display area of ​​the sliding scroll screen and send the size of the display area to the video data processing device, so that the video data processing device decodes the bitstream based on the size of the display area at the current time point and the previous frame time point. It should be noted that the curl state determination device can be realized by software, hardware, firmware, or any combination thereof, and may be realized, for example, as a curl state determination circuit, and the embodiments of the present disclosure do not limit the specific embodiment of the curl state determination device.

[0249] It should be noted that the embodiments of the present disclosure do not limit the type of display device. For example, the display device may be a mobile terminal, a computer, a tablet computer, a smart watch, a television, etc., and the embodiments of the present disclosure are not limited thereto. Similarly, the embodiments of the present disclosure do not limit the type of slide scroll screen. For example, in the embodiments of the present disclosure, the slide scroll screen may be any type of display screen with a variable display area, including but not limited to the type of slide scroll screen shown in FIG. 1. For example, in the embodiments of the present disclosure, the video data processing device included in the display device may be implemented as the video data processing device 40 / 90 / 600, etc., referred to in the present disclosure, and the embodiments of the present disclosure are not limited to the specific embodiment of the video data processing device.

[0250] However, in the embodiments of the present disclosure, the display device may include more or fewer circuits or units, and the connection relationships between each circuit or unit are not limited and can be determined according to actual needs. The specific configuration manner of each circuit is not limited and may be configured by analog devices, digital chips, or other applicable manners based on circuit principles.

[0251] FIG. 19 is a schematic block diagram of another video data processing apparatus in accordance with at least one embodiment of the present disclosure.

[0252] At least one embodiment of the present disclosure further provides a video data processing device 90. As shown in FIG. 19 , the video data processing device 90 includes a processor 910 and a memory 920. The memory 920 includes one or more computer program modules 921. The one or more computer program modules 921 are stored in the memory 920 and configured to be executed by the processor 910, the one or more computer program modules 921 including instructions for performing the video data processing method 10 according to at least one embodiment of the present disclosure, and when executed by the processor 910, can perform one or more steps of the video data processing method 10 according to at least one embodiment of the present disclosure. The memory 920 and the processor 910 can be interconnected by a bus system and / or other type of connection mechanism (not shown).

[0253] For example, the processor 910 may be a central processing unit (CPU), a digital signal processor (DSP), or other type of processing unit having data processing and / or program execution capabilities, such as a field programmable gate array (FPGA), etc. For example, the central processing unit (CPU) may have an X86 or ARM architecture, etc. The processor 910 may be a general-purpose processor or a special-purpose processor, and may control other assemblies in the video data processing device 90 to perform desired functions.

[0254] For example, the memory 920 may include any combination of one or more computer program products, which may include various types of computer-readable storage media, such as volatile and / or non-volatile memory. Volatile memory may include, for example, random access memory (RAM) and / or cache memory. Non-volatile memory may include, for example, read-only memory (ROM), a hard disk, erasable programmable read-only memory (EPROM), portable compact disc read-only memory (CD-ROM), USB memory, flash memory, etc. One or more computer program modules 921 may be stored in the computer-readable storage medium, and various functions of the video data processing device 90 may be realized by the processor 910 operating the one or more computer program modules 921. The computer-readable storage medium may also store various application programs, various data, and various data used and / or generated by the application programs. For specific functions and technical effects of the video data processing device 90, please refer to the above description of the video data processing method 10 / 30, and detailed description thereof will be omitted here.

[0255] FIG. 20 is a schematic block diagram of yet another video data processing apparatus in accordance with at least one embodiment of the present disclosure.

[0256] The terminal devices in the embodiments of the present disclosure may include, but are not limited to, mobile terminals such as mobile phones, notebook computers, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), in-vehicle terminals (e.g., in-vehicle navigation terminals), and fixed terminals such as digital TVs and desktop computers. The video data processing device 600 shown in Fig. 20 is merely an example and should not limit the functions and scope of use of the embodiments of the present disclosure in any way.

[0257] 20, in some examples, a video data processing device 600 includes a processing device (e.g., a central processor, a graphics processor, etc.) 601, which may perform various appropriate operations and processes based on programs stored in a read-only memory (ROM) 602 or programs loaded from a storage device 608 into a random access memory (RAM) 603. The RAM 603 further stores various programs and data necessary for the operation of the computer system. The processing device 601, the ROM 602, and the RAM 603 are connected to each other via a bus 604. An input / output (I / O) interface 605 is also connected to the bus 604.

[0258] Input devices 906, including, for example, a touch panel, touch pad, keyboard, mouse, camera, microphone, accelerometer, gyroscope, etc.; output devices 607, including, for example, a liquid crystal display (LCD), speaker, vibrator, etc.; storage devices 608, including, for example, a magnetic tape, hard disk, etc.; and communication devices 609, including, for example, a network interface card such as a LAN card or modem, can be connected to the I / O interface 605. The communication devices 609 allow the video data processing device 600 to exchange data with other devices via wireless or wired communication or to perform communication processing via an Internet network. A driver 610 is also connected to the I / O interface 605 as needed. A removable medium 611, such as a magnetic disk, optical disk, magneto-optical disk, semiconductor memory, etc., is installed in the driver 610 as needed, and a computer program read from the removable medium 611 is installed in the storage device 608 as needed. While FIG. 20 shows the video data processing device 600 including various devices, it is not required to embody or include all of the devices shown. It should be understood that more or fewer devices may alternatively be embodied or included.

[0259] For example, the video data processing device 600 may further include a peripheral interface (not shown), etc. The peripheral interface may be various types of interfaces, such as a USB interface, a Lightning interface, etc. The communication device 609 can communicate with a network and other devices by wireless communication, where the network is, for example, a wireless network such as the Internet, an intranet, and / or a cellular phone network, a wireless local area network (LAN), and / or a metropolitan area network (MAN). Wireless communication may use any of a number of communication standards, protocols, and technologies, including, but not limited to, Global System for Mobile Communications (GSM), Enhanced Data GSM Environment (EDGE), Wideband Code Division Multiple Access (W-CDMA), Code Division Multiple Access (CDMA), Time Division Multiple Access (TDMA), Bluetooth, Wi-Fi (e.g., based on the IEEE 802.11a, IEEE 802.11b, IEEE 802.11g, and / or IEEE 802.11n standards), Voice over Internet Protocol (VoIP), Wi-MAX, email, protocols for instant messaging and / or short message service (SMS), or any other suitable communication protocol.

[0260] For example, the video data processing device 600 may be any device such as a mobile phone, a tablet computer, a laptop, an e-book, a television, etc., or may be any combination of data processing devices and hardware, and the embodiments of the present disclosure are not limited thereto.

[0261] For example, according to embodiments of the present disclosure, the processes described with reference to the flowcharts above may be implemented as a computer software program. For example, embodiments of the present disclosure include a computer program product including a computer program embodied on a non-transitory computer-readable medium, the computer program including program code for performing the methods illustrated in the flowcharts. In such embodiments, the computer program may be downloaded and installed from a network via the communication device 609, or installed from the storage device 608, or installed from the ROM 602. When the computer program is executed by the processing device 601, it performs the video data processing method 10 disclosed in the embodiments of the present disclosure.

[0262] It should be noted that the computer-readable medium of the present disclosure may be a computer-readable signal medium, a computer-readable storage medium, or any combination of the above. The computer-readable storage medium may be, but is not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of the computer-readable storage medium may include, but are not limited to, an electrical connection having one or more conductors, a portable computer magnetic disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In embodiments of the present disclosure, the computer-readable storage medium may be any tangible medium that contains or stores a program, which may be used by or in combination with an instruction execution system, apparatus, or device. In embodiments of the present disclosure, the computer-readable signal medium may include a data signal propagating in baseband or as part of a carrier wave, and which includes computer-readable program code. Such propagated data signals may take various forms, including, but not limited to, electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium, which is capable of transmitting, propagating, or transmitting a program for use by or in connection with an instruction execution system, apparatus, or device. Program code contained in a computer-readable medium may be transmitted by any suitable medium, including, but not limited to, electrical wire, optical cable, RF (radio frequency), etc., or any suitable combination of the above.

[0263] The computer-readable medium may be included in the video data processing device 600 or may exist independently of the video data processing device 600 .

[0264] FIG. 21 is a schematic block diagram of a non-transitory readable storage medium in accordance with at least one embodiment of the present disclosure.

[0265]

[0023] An embodiment of the present disclosure further provides a non-transitory readable storage medium. Figure 21 is a schematic block diagram of a non-transitory readable storage medium according to at least one embodiment of the present disclosure. As shown in Figure 21, computer instructions 111 are stored on the non-transitory readable storage medium 70, and when executed by a processor, the computer instructions 111 perform one or more steps of the video data processing method 10 described above.

[0266] For example, the non-transitory readable storage medium 70 may be any combination of one or more computer-readable storage media, such as one computer-readable storage medium including computer-readable program code for obtaining a first display mode of a video, another computer-readable storage medium including computer-readable program code for determining a first subset of a current plurality of available inter-prediction modes for a current video block of the video, and yet another computer-readable storage medium including computer-readable program code for selecting a first available inter-prediction mode from the first subset and performing a conversion between a current video block of a current video frame and a bitstream of the video, wherein a usable effective range of a motion vector used by each member of the first subset is determined based on a base region corresponding to the first display mode of the video. Of course, the above program codes may be stored on the same computer-readable medium, and the embodiments of the present disclosure are not limited thereto.

[0267] For example, when the program code is read by a computer, the computer can execute the program code stored in the computer storage medium, for example, to perform the video data processing method 10 according to any embodiment of the present disclosure.

[0268] For example, the storage medium may include a memory card of a smartphone, a storage member of a tablet computer, a hard disk of a personal computer, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM), a portable compact disk read-only memory (CD-ROM), a flash memory, or any combination of the above storage media, or other applicable storage media. For example, the readable storage medium may be the memory 920 in FIG. 19, the relevant description of which can be referred to above, and detailed description thereof will be omitted here.

[0269] In this disclosure, unless expressly specified and limited otherwise, the term "plurality" refers to two or more than two.

[0270] Those skilled in the art will readily appreciate other implementation solutions of the present disclosure after considering the specification and practicing the disclosure disclosed herein. The present disclosure is intended to cover any variations, uses, or adaptations of the present disclosure, which variations, uses, or adaptations comply with the general principles of the present disclosure and include common general knowledge or customary techniques in the art that are not disclosed in the present disclosure. It is intended that the specification and examples be considered as exemplary only, with the true scope and spirit of the present disclosure being indicated by the following claims.

[0271] It should be understood that the present disclosure is not limited to the exact constructions previously described above and illustrated in the drawings, and various modifications and variations can be made without departing from the scope thereof, which is limited only by the appended claims.

Claims

1. 1. A method for processing video data, comprising: determining, for a current video block of a video, to code using a first inter-prediction mode; performing a conversion between the current video block and the video bitstream based on the determination; deriving a motion vector for the current video block in the first inter-prediction mode based on a base region of the video corresponding to a first display mode.

2. The method of claim 1 , wherein a first display area along an expansion direction from a video expansion start position defined by the first display mode is set as the base area for the current video frame of the video.

3. 3. The method of claim 2, wherein, in response to the current video block being located within the first display area, a motion vector of the current video block is within a prediction range of a first motion vector.

4. The method of claim 3 , wherein the prediction range of the first motion vector is determined based on a position of the current video block, a prediction accuracy of the motion vector, and a boundary of the first display area.

5. the current video frame includes the first display area and at least one display sub-area laid out adjacently in order along the expansion direction from left to right; 5. The method of claim 3, wherein, in response to the current video block being located within a first display sub-area to the right of the first display area in the current video frame, a motion vector of the current video block is within a prediction range of a second motion vector.

6. 6. The method of claim 5, wherein the prediction range of the second motion vector is determined based on the position of the current video block, a prediction accuracy of the motion vector, a boundary of the first display region, and a width of the first display sub-region.

7. 7. The method of claim 5, wherein, in response to the current video block being located within a kth first display sub-area to the right of the first display area and k=1, a prediction range of the second motion vector is equal to a prediction range of the first motion vector.

8. 8. The method of claim 5, wherein, in response to the current video block being located within a kth first display sub-area to the right of the first display area, and k is an integer greater than 1, a first right boundary of a prediction range of the first motion vector is different from a second right boundary of a prediction range of the second motion vector.

9. the current video frame includes the first display area and at least one display sub-area laid out adjacent to each other in order along the expansion direction from top to bottom; 9. The method of claim 3, wherein, in response to the current video block being located within a second display sub-area below the first display area, a motion vector of the current video block is within a prediction range of a third motion vector.

10. 10. The method of claim 9, wherein the prediction range of the third motion vector is determined based on the position of the current video block, a prediction accuracy of the motion vector, a boundary of the first display region, and a height of the second display sub-region.

11. 11. The method of claim 9 or 10, wherein, in response to the current video block being located within an m-th second display sub-area below the first display area and m=1, a prediction range of the third motion vector is equal to a prediction range of the first motion vector.

12. 12. The method of claim 9, wherein, in response to the current video block being located within an m-th second display sub-area below the first display area and m being an integer greater than 1, a first lower boundary of a prediction range of the first motion vector is different from a third lower boundary of a prediction range of the third motion vector.

13. 13. The method of claim 2, wherein, in response to the current video block being located outside the base region, a predicted value of a temporal domain candidate motion vector in a prediction candidate list of motion vectors for the current video block is calculated based on a predicted value of a spatial domain candidate motion vector.

14. 14. The method of claim 2, wherein, in response to the current video block being located within the base region, all reference pixels used by the current video block are within the base region.

15. The method according to any one of claims 1 to 14, wherein the first inter prediction mode comprises a merge prediction mode, an advanced motion vector prediction (AMVP) mode, a merge mode with motion vector difference, a bidirectional weighted prediction mode or an affine prediction mode.

16. 1. A method for processing video data, comprising: receiving a video bitstream; determining that a current video block of the video is to be coded using a first inter-prediction mode; and decoding the bitstream based on the determination; deriving a motion vector for the current video block in the first inter-prediction mode based on a base region of the video corresponding to a first display mode.

17. The step of decoding the bitstream comprises: determining a current video frame of the video to be decoded, the current video frame including at least a first display region corresponding to the base region; and decoding the bitstream based on the region to be decoded.

18. The step of determining a decoding target region includes:

18. The method of claim 17, comprising determining the area to be decoded based on at least one of a number of pixels and a number of coding units to be displayed of the current video frame, a number of coding units of the first display area, a number of displayed pixels of a previous video frame, and a number of coding units of a decoded area of ​​the previous video frame.

19. The step of determining a decoding target region includes: in response to the quantity of coding units to be displayed of the current video frame being greater than the quantity of coding units of the decoded region of the previous video frame; or In response to a quantity of coding units to be displayed of the current video frame being equal to a quantity of coding units of a decoded region of a previous video frame, and a quantity of pixels to be displayed of the current video frame being greater than a quantity of displayed pixels of the previous video frame, 20. The method of claim 18, comprising determining that the region to be decoded includes a decoded region of the previous video frame and one new display sub-region.

20. The step of determining a decoding target region includes: in response to a quantity of coding units to be displayed of the current video frame being greater than a quantity of coding units in the first display area and a quantity of coding units to be displayed of the current video frame being less than a quantity of coding units in a decoded area of ​​a previous video frame; or in response to a quantity of coding units to be displayed of the current video frame being greater than a quantity of coding units in the first display area, a quantity of coding units to be displayed of the current video frame being equal to a quantity of coding units in the decoded area of ​​the previous video frame, and a quantity of pixels to be displayed of the current video frame being less than or equal to a quantity of displayed pixels of the previous video frame; 20. A method according to claim 18 or 19, comprising determining that a region of the current video frame to be decoded includes a region of the current video frame to be displayed.

21. 1. A video data processing device, comprising: a decision module configured to decide to code a current video block of the video using a first inter-prediction mode; an execution module configured to perform a conversion between the current video block and the video bitstream based on the determination; A video data processing apparatus that, in the first inter-prediction mode, derives a motion vector for the current video block based on a base region of the video corresponding to a first display mode.

22. A display device, comprising: a video data processing device; and a sliding scroll screen; A display device, wherein the video data processing device is configured to decode the received bitstream according to the method of any one of claims 1 to 20, and to transmit the decoded pixel values ​​to the sliding scrolling screen for display.

23. 23. The display device of claim 22, wherein in response to the sliding scrolling screen including a display area and a non-display area in operation, the video data processing device decodes the bitstream based on the size of the display area at a current time point and a previous frame time point.

24. 24. The display device of claim 22 or 23, further comprising a curl state determination device configured to detect the size of a display area of ​​the sliding scroll screen and transmit the size of the display area to the video data processing device, so that the video data processing device decodes the bitstream based on the size of the display area at a current time point and a previous frame time point.

25. 1. A video data processing device, comprising: a processor; a memory containing one or more computer program modules; A video data processing device, wherein the one or more computer program modules are stored in the memory and configured to be executed by the processor, the one or more computer program modules including instructions for performing the video data processing method of any one of claims 1 to 20.

26. A computer readable storage medium having stored thereon computer instructions which, when executed by a processor, implement the steps of the video data processing method of any one of claims 1 to 20.