A candidate motion information list determination method and device, electronic equipment and storage medium
Patent Information
- Application Number
- CN202610879016.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2020-10-18
- Publication Date
- 2026-09-11
AI Technical Summary
[0004]但是,当运动信息列表中包含的位移矢量不足时,会导致该运动信息列表无法提供有效的预测位移矢量,这会影响到视频的压缩性能,无法有效减少经过压缩的视频的体积,不利于用户的使用体验
本申请实施例通过确定针对当前编解码块的历史运动信息表中候选位移矢量的数量;当历史运动信息表中候选位移矢量的数量大于等于两个时,触发位移矢量预测导出预测位移矢量,从历史运动信息表获取至少一个运动信息,并基于从历史运动信息表中获取的至少一个运动信息对候选运动信息列表进行填充;当候选运动信息列表未填充完整时,构建空域运动信息列表,空域运动信息列表中包括当前编解码块的空域临近块的运动信息,空域临近块至少包括帧内串复制ISC块;从空域运动信息列表中获取至少一个运动信息,并基于从空域运动信息列表中获取的至少一个运动信息对候选运动信息列表进行填充,候选运动信息列表用于为当前编解码块提供候选的预测位移矢量,由此,可以实现在该候选运动信息列表中提供更多更有效的位移矢量,以取得更好的位移矢量预测效果,提升视频的压缩性能,提升用户的使用体验。
Smart Images

Figure CN122741702A_ABST
Abstract
Description
[0001] Case Analysis This application is a divisional application of Chinese Patent Application No. 202011114059.5, filed on October 18, 2020, entitled "A method, apparatus, electronic device and storage medium for determining a candidate motion information list". Technical Field
[0002] This application relates to video encoding and decoding technology, and more particularly to a method, apparatus, electronic device, and storage medium for determining a list of candidate motion information. Background Technology
[0003] In existing technologies for video compression processing, such as VVC (Versatile Video Coding) and AVS3 (Audio Video Coding Standard 3), video codecs typically need to construct a list of motion information to derive the predicted displacement vector.
[0004] However, when the motion information list contains insufficient displacement vectors, it will be unable to provide effective predicted displacement vectors. This will affect the compression performance of the video, making it impossible to effectively reduce the size of the compressed video, which is detrimental to the user experience. Summary of the Invention
[0005] In view of this, embodiments of this application provide a method, apparatus, electronic device, and storage medium for determining a candidate motion information list, which can achieve better displacement vector prediction results and improve video compression performance.
[0006] The technical solution of this application embodiment is implemented as follows: This application provides a method for determining a candidate motion information list, including: Determine a historical motion information table for the current codec block, wherein the historical motion information table is used for at least one of inter-frame prediction, intra-frame prediction, block copy intra-frame prediction (IBC), and string copy intra-frame prediction (ISC); When the number of candidate displacement vectors in the historical motion information table is greater than or equal to two, displacement vector prediction is triggered to derive the predicted displacement vector, and the candidate motion information list is populated based on at least one motion information obtained from the historical motion information table. When the candidate motion information list is not completely filled, a spatial motion information list is constructed. The spatial motion information list includes the motion information of the spatial neighboring blocks of the current codec block. The spatial neighboring blocks include at least the intra-string copy (ISC) block. The candidate motion information list is populated based on at least one motion information obtained from the airspace motion information list; Specifically, the candidate motion information list is filled based on different filling positions in the candidate motion information list; The candidate motion information list is used to provide candidate predicted displacement vectors for the current codec block.
[0007] This application also provides a candidate motion information list determination device, including: The first information processing module is used to determine the historical motion information table for the current codec block, wherein the historical motion information table is used for at least one of inter-frame prediction, intra-frame prediction, block copy intra-frame prediction (IBC), and string copy intra-frame prediction (ISC). The first information filling module is used to obtain the number of candidate displacement vectors in the historical motion information table; When the number of candidate displacement vectors in the historical motion information table is greater than or equal to two, displacement vector prediction is triggered to derive the predicted displacement vector, and the candidate motion information list is populated based on at least one motion information obtained from the historical motion information table. The first information filling module is further configured to construct a spatial motion information list when the candidate motion information list is not completely filled. The spatial motion information list includes motion information of spatial neighboring blocks of the current codec block. The spatial neighboring blocks include at least intra-string copy (ISC) blocks. The second information filling module is used to fill the candidate motion information list based on at least one motion information obtained from the spatial motion information list; wherein, the candidate motion information list is filled based on different filling positions in the candidate motion information list; The candidate motion information list is used to provide candidate predicted displacement vectors for the current codec block.
[0008] This application also provides an electronic device, the electronic device comprising: Memory, used to store executable instructions; A processor, when executing executable instructions stored in the memory, implements a method for determining a list of any candidate motion information preceding the execution.
[0009] This application also provides a computer-readable storage medium storing executable instructions, which, when executed by a processor, implement a method for determining a list of any preceding candidate motion information.
[0010] The embodiments of this application have the following beneficial effects: This application embodiment determines the number of candidate displacement vectors in the historical motion information table for the current codec block. When the number of candidate displacement vectors in the historical motion information table is greater than or equal to two, displacement vector prediction is triggered to derive predicted displacement vectors. At least one motion information is obtained from the historical motion information table, and the candidate motion information list is populated based on the at least one motion information obtained from the historical motion information table. When the candidate motion information list is not completely populated, a spatial motion information list is constructed, which includes motion information of spatially adjacent blocks of the current codec block. The spatially adjacent blocks include at least one intra-string copy (ISC) block. At least one motion information is obtained from the spatial motion information list, and the candidate motion information list is populated based on the at least one motion information obtained from the spatial motion information list. The candidate motion information list is used to provide candidate predicted displacement vectors for the current codec block. Thus, more and more effective displacement vectors can be provided in the candidate motion information list to achieve better displacement vector prediction results, improve video compression performance, and enhance the user experience. Attached Figure Description
[0011] Figure 1 A schematic diagram illustrating the use scenario determined by the candidate motion information list provided in the embodiments of this application; Figure 2 A schematic diagram of the composition structure of the electronic device provided in the embodiments of this application; Figure 3 This is a schematic diagram of an optional video encoding process in an embodiment of this application; Figure 4 This is a schematic diagram of the inter-frame prediction mode in an embodiment of this application; Figure 5 A schematic diagram of candidate motion vectors provided in the embodiments of this application; Figure 6 A schematic diagram of the intra-block copying mode provided in an embodiment of this application; Figure 7 A schematic diagram of the intra-frame string copying mode provided in an embodiment of this application; Figure 8 An optional flowchart illustrating the method for determining a list of candidate motion information provided in an embodiment of this application; Figure 9 An optional flowchart illustrating the method for determining a list of candidate motion information provided in an embodiment of this application; Figure 10 An optional flowchart illustrating the method for determining a list of candidate motion information provided in an embodiment of this application; Figure 11 This is a schematic diagram of a spatial proximity block provided in one embodiment of this application; Figure 12A schematic diagram illustrating a use case of the candidate motion information list determination method provided in this application embodiment; Figure 13 This is a schematic diagram of an optional video compression method in an embodiment of this application; Figure 14 This is a schematic diagram of an optional compressed video presentation in an embodiment of this application; Figure 15 This is an optional flowchart illustrating the method for determining a list of candidate motion information provided in an embodiment of this application.
[0012] It should be noted that the terms "first" and "second" mentioned above are only used to distinguish between different options and do not represent the degree of superiority or inferiority of the options or their priority in the implementation process. Detailed Implementation
[0013] To make the objectives, technical solutions, and advantages of this application clearer, the application will be further described in detail below with reference to the accompanying drawings. The described embodiments should not be regarded as limitations on this application. All other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0014] In the following description, references are made to “some embodiments,” which describe a subset of all possible embodiments. However, it is understood that “some embodiments” may be the same subset or different subsets of all possible embodiments and may be combined with each other without conflict.
[0015] Before providing a further detailed description of the embodiments of this application, the nouns and terms involved in the embodiments of this application will be explained, and the nouns and terms involved in the embodiments of this application shall be interpreted as follows.
[0016] 1) API: Short for Application Programming Interface, it refers to a set of predefined functions or conventions for the connection between different components of a software system. Its purpose is to provide applications and developers with the ability to access a set of routines based on certain software or hardware, without needing to access the source code or understand the details of the internal workings.
[0017] 2) SDK: Software Development Kit, which is a collection of development tools for building application software for specific software packages, software frameworks, hardware platforms, operating systems, etc. In a broad sense, it includes a collection of related documents, examples and tools that assist in the development of a certain type of software.
[0018] 3) P-frame: Inter-frame prediction frame, which can use intra-frame prediction and inter-frame prediction, and can use forward reference prediction video coding method.
[0019] 4) B-frame: Inter-frame prediction frame, which can use intra-frame prediction and inter-frame prediction, and can use forward, backward and bi-directional reference prediction.
[0020] 5) I-frame: Intra-predictive frame, which uses intra-frame information for prediction.
[0021] 6) Video codec standard: A set of agreed-upon rules for decoding video streams.
[0022] 7) Video transcoding refers to converting a compressed video stream into another video stream to adapt to different network bandwidths, different terminal processing capabilities, and different user needs.
[0023] 8) Client: A carrier in a terminal that implements specific functions. For example, a mobile client (APP) is a carrier of specific functions in a mobile terminal, such as performing online live streaming (video streaming) or playing online videos.
[0024] 9) Responding to: used to indicate the conditions or states on which the operation performed depends. When the conditions or states on which it depends are met, one or more operations performed may be performed in real time or with a set delay. Unless otherwise specified, there is no restriction on the order in which the multiple operations are performed.
[0025] The following describes the application environment of the method for determining the candidate motion information list provided in this application. See [link to relevant documentation]. Figure 1 , Figure 1 A schematic diagram illustrating the use scenario for determining the candidate motion information list provided in this application embodiment is shown below. Figure 1The terminals (including terminals 10-1 and 10-2) are equipped with corresponding clients capable of performing different functions. These clients, on the other hand, allow terminals 10-1 and 10-2 to access and browse different video information from their respective servers 200 via network 300 using different business processes. The terminals connect to server 200 via network 300, which can be a wide area network (WAN), a local area network (LAN), or a combination of both, using a wireless link for data transmission. The types of videos accessed by the terminals 10-1 and 10-2 from server 200 via network 300 are not the same. For example, terminals 10-1 and 10-2 can access videos (i.e., videos containing video information or corresponding video links) from server 200 via network 300, or they can access videos from server 400 containing only different types of videos (e.g., short videos or long videos) for browsing. Servers 200 and 400 can store different types of videos. In some embodiments of this application, the processes for storing different types of videos in server 200 can be written in software code of different programming languages, and the code objects can be different types of code entities. For example, in C language software code, a code object can be a function. In JAVA language software code, a code object can be a class, and in iOS Objective-C, it can be a piece of object code. In C++ language software code, a code object can be a class or a function. This application does not distinguish between the compilation environments of different types of videos. However, in this process, in existing technologies such as VVC (Versatile Video Coding) and AVS3 (Audio Video Coding Standard 3), video codecs typically need to construct a list of motion information to derive the predicted displacement vector during video compression processing. However, when the motion information list contains insufficient displacement vectors, it cannot provide effective predicted displacement vectors. Specifically, during displacement vector prediction, the candidate motion information list is constructed solely by building the IntraHMVP table of intra-frame prediction history motion information, and the predicted block displacement vector (BVP) or predicted string vector (SVP) is exported. The maximum length of IntraHMVP is 12, and the maximum length of the candidate motion information list is 7. When the IntraHMVP is insufficient or empty, the candidate motion information list cannot be filled, resulting in insufficient motion information for displacement vector prediction. This will affect the video compression performance, reduce the video compression rate, and negatively impact the user experience of video compression.
[0026] Furthermore, during the process of server 200 sending or receiving different types of video to terminals (terminal 10-1 and / or terminal 10-2) via network 300, the video information occupies a large amount of storage space, thus requiring compression. Therefore, as an example, server 200 is used to determine a candidate motion information list and a historical motion information table. The historical motion information table is used for at least one of inter-frame prediction, intra-frame prediction, block copy intra-frame prediction (IBC), and string copy intra-frame prediction (ISC). It can also determine the number of candidate displacement vectors in the historical motion information table. When the candidate motion information list is not full, at least one motion information is obtained based on the historical motion information table, and the candidate motion information list is filled based on the at least one motion information. The candidate motion information list is used to provide candidate predicted displacement vectors for the current codec block. Of course, in some embodiments of this application, the server 200 can also be used to determine the positional relationship between the historical motion information table and the spatial motion information list; when the candidate motion information list is not full, at least one motion information is obtained based on the historical motion information table and / or the spatial motion information list, and the candidate motion information list is filled based on the at least one motion information; wherein, the candidate motion information list is used to provide candidate predicted displacement vectors for the current codec block. Specifically, during the video compression process, the server 200 can flexibly adjust the candidate motion information list determination process according to different usage environments or user settings. For example, it can flexibly use the intra-frame prediction historical motion information table to fill the candidate motion information list, or select other types of historical motion information tables to fill the candidate motion information list in the inter-frame prediction, block copy intra-frame prediction (IBC), and string copy intra-frame prediction (ISC) processes.
[0027] The structure of the server in this application embodiment will be described in detail below. The server can be implemented in various forms, such as a dedicated terminal with candidate motion information list determination function, such as a gateway, or a server with candidate motion information list determination function, as described above. Figure 1 Server 200 in the middle. Figure 2 This is a schematic diagram of the composition structure of the electronic device provided in the embodiments of this application. It can be understood that... Figure 2 The server architecture shown is merely an example and not the entire architecture; implementation is possible as needed. Figure 2 The structure shown may be part or all of the structure.
[0028] The server provided in this application embodiment includes at least one processor 201, a memory 202, a user interface 203, and at least one network interface 204. Various components in the electronic device are coupled together via a bus system 205. It can be understood that the bus system 205 is used to implement communication between these components. In addition to a data bus, the bus system 205 also includes a power bus, a control bus, and a status signal bus. However, for clarity, in... Figure 2 The general labeled all buses as Bus System 205.
[0029] The user interface 203 may include a monitor, keyboard, mouse, trackball, click wheel, buttons, touchpad, or touch screen.
[0030] It is understood that memory 202 can be volatile memory or non-volatile memory, or both. The memory 202 in this embodiment is capable of storing data to support the operation of the terminal (e.g., 10-1). Examples of this data include any computer programs used to operate on the terminal (e.g., 10-1), such as operating systems and applications. The operating system includes various system programs, such as the framework layer, core library layer, driver layer, etc., used to implement various basic services and handle hardware-based tasks. Applications can include various applications.
[0031] In some embodiments, the candidate motion information list determination device provided in this application can be implemented using a combination of hardware and software. For example, the candidate motion information list determination device provided in this application can be a processor in the form of a hardware decoding processor, programmed to execute the candidate motion information list determination method provided in this application. For instance, the hardware decoding processor can employ one or more application-specific integrated circuits (ASICs), DSPs, programmable logic devices (PLDs), complex programmable logic devices (CPLDs), field-programmable gate arrays (FPGAs), or other electronic components.
[0032] As an example of the candidate motion information list determination device provided in this application embodiment being implemented using a combination of hardware and software, the candidate motion information list determination device provided in this application embodiment can be directly embodied as a combination of software modules executed by processor 201. The software modules can be located in a storage medium, which is located in memory 202. Processor 201 reads the executable instructions included in the software modules in memory 202 and combines them with necessary hardware (e.g., including processor 201 and other components connected to bus 205) to complete the candidate motion information list determination method provided in this application embodiment.
[0033] As an example, processor 201 can be an integrated circuit chip with signal processing capabilities, such as a general-purpose processor, a digital signal processor (DSP), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc., wherein the general-purpose processor can be a microprocessor or any conventional processor, etc.
[0034] As an example of the hardware implementation of the candidate motion information list determination device provided in this application embodiment, the device provided in this application embodiment can be directly executed by a processor 201 in the form of a hardware decoding processor. For example, it can be executed by one or more application specific integrated circuits (ASICs), DSPs, programmable logic devices (PLDs), complex programmable logic devices (CPLDs), field-programmable gate arrays (FPGAs), or other electronic components to implement the candidate motion information list determination method provided in this application embodiment.
[0035] The memory 202 in this embodiment is used to store various types of data to support the operation of the electronic device. Examples of such data include: any executable instructions for operation on the electronic device, such as executable instructions that can be included in the executable instructions, and a program implementing the candidate motion information list determination method of this embodiment.
[0036] In other embodiments, the candidate motion information list determination device provided in this application can be implemented in software. Figure 2A candidate motion information list determination device 2020 stored in memory 202 is shown. This device can be software in the form of programs and plugins, and includes a series of modules. As an example of a program stored in memory 202, it may include the candidate motion information list determination device 2020. The candidate motion information list determination device 2020 includes the following software modules: a first information processing module 2081, a first information filling module 2082, a second information processing module 2083, and a second information filling module 2084. When the software modules in the candidate motion information list determination device 2020 are read into RAM by processor 201 and executed, the candidate motion information list determination method provided in this application embodiment will be implemented. The functions of each software module in the candidate motion information list determination device 2020 are described below: The first information processing module 2081 is used to determine the candidate motion information list and the historical motion information table; The first information filling module 2082 is used to obtain at least one piece of motion information based on the historical motion information table when the candidate motion information list is not full. The first information filling module 2082 is used to fill the candidate motion information list based on the at least one motion information; The second information processing module 2083 is used to determine the positional relationship between the historical motion information table and the airspace motion information list; The second information filling module 2084 is used to obtain at least one motion information based on the historical motion information table and / or the spatial motion information list when the candidate motion information list is not full, and to fill the candidate motion information list based on the at least one motion information. The candidate motion information list is used to provide candidate predicted displacement vectors for the current codec block.
[0037] according to Figure 2 The electronic device shown, in one aspect of this application, also provides a computer program product or computer program including computer instructions stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform various embodiments and combinations of embodiments provided in the various optional implementations of the candidate motion information list determination method described above.
[0038] Before introducing the method for determining the candidate motion information list provided in this application, the video coding process in related technologies is first described, in which... Figure 3This is an optional flowchart illustrating the video encoding process in an embodiment of this application. The video signal refers to an image sequence comprising multiple frames. A frame represents the spatial information of the video signal. Taking YUV mode as an example, a frame includes a luminance sample matrix (Y) and two chrominance sample matrices (Cb and Cr). From the perspective of video signal acquisition methods, there are two types: those captured by a camera and those generated by a computer. Due to differences in statistical characteristics, the corresponding compression encoding methods may also differ.
[0039] In relevant video coding technologies, such as H.265 / HEVC (High Efficient Video Coding), H.266 / VVC (Versatile Video Coding), and AVS (Audio Video Coding Standard) (e.g., AVS3), a hybrid coding framework is used to perform a series of operations and processes on the input raw video signal: 1. Block Partition Structure: The input image is divided into several non-overlapping processing units, each of which undergoes a similar compression operation. This processing unit is called a CTU (Coding Tree Unit) or LCU (Large Coding Unit). Further subdivisions can be made below the CTU to obtain one or more basic coding units, called CUs (Coding Units). Each CU is the most basic element in a coding process. The following describes the various coding methods that can be used for each CU.
[0040] 2. Predictive Coding: This includes intra-frame prediction and inter-frame prediction. The original video signal is predicted from a selected reconstructed video signal to obtain a residual video signal. The encoder needs to select the most suitable predictive coding mode from among many possible modes for the current CU and inform the decoder. Intra-frame prediction refers to predicting signals from already encoded and reconstructed regions within the same image. Inter-frame prediction refers to predicting signals from other encoded images (called reference images) that are different from the current image.
[0041] 3. Transform Coding and Quantization: The residual video signal undergoes transformation operations such as DFT (Discrete Fourier Transform) and DCT (Discrete Cosine Transform) to convert the signal into the transform domain, where the coefficients are called transform coefficients. In the transform domain, the signal undergoes further lossy quantization, losing some information to make the quantized signal more suitable for compression. Some video coding standards may offer more than one transform option; therefore, the encoder needs to select one transform for the current CU and inform the decoder. The fineness of quantization is usually determined by the quantization parameter. A larger QP (Quantization Parameter) value means that coefficients with a wider range of values will be quantized into the same output, which usually results in greater distortion and a lower bitrate. Conversely, a smaller QP value means that coefficients with a smaller range of values will be quantized into the same output, which usually results in less distortion and a higher bitrate.
[0042] 4. Entropy Coding or Statistical Coding: The quantized transform domain signal is statistically compressed and encoded based on the frequency of each value, ultimately outputting a binary (0 or 1) compressed bitstream. Simultaneously, other information generated during encoding, such as the selected mode and motion vectors, also requires entropy coding to reduce the bit rate. Statistical coding is a lossless coding method that effectively reduces the bit rate required to represent the same signal. Common statistical coding methods include Variable Length Coding (VLC) or Content-Adaptive Binary Arithmetic Coding (CABAC).
[0043] 5. Loop Filtering: An encoded image undergoes inverse quantization, inverse transform, and prediction compensation operations (the reverse of operations 2-4 above) to obtain a reconstructed decoded image. Compared to the original image, the reconstructed image differs in some information due to quantization, resulting in distortion. Filtering the reconstructed image, such as deblocking, SAO (Sample Adaptive Offset), or ALF (Adaptive Lattice Filter), can effectively reduce the distortion caused by quantization. Since these filtered reconstructed images will serve as a reference for subsequent encoded images to predict future signals, the above filtering operations are also called loop filtering, or filtering operations within the encoding loop.
[0044] As can be seen from the above encoding process, at the decoding end, for each CU, after obtaining the compressed bitstream, the decoder first performs entropy decoding to obtain various mode information and quantized transform coefficients. Each coefficient undergoes inverse quantization and inverse transform to obtain the residual signal. On the other hand, based on the known encoding mode information, the prediction signal corresponding to that CU can be obtained. After adding the two, the reconstructed signal is obtained. Finally, the reconstructed value of the decoded image needs to undergo a loop filtering operation to generate the final output signal.
[0045] Relevant video coding standards, such as HEVC, VVC, and AVS3, all employ block-based hybrid coding frameworks. They divide the raw video data into a series of coded blocks and combine video coding methods such as prediction, transform, and entropy coding to achieve video data compression. Motion compensation is a commonly used prediction method in video coding. Based on the redundancy characteristics of video content in the temporal or spatial domains, motion compensation derives the predicted value of the current coded block from the already coded regions. These prediction methods include inter-frame prediction, intra-frame block copy prediction, and intra-frame string copy prediction. In specific coding implementations, these prediction methods may be used individually or in combination. For coded blocks using these prediction methods, one or more two-dimensional displacement vectors are typically explicitly or implicitly encoded in the bitstream to indicate the displacement of the current block (or its sibling blocks) relative to one or more reference blocks.
[0046] It is important to note that the displacement vector may have different names depending on the prediction mode and implementation. This article will uniformly describe them as follows: 1) The displacement vector in inter-frame prediction mode is called the motion vector (MV); 2) The displacement vector in IBC (Intra Block Copy) prediction mode is called the block vector (BV); 3) The displacement vector in ISC (Intra String Copy) prediction mode is called the string vector (SV). Intra-frame string copy is also called "string prediction" or "string matching," etc.
[0047] MV refers to the displacement vector used in inter-frame prediction mode, pointing from the current image to the reference image. Its value is the coordinate offset between the current block and the reference block, which are located in two different images. In inter-frame prediction mode, motion vector prediction can be introduced. By predicting the motion vector of the current block, the predicted motion vector corresponding to the current block is obtained. The difference between the predicted motion vector and the actual motion vector corresponding to the current block is encoded and transmitted. Compared with directly encoding and transmitting the actual motion vector corresponding to the current block, this saves bit overhead. In the embodiments of this application, the predicted motion vector refers to the predicted value of the motion vector of the current block obtained through motion vector prediction technology.
[0048] BV refers to the displacement vector used in IBC prediction mode, and its value is the coordinate offset between the current block and the reference block, where both the current block and the reference block are within the current image. In IBC prediction mode, block vector prediction can be introduced. By predicting the block vector of the current block, a predicted block vector corresponding to the current block is obtained. The difference between the predicted block vector and the actual block vector is encoded and transmitted. Compared to directly encoding and transmitting the actual block vector corresponding to the current block, this saves bit overhead. In the embodiments of this application, the predicted block vector refers to the predicted value of the block vector of the current block obtained through block vector prediction technology.
[0049] SV refers to the displacement vector used in ISC prediction mode, and its value is the coordinate offset between the current string and the reference string, where both the current string and the reference string are within the current image. In ISC prediction mode, string vector prediction can be introduced. By predicting the string vector of the current string, a predicted string vector corresponding to the current string is obtained. The difference between the predicted string vector and the actual string vector is encoded and transmitted. Compared to directly encoding and transmitting the actual string vector corresponding to the current string, this saves bit overhead. In the embodiments of this application, the predicted string vector refers to the predicted value of the string vector of the current string obtained through string vector prediction technology.
[0050] The following section introduces several different prediction models: I. Inter-frame prediction mode like Figure 4 As shown, Figure 4 This is a schematic diagram of the inter-frame prediction mode in an embodiment of this application. Inter-frame prediction utilizes the correlation in the temporal domain of the video, using pixels from neighboring encoded images to predict pixels in the current image, thereby effectively removing temporal redundancy and saving bits of encoded residual data. Here, P is the current frame, Pr is the reference frame, B is the current block to be encoded, and Br is the reference block of B. B' and B have the same coordinate position in the image; Br's coordinates are (xr, yr), and B''s coordinates are (x, y). The displacement between the current block to be encoded and its reference block is called the motion vector (MV), i.e.: MV=(xr-x,yr-y).
[0051] Considering the strong correlation between neighboring blocks in the temporal or spatial domains, MV prediction techniques can be used to further reduce the bits required to encode MVs. In H.265 / HEVC, inter-frame prediction includes two MV prediction techniques: Merge and AMVP (Advanced Motion Vector Prediction).
[0052] The Merge mode creates a candidate MV list for the current Prediction Unit (PU), containing five candidate MVs (and their corresponding reference images). It iterates through these five candidate MVs, selecting the one with the lowest rate-distortion cost as the optimal MV. If the codec creates the candidate list in the same way, the encoder only needs to transmit the index of the optimal MV in the candidate list. It's important to note that HEVC's MV prediction technique also has a Skip mode, a special case of the Merge mode. After finding the optimal MV in the Merge mode, if the current block and the reference block are essentially identical, then residual data does not need to be transmitted; only the MV index and a skip flag need to be sent.
[0053] refer to Figure 5 , Figure 5 This is a schematic diagram of candidate motion vectors provided in an embodiment of this application. The MV candidate list established by the Merge mode includes both spatial and temporal domain scenarios. For B-slice (B-frame image), it also includes a combined list method. The spatial domain provides a maximum of four candidate MVs, such as... Figure 5 The spatial domain list shown is constructed in the order A1→B1→B0→A0→B2, where B2 is a substitute; that is, if one or more of A1, B1, B0, and A0 are missing, the motion information of B2 must be used. The time domain provides at most one candidate MV, such as... Figure 5As shown, the MV of the isotopic PU is obtained by stretching and contracting according to the following formula: curMV=td colMV / tb; Where `curMV` represents the MV of the current PU, `colMV` represents the MV of the corresponding PU, `td` represents the distance between the current image and the reference image, and `tb` represents the distance between the corresponding image and the reference image. If the PU at position D0 in the corresponding block is unavailable, it is replaced by the corresponding PU at position D1. For a PU in a B-slice, since there are two MVs, its MV candidate list also needs to provide two MVPs (Motion Vector Predictors). HEVC generates a combination list for the B-slice by pairwise combining the first four candidate MVs in the MV candidate list.
[0054] Similarly, the AMVP mode utilizes the correlation of MVs (Motion Vector Differences) between neighboring blocks in the spatial and temporal domains to build a candidate MV list for the current PU. Unlike the Merge mode, the AMVP mode selects the optimal predicted MV from the candidate MV list and performs differential encoding with the optimal MV obtained through motion search for the current block to be encoded, i.e., encoded MVD = MV - MVP, where MVD is the Motion Vector Difference. The decoder builds the same list and only needs the MVD and MVP's indices in the list to calculate the MV of the current decoded block. The AMVP mode's MV candidate list also includes both spatial and temporal domain cases, but the difference is that the length of the AMVP mode's MV candidate list is only 2.
[0055] HMVP (History-based Motion Vector Prediction) is a new motion vector prediction technique adopted in H.266 / VVC. HMVP is a motion vector prediction method based on historical information. Motion information from historical coded blocks is stored in the HMVP list and used as the MVP for the current CU (Coding Unit). H.266 / VVC adds the HMVP to the candidate list of merge modes, following the spatial and temporal MVPs. HMVP uses a First-In-First-Out (FIFO) queue to store the motion information of previous coded blocks. If a stored candidate motion information is identical to the recently encoded motion information, this duplicate candidate motion information is removed first, then all HMVP candidates are moved forward, and the motion information of the current coding unit is added to the end of the FIFO. If the motion information of the current coding unit is different from any candidate motion information in the FIFO, the latest motion information is added to the end of the FIFO. When adding new motion information to the HMVP list, if the list has reached its maximum length, the first candidate in the FIFO is removed, and the latest motion information is added to the end of the FIFO. The HMVP list is reset (cleared) when a new CTU (Coding Tree Unit) row is encountered. In H.266 / VVC, the HMVP table size S is set to 6. To reduce the number of redundant determination operations, the following simplification is introduced: 1. Set the number of HMVP candidates used to generate the Merge list to (N <= 4)? M: (8-N), where N represents the number of existing candidates in the Merge list, and M represents the number of available HMVP candidates in the Merge list.
[0056] 2. Once the length of the available merge list reaches the maximum allowed length minus 1, the HMVP merge candidate list building process terminates.
[0057] II. IBC Prediction Model IBC is an intra-frame coding tool adopted in the HEVC Screen Content Coding (SCC) extension, which significantly improves the coding efficiency of screen content. AVS3 and VVC also employ IBC technology to enhance the performance of screen content coding. IBC utilizes the spatial correlation of screen content video, using already encoded image pixels in the current image to predict the pixels of the current block to be encoded, effectively saving the bits required to encode pixels. Figure 6 As shown, Figure 6This diagram illustrates the intra-block copy mode provided in an embodiment of this application. In IBC, the displacement between the current block and its reference block is called the BV (block vector). H.266 / VVC employs a BV prediction technique similar to inter-frame prediction to further save the bits required to encode the BV.
[0058] III. ISC Prediction Model ISC technology divides a coded block into a series of pixel strings or unmatched pixels according to a certain scanning order (such as raster scanning, round-trip scanning, and Zig-Zag scanning). Similar to IBC, each string searches for a reference string of the same shape in the currently coded area of the image, derives the predicted value of the current string, and effectively saves bits by encoding the residual between the current string's pixel value and the predicted value instead of directly encoding the pixel value. Figure 7 A schematic diagram of intra-frame string copying provided in this application embodiment is given. The dark gray area represents the encoded area, the 28 white pixels represent string 1, the 35 light gray pixels represent string 2, and the 1 black pixel represents an unmatched pixel. The displacement between string 1 and its reference string is... Figure 7 The displacement between string vector 1 and string 2 and their reference string is... Figure 7 String vector 2 in the string vector.
[0059] Intra-frame string copying technology requires encoding the SV (Structured Value), string length, and a flag indicating whether a matching string exists for each string in the current coding block. Here, SV represents the offset of the string to be encoded from its reference string. String length represents the number of pixels contained in the string. Different implementations encode the string length in various ways; the following are some examples (some examples may be combined): 1) Encode the string length directly in the bitstream; 2) Encode the number of pixels to be processed after processing the string in the bitstream, and the decoder calculates the current string length L=N-N1-N2 based on the current block size N, the number of processed pixels N1, and the number of pixels to be processed N2 obtained from decoding; 3) Encode a flag in the bitstream indicating whether the string is the last string. If it is the last string, the current string length L=N-N1 is calculated based on the current block size N and the number of processed pixels N1. If a pixel does not find a corresponding reference in the referable area, the pixel value of the unmatched pixel will be directly encoded.
[0060] IV. Intra-frame prediction motion vector prediction in AVS3 IBC and ISC are two screen content coding tools in AVS3. Both use the current image as a reference and derive the predicted values of coding units through motion compensation. Considering that IBC and ISC have similar reference areas, BV and SV have high correlation, and coding efficiency can be further improved by allowing predictions between the two. AVS3 uses an IntraHMVP (IntraPrediction History Motion Information Table) similar to HMVP to record the displacement vector information, position information, size information, and repetition count of these two types of coding blocks. The IntraHMVP derives the BVP (Block Vector Predictor) and SVP (String Vector Predictor). Here, BVP is the predicted value of the block vector, and SVP is the predicted value of the string vector. To support parallel coding, if the current largest coding unit is the first largest coding unit in the current row of the chip, the value of CntIntraHmvp in the IntraHMVP is initialized to 0.
[0061] 1. Derivation of the prediction block vector AVS3 adopts Class-based Block Vector Prediction (CBVP), similar to HMVP. This method first uses a History-based Block Vector Prediction (HBVP) list to store information about historical IBC coded blocks. Besides recording the block value (BV) information of historical coded blocks, it also records information such as the position and size of the historical coded blocks. For the current coded block, candidate BVs in the HBVP are classified according to the following criteria: Category 0: The area of a historical coded block is greater than or equal to 64 pixels; Category 1: BV frequency is greater than or equal to 2; Category 2: The coordinates of the top-left corner of the historical encoded block are located to the left of the top-left corner coordinates of the current block; Category 3: The coordinates of the top-left corner of the historical coded block are located above the coordinates of the top-left corner of the current block; Category 4: The coordinates of the top-left corner of the historical encoded block are located to the upper left of the top-left corner of the current block; Category 5: The coordinates of the top-left corner of the historical encoded block are located to the upper right of the top-left corner coordinates of the current block; Category 6: The coordinates of the top-left corner of the historical coded block are located to the lower left of the top-left corner coordinates of the current block; In this process, instances within each category are arranged in reverse order of their encoding sequence (the closer the encoding sequence is to the current block, the higher the order). The BV corresponding to the first historical encoded block is the candidate BV for that category. Then, candidate BVs for each category are added to the CBVP list in the order of category 0 to category 6. When adding a new BV to the CBVP list, it is necessary to determine whether a duplicate BV already exists in the CBVP list. A BV is added to the CBVP list only if no duplicate exists. The encoder selects the best candidate BV from the CBVP list as the BVP and encodes an index in the bitstream representing the index of the category corresponding to the best candidate BV in the CBVP list. The decoder decodes the BVP from the CBVP list based on this index.
[0062] After decoding the current prediction unit, if the prediction type of the current prediction unit is Block Copy Intra-Prediction (IBC), when NumOfIntraHmvpCand is greater than 0, IntraHMVP is updated according to the block copy intra-prediction motion information of the current prediction block in the manner described below. The intra-prediction motion information of the current prediction block includes displacement vector information, position information, size information, and repetition count. The displacement vector information of the block copy intra-prediction block is the block vector; the position information includes the x-coordinate and y-coordinate of the top-left corner of the current prediction block; the size information is the product of the width and height; and the repetition count of the current prediction block is initialized to 0.
[0063] 2. Derivation of the prediction string vector AVS3 encodes an index for each string in an ISC coded block, indicating the position of the string's SVP in the IntraHMVP. Similar to the skip mode in inter-frame prediction, the SV of the current string equals the SVP, so there is no need to encode the residual between SV and SVP.
[0064] After decoding the current prediction unit, if the prediction type of the current prediction unit is String Copy Intra-Prediction (ISC), when NumOfIntraHmvpCand is greater than 0, IntraHMVP is updated according to the String Copy Intra-Prediction Motion Information of the current prediction block in the manner described below. The String Copy Intra-Prediction Motion Information of the current prediction block includes displacement vector information, position information, size information, and repetition count. The displacement vector information of the current string is a string vector; the position information includes the x and y coordinates of the first pixel sample of the string, i.e., (xi, yi); the size information is the string length of this part, i.e., StrLen[i]; the repetition count is initialized to 0.
[0065] 3. Intra-frame prediction historical motion information table update Intra-frame prediction motion information includes displacement vector information, position information, size information, and repetition count. After decoding the current prediction unit, if the prediction type of the current prediction unit is block copy intra-frame prediction or string copy intra-frame prediction, and NumOfIntraHmvpCand is greater than 0, the intra-frame prediction historical motion information table IntraHmvpCandidateList is updated according to the intra-frame prediction motion information of the current prediction block. The displacement vector information, position information, size information, and repetition count of IntraHmvpCandidateList[X] are denoted as intraMvCandX, posCandX, sizeCandX, and cntCandX, respectively; otherwise, the operation defined in this clause is not executed.
[0066] a) Initialize X to 0 and cntCur to 0.
[0067] b) If CntIntraHmvp equals 0, then IntraHmvpCandidateList[CntIntraHmvp] contains the intra-frame predicted motion information of the current prediction unit, and CntIntraHmvp is incremented by 1.
[0068] c) Otherwise, determine whether the intra-predicted motion information of the current prediction block and IntraHmvpCandidateList[X] are the same based on whether intraMvCur and intraMvCandX are equal: 1) If intraMvCur and intraMvCandX are the same, execute step d); otherwise, increment X by 1.
[0069] 2) If X is less than CntIntraHmvp, proceed to step c); otherwise, proceed to step e).
[0070] d) cntCur equals the value of cntCandX plus 1. If sizeCur is less than sizeCandX, then the current sizeCur is equal to sizeCandX.
[0071] e) If X is less than CntIntraHmvp, then: 1) Let i be from X to CntIntraHmvp-1, and let IntraHmvpCandidateList[i] be equal to IntraHmvpCandidateList[i+1]; 2) IntraHmvpCandidateList[CntIntraHmvp-1] equals the intra-frame predicted motion information of the current prediction unit.
[0072] f) Otherwise, if X equals CntIntraHmvp and CntIntraHmvp equals NumOfIntraHmvpCand, then: 1) Let i range from 0 to CntIntraHmvp-1, and set IntraHmvpCandidateList[i] equal to IntraHmvpCandidateList[i+1]; 2) IntraHmvpCandidateList [CntIntraHmvp-1] equals the intra-predicted motion information of the current prediction unit.
[0073] g) Otherwise, if X equals CntIntraHmvp and CntIntraHmvp is less than NumOfIntraHmvpCand, then IntraHmvpCandidateList[CntIntraHmvp] equals the intra-predicted motion information of the current prediction unit, and CntIntraHmvp is incremented by 1.
[0074] In the current AVS3 standard, when performing displacement vector prediction, a candidate motion information list is constructed solely by building an IntraHMVP table, and the prediction block vector (BVP) or prediction string vector (SVP) is then derived. The maximum length of the IntraHMVP is 12, and the maximum length of the candidate motion information list is 7. When the IntraHMVP is insufficient or empty, the candidate motion information list cannot be filled, resulting in insufficient motion information for displacement vector prediction.
[0075] To overcome the above-mentioned shortcomings, continue to refer to Figure 8 The method for determining the candidate motion information list provided in this application can determine the corresponding candidate motion information list based on the spatial domain, and improve the compression performance of the video by combining it with the historical motion information list, thereby improving the user experience of video compression.
[0076] Continuing with the preceding embodiments, the method for determining the candidate motion information list provided in this application will be described. See also... Figure 8 , Figure 8 This is an optional flowchart illustrating the method for determining the candidate motion information list provided in the embodiments of this application. It can be understood that... Figure 8 The steps shown can be performed by various servers that run the candidate motion information list determination device, such as dedicated terminals, servers, or server clusters with candidate motion information list determination capabilities. The following section addresses... Figure 8 The steps shown are explained.
[0077] Step 801: The candidate motion information list determination device determines the candidate motion information list and the historical motion information table.
[0078] The historical motion information table is used for at least one of inter-frame prediction, intra-frame prediction, block copy intra-frame prediction (IBC), and string copy intra-frame prediction (ISC).
[0079] The candidate motion information list determination method provided in this application embodiment can be applied to three-dimensional vision. In high-frequency coding techniques, such as 3D coding techniques like HEVC (High Efficiency Video Coding), 3D coding uses adjacent reconstructed pixels to perform intra-frame prediction of the current block and selects predicted motion vectors from the motion vectors of neighboring blocks to construct a motion vector list for motion compensation inter-frame prediction. In 3D coding, three concepts are used to describe the entire coding process: Coding Unit (CU), Prediction Unit (PU), and Transform Unit (TU). A CU is a macroblock or sub-macroblock, and each CU is 2N. A 2N pixel block (N is a power of 2). Each CU implements the prediction process through a PU, the size of which is limited by the CU and can be a cube (e.g., 2N). 2N, N N), or a rectangle (2N) N, N 2N), in some embodiments of this application, as follows Figure 11 As shown, the size of block ABCDE is the minimum block size defined by the system (4). 4) Of course, in actual use, the method can be flexibly adjusted according to the application environment of the candidate motion information list.
[0080] Step 802: The candidate motion information list determination device determines whether the candidate motion information list has been filled. If yes, proceed to step 803; otherwise, proceed to step 804.
[0081] Step 803: Continue decoding.
[0082] Step 804: When the candidate motion information list is not full, the candidate motion information determination device obtains at least one motion information based on the historical motion information table.
[0083] The candidate motion information list is not full, including: The candidate motion information list is not full after being filled based on the spatial motion information list; or, the candidate motion information list is not full without being filled based on the spatial motion information list.
[0084] Specifically, when the candidate motion information list is not full, the candidate motion information list can be filled in at least one of the following ways: The candidate motion information list is populated based on the historical motion information table; alternatively, it is populated based on the spatial motion information list; or, it is populated based on both the historical motion information table and the spatial motion information list. By using different population processes individually or in combination, it is possible to adapt to different video compression environments, improving the adaptability of the candidate motion information list determination method provided in this application. The processes for determining the candidate motion information list using the historical motion information table, the spatial motion information list, and the combined spatial motion information list and historical motion information table will be described sequentially in subsequent embodiments.
[0085] Step 805: The candidate motion information list determination device populates the candidate motion information list based on at least one motion information.
[0086] The candidate motion information list is used to provide candidate predicted displacement vectors for the current codec block.
[0087] In some embodiments of this application, at least one piece of motion information is obtained based on a historical motion information table, and the candidate motion information list is populated based on the at least one piece of motion information, including: Based on the historical motion information table, the corresponding displacement vector is determined, and the candidate motion information list is filled based on the displacement vector; or, based on the historical motion information table, the corresponding displacement vector is determined, and dynamic filling is performed based on the parity of the index of the filling position in the candidate motion information list; or, based on the historical motion information table, the corresponding displacement vector is determined, and displacement vectors in the prediction mode are filled based on the displacement vector to achieve filling of the candidate motion information list; or, based on the historical motion information table, the corresponding displacement vector is determined, and dynamic filling is performed based on the number of displacement vectors and the index of the filling position in the candidate motion information list. Filling the candidate motion information list based on the displacement vector can be performed in any of the following methods 1)-3): Method 1) Based on the historical motion information table, determine the first and last displacement vectors in the historical motion information table, and fill the candidate motion information list based on the average of the first and last displacement vectors; Method 2) Based on the historical motion information table, determine the first and last displacement vectors in the historical motion information table, and fill the candidate motion information list based on the weighted average of the first and last displacement vectors; Method 3) Based on the historical motion information table, determine the first displacement vector in the historical motion information table, and fill the candidate motion information list based on the first displacement vector.
[0088] In some embodiments of this application, since the parity of the indices of the filling positions in the candidate motion information list differs, dynamic filling can be performed based on the parity of the indices. Specifically, when the index of the filling position in the candidate motion information list is odd, the first and last displacement vectors in the historical motion information table are determined, and the candidate motion information list is filled based on the average of the first and last displacement vectors; or, based on the historical motion information table, the first displacement vector in the historical motion information table is determined, and the candidate motion information list is filled based on the first displacement vector. On the other hand, when the index of the filling position in the candidate motion information list is even, the first and last displacement vectors in the historical motion information table are determined, and the candidate motion information list is filled based on the average of the first and last displacement vectors; or, based on the historical motion information table, the first displacement vector in the historical motion information table is determined, and the candidate motion information list is filled based on the first displacement vector.
[0089] In some embodiments of this application, the corresponding displacement vector is determined based on the historical motion information table, and the displacement vector in the prediction mode is filled based on the displacement vector. This can be achieved in the following ways: Based on the historical motion information table, the horizontal value of the first displacement vector and the vertical value of the last displacement vector in the table are determined. The horizontal direction of the displacement vectors in the prediction mode is filled based on the horizontal value of the first displacement vector; the vertical direction of the displacement vectors in the prediction mode is filled based on the vertical value of the last displacement vector. Of course, since the video compression environment during filling varies, to adapt to different usage environments, it is also possible to determine the vertical direction of the first displacement vector and the horizontal direction of the last displacement vector in the historical motion information table; fill the vertical direction of the displacement vectors in the prediction mode based on the vertical value of the first displacement vector; and fill the horizontal direction of the displacement vectors in the prediction mode based on the horizontal value of the last displacement vector.
[0090] In some embodiments of this application, since the number of candidate displacement vectors in the historical motion information table dynamically changes with different usage environments or compression levels, when the number of candidate displacement vectors in the historical motion information table is greater than the index of the filling position in the candidate motion information list, the position of the displacement vector to be filled in the historical motion information table can be determined. Based on the position of the displacement vector to be filled in the intra-frame predicted historical motion information table, the corresponding displacement vector is determined for dynamic filling. The position of the displacement vector to be filled in the historical motion information table is determined based on the number of candidate displacement vectors and the index of the filling position in the candidate motion information list. Specifically, when the number of candidate BVs, cnt_hbvp_cands, is greater than the index of the filling position in the candidate motion information list, cbvp_index, the (cnt_hbvp_cands%(cbvp_index+1))th BV in IntraHMVP is filled. Further, when the number of candidate displacement vectors in the historical motion information table is less than or equal to the index of the filling position in the candidate motion information list, the candidate motion information list determination method provided in the preceding embodiments can be executed to fill the candidate motion information list.
[0091] Continuing with the preceding embodiments, the method for determining the candidate motion information list provided in this application will be described. See also... Figure 9 , Figure 9 This is an optional flowchart illustrating the method for determining the candidate motion information list provided in the embodiments of this application. It can be understood that... Figure 9 The steps shown can be performed by various servers that run the candidate motion information list determination device, such as dedicated terminals, servers, or server clusters with candidate motion information list determination capabilities. The following section addresses... Figure 9 The steps shown are explained.
[0092] Step 901: Use a historical motion information list to record the motion information of historical prediction units (such as decoded blocks or decoded strings) during the decoding process in a first-in-first-out (FIFO) manner.
[0093] Step 902: When decoding the displacement vector of the target prediction unit (such as the current block or the current string), a candidate motion information list is determined based on the historical motion information list combined with other motion information.
[0094] Step 903: Obtain the position (or index) of the predicted displacement vector of the current prediction unit in the candidate motion information list from the code stream, and determine the predicted displacement vector of the current prediction unit.
[0095] Among them, reference Figure 10 , Figure 10 The following is an optional flowchart illustrating the method for determining a candidate motion information list provided in the embodiments of this application. In some embodiments of this application, the candidate motion information list can be determined based on a historical motion information list combined with other motion information, including the following steps: Step 1001: Determine the target parsing order. Based on the target parsing order, determine the corresponding ISC encoded block by considering the positions of multiple adjacent blocks of the current block.
[0096] Among them, reference Figure 11 , Figure 11This is a schematic diagram of a spatially adjacent block provided in one embodiment of this application. Decoded ISC blocks can be found in the order {A, B, C, D, E}, and the SV information of spatially adjacent ISC blocks can be recorded respectively. The recorded SV information may include the first or last SV information determined by the target parsing order of the encoded block. In this embodiment, a spatially adjacent block of the current codec block refers to a codec block belonging to the same image as the current codec block and spatially adjacent to it. Here, "adjacent" can mean that the distance between the current codec block and the current codec block is less than a threshold, which can be set according to actual conditions. Furthermore, the calculation method for the distance between two codec blocks can be flexibly set. For example, it can be the distance between the upper left corner coordinates of the two codec blocks, the distance between the upper left horizontal (or vertical) coordinates of the two codec blocks, the distance between the center positions of the two codec blocks, or the shortest distance between the two codec blocks, etc. This embodiment of the application does not limit this. In one example, suppose the top-left corner of the current codec block is (x0, y0), and the top-left corner of another codec block is (x1, y1). If the conditions |x0-x1| < a or |y0-y1| < b are met, then the other codec block is determined to be a spatially adjacent block of the current codec block. The values of a and b can be equal or unequal. For example, a = b = 8.
[0097] Spatial proximity blocks include spatially adjacent blocks and spatially non-adjacent blocks. Specifically, a spatially adjacent block of the current codec block refers to a codec block that belongs to the same image as the current codec block and is spatially adjacent to it. "Adjacent" can mean that it shares an edge or vertex with the current codec block. For example, ... Figure 11 As shown, the spatially adjacent blocks of the current codec block F can include codec blocks A, B, C, D, E, etc. The spatially non-adjacent blocks of the current codec block refer to codec blocks that belong to the same image as the current codec block and are spatially adjacent but not adjacent to it. "Adjacent but not adjacent" can mean that the distance between the current codec block and the current codec block is less than a threshold, but there are no overlapping edges or vertices between them. For example, as... Figure 11 As shown, the non-adjacent blocks in the spatial domain of the current codec block F can include codec blocks such as A´, B´, C´, D´, and E´.
[0098] In this embodiment, the spatially adjacent block includes an ISC block, where an "ISC block" refers to an encoding / decoding block that uses an ISC prediction mode for motion vector prediction. Optionally, the spatially adjacent block includes at least one of the following: spatially adjacent ISC blocks and spatially non-adjacent ISC blocks. Specifically, a spatially adjacent ISC block of the current encoding / decoding block is a spatially adjacent block of the current encoding / decoding block that is also an ISC block. A spatially non-adjacent ISC block of the current encoding / decoding block is a spatially non-adjacent block of the current encoding / decoding block that is also an ISC block.
[0099] Optionally, the spatially adjacent block further includes an IBC block, where an "IBC block" refers to an encoding / decoding block that uses the IBC prediction mode for motion vector prediction. Optionally, the spatially adjacent block further includes at least one of the following: spatially adjacent IBC blocks and spatially non-adjacent IBC blocks. Specifically, a spatially adjacent IBC block of the current encoding / decoding block refers to a spatially adjacent block of the current encoding / decoding block that is also an IBC block. A spatially non-adjacent IBC block of the current encoding / decoding block refers to a spatially non-adjacent block of the current encoding / decoding block that is also an IBC block.
[0100] Step 1002: Determine whether the motion information in the candidate motion information list is completely filled. If it is completely filled, proceed to step 1003; otherwise, proceed to step 1004. Step 1003: Execute the decoding process to determine the prediction displacement vector of the current prediction unit.
[0101] Step 1004: Determine the list of candidate motion information in the airspace, and fill the list of candidate motion information based on the corresponding filling information through the first filling process.
[0102] The first filling process can be used to fill based on a list of spatial motion information.
[0103] In some embodiments of this application, the filling information can be motion information of the ISC block, including at least one of the following: 1. The first encoded string SV determined in the ISC block according to the target parsing order, wherein the target parsing order can be the first target parsing order {A,B,C,D,E}; 2. The SV of the last encoded / decoded string determined according to the target parsing order in the ISC block; 3. The average SV of the first and last encoded / decoded strings in the ISC block, determined according to the target parsing order; 4. The weighted average of the SV of the first and last encoded / decoded strings determined according to the target parsing order in the ISC block; 5. SV of multiple encoded / decoded strings in an ISC block.
[0104] It should be noted that the motion information of the ISC blocks listed above is only exemplary and illustrative, and the embodiments of this application are not limited to other implementation methods.
[0105] In some embodiments of this application, the order for finding spatial neighboring blocks and non-neighboring block vectors can be a third target parsing order {A->A',B->B',C->C',D->D',E->E'}, where, for each position X (containing neighboring blocks and non-neighboring blocks, X=A,B,C,D,E), it is first determined whether the neighboring block X is an ISC block. When it is determined whether the neighboring block X is an ISC block, the SV corresponding to the ISC block is obtained, and the next position containing neighboring blocks and non-neighboring blocks is traversed according to the target parsing order until the filling is completed and the traversal stops.
[0106] If it is determined that the neighboring block X is not an ISC block, determine whether X' is an ISC or IBC block. If it is determined that the neighboring block X' is an ISC block, obtain the SV or BV corresponding to X'; otherwise, continue traversing according to the target parsing order.
[0107] In some embodiments of this application, the SV and / or BV information of non-adjacent blocks in the spatial domain can be determined first, and then the SV information of adjacent blocks in the spatial domain can be determined. An optional fourth target resolution order can be {A'->A, B'->B, C'->C, D'->D, E'->E}. For each position X (including neighboring and non-adjacent blocks, X=A, B, C, D, E), first determine whether the non-adjacent block X' is an ISC or IBC block. If the non-adjacent block X' is determined to be an ISC or IBC block, its corresponding SV or BV is obtained, and the traversal continues according to the target resolution order; otherwise, determine whether the neighboring block X is an ISC block. If it is, its corresponding SV is obtained; otherwise, the position is considered unobtainable, and the traversal continues according to the target resolution order.
[0108] Specifically, when an adjacent ISC block in the spatial domain is empty, or the SV of an adjacent ISC block in the spatial domain is unavailable, the candidate motion information list is filled based on the motion information of the already decoded non-adjacent IBC or ISC blocks in the spatial domain.
[0109] Among them, the location reference of non-adjacent IBC or ISC blocks in the airspace Figure 11 According to Table 1, the selection of spatially non-adjacent blocks can be based on the second target resolution order {A',B',C',D',E'}. Specifically, based on the second target resolution order, the BV or ISC block of the spatially non-adjacent IBC block is used according to the first or last SV in the scan order. Based on the SV information of the spatially non-adjacent ISC block, the filling is performed using the SV information of the ISC block, which is the same as the filling process shown in step 1004.
[0110] Where neig_x_pos and neig_y_pos represent the x and y coordinates of the top-left corner sample of the non-adjacent block in the spatial domain, respectively; cur_x_pos and cur_y_pos represent the x and y coordinates of the top-left corner sample of the current block F, respectively; and cu_width and cu_height represent the width and height of the current block F, respectively.
[0111] Table 1
[0112] When checking that a certain spatial block is in ISC mode, the specified SV can be added to the candidate list, or all stored SVs can be added to the candidate list in order.
[0113] Furthermore, optionally, when the candidate list is not yet full, the candidate motion information list can be filled with (0, 0).
[0114] Step 1005: Fill the candidate motion information list based on the corresponding filling information through the second filling process.
[0115] Step 1006: Determine the list of candidate motion information in the airspace, and fill the list of candidate motion information based on the corresponding filling information through the first filling process. If the motion information in the list of candidate motion information is still not completely filled, fill the list of candidate motion information based on the corresponding filling information through the second filling process.
[0116] Of these, steps 1004, 1005, and 1006 can be executed in one of them.
[0117] In some embodiments of this application, when filling the candidate motion information list based on spatial motion vectors and historical motion information tables, the filling positions in the candidate motion information list can be determined first, and then the corresponding filling information can be determined. Specifically, determining the filling positions in the candidate motion information list can include at least one of the following: Adjust the candidate motion information in the airspace before the historical motion information; or, adjust the candidate motion information in the airspace after the historical motion information; or, obtain at least one adjacent or non-adjacent motion information block in the airspace for filling.
[0118] In some embodiments of this application, corresponding spatial vectors can be filled at fixed candidate motion information list positions. For example, if a candidate motion vector of a certain type from historical motion information does not exist in the CBVP list, spatial vectors can be filled into the corresponding category positions in the candidate motion information list in sequence. Specifically, when a candidate motion vector of the i-th type from historical motion information to the left / top / upper left / upper right / lower left of the current block does not exist in the CBVP list, spatial vectors at the corresponding positions (left / top / upper left / upper right / lower left) can be filled.
[0119] In some embodiments of this application, when the spatial vector at the corresponding position does not exist, the candidate motion information list can be filled either by (0,0) encoding or by a second filling process.
[0120] The following example illustrates the method for determining the candidate motion information list provided in this application, using the transmission of short videos via an instant messaging client as an example. Figure 12 This diagram illustrates a usage scenario for the candidate motion information list determination method provided in this application embodiment. The video requiring compression and transmission is a short video. The terminals (including terminals 10-1 and 10-2) are equipped with a client application for software capable of displaying the corresponding short video, such as a short video playback client or plugin. Users can obtain and display the target video through the client. The terminals connect to the short video server 200 via network 300, which can be a wide area network (WAN), a local area network (LAN), or a combination of both, using a wireless link for data transmission. Alternatively, users can upload videos via a mini-program of instant messaging software on the terminal for other users on the network to view. In this process, the terminal needs to compress the uploaded video, or the server needs to compress the stored video to save transmission time, reduce storage space, and improve video compression efficiency. Figure 13 This is an optional video compression diagram in an embodiment of this application, in which the user compresses the target video to be processed through the compression plugin of the instant messaging client. The candidate motion information list determination method provided in this application can be saved in the plugin of the instant messaging client through a corresponding storage medium for the user to call. Figure 14 This is an optional compressed video presentation diagram in an embodiment of this application. When a user obtains video information stored on the server through a video client, the server can compress the target video to be transmitted using the candidate motion information list determination method provided in this application, so as to save the transmission time of the target video and reduce the storage space occupied by the target video.
[0121] Continue to refer to Figure 15 , Figure 15This is an optional flowchart illustrating the candidate motion information list determination method provided in the embodiments of this application. In some embodiments of this application, the target video can be compressed using only the first filling process, only the second filling process, or a combination of the first filling process and the second filling process.
[0122] In the compression process of short videos, the historical motion information table is used to process the intra-frame prediction historical motion information table. Populating the candidate motion information list using the intra-frame prediction historical motion information table can include the following steps: Step 1501: When the number of candidate displacement vectors in the intra-frame prediction history motion information table is greater than or equal to two, trigger displacement vector prediction to derive the predicted displacement vector.
[0123] When the number of candidate displacement vectors in the intra-frame prediction history motion information table is less than two, it is not necessary to encode the index of the category corresponding to the best candidate BV in the CBVP list in the bitstream.
[0124] Step 1502: Fill the data based on the displacement vector in the historical motion information table through the second filling process.
[0125] In some embodiments of this application, the mean of the first and last candidate BV in the historical motion information table IntraHMVP can be filled; the weighted mean of the first and last candidate BV in the historical motion information table IntraHMVP can also be filled; or, the first BV in the historical motion information table IntraHMVP can be filled.
[0126] Step 1503: Dynamically fill in the positions based on the parity of the indices of the positions to be filled in the candidate motion information list.
[0127] In some embodiments of this application, when the index cbvp_index of the position to be filled in the candidate motion information list is odd, the average of the first and last candidate BV in the historical motion information table IntraHMVP is filled; otherwise, the first BV in the historical motion information table IntraHMVP is filled.
[0128] In some embodiments of this application, when the index cbvp_index of the position to be filled in the candidate motion information list is even, the average of the first and last candidate BVs in IntraHMVP is filled; otherwise, the first BV in IntraHMVP is filled.
[0129] Step 1504: Dynamically fill in the parameters based on the displacement vector values.
[0130] In some embodiments of this application, the first BV value in the horizontal direction of the IntraHMVP can be filled in the horizontal direction of the block vector, and the last BV value in the vertical direction of the IntraHMVP can be filled in the vertical direction of the block vector; or, the last BV value in the horizontal direction of the IntraHMVP can be filled in the horizontal direction of the block vector, and the first BV value in the vertical direction of the IntraHMVP can be filled in the vertical direction of the block vector.
[0131] Step 1505: When the number of candidate BVs in IntraHMVP is greater than the index of the position to be filled in the candidate motion information list, fill in the (cnt_hbvp_cands%(cbvp_index+1))th BV in IntraHMVP.
[0132] Since the number of candidate BVs in IntraHMVP changes dynamically as video compression progresses, when the number of candidate BVs in IntraHMVP is greater than the index of the position to be filled in the candidate motion information list, the (cnt_hbvp_cands%(cbvp_index+1))th BV in IntraHMVP can be filled; otherwise, steps 1502 to 1504 in the preceding embodiment can be executed.
[0133] The embodiments of this application have the following beneficial effects: This application embodiment determines a candidate motion information list and the number of candidate displacement vectors in the intra-frame prediction history motion information table. When the candidate motion information list is not full, at least one motion information is obtained based on the intra-frame prediction history motion information table, and the candidate motion information list is filled based on the at least one motion information. The candidate motion information list is used to provide candidate predicted displacement vectors for the current codec block. Thus, more and more effective displacement vectors can be provided in the candidate motion information list to achieve better displacement vector prediction results, improve video compression performance, and enhance the user experience.
[0134] The above description is merely an embodiment of this application and is not intended to limit the scope of protection of this application. Any modifications, equivalent substitutions, and improvements made within the spirit and scope of this application are included within the scope of protection of this application.
Claims
1. A method of determining a list of candidate motion information, characterized by, The method includes: Determine a historical motion information table for the current codec block, wherein the historical motion information table is used for at least one of inter-frame prediction, intra-frame prediction, block copy intra-frame prediction (IBC), and string copy intra-frame prediction (ISC); Obtain the number of candidate displacement vectors in the historical motion information table; When the number of candidate displacement vectors in the historical motion information table is greater than or equal to two, displacement vector prediction is triggered to derive the predicted displacement vector, and the candidate motion information list is populated based on at least one motion information obtained from the historical motion information table. When the candidate motion information list is not completely filled, a spatial motion information list is constructed. The spatial motion information list includes the motion information of the spatial neighboring blocks of the current codec block. The spatial neighboring blocks include at least the intra-string copy (ISC) block. The candidate motion information list is populated based on at least one motion information obtained from the airspace motion information list; Specifically, the candidate motion information list is filled based on different filling positions in the candidate motion information list; The candidate motion information list is used to provide candidate predicted displacement vectors for the current codec block.
2. The method of claim 1, wherein, The method further includes: When the number of candidate displacement vectors is less than two, displacement vector prediction is not performed, and there is no need to encode the index of the position of the category corresponding to the best candidate displacement vector in the candidate motion information list in the bitstream.
3. The method of claim 1, wherein, The step of populating the candidate motion information list based on at least one motion information obtained from the historical motion information table includes: Based on the historical motion information table, the first and last displacement vectors in the historical motion information table are determined, and the candidate motion information list is populated based on the weighted average of the first and last displacement vectors.
4. The method of claim 1, wherein, The step of populating the candidate motion information list based on at least one motion information obtained from the historical motion information table includes: Dynamic filling is performed based on the parity of the index of the filling position in the candidate motion information list; or, Based on the historical motion information table, the corresponding displacement vector is determined, and the displacement vector in the prediction mode is filled in based on the displacement vector.
5. The method of claim 4, wherein, The dynamic filling based on the parity of the index of the filling position in the candidate motion information list includes: When the index of the position to be filled in the candidate motion information list is odd, determine the first and last displacement vectors in the historical motion information table, and fill the candidate motion information list based on the average of the first and last displacement vectors; or, Based on the historical motion information table, the first displacement vector in the historical motion information table is determined, and the candidate motion information list is filled based on the first displacement vector.
6. The method of claim 4, wherein, The dynamic filling based on the parity of the index of the filling position in the candidate motion information list includes: When the index of the position to be filled in the candidate motion information list is even, determine the first and last displacement vectors in the historical motion information table, and fill the candidate motion information list based on the average of the first and last displacement vectors; or, Based on the historical motion information table, the first displacement vector in the historical motion information table is determined, and the candidate motion information list is filled based on the first displacement vector.
7. The method of claim 4, wherein, The step of determining the corresponding displacement vector based on the historical motion information table, and filling the displacement vector in the prediction mode based on the displacement vector, includes: Based on the historical motion information table, determine the horizontal value of the first displacement vector and the vertical value of the last displacement vector in the historical motion information table; The horizontal direction of the displacement vector in the prediction pattern is filled based on the value of the horizontal direction of the first displacement vector; The vertical direction of the displacement vector in the prediction mode is filled based on the value of the vertical direction of the last displacement vector.
8. The method of claim 4, wherein, The step of determining the corresponding displacement vector based on the historical motion information table, and filling the displacement vector in the prediction mode based on the displacement vector, includes: Based on the historical motion information table, determine the vertical value of the first displacement vector and the horizontal value of the last displacement vector in the historical motion information table; The vertical direction of the displacement vector in the prediction mode is filled based on the value of the vertical direction of the first displacement vector; The horizontal direction of the displacement vector in the prediction mode is filled based on the value of the horizontal direction of the last displacement vector.
9. The method of claim 1, wherein, The method further includes: When the number of candidate displacement vectors in the historical motion information table is greater than the index of the filling position in the candidate motion information list, the position of the displacement vector to be filled in the historical motion information table is determined. Based on the position of the displacement vector to be filled in the historical motion information table, the corresponding displacement vector is determined for dynamic filling. The position of the displacement vector to be filled in the historical motion information table is determined based on the number of candidate displacement vectors and the index value of the filling position in the candidate motion information list.
10. The method according to claim 1, characterized in that, The candidate motion information list is used to provide candidate predicted block vectors (BVPs) for IBC blocks; or, The candidate motion information list is used to provide candidate predicted string vectors (SVPs) for the ISC string; or, The candidate motion information list is used to provide candidate BVPs for IBC blocks and candidate SVPs for ISC strings.
11. The method of claim 1, wherein, When the adjacent blocks in the airspace include ISC blocks, the motion information of the ISC blocks includes at least one of the following: The string vector SV of the first encoded / decoded string in the scan order of the ISC block; The SV of the last encoded / decoded string in the scan order of the ISC block; The average of the SV of the first and last encoded strings in the scan order of the ISC block; The weighted average of the SV of the first and last encoded / decoded strings in the scan order of the ISC block; The SV of multiple encoded / decoded strings in the ISC block.
12. The method of claim 1, wherein, The method further includes: The motion information derived from the airspace motion information list is located at a predetermined position in the candidate motion information list; or, When no motion information exists at at least one position in the candidate motion information list, motion information derived from the airspace motion information list is filled into the at least one position; wherein, during the filling process, the motion information is filled into the at least one position in the order of the motion information in the airspace motion information list; or, motion information from a specified position in the airspace motion information list is filled into the at least one position.
13. The method of claim 1, wherein, The process of filling the candidate motion information list based on different fill positions in the candidate motion information list includes: When filling the candidate motion information list, the motion information obtained from the airspace motion information list is in a first position in the candidate motion information list. This first position is located before the motion information obtained from the historical motion information list is in a second position in the candidate motion information list, or the first position is located after the motion information obtained from the historical motion information list is in a third position in the candidate motion information list.
14. A candidate motion information list determination apparatus characterized by comprising: include: The first information processing module is used to determine the historical motion information table for the current codec block, wherein the historical motion information table is used for at least one of inter-frame prediction, intra-frame prediction, block copy intra-frame prediction (IBC), and string copy intra-frame prediction (ISC). The first information filling module is used to obtain the number of candidate displacement vectors in the historical motion information table; When the number of candidate displacement vectors in the historical motion information table is greater than or equal to two, displacement vector prediction is triggered to derive the predicted displacement vector, and the candidate motion information list is populated based on at least one motion information obtained from the historical motion information table. The first information filling module is further configured to construct a spatial motion information list when the candidate motion information list is not completely filled. The spatial motion information list includes motion information of spatial neighboring blocks of the current codec block. The spatial neighboring blocks include at least intra-string copy (ISC) blocks. The second information filling module is used to fill the candidate motion information list based on at least one motion information obtained from the spatial motion information list; wherein, the candidate motion information list is filled based on different filling positions in the candidate motion information list; The candidate motion information list is used to provide candidate predicted displacement vectors for the current codec block.
15. An electronic device, comprising: The electronic device includes: Memory, used to store executable instructions; A processor, when executing executable instructions stored in the memory, implements the candidate motion information list determination method according to any one of claims 1 to 13.
16. A computer-readable storage medium storing executable instructions, wherein the instructions, when executed by a processor, cause the processor to perform operations comprising: When the executable instructions are executed by the processor, they implement the candidate motion information list determination method according to any one of claims 1 to 13.