Video information processing method and system, electronic equipment and storage medium
By analyzing the correlation between the encoding mode and the target encoding mode of the associated encoding video block of the video block to be encoded, the use of the affine prediction encoding mode is optimized, and the problem of low video encoding efficiency in the prior art is solved, and faster encoding speed and higher encoding efficiency are achieved.
Patent Information
- Application Number
- CN202510115543.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-23
- Publication Date
- 2025-05-13
AI Technical Summary
When processing complex video content, existing video encoding technology cannot fully explore and utilize broader and deeper information, resulting in limited encoding speed and ineffective improvement of encoding efficiency.
By obtaining the encoding mode of the associated coded video block of the video block to be encoded, and testing the affine predicted encoding mode based on the correlation between the encoding mode and the target encoding mode, the encoding strategy is determined to optimize the encoding process.
It effectively reduces the ineffective attempts of affine prediction, optimizes the encoding strategy, significantly speeds up the encoding speed, and improves the efficiency of video encoding.
Smart Images

Figure CN119996655A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of video coding, and in particular to a video information processing method, system, electronic device and storage medium. Background Art
[0002] At present, video codecs play a vital role in video communication and storage. They can reduce the transmission bandwidth and storage space requirements by compressing video files (video data). A series of efficient coding tools have been introduced in video codecs. Among them, affine prediction coding has attracted widespread attention due to its excellent performance in processing complex motion scenes. However, the efficiency and accuracy of affine prediction coding comes at the cost of increased computational complexity, making affine prediction one of the key algorithms that affects the speed of video codecs in the video coding process. In practical applications, coding speed is particularly important for real-time video communication and large-scale video processing. Therefore, how to speed up the coding process while ensuring coding quality has become an urgent problem to be solved.
[0003] In the related art, in order to improve the speed of affine prediction coding, the strategy usually adopted is to determine whether affine prediction is needed based on the rate-distortion cost (Rate-Distortion Cost) of the tried coding mode of the current video block to be encoded (unit to be encoded). Although the above method can accelerate the encoding process to a certain extent, it can only use limited information, making the decision-making basis relatively single, which limits the potential for further improving the encoding speed. As a result, when processing complex video content, the encoder cannot fully explore and utilize broader and deeper information, thus encountering a bottleneck in improving the encoding speed. Therefore, there is still a technical problem of low efficiency of video encoding.
[0004] To address the above-mentioned problems, no effective solution has been proposed yet. Summary of the invention
[0005] The embodiments of the present application provide a video information processing method, system, electronic device and storage medium to at least solve the technical problem of low efficiency of video encoding.
[0006] According to one aspect of an embodiment of the present application, a method for processing video information is provided. The method may include: obtaining at least one video block to be encoded from a video file; determining at least one coding mode corresponding to the video block to be encoded, wherein the coding mode is a mode allowed to be adopted by at least one associated coding video block associated with the video block to be encoded during the coding process, or the coding mode is a mode that has been tested during the coding process of the video block to be encoded; based on the correlation between the coding mode and the target coding mode corresponding to the video block to be encoded, testing the affine prediction coding mode to obtain a test result, wherein the performance index of the target coding mode for encoding the video block to be encoded is greater than the performance index threshold, and the test result is used to indicate the probability that the affine prediction coding mode becomes the target coding mode; according to the test result, determining the coding strategy of the video block to be encoded, wherein the coding strategy is used to indicate the rule for encoding the video block to be encoded.
[0007] According to another aspect of the embodiment of the present application, a video encoding method is provided. The method may include: obtaining at least one video block to be encoded from a video file; determining an encoding strategy for the video block to be encoded, wherein the encoding strategy is used to represent a rule for encoding the video block to be encoded, and the encoding strategy is determined according to a test result, the test result is used to represent the probability that the affine prediction encoding mode becomes the target encoding mode corresponding to the video block to be encoded, the test result is obtained by testing the affine prediction encoding mode based on the correlation between at least one encoding mode corresponding to the video block to be encoded and the target encoding mode, the performance index of the target encoding mode for encoding the video block to be encoded is greater than the performance index threshold, the encoding mode is a mode allowed to be used by at least one associated encoding video block associated with the video block to be encoded during the encoding process, or the encoding mode is a mode that has been tested by the video block to be encoded during the encoding process; according to the encoding strategy, the video block to be encoded is encoded.
[0008] According to another aspect of the embodiment of the present application, a video information processing method is also provided. The method may include: obtaining at least one video block to be encoded from a video file by calling a first interface, wherein the first interface includes a first parameter, and the parameter value of the first parameter includes the video block to be encoded; determining at least one encoding mode corresponding to the video block to be encoded, wherein the encoding mode is a mode allowed to be adopted by at least one associated encoding video block associated with the video block to be encoded during the encoding process, or the encoding mode is a mode that has been tested during the encoding process of the video block to be encoded; based on the correlation between the encoding mode and the target encoding mode corresponding to the video block to be encoded, testing the affine prediction encoding mode to obtain a test result, wherein the performance index of the target encoding mode for encoding the video block to be encoded is greater than the performance index threshold, and the test result is used to indicate the probability that the affine prediction encoding mode becomes the target encoding mode; determining the encoding strategy of the video block to be encoded according to the test result, wherein the encoding strategy is used to indicate the rule for encoding the video block to be encoded; outputting the encoding strategy by calling a second interface, wherein the second interface includes a second parameter, and the parameter value of the second parameter includes the encoding strategy.
[0009] According to another aspect of the embodiment of the present application, a video information processing system is also provided. The system may include: a video encoder, used to obtain at least one video block to be encoded from a video file; determine the encoding strategy of the video block to be encoded, wherein the encoding strategy is used to represent the rules for encoding the video block to be encoded, and the encoding strategy is determined according to the test result, the test result is used to represent the probability that the affine prediction encoding mode becomes the target encoding mode corresponding to the video block to be encoded, the test result is obtained by testing the affine prediction encoding mode based on the correlation between at least one encoding mode corresponding to the video block to be encoded and the target encoding mode, the performance index of the target encoding mode for encoding the video block to be encoded is greater than the performance index threshold, the encoding mode is the mode allowed to be used by at least one associated encoding video block associated with the video block to be encoded during the encoding process, or the encoding mode is the mode that has been tested during the encoding process of the video block to be encoded; according to the encoding strategy, encode the video block to be encoded; a server, used to obtain the encoded video block to be encoded.
[0010] According to another aspect of the embodiments of the present application, a computing device is further provided, including: a memory storing an executable program; and a processor for running the program, wherein the method in each embodiment of the present application is executed when the program is running.
[0011] According to another aspect of the embodiment of the present application, an electronic device is provided. The electronic device may include a memory and a processor: the memory is used to store computer executable instructions, and the processor is used to execute the computer executable instructions. When the computer executable instructions are executed by the processor, the method of the embodiment of the present application is implemented.
[0012] According to another aspect of an embodiment of the present application, a processor is further provided. The processor is used to run a program, wherein the above method of the embodiment of the present application is executed when the program is running.
[0013] According to another aspect of the embodiment of the present application, a computer-readable storage medium is further provided. The computer-readable storage medium includes a stored program, wherein when the program is executed, the device where the storage medium is located is controlled to execute the above method of the embodiment of the present application.
[0014] According to another aspect of the embodiment of the present application, a computer program product is also provided. The computer program product includes a computer program, and when the computer program is executed by a processor, the method of the embodiment of the present application is implemented.
[0015] In an embodiment of the present application, if it is necessary to process the information of the video, a video file to be processed can be obtained, and at least one video block to be encoded can be obtained from the video file. The coding mode allowed to be adopted by the associated coding video block associated with the video block to be encoded during the coding process can be determined, or the tested coding mode of the video block to be encoded can be determined as the coding mode corresponding to the video block to be encoded. The probability of whether the affine prediction coding mode can be used as the target coding mode can be tested based on the correlation (correlation degree) between the coding mode and the target coding mode corresponding to the video block to be encoded to obtain the test result. The coding strategy of the video block to be encoded can be determined according to the test result, and the corresponding video block to be encoded in the video file can be encoded by using the coding strategy. In an embodiment of the present application, the affine coding mode can be quickly predicted by the correlation between multiple coding modes, that is, by using the correlation between multiple coding modes between the video block to be encoded and the associated coding video block, the effectiveness of the affine prediction coding mode adopted by the video block to be encoded is predicted by judging the target coding mode of the above-mentioned associated coding video block. If the target coding mode indicates that the affine prediction coding mode may not be a suitable choice, the attempt to use the affine prediction coding mode can be skipped, thereby avoiding unnecessary computational complexity and significantly speeding up the coding process. The above method can effectively reduce invalid attempts of affine prediction, and also achieve the purpose of optimizing the coding strategy to increase the video coding speed, thereby achieving the technical effect of improving the efficiency of video coding and solving the technical problem of low video coding efficiency.
[0016] It is easy to notice that the above general description and the following detailed description are only for the purpose of exemplifying and explaining the present application, and do not constitute a limitation of the present application. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings:
[0018] Figure 1 It is a hardware structure block diagram of a computer terminal (or mobile device) for implementing a video information processing method according to an embodiment of the present application;
[0019] Figure 2 is a structural block diagram of a computing environment of a video information processing method according to an embodiment of the present application;
[0020] Figure 3 is a flow chart of a video information processing method according to an embodiment of the present application;
[0021] Figure 4 is a flowchart of a video encoding method according to an embodiment of the present application;
[0022] Figure 5 is a flowchart of another video information processing method according to an embodiment of the present application;
[0023] Figure 6 is a schematic diagram of a video information processing system according to an embodiment of the present application;
[0024] FIG. 7( a ) is a schematic diagram of a square coding unit division type;
[0025] FIG. 7( b ) is a schematic diagram of a rectangular coding unit division type;
[0026] Figure 8 is a flowchart of a process for determining an affine prediction coding mode according to an embodiment of the present application;
[0027] Fig. 9 is a flowchart of a processing process in an affine advanced predictive coding mode according to an embodiment of the present application;
[0028] Fig.10 is a schematic diagram of a video information processing device according to an embodiment of the present application;
[0029] Fig.11 is a schematic diagram of a video encoding device according to an embodiment of the present application;
[0030] Fig.12 is a schematic diagram of another video information processing device according to an embodiment of the present application;
[0031] Fig.13 is a structural block diagram of a computer terminal according to an embodiment of the present application;
[0032] Fig.14 is a structural block diagram of a computing device according to an embodiment of the present application;
[0033] Fig.15 is a structural block diagram of an electronic device according to an embodiment of the present application;
[0034] Fig.16 It is a block diagram of an electronic device according to a video information processing method of an embodiment of the present application. DETAILED DESCRIPTION
[0035] In order to enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be described clearly and completely below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in the field without creative work should fall within the scope of protection of the present application.
[0036] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any of their variations are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device comprising a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0037] First, some nouns or terms that appear in the description of the embodiments of the present application are subject to the following explanations:
[0038] H.266 / VVC, which stands for Versatile Video Coding (VVC), is the latest generation of video compression standard developed by the ITU and ISO Joint Video Experts Group. The standard was officially released in 2020, providing more efficient data compression and wider applicability;
[0039] Regular Merge module, an affine prediction coding mode. It uses the motion information of the coded block in the spatial domain to derive the motion vectors of each sub-block in the unit to be coded, without performing motion search, and has low computational complexity;
[0040] Regular Advanced Motion Vector Prediction (Regular AMVP) is an affine prediction coding mode that obtains affine motion vectors that meet the conditions through motion search and derives the motion vectors of each sub-block in the unit to be coded. The calculation complexity is relatively high. The 4-parameter mode (4paramaffineamvp) can be used for rotation, scaling and other picture contents, and the 6-parameter mode (6paramaffineamvp) can be used for rotation, scaling, shearing and other picture contents;
[0041] Skip mode, a conventional predictive coding mode;
[0042] Geometry Partitioned Prediction (Geo) mode, a prediction coding mode based on geometric partitioning, also known as GPM mode. It divides the unit to be coded into two parts, each of which independently uses the motion information of the coded block in the spatial domain;
[0043] Multi-Method Motion Vector Derivation (MMVD) mode is a prediction merge mode based on motion vector difference. It is based on the motion information of the coded block in the spatiotemporal domain, and supplements with a small number of predefined motion vector differences to obtain a relatively more accurate motion vector;
[0044] Affine Merge Prediction Coding (affine merge) mode is an affine prediction coding mode. It uses the motion information of the coded block in the spatial domain to derive the motion vector of each sub-block in the unit to be coded, without performing motion search, and has low computational complexity;
[0045] Affine Advanced Motion Vector Prediction (affine amvp) mode is an affine prediction coding mode that obtains affine motion vectors that meet the conditions through motion search and derives the motion vectors of each sub-block in the unit to be encoded. The calculation complexity is relatively high. The 4-parameter mode (4paramaffine amvp) can be used for rotation, scaling and other picture contents, and the 6-parameter mode (6param affine amvp) can be used for rotation, scaling, shearing and other picture contents.
[0046] According to an embodiment of the present application, a video information processing method is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer executable instructions, and although a logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that shown here.
[0047] The method embodiments provided in the embodiments of the present application can be executed in a mobile terminal, a computer terminal or a similar computing device. Figure 1 is a hardware structure block diagram of a computer terminal (or mobile device) for implementing a video information processing method according to an embodiment of the present application, such as Figure 1 As shown, the computer terminal 10 (or mobile device) may include one or more (102a, 102b, ..., 102n are used to illustrate) processors 102 (the processor 102 may include but is not limited to a processing device such as a microprocessor (Microcontroller Unit, referred to as MCU) or a programmable logic device (Field Programmable Gate Array, referred to as FPGA)), a memory 104 for storing data, and a transmission device 106 for communication functions. In addition, it may also include: a display, an input / output interface (I / O interface), a Universal Serial Bus (USB) port (which may be included as one of the ports of the BUS bus), a network interface, a power supply and / or a camera. It can be understood by those skilled in the art that Figure 1 The structure shown is only for illustration and does not limit the structure of the above electronic device. Figure 1 More or fewer components as shown, or with Figure 1 Different configurations are shown.
[0048] Figure 1 The hardware structure block diagram shown can be used not only as an exemplary block diagram of the above-mentioned computer terminal 10 (or mobile device), but also as an exemplary block diagram of the above-mentioned server. In an optional embodiment, Figure 2 The block diagram shows the use of the above Figure 1 The computer terminal 10 (or mobile device) shown is an embodiment of a computing node in the computing environment 201.
[0049] Figure 2 is a structural block diagram of a computing environment of a video information processing method according to an embodiment of the present application, such as Figure 2As shown, computing environment 201 includes multiple computing nodes (such as servers) running on a distributed network (shown as 210-1, 210-2, ...) in the figure. The computing nodes all contain local processing and memory resources, and end users 202 can remotely run applications or store data in computing environment 201. Applications can be provided as multiple services 220-1, 220-2, 220-3 and 220-4 in computing environment 201, representing services "A", "D", "E" and "H" respectively.
[0050] The end user 202 can provide and access services through a web browser or other software application on the client, and in some embodiments, the end user 202's provision and / or request can be provided to the entry gateway 230. The entry gateway 230 can include a corresponding agent to handle the provision and / or request for the service (one or more services provided in the computing environment 201).
[0051] Services are provided or deployed based on various virtualization technologies supported by the computing environment 201. In some embodiments, services can be provided based on virtual machine (VM)-based virtualization, container-based virtualization, and / or similar methods. Virtual machine-based virtualization can be to simulate a real computer by initializing a virtual machine, and execute programs and applications without directly contacting any actual hardware resources. While the virtual machine virtualizes the machine, according to container-based virtualization, a container can be started to virtualize the entire operating system so that multiple workloads can run on a single operating system instance.
[0052] In an embodiment based on container virtualization, several containers of a service can be assembled into a Pod (e.g., a Kubernetes Pod). Figure 2 As shown, service 220-2 can be equipped with one or more Pods 240-1, 240-2, ..., 240-N (collectively referred to as Pods). Pods can include a proxy 245 and one or more containers 242-1, 242-2, ..., 242-M (collectively referred to as containers). One or more containers in a Pod process requests related to one or more corresponding functions of the service, and the proxy 245 generally controls network functions related to the service, such as routing, load balancing, etc. Other services can also be equipped with Pods similar to Pods.
[0053] During operation, executing a user request from end user 202 may require invoking one or more services in computing environment 201, and executing one or more functions of a service may require invoking one or more functions of another service. Figure 2As shown, service "A" 220-1 receives a user request from end user 202 from ingress gateway 230, service "A" 220-1 may call service "D" 220-2, and service "D" 220-2 may request service "E" 220-3 to perform one or more functions.
[0054] The computing environment described above can be a cloud computing environment, where the allocation of resources is managed by the cloud service provider, allowing the development of functions without considering the implementation, adjustment or expansion of servers. The computing environment allows developers to execute code in response to events without building or maintaining complex infrastructure. Services can be divided into a set of functions that can be automatically and independently scaled, rather than expanding a single hardware device to handle potential loads.
[0055] Under the above operating environment, this application provides Figure 3 It should be noted that the video information processing method of this embodiment can be Figure 1 The illustrated embodiment is executed by a mobile terminal. Figure 3 is a flow chart of a video information processing method according to an embodiment of the present application, such as Figure 3 As shown, the method may include the following steps:
[0056] Step S302: Obtain at least one to-be-encoded video block from a video file.
[0057] In the technical solution provided in the above step S302 of the present application, the video file may also be referred to as a video. The video file may include multiple formats, such as Audio Video Interleaved (AVI). It should be noted that the above video file formats are only for illustration and are not specifically limited here. The video block to be encoded is the basic coding unit obtained after the video file is divided, and may also be referred to as a coding unit (CU).
[0058] In this embodiment, if video information processing is required, a video file to be processed may be obtained, and a video block to be encoded may be obtained from the video file.
[0059] Optionally, when compressing, converting formats, streaming, storing, editing, video conferencing, communicating, uploading, and distributing video files, it is necessary to encode the video blocks to be encoded in the video files.
[0060] Optionally, video encoding can start from reading a video file, which can come from various sources, such as hard disk, network stream, camera capture, etc. Obtaining the video blocks to be encoded from the video file is a process involving video file parsing, video stream reading, metadata extraction and video block division.
[0061] Optionally, during the video file parsing process, the encapsulation format of the video file and the encoding standard of the video file, such as H.266 / VVC, etc., can be identified. The above identification of the encapsulation format and the encoding format can be completed by reading the header information or metadata of the video file. The video stream in the video file is decapsulated, that is, the video data is separated from the video file, and the audio data and subtitle data can also be separated.
[0062] Optionally, during the video stream reading process, video frames may be read sequentially from the video stream. Each video frame represents an image of the video and includes pixel information.
[0063] Optionally, in the process of dividing the video file into blocks, a method for dividing the video frame into smaller units can be determined according to the video content of the video file and the requirements of the coding standard, wherein the smaller unit is also the unit to be encoded. In the H.266 / VVC standard, strategies such as quad-tree (Quad-Tree, referred to as QT), binary tree (Binary Tree, referred to as BT), ternary tree (Ternary Tree, referred to as TT) division and nested division can be used. Divide the video frame into multiple units to be encoded. The size of the unit to be encoded can be adaptively adjusted according to the complexity of the content and the motion situation. For example, the static part may be divided into larger blocks, while the complex motion area is divided into smaller blocks.
[0064] Optionally, in the process of obtaining the video block to be encoded, one or more units to be encoded may be selected as the video block to be encoded. The selection criteria may be based on coding efficiency, quality requirements or coding strategies. For the selected video block to be encoded, encoding parameters are initialized, such as affine prediction coding mode, transform type, quantization parameter, etc.
[0065] In the embodiment of the present application, obtaining the video block to be encoded from the video file involves parsing the file format, extracting metadata, decapsulating the video stream, reading frames, dividing blocks, and initializing encoding parameters. The above operations prepare the necessary data and parameters for the subsequent video encoding process. Through the above operations, the video data can be efficiently compressed and processed to meet the needs of storage, transmission, and playback.
[0066] Step S304: determine at least one encoding mode corresponding to the video block to be encoded.
[0067] In the technical solution provided in the above step S304 of the present application, the coding mode is a mode allowed to be adopted by at least one associated coding video block associated with the video block to be coded during the coding process, or the coding mode is a mode that has been tested by the video block to be coded during the coding process. That is, the coding mode corresponding to the video block to be coded may be the coding mode of at least one associated coding video block associated with the video block to be coded. The coded video block of the video block to be coded may also be a unit to be coded that has a correlation with the video block to be coded. For example, the associated coding video block may be an attempted coding unit at the same position as the video block to be coded. In the above case, the coding mode may be a coding mode that meets the requirements of the attempted coding unit at the same position. The associated coding video block may also be an adjacent coding unit of the video block to be coded. In this case, the coding mode may be a coding mode that meets the requirements of the adjacent coding unit. The associated coding video block may also be a parent coding unit of the video block to be coded. In this case, the coding mode may be a coding mode that meets the requirements of the parent coding unit. The coding mode may also be a tested coding mode of the video block to be coded, that is, it may be a coding mode that meets the requirements that has been attempted by the video block to be coded. The coding mode may include various conventional prediction coding modes (also referred to as conventional coding modes).
[0068] Optionally, the regular prediction coding mode may include but is not limited to Regular Merge module, Regular AMVP mode, skip mode, geo mode, mmvd mode.
[0069] In this embodiment, after obtaining the to-be-encoded video block from the video file, the encoding mode corresponding to the to-be-encoded video block may be determined.
[0070] Optionally, during the video encoding process, the purpose of determining the encoding mode corresponding to the video block to be encoded is to infer or determine a satisfactory encoding mode for the video block to be encoded based on the correlation between the video blocks, thereby improving encoding efficiency and compression performance.
[0071] Optionally, if there is an encoded block in the current video frame or the previous video frame that is in the same spatial position as the video block to be encoded, then the encoded block can be used as an associated encoded video block. In particular, when the H.266 / VVC standard allows the use of multiple block partitioning types (for example, quadtree, binary tree, and ternary tree partitioning), even in the same video frame, different partitioning paths may generate encoded blocks in the same position as the block to be encoded. The above phenomenon is more obvious in the case of nested partitioning. For example, a larger block can be divided into smaller sub-blocks first. Some of the above sub-blocks can be further divided. Therefore, the encoding mode and motion mode of the blocks in the same position can provide valuable reference information when judging the prediction mode of the block to be encoded.
[0072] Optionally, the surrounding coded blocks (i.e., adjacent coding units) of the video block to be coded can be used as associated coded video blocks, including blocks that share a boundary with the video block to be coded above or on the left. Since adjacent blocks usually have similar texture, motion, and brightness characteristics, the above-mentioned adjacent coding units are similar to the coding mode with coded video blocks.
[0073] Optionally, for a non-root node block partition structure, the parent coding unit of the video block to be encoded can provide important information about its encoding mode. If the parent coding unit adopts a specific encoding mode, the video block to be encoded can also be applied to the mode, or other related encoding modes.
[0074] It should be noted that the above-mentioned associated coded video blocks having correlation with the video block to be encoded are only illustrative and are not specifically limited here. As long as the process and method of determining the coding mode corresponding to the video block to be encoded is determined by analyzing other units to be encoded that are correlated with the video block to be encoded and using the coding mode corresponding to the unit to be encoded, they are all within the protection scope of the embodiments of the present application.
[0075] Optionally, the coding modes of the coded units at the same position are analyzed, especially those coding modes with good coding effects. If the coding mode of the unit to be coded at the same position is used and the coding effect is good, it can be inferred that the video block to be coded may also be suitable for the above coding mode. Check the coding modes of the adjacent coding units around the video block to be coded. If there are one or more adjacent units using a certain coding mode and the coding quality is high, then this may indicate that the block to be coded is also suitable for video coding using the above coding mode. When the video block to be coded is obtained by segmentation from a larger block, the coding mode of its parent coding unit provides contextual information. If the parent coding unit uses a certain coding mode, then the video block to be coded may also be suitable for the above coding mode or a coding mode related to the coding mode.
[0076] In an embodiment of the present application, by analyzing the coding mode of the coded unit related to the video block to be coded, a context-based coding mode decision mechanism is provided for the video block to be coded, which helps to optimize the coding process, reduce invalid attempts, and thus improve the coding speed and coding performance.
[0077] Step S306 , based on the correlation between the coding mode and the target coding mode corresponding to the video block to be coded, the affine prediction coding mode is tested to obtain a test result.
[0078] In the technical solution provided in the above step S306 of the present application, the performance index of the target coding mode for encoding the video block to be encoded is greater than the performance index threshold. For example, the target coding mode can be a high-performance coding mode among multiple candidate coding modes, that is, a coding mode whose performance meets the requirements. The target coding mode generally refers to those coding modes that can provide good compression efficiency and visual quality, and the performance index of the target coding mode (such as bit rate, encoding time, etc.) meets or exceeds the predetermined performance threshold. The performance index threshold can be set according to specific requirements. For example, while maintaining the video quality, the bit rate needs to be lower than a certain value, or the encoding time cannot exceed a certain time limit. The correlation can be used to indicate the connection and similarity between the coding mode corresponding to the video block to be encoded and the target coding mode. Optionally, if the coding mode corresponding to the video block to be encoded is a coding mode whose performance meets the requirements on some video blocks to be encoded, then the above correlation can be used to predict the performance of the video block to be encoded when the affine prediction coding mode is adopted. The test result can be used to indicate the probability of the affine prediction coding mode becoming the target coding mode.
[0079] Optionally, the main difference between the affine prediction coding mode and the conventional prediction coding mode is the description and processing capabilities of the two for motion, as well as the resulting differences in coding complexity and applicable scenarios. According to the characteristics of the video content in the video file, a suitable coding mode can be intelligently selected to achieve suitable compression efficiency and video quality. The affine prediction coding mode may include but is not limited to the affine merge mode and the affine amvp mode. The above coding modes are only for illustration and are not specifically limited here.
[0080] In this embodiment, after the coding mode of the video block to be coded is determined, the affine prediction coding mode may be tested based on the correlation between the coding mode and the target coding mode corresponding to the video block to be coded to obtain a test result.
[0081] Optionally, the above embodiment is a key decision point for deciding whether to adopt the affine prediction coding mode in the video coding process. The goal of the above embodiment is to evaluate the probability of the affine prediction coding mode becoming the target coding mode, that is, to evaluate whether the affine prediction coding mode can provide sufficiently high coding performance to meet a specific performance indicator threshold.
[0082] Optionally, after determining the coding mode corresponding to the video block to be encoded, based on the correlation between the coding mode and the target coding mode, the affine prediction coding mode can be tested to evaluate the probability that the simulated prediction coding mode becomes the target coding mode. During the test, the affine prediction coding mode can be applied to the video block to be encoded, and the bit rate, encoding time, and visual quality after encoding can be calculated and compared with the performance indicator threshold.
[0083] Optionally, the test result may reflect the encoding performance of the affine prediction encoding mode on the video block to be encoded. If the test result indicates that the affine prediction encoding mode can meet or exceed the performance index threshold, then the affine prediction encoding mode may be the target encoding mode. If the test result indicates that the performance of the affine prediction encoding mode does not meet the performance index threshold, or is significantly different from the performance of the encoding mode, the affine prediction encoding mode may be skipped and other encoding modes may be tested to find an encoding mode that meets the requirements.
[0084] In the embodiment of the present application, the spatial and temporal correlations between video blocks are fully utilized through the above decision-making process, and by analyzing the performance of the coding mode, it is decided whether to continue to conduct in-depth testing on the affine prediction coding mode. The above correlation-based test process helps to quickly identify the coding mode suitable for the current video block and avoid unnecessary testing, thereby optimizing the video coding quality while ensuring coding efficiency. In summary, the affine prediction coding mode test and decision-making mechanism based on the correlation of the coding mode of the video block, by evaluating the probability of the affine prediction coding mode becoming the target coding mode, quickly locates the coding mode that meets the requirements from a large number of coding modes, thereby improving the speed and efficiency of video coding.
[0085] Step S308: Determine the encoding strategy of the video block to be encoded according to the test result.
[0086] In the technical solution provided in the above step S308 of the present application, the coding strategy can be used to represent the rules for encoding the video block to be encoded. The coding strategy can be a series of rules and parameter settings followed by the encoder when encoding the video block. It not only includes selecting a specific prediction coding mode, but may also involve transform types, quantization parameters, loop filter settings, etc. The formulation of the coding strategy aims to balance the encoding quality, bit rate and encoding speed, ensuring that the video content is encoded with as few bits as possible under given bandwidth and storage constraints while maintaining good visual quality.
[0087] In this embodiment, after the affine prediction coding mode is tested based on the correlation between the coding mode and the target coding mode corresponding to the video block to be encoded, the coding strategy of the video block to be encoded can be determined according to the test result.
[0088] Optionally, the final coding strategy of the video block to be encoded is determined according to the test result of the affine prediction coding mode. The above method is crucial to improving video coding efficiency and video quality. By intelligently selecting the coding strategy, the encoder can make decisions that meet the corresponding needs when processing complex scenes.
[0089] Optionally, if the test result of the affine prediction coding mode indicates that its performance index meets or exceeds the target performance index (i.e., the performance index threshold), and has obvious advantages over other tested coding modes, then the affine prediction coding mode can be used as the coding strategy for the current video block. If the test result of the affine prediction coding mode is not ideal, or its performance index fails to reach the target performance index, the encoder will select other coding modes as the coding strategy for the current video block based on the test results and clues of the coding mode. If the test results indicate that no single coding mode is significantly better than other modes, a comprehensive coding strategy can be adopted, that is, multiple coding modes are tried for the current video block, and finally one of them is selected whose performance meets the requirements. The above strategy helps to ensure that the required coding effect can be achieved under various video content.
[0090] In an embodiment of the present application, the determination of the encoding strategy should have a certain degree of flexibility to adapt to the diversity of video content. In actual operation, the encoder may dynamically adjust the encoding strategy according to the characteristics of the video frame, motion complexity, target bit rate, etc., to ensure that the required encoding effect can be achieved in different scenarios. In summary, the above method intelligently determines the encoding strategy of the video block to be encoded based on the test results of the affine prediction coding mode and the characteristics of the video content. The above steps can more efficiently and intelligently select the encoding mode and parameter configuration that meet the requirements, thereby achieving a reduction in bit rate and an increase in encoding speed while ensuring video quality.
[0091] Through the above steps S302 to S308 of the present application, if it is necessary to process the information of the video, the video file to be processed can be obtained, and at least one video block to be encoded can be obtained from the video file. The coding mode allowed to be adopted by the associated coding video block associated with the video block to be encoded during the coding process can be determined, or the tested coding mode of the video block to be encoded can be determined as the coding mode corresponding to the video block to be encoded. The probability of whether the affine prediction coding mode can be used as the target coding mode can be tested based on the correlation between the coding mode and the target coding mode corresponding to the video block to be encoded to obtain the test result. The coding strategy of the video block to be encoded can be determined according to the test result, and the corresponding video block to be encoded in the video file can be encoded by using the coding strategy. In this embodiment, the affine coding mode can be quickly predicted by the correlation between multiple coding modes, that is, by using the correlation between multiple coding modes between the video block to be encoded and the associated coding video block, the effectiveness of the affine prediction coding mode adopted by the video block to be encoded is predicted by judging the target coding mode of the above-mentioned associated coding video block. If the target coding mode indicates that the affine prediction coding mode may not be a suitable choice, the attempt to use the affine prediction coding mode can be skipped, thereby avoiding unnecessary computational complexity and significantly speeding up the coding process. The above method can effectively reduce invalid attempts of affine prediction, and also achieve the purpose of optimizing the coding strategy to increase the video coding speed, thereby achieving the technical effect of improving the efficiency of video coding and solving the technical problem of low video coding efficiency.
[0092] The above method of this embodiment is further introduced below.
[0093] As an optional implementation, a performance index of the associated coded video block for encoding the associated coded video block in a mode allowed to be adopted during the encoding process is greater than a performance index threshold, and a performance index of the to-be-coded video block for encoding the to-be-coded video block in a mode that has been tested during the encoding process is greater than the performance index threshold.
[0094] In this embodiment, the performance index of encoding the associated coded video block using the mode allowed to be used in the coding process is greater than the performance index threshold. The performance index of encoding the to-be-coded video block using the mode that has been tested in the coding process is greater than the performance index threshold.
[0095] Optionally, the coding mode and performance index of the associated coded video block and the performance index of the tested coding mode of the video block to be coded are used to guide the selection of the coding strategy required for the subsequent coding process of the video block to be coded. The above mechanism helps to ensure that the coding strategy can meet a specific performance index threshold, thereby maintaining video quality while optimizing coding efficiency.
[0096] Optionally, the associated coded video block refers to a coded video block that is similar to the video block to be coded in position, size or motion characteristics, including a tried coding unit, a neighboring coding unit or a parent coding unit at the same position. The coding mode of the associated coded video block refers to the prediction coding mode selected for the video block during the encoding process.
[0097] Optionally, the performance indicators may include bit rate (bit rate), encoding time, peak signal-to-noise ratio, structural similarity index, etc., which can be used to evaluate the quality and efficiency of video encoding. The performance indicator threshold is the minimum requirement for encoding performance and is set according to the needs of specific application scenarios. For example, in real-time video communication, the encoding time threshold may be very strict, while in a video on demand scenario, the focus may be on bit rate and video quality.
[0098] Optionally, if the performance index of the encoding mode of the associated encoded video block during encoding is greater than the performance index threshold, this means that the encoding mode provides a high-quality encoding effect when processing the associated video block. In the above case, it can be inferred that the encoding mode may also show good performance when processing a video block to be encoded with similar characteristics, and therefore, the above encoding mode can be used as a potential encoding strategy for the video block to be encoded.
[0099] Optionally, during the encoding process of the video block to be encoded, multiple encoding modes may be tried, and the encoding performance index under each encoding mode may be recorded. If the performance index of a tested encoding mode for encoding the video block to be encoded is greater than the performance index threshold, this indicates that the encoding mode can meet the requirements of encoding quality or efficiency and is an effective encoding selection result.
[0100] Optionally, the encoder has determined the coding mode corresponding to the video block to be coded and tested the affine prediction coding mode. The encoder can compare the performance indicators of the coding mode and the tested coding mode to determine the appropriate coding strategy for the video block to be coded.
[0101] In an embodiment of the present application, if the performance index of the affine merge or affine amp mode is greater than the performance index threshold in the associated coding video block, and in the test of the video block to be encoded, the performance index of the affine prediction coding mode also exceeds the threshold, the encoder will give priority to using the affine prediction coding mode for encoding. This is because the affine prediction coding mode can more accurately describe and predict complex motions, such as rotation, scaling and shearing, thereby providing higher compression efficiency and better video quality. On the contrary, if the performance index of the tested conventional coding mode (e.g., regular amp, skip, etc.) exceeds the performance index threshold, and the affine prediction coding mode fails to achieve a similar performance level, the encoder may select these simpler modes for encoding to avoid affecting the overall encoding speed due to the high complexity of affine prediction. In summary, by comparing the performance index of the associated coding video block and the tested coding mode, a more reasonable coding strategy decision can be made to ensure that the required coding quality and efficiency are achieved while meeting the performance index threshold. The above mechanism makes full use of the spatiotemporal correlation of the video content and improves the intelligence and adaptability of the encoder.
[0102] As an optional implementation, step S308, determining the encoding strategy of the video block to be encoded according to the test result, includes: in response to the test result being that the probability is less than or equal to the probability threshold, determining the encoding strategy as a rule for skipping encoding the video block to be encoded according to the affine prediction encoding mode; in response to the test result being that the probability is greater than the probability threshold, determining the encoding strategy as a rule for allowing encoding the video block to be encoded according to the affine prediction encoding mode.
[0103] In this embodiment, in the process of determining the coding strategy of the video block to be coded according to the test result, the magnitude relationship between the probability and the probability threshold in the test result can be determined. If the probability is less than or equal to the probability threshold, the coding strategy can be determined to be a rule for skipping the coding of the video block to be coded according to the affine prediction coding mode. On the contrary, if the probability is greater than the probability threshold, the coding strategy can be determined to be a rule for allowing the coding of the video block to be coded according to the affine prediction coding mode.
[0104] Optionally, the above embodiment describes the logic of deciding whether to use affine prediction coding mode to encode a specific video block during video encoding. The above process of deciding the encoding strategy is based on the comparison of probability and probability threshold, aiming to optimize encoding efficiency and video quality.
[0105] Optionally, in this embodiment of the present application, the test result reflects the possibility or effectiveness of the affine prediction coding mode for the video block to be encoded, and the probability quantifies this possibility. The calculation of the probability is usually based on the characteristics of the video block (such as motion complexity, texture characteristics, etc.) and the mode selection and performance indicators of the encoded block. For example, if the appropriate coding mode of the same position, adjacent or parent coding unit is the affine mode, and their performance indicators (such as rate-distortion cost) are better than other modes, this will increase the probability that the video block to be encoded is also suitable for the affine prediction coding mode.
[0106] Optionally, the probability threshold is a preset criterion in the encoding decision process, which is used to determine whether the affine prediction coding mode should be tried. If the probability is less than or equal to the probability threshold, it means that the affine prediction coding mode may not be a suitable choice for the video block to be encoded, or the additional cost (such as computing time and bandwidth consumption) of using this mode may not be worthwhile. On the contrary, if the probability is greater than the probability threshold, it means that the affine prediction coding mode may be able to provide satisfactory compression performance and video quality, and it is worth trying in the encoding process.
[0107] Optionally, if the test result indicates that the probability is less than or equal to the probability threshold, the encoding of the current video block using the affine prediction coding mode will be skipped. This means that the encoder will directly consider other conventional coding modes, such as regularamvp, skip, etc., to quickly complete the encoding of the video block to be encoded, avoiding unnecessary complex calculations, thereby saving encoding time and improving encoding efficiency.
[0108] Optionally, if the probability is greater than the probability threshold, the encoder will allow the use of the affine prediction coding mode to attempt to encode the current video block. In the above case, relevant calculations of the affine prediction will be performed to determine whether the affine prediction coding mode can provide coding performance that meets the requirements. If the affine prediction coding mode can significantly improve compression efficiency, reduce bit rate, or reduce residuals while maintaining video quality, then the above affine prediction coding mode will be selected as a suitable coding strategy for the current video block.
[0109] In the embodiment of the present application, the above-mentioned probability and threshold-based coding strategy determination method makes full use of the spatiotemporal correlation of the video content, so that the video encoder (which can be referred to as the encoder) can intelligently select the coding mode according to the characteristics of the video block. It helps to balance the computational complexity and coding efficiency, avoid the use of the affine prediction coding mode with high computational cost when it is unnecessary, and when the video block motion is complex or there is significant deformation, it can ensure that a more advanced coding mode is used to improve the coding quality. Through the above method, the video encoder can not only handle static or translational motion scenes, but also efficiently cope with complex motions such as rotation, scaling and shearing, so as to achieve better coding performance on various video contents, which is an important means to improve video coding efficiency and video quality in practical applications.
[0110] As an optional implementation, step S304, determining at least one encoding mode corresponding to the video block to be encoded, includes: determining a target position of the video block to be encoded in the video file, and a target block size of the video block to be encoded; obtaining from the video file a video block encoded at the same position that has the same position as the target position, the same block size as the target block size, and has been encoded; and determining the encoding mode of the video block encoded at the same position.
[0111] In this embodiment, in the process of determining the coding mode of the video block to be coded, the target position of the video block to be coded in the video file and the target block size of the video block to be coded can be determined. The same position coded video block with the same position as the target position and the same block size as the target block size and which has been coded can be obtained from the video file. The coding mode of the same position coded video block can be determined as the coding mode. The same position coded video block can also be called a same position coding unit.
[0112] Optionally, in the video encoding process, determining the encoding mode corresponding to the video block to be encoded is one of the key steps to optimize the encoding decision and improve the encoding efficiency. The above implementation illustrates how to find and use the encoding mode of the encoded video block at the same position based on the position and size of the video block to be encoded to guide the encoding strategy of the current block.
[0113] Optionally, determine the target position and target block size of the video block to be encoded in the video sequence. The target position refers to the spatial coordinates of the video block in the current frame, and the target block size refers to the width and height of the video block. The above attributes are the basis for selecting the video block to be encoded at the same position, because the selection of the encoding mode is often closely related to the position and size of the video block.
[0114] Optionally, a video block having the same position and size as the video block to be encoded is searched from the video file (or encoded frame), i.e., a video block encoded at the same position. Since the video encoder usually adopts a block-based encoding strategy, the video blocks at the same position and size may experience similar motion or deformation in the previous and next frames. Therefore, the encoding mode of the video block encoded at the same position can be used as a good clue to predict the encoding mode of the current block.
[0115] Optionally, after finding the same-position coded video block, the coding mode used is analyzed. The coding mode may include but is not limited to affine merge, affine amvp, regular merge, regular amvp, skip, geo and mmvd modes. The coding mode of the same-position coded video block provides information about the possible appropriate coding strategy for the current block, because the video content is usually spatially correlated, that is, adjacent or same-position blocks may have similar motion characteristics or visual features.
[0116] Optionally, the coding mode of the video block coded at the same position is determined as the coding mode corresponding to the current video block to be coded. The coding mode can be used as a preliminary basis for coding decisions, helping the encoder to skip coding modes that perform poorly under similar conditions, thereby reducing invalid attempts and improving coding speed.
[0117] Optionally, in the subsequent coding decision process, the encoder can evaluate whether the video block to be encoded is suitable for trying the affine prediction coding mode based on the coding mode. If the coding mode indicates that the video block encoded at the same position uses affine prediction, the current video block may also be suitable for affine prediction and will be allowed to try the affine prediction mode. On the contrary, if the coding mode points to other modes, or the affine prediction mode does not perform well on the video block encoded at the same position, the encoder may skip affine prediction and test other coding modes instead.
[0118] In the embodiment of the present application, by determining the coding mode of the video block coded at the same position of the video block to be coded, clues based on spatiotemporal correlation are provided for the coding decision of the current video block, so that the encoder can select the coding mode more intelligently, avoiding unnecessary calculations and attempts, thereby improving the coding speed and efficiency. At the same time, the above method also reflects the emphasis on content adaptation and optimized coding selection in the video coding and decoding process to achieve better compression performance and video quality.
[0119] As an optional implementation, based on the correlation between the coding mode and the target coding mode corresponding to the video block to be encoded, the affine prediction coding mode is tested to obtain a test result, including: in response to the correlation indicating that the coding mode of the video block encoded at the same position is different from the affine prediction coding mode corresponding to the target coding mode, determining that the test result is a probability less than or equal to a probability threshold.
[0120] In this embodiment, in the process of testing the affine prediction coding mode based on the correlation between the coding mode and the target coding mode corresponding to the video block to be encoded, if the correlation indicates that the coding mode of the video block encoded at the same position is different from the affine prediction coding model corresponding to the target coding mode, it is determined that the test result is that the probability is less than or equal to the probability threshold.
[0121] Optionally, in the step of testing the affine prediction coding mode based on the correlation, the encoder tests whether the affine prediction coding mode is suitable for the current video block according to the correlation between the coding mode of the video block encoded at the same position and the target coding mode that may be adopted by the video block to be encoded, and obtains a probability value. The above probability value reflects the possibility or effectiveness of the current video block adopting the affine prediction coding mode.
[0122] Optionally, the encoder can evaluate the correlation between the coding mode of the same-position coded video block and the affine prediction coding mode. For example, if the same-position coded video block uses the affine merge or affine amp mode and the coding effect is good (i.e., the rate distortion cost is low), the probability of the current video block using the affine prediction coding mode will be increased.
[0123] Optionally, the probability in the test result is an estimate of whether the current video block will obtain good coding performance by adopting the affine prediction coding mode. If the correlation indicates that the coding mode of the video block encoded at the same position is inconsistent with the affine prediction coding mode, it means that under similar conditions, the affine mode may not be proven to meet the requirements. At this time, the encoder will determine that the test result is less than or equal to the probability threshold, indicating that the expected effect of the affine prediction coding mode for the current video block may not be good. The encoder presets a probability threshold to determine whether to try the affine prediction coding mode. If the probability is less than or equal to the probability threshold, the encoder will skip the affine prediction coding mode and try other more likely to be effective coding modes, such as regular AVP, skip, geo, etc., to reduce invalid calculations and attempts and increase the encoding speed.
[0124] For example, the encoder determines that there is a tried coding unit at the same position as the video block to be encoded, which means that at the position of the current video block, a block of the same size has been encoded before and its appropriate coding mode has been saved. It can be further checked whether the coding mode that meets the requirements of the coding unit at the same position is affine merge mode or affine amvp mode, both of which are instances of affine prediction coding mode. If the coding mode that meets the requirements is not one of the two affine modes, it implies that the affine mode may not be a suitable choice under similar video content and position conditions.
[0125] For another example, at this time, based on the existence of the above-mentioned co-located coded video blocks and the inspection of the coding modes that meet the requirements, the encoder can infer that the probability of using the affine prediction coding mode for the current video block to be coded is low. This is because the coding modes are spatially correlated, and the appropriate coding mode of the co-located coded video blocks is instructive for the coding decision of the current block. In the above case, the encoder can determine not to try the affine prediction coding mode, directly skip the 4param affine amvp and 6param affine amvp modes, and consider other coding modes instead. The above decision saves computing resources, avoids invalid attempts at the affine prediction model, and thus improves the overall performance of the encoder.
[0126] In the embodiment of the present application, the content adaptability and efficiency optimization in the video encoding algorithm are reflected through the above-mentioned decision logic. By utilizing the information of the encoded blocks, the attempts of complex encoding modes are reduced, thereby improving the encoding speed while maintaining the video quality.
[0127] As an optional implementation, at least one associated coded video block includes: at least one adjacent coded video block adjacent to the video block to be coded in the video file, and a parent coded video block of the video block to be coded in the video file, at least one coding mode includes: a tested coding mode of the video block to be coded, a coding mode of an adjacent coded video block, and a coding mode of a parent coded video block, and determining at least one coding mode corresponding to the video block to be coded from the video file, further comprising: in response to the correlation indicating that the coding mode of the coded video block at the same position is the same as the affine prediction coding mode, or, in response to a failure to obtain the coded video block at the same position, performing the following steps in a target order: a determining step of determining the tested coding mode; a first obtaining step of obtaining the coding mode of the adjacent coded video block; and a second obtaining step of obtaining the coding mode of the parent coded video block.
[0128] In this embodiment, the associated coded video blocks may include: adjacent coded video blocks (adjacent coding units) adjacent to the video block to be coded in the video file and the parent coded video block (parent coding unit) of the video block to be coded in the video file. The coding mode may include: the tested coding mode of the video block to be coded, the coding mode of the adjacent coded video block and the coding mode of the parent coded video block. In the process of determining the coding mode corresponding to the video block to be coded from the video file, if the correlation indicates that the coding mode of the coded video block at the same position is the same as the affine prediction coding mode, or the acquisition of the coded video block at the same position fails, the following steps may be performed in the target order: a determination step, in which the tested coding mode is determined; a first acquisition step, in which the coding mode of the adjacent coded video block may be acquired; a second acquisition step, in which the coding mode of the parent coded video block may be acquired.
[0129] Optionally, the above embodiment describes a strategy for deciding whether to try an affine prediction coding mode in video coding, the core of which is to use information of encoded video blocks to improve the efficiency and accuracy of coding decisions.
[0130] Optionally, the associated coded video blocks include at least one adjacent coded video block adjacent to the video block to be coded, and a parent coded video block of the video block to be coded. The coding modes of these associated coded video blocks (i.e., the modes used when they are coded) have reference value for the coding decision of the current video block.
[0131] Optionally, the coding mode includes the coding modes that have been tried for the current video block (tested coding modes), the coding modes of adjacent coded video blocks, and the coding modes of parent coded video blocks. The above coding mode set provides clues to possible suitable coding modes for the current video block.
[0132] Optionally, if the acquisition of the same-position coded video block fails, in the determination step, the coding modes that have been tried for the current video block can be checked. If the tested coding modes of the video block to be coded include affine merge, geo, or mmvd, and the coding effect of the above modes is better (such as the rate-distortion cost is lower than other modes), it can be preliminarily inferred that the affine prediction coding mode may be applicable to the current video block.
[0133] Optionally, if the coding mode corresponding to the unit to be coded determined in the above determination step does not match affinemerge, geo or mmvd, that is, if the coding mode is not any of the above three modes of affine merge, geo or mmvd, then the first acquisition step can be entered, that is, the coding mode of at least one coded video block adjacent to the video block to be coded is acquired. The coding mode of the adjacent coding block can provide further clues about the selection of the coding mode of the current video block, because video blocks at similar positions often have similar motion characteristics, and the coding mode of a block can predict the coding mode of its neighboring blocks to a certain extent.
[0134] Optionally, in the first acquisition step, it can be determined whether an adjacent coding unit of the current coding unit exists. If the adjacent coding unit exists, it can be determined whether the coding mode of the adjacent coding unit is affine merge or affineamvp mode. If the coding mode of the adjacent coding unit does not match the above-mentioned affine merge or affineamvp mode, that is, the coding mode of the adjacent coding unit is not any of the above-mentioned affine merge or affineamvp modes, then the second acquisition step can be entered, that is, the coding mode of the parent coding video block of the video block to be encoded is obtained. The coding mode of the parent coding video block can also provide information about the selection of the sub-block coding mode, because the sub-coding video block often inherits the motion characteristics or coding mode of the parent coding video block. If the parent coding video block uses the affine prediction coding mode and the effect is good, this may indicate that the sub-coding video block is also suitable for affine prediction coding.
[0135] Optionally, in the second acquisition step, it can be determined whether the current coding unit has a parent coding unit. If the parent coding unit exists, it can be determined whether the coding mode of the parent coding unit is affine merge or affine amp mode. If the coding mode is any of the above affine merge or affine amp modes, the 4-parameter affine amp mode can be tried, and the 6-parameter affine amp mode can be further tried.
[0136] In an embodiment of the present application, an attempt is made to obtain the coding mode of the video block coded at the same position. If the mode has a high correlation with the affine prediction coding mode, it will be determined that the current video block is likely to use affine prediction; if the acquisition of the video block coded at the same position fails, or its coding mode does not match the affine prediction, the encoder will instead consider the coding modes of the adjacent coded video blocks and the parent coded video blocks to find the correlation with the affine prediction coding mode. Through this process, the encoder can intelligently decide whether to try the affine prediction coding mode for the current video block based on the information of the coded video block, thereby reducing invalid attempts and improving coding efficiency. In summary, the above steps reflect the pursuit of content adaptability and efficiency in video coding technology, and guide the coding decision of the current video block by maximizing the use of the information of the coded video block, thereby reducing the coding time and improving the overall coding performance while ensuring the video quality.
[0137] As an optional implementation, the following steps are performed in a target order, including: executing the current step in the target order, wherein the current step is any one of a determination step, a first acquisition step, and a second acquisition step; in response to failure to execute the current step, or in response to the correlation indicating that the coding mode in the current step is different from the affine prediction coding mode, executing the next step of the current step in the target order; based on the correlation between the coding mode and the target coding mode corresponding to the video block to be encoded, testing the affine prediction coding mode to obtain a test result, including: in response to the correlation indicating that the coding mode in the last step in the target order is different from the affine prediction coding mode corresponding to the target coding mode, determining that the test result is a probability less than or equal to a probability threshold; the method also includes: in response to failure to obtain the associated coded video block in the last step in the target order, determining that the test result is a probability less than or equal to a probability threshold.
[0138] In this embodiment, the following steps can be performed in the target order: if the current step in the target order is any one of the determination step, the first acquisition step and the second acquisition step, it can be detected whether the current step fails to execute, and the correlation can also indicate whether the coding mode in the current step is the same as the affine prediction coding mode. If the current step fails to execute, or the correlation indicates that the coding mode in the current step is different from the affine prediction coding mode, the next step of the current step in the target order can be executed. If the correlation indicates that the coding mode in the last step in the target order is different from the affine prediction coding mode corresponding to the target coding mode, it can be determined that the test result is less than or equal to the threshold with probability. If the acquisition of the associated coded video block in the last step in the target order fails, it can be determined that the test result is less than or equal to the probability threshold with probability.
[0139] Optionally, the above embodiment describes an implementation method for determining whether to try an affine prediction coding mode based on the correlation of multiple coding modes in video coding. The above method checks the coding mode information related to the video block to be encoded step by step through a series of target order steps, and determines the test probability of the affine prediction coding mode based on the analysis of the above information.
[0140] Optionally, the target sequence includes three key steps: a determination step, a first acquisition step, and a second acquisition step, corresponding to checking the tested coding mode, acquiring the coding mode of the adjacent coded video block, and acquiring the coding mode of the parent coded video block, respectively.
[0141] Optionally, when executing a current step in the target order, the encoder will first try to complete the step. For example, in the determination step, the encoder checks whether certain encoding modes have been tried for the current video block. If the current step cannot be successfully executed (for example, the acquisition of adjacent or parent encoded video blocks fails), the encoder automatically turns to the next step in the target order.
[0142] Optionally, the encoder can evaluate the correlation between the coding mode and the affine prediction coding mode in each current step executed. If the correlation indicates that the coding mode is different from the affine prediction mode, the encoder will consider whether to skip the test of the affine prediction mode. Specifically, if in the last step of the target order (the second acquisition step or the step directly executed after the failure of the first acquisition step), the coding mode (such as the appropriate coding mode of the parent coding block) is inconsistent with the affine prediction mode, the encoder will determine that the probability of testing the affine prediction mode is less than or equal to a set probability threshold. This means that the encoder believes that the current video block is less likely to adopt the affine prediction coding mode, and therefore, it may decide to skip the attempt of the affine prediction mode to reduce invalid computational costs and increase encoding speed.
[0143] Optionally, the probability threshold is a key parameter for the encoder to decide whether to skip the affine prediction coding mode. If the probability of the affine prediction coding mode for the video block to be encoded is considered to be lower than or equal to this threshold based on the evaluation result of the correlation, the encoder will skip the test of the affine prediction mode and directly try other coding modes.
[0144] Optionally, if the last step in the target order fails to obtain the associated coded video block (for example, the parent coded video block cannot be found), the encoder will also skip the test of the affine prediction coding mode and determine that its probability is less than or equal to the probability threshold. This is because there is insufficient information to evaluate the effectiveness of affine prediction, and the encoder chooses to avoid making costly affine prediction attempts to ensure the efficiency of the encoding process.
[0145] In an embodiment of the present application, by utilizing the correlation between the current video block and the coding modes of the same position, adjacent and parent coding blocks, it is intelligently determined whether to try the affine prediction coding mode. It reflects the pursuit of content adaptability and efficiency optimization in video coding technology, and can reduce encoding time and improve overall encoding performance while maintaining video quality. By gradually checking and evaluating relevant coding modes, the encoder can more accurately determine the applicability of the affine prediction coding mode, avoiding unnecessary calculations, thereby achieving faster encoding speeds and better encoding effects in practical applications.
[0146] As an optional implementation, the target order is: executing the determination step, the first acquisition step, and the second acquisition step in sequence, executing the current step in the target order, including: taking the determination step as the current step to determine the tested coding mode; in response to the correlation indicating that the coding mode in the current step is different from the affine prediction coding mode, executing the next step of the current step in the target order, including: in response to the correlation indicating that the tested coding mode is different from at least one of the following coding modes, executing the next first acquisition step of the determination step in the target order to obtain the coding mode of the adjacent coded video block: the affine prediction coding mode, the geometric prediction coding mode, and the merged prediction mode based on the motion vector difference; in response to the failure of the current step to execute, or in response to the correlation indicating that the coding mode in the current step is different from the affine prediction coding mode, executing the next step of the current step in the target order, including: taking the first acquisition step as the current step, in response to the failure to obtain the adjacent coded video block, or in response to the correlation indicating that the coding mode of the adjacent coded video block is different from the affine prediction coding mode, executing the next second acquisition step of the first acquisition step in the target order to obtain the coding mode of the parent coded video block.
[0147] In this embodiment, the target order is to perform the determination step, the first acquisition step and the second acquisition step in sequence. If the current step is the determination step, the tested coding mode is determined. If the correlation indicates that the coding mode in the current step is different from the affine prediction coding mode, the next step can be executed. When the correlation indicates that the tested coding mode is different from one of the affine prediction coding mode, the geometric prediction coding mode and the motion vector coding mode, the first acquisition step can be executed to obtain the coding mode of the adjacent coded video block. If the current step is the first acquisition step, and the acquisition of the adjacent coded video block fails, or the correlation indicates that the coding mode of the adjacent coded video block is different from the affine prediction coding mode, the second acquisition step can be executed to obtain the coding mode of the parent coded video block. Among them, the affine prediction coding mode can be affine merge. The geometric prediction coding mode can be geo mode. The merge prediction mode based on motion vector difference can be mmvd mode.
[0148] Optionally, the above embodiment aims to determine whether to try the affine prediction coding mode for the video block to be coded by analyzing the information of the coded video block.
[0149] Optionally, the target order defines the decision process of the encoder when processing the video block to be encoded, and three key steps are performed in sequence: a determination step, a first acquisition step, and a second acquisition step. Each step determines whether to continue to execute the subsequent step or directly skip the attempt of the affine prediction coding mode based on the execution result or failure of the previous step.
[0150] Optionally, in the determination step, the encoder may check whether the tested coding mode of the video block to be encoded is similar or related to the affine prediction coding mode (e.g., affine merge) or other specific modes (e.g., geo, mmvd). If the tested coding mode is significantly different from one of these modes, it means that the current video block may not be suitable for affine prediction coding, and the encoder will skip the affine prediction mode and directly execute the next step.
[0151] Optionally, the first acquisition step involves acquiring the coding mode of the coded video block adjacent to the video block to be encoded. If the acquisition fails, or according to the correlation evaluation, the coding mode of the adjacent coded video block (e.g., affine merge, affine amvp, geo, mmvd) is inconsistent with the affine prediction coding mode, this may indicate that affine prediction coding is not a suitable choice for the current video block. At this time, the encoder will not attempt affine prediction, but will perform the next step in the target order.
[0152] Optionally, in the second acquisition step, if the first two steps fail to provide evidence supporting the affine prediction coding mode, the encoder will attempt to obtain the coding mode of the parent coded video block of the video block to be encoded. The coding mode of the parent coded video block can serve as another important reference point to evaluate the applicability of affine prediction coding for the current video block. If the acquisition of the parent block coding mode fails, or the coding mode of the parent coded video block is not related to the affine prediction coding mode, the encoder will also skip the attempt at affine prediction.
[0153] Optionally, in each step, the encoder can decide whether to continue with the subsequent steps or immediately skip the affine prediction coding mode based on the correlation evaluation. The correlation evaluation involves analyzing the relationship between the coding mode of the current video block and the affine prediction coding mode to determine whether affine prediction may be a suitable coding strategy. If the evaluation result indicates that the correlation with the affine prediction coding mode is low, the encoder will not try affine prediction, but will directly switch to other coding modes, etc., to save computing resources and increase encoding speed.
[0154] In an embodiment of the present application, the above-mentioned entire decision-making process starts with checking the tested coding mode of the video block, gradually obtaining the coding mode of the adjacent and parent coded video blocks, and intelligently deciding whether to try affine predictive coding through a series of evaluations. This method effectively utilizes the spatial correlation of video content and the potential connection between different coding modes to reduce invalid attempts and improve the efficiency and accuracy of coding decisions. Through the above-mentioned logical flow, the encoder can reduce encoding time and improve coding efficiency while maintaining video quality. This process reflects the pursuit of content adaptability and efficiency optimization in video coding and decoding technology. By minimizing the waste of computing resources and maximizing coding performance, the video encoder can make reasonable coding decisions more quickly when processing complex video content, thereby achieving more efficient data compression and transmission.
[0155] As an optional implementation, in response to the correlation indicating that the coding mode in the last step of the target sequence is different from the affine prediction coding mode corresponding to the target coding mode, determining that the test result is less than or equal to a probability threshold with probability, including: in response to the correlation indicating that the coding mode of the adjacent coded video block obtained is different from the affine prediction coding mode, and the correlation indicating that the coding mode of the parent coded video block obtained is different from the affine prediction coding mode, determining that the test result is less than or equal to the probability threshold with probability.
[0156] In this embodiment, the coding mode in the last step in the correlation representation target sequence is different from the affine prediction coding mode. In the process of determining the test result, if the correlation representation obtains the coding mode of the parent coding video block that is different from the affine prediction coding mode, and the correlation representation obtains the coding mode of the parent coding video block that is different from the affine prediction coding mode, it can be determined that the test result has a probability less than or equal to a probability threshold.
[0157] Optionally, in this embodiment, the encoder finally decides whether to try the affine prediction coding mode based on the coding mode correlation evaluation result in the last step in the target sequence. The target sequence here refers to a process of sequentially performing a determination step, a first acquisition step, and a second acquisition step, wherein each step involves acquiring coding mode information related to the video block to be encoded.
[0158] Optionally, the correlation indicates that the coding mode of the parent coded video block is different from the affine prediction coding mode, and in the last step (second acquisition step) of the target sequence, the encoder attempts to acquire the coding mode of the parent coded video block of the video block to be coded. If, according to the correlation evaluation, the coding mode of the parent coded block is significantly different from the affine prediction coding mode, this means that the possibility of using affine prediction coding for the current video block is low.
[0159] Optionally, the correlation indicates that the coding mode of the adjacent coded video block is different from the affine prediction coding mode. If the first acquisition step (acquiring the coding mode of the adjacent coded video block) is successfully executed, the encoder will evaluate the correlation between the coding mode of the adjacent coding block and the affine prediction coding mode. If the coding mode of the adjacent coding block is also inconsistent with the affine prediction coding mode, this further reduces the expected effect of the current video block adopting the affine prediction coding mode.
[0160] Optionally, if the evaluation results of the last step in the target order indicate that the coding modes of both the parent coded video block and the adjacent coded video block have low correlation with the affine prediction coding mode, the encoder determines that the probability of testing the affine prediction coding mode is less than or equal to a preset probability threshold. This probability threshold is a decision point for the encoder when analyzing the correlation of coding modes, and its value can be adjusted according to different performance requirements and coding scenarios.
[0161] Optionally, once the probability of testing the affine prediction coding mode is determined to be less than or equal to the probability threshold, the encoder will skip the attempt of the affine prediction coding mode and directly turn to other coding modes, such as regular AVP, SKIP, GEO or MMVD mode, to reduce invalid computational costs and improve encoding speed.
[0162] In an embodiment of the present application, by utilizing the coding mode information of the parent coding block and the adjacent coding block, the possibility of using the affine prediction coding mode for the video block to be encoded is comprehensively analyzed. It reflects the pursuit of content adaptability and efficiency optimization in video coding technology, aiming to maximize coding performance by minimizing the waste of computing resources. By setting the probability threshold, the encoder can intelligently decide whether to try the affine prediction coding mode while maintaining the video quality, thereby reducing the encoding time and improving the encoding efficiency. In practical applications, the above method can help the encoder make reasonable coding decisions more quickly and avoid unnecessary attempts when the video content and motion characteristics do not match the affine prediction coding mode. It not only saves valuable computing resources, but also ensures the efficiency and real-time performance of video coding, which is particularly suitable for processing large amounts of video data or for scenarios with strict requirements on encoding speed.
[0163] As an optional embodiment, in response to the correlation indicating that the coding mode in the last step in the target sequence is different from the affine prediction coding mode, determining that the test result is a probability less than or equal to a probability threshold includes: in response to a failure to obtain an adjacent coded video block, and the correlation indicating that the coding mode of the obtained parent coded video block is different from the affine prediction coding mode, determining that the test result is a probability less than or equal to the probability threshold.
[0164] In this embodiment, the coding mode in the last step in the correlation representation target sequence is different from the affine prediction coding mode. In the process of determining the test result, if the acquisition of the adjacent coded video block fails, and the correlation representation obtains the coding mode of the parent coded video block that is different from the affine prediction coding mode, it can be determined that the test result has a probability less than or equal to a probability threshold.
[0165] Optionally, in this embodiment, the possibility of adopting the affine prediction coding mode (such as affine merge or affine amvp) for the video block to be encoded can be evaluated based on a series of steps, and a decision can be made whether to perform the corresponding test. The above decision process pays special attention to the correlation between the coding mode of the last step in the target sequence and the affine prediction coding mode.
[0166] Optionally, the last step in the target sequence usually involves obtaining the coding mode information of the parent coded video block as a final reference for evaluating the applicability of the affine prediction coding mode. If no evidence supporting the affine prediction coding mode is found in the previous steps (such as determining the tested coding mode, obtaining the coding mode of the adjacent coded video block, etc.), then the evaluation result of this last step will play a decisive role in whether to test the affine prediction coding mode.
[0167] Optionally, in the first acquisition step, the encoder attempts to acquire the coding mode of a coded video block adjacent to the video block to be coded. If this attempt fails, it may be because the adjacent video block has not been coded yet, or its coding mode information is not available at the current moment. In this case, the encoder cannot obtain clues about the applicability of the affine prediction coding mode from the adjacent video blocks, and therefore will rely on other sources of information, such as the coding mode of the parent coded video block.
[0168] Optionally, in the second acquisition step, if the encoder successfully acquires the coding mode of the parent coded video block, and the correlation evaluation indicates that the coding mode of the parent block does not match the affine prediction coding mode, this means that the current video block may inherit the motion characteristics of the parent block, which are not suitable for affine prediction coding. For example, if the parent block uses skip mode, geo mode, or mmvd mode, this may suggest that the motion of the current block is more inclined to translation or local change rather than affine motion such as rotation, scaling, or shearing.
[0169] Optionally, if in the last step of the target order, the result of the correlation evaluation between the coding mode (i.e., the coding mode of the parent coded video block) and the affine prediction coding mode indicates that the two do not match, and the previous attempt to obtain the adjacent coded video block failed, then the encoder will determine that the probability of testing the affine prediction coding mode is less than or equal to a preset probability threshold. This threshold is a key parameter used by the encoder to decide whether to try the affine prediction coding mode, which is intended to avoid invalid attempts when the video content does not match the mode, thereby improving overall coding efficiency.
[0170] Optionally, once the encoder determines that the probability of testing the affine prediction coding mode is less than or equal to the set probability threshold, it will choose to skip the test of the affine prediction coding mode and try other coding modes instead. This decision is based on intelligent analysis of video content and motion characteristics, which can significantly reduce encoding time and improve encoding speed while maintaining video quality.
[0171] In an embodiment of the present application, by setting a decision logic, the encoder can intelligently decide whether to test the affine prediction coding mode based on the coding mode information of the parent coded video block and the failure of trying to obtain the coding mode of the adjacent coded video block. This method reflects the pursuit of content adaptability and efficiency optimization in video coding technology, and can help the encoder make coding decisions more efficiently when processing large amounts of video data, avoid unnecessary computing costs, and ensure the real-time and efficiency of video coding. Through such a logical flow, the encoder can maximize its performance, especially in scenarios where the video content does not match the affine prediction coding mode, and can quickly switch to a more suitable coding strategy.
[0172] As an optional embodiment, in response to a failure to obtain the associated coded video block in the last step in the target sequence, determining that the test result is a probability less than or equal to a probability threshold includes: in response to the correlation indicating that the coding mode of the adjacent coded video block obtained is different from the affine prediction coding mode, and the acquisition of the parent coded video block fails, determining that the test result is a probability less than or equal to the probability threshold.
[0173] In this embodiment, when the associated coded video block in the last step in the target sequence fails to be obtained, in the process of determining the test result, if the correlation indicates that the coding mode of the acquired adjacent coded video block is different from the affine prediction coding mode, and the parent coded video block fails to be obtained, the test result is determined to be a probability less than or equal to the probability threshold.
[0174] Optionally, this embodiment describes how a video encoder, when trying to determine a suitable coding mode for a video block to be coded, evaluates the test probability of an affine prediction coding mode based on obtaining coding mode information of an associated coded video block (such as an adjacent coded video block or a parent coded video block). Specifically, when the encoder fails to successfully obtain the coding mode of the parent coded video block, and the coding mode of the obtained adjacent coded video block does not match the affine prediction coding mode, the encoder will make a probability evaluation decision.
[0175] Optionally, during the execution of the target sequence, the encoder attempts to obtain the coding mode information of the co-located, adjacent, and parent coded video blocks. If the last step of trying to obtain the coding mode of the parent coded video block fails, the encoder will not be able to use the coding mode of the parent block to assist in determining the applicability of the affine prediction coding mode. This may occur in the early stages of video coding, when the parent block has not yet been encoded, or when the parent block information is not available due to other reasons (such as the special structure of the coding unit).
[0176] Optionally, even if the coding mode information of the parent coded video block is not available, the encoder can still obtain the coding mode information from the adjacent coded video block. If the coding mode of the adjacent coded video block is significantly different from the affine prediction coding mode (such as affine AVP or affine merge), this means that the motion characteristics of the current video block may not be suitable for affine prediction coding. For example, if the adjacent block uses regular AVP, skip, geo or mmvd mode, this may mean that the movement of objects in the video is more consistent with the translation, local deformation or motion based on motion vector difference described by these modes, rather than the rotation, scaling or shearing motion that affine prediction coding is good at.
[0177] Optionally, in the case where the acquisition of the coding mode information of the parent coded video block fails and the coding mode of the acquired adjacent coded video block does not match the affine prediction coding mode, the encoder will determine that the probability of testing the affine prediction coding mode is less than or equal to a preset probability threshold. This probability threshold is a decision point for the encoder to perform probability evaluation based on the existing coding modes of adjacent coded blocks when there is a lack of sufficient coding mode information. If the probability is less than or equal to the threshold, the encoder will consider that the affine prediction coding mode may not be a suitable choice for the current video block.
[0178] Optionally, once it is determined that the probability of testing the affine prediction coding mode is less than or equal to the probability threshold, the encoder will skip the attempt of the affine prediction coding mode and try other coding modes instead. This decision logic can help the encoder make suboptimal choices based on limited clues when there is a lack of key coding mode information, avoiding high-cost affine prediction coding attempts in complex or uncertain coding scenarios, thereby improving overall coding efficiency and speed.
[0179] In an embodiment of the present application, by setting a flexible decision-making process, the video encoder can intelligently evaluate the test probability of the affine prediction coding mode based on the acquired coding mode information of the adjacent coded video blocks and the failure when trying to obtain the parent coded video block information. This method is particularly important when dealing with the boundary conditions of video coding. When the encoder cannot obtain enough information, it can make reasonable decisions based on existing clues to reduce invalid attempts and improve coding performance. Through such a logical process, the encoder can significantly reduce the encoding time while maintaining the video quality, improve the real-time and efficiency of video encoding, and is suitable for various video encoding scenarios, especially when the video content does not fully match the affine prediction coding mode.
[0180] As an optional implementation, in response to a failure to obtain the associated coded video block in the last step in the target sequence, determining that the test result is a probability less than or equal to a probability threshold includes: in response to a failure to obtain an adjacent coded video block and a failure to obtain a parent coded video block, determining that the test result is a probability less than or equal to the probability threshold.
[0181] In this embodiment, in the process of determining the test result when the associated coded video block in the last step in the target sequence fails to be obtained, if the adjacent coded video block fails and the parent coded video block fails to be obtained, the test result can be determined to be less than or equal to the probability threshold.
[0182] Optionally, the above embodiment describes the decision logic when the video encoder encounters a situation where the coding mode information of the adjacent coded video blocks and the parent coded video block cannot be obtained when trying to determine the appropriate coding mode of the to-be-coded video block.
[0183] Optionally, during the video encoding process, the encoder attempts to obtain coded video block information related to the video block to be encoded in a target order, including coding modes of adjacent coded video blocks and parent coded video blocks. Such information helps the encoder evaluate the applicability of different coding modes to the current video block, especially the applicability of the affine prediction coding mode.
[0184] Optionally, in the target order, the encoder usually first attempts to obtain the coding mode information of the adjacent coded video block. If this attempt fails, it may be because the adjacent block has not been encoded yet, or the information of the adjacent block is not available for some reason (such as the block does not exist due to different block segmentation paths). This makes it impossible for the encoder to assist in determining the appropriate coding mode for the current video block based on the coding mode of the adjacent block.
[0185] Optionally, if after failing to obtain the information of the adjacent coded video blocks, the encoder attempts to obtain the coding mode information of the parent coded video block, but also fails, this means that the encoder cannot obtain information about the possible appropriate coding mode of the current video block from the perspective of spatial proximity or from the perspective of hierarchy. The lack of parent block information may be due to the fact that the current block is at the beginning of the video frame and the parent block has not been encoded, or the special division method of the current block makes the parent block information unavailable, that is, the coding mode of the parent coding unit is not obtained, which may be caused by the different coding decision orders of the parent coding unit and the corresponding child coding unit.
[0186] Optionally, when the encoder fails to obtain the coding mode information of both the adjacent coded video block and the parent coded video block, it will determine that the probability of testing the affine prediction coding mode is less than or equal to a preset probability threshold based on the currently available limited information. This decision reflects that in the case of insufficient information, the encoder tends to adopt a conservative strategy to reduce high-cost affine prediction coding mode attempts. The preset probability threshold can be a low value, indicating that in the absence of key information, the affine prediction coding mode is determined to be unlikely to be a suitable coding mode for the current video block.
[0187] Optionally, once it is determined that the probability of testing the affine prediction coding mode is less than or equal to the probability threshold, the encoder will choose to skip the attempt of the affine prediction coding mode and directly switch to other coding modes, such as regular AMVP, skip, geo, or mmvd mode. This decision logic ensures that the encoder does not make costly affine prediction coding attempts without sufficient information, thereby improving the efficiency and speed of the entire encoding process.
[0188] In an embodiment of the present application, a decision process is set to handle the situation where the associated coded video block information fails to be obtained in video coding. In the case where attempts to obtain the coding mode information of adjacent and parent coded video blocks fail, the encoder will evaluate the test probability of the affine prediction coding mode based on a preset probability threshold, and make corresponding attempts or skip decisions. This method is particularly important when dealing with edge cases of video coding. When the encoder cannot obtain enough information for accurate evaluation, it can make conservative choices based on existing clues, avoid invalid attempts, and ensure the overall performance and efficiency of video coding. Through the above-mentioned decision logic, the encoder can reduce encoding time while maintaining video quality, improve the real-time performance and efficiency of video encoding, and is suitable for a variety of scenarios of video encoding.
[0189] As an optional implementation, the target sequence is: executing the second acquisition step, the first acquisition step, and the determination step in sequence; or, executing the first acquisition step, the determination step, and the second acquisition step in sequence.
[0190] In this embodiment, the target order may be to sequentially execute the second acquisition step, the first acquisition step, and the determination step, or may be to sequentially execute the first acquisition step, the determination step, and the second acquisition step.
[0191] Optionally, different target order implementations describe how a video encoder adjusts the order of obtaining coding mode information of coding blocks at different positions when trying to determine a suitable coding mode for a video block to be encoded.
[0192] Optionally, in the target order in which the second acquisition step is prioritized, the encoding mode of the parent encoded video block can be obtained through the second acquisition step. In the process of executing the second acquisition step, the encoder attempts to obtain the encoding mode of the parent encoded video block of the video block to be encoded. The above steps make full use of the hierarchical relevance of the video content. The encoding mode of the parent encoded video block often has a greater impact on the encoding mode of the child encoded video block. Therefore, it is possible to judge in advance whether the affine prediction encoding mode may be applicable to the current video block. If the encoding mode of the parent encoding block is significantly different from the affine prediction encoding mode, it may imply that the motion characteristics of the current video block are not suitable for affine prediction encoding. At this time, the encoder can skip the attempt of the affine prediction encoding mode and proceed directly to the next link. If the acquisition of the encoding mode of the parent encoded video block fails or the information is insufficient to make a decision, the encoder will enter the next link.
[0193] Optionally, after the second acquisition step, a first acquisition step may be performed, in which the encoder attempts to obtain the coding mode of the adjacent coded video block of the video block to be encoded. The above steps further utilize the spatial correlation of the video content, and the coding mode of the adjacent block can provide additional clues about the motion characteristics of the current block. If the coding mode of the adjacent coded video block does not match the affine prediction coding mode, the encoder may further reduce the probability of testing the affine prediction coding mode. If the coding mode of the adjacent coded block cannot be obtained, or the information is still insufficient to make a decision, the encoder will proceed to the next step.
[0194] Optionally, after the first acquisition step, a determination step may be performed to analyze the tested coding modes. In this process, the encoder analyzes the coding modes that have been tried for the video block to be encoded and the results thereof. If no mode related to or having a high degree of match with the affine prediction coding mode is found in the tested coding modes, the encoder may determine that the probability of testing the affine prediction coding mode is low, and choose to skip the attempt of the affine prediction coding mode.
[0195] Optionally, in the target order in which the first acquisition step is prioritized, the encoding mode of the adjacent encoded video block can be acquired through the first acquisition step. In the process of executing the first acquisition step, the encoder can try to acquire the encoding mode of the encoded video block adjacent to the video block to be encoded, and use the spatial proximity information to preliminarily determine the applicability of the affine prediction encoding mode. In the process of executing the determination step, the encoder analyzes the encoding modes and results that have been tried for the video block to be encoded, and further evaluates the potential value of the affine prediction encoding mode. In the process of executing the second acquisition step, the encoder attempts to acquire the encoding mode of the parent encoded video block, and uses the hierarchical structure information as the final decision basis.
[0196] In the embodiment of the present application, the key to the above two target order implementation methods is that by adjusting the order of obtaining the encoding mode information of the parent encoded video block, the adjacent encoded video block, and analyzing the tested encoding mode, the encoder can more flexibly evaluate the applicability of the affine prediction encoding mode according to the context information of the current video block. The first order (the second acquisition step is preferred) may be more suitable for the scene when the hierarchical relationship of the video block has a strong implication on the encoding mode; while the second order (the first acquisition step is preferred) may be more effective when the spatial proximity information plays a dominant role in the encoding decision.
[0197] In summary, through the above logical process, the encoder can reduce encoding time and improve encoding efficiency while maintaining video quality. Each order has its advantages and applicable scenarios. The encoder can choose the most suitable order according to the needs of specific video content and encoding strategy to achieve encoding performance that meets the requirements.
[0198] The present application also provides a video encoding method. Figure 4 is a flowchart of a video encoding method according to an embodiment of the present application, such as Figure 4 As shown, the method may include the following steps:
[0199] Step S402: Obtain at least one to-be-encoded video block from a video file.
[0200] Step S404, determining a coding strategy for the video block to be encoded, wherein the coding strategy is used to represent a rule for encoding the video block to be encoded, and the coding strategy is determined according to a test result, the test result is used to represent a probability that the affine prediction coding mode becomes a target coding mode corresponding to the video block to be encoded, the test result is obtained by testing the affine prediction coding mode based on a correlation between at least one coding mode corresponding to the video block to be encoded and the target coding mode, the performance index of the target coding mode for encoding the video block to be encoded is greater than a performance index threshold, the coding mode is a mode allowed to be adopted by at least one associated coding video block associated with the video block to be encoded during the encoding process, or the coding mode is a mode that has been tested during the encoding process of the video block to be encoded.
[0201] Step S406: Encode the video block to be encoded according to the encoding strategy.
[0202] In the technical solution provided in the above step S406 of the present application, the determined encoding strategy can be used to encode the video block to be encoded.
[0203] Optionally, if the encoding mode corresponding to the encoding strategy is the affine merge mode, the video block to be encoded can be encoded by the following steps: in the process of selecting motion control points, two or three motion control points are selected for a coding unit to be encoded. In the 4-parameter (4param) affine model, the points at the upper left corner and the upper right corner are selected; in the 6-parameter (6param) affine model, in addition to the upper left and upper right corner points, the lower left corner point is also considered. The motion vectors of these control points will be used to describe the affine motion of the entire coding unit; in the process of obtaining the MV candidates of the control points, the motion information of the spatially encoded blocks is used according to multiple derivation models to derive the MV candidates of each control point in the coding unit; in the process of calculating the affine transformation, a suitable control point motion vector (Motion Vector, referred to as MV) candidate can be selected to calculate the motion vector of each pixel in the entire coding unit. The above calculation can be based on the affine transformation model and can accurately capture the complex motions such as rotation, scaling and shearing in the video content; in the process of generating the prediction block, the calculated affine transformation motion vector is used to generate the prediction block from the reference frame, and motion compensation is performed to reduce the residual; in the process of motion information encoding, the coding parameters such as the motion vector control point index information are packaged to generate a bit stream signal. In the process of residual signal encoding, the motion compensated residual signal is transformed, quantized and entropy encoded to further reduce the size of the encoded bit stream.
[0204] Optionally, if the encoding mode corresponding to the encoding strategy is the affine amp mode, the video block to be encoded can be encoded by the following steps: in the process of determining the control points, two or three control points (Control Points) are determined for a coding unit to be encoded, depending on whether a 4-parameter affine model or a 6-parameter affine model is used. The 4-parameter mode usually selects the upper left corner and the upper right corner as control points, while the 6-parameter mode additionally includes the lower left corner as the third control point; in the process of motion vector search, an affine motion vector search is performed to find the control point motion vector that best describes the pixel motion within the coding unit; in the process of affine transformation calculation, the motion vector of each pixel in the entire coding unit is calculated using the motion vector of the control point and the affine transformation model. The affine transformation model can describe more complex motions, such as rotation, scaling, and shearing, which makes the prediction more accurate and can better remove temporal redundancy; in the process of prediction block generation and residual calculation, the determined motion vector is used to generate a prediction block from the reference frame. Then the residual between the current coding unit and the predicted block is calculated, that is, the difference in pixel values; in the process of motion vector prediction, once the appropriate control point motion vector is determined, the appropriate control point motion vector is predicted based on the encoded information, so that only the control point motion vector difference needs to be transmitted; in the process of residual signal and motion vector difference encoding, the residual signal is transformed, quantized and entropy encoded, and the control point motion vector difference is encoded at the same time to reduce the size of the encoded bit stream. This step is the general process of residual coding in video coding, which aims to efficiently represent the remaining information.
[0205] Optionally, if the coding mode corresponding to the coding strategy is the regular merge mode, the video block to be coded can be coded by the following steps: In the process of motion vector inheritance, the motion vector of the current coding unit does not need to be obtained through time-consuming motion search, but is directly inherited from its spatially adjacent coded block or temporally coded co-located block. This means that the current block uses the same motion vector as the adjacent block or temporally coded co-located block to predict the position of the current block in the reference frame; since the motion vector is inherited, there is no need to transmit motion vector information additionally during the coding process. This helps to reduce the amount of data in the bitstream and improve the coding speed and efficiency; in the process of generating the prediction block, the inherited motion vector is used to generate the prediction block from the reference frame. The prediction block is obtained through the motion compensation process; in the process of motion information coding, the coding parameters such as the motion vector index information are packaged to generate a bitstream signal for transmission or storage; in the residual coding process, the residual between the current coding unit and the prediction block is calculated. The residual is the difference between the pixel values of the current block and the prediction block, which is used to represent information that cannot be predicted after motion compensation. The above residual will be further encoded, usually including transformation, quantization and entropy coding, to reduce the number of bits stored and transmitted.
[0206] Optionally, if the coding mode corresponding to the coding strategy is skip mode, the video block to be coded can be encoded through the following steps: In the process of motion vector inheritance, the motion vector of the current coding unit does not need to be obtained through time-consuming motion search, but is directly inherited from its spatially adjacent coded block or temporal co-located block. This means that the current block uses the same motion vector as the adjacent block or temporal co-located block to predict the position of the current block in the reference frame; since the motion vector is inherited, there is no need to transmit additional motion vector information during the encoding process. This helps to reduce the amount of data in the bitstream and improve the encoding speed and efficiency; in the process of prediction block generation, the inherited motion vector is used to generate a prediction block from the reference frame. The prediction block is generated through a motion compensation process; in the process of motion information encoding, the encoding parameters such as motion vector index information are packaged to generate a bitstream signal for transmission or storage.
[0207] Optionally, if the coding mode corresponding to the coding strategy is geo mode, the video block to be coded can be encoded by the following steps: in the process of geometric segmentation, a larger coding unit can be divided into two or more sub-regions, and the above segmentation method can be flexibly adjusted according to the shape and size of the prediction area; in the process of motion vector prediction, each sub-region can independently obtain motion information, that is, each sub-region can have its own motion vector. The motion vector of each sub-region can be inherited from the motion vector of the adjacent coded block; in the process of motion compensation, the motion vector of each sub-region is used to obtain the prediction block from the reference frame, and compared with the pixel value of the current sub-region, and motion compensation is performed to reduce the prediction residual; in the process of motion information encoding, the coding parameters such as motion vector index information are packaged to generate a bit stream signal for transmission or storage. In the process of residual coding, if there is a residual (that is, the difference between the pixel value of the current sub-region and the pixel value of the prediction block), the above residual will be encoded. Residual coding usually involves transformation, quantization and entropy coding processes.
[0208] Optionally, if the encoding mode corresponding to the encoding strategy is the mmvd mode, the video block to be encoded can be encoded by the following steps: in the process of motion vector prediction, the motion information of the block encoded in the spatiotemporal domain is used, combined with the characteristics of the current block to be encoded, to generate an initial motion vector prediction; in the process of motion vector difference, try to fine-tune the prediction by adding a small amount of predefined motion vector difference (Δmv) to obtain a more accurate motion vector. This step searches a set of possible Δmv values to find the appropriate one to minimize the prediction error; in the process of motion vector encoding, only the Δmv value is encoded instead of the complete motion vector. Because the initial predicted motion vector has been encoded in the motion information of the adjacent block, only transmitting the differential value can greatly reduce the amount of information in the encoded bitstream; in the process of prediction block generation, the adjusted motion vector is used to generate a prediction block from the reference frame to perform motion compensation and reduce the residual; in the process of motion information encoding, the encoding parameters such as motion vector index information are packaged to generate a bitstream signal for transmission or storage; in the process of residual signal encoding, the residual after motion compensation is transformed, quantized and entropy encoded to further reduce the amount of data transmitted.
[0209] Through the above steps S402 to S406 of the present application, at least one video block to be encoded is obtained from the video file; the encoding strategy of the video block to be encoded is determined; and the video block to be encoded is encoded according to the encoding strategy, thereby achieving the technical effect of improving the efficiency of video encoding and solving the technical problem of low video encoding efficiency.
[0210] The present application also provides another video information processing method. Figure 5 is a flow chart of a video information processing method according to an embodiment of the present application, such as Figure 5 As shown, the method may include the following steps:
[0211] Step S502: acquiring at least one to-be-encoded video block from a video file by calling a first interface, wherein the first interface includes a first parameter, and a parameter value of the first parameter includes the to-be-encoded video block.
[0212] Step S504, determining at least one coding mode corresponding to the video block to be coded, wherein the coding mode is a mode allowed to be adopted by at least one associated coded video block associated with the video block to be coded during the coding process, or the coding mode is a mode that has been tested during the coding process of the video block to be coded.
[0213] Step S506, based on the correlation between the coding mode and the target coding mode corresponding to the video block to be encoded, the affine prediction coding mode is tested to obtain a test result, wherein the performance index of the target coding mode for encoding the video block to be encoded is greater than the performance index threshold, and the test result is used to indicate the probability that the affine prediction coding mode becomes the target coding mode.
[0214] Step S508: Determine the coding strategy of the video block to be coded according to the test result, wherein the coding strategy is used to indicate a rule for coding the video block to be coded.
[0215] Step S510: outputting the encoding strategy by calling the second interface, wherein the second interface includes a second parameter, and a parameter value of the second parameter includes the encoding strategy.
[0216] Through the above steps S502 to S512 of the present application, at least one video block to be encoded is obtained from the video file by calling the first interface; at least one encoding mode corresponding to the video block to be encoded is determined; based on the correlation between the encoding mode and the target encoding mode corresponding to the video block to be encoded, the affine prediction encoding mode is tested to obtain a test result; according to the test result, the encoding strategy of the video block to be encoded is determined; and the encoding strategy is output by calling the second interface, thereby achieving the technical effect of improving the efficiency of video encoding and solving the technical problem of low video encoding efficiency.
[0217] According to an embodiment of the present application, an embodiment of a video information processing system is also provided. Figure 6 is a schematic diagram of a video information processing system according to an embodiment of the present application, such as Figure 6 As shown, the video information processing system 600 may include: a video encoder 601 and a server 602 .
[0218] The video encoder 601 is used to obtain at least one to-be-encoded video block from a video file; determine an encoding strategy for the to-be-encoded video block, wherein the encoding strategy is used to represent a rule for encoding the to-be-encoded video block, and the encoding strategy is determined according to a test result, and the test result is used to represent the probability that the affine prediction encoding mode becomes the target encoding mode corresponding to the to-be-encoded video block, and the test result is obtained by testing the affine prediction encoding mode based on the correlation between at least one encoding mode corresponding to the to-be-encoded video block and the target encoding mode, and the performance index of the target encoding mode for encoding the to-be-encoded video block is greater than a performance index threshold, and the encoding mode is a mode allowed to be adopted by at least one associated encoding video block associated with the to-be-encoded video block during the encoding process, or the encoding mode is a mode that has been tested during the encoding process of the to-be-encoded video block; and encode the to-be-encoded video block according to the encoding strategy.
[0219] The server 602 is used to obtain the encoded video blocks to be encoded.
[0220] In this embodiment, a video information processing method system is provided, wherein at least one video block to be encoded is obtained from a video file through a video encoder 601; an encoding strategy of the video block to be encoded is determined; and the video block to be encoded is encoded according to the encoding strategy. The encoded video block to be encoded is obtained through a server 602. Thus, the technical effect of improving the efficiency of video encoding is achieved, and the technical problem of low efficiency of video encoding is solved.
[0221] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application, for example, the data for verification, are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with the relevant laws, regulations and standards of relevant countries and regions, and provide corresponding operation entrances for users to choose to authorize or refuse.
[0222] At present, the past decade has witnessed the rapid development of the video industry, and emerging businesses such as short videos, low-latency live broadcasts, and video conferencing have gradually emerged. Compared with media such as pictures and texts, videos can bring a more intuitive experience and have gradually become the most important way for people to spread information in work, study, entertainment, and leisure. In addition, the expansion of video services is not limited to this and will maintain a high-speed development momentum in the future.
[0223] In addition, higher-definition and more realistic 4K, 8K, and HDR videos continue to emerge. Behind this is the exponential growth of data volume, and the requirements for transmission bandwidth and storage space are getting higher and higher. Video coding, as a core part of multimedia technology, can significantly reduce the size of video files. Video coding is divided into lossy coding and lossless coding, among which lossy coding is more concerned because it can provide higher compression rates. An excellent lossy coding scheme is to reduce the bit rate of the video as much as possible without significantly reducing the perception of the human eye. Higher compression rates mean lower bandwidth costs and higher-definition and smoother visual experiences.
[0224] Therefore, major domestic companies have begun to develop video encoders with independent intellectual property rights, hoping to create technical barriers and use better encoding quality and lower transcoding prices to attract more video-related businesses.
[0225] Affine prediction in video codecs is an inter-frame prediction algorithm based on affine transformation. It uses an affine model to more accurately describe and predict object motion, and can effectively remove redundant temporal information in rotation, scaling, shearing, etc. Due to its high algorithm complexity, the affine prediction algorithm is one of the important algorithms that affect the encoder speed.
[0226] In order to reduce the invalid attempts of the affine prediction algorithm according to the picture content, the affine prediction coding fast algorithm in the related art usually uses whether the appropriate coding mode of the parent coding unit is a certain specified mode, or the rate distortion cost of the mode that the coding unit has tried. Although the above method can speed up the encoding process to a certain extent, it can only use limited information and ignores the richer correlation between coding modes, which limits the potential for further improving the encoding speed. Therefore, there is still a technical problem of low efficiency of video coding.
[0227] Furthermore, the present application provides a fast affine prediction coding algorithm, and proposes a fast affine prediction coding algorithm based on the correlation of multiple coding modes, which reduces invalid attempts of affine prediction and improves the encoder speed by using the appropriate coding modes of the coding units that have been tried at the same position, the appropriate coding modes that have been tried by the coding units to be coded, the appropriate coding modes of the adjacent coding units, and the appropriate coding modes of the parent coding units. This achieves the technical effect of improving the efficiency of video coding and solves the technical problem of low video coding efficiency.
[0228] The above method of this embodiment is further introduced below.
[0229] FIG7( a ) is a schematic diagram of a square coding unit division type. As shown in FIG7( a ), each square coding unit can be divided into quadtrees, vertical binary divisions, horizontal binary divisions, vertical trifurcations, and horizontal trifurcations.
[0230] FIG7( b ) is a schematic diagram of a rectangular coding unit division type. As shown in FIG7( b ), for the rectangular sub-blocks of binary division and trifurcated division, vertical binary division, horizontal binary division, vertical trifurcated division and horizontal trifurcated division can be nested.
[0231] Optionally, for each coding unit, the H.266 / VVC standard proposes a variety of coding algorithms to efficiently encode rich and diverse picture content. Among them, the affine prediction algorithm uses an affine model to more accurately describe and predict object motion, which can effectively remove redundant temporal information in picture content such as rotation, scaling, and shearing. There are two affine motion models in the VVC standard: a 4-parameter model and a 6-parameter model. The 4-parameter model can be used for picture content such as rotation and scaling. By using the upper left corner motion vector and the upper right corner motion vector of the coding unit as motion control points, the motion vector of each 4×4 small block in the coding unit is obtained according to the affine motion model, and motion compensation is performed on each 4×4 small block. The 6-parameter model uses the lower left corner motion vector as the third motion control point compared to the 4-parameter model, which can be used for picture content such as rotation, scaling, and shearing.
[0232] Optionally, the prediction algorithm based on the affine motion model includes an affine merge mode and an affine amp mode. The affine merge mode uses the motion vector information of adjacent coded units to perform affine motion compensation of the unit to be coded. Since no motion search is performed, the computational complexity is low. The affine amp mode obtains a suitable affine motion vector through a motion search based on gradient descent, and the computational complexity is high.
[0233] In order to reduce invalid attempts of affine AVP mode according to the content of the picture to be encoded, the existing affine prediction coding fast algorithm usually uses whether the appropriate coding mode of the parent coding unit is a certain specified mode, or the rate-distortion cost of the mode that has been tried by the unit to be encoded, to skip the affine prediction. This algorithm proposes a fast affine prediction coding algorithm based on the correlation of multiple coding modes. Through the appropriate coding mode of the coding unit that has been tried at the same position, the appropriate coding mode that has been tried by the unit to be encoded, the appropriate coding mode of the adjacent coding unit and the appropriate coding mode of the parent coding unit, the invalid attempts of affine prediction are further reduced and the encoder speed is improved. At present, this solution has been integrated into the S266 encoder and has been put into large-scale online application.
[0234] Figure 8 is a flowchart of a process for determining an affine prediction coding mode according to an embodiment of the present application, such as Figure 8 As shown, the process may include the following steps:
[0235] Step S801, testing the skip mode, geo mode, mmvd mode and regular merge mode.
[0236] In this embodiment, for the input unit to be encoded, the encoder tries the skip mode, geo mode, mmvd mode and regular merge mode respectively. After comparing the rate distortion cost of each encoding mode, the minimum rate distortion cost is saved, and the mode corresponding to the minimum rate distortion cost is recorded as the appropriate encoding mode.
[0237] Step S802, testing the affine merge mode.
[0238] In this embodiment, for the input unit to be encoded, the encoder tries the affine merge mode, compares the rate distortion cost of the affine merge mode with the minimum rate distortion cost of the above multiple modes, updates the minimum rate distortion cost, and records the encoding mode corresponding to the minimum rate distortion cost as the appropriate encoding mode.
[0239] Step S803, testing the regular AVP mode.
[0240] In this embodiment, for the input unit to be encoded, the encoder tries the regular AVP mode, compares the rate-distortion cost of the regular AVP mode with the minimum rate-distortion cost of the affine merge link, updates the minimum rate-distortion cost, and records the encoding mode corresponding to the minimum rate-distortion cost as the appropriate encoding mode.
[0241] Step S804, testing the affine amvp mode.
[0242] In this embodiment, for the input unit to be encoded, the encoder tries the affine AVP mode, compares the rate-distortion cost of the affine AVP mode with the minimum rate-distortion cost of the regular AVP link, updates the minimum rate-distortion cost, and records the encoding mode corresponding to the minimum rate-distortion cost as the appropriate encoding mode.
[0243] Step S805, testing other prediction modes.
[0244] In this embodiment, for the input unit to be encoded, the encoder tries other prediction modes, such as an adaptive motion vector accuracy mode, compares the rate-distortion cost of other prediction modes with the minimum rate-distortion cost of the above steps, updates the minimum rate-distortion cost, and records the encoding mode corresponding to the minimum rate-distortion cost as the appropriate encoding mode.
[0245] Fig. 9 is a flowchart of a processing process in an affine advanced predictive coding mode according to an embodiment of the present application, such as Fig. 9 As shown, the process may include the following steps:
[0246] Step S901, determine whether there is an attempted coding unit at the same position.
[0247] In this embodiment, according to the block division type in the above-mentioned H.266 / VVC standard, there may be coding units with the same spatial position and the same block size under different nested block division paths, which are recorded as the same-position coding units. For the input unit to be encoded, it is determined whether there is a coding unit that has been tried at the same position, including determining whether there is a coding unit at the same position and whether the coding unit at the same position has been tried. If the coding unit that has been tried at the same position exists, enter S909, otherwise enter S902.
[0248] Step S902, determining whether the unit to be encoded is in affine merge, geo or mmvd mode.
[0249] In this embodiment, it is determined whether the appropriate coding mode for the unit to be coded is the affine merge mode, the geo mode or the mmvd mode. If yes, the process proceeds to S907; otherwise, the process proceeds to S903.
[0250] Step S903, determine whether an adjacent coding unit exists.
[0251] In this embodiment, it is determined whether there are adjacent coding units around the unit to be encoded. Here, the adjacent coding unit is on the upper side of the upper boundary of the unit to be encoded and shares the upper boundary of the unit to be encoded with the unit to be encoded, or is on the left side of the left boundary of the unit to be encoded and shares the left boundary of the unit to be encoded with the unit to be encoded. If the adjacent coding unit exists, proceed to S904; otherwise, proceed to S907.
[0252] Step S904, determining whether the adjacent coding unit is in affine merge or affine amvp mode.
[0253] In this embodiment, it is determined whether there is at least one adjacent coding unit, and its suitable coding mode is the affine merge mode or the affine amvp mode. If yes, proceed to S907; otherwise, proceed to S905.
[0254] Step S905, determine whether there is a parent coding unit.
[0255] In this embodiment, it is determined whether the parent coding unit exists. If so, the process proceeds to S906; otherwise, the process ends.
[0256] Step S906, determining whether the parent coding unit is in affine merge or affine amvp mode.
[0257] In this embodiment, it is determined whether the appropriate coding mode of the parent coding unit is the affine merge or affineamvp mode. If yes, proceed to S907; otherwise, end.
[0258] Step S907, try the 4-parameter affine amvp model.
[0259] In this embodiment, the 4param affine AVP mode is tried, the rate-distortion cost of the 4param affine AVP mode is compared with the minimum rate-distortion cost of the regular AVP mode, the minimum rate-distortion cost is updated, and the mode corresponding to the minimum rate-distortion cost is recorded as the appropriate coding mode.
[0260] Step S908, try the 6-parameter affine amvp model.
[0261] In this embodiment, the 6param affine amp mode is tried, the rate distortion cost of the 6param affine amp mode is compared with the minimum rate distortion cost of the above 4-parameter affine amp, the minimum rate distortion cost is updated, and the mode corresponding to the minimum rate distortion cost is recorded as the appropriate coding mode.
[0262] Step S909, determining whether the attempted coding unit at the same position is in affine merge or affine amvp mode.
[0263] In this embodiment, it is determined whether the appropriate coding mode of the attempted coding unit at the same position is the affinemerge mode or the affine amp mode. If yes, proceed to S902; otherwise, skip the 4paramaffine amp mode and the 6param affine amp mode attempts of the unit to be coded and directly end.
[0264] In an embodiment of the present application, the correlation between multiple encoding modes is utilized. By using suitable encoding modes of attempted encoding units at the same position, suitable encoding modes of attempted units to be encoded, suitable encoding modes of adjacent encoded units, and suitable encoding modes of parent encoding units, invalid attempts at affine prediction in units to be encoded are further reduced, thereby improving the encoder speed.
[0265] Optionally, if a coding unit that has been tried at the same position exists, and the appropriate coding mode of the coding unit that has been tried at the same position is not the affine merge mode or the affine amvp mode, the unit to be encoded skips trying the 4param affineamvp mode and the 6param affine amvp mode.
[0266] Optionally, if the attempted coding unit at the same position does not exist, the suitable coding mode for the unit to be coded is not affinemerge mode, geo mode or mmvd mode, the adjacent coding unit does not exist, and the parent coding unit does not exist, the unit to be coded skips trying 4param affine amvp mode and 6param affine amvp mode.
[0267] Optionally, if the attempted coding unit at the same position does not exist, the suitable coding mode for the unit to be coded is not affinemerge mode, geo mode or mmvd mode, the adjacent coding unit exists, and the suitable coding mode for the non-existent adjacent coding unit is affine merge mode or affine amvp mode, and the parent coding unit does not exist, the unit to be coded skips trying 4param affine amvp mode and 6param affine amvp mode.
[0268] Optionally, if the attempted coding unit at the same position does not exist, the appropriate coding mode for the unit to be coded is not affinemerge mode, geo mode or mmvd mode, the adjacent coding unit does not exist, the parent coding unit exists, and the appropriate coding mode for the parent block coding unit is not affine merge mode or affine amp mode, then the unit to be coded skips trying 4param affine amvp mode and 6param affine amvp mode.
[0269] Optionally, if the attempted coding unit at the same position does not exist, the suitable coding mode for the unit to be coded is not affinemerge mode, geo mode or mmvd mode, the adjacent coding unit exists, and the suitable coding mode for the non-existent adjacent coding unit is affine merge mode or affine amp mode, the parent coding unit exists, and the suitable coding mode for the parent block coding unit is not affine merge mode or affine amp mode, then the unit to be coded skips trying 4paramaffine amp mode and 6param affine amp mode.
[0270] Optionally, if an attempted coding unit at the same position exists, the appropriate coding mode of the attempted coding unit at the same position is affine merge mode or affine amvp mode, the appropriate coding mode of the unit to be coded is not affinemerge mode or geo mode or mmvd mode, the adjacent coding unit does not exist, and the parent coding unit does not exist, then the unit to be coded skips trying 4param affine amvp mode and 6param affine amvp mode.
[0271] In an embodiment of the present application, the order of using information such as the suitable coding mode that has been tried for the coding unit to be coded, the suitable coding mode of the adjacent coding unit, and the suitable coding mode of the parent coding unit can be adjusted. For example, first determine whether the parent coding unit exists and whether the suitable coding mode of the parent coding unit is the affine merge mode or the affineamvp mode, then determine whether the adjacent coding unit exists and whether the suitable coding mode of at least one adjacent coding unit is the affine merge or the affineamvp mode, and then determine whether the suitable coding mode that has been tried for the coding unit to be coded is the affine merge or the geo mode or the mmvd mode. Similarly, another feasible solution is to first determine whether the adjacent coding unit exists and whether the suitable coding mode of at least one adjacent coding unit is the affinemerge or the affineamvp mode, then determine whether the suitable coding mode that has been tried for the coding unit to be coded is the affinemerge or the geo mode or the mmvd mode, and then determine whether the parent coding unit exists and whether the suitable mode of the parent coding unit is the affine merge mode or the affineamvp mode.
[0272] Optionally, the correlations among multiple coding modes are utilized, including the correlation of appropriate coding modes between spatially co-located coding units and units to be encoded, the correlation between different coding modes in units to be encoded, the correlation of appropriate coding modes between spatially adjacent coding units and units to be encoded, and the correlation of appropriate coding modes between child coding units and parent coding units, thereby further effectively reducing invalid attempts at affine prediction and improving the encoding speed of the encoder.
[0273] In an embodiment of the present application, a fast affine prediction coding algorithm based on the correlation of multiple coding modes is proposed to improve the efficiency of affine prediction attempts, thereby further improving the encoding speed of the encoder. According to the correlation between the spatially co-located coding unit and the unit to be encoded, the appropriate coding mode of the spatially co-located attempted coding unit is used to determine the possibility that the appropriate coding mode of the unit to be encoded is an affine prediction mode, thereby skipping unnecessary affine prediction algorithm attempts. According to the correlation between different coding modes in the unit to be encoded, the rate-distortion cost of the affine merge mode, the geo mode, and the mmvd mode is used to determine the possibility that the appropriate coding mode of the unit to be encoded is an affine prediction mode, and skip unnecessary affine prediction algorithm attempts. At the same time, according to the correlation of the appropriate coding mode between the spatially adjacent coding unit and the unit to be encoded, the possibility that the appropriate coding mode of the unit to be encoded is an affine prediction mode is determined, and unnecessary affine prediction algorithm attempts are skipped. According to the correlation of the appropriate coding mode between the child coding unit and the parent coding unit, the possibility that the appropriate coding mode of the unit to be encoded is an affine prediction mode is determined, and unnecessary affine prediction algorithm attempts are skipped. The technical effect of improving the efficiency of video encoding is achieved, and the technical problem of low video encoding efficiency is solved.
[0274] According to an embodiment of the present application, there is also provided a method for implementing the above Figure 3 The video information processing method shown is a video information processing device.
[0275] Fig.10 is a schematic diagram of a video information processing device according to an embodiment of the present application, such as Fig.10 As shown, the video information processing device 800 may include: a first acquisition unit 1002 , a first determination unit 1004 , a first testing unit 1006 and a first determination unit 1008 .
[0276] The first acquisition unit 1002 is configured to acquire at least one to-be-encoded video block from a video file.
[0277] The first determining unit 1004 is configured to determine at least one coding mode corresponding to the to-be-coded video block.
[0278] The first testing unit 1006 is configured to test the affine prediction coding mode based on the correlation between the coding mode and the target coding mode corresponding to the video block to be coded, and obtain a test result.
[0279] The first determining unit 1008 is configured to determine a coding strategy for the to-be-coded video block according to the test result.
[0280] Here, the first acquisition unit 1002, the first determination unit 1004, the first test unit 1006 and the first determination unit 1008 correspond to steps S302 to S308 in the above embodiment, and the four units are the same as the examples and application scenarios implemented by the corresponding steps, but are not limited to the contents disclosed in the above embodiment. It should be noted that the above units can be hardware components or software components stored in a memory (e.g., memory 1304) and processed by one or more processors (e.g., processors 1302a, 1302b..., 1302n), and the above units can also be run as part of the device in the computer terminal a provided in the following embodiment.
[0281] According to an embodiment of the present application, there is also provided a method for implementing the above Figure 4 A video encoding method and a video encoding device are shown.
[0282] Fig.11 is a schematic diagram of a video encoding device according to an embodiment of the present application, such as Fig.11 As shown, the video encoding device 1100 may include: a second acquisition unit 1102 , a third determination unit 1104 and an encoding unit 1106 .
[0283] The second acquisition unit 1102 is configured to acquire at least one to-be-encoded video block from a video file.
[0284] The third determining unit 1104 is configured to determine a coding strategy for the video block to be coded.
[0285] The encoding unit 1106 is used to encode the video block to be encoded according to the encoding strategy.
[0286] It should be noted that the second acquisition unit 1102, the third determination unit 1104 and the encoding unit 1106 correspond to steps S402 to S406 in the above embodiment, and the three units and the corresponding steps implement the same examples and application scenarios, but are not limited to the contents disclosed in the above embodiment. It should be noted that the above units can be hardware components or software components stored in a memory (e.g., memory 1304) and processed by one or more processors (e.g., processors 1302a, 1302b..., 1302n), and the above units can also be part of the device and can be run in the computer terminal a provided in the following embodiment.
[0287] According to an embodiment of the present application, there is also provided a method for implementing the above Figure 5 The video information processing method shown is a video information processing device.
[0288] Fig.12is a schematic diagram of another video information processing device according to an embodiment of the present application, such as Fig.12 As shown, the video information processing device 1200 may include: a third acquisition unit 1202 , a fourth determination unit 1204 , a second test unit 1206 , a fifth determination unit 1208 and an output unit 1210 .
[0289] The third acquisition unit 1202 is configured to acquire at least one to-be-encoded video block from the video file by calling the first interface.
[0290] The fourth determining unit 1204 is configured to determine at least one coding mode corresponding to the to-be-coded video block.
[0291] The second testing unit 1206 is configured to test the affine prediction coding mode based on the correlation between the coding mode and the target coding mode corresponding to the video block to be coded, and obtain a test result.
[0292] The fifth determining unit 1208 is configured to determine the encoding strategy of the video block to be encoded according to the test result.
[0293] The output unit 1210 is configured to output the encoding strategy by calling the second interface.
[0294] It should be noted that the third acquisition unit 1202, the fourth determination unit 1204, the second test unit 1206, the fifth determination unit 1208 and the output unit 1210 correspond to steps S502 to S510 in the above embodiment, and the five units are the same as the examples and application scenarios implemented by the corresponding steps, but are not limited to the contents disclosed in the above embodiment. It should be noted that the above units can be hardware components or software components stored in a memory (e.g., memory 1304) and processed by one or more processors (e.g., processors 1302a, 1302b..., 1302n), and the above units can also be part of the device and can be run in the computer terminal a provided in the following embodiment.
[0295] In the video information processing device, if it is necessary to process the video information, the video file to be processed can be obtained, and at least one video block to be encoded can be obtained from the video file. The coding mode allowed to be adopted by the associated coding video block associated with the video block to be encoded during the coding process can be determined, or the tested coding mode of the video block to be encoded can be determined as the coding mode corresponding to the video block to be encoded. The probability of whether the affine prediction coding mode can be used as the target coding mode can be tested based on the correlation (correlation degree) between the coding mode and the target coding mode corresponding to the video block to be encoded to obtain the test result. The coding strategy of the video block to be encoded can be determined according to the test result, and the corresponding video block to be encoded in the video file can be encoded by using the coding strategy. In the embodiment of the present application, the affine coding mode can be quickly predicted by the correlation between multiple coding modes, that is, by using the correlation between multiple coding modes between the video block to be encoded and the associated coding video block, the effectiveness of the affine prediction coding mode adopted by the video block to be encoded is predicted by judging the target coding mode of the above-mentioned associated coding video block. If the target coding mode indicates that the affine prediction coding mode may not be a suitable choice, the attempt to use the affine prediction coding mode can be skipped, thereby avoiding unnecessary computational complexity and significantly speeding up the coding process. The above method can effectively reduce invalid attempts of affine prediction, and also achieve the purpose of optimizing the coding strategy to increase the video coding speed, thereby achieving the technical effect of improving the efficiency of video coding and solving the technical problem of low video coding efficiency.
[0296] The embodiment of the present application may provide a computer terminal, which may be any computer terminal device in a computer terminal group. Optionally, in this embodiment, the computer terminal may also be replaced by a terminal device such as a mobile terminal.
[0297] Optionally, in this embodiment, the computer terminal may be located in at least one network device among a plurality of network devices of the computer network.
[0298] In this embodiment, the computer terminal can execute the program code of the steps in the above embodiment in the video information processing method.
[0299] Optionally, Fig.13 is a structural block diagram of a computer terminal according to an embodiment of the present application. Fig.13 As shown, the computer terminal a may include: one or more (only one is shown in the figure) processors 1302 , a memory 1304 and a transmission device 1306 .
[0300] Among them, the memory can be used to store software programs and modules, such as the program instructions / modules corresponding to the video information processing method and device in the embodiment of the present application. The processor executes various functional applications and data processing by running the software programs and modules stored in the memory, that is, realizing the above-mentioned video information processing method. The memory may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory may further include a memory remotely arranged relative to the processor, and these remote memories may be connected to the computer terminal a via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0301] According to an embodiment of the present application, a method for processing video information is provided. If it is necessary to process the information of the video, a video file to be processed can be obtained, and at least one video block to be encoded can be obtained from the video file. The coding mode allowed to be adopted by the associated coding video block associated with the video block to be encoded during the coding process can be determined, or the tested coding mode of the video block to be encoded can be determined as the coding mode corresponding to the video block to be encoded. The probability of whether the affine prediction coding mode can be used as the target coding mode can be tested based on the correlation (correlation degree) between the coding mode and the target coding mode corresponding to the video block to be encoded to obtain the test result. The coding strategy of the video block to be encoded can be determined according to the test result, and the corresponding video block to be encoded in the video file can be encoded by using the coding strategy. In the embodiment of the present application, the affine coding mode can be quickly predicted by the correlation between multiple coding modes, that is, by using the correlation between multiple coding modes between the video block to be encoded and the associated coding video block, the effectiveness of the affine prediction coding mode adopted by the video block to be encoded is predicted by judging the target coding mode of the associated coding video block. If the target coding mode indicates that the affine prediction coding mode may not be a suitable choice, the attempt to use the affine prediction coding mode can be skipped, thereby avoiding unnecessary computational complexity and significantly speeding up the coding process. The above method can effectively reduce invalid attempts of affine prediction, and also achieve the purpose of optimizing the coding strategy to increase the video coding speed, thereby achieving the technical effect of improving the efficiency of video coding and solving the technical problem of low video coding efficiency.
[0302] It can be understood by those skilled in the art that Fig.13 The structure shown is for illustration only, and the computer terminal a may also be a smart phone (such as an Android phone, an iOS phone, etc.), a tablet computer, a PDA, a mobile Internet device (Mobile Internet Devices, referred to as MID), a PAD, or other terminal devices. Fig.13It does not limit the structure of the computer terminal a. For example, the computer terminal a may also include Fig.13 More or fewer components (such as network interfaces, display devices, etc.) shown in, or having Fig.13 Different configurations are shown.
[0303] A person of ordinary skill in the art can understand that all or part of the steps in the various methods of the above embodiments can be completed by instructing the hardware related to the terminal device through a program, and the program can be stored in a computer-readable storage medium, and the storage medium may include: a flash drive, a read-only memory (ROM), a random access memory (RAM), a disk or an optical disk, etc.
[0304] An embodiment of the present application may provide a computing device. Fig.14 is a structural block diagram of a computing device according to an embodiment of the present application. Fig.14 As shown, the computing device 140 may include: one or more (only one is shown in the figure) processors 1402, a memory 1404, a storage controller, and a peripheral interface.
[0305] The above-mentioned computing device can be understood as an integrated intelligent terminal, including but not limited to a server, a desktop computer, a personal computer (PC for short), a model all-in-one machine, etc., and the computing device can be pre-installed with the model described in the above-mentioned embodiments of the present application.
[0306] Among them, the memory can be used to store software programs and modules, such as program instructions / modules corresponding to the methods and devices in the embodiments of the present application, and the processor executes various functional applications and data processing by running the software programs and modules stored in the memory, that is, implementing the methods in the above embodiments. The memory may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory may further include a memory remotely arranged relative to the processor, and these remote memories may be connected to the terminal via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0307] The processor may call the executable program stored in the memory through the transmission device to execute the method described in any one of the above embodiments.
[0308] An embodiment of the present application may provide an electronic device. Fig.151 is a block diagram of an electronic device according to an embodiment of the present application. As shown in the figure, the electronic device may include: an input / output device 1502; a memory 1504 and a processor 1506, wherein the processor 1506 is connected to the input / output device 1502 and the memory 1504 via a bus 1508.
[0309] Among them, the memory can be used to store software programs and modules, such as program instructions / modules corresponding to the methods and devices in the embodiments of the present application, and the processor executes various functional applications and data processing by running the software programs and modules stored in the memory, that is, implementing the methods in the above embodiments. The memory may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory may further include a memory remotely arranged relative to the processor, and these remote memories may be connected to the terminal via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0310] The processor may call the executable program stored in the memory through the transmission device to execute the method described in any one of the above embodiments.
[0311] The embodiment of the present application further provides a computer-readable storage medium. Optionally, in this embodiment, the computer-readable storage medium can be used to store the program code executed by the video information processing method provided in the first embodiment.
[0312] Optionally, in this embodiment, the computer-readable storage medium may be located in any one of the computer terminals in a computer terminal group in a computer network, or in any one of the mobile terminals in a mobile terminal group.
[0313] Optionally, in this embodiment, the computer-readable storage medium is configured to store program codes for executing the steps in the above embodiments.
[0314] An embodiment of the present application may provide an electronic device, which may include a memory and a processor.
[0315] Fig.16It is a block diagram of an electronic device according to a method for processing information of a video according to an embodiment of the present application. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workbenches, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital processing, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present application described and / or required herein.
[0316] like Fig.16 As shown, the device 1600 includes a computing unit 1601, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 1602 or a computer program loaded from a storage unit 1608 into a random access memory (RAM) 1603. In the RAM 1603, various programs and data required for the operation of the device 1600 can also be stored. The computing unit 1601, the ROM 1602, and the RAM 1603 are connected to each other via a bus 1604. An input / output (I / O) interface 1605 is also connected to the bus 1604.
[0317] A number of components in the device 1600 are connected to the I / O interface 1605, including: an input unit 1606, such as a keyboard, a mouse, etc.; an output unit 1604, such as various types of displays, speakers, etc.; a storage unit 1608, such as a disk, an optical disk, etc.; and a communication unit 1609, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 1609 allows the device 1600 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.
[0318] The computing unit 1601 may be a variety of general and / or special processing components with processing and computing capabilities. Some examples of the computing unit 1601 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, digital signal processors (DSPs), and any appropriate processors, controllers, microcontrollers, etc. The computing unit 1601 performs the various methods and processes described above, such as the information processing method of the video. For example, in some embodiments, the information processing method of the video may be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as a storage unit 1608. In some embodiments, part or all of the computer program may be loaded and / or installed on the device 1600 via the ROM 1602 and / or the communication unit 1609. When the computer program is loaded into the RAM 1603 and executed by the computing unit 1601, one or more steps of the information processing method of the video described above may be performed. Alternatively, in other embodiments, the computing unit 1601 may be configured to execute the video information processing method in any other appropriate manner (eg, by means of firmware).
[0319] According to another aspect of the embodiment of the present application, a computer program product is also provided. The computer program product includes a computer program, and when the computer program is executed by a processor, the video information processing method of the embodiment of the present application is implemented.
[0320] Various implementations of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chips (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various implementations can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.
[0321] The program code for implementing the method of the present application can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer or other programmable data processing device, so that the program code, when executed by the processor or controller, enables the functions / operations specified in the flow chart and / or block diagram to be implemented. The program code can be executed entirely on the machine, partially on the machine, partially on the machine and partially on a remote machine as a stand-alone software package, or entirely on a remote machine or server.
[0322] Various embodiments of the systems and techniques described above herein may be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard parts (ASSPs), system-on-a-chip systems (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include: being implemented in one or more computer programs that may be executed and / or interpreted on a programmable system including at least one programmable processor that may be a dedicated or general purpose programmable processor that may receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.
[0323] The program code for implementing the method of the present application can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer or other programmable data processing device, so that the program code, when executed by the processor or controller, enables the functions / operations specified in the flow chart and / or block diagram to be implemented. The program code can be executed entirely on the machine, partially on the machine, partially on the machine and partially on a remote machine as a stand-alone software package, or entirely on a remote machine or server.
[0324] In the context of the present application, a machine-readable medium may be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, device, or equipment. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium may include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory for short), an optical fiber, a portable compact disk read-only memory (CD-ROM for short), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0325] To provide interaction with a user, the systems and techniques described herein may be implemented on a computer having: a display device (e.g., a cathode ray tube (CRT) or a liquid crystal display (LCD), a monitor for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a path ball), through which the user may provide input to the computer. Other types of devices may also be used to provide interaction with the user; for example, the feedback provided to the user may be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user may be received in any form (including acoustic input, voice input, or tactile input).
[0326] The systems and techniques described herein may be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer with a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system may be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a Local Area Network (LAN), a Wide Area Network (WAN), and the Internet.
[0327] A computer system may include a client and a server. The client and the server are generally remote from each other and usually interact through a communication network. The relationship of client and server is generated by computer programs running on respective computers and having a client-server relationship with each other. The server may be a cloud server, a server of a distributed system, or a server combined with a blockchain.
[0328] It should be noted that the serial numbers of the above-mentioned embodiments of the present application are only for description and do not represent the advantages or disadvantages of the embodiments.
[0329] In the above embodiments of the present application, the description of each embodiment has its own emphasis. For parts that are not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0330] In the several embodiments provided in this application, it should be understood that the disclosed technical content can be implemented in other ways. Among them, the device embodiments described above are only schematic, for example, the division of units is only a logical function division, and there may be other division methods in actual implementation, for example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of units or modules, which can be electrical or other forms.
[0331] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0332] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of software functional units.
[0333] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product, which is stored in a storage medium and includes several instructions for a computer device (which can be a personal computer, server or network device, etc.) to perform all or part of the steps of the various embodiments of the present application. The aforementioned storage medium includes: various media that can store program codes, such as USB flash drives, read-only memories, random access memories, mobile hard disks, magnetic disks or optical disks.
[0334] The above are only preferred implementations of the present application. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present application. These improvements and modifications should also be regarded as the scope of protection of the present application.
Claims
1. A video information processing method, characterized in that: include: Obtain at least one to-be-encoded video block from a video file; Determine at least one coding mode corresponding to the video block to be coded, wherein the coding mode is a mode allowed to be adopted by at least one associated coded video block associated with the video block to be coded during the coding process, or the coding mode is a mode that has been tested during the coding process of the video block to be coded; Based on the correlation between the coding mode and the target coding mode corresponding to the video block to be encoded, the affine prediction coding mode is tested to obtain a test result, wherein a performance index of the target coding mode for encoding the video block to be encoded is greater than a performance index threshold, and the test result is used to indicate the probability that the affine prediction coding mode becomes the target coding mode; According to the test result, a coding strategy for the video block to be coded is determined, wherein the coding strategy is used to represent a rule for coding the video block to be coded.
2. The method according to claim 1, characterized in that A performance index of a mode allowed to be used in the encoding process of the associated coded video block for encoding the associated coded video block is greater than the performance index threshold, and a performance index of a mode that has been tested in the encoding process of the video block to be encoded for encoding the video block to be encoded is greater than the performance index threshold.
3. The method according to claim 1, characterized in that Determining the encoding strategy of the to-be-encoded video block according to the test result includes: In response to the test result being that the probability is less than or equal to a probability threshold, determining that the encoding strategy is a rule of skipping encoding the to-be-encoded video block according to the affine prediction encoding mode; In response to the test result being that the probability is greater than the probability threshold, the encoding strategy is determined to be a rule that allows the to-be-encoded video block to be encoded according to the affine prediction encoding mode.
4. The method according to claim 1, characterized in that: Determining at least one encoding mode corresponding to the to-be-encoded video block includes: Determining a target position of the video block to be encoded in the video file and a target block size of the video block to be encoded; Acquire, from the video file, a video block having the same position as the target position, the same block size as the target block size, and the same position as the encoded video block; A coding mode for the co-located coded video block is determined.
5. The method according to claim 4, characterized in that Based on the correlation between the coding mode and the target coding mode corresponding to the to-be-coded video block, the affine prediction coding mode is tested to obtain a test result, including: In response to the correlation indicating that the coding mode of the co-located encoded video block is different from the affine prediction coding mode corresponding to the target coding mode, the test result is determined to be that the probability is less than or equal to a probability threshold.
6. The method according to claim 4, characterized in that At least one of the associated coded video blocks includes: at least one adjacent coded video block adjacent to the video block to be coded in the video file, and a parent coded video block of the video block to be coded in the video file; at least one coding mode includes: a tested coding mode of the video block to be coded, a coding mode of the adjacent coded video block, and a coding mode of the parent coded video block; and determining at least one coding mode corresponding to the video block to be coded from the video file, further comprising: In response to the correlation indicating that the coding mode of the co-located coded video block is the same as the affine prediction coding mode, or in response to failure to obtain the co-located coded video block, performing the following steps in target order: A determination step, determining the tested coding mode; A first acquisition step is to acquire the coding mode of the adjacent coded video blocks; The second acquisition step is to acquire the coding mode of the parent coded video block.
7. The method according to claim 6, characterized in that Perform the following steps in order of purpose, including: Execute a current step in the target sequence, wherein the current step is any one of the determining step, the first obtaining step, and the second obtaining step; In response to failure of execution of the current step, or in response to the correlation indicating that the encoding mode in the current step is different from the affine prediction encoding mode, executing a next step of the current step in the target sequence; Based on the correlation between the coding mode and the target coding mode corresponding to the to-be-coded video block, testing the affine prediction coding mode to obtain a test result, including: in response to the correlation indicating that the coding mode in the last step in the target sequence is different from the affine prediction coding mode corresponding to the target coding mode, determining that the test result is that the probability is less than or equal to a probability threshold; The method further includes: in response to failure to obtain the associated coded video block in the last step in the target sequence, determining that the test result is that the probability is less than or equal to the probability threshold.
8. The method according to claim 7, characterized in that The target sequence is: executing the determining step, the first obtaining step, and the second obtaining step in sequence, and executing the current step in the target sequence includes: Taking the determining step as the current step to determine the tested coding mode; In response to the correlation indicating that the coding mode in the current step is different from the affine prediction coding mode, performing the next step of the current step in the target order, including: in response to the correlation indicating that the tested coding mode is different from at least one of the following coding modes, performing the next first acquisition step of the determining step in the target order to acquire the coding mode of the adjacent coded video block: the affine prediction coding mode, the geometric prediction coding mode and the merge prediction mode based on motion vector difference; In response to failure to execute the current step, or in response to the correlation indicating that the coding mode in the current step is different from the affine prediction coding mode, executing the next step of the current step in the target order, including: taking the first acquisition step as the current step, in response to failure to acquire the adjacent coded video block, or in response to the correlation indicating that the coding mode of the adjacent coded video block is different from the affine prediction coding mode, executing the second acquisition step next to the first acquisition step in the target order to acquire the coding mode of the parent coded video block.
9. The method according to claim 7, characterized in that: In response to the correlation indicating that the coding mode in the last step in the target sequence is different from the affine prediction coding mode corresponding to the target coding mode, determining that the test result is that the probability is less than or equal to the probability threshold comprises: In response to the correlation indicating that the encoding mode of the adjacent encoded video block obtained is different from the affine prediction encoding mode, and the correlation indicating that the encoding mode of the parent encoded video block obtained is different from the affine prediction encoding mode, it is determined that the test result is that the probability is less than or equal to the probability threshold.
10. The method according to claim 7, characterized in that In response to the correlation indicating that the coding mode in the last step in the target sequence is different from the affine prediction coding mode, determining that the test result is that the probability is less than or equal to the probability threshold includes: In response to failure to obtain the adjacent coded video block, and the correlation indicates that the obtained coding mode of the parent coded video block is different from the affine prediction coding mode, determining the test result as the probability is less than or equal to the probability threshold.
11. The method according to claim 7, characterized in that In response to a failure in acquiring the associated coded video block in the last step in the target sequence, determining that the test result is that the probability is less than or equal to the probability threshold comprises: In response to the correlation indicating that the encoding mode of the acquired adjacent encoded video block is different from the affine prediction encoding mode and acquisition of the parent encoded video block fails, determining that the test result is that the probability is less than or equal to the probability threshold.
12. The method according to claim 7, characterized in that In response to a failure in acquiring the associated coded video block in the last step in the target sequence, determining that the test result is that the probability is less than or equal to the probability threshold comprises: In response to a failure in acquiring the adjacent coded video block and a failure in acquiring the parent coded video block, determining that the test result is that the probability is less than or equal to the probability threshold.
13. The method according to claim 6, characterized in that The target sequence is: executing the second acquisition step, the first acquisition step, and the determination step in sequence; or, executing the first acquisition step, the determination step, and the second acquisition step in sequence.
14. A video encoding method, characterized in that: include: Obtain at least one to-be-encoded video block from a video file; Determining a coding strategy for the video block to be encoded, wherein the coding strategy is used to represent a rule for encoding the video block to be encoded, and the coding strategy is determined according to a test result, the test result is used to represent a probability that an affine prediction coding mode becomes a target coding mode corresponding to the video block to be encoded, the test result is obtained by testing the affine prediction coding mode based on a correlation between at least one coding mode corresponding to the video block to be encoded and the target coding mode, a performance index of the target coding mode for encoding the video block to be encoded is greater than a performance index threshold, the coding mode is a mode allowed to be adopted by at least one associated coding video block associated with the video block to be encoded during the encoding process, or the coding mode is a mode that has been tested during the encoding process of the video block to be encoded; The to-be-encoded video block is encoded according to the encoding strategy.
15. A video information processing method, characterized in that: include: Obtain at least one to-be-encoded video block from a video file by calling a first interface, wherein the first interface includes a first parameter, and a parameter value of the first parameter includes the to-be-encoded video block; Determine at least one coding mode corresponding to the video block to be coded, wherein the coding mode is a mode allowed to be adopted by at least one associated coded video block associated with the video block to be coded during the coding process, or the coding mode is a mode that has been tested during the coding process of the video block to be coded; Based on the correlation between the coding mode and the target coding mode corresponding to the video block to be encoded, the affine prediction coding mode is tested to obtain a test result, wherein a performance index of the target coding mode for encoding the video block to be encoded is greater than a performance index threshold, and the test result is used to indicate the probability that the affine prediction coding mode becomes the target coding mode; Determining, according to the test result, a coding strategy for the video block to be encoded, wherein the coding strategy is used to represent a rule for encoding the video block to be encoded; The encoding strategy is output by calling a second interface, wherein the second interface includes a second parameter, and a parameter value of the second parameter includes the encoding strategy.
16. A video information processing system, characterized in that: include: A video encoder, used for obtaining at least one video block to be encoded from a video file; Determining a coding strategy for the video block to be encoded, wherein the coding strategy is used to represent a rule for encoding the video block to be encoded, and the coding strategy is determined according to a test result, the test result is used to represent a probability that an affine prediction coding mode becomes a target coding mode corresponding to the video block to be encoded, the test result is obtained by testing the affine prediction coding mode based on a correlation between at least one coding mode corresponding to the video block to be encoded and the target coding mode, a performance index of the target coding mode for encoding the video block to be encoded is greater than a performance index threshold, the coding mode is a mode allowed to be adopted by at least one associated coding video block associated with the video block to be encoded during the encoding process, or the coding mode is a mode that has been tested during the encoding process of the video block to be encoded; encoding the video block to be encoded according to the coding strategy; The server is used to obtain the encoded video block to be encoded.
17. A computing device, characterized in that include: A memory storing an executable program; A processor, configured to run the program, wherein the program executes the method according to any one of claims 1 to 15 when running.
18. An electronic device, characterized in that: include: A memory storing an executable program; A processor, connected to the memory via a bus, and configured to run the program, wherein the program executes the method described in any one of claims 1 to 15 when running.
19. A computer-readable storage medium, characterized in that: The computer-readable storage medium includes a stored executable program, wherein when the executable program is executed, the device where the computer-readable storage medium is located is controlled to execute the method according to any one of claims 1 to 15.
20. A computer program product, characterized in that The invention comprises a computer program which, when executed by a processor, implements the method according to any one of claims 1 to 15.