Video frame coding and decoding method and device, equipment and storage medium
By identifying the original prediction unit that meets the intra-frame copy condition in video coding and combining it with artificial intelligence technology, the prediction accuracy problem of the IBC method under complex coding unit and noisy conditions is solved, thus improving the video coding and decoding performance.
Patent Information
- Application Number
- CN202410550756.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-05-06
- Publication Date
- 2025-11-11
AI Technical Summary
Existing video coding technologies struggle to accurately determine prediction units when dealing with complex coding units or in noisy conditions, resulting in low coding performance.
By identifying the original prediction units that meet the intra-frame copy condition in the encoded image units, and further selecting the target prediction units from the matching unit set, and combining artificial intelligence technology to identify image patch similarity, the accuracy of the prediction units is improved.
It improves the accuracy of intra-frame prediction, reduces the amount of encoded data, lowers the probability of encoding reconstruction distortion, and enhances video encoding and decoding performance.
Smart Images

Figure CN120935368A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, and particularly to the field of video encoding and decoding technology. It provides a video frame encoding, decoding method, device, equipment, and storage medium. Background Art
[0002] Video is a sequence of continuous images, consisting of continuous frames, and one frame is an image. Video encoding refers to the method of converting the original video format file into another video format file through compression technology. There is a large amount of redundant information in the original video format file, and the purpose of compression is to remove the redundant information in the data. Prediction coding is one of the core technologies of video encoding, which means using the encoded sample values, according to a certain model or method, to predict the current sample value, and encoding the difference between the true value and the predicted value of the sample. The video encoder performs transformation, quantization, and entropy coding on the predicted residual instead of the original pixel values, thereby greatly improving the encoding efficiency.
[0003] For video, similar image textures in one frame may appear multiple times. For example, Figure 1 in a frame containing a presentation document as shown, the same text will appear multiple times. For example, Figure 1 the letter "m" and the character "的" appear multiple times as shown, resulting in spatial redundancy within the same frame image. When compressing the intra-frame image data, the already encoded similar image textures can be used to perform efficient prediction coding on the current image block to be encoded. The intra-frame prediction method based on Intra Block Copy (IBC) is used to predict the repeated textures. The IBC method searches for the already encoded area given in the current frame for the current Coding Unit (CU), and takes the block closest to it as the Prediction Unit (PU) of the current CU. Based on this PU, the prediction residual of the current CU is obtained, and further transformation, quantization, and entropy coding processes are carried out to complete the entire encoding process of the current CU.
[0004] Therefore, accurate intra-frame prediction has a great impact on the entire encoding process of the current CU. However, for a more complex CU or due to the presence of noise, the IBC method may not be able to accurately determine the PU, which will lead to low encoding performance of the current CU. Summary of the Invention
[0005] Embodiments of this application provide a video frame encoding, decoding method, device, equipment, and storage medium for improving intra-frame prediction performance, thereby improving video encoding and decoding performance.
[0006] On the one hand, a video frame encoding method is provided, the method comprising:
[0007] From at least one image unit already encoded within the target video frame, determine the original prediction unit whose similarity to the current coding unit meets the intra-frame copy condition;
[0008] From the at least one image unit, determine a set of matching units whose similarity to the original prediction unit meets the preset matching conditions;
[0009] Based on the original prediction unit and the matching unit set, the target prediction unit of the current coding unit is determined;
[0010] Based on the difference information between the target prediction unit and the current coding unit, the current coding unit is encoded to obtain the target coding result of the current coding unit.
[0011] On the one hand, a video frame decoding method is provided, the method comprising:
[0012] Obtain the target encoding result of the current decoding unit within the target video frame. The target encoding result includes data encoding result, prediction unit indication information, and position indication information. The prediction unit indication information is used to indicate the target prediction unit that the current decoding unit will ultimately use during encoding. The position indication information is used to indicate the position of the original prediction unit of the current decoding unit.
[0013] Based on the location indication information, the original prediction unit is determined from at least one decoded image unit within the target video frame;
[0014] The target prediction unit is determined based on the prediction unit indication information and the original prediction unit;
[0015] Based on the decoding result of the target prediction unit, the data encoding result is decoded to obtain the decoding result of the current decoding unit.
[0016] On one hand, a video frame encoding apparatus is provided, the apparatus comprising:
[0017] An intra-frame copy prediction unit is used to determine, from at least one encoded image unit within a target video frame, an original prediction unit whose similarity to the current coding unit meets the intra-frame copy condition.
[0018] A matching unit is used to determine, from the at least one image unit, a set of matching units whose similarity to the original prediction unit meets a preset matching condition;
[0019] A prediction unit selection unit is used to determine the target prediction unit of the current coding unit based on the original prediction unit and the matching unit set.
[0020] The encoding execution unit is used to encode the current encoding unit based on the difference information between the target prediction unit and the current encoding unit, so as to obtain the target encoding result of the current encoding unit.
[0021] On one hand, a video frame decoding apparatus is provided, the apparatus comprising:
[0022] The encoding result acquisition unit is used to obtain the target encoding result of the current decoding unit within the target video frame. The target encoding result includes data encoding result, prediction unit indication information and position indication information. The prediction unit indication information is used to indicate the target prediction unit that the current decoding unit will ultimately use when encoding. The position indication information is used to indicate the position of the original prediction unit of the current decoding unit.
[0023] A prediction unit determining unit is configured to determine the original prediction unit from at least one decoded image unit within the target video frame based on the location indication information; and to determine the target prediction unit based on the prediction unit indication information and the original prediction unit.
[0024] The decoding execution unit is used to decode the data encoding result based on the decoding result of the target prediction unit to obtain the decoding result of the current decoding unit.
[0025] On one hand, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of any of the above methods.
[0026] On the one hand, a computer storage medium is provided, on which a computer program is stored, which, when executed by a processor, implements the steps of any of the above methods.
[0027] On one hand, a computer program product is provided, comprising a computer program stored in a computer-readable storage medium. A processor of a computer device reads the computer program from the computer-readable storage medium, and executes the computer program, causing the computer device to perform the steps of any of the methods described above.
[0028] In this embodiment of the application, when determining the original prediction unit that meets the intra-frame copy condition from the encoded image units, the applicant further determines the matching unit set that meets the preset matching condition with the original prediction unit from at least one image unit, and determines the target prediction unit of the current coding unit based on the original prediction unit and the matching unit set. Finally, the current coding unit is encoded based on the difference information between the target prediction unit and the current coding unit.
[0029] By further determining at least one matching unit based on the original prediction unit, and combining the original prediction unit and these matching units to determine the target prediction unit for the current coding unit, it is equivalent to further processing based on the currently determined original prediction units. This helps to improve the correlation between the finally determined target prediction unit and the current coding unit, thereby improving the accuracy of determining the final prediction unit for the current coding unit and thus improving coding performance. For example, the greater the correlation between the target prediction unit and the current coding unit, the smaller the difference between them, which reduces the amount of data to be encoded and helps to reduce the probability of coding reconstruction distortion. Therefore, even for more complex coding units or under the influence of noise, the prediction unit can be determined more accurately, improving the performance of intra-frame prediction and contributing to the improvement of video coding and decoding performance. Attached Figure Description
[0030] To more clearly illustrate the technical solutions in the embodiments of this application or related technologies, the drawings used in the description of the embodiments or related technologies will be briefly introduced below. Obviously, the drawings described below are only embodiments of this application. For those skilled in the art, other drawings can be obtained based on the provided drawings without creative effort.
[0031] Figure 1 Example diagram of similar image textures in a frame provided for embodiments of this application;
[0032] Figure 2 This is a schematic diagram illustrating the principle of IBC-based intra-frame prediction methods in related technologies.
[0033] Figure 3 This is a schematic diagram illustrating an application scenario provided in the embodiments of this application;
[0034] Figure 4 A flowchart illustrating the video frame encoding method provided in this application embodiment;
[0035] Figure 5 Example diagrams for determining the original prediction unit provided in embodiments of this application;
[0036] Figure 6Aand Figure 6B A schematic diagram of the search and matching process provided in the embodiments of this application;
[0037] Figure 7 A schematic diagram illustrating the search and matching of the original prediction unit provided in an embodiment of this application;
[0038] Figure 8 A schematic flowchart illustrating the image fusion processing procedure provided in an embodiment of this application;
[0039] Figure 9A and Figure 9B A flowchart illustrating the target selection prediction unit provided in an embodiment of this application;
[0040] Figure 10 A flowchart illustrating the video frame decoding method provided in this application embodiment;
[0041] Figure 11 This is a schematic diagram of the decoding process provided in an embodiment of this application;
[0042] Figure 12 An example diagram illustrating the interaction between the encoding and decoding ends provided in an embodiment of this application;
[0043] Figure 13 A schematic diagram of a video frame encoding apparatus provided in an embodiment of this application;
[0044] Figure 14 This is a schematic diagram of a video frame decoding device provided in an embodiment of this application;
[0045] Figure 15 This is a schematic diagram of the composition structure of a computer device provided in an embodiment of this application;
[0046] Figure 16 This is a schematic diagram of the composition structure of another computer device using an embodiment of this application. Detailed Implementation
[0047] To make the objectives, technical solutions, and advantages of this application clearer, the technical solutions in the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of them. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application. Unless otherwise specified, the embodiments and features in the embodiments of this application can be arbitrarily combined with each other. Furthermore, although a logical order is shown in the flowchart, in some cases, the steps shown or described may be performed in a different order than that shown here.
[0048] To facilitate understanding of the technical solutions provided in the embodiments of this application, some key terms used in the embodiments of this application will be explained below:
[0049] Video encoding refers to the method of converting an original video format file into another video format file through compression technology. The original video format file contains a large amount of redundant information, and the purpose of compression is to remove this redundant information. Redundant information includes spatial redundancy and temporal redundancy. Spatial redundancy refers to strong correlation between image units within a frame, such as the repeated appearance of similar image textures within a frame. Temporal redundancy refers to strong correlation between adjacent frames. The techniques for removing spatial and temporal redundancy are intra-frame prediction and inter-frame prediction, respectively. The technical solution of this application mainly involves intra-frame prediction technology. Predictive coding is one of the core technologies of video encoding. It refers to using already encoded sample values, according to a certain model or method, to predict the current sample value, and encoding the difference between the actual sample value and the predicted value. The video encoder transforms, quantizes, and entropy-encodes the predicted residual rather than the original pixel value, thereby significantly improving coding efficiency.
[0050] Video Decoding: Decoding is the reverse process of encoding. During intra-frame encoding, video frames are typically encoded in a specific order, such as top-to-bottom or left-to-right. Decoding a video frame follows the same order. The decoding process can also be called the encoding-reconstruction process. The video encoder transforms, quantizes, and entropy-encodes the predicted residual rather than the original pixel values. The decoding process reconstructs the current decoding unit based on the residual between the target prediction unit (which has already been decoded by the time the current decoding unit is reached) and the current decoding unit.
[0051] Coding Tree Unit (CTU): In the High Efficiency Video Coding (HEVC) (H.265) standard, a frame of image is divided into multiple slices, each slice is encoded and decoded independently, and each slice is divided into multiple tree-structured coding units (CTUs). CTUs are further divided into CUs, PUs, and Transform Units (TUs), which separates encoding, prediction, and transformation, making the processing more flexible.
[0052] A coding unit (CU), also known as a coding block or coded image unit, is essentially an image block that belongs to a portion of a frame. In H.265, the CU is defined as the most basic square coding unit. The size of a CU in H.265 can range from 8×8 to a maximum of 64×64, and its size is represented by depth: Depth = 0 is 64×64, and Depth = 3 is 8×8. Therefore, the length and width of a CU with Depth = k and Depth = k+1 differ by a factor of two. Thus, a CU of size 2N×2N with Depth = k can be divided into four CUs of size N×N with Depth = k+1, which can be represented by a quadtree structure. Furthermore, each CU contains one or more PU and TU structures; therefore, we can consider the CU as the outermost frame, and the size of the PU and TU cannot exceed that of the CU.
[0053] A prediction unit (PU), also known as a prediction block or prediction image unit, is essentially an image block that has already been encoded before the current coding unit. In H.265, the PU is the basic unit for intra-frame and inter-frame prediction, with sizes ranging from 4×4 to 64×64. Besides the symmetric motion partitioning methods similar to H.264 (2N×2N, N×N, 2N×N, and N×2N), H.265 also provides asymmetric motion partitioning (AMP), including 2N×nU, 2N×nD, nL×2N, and nR×2N. The uppercase letters represent the positions of the shorter-side partitions. When encoding a coding unit, the residual between the coding unit and the prediction unit is encoded to improve coding efficiency.
[0054] Intra-frame prediction: Predictive coding leverages the correlation (Normalized Cross-Correlation, NCC) between discrete signals, using one or more preceding signals to predict the next signal, and then encoding the difference between the actual and predicted values (prediction error). If the prediction is accurate, the error will be small. Under the same precision requirements, fewer bits can be used for encoding, achieving data compression. Predictive coding can include intra-frame prediction and inter-frame prediction; this application primarily focuses on intra-frame prediction. Intra-frame prediction first appeared in H.264 video coding. Intra-frame prediction is the process of searching for the prediction unit most similar to the current coding unit within a frame. Therefore, when encoding a coding unit, the residual between the coding unit and the prediction unit is encoded to improve coding efficiency. Because intra-frame prediction can efficiently compress intra-frame coding, it continues to be used in H.265. H.264 uses 4*4 (for complex images) and 16*16 (for smooth images), for a total of 9 prediction modes. The two modes can be used in combination to flexibly apply different processing methods and sizes to areas with varying image quality and texture details. H.265 provides up to 33 directional and 2 non-directional prediction modes, totaling 35, making predictions more accurate.
[0055] Difference information, also known as residuals or prediction errors, refers to the differences between two image patches. For example, an image patch can be represented by the pixel values it contains. The difference information can be determined based on the differences between the pixel values of the two image patches. For instance, it could be the sum of the squares of the differences between the pixel values of the two image patches, or the sum of the absolute errors between the pixel values, etc.
[0056] The Sum of Squared Errors (SSE), also known as the Sum of Squared Differences (SSD), is a measure of how well a data prediction model fits. Given a set of data points and the model's predictions, the SSE is the sum of the squares of all prediction errors. Its formula is as follows:
[0057]
[0058] Where y is the actual observed value, These are the model's predicted values.
[0059] Sum of Absolute Error (SAE): Also known as Sum of Absolute Difference (SAD). SAD is the sum of the absolute values of the differences between corresponding pixels in an image, and is a region-based local matching criterion.
[0060] The embodiments of this application relate to video coding technology, and are mainly designed based on intra-frame prediction technology in video coding.
[0061] In related technologies, intra-frame prediction methods based on IBC are used for predicting repetitive textures; see [link to relevant documentation]. Figure 2 The diagram shows the principle of the intra-frame prediction method based on IBC in related technologies. The IBC method searches the already encoded region of the current CU in the current frame, takes the block that is closest to it as the PU of the current CU, obtains the prediction residual of the current CU based on the PU, and further performs transformation, quantization and entropy coding processes to complete the entire coding process of the current CU.
[0062] However, for more complex CUs or due to the presence of noise, the IBC method may not be able to accurately determine the PU, which will lead to poor coding performance for the current CU. For example, if the prediction residual between the selected PU and the current CU is large, the coding result of the current CU will still require a lot of bits, resulting in poor coding performance, or it will consume a lot of resources, resulting in low coding efficiency.
[0063] Based on this, embodiments of this application provide a video frame coding method that improves upon the intra-frame copy prediction method. In this method, at least one matching unit is further determined based on the original prediction unit, and the target prediction unit for the current coding unit is determined by combining the original prediction unit and these matching units. This is equivalent to further processing based on the currently determined original prediction unit, which helps improve the correlation between the finally determined target prediction unit and the current coding unit, thereby improving the accuracy of determining the final prediction unit for the current coding unit and thus improving coding performance. For example, the greater the correlation between the target prediction unit and the current coding unit, the smaller the difference between them, reducing the amount of data required for coding and helping to reduce the probability of coding reconstruction distortion. Therefore, even for more complex coding units or under the influence of noise, the prediction unit can be determined more accurately, improving the performance of intra-frame prediction and contributing to improved video coding and decoding performance.
[0064] Furthermore, this application also provides a video frame decoding method. In this method, the target encoding result of the current decoding unit includes data encoding result, prediction unit indication information, and position indication information. The position indication information determines the original prediction unit of the current decoding unit. Combining the prediction unit indication information with the original prediction unit determines the target prediction unit ultimately used by the current decoding unit during encoding. Then, based on the decoding result of the target prediction unit, the data encoding result is decoded. Since the target prediction unit is a prediction unit with better encoding performance than the prediction unit obtained by the IBC method in related technologies, the amount of data that needs to be processed during decoding is correspondingly less. Alternatively, less encoding reconstruction distortion results in more accurate decoding results, thus improving decoding performance.
[0065] Furthermore, in this embodiment of the application, for at least one matching unit determined by the original prediction unit, there is no need to transmit the corresponding location information, but only a minimum of one bit is needed to indicate the final target prediction unit in combination with the original prediction unit, thus saving transmission resource overhead.
[0066] The technical solutions in this application embodiment can be implemented by combining artificial intelligence (AI) technology.
[0067] AI (Artificial Intelligence) is the theory, methods, technology, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to achieve optimal results. In other words, artificial intelligence is a comprehensive technology within computer science that attempts to understand the essence of intelligence and produce a new kind of intelligent machine that can react in a way similar to human intelligence. Artificial intelligence studies the design principles and implementation methods of various intelligent machines, enabling them to possess the functions of perception, reasoning, and decision-making.
[0068] Artificial intelligence (AI) is a comprehensive discipline encompassing a wide range of fields, including both hardware and software technologies. Fundamental AI technologies generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing, pre-trained model technology, operating / interactive systems, and mechatronics. Among these, pre-trained models, also known as large-scale models or foundational models, can be widely applied to downstream tasks across various AI fields after fine-tuning. AI software technologies primarily include computer vision (CV), speech processing, natural language processing, and machine learning / deep learning.
[0069] Computer vision is the science that studies how to enable machines to "see." More specifically, it refers to machine vision, which uses cameras and computers to replace human eyes in recognizing and measuring targets, and further processes images to create images more suitable for human observation or transmission to instruments. As a scientific discipline, computer vision studies related theories and technologies, attempting to build artificial intelligence systems capable of extracting information from images or multidimensional data. Large model technology has brought significant changes to the development of computer vision technology. Pre-trained models in the vision field, such as Swin-transformer, ViT, V-MOE, and MAE, can be quickly and widely applied to downstream tasks after fine-tuning. Computer vision technologies typically include image processing, image recognition, image semantic understanding, image retrieval, OCR, video processing, video semantic understanding, video content / behavior recognition, 3D object reconstruction, 3D technology, virtual reality, augmented reality, simultaneous localization and mapping (SLAM), and common biometric recognition technologies such as facial recognition and fingerprint recognition.
[0070] The technical solutions in this application mainly involve computer vision (CV) technology. For example, in the technical solutions of this application, CV technology can be combined to identify the similarities between two image blocks, thereby assisting in selecting a prediction unit that is closer to the current coding unit.
[0071] The following is a brief introduction to the application scenarios to which the technical solutions of the embodiments of this application are applicable. It should be noted that the application scenarios described below are only for illustrating the embodiments of this application and are not intended to limit the scope. In specific implementation, the technical solutions provided by the embodiments of this application can be flexibly applied according to actual needs.
[0072] The solution provided in this application can be applied to any scenario related to video frame encoding and decoding, such as video calls, short videos, video websites, or remote conferencing, for video compression, improving video compression performance and saving video transmission bandwidth. Figure 3 The diagram shown is an application scenario provided by an embodiment of this application. In this scenario, an encoding device 101 and a decoding device 102 may be included.
[0073] Encoding device 101 is used to execute the steps of the video frame encoding method to achieve video frame encoding, and decoding device 102 is used to execute the steps of the video frame decoding method to decode video frames. Encoding device 101 or decoding device 102 can be implemented using a terminal device or a server. For example, both encoding device 101 and decoding device 102 can be terminal devices; or encoding device 101 can be a server and decoding device 102 can be a terminal device; or encoding device 101 can be a terminal device and decoding device 102 can be a server; or both encoding device 101 and decoding device 102 can be servers.
[0074] The aforementioned terminal devices can be, for example, mobile phones, tablets (PADs), laptops, desktop computers, smart home appliances, aircraft, XR devices (such as augmented reality (AR) devices, mixed reality (MR) devices, or virtual reality (VR) devices), smart in-vehicle devices, and smart wearable devices, etc.
[0075] The aforementioned servers can be, for example, independent physical servers, server clusters or distributed systems composed of multiple physical servers, or cloud servers that provide basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks (CDN), and big data and artificial intelligence platforms, but are not limited to these.
[0076] It should be noted that, Figure 3 The examples shown are merely illustrative. In reality, the number of encoding end devices 101 and decoding end devices 102 is not limited. For example, one encoding end device 101 can correspond to multiple decoding end devices 102, or multiple encoding end devices 101 can correspond to one decoding end device 102, etc. No specific limitation is made in the embodiments of this application.
[0077] In this embodiment, the encoding device 101 and the decoding device 102 can be connected through one or more networks 103. The network 103 can be a wired network or a wireless network. For example, the wireless network can be a mobile cellular network or a Wireless-Fidelity (WIFI) network. Of course, it can also be other possible networks, and this embodiment does not limit them.
[0078] The video frame encoding and decoding method provided by the exemplary embodiments of this application will be described below with reference to the accompanying drawings and the application scenarios described above. It should be noted that the application scenarios described above are only shown to facilitate understanding of the spirit and principles of this application, and the embodiments of this application are not limited in any way in this respect.
[0079] See Figure 4 The diagram shown is a flowchart illustrating a video frame encoding method provided in an embodiment of this application. This method can be executed by a computer device, such as a computer that can be... Figure 3 The encoding end device 101 shown is described below. The specific implementation process of this method is as follows:
[0080] Step 401: From at least one encoded image unit within the target video frame, determine the original prediction unit whose similarity to the current encoded unit meets the intra-frame copy condition.
[0081] In this embodiment, the target video frame can be any frame in the video to be encoded. A frame is an image, and an image can be divided into multiple image units. During encoding, each image unit within the target video frame is encoded separately. For example, each image unit within the frame can be encoded in the order from top to bottom and from left to right within the image.
[0082] Within the target video frame, the current coding unit refers to the currently encoded image block, which typically comprises multiple pixels. For example, a typical coding unit is a square image region composed of the pixel values of multiple pixels. Figure 5 As shown, this is an example diagram for determining the original prediction unit. The current coding unit C is shown as an 8x8 image region. The gray area represents the already encoded image region. Then, it is necessary to determine the original prediction unit PU0 in the already encoded image units that has a similarity with the current coding unit that meets the intra-frame copy condition.
[0083] Specifically, the original prediction unit is an image block of the same size as the current coding unit. The image blocks that meet the intra-frame copy condition can be determined by calculating the similarity between each image block in at least one encoded image block and the current coding unit, and then used as the original prediction unit.
[0084] In one possible implementation, the intra-frame copy condition can be the highest similarity, that is, after calculating the similarity between each image unit and the current coding unit, the image unit with the highest similarity is selected as the original prediction unit based on the similarity.
[0085] In another possible implementation, the intra-frame copy condition can also be a similarity threshold of not less than a preset threshold. That is, after calculating the similarity between each image unit and the current coding unit, image units with a similarity threshold of not less than the preset threshold are selected as the original prediction units. If multiple image units meet the condition, the one with the highest similarity can be selected as the original prediction unit.
[0086] In this embodiment of the application, the original prediction unit of the current coding unit can be determined from at least one encoded image unit using the IBC prediction method.
[0087] In one implementation, at least one image unit can be all the image units that have been encoded within the target video frame. In this way, the prediction unit that is most suitable for the current coding unit can be found from within the target video frame, which helps to improve the accuracy of the prediction unit.
[0088] In another implementation, at least one image unit can be an image unit within a certain area near the current coding unit within the target video frame. This reduces the search range during intra-frame prediction, helps to reduce the complexity of intra-frame prediction, and improves coding efficiency.
[0089] For example, the search matching range of the prediction unit of the coding unit can be preset, such as a 64x64 image region around the current coding unit. Then at least one image unit is an image unit that has been encoded within the search matching range.
[0090] For example, when encoding video frames, a frame is usually divided into multiple CTUs, and a CTU can be further divided into multiple coding units. Therefore, at least one image unit can be an image unit that has already been encoded within the CTU where the current coding unit is located, or it can be an image unit that has already been encoded within a certain area of CTUs near the current coding unit.
[0091] In this embodiment of the application, the matching of the original prediction units can be achieved in the following two ways.
[0092] In one implementation, matching can be performed sequentially according to the order of at least one image unit within the target video frame to select the original prediction unit that meets the intra-frame copy condition.
[0093] For example, see Figure 6A The diagram shown is a schematic representation of a search and matching process provided in an embodiment of this application. See also... Figure 6A As shown, it can start from the first image unit in at least one image unit, such as Figure 6AAs shown in step ①, the image unit is matched with the current coding unit; then, according to the step size set for the search match, for example, 1 pixel, it can be moved 1 pixel to the right, as shown in step ①. Figure 6A As shown in step ②, the image unit of this region is matched with the current coding unit; then it is moved one pixel to the right, as shown in step ③. Figure 6A As shown in step ②, the image units in the region are matched with the current coding unit until all image units are matched. Of course, the step size can also be other values, and this embodiment does not limit this, and the search and matching process can also be performed in parallel.
[0094] This implementation method comprehensively considers all image units and selects image units that are as close as possible to the current coding unit as prediction units, thereby improving the accuracy of prediction units and thus improving coding efficiency.
[0095] In another implementation, considering that the correlation between image blocks in neighboring regions may be higher, in order to improve the prediction efficiency of the original prediction unit, a search can be performed in the vicinity of the current coding unit first, and the similarity between each image unit and the current coding unit can be calculated. If there is an image unit that meets the intra-frame copy condition, it is used as the original prediction unit. If there is no image unit that meets the intra-frame copy condition, the search matching range can be expanded to search and match in image units that are further away from the current coding unit until an image unit that meets the intra-frame copy condition is found and used as the original prediction unit.
[0096] See Figure 6B The diagram shown illustrates another search and matching process provided in an embodiment of this application. See also... Figure 6B As shown, the process includes two search and matching steps. In the first search and matching step, a search and matching can be performed within a relatively small image region. If no matching image unit is found, then... Figure 6B The search area is expanded to include a second search and matching operation within a larger image region. During each search and matching process, techniques such as... Figure 6A The procedure is as shown, so it will not be elaborated upon further.
[0097] This implementation method can minimize the computational load when searching for matching in the original prediction unit, thereby improving coding efficiency.
[0098] In this embodiment of the application, the similarity between the current coding unit and the image unit can be characterized in the following ways:
[0099] (1) The SSE between the image unit and the current coding unit: If the intra-frame copy condition is to search for the image unit with the highest similarity to the current coding unit, then the SSE between each image unit and the current coding unit can be calculated separately, and the image unit with the smallest SSE between it and the current coding unit can be selected as the original prediction unit. The SSE between the original prediction unit and the current coding unit can be expressed as follows:
[0100]
[0101] Where C represents the current coding unit, C(i,j) represents the data of the pixel in the i-th row and j-th column of the current coding unit, such as the pixel value, PU0 represents the original prediction unit, and PU0(i,j) represents the data of the pixel in the i-th row and j-th column of the original prediction unit, such as the pixel value.
[0102] (2) The SAD between the image unit and the current coding unit is similar to that of SSE. The image unit with the smallest SAD between it and the current coding unit can be selected as the original prediction unit.
[0103] (3) Calculate the similarity between the current coding unit and each image unit using a similarity calculation method. For example, cosine similarity can be used. Alternatively, a similarity calculation model can be pre-trained using machine learning. The current coding unit and each image unit can be used as input. After feature extraction of the current coding unit and image units, the trained similarity calculation model is used to calculate the similarity between the pairwise image features. The original prediction unit is then selected based on this similarity. Using machine learning for training helps improve the applicability of this similarity calculation method in intra-frame prediction scenarios, extracting more suitable feature information for the current scenario, thus more accurately calculating the similarity between the current coding unit and image units, and improving the accuracy of intra-frame prediction.
[0104] In this embodiment of the application, during the encoding process, when calculating similarity parameters such as SSE, SAD, or other similarity metrics between the current encoding unit and the encoded image units, the data used by the current encoding unit is the real data of that encoding unit in the target video frame. The encoded image units are typically reconstructed using the encoded data of that image unit, i.e., data restored based on the encoding result of that image unit. See [link to relevant documentation]. Figure 5 As shown, Figure 5The image areas represented by gray are all encoded reconstructed data, while the current encoding unit uses real data. The main purpose of this operation is to take into account that the decoding end does not know the real data of each image unit, but can only obtain the encoded reconstructed data (i.e., the decoding data). By using the encoded reconstructed data during encoding, the data used by the encoding end and the decoding end during the encoding and decoding process can be strictly consistent, reducing the probability of decoding errors and other situations.
[0105] Step 402: From at least one image unit, determine a set of matching units whose similarity to the original prediction unit meets the preset matching conditions.
[0106] In this embodiment, although a similar original prediction unit can be selected from the target video frame, there will always be a certain difference between the original prediction unit and the current coding unit. Under otherwise unchanged conditions, the smaller the difference, the better the intra-frame prediction performance. Therefore, in practical applications, it is not limited to seeking prediction units in the target video frame; the constraint of the target video frame can be overcome, and prediction units with better intra-frame prediction performance can be sought in scenes outside the target video frame. Furthermore, this embodiment also considers the consistency between the decoding and encoding ends to minimize changes to both. Specifically, for the encoding end, the original prediction unit can be determined through the aforementioned step 401, and the determined original prediction unit will be indicated to the decoding end. Therefore, the information that both the encoding and decoding ends can know is the original prediction unit. Based on this, this embodiment proposes a new method for determining prediction units based on the original prediction unit.
[0107] Specifically, after determining the original prediction unit through step 401, this embodiment of the application uses the original prediction unit as a reference and searches again from at least one image unit for at least one matching unit that is close to the original prediction unit. It should be noted that the matching unit here refers to the matching unit that meets the preset matching conditions with the original prediction unit, and it is essentially still an image unit composed of multiple pixels in the target video frame.
[0108] In one possible implementation, at least one matching unit that is close to the original prediction unit can be searched among all the image units included in at least one image unit. This allows for a comprehensive consideration of all image units to select matching units that are as close as possible to the original prediction unit, thereby improving the accuracy of the subsequently obtained prediction units and thus enhancing coding performance.
[0109] In another possible implementation, considering that expanding the search range will directly affect the data processing efficiency of the encoding and decoding ends, the search range can be preset based on device performance considerations. After the original prediction unit is determined, an image region range with a preset range size based on the original prediction unit can be determined from at least one image unit. Then, from the image units included in the image region range, a set of matching units whose similarity with the original prediction unit meets the preset matching conditions can be determined. The set of matching units includes PU1 to PUn, where N is a positive integer.
[0110] For example, a set of matching units is selected from the image units of size SxS that have already been encoded around the original prediction unit. The value of SxS can be determined by considering both accuracy and the encoding and decoding performance of the device; for example, SxS can be 16x16. The number N of matching units in the set can be set empirically or experimentally; for example, N can be 3 or other possible values, which are not limited in this embodiment.
[0111] In real-world scenarios, the search range can also be dynamically adjusted. For example, the size of the search range can be determined based on the processing performance of the current device. When the device's processing performance is good, the search range can be expanded appropriately, while when the device's processing performance is poor, the search range can be reduced appropriately.
[0112] In this embodiment, the set of matching units whose similarity meets the preset matching conditions may include at least one matching unit with the highest similarity. For example, if an upper limit N for the number of matching units is specified in advance, then the N image units with the highest similarity are selected to form the matching unit set. Alternatively, the set of matching units whose similarity meets the preset matching conditions may include image units with a similarity greater than a preset similarity threshold.
[0113] For example, with N=3, see Figure 7 The diagram illustrates the search and matching process for the original prediction unit. It shows that after determining the corresponding original prediction unit PU0 based on the current coding unit C, the original prediction unit PU0 can be used as a benchmark to search for a set of matching units whose similarity meets preset matching conditions, such as... Figure 7 The figures PU1 to PU3 are shown.
[0114] Similar to the search and matching process for the current coding unit, the search and matching process for the original prediction unit can be referred to in step 401, and therefore will not be elaborated further. The similarity can also be measured using SAD or SSE. Taking SSE as an example, the SSE of the original prediction unit PU0 and the matching unit PU1 can be expressed as follows:
[0115]
[0116] Wherein, PU1(i,j) represents the data of the pixel point in the i-th row and j-th column of the matching unit PU1, such as the pixel value.
[0117] Step 403: Determine the target prediction unit for the current coding unit based on the original prediction unit and the matching unit set.
[0118] After obtaining the set of matching units, the target prediction unit of the current coding unit can be determined based on the original prediction unit and the set of matching units.
[0119] In one possible implementation, image fusion processing can be performed based on the original prediction unit and the matching unit set to obtain a fused prediction unit. The fused prediction unit is not an image unit existing in the target video frame, but a new image unit obtained by combining the original prediction unit and each matching unit. The fused prediction unit combines the features of the original prediction unit and each matching unit, and may have a smaller difference from the current coding unit, or it may have a larger difference from the current coding unit. Therefore, after obtaining the fused prediction unit, it is also necessary to determine the target prediction unit suitable for the current coding unit from the original prediction unit and the fused prediction unit.
[0120] In this embodiment of the application, when performing image fusion processing on the original prediction unit and the matching unit set, considering that different image units have different degrees of difference (or similarity) with the current coding unit, in order to retain as much similarity as possible with the current coding unit (or the original prediction unit), different weight values can be assigned to different image units, and then the image fusion processing process can be realized according to the weight values.
[0121] See Figure 8 The diagram shown is a flowchart illustrating the image fusion processing procedure provided in this embodiment. First, for each image unit—that is, each matching unit in the original prediction unit and matching unit set—a weight value is determined based on its similarity to the current coding unit. This weight value represents the proportion of the corresponding image unit's features in the fusion prediction unit. Then, based on the obtained weight values, the original prediction unit and each matching unit are weighted to obtain the fusion prediction unit. For example, the weighting method can be a weighted average. Taking a weighted average image fusion process as an example, PU1 to PU2 are obtained by matching the original prediction unit PU0. n Then, the other prediction unit of the current coding unit can be obtained by weighted averaging, and the fused prediction unit can be represented as:
[0122] PU w = (PU0*W0+PU1*W1+…+PU n *Wn )
[0123] W0+W1+…+W n =1
[0124] Among them, W0, W1, ..., W n These are the weight values for each image unit, such as W0 representing the weight value of PU0, and so on.
[0125] In one implementation, considering that the greater the similarity between an image unit and the current coding unit, the higher its weight value, that is, the weight value of each image unit is proportional to its corresponding similarity. In this way, the image unit with greater similarity to the current coding unit has a larger weight value, thereby allowing the features in that image unit to occupy a larger proportion in the obtained fusion prediction unit.
[0126] In another implementation, since the original prediction unit has the highest similarity to the current coding unit, it can be assigned the highest weight value. For example, the weight value of the original prediction unit is fixed, which is equivalent to the total weight value of the matching unit set being fixed. The weight value of each matching unit in the matching unit set can be determined based on the similarity between each matching unit and the original prediction unit. Then, based on the obtained weight values, the original prediction unit and each matching unit are weighted to obtain the fused prediction unit.
[0127] Taking SSE (Short-Side Similarity) as an example, the larger the SSE, the smaller the similarity; conversely, the smaller the SSE, the larger the similarity. See also... Figure 8 As shown, taking a set of matching units containing 3 matching units as an example, in terms of weight design, since PU0 is the most similar block to the current coding unit, PU0 is given the highest weight, for example, W0 = 1 / 2. W1, W2 and W3 are calculated based on the matching error SSE between them and PU0. The smaller the matching error, the larger the weight value.
[0128] In one example, the weight values of each matching unit can be calculated as follows:
[0129] Wi=((SSEa-SSEi) / (2*SSEa)) / 2
[0130] SSEa = SSE1 + SSE2 + SSE3
[0131] SSEa represents the sum of SSE between PU1 to PU3 and the original prediction unit, respectively. Wi represents the weight value of the i-th matching unit, where i can be 1, 2, or 3.
[0132] In this embodiment, by assigning different weight values to the original prediction unit and each matching unit, the features in each image unit can be fused into the fusion prediction unit according to different degrees of fusion. For example, the larger the weight value of the original prediction unit, the more content of the original prediction unit is included in the fusion prediction unit, which helps to increase the probability that the fusion prediction unit is more similar to the current coding unit, thereby improving the accuracy of the final prediction unit and improving coding performance.
[0133] In this embodiment of the application, in addition to weighted processing, other image fusion processing methods can also be used, such as pooling multiple image units or performing convolutional feature processing on multiple image units through convolutional neural networks (CNNs). This embodiment of the application does not limit the method.
[0134] After obtaining the fused prediction unit, for the current coding unit, there are two candidate prediction units: the original prediction unit and the fused prediction unit. One of these two units can be selected as the final prediction unit, i.e., the target prediction unit. For example, the best prediction unit for the current coding unit can be selected in the following two ways:
[0135] In one possible implementation, a prediction unit that is more similar to the current coding unit can be selected based on the similarity between the original prediction unit and the fused prediction unit and the current coding unit, respectively.
[0136] Specifically, the similarity between each of the original prediction units and the fused prediction units and the current coding unit is obtained, and the one with the highest similarity is determined as the target prediction unit. The similarity can be measured using parameters such as SSE or SAD.
[0137] For example, see Figure 9A As shown, taking SSE as the similarity denoted as SSE, the original prediction unit PU0 and the fused prediction unit PU can be calculated separately. w The similarity between each unit and the current coding unit, i.e., SSE, is used to obtain SSE0 and SSE. w SSE0 represents the similarity between the original prediction unit PU0 and the current coding unit. w Indicates the fusion prediction unit PU w The similarity between the current coding unit and SSE0 is then compared with that of SSE0. w This is to determine the target prediction unit. The selection method can then be expressed as follows:
[0138]
[0139] Wherein, SSEw is the SSE between the current coding unit and the fusion prediction unit PUw, which can be calculated using the following formula:
[0140]
[0141] Among them, PU w (i, j) represents the fusion prediction unit PU w The data of the pixel in the i-th row and j-th column, such as the pixel value.
[0142] According to the formula above, if SSE0 ≤ SSE w This indicates that the similarity between the original prediction unit PU0 and the current coding unit is not less than that between the fused prediction unit PU0 and the current coding unit. w The similarity between the original prediction unit and the current coding unit determines whether to select the original prediction unit PU0 as the target prediction unit. Conversely, if SSE0 > SSE0, the original prediction unit PU0 is selected as the target prediction unit. w This indicates that the similarity between the original prediction unit PU0 and the current coding unit is less than that between the fused prediction unit PU0 and the current coding unit. w The similarity between the current coding unit and the selected fusion prediction unit (PU) is determined by the similarity between the current coding unit and the PU. w As a target prediction unit.
[0143] By implementing the above methods, selecting a prediction unit that is closer to the current coding unit results in a smaller difference between the final target prediction unit and the current coding unit, less difference information to be encoded, fewer bits occupied by the final encoding result, and a smaller difference, which helps to reduce the distortion probability during encoding reconstruction and improve encoding and decoding performance.
[0144] In another possible implementation, a prediction unit more similar to the current coding unit can be selected based on the coding performance of the original prediction unit and the fused prediction unit. To measure coding performance, this embodiment employs a rate-distortion optimization process at the coding end to obtain the optimal prediction unit, i.e., comparing the coding costs when using the original prediction unit and the fused prediction unit as the target prediction unit. Coding cost reflects coding performance; generally, coding cost can characterize at least one of the transmission resources occupied by the coding result and coding distortion. Less transmission resources, lower coding distortion, and better coding performance.
[0145] Since the processes for determining the coding cost are similar for both the original prediction unit and the fused prediction unit, this section uses the original prediction unit as an example to illustrate the process of obtaining the coding cost. See also... Figure 9B As shown, this application embodiment provides another schematic diagram of the selection process for a prediction unit.
[0146] Specifically, such as Figure 9B As shown, the original prediction unit is used as the target prediction unit. The residual between the original prediction unit PU0 and the current coding unit C is calculated. Based on the difference between the original prediction unit PU0 and the current coding unit C, the current coding unit C is attempted to be encoded to obtain the prediction coding result. Then, decoding is performed based on the prediction coding result, that is, the current coding unit is reconstructed based on the prediction coding result. The reconstruction result (or decoding result) is the content of the current coding unit that the decoder can obtain. By comparing the reconstruction result with the actual content of the current coding unit, the reconstruction distortion of the current coding unit can be determined. In other words, by comparing the reconstruction result with the actual content of the current coding unit, the difference between the decoding result and the current coding unit can be determined. Similarly, the reconstruction distortion of the fused prediction unit can be calculated in the same way.
[0147] The predictive coding result reveals the number of bits it occupies, i.e., the number of bits required to transmit it. Generally, fewer bits result in better compression. The difference information reveals the level of encoding reconstruction distortion in the current coding unit. Generally, fewer differences indicate lower encoding reconstruction distortion and better coding performance. Therefore, the coding cost when using the original prediction unit as the target prediction unit can be obtained based on at least one of the number of bits occupied by the predictive coding result and the difference information.
[0148] For example, the encoding cost can be represented as follows:
[0149] J=D+λR
[0150] Where D represents the encoding reconstruction distortion, R represents the number of bits required for encoding, and λ is the Lagrange daily number.
[0151] So, let's continue as follows Figure 9B As shown, after determining the coding cost when using the original prediction unit as the target prediction unit, and the coding cost when using the fused prediction unit as the target prediction unit, the two coding costs can be compared, and the one with the lowest coding cost between the original prediction unit and the fused prediction unit can be determined as the target prediction unit. This selection method can then be expressed as follows:
[0152]
[0153] Among them, J0 and J w These represent the use of the original prediction unit PU0 and the fused prediction unit PU, respectively. w The encoding cost after predictive encoding. That is, if J0 ≤ J wIf J0 > J0, it indicates that the encoding cost of the original prediction unit PU0 is lower, and therefore the original prediction unit PU0 is determined as the target prediction unit. Conversely, if J0 > J0, then the original prediction unit PU0 is determined as the target prediction unit. w This indicates that the fusion prediction unit PU w If the encoding cost is lower, then the fusion prediction unit PU is determined. w As a target prediction unit.
[0154] This implementation method allows for the selection of prediction units with lower encoding costs. Lower encoding costs result in better final encoding performance, such as lower transmission resource consumption or lower encoding reconstruction distortion, thereby improving encoding performance.
[0155] In this embodiment, besides determining the target prediction unit based on the original prediction unit and the matching unit set as described above, other methods can also be used. For example, in another possible implementation, since both the original prediction unit and the matching unit set are determined based on image similarity, in order to improve the accuracy of the finally determined prediction unit, the original prediction unit and the matching unit set can be sorted from other dimensions, and the similarity and other dimensions' sorting can be combined to determine the target prediction unit of the current coding unit.
[0156] Step 404: Based on the difference information between the target prediction unit and the current coding unit, encode the current coding unit to obtain the target coding result of the current coding unit.
[0157] In this embodiment of the application, after determining the target prediction unit, predictive coding can be used to encode the current coding unit based on the difference information between the target prediction unit and the current coding unit. For example, the prediction residual (i.e., difference information) of the current coding unit is obtained based on the target prediction unit, and further transformation, quantization, and entropy coding processes are performed to complete the entire coding process of the current coding unit.
[0158] When generating the target encoding result of the current coding unit, since it is necessary to notify the decoding end which target prediction unit is currently being used, or to notify the decoding end that the intra-block copy prediction method was used during the encoding process, corresponding prediction unit indication information will also be generated based on the final selected target prediction unit. This prediction unit indication information is used to indicate the final determined target prediction unit.
[0159] In one possible implementation, if the target prediction unit is determined from the aforementioned original prediction unit and fused prediction unit, the prediction unit indication information includes one bit. If the target prediction unit is the original prediction unit, the value of the one bit is determined to be a first value indicating the original prediction unit, such as 0 or 1; if the target prediction unit is the fused prediction unit, the value of the one bit is determined to be a second value indicating the fused prediction unit, such as 1 or 0.
[0160] For example, to enable the decoder to determine the intra-block copy prediction (IPC) scheme used by the current coding unit, a 1-bit syntax element can be added to represent the optimal IPC scheme for the current coding unit. This syntax element could be named `muiti_prediction_flag`. If `muiti_prediction_flag = 1`, it indicates that the PU (Programmable Logic Optimization) scheme is used. w Prediction was performed on the current block. The decoder first obtained the original prediction unit PU0 based on the block vector information predicted by IBC. Based on the original prediction unit PU0, PU1 to PU3 were further searched at the decoder, and then the fused prediction unit PU of the current coding unit was calculated. w Conversely, if muiti_prediction_flag = 0, it means that the original prediction unit PU0 was used to predict the current coding unit, and there is no need to further search PU1 to PU3.
[0161] In this implementation, by adding a 1-bit syntax element to represent the best intra-block copy prediction method for the current coding unit, the additional matching blocks PU1 to PU3 can be obtained at the encoder and decoder. Therefore, no additional bits are needed to identify the position information of these matching blocks, thus saving bit overhead.
[0162] In another possible implementation, if the target prediction unit is selected from the set of original prediction units and matching units, the positional information of these image units can be used as the prediction unit indication information. For example, the positional information is the relative positional information between the image unit and the current coding unit.
[0163] Furthermore, in order for the decoder to know the original prediction unit of the current coding unit, it also needs to obtain position indication information to indicate the location of the original prediction unit. For example, this position indication information can be the block vector (BV) information of the current coding unit.
[0164] Furthermore, based on the data encoding result, prediction unit indication information, and position indication information, the target encoding result for the current encoding unit is generated. That is, for the current encoding unit, its final encoding result can include three aspects of information: data encoding result, position indication information, and prediction unit indication information. This assists the decoding end in accurately determining the target prediction unit corresponding to the current encoding unit based on the prediction unit indication information and position indication information during decoding, thereby accurately implementing the decoding process and improving decoding performance.
[0165] It should be noted that the above process mainly describes the encoding process of a single image unit in a video frame. In real-world scenarios, however, all image units within a video frame can be encoded using the same method. The encoded result can be transmitted over a network to the decoding end for video playback after decoding.
[0166] Therefore, this application also provides a video frame decoding method, see [link to relevant documentation]. Figure 10 The diagram shown is a flowchart illustrating a video frame decoding method provided in an embodiment of this application. This method can be executed by a computer device, which can be... Figure 3 The specific implementation process of this method for the decoding device shown is as follows:
[0167] Step 1001: Obtain the target encoding result of the current decoding unit within the target video frame. The target encoding result includes the data encoding result, prediction unit indication information, and position indication information. The prediction unit indication information is used to indicate the target prediction unit that the current decoding unit will ultimately use during encoding, and the position indication information is used to indicate the position of the original prediction unit of the current decoding unit.
[0168] The decoding end can obtain the data packets of the current video through the network. Each data packet can include data from one or more video frames, and this data includes the target encoding result, i.e., through the aforementioned... Figure 4 The target encoding result is obtained from a portion of the encoding methods. For example, in a video conference, a terminal device can receive video data packets sent by the peer device; or, when playing video, the terminal device can also pull video data packets from the server; in a video call, the terminal device can also receive video data packets sent by the peer device.
[0169] The content included in the target encoding result has been... Figure 4 Some of the encoding methods have been introduced, so please refer to the above introduction; they will not be repeated here.
[0170] Step 1002: Based on the position indication information, determine the original prediction unit from at least one decoded image unit within the target video frame.
[0171] The position indication information is used to indicate the location of the original prediction unit of the current decoding unit. For example, the position indication information can indicate the position information of the original prediction unit relative to the current decoding unit. For example, the position indication information is (8, 9), which indicates that the original prediction unit is located at a position 8 pixels away from the current decoding unit on the X-axis and 9 pixels away on the Y-axis. Thus, the original prediction unit corresponding to the current decoding unit can be determined from at least one decoded image unit in the target video frame.
[0172] Step 1003: Determine the target prediction unit based on the prediction unit indication information and the original prediction unit.
[0173] In traditional predictive coding, the information commonly known to both the encoder and decoder is the original prediction unit of the current encoding (decoding) unit. Therefore, in this embodiment, during encoding, at least one matching unit is matched based on the original prediction unit. Similarly, the decoder can also match at least one matching unit based on the original prediction unit. The matching units obtained by the encoder and decoder are the same, thereby ensuring that the data of the encoder and decoder remain consistent.
[0174] The prediction unit indication information is used to indicate which specific prediction unit is being targeted, or in other words, to indicate the intra-block copy prediction method used by the encoder during encoding. Therefore, by combining the prediction unit indication information, the target decoding unit ultimately used by the current decoding unit during encoding can be determined.
[0175] See Figure 11 The diagram shown is a schematic representation of the decoding process provided in an embodiment of this application. See also... Figure 11 As shown, firstly, based on the position indication information, the original prediction unit PU0 corresponding to the current decoding unit can be determined. Then, based on the prediction unit indication information, the target decoding unit ultimately used by the current decoding unit during encoding can be determined. For example, the prediction unit indication information includes one bit, which indicates the intra-block copy prediction method used by the encoder during encoding. If this bit is the first value indicating the original prediction unit PU0, such as... Figure 11 The first value in the bit is 0, indicating that when the traditional intra-block copy prediction method is used, the original prediction unit PU0 is determined as the target prediction unit PU. If this bit is the second value indicating the fused prediction unit, such as... Figure 11 If the second value in the formula is 1, then from at least one image unit, a set of matching units whose similarity to the original prediction unit PU0 meets the preset matching conditions is determined, and the original prediction unit PU0 is matched with the set of matching units (e.g., ...). Figure 11 As shown in PU1 to PU3, the target prediction unit of the current coding unit is determined, for example, the fusion prediction unit PU is used.w The target prediction unit (PU) has been identified.
[0176] Step 1004: Based on the decoding result of the target prediction unit, decode the data encoding result to obtain the decoding result of the current decoding unit.
[0177] Specifically, the decoding process is the inverse of the encoding process. It involves reconstructing the data encoding corresponding to the current decoding unit based on the residual and target prediction unit indicated in the data encoding result, in order to obtain the decoding result of the current decoding unit.
[0178] In this implementation, multiple matching units for IBC prediction can be obtained using less bit overhead (such as the 1-bit identifier bit mentioned above), which improves the overall prediction performance of IBC, thereby improving the performance of intra-frame prediction and video coding performance.
[0179] The following example illustrates the solution of this application embodiment. See [link to example]. Figure 12 The diagram shown is an example of the interaction between an encoding end and a decoding end provided in an embodiment of this application.
[0180] Step 1201: The encoder performs initial IBC prediction, that is, for the input current coding unit C, the IBC method is used to search for its most similar prediction unit PU0. For example, PU0 is obtained by performing block matching search in a given region that has already been encoded around the current coding unit, that is, searching for the block with the smallest SSE or SAD of the current block as the prediction unit PU0.
[0181] Step 1202: The encoder performs a block matching search, that is, based on the prediction unit PU0 obtained by IBC prediction, using a process similar to block matching search, it searches for several matching units PU1 to PUn that best match the prediction unit PU0 in the already encoded area around the prediction unit PU0 (such as the range SxS around PU0).
[0182] Step 1203: The encoder performs weighted prediction, that is, given PU0 and PU1~PUn, the other prediction block PU of the current coding unit C is predicted. w It can be obtained by weighted averaging. Among them, PU0 is the most similar block to the current coding unit C, so PU0 is given the highest weight. The weight values of the other matching units can be calculated based on the matching error SSE between them and PU0. The smaller the matching error, the larger the weight.
[0183] Step 1204: The encoder performs the best prediction unit selection, that is, from prediction unit PU0 and the weighted prediction unit PU0. w The best prediction unit for the current coding unit C is obtained by selecting from the options provided.
[0184] Step 1205: The encoding end performs encoding, that is, based on the best prediction unit, generates the target encoding result for the current encoding unit C. The target encoding result includes BV information, the best IBC prediction method indication information, and the data encoding result.
[0185] The optimal IBC prediction method indicator is represented by a 1-bit syntax element. If it is 1, it means that PUw was used to predict the current coding unit. Conversely, if it is 0, it means that prediction unit PU0 was used to predict the current block.
[0186] Step 1206: The encoding end sends the target encoding result to the decoding end.
[0187] Step 1207: The decoder determines the prediction unit PU0 based on the BV information.
[0188] Step 1208: If the optimal IBC prediction method indication information is 1, the decoder further searches based on prediction unit PU0 to obtain PU1 to PU3, thereby calculating prediction unit PUw, and performing decoding based on prediction unit PUw.
[0189] Step 1209: If the optimal IBC prediction method indication information is 1, the decoder does not need to further search PU1 to PU3, and directly performs decoding based on prediction unit PU0.
[0190] In summary, this application proposes an intra-frame block copy prediction method based on multiple matching blocks. By utilizing the best-matching block of the current block, a matching search is performed within a certain neighborhood to obtain several matching blocks of the best-matching block. Then, a weighted average is taken of the best-matching block and several matching blocks to obtain the final predicted block of the current coding block. This allows the use of more matching blocks to further improve the accuracy of intra-frame block copy prediction, thereby enhancing the performance of intra-frame prediction. Simultaneously, it saves syntax element information used to represent the positions of multiple matching blocks, saving bit overhead and reducing transmission resource consumption.
[0191] Please see Figure 13 Based on the same inventive concept, embodiments of this application also provide a video frame encoding device 130, which includes:
[0192] Intra-frame copy prediction unit 1301 is used to determine, from at least one image unit already encoded within the target video frame, an original prediction unit whose similarity to the current coding unit meets the intra-frame copy condition.
[0193] Matching unit 1302 is used to determine a set of matching units from at least one image unit whose similarity to the original prediction unit meets a preset matching condition;
[0194] The prediction unit selection unit 1303 is used to determine the target prediction unit of the current coding unit based on the original prediction unit and the matching unit set;
[0195] The encoding execution unit 1304 is used to encode the current encoding unit based on the difference information between the target prediction unit and the current encoding unit, so as to obtain the target encoding result of the current encoding unit.
[0196] In one possible implementation, the prediction unit selection unit 1303 is specifically used for:
[0197] Image fusion processing is performed based on the original prediction units and the matching unit set to obtain fused prediction units;
[0198] The target prediction unit is determined from the original prediction unit and the fused prediction unit.
[0199] In one possible implementation, the prediction unit selection unit 1303 is specifically used for:
[0200] Based on the similarity between each matching unit in the matching unit set and the original prediction unit, a weight value is obtained for each matching unit. Then, based on these weight values, the original prediction unit and each matching unit are weighted to obtain the fused prediction unit; or...
[0201] Based on the similarity between each matching unit in the original prediction unit and the current coding unit, corresponding weight values are obtained. Then, the original prediction unit and each matching unit are weighted according to the obtained weight values to obtain the fused prediction unit.
[0202] In one possible implementation, the prediction unit selection unit 1303 is specifically used for:
[0203] The similarity between the original prediction unit and the fused prediction unit and the current coding unit is obtained respectively.
[0204] The target prediction unit is determined from the original prediction unit and the fused prediction unit that has the highest similarity to the current coding unit.
[0205] In one possible implementation, the prediction unit selection unit 1303 is specifically used for:
[0206] Determine the coding cost when the original prediction unit is used as the target prediction unit, and determine the coding cost when the fused prediction unit is used as the target prediction unit; wherein the coding cost is used to characterize at least one of the transmission resources occupied by the coding result and the coding distortion.
[0207] The target prediction unit is determined by the one with the lowest encoding cost between the original prediction unit and the fused prediction unit.
[0208] In one possible implementation, the prediction unit selection unit 1303 is specifically used for:
[0209] Based on the difference information between the original prediction unit and the current coding unit, the current coding unit is encoded to obtain the prediction coding result;
[0210] Decoding is performed based on the predictive coding results, and the difference information between the decoding result of the current coding unit and the current coding unit is determined.
[0211] Based on at least one of the number of bits occupied by the prediction coding result and the difference information, the coding cost when the original prediction unit is used as the target prediction unit is obtained.
[0212] In one possible implementation, the encoding execution unit 1304 is specifically used for:
[0213] Based on the difference information between the target prediction unit and the current coding unit, the current coding unit is encoded to obtain the data coding result;
[0214] Based on the target prediction unit, corresponding prediction unit indication information is generated. The prediction unit indication information is used to indicate the finally determined target prediction unit.
[0215] Obtain location indication information to indicate the location of the original prediction unit;
[0216] Based on the data encoding results, prediction unit indication information, and position indication information, the target encoding result of the current encoding unit is generated.
[0217] In one possible implementation, if the target prediction unit is determined from the original prediction unit and the fused prediction unit, the prediction unit indication information includes one bit; then the encoding execution unit 1304 is specifically used for:
[0218] If the target prediction unit is the original prediction unit, determine the value of a bit to indicate the first value of the original prediction unit;
[0219] If the target prediction unit is a fusion prediction unit, determine the value of a bit to indicate the second value of the fusion prediction unit.
[0220] In one possible implementation, the matching unit 1302 is specifically used for:
[0221] From at least one image unit, determine an image region range of a predetermined size with the original prediction unit as the reference point;
[0222] From the image units included in the image region, determine the set of matching units whose similarity to the original prediction unit meets the preset matching conditions.
[0223] By using the above-described apparatus, at least one matching unit is further determined based on the original prediction unit, and the target prediction unit of the current coding unit is determined by combining the original prediction unit and these matching units. This is equivalent to further processing based on the currently determined original prediction unit, which helps to improve the correlation between the finally determined target prediction unit and the current coding unit, thereby improving the accuracy of determining the final prediction unit for the current coding unit and thus improving coding performance.
[0224] Please see Figure 14 Based on the same inventive concept, embodiments of this application also provide a video frame decoding device 140, which includes:
[0225] The encoding result acquisition unit 1401 is used to obtain the target encoding result of the current decoding unit within the target video frame. The target encoding result includes the data encoding result, prediction unit indication information and position indication information. The prediction unit indication information is used to indicate the target prediction unit that the current decoding unit will ultimately use when encoding, and the position indication information is used to indicate the position of the original prediction unit of the current decoding unit.
[0226] The prediction unit determination unit 1402 is configured to determine an original prediction unit from at least one decoded image unit within a target video frame based on position indication information; and to determine a target prediction unit based on the prediction unit indication information and the original prediction unit.
[0227] The decoding execution unit 1403 is used to decode the data encoding result based on the decoding result of the target prediction unit to obtain the decoding result of the current decoding unit.
[0228] In one possible implementation, the prediction unit indicates that the information includes one bit;
[0229] The prediction unit determines unit 1402, which is specifically used for:
[0230] If a bit represents the first value indicating the original prediction unit, the original prediction unit is determined as the target prediction unit; or,
[0231] If a bit is a second value indicating the fusion prediction unit, then from at least one image unit, a set of matching units whose similarity to the original prediction unit meets the preset matching conditions is determined, and the target prediction unit of the current coding unit is determined based on the original prediction unit and the set of matching units.
[0232] The target prediction unit used in the above device has better coding performance than the prediction unit obtained by the IBC method of related technologies. Accordingly, the amount of data that needs to be processed during decoding is less. Or, since there is less coding reconstruction distortion, the corresponding decoding result is more accurate, which helps to improve decoding performance.
[0233] This device can be used to execute the methods shown in the various embodiments of this application. Therefore, the functions that each functional module of this device can achieve can be referred to the description of the foregoing embodiments, and will not be repeated here.
[0234] Please see Figure 15 Based on the same technical concept, embodiments of this application also provide a computer device. In one embodiment, the computer device can be... Figure 1 The server shown or Figure 2 The cloud-related device shown is a computer device such as... Figure 15 As shown, it includes a memory 1501, a communication module 1503, and one or more processors 1502.
[0235] The memory 1501 is used to store computer programs executed by the processor 1502. The memory 1501 may mainly include a program storage area and a data storage area. The program storage area may store the operating system and programs required to run instant messaging functions, etc.; the data storage area may store various instant messaging information and operation instruction sets, etc.
[0236] Memory 1501 may be volatile memory, such as random-access memory (RAM); memory 1501 may also be non-volatile memory, such as read-only memory, flash memory, hard disk drive (HDD), or solid-state drive (SSD); or memory 1501 may be any other medium capable of carrying or storing desired program code in the form of instructions or data structures and accessible by a computer, but is not limited thereto. Memory 1501 may be a combination of the above-described memories.
[0237] Processor 1502 may include one or more central processing units (CPUs) or digital processing units, etc. Processor 1502 is used to implement the above-described video frame encoding and video frame decoding methods when calling computer programs stored in memory 1501.
[0238] The communication module 1503 is used to communicate with terminal devices and other servers.
[0239] This application embodiment does not limit the specific connection medium between the memory 1501, communication module 1503, and processor 1502. This application embodiment... Figure 15 The memory 1501 and the processor 1502 are connected via a bus 1504, and the bus 1504 is in Figure 15 The diagram uses thick lines to describe the connections between other components; these are for illustrative purposes only and should not be considered limiting. The 1504 bus can be divided into address bus, data bus, control bus, etc. For ease of description, Figure 15 It is described using only a thick line, but does not indicate that there is only one bus or one type of bus.
[0240] The memory 1501 stores a computer storage medium, which stores computer-executable instructions. The computer-executable instructions are used to implement the video frame encoding and video frame decoding methods of the embodiments of this application. The processor 1502 is used to execute the video frame encoding and video frame decoding methods of the above embodiments.
[0241] In another embodiment, the computer device can also be a terminal device, such as... Figure 1 The terminal device shown. In this embodiment, the structure of the computer device can be as follows. Figure 16 As shown, it includes components such as: communication component 1610, memory 1620, display unit 1630, camera 1640, sensor 1650, audio circuit 1660, Bluetooth module 1670, processor 1680, etc.
[0242] The communication component 1610 is used to communicate with the server. In some embodiments, it may include a Circuit-Based Wireless Fidelity (WiFi) module, which is a short-range wireless transmission technology. Computer devices can use WiFi modules to help users send and receive information.
[0243] The memory 1620 can be used to store software programs and data. The processor 1680 executes various functions of the terminal device and data processing by running the software programs or data stored in the memory 1620. The memory 1620 may include high-speed random access memory, and may also include non-volatile memory, such as at least one disk storage device, flash memory device, or other volatile solid-state storage device. The memory 1620 stores an operating system that enables the terminal device to run. In this application, the memory 1620 may store the operating system and various application programs, and may also store code that executes the video frame encoding and video frame decoding methods of the embodiments of this application.
[0244] The display unit 1630 can also be used to display information input by the user or information provided to the user, as well as various menus of the terminal device, forming a graphical user interface (GUI). Specifically, the display unit 1630 may include a display screen 1632 disposed on the front of the terminal device. The display screen 1632 may be configured as a liquid crystal display, a light-emitting diode, or the like. The display unit 1630 can be used to display various search result pages or sub-result pages in the embodiments of this application.
[0245] The display unit 1630 can also be used to receive input digital or character information and generate signal inputs related to user settings and function control of the terminal device. Specifically, the display unit 1630 may include a touch screen 1631 disposed on the front of the terminal device, which can collect touch operations of the user on or near it, such as clicking buttons, dragging scroll boxes, etc.
[0246] The touchscreen 1631 can be placed on top of the display screen 1632, or the touchscreen 1631 and the display screen 1632 can be integrated to realize the input and output functions of the terminal device. After integration, it can be referred to as a touch display screen. In this application, the display unit 1630 can display the application and the corresponding operation steps.
[0247] Camera 1640 can be used to capture still images, which users can then post comments on via the application. There can be one or multiple cameras 1640. An object is projected onto a photosensitive element through a lens, generating an optical image. This photosensitive element can be a charge-coupled device (CCD) or a complementary metal-oxide-semiconductor (CMOS) phototransistor. The photosensitive element converts the light signal into an electrical signal, which is then transmitted to the processor 1680 for conversion into a digital image signal.
[0248] The terminal device may also include at least one sensor 1650, such as an accelerometer 1651, a proximity sensor 1652, a fingerprint sensor 1653, and a temperature sensor 1654. The terminal device may also be equipped with other sensors such as a gyroscope, barometer, hygrometer, thermometer, infrared sensor, light sensor, and motion sensor.
[0249] Audio circuitry 1660, speaker 1661, and microphone 1662 provide an audio interface between the user and the terminal device. Audio circuitry 1660 converts received audio data into electrical signals, which are then transmitted to speaker 1661, where they are converted into sound signals for output. The terminal device can also be equipped with volume buttons for adjusting the volume of the sound signal. Conversely, microphone 1662 converts collected sound signals into electrical signals, which are then received by audio circuitry 1660, converted back into audio data, and output to communication component 1610 for transmission to, for example, another terminal device, or to memory 1620 for further processing.
[0250] The Bluetooth module 1670 is used to interact with other Bluetooth devices that also have a Bluetooth module via the Bluetooth protocol. For example, a terminal device can establish a Bluetooth connection with a wearable computer device (such as a smartwatch) that also has a Bluetooth module through the Bluetooth module 1670, thereby exchanging data.
[0251] The processor 1680 is the control center of the terminal device, connecting various parts of the terminal through various interfaces and lines. It executes software programs stored in the memory 1620 and calls data stored in the memory 1620 to perform various functions and process data. In some embodiments, the processor 1680 may include one or more processing units; the processor 1680 may also integrate an application processor and a baseband processor, wherein the application processor mainly handles the operating system, user interface, and applications, and the baseband processor mainly handles wireless communication. It is understood that the baseband processor may not be integrated into the processor 1680. In this application, the processor 1680 can run the operating system, applications, user interface display and touch response, as well as the video frame encoding and decoding methods of this embodiment. Furthermore, the processor 1680 is coupled to the display unit 1630.
[0252] Based on the same inventive concept, embodiments of this application also provide a storage medium storing a computer program that, when run on a computer, causes the computer to perform the steps in the video frame encoding and video frame decoding methods described above according to various exemplary embodiments of this application.
[0253] In some possible implementations, various aspects of the video frame encoding and video frame decoding methods provided in this application can also be implemented in the form of a computer program product, which includes a computer program. When the program product is run on a computer device, the computer program is used to cause the computer device to perform the steps in the video frame encoding and video frame decoding methods according to the various exemplary embodiments of this application described above. For example, the computer device can perform the steps of each embodiment.
[0254] The program product may employ any combination of one or more readable media. A readable medium may be a readable signal medium or a readable storage medium. A readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples (a non-exhaustive list) of readable storage media include: electrical connections having one or more wires, portable disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.
[0255] The program product of the embodiments of this application may employ a portable compact disc read-only memory (CD-ROM) and include a computer program, and may run on a computer device. However, the program product of this application is not limited thereto. In this application, the readable storage medium may be any tangible medium that contains or stores a program, and the computer program included therein may be used by or in conjunction with a command execution system, apparatus, or device.
[0256] A readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, carrying a readable computer program. This propagated data signal may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A readable signal medium may also be any readable medium other than a readable storage medium, capable of sending, propagating, or transmitting a program for use by or in conjunction with a command execution system, apparatus, or device.
[0257] Computer programs contained on readable media may be transmitted using any suitable medium, including but not limited to wireless, wired, optical fiber, RF, etc., or any suitable combination thereof.
[0258] Computer programs for performing the operations of this application can be written in any combination of one or more programming languages, including object-oriented programming languages such as Java and C++, as well as conventional procedural programming languages such as the "C" language or similar programming languages.
[0259] It should be noted that although several units or sub-units of the device have been mentioned in the detailed description above, this division is merely exemplary and not mandatory. In fact, according to embodiments of this application, the features and functions of two or more units described above can be embodied in one unit. Conversely, the features and functions of one unit described above can be further divided and embodied by multiple units.
[0260] Furthermore, although the operations of the method of this application are described in a specific order in the accompanying drawings, this does not require or imply that these operations must be performed in that specific order, or that all the operations shown must be performed to achieve the desired result. Additionally or alternatively, certain steps may be omitted, multiple steps may be combined into one step, and / or one step may be broken down into multiple steps.
[0261] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0262] Although preferred embodiments of this application have been described, those skilled in the art, upon learning the basic inventive concept, can make other changes and modifications to these embodiments. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments as well as all changes and modifications falling within the scope of this application.
[0263] Obviously, those skilled in the art can make various modifications and variations to this application without departing from the spirit and scope of this application. Therefore, if such modifications and variations fall within the scope of the claims of this application and their equivalents, this application also intends to include such modifications and variations.
Claims
1. A video frame encoding method, characterized in that, The method includes: From at least one encoded image unit within the target video frame, determine the original prediction unit whose similarity to the current encoded unit meets the intra-frame copy condition; From the at least one image unit, determine a set of matching units whose similarity to the original prediction unit meets preset matching conditions; Based on the original prediction unit and the matching unit set, the target prediction unit of the current coding unit is determined; Based on the difference information between the target prediction unit and the current coding unit, the current coding unit is encoded to obtain the target coding result of the current coding unit.
2. The method as described in claim 1, characterized in that, The step of determining the target prediction unit of the current coding unit based on the original prediction unit and the matching unit set includes: Image fusion processing is performed based on the original prediction unit and the matching unit set to obtain a fused prediction unit; The target prediction unit is determined from the original prediction unit and the fused prediction unit.
3. The method as described in claim 2, characterized in that, The step of performing image fusion processing based on the original prediction unit and the plurality of matching units to obtain a fused prediction unit includes: Based on the similarity between each matching unit in the matching unit set and the original prediction unit, a weight value for each matching unit is obtained. Then, based on the obtained weight values, a weighted average is applied between the original prediction unit and each matching unit to obtain the fused prediction unit; or, Based on the similarity between the original prediction unit and each matching unit in the matching unit set and the current coding unit, corresponding weight values are obtained. Based on the obtained weight values, the original prediction unit and each matching unit are weighted to obtain the fusion prediction unit.
4. The method as described in claim 2, characterized in that, Determining the target prediction unit from the original prediction unit and the fused prediction unit includes: The similarity between the original prediction unit and the fused prediction unit and the current coding unit is obtained respectively. The target prediction unit is determined by the one with the highest similarity to the current coding unit among the original prediction unit and the fused prediction unit.
5. The method as described in claim 2, characterized in that, Determining the target prediction unit from the original prediction unit and the fused prediction unit includes: The encoding cost is determined when the original prediction unit is used as the target prediction unit, and the encoding cost is determined when the fused prediction unit is used as the target prediction unit; wherein the encoding cost is used to characterize at least one of the transmission resources occupied by the encoding result and the encoding distortion. The target prediction unit is determined as the one with the lowest encoding cost among the original prediction unit and the fused prediction unit.
6. The method as described in claim 5, characterized in that, The determination of the encoding cost when using the original prediction unit as the target prediction unit includes: Based on the difference information between the original prediction unit and the current coding unit, the current coding unit is encoded to obtain the prediction coding result; Decoding is performed based on the predicted coding result, and the difference information between the decoding result of the current coding unit and the current coding unit is determined. Based on at least one of the number of bits occupied by the predicted coding result and the difference information, the coding cost when the original prediction unit is used as the target prediction unit is obtained.
7. The method according to any one of claims 1 to 6, characterized in that, Based on the difference information between the target prediction unit and the current coding unit, the current coding unit is encoded to obtain the target coding result of the current coding unit, including: Based on the difference information between the target prediction unit and the current coding unit, the current coding unit is encoded to obtain the data coding result; Based on the target prediction unit, corresponding prediction unit indication information is generated, which is used to indicate the finally determined target prediction unit; Obtain location indication information to indicate the position of the original prediction unit; Based on the data encoding result, the prediction unit indication information, and the position indication information, the target encoding result of the current encoding unit is generated.
8. The method as described in claim 7, characterized in that, If the target prediction unit is determined from the original prediction unit and the fused prediction unit, then the prediction unit indication information includes one bit; The step of generating corresponding first indication information based on the target prediction unit includes: If the target prediction unit is the original prediction unit, the value of the one bit is determined to be a first value indicating the original prediction unit; If the target prediction unit is the fusion prediction unit, the value of the one bit is determined to be a second value indicating the fusion prediction unit.
9. The method according to any one of claims 1 to 6, characterized in that, Determining a set of matching units from the at least one image unit whose similarity to the original prediction unit meets preset matching conditions includes: From the at least one image unit, determine an image region range of a preset size with the original prediction unit as the reference point; From the image units included in the image region, determine a set of matching units whose similarity to the original prediction unit meets preset matching conditions.
10. A video frame decoding method, characterized in that, The method includes: Obtain the target encoding result of the current decoding unit within the target video frame. The target encoding result includes data encoding result, prediction unit indication information, and position indication information. The prediction unit indication information is used to indicate the target prediction unit that the current decoding unit will ultimately use during encoding. The position indication information is used to indicate the position of the original prediction unit of the current decoding unit. Based on the location indication information, the original prediction unit is determined from at least one decoded image unit within the target video frame; The target prediction unit is determined based on the prediction unit indication information and the original prediction unit; Based on the decoding result of the target prediction unit, the data encoding result is decoded to obtain the decoding result of the current decoding unit.
11. The method as described in claim 10, characterized in that, The prediction unit indicates that the information includes one bit; Based on the first indication information and the original prediction unit, the target prediction unit is determined, including: If the stated bit is a first value indicating the original prediction unit, then the original prediction unit is determined as the target prediction unit; or... If the bit is a second value indicating the fusion prediction unit, then from the at least one image unit, a set of matching units whose similarity to the original prediction unit meets the preset matching conditions is determined, and the target prediction unit of the current coding unit is determined based on the original prediction unit and the set of matching units.
12. A video frame encoding apparatus, characterized in that, The device includes: An intra-frame copy prediction unit is used to determine, from at least one encoded image unit within a target video frame, an original prediction unit whose similarity to the current coding unit meets the intra-frame copy condition. A matching unit is used to determine, from the at least one image unit, a set of matching units whose similarity to the original prediction unit meets preset matching conditions; A prediction unit selection unit is used to determine the target prediction unit of the current coding unit based on the original prediction unit and the matching unit set. The encoding execution unit is used to encode the current encoding unit based on the difference information between the target prediction unit and the current encoding unit, so as to obtain the target encoding result of the current encoding unit.
13. A video frame decoding device, characterized in that, The device includes: The encoding result acquisition unit is used to obtain the target encoding result of the current decoding unit within the target video frame. The target encoding result includes data encoding result, prediction unit indication information and position indication information. The prediction unit indication information is used to indicate the target prediction unit that the current decoding unit will ultimately use when encoding. The position indication information is used to indicate the position of the original prediction unit of the current decoding unit. A prediction unit determining unit is configured to determine the original prediction unit from at least one decoded image unit within the target video frame based on the location indication information; and to determine the target prediction unit based on the prediction unit indication information and the original prediction unit. The decoding execution unit is used to decode the data encoding result based on the decoding result of the target prediction unit to obtain the decoding result of the current decoding unit.
14. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 9 and 10 to 11.
15. A computer storage medium having a computer program stored thereon, characterized in that, When executed by a processor, the computer program implements the steps of the method described in any one of claims 1 to 9 and 10 to 11.
16. A computer program product, comprising a computer program, characterized in that, When executed by a processor, the computer program implements the steps of the method described in any one of claims 1 to 9 and 10 to 11.