Video coding method and device and electronic equipment
By quantizing the forward predicted coded frames in the video encoding method and determining the distortion impact range, the problem of poor encoding of simple scene video frames in the prior art is solved, and more efficient and high-quality video encoding is achieved.
Patent Information
- Application Number
- CN202311781268.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-22
- Publication Date
- 2025-06-24
AI Technical Summary
When the existing video encoding method handles video frames with simple scenes, the quantization parameter settings are not accurate enough, resulting in poor encoding effects.
By acquiring multiple target video frames and video frame image groups of video frames to be encoded, each forward predicted coded frame is quantized, the current forward coded frame is determined and the distortion impact range is analyzed, and the quantization index data is updated to optimize the encoding process.
The encoding effect of simple scene video frames is improved, and the quality and efficiency of video encoding are improved by accurately determining the distortion impact range of forward prediction coded frames and adaptively adjusting the quantization index data.
Smart Images

Figure CN120201193A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of artificial intelligence technology, and in particular, to a video encoding method, apparatus, and electronic device. Background Art
[0002] Artificial intelligence is a technology that simulates human intelligence and aims to create intelligent machines that can autonomously learn and process information. Artificial intelligence software technologies mainly include several major directions such as computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning. Specifically, computer vision technology (Computer Vision, CV) is a science that studies how to enable machines to "see". Further, it refers to using cameras and computers to replace the human eye to identify and measure targets, etc., for machine vision, and further perform graphic processing to make the computer process into an image that is more suitable for human eye observation or transmission to an instrument for detection. Computer vision technology usually includes technologies such as video processing, video semantic understanding, video content / behavior recognition, and three-dimensional object reconstruction.
[0003] During the video encoding process, the rate-distortion cost of different encoding modes can be calculated to find the mode with the smallest rate-distortion cost of the encoding mode. Before determining the rate-distortion cost, the distortion of the previous frame of the P frame (Predictive-coded Picture) and the distortion of the current P frame can be statistically analyzed, and then the degree to which the distortion of the P frame is affected by the distortion of the previous frame can be weighted statistically, and the quantization parameter (Quantizer Parameter, QP) and the allocated bit rate that should be specifically set for the current P frame can be analyzed based on this value. Thus, the rate-distortion cost of different encoding modes can be determined by combining the set quantization parameter and the allocated bit rate. However, for test sequences with different video contents, the importance of the P frames in the GOP (Group of Picture) is different. For videos with strong reference, that is, videos with simple scenes, the distortion of the P frame in the previous GOP not only affects the current GOP, but also affects the distortion of all frames in the next GOP. Therefore, in the above existing quantization parameter determination method, the quantization parameter setting for video frames with simple scenes is not accurate enough, resulting in poor encoding effects for video frames with simple scenes. Summary of the Invention
[0004] In view of the above technical problems, the present disclosure proposes a video encoding method, apparatus, and electronic device.
[0005] According to an aspect of an embodiment of the present disclosure, a video encoding method is provided, including:
[0006] Obtain multiple target video frames of the video to be encoded, as well as multiple groups of video frame images corresponding to the video to be encoded; each group of video frame images includes multiple sub - groups of images; each sub - group of images in at least one sub - group of images includes a forward - predicted encoded frame;
[0007] Perform quantization analysis on each forward - predicted encoded frame to obtain first quantization index data of each forward - predicted encoded frame; the first quantization index data is used to indicate the spatial detail compression condition of the video frame;
[0008] Sequentially determine a current forward - encoded frame from each group of video frame images, where the current forward - encoded frame is the forward - predicted encoded frame with the earliest time sequence among the forward - predicted encoded frames in each group of video frame images that are not affected by the distortion of the analyzed encoded frames, and the analyzed encoded frames are the forward - predicted encoded frames for which the distortion influence range analysis has been performed;
[0009] Perform distortion influence range analysis on the current forward - encoded frame to obtain influence range index data corresponding to the current forward - encoded frame; the influence range index data corresponding to the current forward - encoded frame represents the range of video frames affected by the distortion of the current forward - encoded frame during the encoding process of the video to be encoded;
[0010] Based on the influence range index data corresponding to the current forward - encoded frame, update the first quantization index data of the forward - predicted encoded frames within the target sub - group of images corresponding to the current forward - encoded frame to obtain updated quantization index data corresponding to the current forward - encoded frame; the target sub - group of images corresponding to the current forward - encoded frame is the sub - group of images belonging to the influence range indicated by the influence range index data corresponding to the current forward - encoded frame;
[0011] Based on the updated quantization index data corresponding to multiple forward - predicted encoded frames within each group of video frame images and the first quantization index data corresponding to the un - updated forward video frames within each group of video frame images, perform encoding processing on the video to be encoded to obtain target video encoding data.
[0012] According to another aspect of the embodiments of the present disclosure, there is provided a video encoding apparatus, including:
[0013] A data acquisition module, configured to obtain multiple target video frames of the video to be encoded, as well as multiple groups of video frame images corresponding to the video to be encoded; each group of video frame images includes multiple sub - groups of images; each sub - group of images in at least one sub - group of images includes a forward - predicted encoded frame;
[0014] A forward quantization analysis module, configured to perform quantization analysis on each forward - predicted encoded frame to obtain first quantization index data of each forward - predicted encoded frame; the first quantization index data is used to indicate the degree of spatial detail compression of the video frame;
[0015] A current frame determination module, configured to sequentially determine a current forward coding frame from each group of video frame images, where the current forward coding frame is the forward prediction coding frame with the earliest time sequence among the forward prediction coding frames in each group of video frame images that are not affected by the distortion of the analyzed coding frames, and the analyzed coding frames are the forward prediction coding frames for which the distortion influence range analysis is performed;
[0016] An influence range analysis module, configured to perform distortion influence range analysis on the current forward coding frame to obtain influence range index data corresponding to the current forward coding frame; the influence range index data corresponding to the current forward coding frame represents the range of video frames affected by the distortion of the current forward coding frame during the coding process of the video to be coded;
[0017] An update module, configured to update the first quantization index data of the forward prediction coding frames in the target sub-image group corresponding to the current forward coding frame based on the influence range index data corresponding to the current forward coding frame to obtain updated quantization index data corresponding to the current forward coding frame; the target sub-image group corresponding to the current forward coding frame is the sub-image group belonging to the influence range indicated by the influence range index data corresponding to the current forward coding frame;
[0018] A coding processing module, configured to perform coding processing on the video to be coded based on the updated quantization index data corresponding to multiple forward prediction coding frames in each group of video frame images and the first quantization index data corresponding to the non-updated forward video frames in each group of video frame images to obtain target video coding data.
[0019] Optionally, the update module includes:
[0020] A first merging processing module, configured to perform merging processing on the target sub-image group corresponding to the current forward coding frame based on the influence range index data corresponding to the current forward coding frame to obtain a merged sub-image group corresponding to the current forward coding frame;
[0021] A position indication acquisition module, configured to determine position indication information corresponding to each first prediction coding frame among the multiple first prediction coding frames corresponding to the current forward coding frame; the multiple first prediction coding frames corresponding to the current forward coding frame are the multiple forward prediction coding frames included in the merged sub-image group corresponding to the current forward coding frame, and the position indication information corresponding to each first prediction coding frame is used to indicate the position of each first prediction coding frame among the multiple first prediction coding frames;
[0022] A first update module, configured to update at least one first quantization metric data corresponding to a to-be-updated predicted coded frame among a plurality of first predicted coded frames corresponding to the current forward coded frame based on the influence range metric data corresponding to the current forward coded frame and the position indication information corresponding to each first predicted coded frame, so as to obtain updated quantization metric data corresponding to the current forward coded frame.
[0023] Optionally, the first update module includes:
[0024] A first weight acquisition module, configured to determine first weight information corresponding to each to-be-updated predicted coded frame based on the position indication information corresponding to each to-be-updated predicted coded frame and a preset mapping relationship; the preset mapping relationship is a mapping relationship between a plurality of position indication information and a plurality of preset weight information;
[0025] A first increment analysis module, configured to perform quantization increment analysis on each to-be-updated predicted coded frame based on the first weight information corresponding to each to-be-updated predicted coded frame and the influence range metric data corresponding to the current forward coded frame, so as to obtain first increment metric data corresponding to each to-be-updated predicted coded frame;
[0026] A second update module, configured to update the first quantization metric data corresponding to each to-be-updated predicted coded frame based on the first increment metric data corresponding to each to-be-updated predicted coded frame, so as to obtain first updated metric data corresponding to each to-be-updated predicted coded frame.
[0027] Optionally, the first update module includes:
[0028] A second weight acquisition module, configured to determine second weight information corresponding to the second predicted coded frame based on the position indication information corresponding to the second predicted coded frame and a preset mapping relationship;
[0029] A second increment analysis module, configured to perform quantization increment analysis on the second predicted coded frame based on the second weight information corresponding to the second predicted coded frame and the influence range metric data corresponding to the current forward coded frame, so as to obtain second increment metric data corresponding to the second predicted coded frame;
[0030] A third update module, configured to update the first quantization metric data corresponding to the second predicted coded frame among the plurality of first predicted coded frames based on the second increment metric data, so as to obtain second updated metric data corresponding to the second predicted coded frame.
[0031] Optionally, the update module includes:
[0032] A quantity acquisition module, configured to determine the number of frames to be updated corresponding to the current forward encoded frame and the number of affected frames corresponding to the influence range index data, where the number of frames to be updated is the number of forward prediction encoded frames in each group of video frame images that are not affected by the distortion of the analyzed encoded frames; the number of affected frames is the number of forward prediction encoded frames belonging to the influence range corresponding to the influence range index data;
[0033] A comparison processing module, configured to perform comparison processing on the number of affected frames and the number of frames to be updated to obtain a target comparison result;
[0034] Correspondingly, the first merging processing module includes:
[0035] A second merging processing module, configured to, when the target comparison result indicates that the number of affected frames is less than or equal to the number of frames to be updated, perform merging processing on the target sub-image group corresponding to the current forward encoded frame based on the influence range index data corresponding to the current forward encoded frame, to obtain a merged sub-image group corresponding to the current forward encoded frame.
[0036] Optionally, the apparatus further includes:
[0037] A complexity analysis module, configured to perform inter-frame complexity analysis on each forward prediction encoded frame based on each forward prediction encoded frame and the previous video frame of each forward prediction encoded frame, to obtain inter-frame complexity index data corresponding to each forward prediction encoded frame, where the inter-frame complexity index data corresponding to each forward prediction encoded frame represents the degree of difference between each forward prediction encoded frame and the previous video frame of each forward prediction encoded frame;
[0038] Correspondingly, the influence range analysis module includes:
[0039] A distortion influence analysis module, configured to perform distortion influence analysis on the current forward encoded frame based on the inter-frame complexity index data corresponding to the current forward encoded frame, to obtain influence range index data corresponding to the current forward encoded frame.
[0040] Optionally, the distortion influence analysis module includes:
[0041] An inter-frame correlation analysis module, configured to perform inter-frame correlation analysis on the current forward encoded frame based on the inter-frame complexity index data corresponding to the current forward encoded frame, to obtain inter-frame correlation index data corresponding to the current forward encoded frame; the inter-frame correlation index data corresponding to the current forward encoded frame represents the degree of correlation between the current forward encoded frame and the previous video frame of the current forward encoded frame;
[0042] A matching processing module, configured to perform matching processing on a preset metric value range and the inter-frame correlation metric data corresponding to the current forward-encoded frame to obtain a target matching result;
[0043] A metric data determination module, configured to determine the influence range metric data corresponding to the current forward-encoded frame from the two endpoint metric data corresponding to the preset metric value range and the inter-frame correlation metric data based on the target matching result.
[0044] Optionally, the metric data determination module includes:
[0045] A first metric determination module, configured to perform floor operation on the inter-frame correlation metric data when the target matching result indicates that the inter-frame correlation metric data belongs to the preset metric value range, to obtain the first range metric data corresponding to the current forward-encoded frame.
[0046] Optionally, the metric data determination module further includes:
[0047] An endpoint metric acquisition module, configured to determine the adjacent endpoint metric data from the two endpoint metric data corresponding to the preset metric value range when the target matching result indicates that the inter-frame correlation metric data does not belong to the preset metric value range, where the adjacent endpoint metric data is the metric data closest to the inter-frame correlation metric data among the two endpoint metric data;
[0048] A second metric determination module, configured to use the adjacent endpoint metric data as the second range metric data corresponding to the current forward-encoded frame.
[0049] Optionally, the encoding processing module includes:
[0050] A bidirectional quantization analysis module, configured to perform quantization analysis on multiple bidirectional prediction encoded frames in each video frame picture group based on the updated quantization metric data corresponding to multiple forward prediction encoded frames in each video frame picture group and the first quantization metric data corresponding to the non-updated forward video frames in each video frame picture group, to obtain the second quantization metric data corresponding to each of the multiple bidirectional prediction encoded frames in each video frame picture group;
[0051] A target dimension determination module, configured to determine the target encoding dimension corresponding to each target video frame from multiple preset encoding dimensions based on the updated quantization metric data corresponding to multiple forward prediction encoded frames in each video frame picture group, the first quantization metric data corresponding to the non-updated forward video frames in each video frame picture group, and the second quantization metric data corresponding to each of the multiple bidirectional prediction encoded frames in each video frame picture group;
[0052] A target data determination module, configured to determine the target quantization metric data corresponding to each target video frame from the updated quantization metric data corresponding to multiple forward prediction coded frames within each group of video frame images, the first quantization metric data corresponding to the non-updated forward video frames within each group of video frame images, and the second quantization metric data corresponding to each of the multiple bidirectional prediction coded frames within each group of video frame images;
[0053] A video frame encoding module, configured to perform video frame encoding processing on each target video frame based on the target encoding dimension corresponding to each target video frame and the target quantization metric data corresponding to each target video frame, to obtain the target encoded data corresponding to each target video frame.
[0054] Optionally, the target dimension determination module includes:
[0055] A rate-distortion analysis module, configured to perform rate-distortion analysis on each target video frame based on the updated quantization metric data corresponding to multiple forward prediction coded frames within each group of video frame images, the first quantization metric data corresponding to the non-updated forward video frames within each group of video frame images, and the second quantization metric data corresponding to each of the multiple bidirectional prediction coded frames within each group of video frame images, to obtain the rate-distortion metric data corresponding to each target video frame at each preset encoding dimension;
[0056] An encoding dimension determination module, configured to use the preset encoding dimension with the smallest rate-distortion metric data corresponding to each target video frame among the rate-distortion metric data corresponding to each target video frame at multiple preset encoding dimensions as the target encoding dimension corresponding to each target video frame.
[0057] According to another aspect of the embodiments of the present disclosure, there is provided an electronic device, including: a processor; a memory for storing instructions executable by the processor; wherein, the processor is configured to execute the instructions to implement the above video encoding method.
[0058] According to another aspect of the embodiments of the present disclosure, there is provided a computer-readable storage medium, when the instructions in the storage medium are executed by a processor of an electronic device, enabling the electronic device to execute the above video encoding method.
[0059] According to another aspect of the embodiments of the present disclosure, there is provided a computer program product containing instructions, when it runs on a computer, enabling the computer to execute the above video encoding method.
[0060] The technical solutions provided by the embodiments of the present disclosure at least bring the following beneficial effects:
[0061] By obtaining multiple target video frames of the video to be encoded and multiple groups of video frame images corresponding to the video to be encoded, and performing quantization analysis on each forward prediction encoded frame to obtain the first quantization index data of each forward prediction encoded frame, the quantization analysis of each forward prediction encoded frame can be achieved. Then, from each group of video frame images, the current forward encoded frame is sequentially determined. The current forward encoded frame is the forward prediction encoded frame with the earliest time sequence among the forward prediction encoded frames that are not affected by the distortion of the already analyzed encoded frames in each group of video frame images. The already analyzed encoded frames are the forward prediction encoded frames for which the distortion influence range analysis has been performed. The distortion influence range analysis is performed on the current forward encoded frame to obtain the influence range index data corresponding to the current forward encoded frame. The influence range index data corresponding to the current forward encoded frame represents the range of video frames affected by the distortion of the current forward encoded frame during the encoding process of the video to be encoded, and the accurate determination of the distortion influence range of the current forward encoded frame on subsequent forward prediction encoded frames can be achieved. Next, in combination with the influence range index data corresponding to the current forward encoded frame, the first quantization index data of the forward prediction encoded frames within the target sub-image group corresponding to the current forward encoded frame is updated to obtain the updated quantization index data corresponding to the current forward encoded frame. The target sub-image group corresponding to the current forward encoded frame is the sub-image group belonging to the influence range indicated by the influence range index data corresponding to the current forward encoded frame, and the adaptive adjustment of the first quantization index data of the forward prediction encoded frames in the video to be encoded can be achieved. Furthermore, the encoding effect of the video frames in simple scenes can be effectively improved. Then, in combination with the updated quantization index data corresponding to multiple forward prediction encoded frames within each group of video frame images and the first quantization index data corresponding to the unupdated forward video frames within each group of video frame images, the video to be encoded is encoded to obtain the target video encoding data, and the effective improvement of the encoding effect of the video frames in simple scenes in the video can be achieved.
[0062] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and do not limit the present disclosure. Brief Description of the Drawings
[0063] The accompanying drawings herein are incorporated into the specification and constitute a part of this specification, showing embodiments consistent with the present disclosure, and are used together with the specification to explain the principles of the present disclosure and do not constitute an improper limitation of the present disclosure.
[0064] Figure 1 is a schematic diagram of an application system shown according to an exemplary embodiment;
[0065] Figure 2 is a flowchart of a video encoding method shown according to an exemplary embodiment;
[0066] Figure 3It is a block diagram of a video encoding device shown according to an exemplary embodiment;
[0067] Figure 4 It is a block diagram of an electronic device for encoding multiple target video frames in a video to be encoded shown according to an exemplary embodiment;
[0068] Figure 5 It is a block diagram of another electronic device for encoding multiple target video frames in a video to be encoded shown according to an exemplary embodiment. Detailed implementation manners
[0069] Various exemplary embodiments, features, and aspects of the present application will be described in detail below with reference to the accompanying drawings. The same reference numerals in the drawings denote elements having the same or similar functions. Although various aspects of the embodiments are shown in the drawings, the drawings do not have to be drawn to scale unless otherwise specified.
[0070] The term "exemplary" used herein means "serving as an example, embodiment, or illustration". Any embodiment described as "exemplary" herein does not have to be construed as superior or better than other embodiments.
[0071] In addition, for a better description of the present application, numerous specific details are given in the following detailed implementation manners. Those skilled in the art should understand that the present application can also be implemented without some specific details. In some instances, methods, means, elements, and circuits well known to those skilled in the art are not described in detail so as to highlight the gist of the present application.
[0072] Artificial Intelligence (AI) is a theory, method, technology, and application system that uses a digital computer or a machine controlled by a digital computer to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science that attempts to understand the essence of intelligence and produce a new intelligent machine that can react in a way similar to human intelligence. Artificial intelligence also studies the design principles and implementation methods of various intelligent machines to enable the machines to have the functions of perception, reasoning, and decision-making.
[0073] Artificial intelligence technology is an interdisciplinary subject with a wide range of fields involved, including both hardware-level and software-level technologies. The basic technologies of artificial intelligence generally include sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, pre-trained model technology, operation / interaction systems, mechatronics, etc. Among them, the pre-trained model, also known as the large model or the foundation model, can be widely applied to downstream tasks in various directions of artificial intelligence after fine-tuning. Artificial intelligence software technology mainly includes several major directions such as computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning.
[0074] Computer Vision Technology (CV) Computer vision is a science that studies how to enable machines to "see". More specifically, it refers to using cameras and computers to replace human eyes for tasks such as object recognition and measurement in machine vision, and further performing graphic processing to make the computer-processed images more suitable for human eyes to observe or be transmitted to instruments for detection. As a scientific discipline, computer vision studies related theories and technologies, and attempts to build artificial intelligence systems that can obtain information from images or multi-dimensional data. Computer vision technology usually includes technologies such as image processing, image recognition, image semantic understanding, image retrieval, video processing, video semantic understanding, video content / behavior recognition, three-dimensional object reconstruction, 3D technology, virtual reality, augmented reality, simultaneous localization and mapping, etc., and also includes common biometric recognition technologies such as face recognition and fingerprint recognition.
[0075] Cloud technology refers to a hosting technology that unifies a series of resources such as hardware, software, and networks within a wide area network or a local area network to achieve data computing, storage, processing, and sharing.
[0076] Cloud technology is the general term for network technology, information technology, integration technology, management platform technology, application technology, etc. based on the cloud computing business model. It can form a resource pool, be used on demand, and is flexible and convenient. Cloud computing technology will become an important support. The back-end services of technical network systems require a large amount of computing and storage resources, such as video websites, picture websites, and more portal websites. With the highly developed application of the Internet industry, in the future, each item may have its own identification mark and needs to be transmitted to the back-end system for logical processing. Data at different levels will be processed separately, and various industry data requires a powerful system back-end support, which can only be achieved through cloud computing.
[0077] Cloud computing is a computing model that distributes computing tasks across a resource pool composed of a large number of computing devices, enabling various application systems to obtain computing power, storage space, and information services as needed. The network that provides resources is called the "cloud". The resources in the "cloud" seem to users to be infinitely scalable, and can be obtained at any time, used on demand, expanded at any time, and paid according to usage.
[0078] As a basic capability provider of cloud computing, a cloud computing resource pool (abbreviated as cloud platform, generally called IaaS (Infrastructure as a Service) platform) will be established, and various types of virtual resources will be deployed in the resource pool for external customers to select and use. The cloud computing resource pool mainly includes: computing devices (virtual machines containing operating systems), storage devices, and network devices.
[0079] According to the logical function division, the PaaS (Platform as a Service) layer can be deployed on the IaaS (Infrastructure as a Service) layer, and the SaaS (Software as a Service) layer can be deployed on top of the PaaS layer. The SaaS layer can also be directly deployed on the IaaS. PaaS is a platform for software operation, such as databases, web containers, etc. SaaS is various business software, such as web portals, SMS mass senders, etc. Generally speaking, SaaS and PaaS are upper layers relative to IaaS.
[0080] In recent years, with the research and progress of artificial intelligence technology, artificial intelligence technology has been widely used in multiple fields. The solution provided in the embodiments of this application involves technologies such as computer vision, and will be specifically described through the following embodiments:
[0081] Please refer to Figure 1 , Figure 1 which is a schematic diagram of an application system shown according to an exemplary embodiment. The application system can be used for the video encoding method of this application. As Figure 1 shown, the application system can at least include a server 01 and a terminal 02.
[0082] In the embodiments of this application, the server 01 can be used to perform encoding processing on the video to be encoded. Specifically, the above-mentioned server 01 can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, as well as big data and artificial intelligence platforms.
[0083] In the embodiments of the present application, the terminal 02 can be used to generate a video to be encoded. The above-mentioned terminal 02 can include physical devices such as smart phones, desktop computers, tablet computers, laptop computers, smart speakers, vehicle-mounted terminals, digital assistants, augmented reality (AR) / virtual reality (VR) devices, and smart wearable devices, or can also include software running on the physical devices, such as application programs, etc. The operating systems running on the above-mentioned terminal 02 in the embodiments of the present application can include but are not limited to Android system, IOS system, Linux, Windows, etc.
[0084] In addition, it should be noted that Figure 1 The application environment shown is only one provided by the present disclosure. In actual applications, other application environments may also be included. For example, the encoding process for the video to be encoded can also be implemented on the terminal 02.
[0085] In the embodiments of this specification, the above-mentioned terminal 02 and the server 01 can be directly or indirectly connected through wired or wireless communication methods, and the present application does not make any limitations in this regard.
[0086] It should be noted that the step sequence shown in the following figures is a possible one, and in fact, it is not necessary to strictly follow this sequence. Some steps can be executed in parallel without depending on each other.
[0087] Specifically, Figure 2 is a flowchart of a video encoding method shown according to an exemplary embodiment. As Figure 2 shown, this video encoding method can be used in electronic devices such as terminals or servers, and specifically can include the following steps:
[0088] S201: Obtain multiple target video frames of the video to be encoded, and multiple groups of video frame images corresponding to the video to be encoded.
[0089] In a specific embodiment, the video to be encoded can refer to a video that needs to be encoded. The video to be encoded can include multiple target video frames.
[0090] In a specific embodiment, any group of video frame images corresponding to the video to be encoded may refer to a group of video frames with continuous pictures. The multiple groups of video frame images may be obtained by performing group-of-pictures (GOP) partitioning on multiple target video frames in the video to be encoded. Specifically, based on a preset GOP length, the multiple target video frames in the video to be encoded may be grouped to obtain multiple groups of video frame images corresponding to the video to be encoded; wherein, the length of each group of video frame images may be equal to the above-mentioned preset GOP length. The preset GOP length may be set according to actual application needs, and the present disclosure does not make any limitation.
[0091] In a specific embodiment, each group of video frame images may include multiple sub-GOPs. Each sub-GOP in at least one sub-GOP may include a forward-predicted coded frame. Further, the multiple sub-GOPs of each group of video frame images may include a first sub-GOP and at least one second sub-GOP (i.e., the above-mentioned at least one sub-GOP). Among them, the first sub-GOP may refer to a sub-GOP that contains one intra-coded frame and multiple bi-predicted interpolated coded frames. The second sub-GOP may refer to a sub-GOP that contains at least one forward-predicted coded frame. Specifically, the second sub-GOP may include one bi-predicted interpolated coded frame, and may also include at least one bi-predicted interpolated coded frame.
[0092] S203: Perform quantization analysis on each forward-predicted coded frame to obtain first quantization index data of each forward-predicted coded frame.
[0093] In a specific embodiment, the first quantization index data of any forward-predicted coded frame may be used to indicate the spatial detail compression condition of the video frame. Further, the first quantization index data of any forward-predicted coded frame may be used to indicate the spatial detail compression condition of the above-mentioned any forward-predicted coded frame. The first quantization index data of any forward-predicted coded frame may include the quantization parameter of the above-mentioned any forward-predicted coded frame.
[0094] In a specific embodiment, the complexity of the current video frame corresponding to any forward prediction coding frame can be determined first; then, based on the complexity of the current video frame corresponding to any of the above forward prediction coding frames, quality analysis can be performed on any of the above forward prediction coding frames to obtain the video output quality data corresponding to any of the above forward prediction coding frames, where the video output quality data corresponding to any of the above forward prediction coding frames can characterize the output quality of any of the above forward prediction coding frames; then, based on the video output quality data corresponding to any of the above forward prediction coding frames, quantization prediction can be performed on any of the above forward prediction coding frames to obtain the first quantization index data of any of the above forward prediction coding frames. Specifically, the SATD (Sum of Absolute Transformed Difference) corresponding to any forward prediction coding frame and the SATD corresponding to the historical forward coding frame corresponding to any of the above forward prediction coding frames can be obtained first; based on the SATD corresponding to any of the above forward prediction coding frames and the SATD corresponding to the historical forward coding frame corresponding to any of the above forward prediction coding frames, the cumulative historical frame complexity corresponding to any of the above forward prediction coding frames can be obtained; based on the cumulative historical frame complexity corresponding to any of the above forward prediction coding frames, the complexity of the current video frame corresponding to any of the above forward prediction coding frames can be determined. Wherein, the historical forward coding frame corresponding to any of the above forward prediction coding frames can refer to the forward prediction coding frame before any of the above forward prediction coding frames in time sequence. The cumulative historical frame complexity corresponding to any of the above forward prediction coding frames can refer to the cumulative complexity of the past video frames corresponding to any of the above forward prediction coding frames.
[0095] Further, the cumulative historical frame complexity corresponding to any forward prediction coding frame can be obtained through the following formula:
[0096] cplx_sum(i) = 0.5 * cplx_sum(i - 1) + SATD(i)
[0097] Wherein, cplx_sum(i) is the cumulative historical frame complexity corresponding to the i-th forward prediction coding frame in the video to be encoded; cplx_sum(i - 1) is the cumulative historical frame complexity corresponding to the (i - 1)-th forward prediction coding frame in the video to be encoded; SATD(i) is the SATD corresponding to the i-th forward prediction coding frame in the video to be encoded.
[0098] Correspondingly, the complexity of the current video frame corresponding to any forward prediction coding frame can be obtained through the following formula:
[0099] cplx_blur(i) = cplx_sum(i) / cplx_count
[0100] Among them, cplx_blur(i) is the current video frame complexity corresponding to the i-th forward prediction encoded frame in the video to be encoded; cplx_sum(i) is the cumulative historical frame complexity corresponding to the i-th forward prediction encoded frame in the video to be encoded; cplx_count is the number of video frames of the historical forward encoded frames corresponding to the i-th forward prediction encoded frame.
[0101] Next, the video output quality data corresponding to any one of the above forward prediction encoded frames can be obtained through the following formula:
[0102] qscale(i) = cplx_blur(i) 1-qc
[0103] Among them, qscale(i) is the video output quality data corresponding to the i-th forward prediction encoded frame in the video to be encoded; cplx_blur(i) is the current video frame complexity corresponding to the i-th forward prediction encoded frame in the video to be encoded; qc is a preset compression factor. Specifically, qc can be set according to actual application needs, and the present disclosure does not make a limitation. Exemplarily, in the X265 encoder, qc can be 0.6.
[0104] Then, the first quantization index data of any forward prediction encoded frame can be obtained through the following formula:
[0105] QP(i) = 12.0 + 6.0 * (double)log2(qscale(i) / 0.85)
[0106] Among them, QP(i) is the first quantization index data of the i-th forward prediction encoded frame in the video to be encoded; qscale(i) is the video output quality data corresponding to the i-th forward prediction encoded frame in the video to be encoded.
[0107] S205: Sequentially determine the current forward encoded frame from each group of video frame images.
[0108] In a specific embodiment, the current forward encoded frame in each group of video frame images may refer to the forward prediction encoded frame with the earliest time sequence among the forward prediction encoded frames not affected by the distortion of the analyzed encoded frames in each group of video frame images. Among them, the analyzed encoded frames may refer to the forward prediction encoded frames for which the distortion influence range analysis has been performed.
[0109] In a specific embodiment, for any group of video frame images, before the first analysis of the distortion influence range, that is, when the distortion influence range analysis has not been performed, the forward prediction coding frame with the earliest time sequence in any of the above-mentioned groups of video frame images can be used as the current forward coding frame of any of the above-mentioned groups of video frame images; it can be understood that since before the first analysis of the distortion influence range, that is, when the distortion influence range analysis has not been performed, the analyzed coding frames are empty, then multiple forward prediction coding frames in any of the above-mentioned groups of video frame images can belong to the forward prediction coding frames that are not affected by the distortion of the analyzed coding frames.
[0110] In a specific embodiment, for any group of video frame images, when at least one distortion influence range analysis has been performed, the current analyzed coding frame can be determined first, and then based on the current analyzed coding frame and the influence range index data corresponding to the coding frame analyzed most recently in the current analyzed coding frame, the current candidate coding frame can be determined. Correspondingly, the forward prediction coding frame with the earliest time sequence in the current candidate coding frame can be used as the current forward coding frame. Specifically, based on the influence range index data corresponding to the coding frame analyzed most recently in the current analyzed coding frame, at least one distortion influence coding frame corresponding to the coding frame analyzed most recently can be determined, where at least one distortion influence coding frame corresponding to the coding frame analyzed most recently can refer to the forward prediction coding frames affected by the distortion of the coding frame analyzed most recently; then, the forward prediction coding frame with the latest time sequence can be selected from the at least one distortion influence coding frame as the end distortion influence coding frame; then, the forward prediction coding frames in any of the above-mentioned groups of video frame images that are after the end distortion influence coding frame can be used as the current candidate coding frame; it can be understood that the current candidate coding frame does not include the end distortion influence coding frame.
[0111] Exemplarily, assuming that multiple forward prediction coding frames in any group of video frame images are arranged in time sequence as forward prediction coding frame P1, forward prediction coding frame P2, forward prediction coding frame P3, forward prediction coding frame P4, forward prediction coding frame P5, and forward prediction coding frame P6, when the analyzed coding frame is forward prediction coding frame P1 and the influence range index data corresponding to the forward prediction coding frame P1 is 2, it can be determined that the current candidate coding frames include forward prediction coding frame P3, forward prediction coding frame P4, forward prediction coding frame P5, and forward prediction coding frame P6. Correspondingly, it can be determined that the current forward coding frame is forward prediction coding frame P3.
[0112] S207: Perform a distortion influence range analysis on the current forward coding frame to obtain the influence range index data corresponding to the current forward coding frame.
[0113] In a specific embodiment, the influence range index data corresponding to the current forward-coded frame can characterize the range of video frames affected by the distortion of the current forward-coded frame during the encoding process of the video to be encoded. Exemplarily, assuming that the influence range index data corresponding to the current forward-coded frame (if it is the forward prediction coded frame P3) is 2, and in any video frame group, multiple forward prediction coded frames are arranged in sequence as forward prediction coded frame P1, forward prediction coded frame P2, forward prediction coded frame P3, forward prediction coded frame P4, forward prediction coded frame P5, and forward prediction coded frame P6, then it can be determined that the sub-image group corresponding to the above forward prediction coded frame P3 and the sub-image group corresponding to the forward prediction coded frame P4 belong to the range of video frames affected by the distortion of the above current forward-coded frame.
[0114] In a specific embodiment, before performing quantization analysis on each forward prediction coded frame to obtain the first quantization index data of each forward prediction coded frame, the method may further include:
[0115] Performing inter-frame complexity analysis on each forward prediction coded frame based on each forward prediction coded frame and the previous video frame of each forward prediction coded frame to obtain the inter-frame complexity index data corresponding to each forward prediction coded frame;
[0116] Correspondingly, performing distortion influence range analysis on the current forward-coded frame to obtain the influence range index data corresponding to the current forward-coded frame may include:
[0117] Performing distortion influence analysis on the current forward-coded frame based on the inter-frame complexity index data corresponding to the current forward-coded frame to obtain the influence range index data corresponding to the current forward-coded frame.
[0118] In a specific embodiment, the previous video frame of any forward prediction coded frame may refer to the previous forward prediction coded frame of the above any forward prediction coded frame. It can be understood that the previous video frame of the above any forward prediction coded frame may be a forward prediction coded frame and is before the above any forward prediction coded frame in terms of time sequence.
[0119] In a specific embodiment, the inter-frame complexity index data corresponding to each forward prediction coded frame can characterize the degree of difference between each forward prediction coded frame and the previous video frame of each forward prediction coded frame.
[0120] In a specific embodiment, the inter-frame complexity index data corresponding to each forward prediction coding frame can be obtained by performing motion analysis on each forward prediction coding frame and the previous video frame of each forward prediction coding frame. Specifically, based on a preset coding block size, any forward prediction coding frame and the previous video frame of any forward prediction coding frame can be segmented respectively to obtain a plurality of first coding blocks corresponding to any forward prediction coding frame, and a plurality of second coding blocks corresponding to the previous video frame of any forward prediction coding frame; similarity matching is performed on each first coding block and the plurality of second coding blocks to obtain a similar coding block corresponding to each first coding block, where the similar coding block corresponding to any first coding block may refer to the second coding block that is most similar to any first coding block among the plurality of second coding blocks; residual analysis is performed on each first coding block and the similar coding block corresponding to each first coding block to obtain the complexity index data corresponding to each first coding block; and then the complexity index data corresponding to each of the plurality of first coding blocks are superimposed to obtain the inter-frame complexity index data corresponding to any forward prediction coding frame. Further, by performing residual analysis on any first coding block and the plurality of second coding blocks, the second coding block with the smallest residual result in the residual analysis result corresponding to any first coding block can be used as the similar coding block corresponding to any first coding block.
[0121] In a specific embodiment, the influence range index data corresponding to the current forward coding frame can be obtained by the following formula:
[0122] g_size = floor(clip(1, 4, w * (inter_complex) -1 ))
[0123] where g_size is the influence range index data corresponding to the current forward coding frame; inter_complex is the inter-frame complexity index data corresponding to the current forward coding frame; w is a preset index parameter; floor() is a preset floor function; clip() is a preset clipping function.
[0124] In a specific embodiment, the above-mentioned distortion influence analysis of the current forward coding frame based on the inter-frame complexity index data corresponding to the current forward coding frame to obtain the influence range index data corresponding to the current forward coding frame may include:
[0125] Based on the inter-frame complexity index data corresponding to the current forward coding frame, inter-frame correlation analysis is performed on the current forward coding frame to obtain the inter-frame correlation index data corresponding to the current forward coding frame;
[0126] Matching processing is performed on the preset index value range and the inter-frame correlation index data corresponding to the current forward coding frame to obtain a target matching result;
[0127] Based on the target matching result, determine the influence range index data corresponding to the current forward coding frame from the two endpoint index data corresponding to the preset index value range and the inter-frame correlation index data.
[0128] In a specific embodiment, the inter-frame correlation index data corresponding to the current forward coding frame can represent the degree of correlation between the current forward coding frame and the previous video frame of the current forward coding frame.
[0129] In a specific embodiment, the inter-frame correlation index data corresponding to the current forward coding frame can be obtained through the following formula:
[0130] r = w * (inter_complex) -1
[0131] where r is the inter-frame correlation index data corresponding to the current forward coding frame; inter_complex is the inter-frame complexity index data corresponding to the current forward coding frame; w is a preset index parameter. Specifically, the above preset index parameter can be set according to actual application needs; optionally, the above preset index parameter can be 0.8.
[0132] In a specific embodiment, the target matching result can be used to indicate whether the inter-frame correlation index data belongs to the preset index value range. The target matching result can include a first matching result or a second matching result. Among them, the first matching result can be used to indicate that the inter-frame correlation index data belongs to the preset index value range. The second matching result can be used to indicate that the inter-frame correlation index data does not belong to the preset index value range.
[0133] In a specific embodiment, the preset index value range can be used to indicate the value range of the influence range index data. The two endpoint index data corresponding to the preset index value range can represent the numerical values of the two endpoints of the preset index value range. Specifically, the preset index value range can be set according to actual application needs. Optionally, the preset index value range can include [1, 4] or [1, 3], etc. Exemplarily, when the preset index value range is [1, 4], the two endpoint index data corresponding to the above preset index value range can be 1 and 4 respectively.
[0134] In a specific embodiment, the influence range index data corresponding to the current forward coding frame can include at least one of the first range index data and the second range index data.
[0135] In a specific embodiment, when the influence range index data corresponding to the current forward coding frame includes first range index data, determining the influence range index data corresponding to the current forward coding frame from the two endpoint index data and the inter-frame correlation index data corresponding to the preset index value range based on the target matching result may include:
[0136] When the target matching result indicates that the inter-frame correlation index data belongs to the preset index value range, round down the inter-frame correlation index data to obtain the first range index data corresponding to the current forward coding frame.
[0137] In a specific embodiment, when the target matching result is the first matching result, rounding down the inter-frame correlation index data can obtain the first range index data corresponding to the current forward coding frame; correspondingly, based on the first range index data corresponding to the current forward coding frame, update the first quantization index data of the forward prediction coding frame in the target sub-image group corresponding to the current forward coding frame to obtain the updated quantization index data corresponding to the current forward coding frame.
[0138] In a specific embodiment, when the influence range index data corresponding to the current forward coding frame includes second range index data, the method may further include:
[0139] When the target matching result indicates that the inter-frame correlation index data does not belong to the preset index value range, determine the adjacent endpoint index data from the two endpoint index data corresponding to the preset index value range;
[0140] Use the adjacent endpoint index data as the second range index data corresponding to the current forward coding frame.
[0141] In a specific embodiment, the adjacent endpoint index data may be the index data closest to the inter-frame correlation index data among the two endpoint index data.
[0142] In a specific embodiment, when the target matching result is the second matching result, the endpoint index data closest to the inter-frame correlation index data among the two endpoint index data corresponding to the preset index value range may be used as the adjacent endpoint index data. Exemplarily, when the preset index value range is [1, 4] and the inter-frame correlation index data is 5, the adjacent endpoint index data can be determined to be 4.
[0143] In a specific embodiment, the adjacent endpoint metric data can be used as the second range metric data corresponding to the current forward-coded frame. Correspondingly, based on the second range metric data corresponding to the current forward-coded frame, the first quantization metric data of the forward-predicted coded frames within the target sub-image group corresponding to the current forward-coded frame can be updated to obtain the updated quantization metric data corresponding to the current forward-coded frame.
[0144] In the above embodiment, by performing an inter-frame correlation analysis on the current forward-coded frame based on the inter-frame complexity metric data corresponding to the current forward-coded frame, the inter-frame correlation metric data corresponding to the current forward-coded frame is obtained. A matching process is performed on the preset metric value range and the inter-frame correlation metric data corresponding to the current forward-coded frame to obtain a target matching result. Based on the target matching result, the influence range metric data corresponding to the current forward-coded frame is determined from the two endpoint metric data corresponding to the preset metric value range and the inter-frame correlation metric data, so as to accurately predict the distortion influence range of the current forward-coded frame, facilitating the accurate optimization of the quantization metric data through the above influence range metric data.
[0145] S209: Based on the influence range metric data corresponding to the current forward-coded frame, the first quantization metric data of the forward-predicted coded frames within the target sub-image group corresponding to the current forward-coded frame is updated to obtain the updated quantization metric data corresponding to the current forward-coded frame.
[0146] In a specific embodiment, the target sub-image group corresponding to the current forward-coded frame may refer to the sub-image group within the influence range indicated by the influence range metric data corresponding to the current forward-coded frame. Specifically, the target sub-image group may include at least one second sub-image group.
[0147] In a specific embodiment, the updated quantization metric data corresponding to the current forward-coded frame may refer to the updated first quantization metric data of the updated coded frame corresponding to the current forward-coded frame. Among them, the coded frame to be updated corresponding to the current forward-coded frame may refer to the forward-predicted coded frames that need to be updated within the target sub-image group corresponding to the current forward-coded frame. Specifically, the coded frames to be updated may include at least one forward-predicted coded frame within the target sub-image group corresponding to the current forward-coded frame.
[0148] In a specific embodiment, the above step S209 may include:
[0149] Based on the influence range metric data corresponding to the current forward-coded frame, the target sub-image group corresponding to the current forward-coded frame is merged to obtain the merged sub-image group corresponding to the current forward-coded frame;
[0150] Determine the position indication information corresponding to each of the multiple first predicted coding frames corresponding to the current forward coding frame;
[0151] Based on the influence range metric data corresponding to the current forward coding frame and the position indication information corresponding to each of the multiple first predicted coding frames, update the first quantization metric data corresponding to at least one predicted coding frame to be updated among the multiple first predicted coding frames corresponding to the current forward coding frame, to obtain the updated quantization metric data corresponding to the current forward coding frame.
[0152] In a specific embodiment, the merged sub-image group may refer to the merged sub-image group corresponding to the current forward coding frame. The merged sub-image group may include multiple forward predicted coding frames.
[0153] In a specific embodiment, the multiple first predicted coding frames corresponding to the current forward coding frame may refer to the multiple forward predicted coding frames included in the merged sub-image group corresponding to the current forward coding frame. The position indication information corresponding to each first predicted coding frame may be used to indicate the position of each first predicted coding frame among the multiple first predicted coding frames. Exemplarily, assuming that the multiple first predicted coding frames are arranged in time sequence as forward predicted coding frame P1, forward predicted coding frame P2, forward predicted coding frame P3, forward predicted coding frame P4, forward predicted coding frame P5, and forward predicted coding frame P6, when the first predicted coding frame is "forward predicted coding frame P3", it can be determined that the position indication information corresponding to the above first predicted coding frame may be 3.
[0154] In a specific embodiment, a video frame sequence corresponding to the multiple first predicted coding frames may be generated in the order from the earliest to the latest in time sequence according to the time sequence relationship between the multiple first predicted coding frames; then, based on the position of any first predicted coding frame in the above video frame sequence, the position indication information corresponding to the above any first predicted coding frame may be generated.
[0155] In a specific embodiment, the updated quantization metric data corresponding to the current forward coding frame may include the first updated metric data corresponding to each predicted coding frame to be updated. Wherein, any predicted coding frame to be updated may refer to any one of the first predicted coding frames in the coding frames to be updated corresponding to the current forward coding frame. The first updated metric data corresponding to any predicted coding frame to be updated may refer to the updated first quantization metric data of the above any predicted coding frame to be updated. Specifically, at least one predicted coding frame to be updated may be the preset updated number of first predicted coding frames with earlier time sequence among the multiple first predicted coding frames corresponding to the current forward coding frame. Further, the preset updated number may be set according to actual application requirements. Exemplarily, the preset updated number may include 1, 2, or 3, etc.
[0156] In a specific embodiment, when at least one predicted coding frame to be updated corresponding to the current forward coding frame is a plurality of first predicted coding frames corresponding to the current forward coding frame, updating the first quantization index data corresponding to at least one predicted coding frame to be updated among the plurality of first predicted coding frames corresponding to the current forward coding frame based on the influence range index data corresponding to the current forward coding frame and the position indication information corresponding to each first predicted coding frame to obtain the updated quantization index data corresponding to the current forward coding frame may include:
[0157] Determine the first weight information corresponding to each predicted coding frame to be updated based on the position indication information corresponding to each predicted coding frame to be updated and the preset mapping relationship;
[0158] Perform quantization increment analysis on each predicted coding frame to be updated based on the first weight information corresponding to each predicted coding frame to be updated and the influence range index data corresponding to the current forward coding frame to obtain the first increment index data corresponding to each predicted coding frame to be updated;
[0159] Update the first quantization index data corresponding to each predicted coding frame to be updated based on the first increment index data corresponding to each predicted coding frame to be updated to obtain the first updated index data corresponding to each predicted coding frame to be updated.
[0160] In a specific embodiment, the preset mapping relationship may be a mapping relationship between a plurality of position indication information and a plurality of preset weight information. Specifically, in the above preset mapping relationship, the closer the position of the corresponding forward predicted coding frame indicated by the position indication information is to the front, the larger the absolute value of the preset weight information corresponding to the position indication information may be. Further, the preset mapping relationship can be set according to actual application needs. Exemplarily, when the first predicted coding frame P1 is the first first predicted coding frame among the plurality of first predicted coding frames (i.e., the corresponding position indication information is "1"), the preset weight information corresponding to the first predicted coding frame P1 may be 0.5; when the first predicted coding frame P2 is the second first predicted coding frame among the plurality of first predicted coding frames (i.e., the corresponding position indication information is "2"), the preset weight information corresponding to the first predicted coding frame P2 may be -0.2; when the first predicted coding frame P3 is the third first predicted coding frame among the plurality of first predicted coding frames (i.e., the corresponding position indication information is "3"), the preset weight information corresponding to the first predicted coding frame P3 may be 0; correspondingly, when the position indication information is greater than 2, the corresponding preset weight information may be 0.
[0161] In a specific embodiment, the first weight information corresponding to any predicted coding frame to be updated can be used to characterize the required update degree of the first quantization index data of any predicted coding frame to be updated.
[0162] In a specific embodiment, in the above preset mapping relationship, the weight information corresponding to the position indication information of any to-be-updated prediction coding frame may be searched, and the corresponding weight information may be used as the first weight information corresponding to any to-be-updated prediction coding frame.
[0163] In a specific embodiment, the first incremental index data corresponding to any to-be-updated prediction coding frame may represent the change amount of the first quantization index data corresponding to any to-be-updated prediction coding frame.
[0164] In a specific embodiment, the first incremental index data corresponding to any to-be-updated prediction coding frame may be obtained by the following formula:
[0165] delta_qp = weight_i * g_size
[0166] Wherein, delta_qp is the first incremental index data corresponding to any to-be-updated prediction coding frame (i.e., the i-th to-be-updated prediction coding frame); weight_i is the first weight information corresponding to the i-th to-be-updated prediction coding frame; g_size is the influence range index data corresponding to the current forward coding frame. It can be understood that when the positions of the to-be-updated prediction coding frames are the same, the larger the influence range index data corresponding to the current forward coding frame, the greater the adjustment required for the first quantization index data of the to-be-updated prediction coding frame.
[0167] In a specific embodiment, the first update index data corresponding to any to-be-updated prediction coding frame may be obtained by the following formula:
[0168] QP_new_i = QP - delta_qp
[0169] Wherein, QP_new_i is the first update index data corresponding to the i-th to-be-updated prediction coding frame; QP is the first incremental index data corresponding to the i-th to-be-updated prediction coding frame; delta_qp is the first incremental index data corresponding to the i-th to-be-updated prediction coding frame.
[0170] In a specific embodiment, the updated quantization index data corresponding to the current forward coding frame may include the second update index data corresponding to the second prediction coding frame. Wherein, the second prediction coding frame may be the first prediction coding frame with the earliest time sequence among the multiple first prediction coding frames corresponding to the current forward coding frame. The second update index data may refer to the updated first quantization index data of the second prediction coding frame.
[0171] In a specific embodiment, the influence range index data corresponding to the current forward coding frame, the position indication information corresponding to any to-be-updated prediction coding frame, and the first quantization index data corresponding to any to-be-updated prediction coding frame may be input into a preset quantization update model for quantization update processing to obtain the updated quantization index data corresponding to any to-be-updated prediction coding frame. Among them, the preset quantization update model may be used to perform quantization update on the first quantization index data corresponding to any to-be-updated prediction coding frame. Specifically, the sample position indication information corresponding to the sample coding frame, the sample quantization index data corresponding to the sample coding frame, and the influence range index data corresponding to the sample coding frame may be obtained from the sample dataset, and the sample position indication information corresponding to the sample coding frame, the sample quantization index data corresponding to the sample coding frame, and the influence range index data corresponding to the sample coding frame are input into a preset machine learning model for quantization update processing to obtain the updated quantization index data corresponding to the sample coding frame; based on the updated quantization index data corresponding to the sample coding frame, video frame coding processing is performed on the sample coding frame to obtain the sample coding data corresponding to the sample coding frame; based on the sample coding data corresponding to the sample coding frame, target loss information may be determined, where the target loss information may characterize the compression quality of the sample coding data corresponding to the sample coding frame; the preset machine learning model may be trained based on the target loss information to obtain the preset quantization update model.
[0172] In a specific embodiment, when at least one to-be-updated prediction coding frame corresponding to the current forward coding frame is a second prediction coding frame, the updating the first quantization index data corresponding to at least one to-be-updated prediction coding frame among the multiple first prediction coding frames corresponding to the current forward coding frame based on the influence range index data corresponding to the current forward coding frame and the position indication information corresponding to each first prediction coding frame to obtain the updated quantization index data corresponding to the current forward coding frame may include:
[0173] Determine the second weight information corresponding to the second prediction coding frame based on the position indication information corresponding to the second prediction coding frame and the preset mapping relationship;
[0174] Perform quantization increment analysis on the second prediction coding frame based on the second weight information corresponding to the second prediction coding frame and the influence range index data corresponding to the current forward coding frame to obtain the second increment index data corresponding to the second prediction coding frame;
[0175] Update the first quantization index data corresponding to the second prediction coding frame among the multiple first prediction coding frames based on the second increment index data to obtain the second updated index data corresponding to the second prediction coding frame.
[0176] In a specific embodiment, the second weight information may characterize the required update degree of the first quantization index data corresponding to the second predictive coding frame.
[0177] In a specific embodiment, in the above preset mapping relationship, the weight information corresponding to the position indication information of the second predictive coding frame may be searched for, and the weight information corresponding to the second predictive coding frame may be used as the second weight information.
[0178] In a specific embodiment, the second incremental index data may characterize the change amount of the first quantization index data corresponding to the second predictive coding frame. Specifically, the determination process of the second incremental index data may refer to the determination process of the first incremental index data, and details are not described herein again in the present disclosure.
[0179] In a specific embodiment, the determination process of the second update index data may refer to the determination process of the first update index data, and details are not described herein again in the present disclosure.
[0180] In a specific embodiment, updating the first quantization index data of the forward predictive coding frames within the target sub-image group corresponding to the current forward coding frame based on the influence range index data corresponding to the current forward coding frame to obtain the updated quantization index data corresponding to the current forward coding frame may further include:
[0181] Determining the number of frames to be updated corresponding to the current forward coding frame and the number of influencing frames corresponding to the influence range index data;
[0182] Performing a comparison process on the number of influencing frames and the number of frames to be updated to obtain a target comparison result;
[0183] Correspondingly, merging the target sub-image group corresponding to the current forward coding frame based on the influence range index data corresponding to the current forward coding frame to obtain the merged sub-image group corresponding to the current forward coding frame may include:
[0184] In the case where the target comparison result indicates that the number of influencing frames is less than or equal to the number of frames to be updated, merging the target sub-image group corresponding to the current forward coding frame based on the influence range index data corresponding to the current forward coding frame to obtain the merged sub-image group corresponding to the current forward coding frame.
[0185] In a specific embodiment, the number of frames to be updated corresponding to each video frame image group may refer to the number of forward prediction coding frames in each video frame image group that are currently not affected by the distortion of the analyzed and encoded frames. Specifically, the current forward coding frame, the last forward prediction coding frame in the corresponding video frame image group, and the intermediate forward coding frame corresponding to the current forward coding frame can be obtained, and based on the number of video frames of the above-mentioned current forward coding frame, the last forward prediction coding frame in the corresponding video frame image group, and the intermediate forward coding frame corresponding to the current forward coding frame, the number of frames to be updated can be obtained; wherein, the intermediate forward coding frame corresponding to the current forward coding frame may be a forward prediction coding frame between the current forward coding frame and the last forward prediction coding frame corresponding to the current forward coding frame.
[0186] In a specific embodiment, the number of affected frames corresponding to the impact range index data may refer to the number of forward prediction coding frames belonging to the impact range corresponding to the above-mentioned impact range index data. Specifically, the corresponding number of affected frames can be generated based on the impact range index data. Further, in the case where the impact range index data indicates the number of affected frames, the impact range index data can be used as the number of affected frames corresponding to the impact range index data.
[0187] In a specific embodiment, the target comparison result can be used to indicate the magnitude relationship between the number of affected frames and the number of frames to be updated. The target comparison result may include a first comparison result or a second comparison result. Among them, the first comparison result can be used to indicate that the number of affected frames is less than or equal to the number of frames to be updated. The second comparison result can be used to indicate that the number of affected frames is greater than the number of frames to be updated.
[0188] In a specific embodiment, when the target comparison result is the first comparison result, the target sub-image group corresponding to the current forward coding frame can be merged based on the impact range index data corresponding to the current forward coding frame to obtain the merged sub-image group corresponding to the current forward coding frame. Correspondingly, the position indication information corresponding to each first prediction coding frame among the multiple first prediction coding frames corresponding to the current forward coding frame can be determined; based on the impact range index data corresponding to the current forward coding frame and the position indication information corresponding to each first prediction coding frame, the first quantization index data corresponding to at least one to-be-updated prediction coding frame among the multiple first prediction coding frames corresponding to the current forward coding frame can be updated to obtain the updated quantization index data corresponding to the current forward coding frame.
[0189] In a specific embodiment, when the target comparison result corresponding to any video frame image group is the second comparison result, the remaining to-be-updated coding frames in the above-mentioned any video frame image group may not be updated.
[0190] In the above embodiments, based on the influence range index data corresponding to the current forward-encoded frame, the target sub-image group corresponding to the current forward-encoded frame is merged to obtain the merged sub-image group corresponding to the current forward-encoded frame. The position indication information corresponding to each first predicted encoded frame among the multiple first predicted encoded frames corresponding to the current forward-encoded frame is determined. Based on the influence range index data corresponding to the current forward-encoded frame and the position indication information corresponding to each first predicted encoded frame, the first quantization index data corresponding to at least one to-be-updated predicted encoded frame among the multiple first predicted encoded frames corresponding to the current forward-encoded frame is updated to obtain the updated quantization index data corresponding to the current forward-encoded frame, which can realize the adaptive adjustment of the first quantization index data of the forward-predicted encoded frames in the video to be encoded, and further can effectively improve the encoding effect of the video frames in simple scenes. It can be understood that the encoding effect of the above video frames can mean that, when the image distortion situation is the same, the memory occupancy of the target video encoding data can be reduced; or, when the target video encoding data remains unchanged, the image quality of the video frames can be improved and the distortion can be reduced.
[0191] S211: Based on the updated quantization index data corresponding to multiple forward-predicted encoded frames in each video frame group and the first quantization index data corresponding to the un-updated forward video frames in each video frame group, the video to be encoded is encoded to obtain the target video encoding data.
[0192] In a specific embodiment, the target video encoding data may refer to the encoding result of the video to be encoded. The target video encoding data may include the target encoding data corresponding to each of the multiple target video frames. The target encoding data corresponding to any target video frame may refer to the encoding result of any of the above target video frames.
[0193] In a specific embodiment, the above step S211 may include:
[0194] Based on the updated quantization index data corresponding to multiple forward-predicted encoded frames in each video frame group and the first quantization index data corresponding to the un-updated forward video frames in each video frame group, quantization analysis is performed on multiple bidirectional predicted encoded frames in each video frame group to obtain the second quantization index data corresponding to each of the multiple bidirectional predicted encoded frames in each video frame group;
[0195] Based on the updated quantization index data corresponding to multiple forward-predicted encoded frames in each video frame group, the first quantization index data corresponding to the un-updated forward video frames in each video frame group, and the second quantization index data corresponding to each of the multiple bidirectional predicted encoded frames in each video frame group, the target encoding dimension corresponding to each target video frame is determined from multiple preset encoding dimensions;
[0196] Determine the target quantization metric data corresponding to each target video frame from the updated quantization metric data corresponding to multiple forward prediction encoded frames within each group of video frame images, the first quantization metric data corresponding to the non-updated forward video frames within each group of video frame images, and the second quantization metric data corresponding to each of the multiple bidirectional prediction encoded frames within each group of video frame images;
[0197] Based on the target coding dimension corresponding to each target video frame and the target quantization metric data corresponding to each target video frame, perform video frame coding processing on each target video frame to obtain the target coding data corresponding to each target video frame.
[0198] In a specific embodiment, the second quantization metric data corresponding to any bidirectional prediction encoded frame may refer to the quantization metric data of the above-mentioned any bidirectional prediction encoded frame.
[0199] In a specific embodiment, the quantization metric data corresponding to the intra-coded frame in any group of video frame images may be obtained first; then, two associated reference coded frames corresponding to each bidirectional prediction encoded frame in any group of video frame images may be determined from the multiple reference coded frames corresponding to any group of video frame images; then, based on the quantization metric data of the two associated reference coded frames corresponding to each bidirectional prediction encoded frame, perform quantization analysis on each bidirectional prediction encoded frame to obtain the second quantization metric data corresponding to each bidirectional prediction encoded frame. Among them, the multiple reference coded frames corresponding to any group of video frame images may refer to the intra-coded frames and forward prediction encoded frames included in the above-mentioned any group of video frame images. The two associated reference coded frames corresponding to any bidirectional prediction encoded frame may refer to the reference coded frames on both sides of the above-mentioned any bidirectional prediction encoded frame and close to the above-mentioned any bidirectional prediction encoded frame.
[0200] In a specific embodiment, the target coding dimension corresponding to any target video frame may refer to the coding dimension used for the coding result obtained by coding the above-mentioned any target video frame. The multiple preset coding dimensions may characterize the coding modes adopted by the encoder. Specifically, the multiple preset coding dimensions may be set according to actual application needs, and the present disclosure does not make any limitations.
[0201] In a specific embodiment, the above-mentioned determination of the target coding dimension corresponding to each target video frame from the multiple preset coding dimensions based on the updated quantization metric data corresponding to multiple forward prediction encoded frames within each group of video frame images, the first quantization metric data corresponding to the non-updated forward video frames within each group of video frame images, and the second quantization metric data corresponding to each of the multiple bidirectional prediction encoded frames within each group of video frame images may include:
[0202] Based on the updated quantization metric data corresponding to multiple forward prediction encoded frames within each video frame image group, the first quantization metric data corresponding to the non-updated forward video frames within each video frame image group, and the second quantization metric data corresponding to each of the multiple bidirectional prediction encoded frames within each video frame image group, perform rate-distortion analysis on each target video frame to obtain the rate-distortion metric data corresponding to each target video frame in each preset coding dimension;
[0203] Take the preset coding dimension with the smallest rate-distortion metric data among the rate-distortion metric data corresponding to each target video frame in each of the multiple preset coding dimensions as the target coding dimension corresponding to each target video frame.
[0204] In a specific embodiment, the rate-distortion metric data corresponding to any target video frame in any preset coding dimension can represent the rate-distortion cost of encoding and processing any target video frame through any of the above preset coding dimensions.
[0205] In a specific embodiment, the target quantization metric data corresponding to any target video frame can refer to the quantization metric data used in the encoding process of any of the above target video frames. Specifically, when any of the above target video frames is a bidirectional prediction encoded frame, the second quantization metric data corresponding to any of the above target video frames can be used as the target quantization metric data corresponding to any of the above target video frames; when any of the above target video frames is a forward prediction encoded frame and the corresponding quantization metric data is updated, the updated quantization metric data corresponding to any of the above target video frames can be used as the target quantization metric data corresponding to any of the above target video frames; when any of the above target video frames is a forward prediction encoded frame and the corresponding quantization metric data is not updated, the first quantization metric data corresponding to any of the above target video frames can be used as the target quantization metric data corresponding to any of the above target video frames.
[0206] In a specific embodiment, based on the target quantization metric data corresponding to any of the above target video frames, the bitrate data corresponding to any of the above target video frames in any preset coding dimension and the image quality data corresponding to any of the above target video frames in any preset coding dimension can be determined; then, based on the bitrate data corresponding to any of the above target video frames in any preset coding dimension and the image quality data corresponding to any of the above target video frames in any preset coding dimension, the rate-distortion metric data corresponding to any of the above target video frames in any preset coding dimension can be determined. Specifically, the rate-distortion metric data corresponding to any of the above target video frames in any preset coding dimension can be obtained through the following formula:
[0207] J m =(qscale / 0.85)*R m +D m
[0208] Among them, J m is the rate-distortion metric data corresponding to any target video frame in the m-th preset coding dimension; R m is the bitrate data corresponding to any of the above target video frames in the m-th preset coding dimension; D m is the image quality data corresponding to any of the above target video frames in the m-th preset coding dimension; qscale is the video output quality data corresponding to any of the above target video frames.
[0209] In a specific embodiment, based on the target coding dimension corresponding to any target video frame and the corresponding target quantization metric data, the target bitrate data corresponding to any target video frame in the target coding dimension and the target image quality data corresponding to any target video frame in the target coding dimension can be determined; then, based on a preset encoder, by combining the target bitrate data corresponding to any target video frame in the target coding dimension and the target image quality data corresponding to any target video frame in the target coding dimension, video frame encoding processing can be performed on any target video frame from the target coding dimension corresponding to any target video frame to obtain the target coding data corresponding to any target video frame.
[0210] In the above embodiments, by obtaining a plurality of target video frames of the video to be encoded and a plurality of groups of video frame images corresponding to the video to be encoded, and performing quantization analysis on each forward prediction encoded frame to obtain first quantization index data of each forward prediction encoded frame, quantization analysis of each forward prediction encoded frame can be achieved. Then, from each group of video frame images, the current forward encoded frame is sequentially determined. The current forward encoded frame is the forward prediction encoded frame with the earliest time sequence among the forward prediction encoded frames that are not affected by the distortion of the analyzed encoded frames in each group of video frame images. The analyzed encoded frames are the forward prediction encoded frames for which the distortion influence range analysis is performed. The distortion influence range analysis is performed on the current forward encoded frame to obtain influence range index data corresponding to the current forward encoded frame. The influence range index data corresponding to the current forward encoded frame represents the range of video frames affected by the distortion of the current forward encoded frame during the encoding process of the video to be encoded. The accurate determination of the distortion influence range of the current forward encoded frame on subsequent forward prediction encoded frames can be achieved. Next, in combination with the influence range index data corresponding to the current forward encoded frame, the first quantization index data of the forward prediction encoded frames within the target sub-image group corresponding to the current forward encoded frame is updated to obtain updated quantization index data corresponding to the current forward encoded frame. The target sub-image group corresponding to the current forward encoded frame is the sub-image group that belongs to the influence range indicated by the influence range index data corresponding to the current forward encoded frame. The adaptive adjustment of the first quantization index data of the forward prediction encoded frames in the video to be encoded can be achieved, and thus the encoding effect of video frames in a simple scene can be effectively improved. Then, in combination with the updated quantization index data corresponding to multiple forward prediction encoded frames within each group of video frame images and the first quantization index data corresponding to the unupdated forward video frames within each group of video frame images, the video to be encoded is encoded to obtain target video encoded data, and the encoding effect of video frames in a simple scene can be effectively improved.
[0211] Figure 3 is a block diagram of a video encoding apparatus shown according to an exemplary embodiment. As Figure 3 shown, the apparatus may include:
[0212] A data acquisition module 310, which may be used to obtain a plurality of target video frames of the video to be encoded and a plurality of groups of video frame images corresponding to the video to be encoded; each group of video frame images includes a plurality of sub-image groups; each sub-image group in at least one sub-image group includes forward prediction encoded frames;
[0213] A forward quantization analysis module 320, which may be used to perform quantization analysis on each forward prediction encoded frame to obtain first quantization index data of each forward prediction encoded frame; the first quantization index data is used to indicate the spatial detail compression degree of the video frame;
[0214] The current frame determination module 330 can be used to sequentially determine the current forward encoded frame from each group of video frame images. The current forward encoded frame is the forward prediction encoded frame with the earliest time sequence among the forward prediction encoded frames in each group of video frame images that are not affected by the distortion of the analyzed encoded frames. The analyzed encoded frames are the forward prediction encoded frames for which the distortion influence range analysis is performed.
[0215] The influence range analysis module 340 can be used to perform a distortion influence range analysis on the current forward encoded frame to obtain the influence range index data corresponding to the current forward encoded frame. The influence range index data corresponding to the current forward encoded frame represents the range of video frames affected by the distortion of the current forward encoded frame during the encoding process of the video to be encoded.
[0216] The update module 350 can be used to update the first quantization index data of the forward prediction encoded frames within the target sub-image group corresponding to the current forward encoded frame based on the influence range index data corresponding to the current forward encoded frame, to obtain the updated quantization index data corresponding to the current forward encoded frame. The target sub-image group corresponding to the current forward encoded frame is the sub-image group belonging to the influence range indicated by the influence range index data corresponding to the current forward encoded frame.
[0217] The encoding processing module 360 can be used to perform encoding processing on the video to be encoded based on the updated quantization index data corresponding to multiple forward prediction encoded frames within each group of video frame images and the first quantization index data corresponding to the unupdated forward video frames within each group of video frame images, to obtain the target video encoding data.
[0218] In a specific embodiment, the above update module 350 may include:
[0219] The first merging processing module can be used to perform a merging processing on the target sub-image group corresponding to the current forward encoded frame based on the influence range index data corresponding to the current forward encoded frame, to obtain the merged sub-image group corresponding to the current forward encoded frame.
[0220] The position indication obtaining module can be used to determine the position indication information corresponding to each first prediction encoded frame among the multiple first prediction encoded frames corresponding to the current forward encoded frame. The multiple first prediction encoded frames corresponding to the current forward encoded frame are the multiple forward prediction encoded frames included in the merged sub-image group corresponding to the current forward encoded frame, and the position indication information corresponding to each first prediction encoded frame is used to indicate the position of each first prediction encoded frame among the multiple first prediction encoded frames.
[0221] The first update module can be used to update the first quantization metric data corresponding to at least one to-be-updated prediction coding frame among multiple first prediction coding frames corresponding to the current forward coding frame based on the influence range metric data corresponding to the current forward coding frame and the position indication information corresponding to each first prediction coding frame, so as to obtain the updated quantization metric data corresponding to the current forward coding frame.
[0222] In a specific embodiment, the above-mentioned first update module may include:
[0223] The first weight acquisition module can be used to determine the first weight information corresponding to each to-be-updated prediction coding frame based on the position indication information corresponding to each to-be-updated prediction coding frame and a preset mapping relationship; the preset mapping relationship is the mapping relationship between multiple position indication information and multiple preset weight information;
[0224] The first increment analysis module can be used to perform quantization increment analysis on each to-be-updated prediction coding frame based on the first weight information corresponding to each to-be-updated prediction coding frame and the influence range metric data corresponding to the current forward coding frame, so as to obtain the first increment metric data corresponding to each to-be-updated prediction coding frame;
[0225] The second update module can be used to update the first quantization metric data corresponding to each to-be-updated prediction coding frame based on the first increment metric data corresponding to each to-be-updated prediction coding frame, so as to obtain the first updated metric data corresponding to each to-be-updated prediction coding frame.
[0226] In a specific embodiment, the above-mentioned first update module may include:
[0227] The second weight acquisition module can be used to determine the second weight information corresponding to the second prediction coding frame based on the position indication information corresponding to the second prediction coding frame and a preset mapping relationship;
[0228] The second increment analysis module can be used to perform quantization increment analysis on the second prediction coding frame based on the second weight information corresponding to the second prediction coding frame and the influence range metric data corresponding to the current forward coding frame, so as to obtain the second increment metric data corresponding to the second prediction coding frame;
[0229] The third update module can be used to update the first quantization metric data corresponding to the second prediction coding frame among multiple first prediction coding frames based on the second increment metric data, so as to obtain the second updated metric data corresponding to the second prediction coding frame.
[0230] In a specific embodiment, the above-mentioned update module 350 may include:
[0231] A quantity acquisition module, which can be used to determine the number of frames to be updated corresponding to the current forward-encoded frame and the number of affected frames corresponding to the impact range index data. The number of frames to be updated is the number of forward-predicted encoded frames in each group of video frame images that are not affected by the distortion of the analyzed encoded frames; the number of affected frames is the number of forward-predicted encoded frames belonging to the impact range corresponding to the impact range index data.
[0232] A comparison processing module, which can be used to perform comparison processing on the number of affected frames and the number of frames to be updated to obtain a target comparison result.
[0233] Correspondingly, the above first merging processing module may include:
[0234] A second merging processing module, which can be used to, when the target comparison result indicates that the number of affected frames is less than or equal to the number of frames to be updated, perform merging processing on the target sub-image group corresponding to the current forward-encoded frame based on the impact range index data corresponding to the current forward-encoded frame, to obtain a merged sub-image group corresponding to the current forward-encoded frame.
[0235] In a specific embodiment, the above device may further include:
[0236] A complexity analysis module, which can be used to perform inter-frame complexity analysis on each forward-predicted encoded frame based on each forward-predicted encoded frame and the previous video frame of each forward-predicted encoded frame, to obtain inter-frame complexity index data corresponding to each forward-predicted encoded frame. The inter-frame complexity index data corresponding to each forward-predicted encoded frame represents the degree of difference between each forward-predicted encoded frame and the previous video frame of each forward-predicted encoded frame.
[0237] Correspondingly, the above impact range analysis module 340 may include:
[0238] A distortion impact analysis module, which can be used to perform distortion impact analysis on the current forward-encoded frame based on the inter-frame complexity index data corresponding to the current forward-encoded frame, to obtain impact range index data corresponding to the current forward-encoded frame.
[0239] In a specific embodiment, the above distortion impact analysis module may include:
[0240] An inter-frame correlation analysis module, which can be used to perform inter-frame correlation analysis on the current forward-encoded frame based on the inter-frame complexity index data corresponding to the current forward-encoded frame, to obtain inter-frame correlation index data corresponding to the current forward-encoded frame. The inter-frame correlation index data corresponding to the current forward-encoded frame represents the degree of correlation between the current forward-encoded frame and the previous video frame of the current forward-encoded frame.
[0241] A matching processing module, which can be used to perform matching processing on a preset index value range and the inter-frame correlation index data corresponding to the current forward-encoded frame to obtain a target matching result;
[0242] An index data determination module, which can be used to determine the influence range index data corresponding to the current forward-encoded frame from the two endpoint index data corresponding to the preset index value range and the inter-frame correlation index data based on the target matching result.
[0243] In a specific embodiment, the above index data determination module may include:
[0244] A first index determination module, which can be used to round down the inter-frame correlation index data to obtain the first range index data corresponding to the current forward-encoded frame when the target matching result indicates that the inter-frame correlation index data belongs to the preset index value range.
[0245] In a specific embodiment, the above index data determination module may further include:
[0246] An endpoint index acquisition module, which can be used to determine the adjacent endpoint index data from the two endpoint index data corresponding to the preset index value range when the target matching result indicates that the inter-frame correlation index data does not belong to the preset index value range, and the adjacent endpoint index data is the index data closest to the inter-frame correlation index data among the two endpoint index data;
[0247] A second index determination module, which can be used to use the adjacent endpoint index data as the second range index data corresponding to the current forward-encoded frame.
[0248] In a specific embodiment, the above encoding processing module 360 may include:
[0249] A bidirectional quantization analysis module, which can be used to perform quantization analysis on multiple bidirectional prediction encoded frames in each video frame picture group based on the updated quantization index data corresponding to multiple forward prediction encoded frames in each video frame picture group and the first quantization index data corresponding to the non-updated forward video frames in each video frame picture group to obtain the second quantization index data corresponding to each of the multiple bidirectional prediction encoded frames in each video frame picture group;
[0250] A target dimension determination module, which can be used to determine the target encoding dimension corresponding to each target video frame from multiple preset encoding dimensions based on the updated quantization index data corresponding to multiple forward prediction encoded frames in each video frame picture group, the first quantization index data corresponding to the non-updated forward video frames in each video frame picture group, and the second quantization index data corresponding to each of the multiple bidirectional prediction encoded frames in each video frame picture group;
[0251] A target data determination module, which can be used to determine the target quantization metric data corresponding to each target video frame from the updated quantization metric data corresponding to multiple forward prediction coded frames within each video frame image group, the first quantization metric data corresponding to the non-updated forward video frames within each video frame image group, and the second quantization metric data corresponding to each of the multiple bidirectional prediction coded frames within each video frame image group;
[0252] A video frame encoding module, which can be used to perform video frame encoding processing on each target video frame based on the target encoding dimension corresponding to each target video frame and the target quantization metric data corresponding to each target video frame, to obtain the target encoded data corresponding to each target video frame.
[0253] In a specific embodiment, the above-mentioned target dimension determination module may include:
[0254] A rate-distortion analysis module, which can be used to perform rate-distortion analysis on each target video frame based on the updated quantization metric data corresponding to multiple forward prediction coded frames within each video frame image group, the first quantization metric data corresponding to the non-updated forward video frames within each video frame image group, and the second quantization metric data corresponding to each of the multiple bidirectional prediction coded frames within each video frame image group, to obtain the rate-distortion metric data corresponding to each target video frame at each preset encoding dimension;
[0255] An encoding dimension determination module, which can be used to use the preset encoding dimension with the smallest rate-distortion metric data corresponding to each target video frame among the rate-distortion metric data corresponding to each target video frame at multiple preset encoding dimensions as the target encoding dimension corresponding to each target video frame.
[0256] Regarding the device in the above embodiment, the specific manners in which each module and unit perform operations have been described in detail in the embodiment of the method, and will not be elaborated here.
[0257] Figure 4 is a block diagram of an electronic device for encoding multiple target video frames in a video to be encoded according to an exemplary embodiment. The electronic device may be a server, and its internal structure diagram may be as Figure 4 shown. The electronic device includes a processor, a memory, and a network interface connected through a system bus. Among them, the processor of the electronic device is used to provide computing and control capabilities. The memory of the electronic device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the electronic device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, a video encoding method is implemented.
[0258] Figure 5 It is a block diagram of another electronic device for encoding multiple target video frames in a video to be encoded according to an exemplary embodiment. The electronic device may be a terminal, and its internal structure diagram may be as follows Figure 5 shown. The electronic device includes a processor, a memory, a network interface, a display screen, and an input device connected through a system bus. Among them, the processor of the electronic device is used to provide computing and control capabilities. The memory of the electronic device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the electronic device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, a video encoding method is implemented. The display screen of the electronic device may be a liquid crystal display screen or an electronic ink display screen. The input device of the electronic device may be a touch layer covering the display screen, or a button, a trackball, or a touchpad provided on the housing of the electronic device, or an external keyboard, touchpad, or mouse, etc.
[0259] Those skilled in the art can understand that Figure 4 or Figure 5 the structure shown in is only a block diagram of some structures related to the solution of the present disclosure, and does not constitute a limitation on the electronic device to which the solution of the present disclosure is applied. The specific electronic device may include more or fewer components than those shown in the figure, or combine some components, or have a different component layout.
[0260] In an exemplary embodiment, there is also provided an electronic device, including: a processor; a memory for storing executable instructions of the processor; wherein, the processor is configured to execute the instructions to implement the video encoding method in the embodiment of the present disclosure.
[0261] In an exemplary embodiment, there is also provided a computer-readable storage medium. When the instructions in the storage medium are executed by the processor of the electronic device, the electronic device can execute the video encoding method in the embodiment of the present disclosure.
[0262] In an exemplary embodiment, there is also provided a computer program product containing instructions. When it runs on a computer, the computer executes the video encoding method in the embodiment of the present disclosure.
[0263] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. This computer program can be stored in a non-volatile computer-readable storage medium. When this computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium used in the embodiments provided in this application can include non-volatile and / or volatile memories. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or an external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and Rambus dynamic RAM (RDRAM), etc.
[0264] In the embodiments of this application, the term "module" or "unit" refers to a computer program with a predetermined function or a part of a computer program, which works together with other related parts to achieve a predetermined goal, and can be fully or partially implemented by using software, hardware (such as a processing circuit or a memory), or a combination thereof. Similarly, one processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be a part of an overall module or unit that includes the functions of that module or unit.
[0265] It can be understood that in the specific implementation manners of this application, when it comes to relevant data such as user information, when the above embodiments of this application are applied to specific products or technologies, user permission or consent needs to be obtained, and the collection, use, and processing of relevant data need to comply with the relevant laws, regulations, and standards of relevant countries and regions.
[0266] After considering the specification and practicing the invention disclosed herein, those skilled in the art will readily conceive of other embodiments of the present disclosure. This application is intended to cover any variations, uses, or adaptations of the present disclosure, which follow the general principles of the present disclosure and include known common knowledge or conventional technical means in the technical field not disclosed in the present disclosure. The specification and embodiments are only to be considered as exemplary, and the true scope and spirit of the present disclosure are pointed out by the following claims.
[0267] It should be understood that the present disclosure is not limited to the exact structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present disclosure is only limited by the appended claims.
Claims
1. A video encoding method, characterized in that, The method includes: Obtaining a plurality of target video frames of the video to be encoded, and a plurality of video frame image groups corresponding to the video to be encoded; each video frame image group includes a plurality of sub-image groups; each sub-image group in at least one sub-image group includes a forward prediction coding frame; Performing quantization analysis on each forward prediction coding frame to obtain first quantization index data of each forward prediction coding frame; the first quantization index data is used to indicate the spatial detail compression condition of the video frame; Sequentially determining a current forward coding frame from each video frame image group, where the current forward coding frame is the forward prediction coding frame with the earliest time sequence among the forward prediction coding frames in each video frame image group that are not affected by the distortion of the analyzed coding frames, and the analyzed coding frames are the forward prediction coding frames for which the distortion influence range analysis has been performed; Performing distortion influence range analysis on the current forward coding frame to obtain influence range index data corresponding to the current forward coding frame; the influence range index data corresponding to the current forward coding frame characterizes the range of video frames affected by the distortion of the current forward coding frame during the encoding process of the video to be encoded; Based on the influence range index data corresponding to the current forward coding frame, updating the first quantization index data of the forward prediction coding frames in the target sub-image group corresponding to the current forward coding frame to obtain updated quantization index data corresponding to the current forward coding frame; the target sub-image group corresponding to the current forward coding frame is the sub-image group belonging to the influence range indicated by the influence range index data corresponding to the current forward coding frame; Based on the updated quantization index data corresponding to the multiple forward prediction coding frames in each video frame image group and the first quantization index data corresponding to the un-updated forward video frames in each video frame image group, performing encoding processing on the video to be encoded to obtain target video encoding data.
2. The method according to claim 1, characterized in that, The step of updating the first quantization index data of the forward prediction coding frames in the target sub-image group corresponding to the current forward coding frame based on the influence range index data corresponding to the current forward coding frame to obtain the updated quantization index data corresponding to the current forward coding frame includes: Based on the influence range index data corresponding to the current forward coding frame, performing merging processing on the target sub-image group corresponding to the current forward coding frame to obtain a merged sub-image group corresponding to the current forward coding frame; Determining position indication information corresponding to each first prediction coding frame among the multiple first prediction coding frames corresponding to the current forward coding frame; the multiple first prediction coding frames corresponding to the current forward coding frame are the multiple forward prediction coding frames included in the merged sub-image group corresponding to the current forward coding frame, and the position indication information corresponding to each first prediction coding frame is used to indicate the position of each first prediction coding frame among the multiple first prediction coding frames; Based on the influence range index data corresponding to the current forward coding frame and the position indication information corresponding to each first predicted coding frame, update the first quantization index data corresponding to at least one to-be-updated predicted coding frame among the multiple first predicted coding frames corresponding to the current forward coding frame, to obtain the updated quantization index data corresponding to the current forward coding frame.
3. The method according to claim 2, wherein The updated quantization index data corresponding to the current forward coding frame includes the first updated index data corresponding to each to-be-updated predicted coding frame; when at least one to-be-updated predicted coding frame corresponding to the current forward coding frame is the multiple first predicted coding frames corresponding to the current forward coding frame, the updating the first quantization index data corresponding to at least one to-be-updated predicted coding frame among the multiple first predicted coding frames corresponding to the current forward coding frame based on the influence range index data corresponding to the current forward coding frame and the position indication information corresponding to each first predicted coding frame, to obtain the updated quantization index data corresponding to the current forward coding frame, includes: Determine the first weight information corresponding to each to-be-updated predicted coding frame based on the position indication information corresponding to each to-be-updated predicted coding frame and a preset mapping relationship; the preset mapping relationship is a mapping relationship between multiple position indication information and multiple preset weight information; Perform quantization increment analysis on each to-be-updated predicted coding frame based on the first weight information corresponding to each to-be-updated predicted coding frame and the influence range index data corresponding to the current forward coding frame, to obtain the first increment index data corresponding to each to-be-updated predicted coding frame; Update the first quantization index data corresponding to each to-be-updated predicted coding frame based on the first increment index data corresponding to each to-be-updated predicted coding frame, to obtain the first updated index data corresponding to each to-be-updated predicted coding frame.
4. The method according to claim 2, wherein The updated quantization index data corresponding to the current forward coding frame includes the second updated index data corresponding to a second predicted coding frame, and the second predicted coding frame is the first predicted coding frame with the earliest time sequence among the multiple first predicted coding frames corresponding to the current forward coding frame; When at least one to-be-updated predicted coding frame corresponding to the current forward coding frame is the second predicted coding frame, the updating the first quantization index data corresponding to at least one to-be-updated predicted coding frame among the multiple first predicted coding frames corresponding to the current forward coding frame based on the influence range index data corresponding to the current forward coding frame and the position indication information corresponding to each first predicted coding frame, to obtain the updated quantization index data corresponding to the current forward coding frame, includes: Determine the second weight information corresponding to the second predicted coding frame based on the position indication information corresponding to the second predicted coding frame and the preset mapping relationship; Perform quantization increment analysis on the second predicted coding frame based on the second weight information corresponding to the second predicted coding frame and the influence range index data corresponding to the current forward coding frame, to obtain the second increment index data corresponding to the second predicted coding frame; Based on the second incremental index data, update the first quantization index data corresponding to the second predicted coding frame among the multiple first predicted coding frames to obtain the second updated index data corresponding to the second predicted coding frame.
5. The method according to claim 2, wherein The method of updating the first quantization index data of the forward predicted coding frames within the target sub-image group corresponding to the current forward coding frame based on the influence range index data corresponding to the current forward coding frame to obtain the updated quantization index data corresponding to the current forward coding frame further includes: Determine the number of frames to be updated corresponding to the current forward coding frame and the number of affected frames corresponding to the influence range index data. The number of frames to be updated is the number of forward predicted coding frames in each video frame image group that are not affected by the distortion of the analyzed coded frames. The number of affected frames is the number of forward predicted coding frames belonging to the influence range corresponding to the influence range index data. Perform a comparison process on the number of affected frames and the number of frames to be updated to obtain a target comparison result. The method of merging the target sub-image group corresponding to the current forward coding frame based on the influence range index data corresponding to the current forward coding frame to obtain the merged sub-image group corresponding to the current forward coding frame includes: When the target comparison result indicates that the number of affected frames is less than or equal to the number of frames to be updated, merge the target sub-image group corresponding to the current forward coding frame based on the influence range index data corresponding to the current forward coding frame to obtain the merged sub-image group corresponding to the current forward coding frame.
6. The method according to claim 1, characterized in that Before quantifying and analyzing each forward predicted coding frame to obtain the first quantization index data of each forward predicted coding frame, the method further includes: Based on each forward predicted coding frame and the previous video frame of each forward predicted coding frame, perform an inter-frame complexity analysis on each forward predicted coding frame to obtain the inter-frame complexity index data corresponding to each forward predicted coding frame. The inter-frame complexity index data corresponding to each forward predicted coding frame characterizes the degree of difference between each forward predicted coding frame and the previous video frame of each forward predicted coding frame. The method of analyzing the distortion influence range of the current forward coding frame to obtain the influence range index data corresponding to the current forward coding frame includes: Based on the inter-frame complexity index data corresponding to the current forward coding frame, perform a distortion influence analysis on the current forward coding frame to obtain the influence range index data corresponding to the current forward coding frame.
7. The method according to claim 6, characterized in that, The method of performing a distortion influence analysis on the current forward coding frame based on the inter-frame complexity index data corresponding to the current forward coding frame to obtain the influence range index data corresponding to the current forward coding frame includes: Perform inter-frame correlation analysis on the current forward-coded frame based on the inter-frame complexity metric data corresponding to the current forward-coded frame to obtain the inter-frame correlation metric data corresponding to the current forward-coded frame; the inter-frame correlation metric data corresponding to the current forward-coded frame represents the degree of correlation between the current forward-coded frame and the previous video frame of the current forward-coded frame. Perform a matching process on the preset metric value range and the inter-frame correlation metric data corresponding to the current forward-coded frame to obtain a target matching result. Based on the target matching result, determine the influence range metric data corresponding to the current forward-coded frame from the two endpoint metric data corresponding to the preset metric value range and the inter-frame correlation metric data.
8. The method according to claim 7, wherein The influence range metric data corresponding to the current forward-coded frame includes first range metric data; the determining the influence range metric data corresponding to the current forward-coded frame from the two endpoint metric data corresponding to the preset metric value range and the inter-frame correlation metric data based on the target matching result includes: In the case where the target matching result indicates that the inter-frame correlation metric data belongs to the preset metric value range, round down the inter-frame correlation metric data to obtain the first range metric data corresponding to the current forward-coded frame.
9. The method according to claim 8, wherein The influence range metric data corresponding to the current forward-coded frame includes second range metric data; the method further includes: In the case where the target matching result indicates that the inter-frame correlation metric data does not belong to the preset metric value range, determine the adjacent endpoint metric data from the two endpoint metric data corresponding to the preset metric value range, and the adjacent endpoint metric data is the metric data closest to the inter-frame correlation metric data among the two endpoint metric data. Use the adjacent endpoint metric data as the second range metric data corresponding to the current forward-coded frame.
10. The method according to claim 1, characterized in that, The target video coding data includes the target coding data corresponding to each of the multiple target video frames; the encoding the video to be encoded to obtain the target video coding data based on the updated quantization metric data corresponding to multiple forward prediction coded frames in each video frame group of pictures and the first quantization metric data corresponding to the non-updated forward video frames in each video frame group of pictures includes: Perform quantization analysis on multiple bidirectional prediction coded frames in each video frame group of pictures based on the updated quantization metric data corresponding to multiple forward prediction coded frames in each video frame group of pictures and the first quantization metric data corresponding to the non-updated forward video frames in each video frame group of pictures to obtain the second quantization metric data corresponding to each of the multiple bidirectional prediction coded frames in each video frame group of pictures. Based on the updated quantization metric data corresponding to multiple forward prediction coded frames in each video frame group of pictures, the first quantization metric data corresponding to the non-updated forward video frames in each video frame group of pictures, and the second quantization metric data corresponding to each of the multiple bidirectional prediction coded frames in each video frame group of pictures, determine the target coding dimension corresponding to each target video frame from multiple preset coding dimensions. Determine the target quantization index data corresponding to each target video frame from the updated quantization index data corresponding to multiple forward prediction coded frames within each group of video frame images, the first quantization index data corresponding to the non-updated forward video frames within each group of video frame images, and the second quantization index data corresponding to each of the multiple bidirectional prediction coded frames within each group of video frame images; Based on the target coding dimension corresponding to each target video frame and the target quantization index data corresponding to each target video frame, perform video frame coding processing on each target video frame to obtain the target coding data corresponding to each target video frame.
11. The method according to claim 10, characterized in that, The determining, from multiple preset coding dimensions, the target coding dimension corresponding to each target video frame based on the updated quantization index data corresponding to multiple forward prediction coded frames within each group of video frame images, the first quantization index data corresponding to the non-updated forward video frames within each group of video frame images, and the second quantization index data corresponding to each of the multiple bidirectional prediction coded frames within each group of video frame images, includes: Perform rate-distortion analysis on each target video frame based on the updated quantization index data corresponding to multiple forward prediction coded frames within each group of video frame images, the first quantization index data corresponding to the non-updated forward video frames within each group of video frame images, and the second quantization index data corresponding to each of the multiple bidirectional prediction coded frames within each group of video frame images, to obtain the rate-distortion index data corresponding to each target video frame in each preset coding dimension; Use the preset coding dimension with the smallest rate-distortion index data among the rate-distortion index data corresponding to each target video frame in multiple preset coding dimensions as the target coding dimension corresponding to each target video frame.
12. A video encoding device, characterized in that, The apparatus includes: A data acquisition module, configured to acquire multiple target video frames of a video to be coded, and multiple groups of video frame images corresponding to the video to be coded; each group of video frame images includes multiple sub-groups of images; each sub-group of images in at least one sub-group of images includes forward prediction coded frames; A forward quantization analysis module, configured to perform quantization analysis on each forward prediction coded frame to obtain the first quantization index data of each forward prediction coded frame; the first quantization index data is used to indicate the spatial detail compression degree of the video frame; A current frame determination module, configured to sequentially determine a current forward coded frame from each group of video frame images, where the current forward coded frame is the forward prediction coded frame with the earliest time sequence among the forward prediction coded frames in each group of video frame images that are not affected by the distortion of the analyzed coded frames, and the analyzed coded frames are the forward prediction coded frames for which the distortion influence range analysis has been performed; An influence range analysis module, configured to perform distortion influence range analysis on the current forward coded frame to obtain the influence range index data corresponding to the current forward coded frame; the influence range index data corresponding to the current forward coded frame represents the range of video frames affected by the distortion of the current forward coded frame during the coding process of the video to be coded; An update module, configured to update first quantization index data of forward prediction coding frames within a target sub-image group corresponding to the current forward coding frame based on influence range index data corresponding to the current forward coding frame, so as to obtain updated quantization index data corresponding to the current forward coding frame; the target sub-image group corresponding to the current forward coding frame is a sub-image group belonging to an influence range indicated by the influence range index data corresponding to the current forward coding frame; A coding processing module, configured to perform coding processing on the video to be coded based on the updated quantization index data corresponding to multiple forward prediction coding frames within each video frame image group and the first quantization index data corresponding to unupdated forward video frames within each video frame image group, so as to obtain target video coding data.
13. An electronic device, characterized in that, Comprising: A processor; A memory for storing processor-executable instructions; Wherein, the processor is configured to execute the executable instructions to implement the video coding method according to any one of claims 1 to 11.