Data processing method, device and equipment and readable storage medium
By performing machine learning training on filtering models in video encoding applications, the problem of video image distortion in the prior art is solved, and more efficient video encoding is achieved, reducing image distortion and improving image quality.
Patent Information
- Application Number
- CN202510353186.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-21
- Publication Date
- 2025-05-30
AI Technical Summary
The existing video encoding algorithm adopts lossy compression method, resulting in serious distortion of video images. Traditional loop filters rely on manual design of filter coefficients, and the accuracy is not high, so they cannot effectively reduce distortion.
Through a data processing method, machine learning technology is used to train the filtering model in video encoding applications to generate a more efficient filtering model and reduce the distortion of video images. The specific steps include inputting the sample video data into the video encoding application, generating training data, training the filter model based on the training data, updating and deploying a new filter model until the filter quality requirements are met.
Improve filtering performance, reduce image distortion of encoded videos, improve image quality of encoded videos without relying on manual experience.
Smart Images

Figure CN120075440A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, and particularly to a data processing method, apparatus, device, and readable storage medium. Background Art
[0002] Currently, in the process of encoding and compressing a video, all video encoding and compression algorithms adopt a lossy compression method, that is, there is a certain distortion between the image of the encoded and compressed video and the image of the original video. Moreover, in the case of a high compression ratio, the distortion and distortion of the image are even more serious.
[0003] Existing video coding standards have introduced a loop filter to filter the image of the compressed video, thereby reducing the degree of distortion and enabling the image quality of the compressed video to be infinitely close to the image quality of the original video. However, since the traditional loop filter mainly designs the filter coefficients manually and relies too much on manual experience, the accuracy is not high and it cannot well reduce the degree of distortion. Summary of the Invention
[0004] Embodiments of this application provide a data processing method, apparatus, device, and readable storage medium, which can improve the filtering performance, reduce the image distortion degree of the encoded video, and improve the image quality of the encoded video.
[0005] One aspect of the embodiments of this application provides a data processing method, including:
[0006] Input sample video data into a video coding application including a first filtering model updated and deployed for the kth time, and generate first training data through the video coding application updated and deployed for the kth time and the sample video data; the first training data includes a sample original video frame as a training label and a first sample frame to be filtered and reconstructed corresponding to the sample original video frame; the first sample frame to be filtered and reconstructed refers to a reconstructed frame that has not been filtered by the first filtering model during the process of reconstructing the sample original video frame through the video coding application updated and deployed for the kth time; the sample original video frame is a video frame in the sample video data; k is a positive integer;
[0007] Train a filtering model to be trained in the video coding application updated and deployed for the kth time based on the sample original video frame and the first sample frame to be filtered and reconstructed, obtain a second filtering model in a training convergence state, and update and deploy the second filtering model in the video coding application updated and deployed for the kth time to obtain a video coding application updated and deployed for the (k + 1)th time;
[0008] When the video coding application updated and deployed for the (k + 1)th time meets the filtering quality requirement condition, determine the video coding application updated and deployed for the (k + 1)th time as the target video coding application for video coding processing of video data.
[0009] In one aspect, an embodiment of the present application provides another data processing method, including:
[0010] Input video data into a target video encoding application, and perform video encoding processing on the video data through the target video encoding application to obtain a video compression bitstream corresponding to the video data; the target video encoding application refers to the video encoding application updated in the (k + 1)-th deployment that meets the filtering quality requirement conditions; the video encoding application updated in the (k + 1)-th deployment includes a second filtering model in a training convergence state; the second filtering model is obtained by training the filtering model to be trained in the video encoding application including the first filtering model updated in the k-th deployment based on the sample original video frames as training labels in the first training data and the first sample reconstruction frame to be filtered corresponding to the sample original video frames; the first training data is generated by the video encoding application updated in the k-th deployment and the sample video data; the first sample reconstruction frame to be filtered refers to the reconstruction frame that has not been filtered by the first filtering model during the process of reconstructing the sample original video frames through the video encoding application updated in the k-th deployment; the sample original video frames are the video frames in the sample video data; k is a positive integer;
[0011] Send the video compression bitstream to a receiving device so that the receiving device performs decoding processing on the video compression bitstream.
[0012] In one aspect, an embodiment of the present application provides a data processing device, including:
[0013] A training data generation module, configured to input sample video data into the video encoding application including the first filtering model updated in the k-th deployment;
[0014] The training data generation module is further configured to generate first training data through the video encoding application updated in the k-th deployment and the sample video data; the first training data includes sample original video frames as training labels and first sample reconstruction frames to be filtered corresponding to the sample original video frames; the first sample reconstruction frame to be filtered refers to the reconstruction frame that has not been filtered by the first filtering model during the process of reconstructing the sample original video frames through the video encoding application updated in the k-th deployment; the sample original video frames are the video frames in the sample video data; k is a positive integer;
[0015] A model training module, configured to train the filtering model to be trained in the video encoding application updated in the k-th deployment based on the sample original video frames and the first sample reconstruction frames to be filtered, and obtain a second filtering model in a training convergence state;
[0016] An application update module, configured to update and deploy the second filtering model in the video encoding application updated in the k-th deployment to obtain the video encoding application updated in the (k + 1)-th deployment;
[0017] A target application determination module, configured to determine the video encoding application updated in the (k + 1)-th deployment as the target video encoding application for performing video encoding processing on video data when the video encoding application updated in the (k + 1)-th deployment meets the filtering quality requirement condition.
[0018] In one embodiment, the first filtering model includes a first intra-frame filtering model for filtering a to-be-filtered reconstructed frame belonging to the intra-frame prediction type; the first intra-frame filtering model is trained based on the initial intra-frame filtering model in the video encoding application updated in the (k - 1)-th deployment, and the video encoding application updated in the k-th deployment further includes an untrained initial inter-frame filtering model; the initial inter-frame filtering model is used for filtering a to-be-filtered reconstructed frame belonging to the inter-frame prediction type, and the to-be-trained filtering model in the video encoding application updated in the k-th deployment includes the initial inter-frame filtering model and the first intra-frame filtering model;
[0019] The model training module includes:
[0020] A video label acquisition unit, configured to acquire a to-be-filtered inter-frame reconstructed frame belonging to the inter-frame prediction type in a first sample to-be-filtered reconstructed frame, and use the sample original video frame corresponding to the to-be-filtered inter-frame reconstructed frame in the sample original video frame as a first video frame label;
[0021] The video label acquisition unit is further configured to acquire a first to-be-filtered intra-frame reconstructed frame belonging to the intra-frame prediction type in the first sample to-be-filtered reconstructed frame, and use the sample original video frame corresponding to the first to-be-filtered intra-frame reconstructed frame in the sample original video frame as a second video frame label;
[0022] A model training unit, configured to train the initial inter-frame filtering model based on the to-be-filtered inter-frame reconstructed frame and the first video frame label to obtain a first inter-frame filtering model in a training convergence state;
[0023] The model training unit is further configured to train the first intra-frame filtering model based on the to-be-filtered intra-frame reconstructed frame and the second video frame label to obtain a second intra-frame filtering model in a training convergence state;
[0024] A model determination unit, configured to determine the first inter-frame filtering model and the second intra-frame filtering model as a second filtering model.
[0025] In one embodiment, the application update module includes:
[0026] An intra-frame model replacement unit, configured to replace and update the first intra-frame filtering model in the video encoding application updated in the k-th deployment with the second intra-frame filtering model;
[0027] An inter-frame model replacement unit for replacing and updating the initial inter-frame filtering model in the video coding application updated in the k-th deployment with a first inter-frame filtering model.
[0028] In one embodiment, the first filtering model includes a first intra-frame filtering model for filtering a reconstructed frame to be filtered belonging to the intra-frame prediction type; the first intra-frame filtering model is trained based on the initial intra-frame filtering model in the video coding application updated in the (k - 1)-th deployment, and the video coding application updated in the k-th deployment further includes an untrained first-type initial inter-frame filtering model and a second-type initial inter-frame filtering model; the first-type initial inter-frame filtering model is used for filtering a reconstructed frame to be filtered belonging to the first inter-frame prediction type, and the second-type initial inter-frame filtering model is used for filtering a reconstructed frame to be filtered belonging to the second inter-frame prediction type;
[0029] The model training module includes:
[0030] A model to be trained determination unit for determining the first intra-frame filtering model and the first-type initial inter-frame filtering model in the video coding application updated in the k-th deployment as the filtering model to be trained in the video coding application updated in the k-th deployment;
[0031] A video frame label determination unit for obtaining a first-type reconstructed frame to be filtered belonging to the first inter-frame prediction type in the first sample reconstructed frame to be filtered, and taking the sample original video frame corresponding to the first-type reconstructed frame to be filtered in the sample original video frame as the third video frame label;
[0032] The video frame label determination unit is further configured to obtain a second reconstructed frame to be filtered belonging to the intra-frame prediction type in the first sample reconstructed frame to be filtered, and take the sample original video frame corresponding to the second reconstructed frame to be filtered in the sample original video frame as the fourth video frame label;
[0033] An inter-frame model training unit for training the first-type initial inter-frame filtering model based on the first-type reconstructed frame to be filtered and the third video frame label to obtain a first-type inter-frame filtering model in a training convergence state;
[0034] An intra-frame model training unit for training the first intra-frame filtering model based on the second reconstructed frame to be filtered and the fourth video frame label to obtain a second intra-frame filtering model in a training convergence state;
[0035] A filtering model determination unit for determining the first-type inter-frame filtering model and the second intra-frame filtering model as the second filtering model.
[0036] In one embodiment, the data processing device further includes:
[0037] A data generation module, configured to generate second training data by using the video encoding application updated in the (k + 1)-th deployment and sample video data when the video encoding application updated in the (k + 1)-th deployment does not meet the filtering quality requirement condition; the second training data includes sample original video frames as training labels, and second sample frames to be filtered and reconstructed corresponding to the sample original video frames, where the second sample frames to be filtered and reconstructed refer to video frames that are not filtered by the second filtering model during the process of reconstructing the sample original video frames by using the video encoding application updated in the (k + 1)-th deployment.
[0038] A filtering model training module, configured to train a filtering model to be trained in the video encoding application updated in the (k + 1)-th deployment based on the sample original video frames and the second sample frames to be filtered and reconstructed, to obtain a third filtering model in a training convergence state.
[0039] A deployment model module, configured to update and deploy the third filtering model in the video encoding application updated in the (k + 1)-th deployment, to obtain a video encoding application updated in the (k + 2)-th deployment.
[0040] A target application determination module, configured to determine the video encoding application updated in the (k + 2)-th deployment as a target video encoding application when the video encoding application updated in the (k + 2)-th deployment meets the filtering quality requirement condition.
[0041] In one embodiment, the filtering model to be trained in the video encoding application updated in the (k + 1)-th deployment includes a second-type initial inter-frame filtering model.
[0042] The filtering model training module includes:
[0043] A label determination unit, configured to obtain second-type frames to be filtered and reconstructed belonging to a second inter-frame prediction type in the second sample frames to be filtered and reconstructed, and use the sample original video frames corresponding to the second-type frames to be filtered and reconstructed in the sample original video frames as fifth video frame labels.
[0044] A type inter-frame model training unit, configured to train the second-type initial inter-frame filtering model based on the second-type frames to be filtered and reconstructed and the fifth video frame labels, to obtain a second-type inter-frame filtering model in a training convergence state.
[0045] A filtering model determination unit, configured to determine the second-type inter-frame filtering model as the third filtering model.
[0046] In one embodiment, the model training module includes:
[0047] A filtered frame output unit, configured to input a first sample frame to be filtered and reconstructed into a first filtering model, and output a sample filtered and reconstructed frame corresponding to the first sample frame to be filtered and reconstructed through the first filtering model.
[0048] A parameter adjustment unit, configured to determine an error value between a sample filtered reconstruction frame and a sample original video frame;
[0049] The parameter adjustment unit is further configured to adjust model parameters of a first filtering model through the error value to obtain a first filtering model with adjusted model parameters;
[0050] The parameter adjustment unit is further configured to, when the first filtering model with adjusted model parameters meets a model convergence condition, determine the first filtering model with adjusted model parameters as a second filtering model in a training convergence state.
[0051] In one embodiment, the parameter adjustment unit includes:
[0052] A function acquisition subunit, configured to acquire a loss function corresponding to a video coding application updated in the k-th deployment;
[0053] An error value determination subunit, configured to acquire an original image quality corresponding to a sample original video frame based on the loss function, and use the original image quality as an image quality label;
[0054] The error value determination subunit is further configured to acquire a filtered image quality corresponding to a sample filtered reconstruction frame based on the loss function, and determine an error value between the sample filtered video frame and the sample original video frame through the loss function, the image quality label, and the filtered image quality.
[0055] In one embodiment, the loss function includes an absolute value loss function and a mean squared error loss function;
[0056] The error value determination subunit is further specifically configured to determine a first error value between the sample filtered reconstruction frame and the sample original video frame through the mean squared error loss function, the image quality label, and the filtered image quality;
[0057] The error value determination subunit is further specifically configured to determine a second error value between the sample filtered reconstruction frame and the sample original video frame through the absolute value loss function, the image quality label, and the filtered image quality;
[0058] The error value determination subunit is further specifically configured to perform arithmetic processing on the first error value and the second error value, and determine the result of the arithmetic as the error value between the sample filtered reconstruction frame and the sample original video frame.
[0059] In one embodiment, the data processing device further includes:
[0060] A filtered video output module, configured to input sample video data into a video coding application updated in the (k + 1)-th deployment, and output to-be-detected filtered video data through the video coding application updated in the (k + 1)-th deployment;
[0061] A video quality acquisition module, configured to acquire the filtered video quality corresponding to the video data to be detected for filtering, and the original video quality corresponding to the sample video data;
[0062] An application detection module, configured to detect the video encoding application updated in the (k + 1)-th deployment according to the filtered video quality and the original video quality.
[0063] In one embodiment, the application detection module includes:
[0064] A difference quality determination unit, configured to determine the difference video quality between the filtered video quality and the original video quality;
[0065] A detection result determination unit, configured to determine that the second video encoding application meets the filtered quality requirement condition if the difference video quality is less than the difference quality threshold;
[0066] The detection result determination unit is further configured to determine that the second video encoding application does not meet the filtered quality requirement condition if the difference video quality is greater than the difference quality threshold.
[0067] On the one hand, an embodiment of the present application provides another data processing device, including:
[0068] A bitstream generation module, configured to input video data into a target video encoding application, perform video encoding processing on the video data through the target video encoding application to obtain a video compression bitstream corresponding to the video data; the target video encoding application refers to the video encoding application updated in the (k + 1)-th deployment that meets the filtered quality requirement condition; the video encoding application updated in the (k + 1)-th deployment includes a second filtering model in a training convergence state; the second filtering model is obtained by training the filtering model to be trained in the video encoding application including the first filtering model updated in the k-th deployment based on the sample original video frames as training labels in the first training data and the first sample frames to be filtered and reconstructed corresponding to the sample original video frames; the first training data is generated by the video encoding application updated in the k-th deployment and the sample video data; the first sample frames to be filtered and reconstructed refer to the reconstructed frames that are not filtered by the first filtering model during the process of reconstructing the sample original video frames through the video encoding application updated in the k-th deployment; the sample original video frames are the video frames in the sample video data; k is a positive integer;
[0069] A bitstream sending module, configured to send the video compression bitstream to a receiving device so that the receiving device decodes the video compression bitstream.
[0070] On the one hand, an embodiment of the present application provides a computer device, including: a processor and a memory;
[0071] The memory stores a computer program, which, when executed by a processor, causes the processor to execute the method in the embodiments of the present application.
[0072] One aspect of the embodiments of the present application provides a computer-readable storage medium storing a computer program, the computer program including program instructions that, when executed by a processor, execute the method in the embodiments of the present application.
[0073] One aspect of the present application provides a computer program product or a computer program, the computer program product or the computer program including computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, causing the computer device to execute the method provided in one aspect of the embodiments of the present application.
[0074] In an embodiment of the present application, sample video data can be input into a video encoding application updated and deployed for the k-th time (e.g., the 1st time, the 2nd time, the 3rd time, ...), which is a video encoding application deployed with a first filtering model. Through the video encoding application updated and deployed for the k-th time, first training data corresponding to the sample video data can be output; based on the first training data, the first filtering model can be retrained to obtain a second filtering model, and the second filtering model can be used to update and deploy the video encoding application updated and deployed for the k-th time again, so as to obtain a video encoding application updated and deployed for the (k + 1)-th time; at this time, the video encoding application updated and deployed for the (k + 1)-th time can be detected. When the video encoding application updated and deployed for the (k + 1)-th time meets the filtering quality requirement conditions, the video encoding application updated and deployed for the (k + 1)-th time can be determined as the target video encoding application for video encoding processing of video data. It should be understood that deploying a filtering model (e.g., the first filtering model) to a video encoding application is a deployment update of the video encoding application. After generating training data based on the video encoding application, the filtering model can be trained in a machine learning manner, and then the trained filtering model can be deployed and updated to the video encoding application to obtain an updated video encoding application; subsequently, the present application can continue to generate new training data based on the updated video encoding application, retrain the filtering model based on the new training data, and then deploy and update the trained filtering model to the video encoding application again to obtain an updated video encoding application again until the video encoding meets the filtering quality requirement conditions. That is to say, after deploying the filtering model to the video encoding application, the present application can iteratively update the video encoding application, repeat the training process of the filtering model by updating the training data, and continuously deploy and update the video encoding application, which can improve the consistency between the training effect and the testing effect of the filtering model, improve the encoding efficiency, and can also improve the filtering quality of the filtering model deployed in the video encoding application without relying on manual experience, and reduce the distortion degree of the encoded video. In summary, the present application can improve the filtering performance, reduce the image distortion degree of the encoded video, and improve the image quality of the encoded video. BRIEF DESCRIPTION OF THE DRAWINGS
[0075] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0076] Figure 1 It is a schematic structural diagram of a network architecture provided by an embodiment of the present application;
[0077] Figure 2 It is a schematic flowchart of a data processing method provided by an embodiment of the present application;
[0078] Figure 3 It is a schematic diagram of a coding reference relationship provided by an embodiment of the present application;
[0079] Figure 4 It is a schematic diagram of a coding reference relationship provided by an embodiment of the present application;
[0080] Figure 5 It is a system architecture diagram of iterative training provided by an embodiment of the present application;
[0081] Figure 6 It is a schematic flowchart of a process for training a filtering model provided by an embodiment of the present application;
[0082] Figure 7 It is a schematic diagram of a model training scenario provided by an embodiment of the present application;
[0083] Figure 8 It is a schematic diagram of a model training scenario provided by an embodiment of the present application;
[0084] Figure 9 It is a schematic flowchart of a process for training a filtering model provided by an embodiment of the present application;
[0085] Figure 10 It is another schematic diagram of a model training scenario provided by an embodiment of the present application;
[0086] Figure 11 It is a schematic flowchart of a data processing method provided by an embodiment of the present application;
[0087] Figure 12 It is a schematic structural diagram of a data processing device provided by an embodiment of the present application;
[0088] Figure 13 It is a schematic structural diagram of a data processing device provided by an embodiment of the present application;
[0089] Figure 14 It is a schematic structural diagram of a computer device provided by an embodiment of the present application. Detailed implementation manners
[0090] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts shall fall within the protection scope of the present application.
[0091] Please refer to Figure 1 , Figure 1 which is a schematic structural diagram of a network architecture provided by an embodiment of the present application. As Figure 1 shown, the network architecture may include a service server 1000 and a cluster of terminal devices. The cluster of terminal devices may include one or more terminal devices, and the number of terminal devices will not be limited here. As Figure 1 shown, the multiple terminal devices may specifically include terminal device 100a, terminal device 100b, terminal device 100c, …, terminal device 100n. As Figure 1 shown, terminal device 100a, terminal device 100b, terminal device 100c, …, terminal device 100n may be respectively network-connected to the above-mentioned service server 1000, so that each terminal device can perform data interaction with the service server 1000 through this network connection. Among them, the network connection here does not limit the connection method, and it can be directly or indirectly connected through a wired communication method, or can be directly or indirectly connected through a wireless communication method, or can also be connected through other methods, and the present application does not make any restrictions here.
[0092] Each terminal device can be integrally installed with a target application. When the target application runs on each terminal device, it can perform data interaction with the service server 1000 as shown above Figure 1 . Among them, the target application may include applications with functions of displaying data information such as text, images, audio, and video. Among them, the application may include social applications, multimedia applications (for example, video applications), entertainment applications (for example, game applications), educational applications, live broadcast applications, etc. that have video encoding functions. Of course, the application may also be other applications with functions of displaying data information and video encoding functions, and examples will not be given one by one here. Among them, the application may be an independent application or an embedded sub-application integrated in a certain application (for example, social applications, educational applications, and multimedia applications, etc.), and no limitation is made here.
[0093] For ease of understanding, an embodiment of the present application may select one terminal device from the multiple terminal devices shown in Figure 1 as the target terminal device. For example, an embodiment of the present application may use the terminal device 100a shown in Figure 1 as the target terminal device, and the target terminal device may be integrally installed with a target application having a video encoding function. At this time, the target terminal device can realize data interaction with the service server 1000 through the service data platform corresponding to the application client.
[0094] It should be understood that the computer devices with video encoding functions in the embodiments of the present application (e.g., terminal device 100a, service server 1000) can, through cloud technology, implement data encoding and data transmission of multimedia data (e.g., video data). Among them, cloud technology refers to a hosting technology that unifies a series of resources such as hardware, software, and networks within a wide area network or a local area network to achieve data computing, storage, processing, and sharing.
[0095] Cloud technology can be the general term for network technology, information technology, integration technology, management platform technology, application technology, etc. based on the cloud computing business model. It can form a resource pool, be used as needed, and is flexible and convenient. Cloud computing technology will become an important support. The back-end services of the technical network system require a large amount of computing and storage resources, such as video websites, picture websites, and more portal websites. With the highly developed application of the Internet industry, in the future, each item may have its own identification mark and needs to be transmitted to the back-end system for logical processing. Data at different levels will be processed separately, and various industry data requires a powerful system back-end support, which can only be achieved through cloud computing.
[0096] For example, the data processing method provided in the embodiments of the present application can be applied to scenarios with high resolution and high frame rate such as video viewing scenarios, video call scenarios, video transmission scenarios, cloud conference scenarios, and live broadcast scenarios. Among them, cloud conference is an efficient, convenient, and low-cost conference form based on cloud computing technology. Users only need to perform simple and easy operations through the Internet interface to quickly and efficiently synchronously share voice, data files, and videos with teams and users everywhere, and complex technologies such as data transmission and processing during the conference are operated by cloud conference service providers to help users. Currently, cloud conferences mainly focus on service contents with the SaaS (Software as a Service) model as the main body, including service forms such as telephone, network, and video. A video conference based on cloud computing is called a cloud conference. In the era of cloud conferences, the transmission, processing, and storage of data are all processed by the computer resources of video conference manufacturers. Users no longer need to purchase expensive hardware and install cumbersome software at all. They only need to open a browser and log in to the corresponding interface to conduct efficient remote conferences. The cloud conference system supports dynamic cluster deployment of multiple servers and provides multiple high-performance servers, greatly improving the stability, security, and usability of conferences. In recent years, video conferences have been welcomed by many users because they can significantly improve communication efficiency, continuously reduce communication costs, and bring about an upgrade in internal management level, and have been widely applied in various fields such as transportation, transportation, finance, operators, education, and enterprises. There is no doubt that after video conferences use cloud computing, they have stronger attractiveness in terms of convenience, speed, and ease of use, and will surely stimulate the arrival of a new upsurge in video conference applications.
[0097] It should be understood that a computer device with video encoding capabilities can encode video data through a video encoder to obtain a video bitstream corresponding to the video data, thereby improving the transmission efficiency of the video data. Among them, the video encoder can be an AV1 video encoder, an H.266 video encoder, an AVS3 video encoder, etc., and no further examples will be given here. Among them, the video encoder needs to conform to the corresponding video coding compression standard. For example, the video compression standard of the AV1 video encoder is the first-generation video coding standard developed by the Alliance for Open Media (AOM).
[0098] To facilitate understanding of the process of the video encoder encoding video data, the following will elaborate on the specific process of video encoding. The process of video encoding can at least include the following steps 1 - step 5:
[0099] Step 1: Perform block partitioning (block partition structure) on the video frame. It can be understood that after inputting the video data into the video encoder, the video encoder can encode each video frame of the video data. For a certain video frame, it can be divided into several non-overlapping processing units according to the size of the video frame, and each processing unit will perform similar compression operations. Among them, this processing unit can be called a Coding Tree Unit (CTU) or a Largest Code Unit (LCU). For each CTU, it can be further divided more finely to obtain one or more basic coding units, and this unit can be called a Coding Unit (CU). Among them, each CU is the most basic element in a coding process. The following steps 2 - step 4 can describe various coding methods that may be adopted for each CU.
[0100] Step 2: Perform predictive coding on the CU. It can be understood that after dividing a video frame to obtain one or more coding units (i.e., CUs), predictive coding can be performed on each coding unit. Among them, the prediction methods for coding units can include intra-frame prediction (where the predicted signal comes entirely from the already encoded and reconstructed regions within the same image) and inter-frame prediction (where the predicted signal comes from other images that have been encoded and are different from the current image (i.e., other video frames, which can be called reference images or reference frames)). It should be understood that after the original video signal (which can be understood as the original video frame) is predicted by the selected reconstructed video signal (which can be a video frame that has undergone decoding and reconstruction processing at the decoding end), a residual video signal (i.e., the residual value) can be obtained. The encoding end in video coding applications can determine the most suitable one among many possible predictive coding methods for the current CU and inform the decoding end in the video coding application of the predictive coding method for the current CU, so that the decoding end can perform decoding and reconstruction processing on the current video frame to generate a reconstructed frame (this reconstructed frame can be used as a reference frame when the encoding end performs predictive coding on an original video frame).
[0101] Step 3: Perform transform coding and quantization on the residual video signal: The obtained residual video signal can be subjected to transform processing (such as discrete cosine transform (DCT)). Through the transform processing, the residual video signal can be converted into the transform domain (which can be called transform coefficients). Subsequently, the signal in the transform domain can be further subjected to a lossy quantization operation, losing certain information, so that the quantized signal is conducive to compressed representation. Among some video coding standards, one or more transform methods can be selected for the transform processing. Therefore, the encoding end also needs to select one of the transforms for the current coding CU and inform the decoding end. The fineness of quantization is usually determined by the quantization parameter (QP). A larger QP value means that a larger range of coefficients will be quantized to the same output, so it usually brings greater distortion and a lower bit rate; on the contrary, a smaller QP value means that a smaller range of coefficients will be quantized to the same output, so it usually brings less distortion and a corresponding higher bit rate.
[0102] Step 4: Perform entropy coding or statistical coding on the quantized signal: It can be understood that the quantized transform domain signal obtained above can be statistically compressed and coded according to the frequencies of each value, and finally a binary (including value 0 or value 1) compressed bitstream is output. At the same time, other information generated by the coding, such as the selected mode, motion vector, etc., also needs to be entropy coded to reduce the bit rate. Among them, statistical coding is a lossless coding method that can effectively reduce the bit rate required to represent the same signal. Common statistical coding methods include variable length coding (VLC) or context-adaptive binary arithmetic coding (CABAC).
[0103] Step 5: Loop filtering operation: It can be understood that for the already encoded image, inverse quantization, inverse transformation, and prediction compensation operations can be performed on it (which can be understood as the reverse operations of the above steps 2 to 4), and thus a reconstructed decoded image can be obtained. Among them, compared with the original image, due to the influence of quantization, some information will be different from the original image, resulting in distortion. The distortion caused by quantization can be effectively reduced by performing filtering operations on the reconstructed image, such as deblocking filter (DBF), SAO (Sample Adaptive Offset), or ALF (Adaptive Loop Filter), etc. The reconstructed image after filtering will be used as a reference for encoding a subsequent video frame and for predicting future signals. Therefore, the above filtering operation is also called loop filtering, that is, the filtering operation within the coding loop.
[0104] It can be understood that to improve the video coding efficiency and the filtering performance of the filters deployed in the video encoder, the present application proposes a method for iteratively training the filters (which can also be called filtering models) and iteratively updating the video encoder (which can be called video coding applications). The specific method can be seen in the description of the corresponding embodiments Figure 2 described later.
[0105] It can be understood that the method provided by the embodiments of the present application can be executed by a computer device, which includes but is not limited to a terminal device or a service server. Among them, the service server can be an independent physical server, or a server cluster or a distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms.
[0106] Optionally, it can be understood that the above computer device (such as the above service server 1000, terminal device 100a, terminal device 100b, etc.) can be a node in a distributed system. Among them, the distributed system can be a blockchain system, and the blockchain system can be a distributed system formed by connecting the multiple nodes in the form of network communication. Among them, the nodes can form a peer-to-peer (P2P) network, and the P2P protocol is an application layer protocol running on top of the Transmission Control Protocol (TCP). In the distributed system, any form of computer device, such as a service server, a terminal device and other electronic devices, can become a node in the blockchain system by joining the peer-to-peer network. For the convenience of understanding, the concept of blockchain will be described below: Blockchain is a new application mode of computer technologies such as distributed data storage, peer-to-peer transmission, consensus mechanism, and encryption algorithm, mainly used to sort data in chronological order, encrypt it into a ledger, make it impossible to be tampered with and forged, and at the same time, data verification, storage, and update can be performed. When the computer device is a blockchain node, due to the tamper-proof and anti-counterfeiting characteristics of the blockchain, the data in the present application (such as video data, encoded video data, relevant parameters during the encoding process, etc.) can have authenticity and security, so that the results obtained after performing relevant data processing based on these data are more reliable.
[0107] Further, please refer to Figure 2 , Figure 2 is a schematic flowchart of a data processing method provided by the embodiments of the present application. Among them, the method can be executed by a terminal device (such as any terminal device in the terminal device cluster corresponding to the above Figure 1 embodiment, such as terminal device 100a); the method can also be executed by a service server (such as the service server 1000 in the corresponding embodiment of the above Figure 1 embodiment); the method can also be jointly executed by the terminal device and the service server. Taking the method being executed by the service server as an example, as Figure 2As shown, the method process may at least include the following steps S101 to step S103:
[0108] Step S101: Input the sample video data into the video coding application including the first filtering model updated in the k-th deployment, and generate first training data through the video coding application updated in the k-th deployment and the sample video data; the first training data includes the sample original video frames as training labels and the first sample reconstruction frames to be filtered corresponding to the sample original video frames; the first sample reconstruction frames to be filtered refer to the reconstruction frames that are not filtered by the first filtering model during the process of reconstructing the sample original video frames through the video coding application updated in the k-th deployment; the sample original video frames are the video frames in the sample video data; k is a positive integer.
[0109] In this application, the sample video data may refer to the video data used to train the filtering model. A computer device (e.g., a terminal device) can obtain the video data collected by an image collector (e.g., the camera of the terminal device) in a data transmission scenario, and this video data can be used as the sample video data. The terminal device can input the sample video data into the video coding application. In other words, when the video coding application is deployed to the computer device, the computer device has the video coding function. The computer device can collect video data through the image collector, and the video coding application can also obtain this video data. It can be understood that the video data here can also be any video data that needs to be encoded in other scenarios. For example, the video data can be the video data collected by the camera of the terminal device in an audio-video call scenario, the video data saved in the album of the terminal device, or the video data downloaded by the terminal device from the network. This is not limited here.
[0110] It can be understood that a video encoding application can be used to perform encoding processing on video data (e.g., sample video data) to obtain a video bitstream corresponding to the video data. When the video encoding application encodes the video data, it usually encodes each video frame in the video data. For example, the video encoding application can obtain a video frame to be encoded from the sample video data, and this video frame can be called a target video frame. The video encoding application can perform image block partitioning on the target video frame to obtain one or more image blocks of the target video frame, and can obtain a coding unit to be encoded from the one or more image blocks and perform prediction processing on the coding unit to be encoded. Among them, when encoding the target video frame, the frame type to which the target video frame belongs can be obtained, and when performing prediction processing on the coding unit to be encoded, the coding unit to be encoded can be predicted based on the frame type to which the target video frame belongs. The frame type can include an intra prediction type and an inter prediction type. For example, the frame type to which an intra picture (I frame for short) belongs is the intra prediction type, and the frame types to which a bi-directional interpolated prediction frame (B frame for short) and a predictive-frame (P frame for short) belong are the inter prediction types. When the frame type of the target video frame is the intra prediction type, when predicting the coding unit to be encoded, only the already encoded and reconstructed area within the target video frame needs to be referred to (it should be noted that the intra prediction type here can refer to the full intra prediction type, that is, when predicting all coding units to be encoded within the target video frame, only the area within the target video frame is referred to); when the frame type of the target video frame is the inter prediction type, when predicting the coding unit to be encoded, other already encoded video frames different from the target video frame (which can be called reference video frames) need to be referred to. Among them, when selecting the reference video frame of the target video frame, it can be selected according to the frame type of the target video frame. For example, when the target video frame is a bi-directional interpolated prediction frame, the reference video frames for this target video frame can be the previous frame and the next frame of this target video frame; when the target video frame is a predictive-frame, the reference video frame for this target video frame can be the previous frame of this target video frame. Among them, the specific implementation steps for encoding the video frame can be referred to the description in the corresponding embodiment above Figure 1 as described in the corresponding embodiment.
[0111] For ease of understanding the encoding dependency relationship between video frames, please also refer to Figure 3 , Figure 3 which is a schematic diagram of an encoding reference relationship provided by an embodiment of the present application. As shown in Figure 3In the encoding configuration shown, the encoding order of video frames is the same as the display order. The arrows point to the reference frames, and the numbers are used to represent the encoding order. That is, the 0th frame (i.e., the I-frame) can be encoded first, and then the B-frame of the 1st frame can be encoded. When encoding the B-frame of the 1st frame, the I-frame of the 0th frame that has been encoded and reconstructed needs to be referred to. Subsequently, the B-frame of the 2nd frame can be encoded. When encoding the B-frame of the 2nd frame, the B-frame of the 1st frame that has been encoded and reconstructed needs to be referred to. And so on, until the encoding process of all video frames is completed.
[0112] For ease of understanding the encoding dependency relationship between video frames, please also refer to Figure 4 , Figure 4 which is a schematic diagram of an encoding reference relationship provided by an embodiment of the present application. In the encoding configuration as shown in Figure 4 the encoding order of video frames is not the same as the display order. The arrows point to the reference frames, and the numbers are used to represent the encoding order. That is, the 0th frame (i.e., the I-frame) can be encoded first, and then the B-frame of the 1st frame can be encoded. When encoding the B-frame of the 1st frame, the I-frame of the 0th frame that has been encoded and reconstructed needs to be referred to. Subsequently, the B-frame of the 2nd frame can be encoded. When encoding the B-frame of the 2nd frame, the arrows point to the 0th frame and the 1st frame, so the 1st frame and the 0th frame that have been encoded and reconstructed need to be referred to. And so on, until the encoding process of all video frames is completed.
[0113] It can be understood that, as described above, when encoding video frames according to the encoding order of video frames in the encoding configuration, for a certain video frame, an encoding process as shown in Figure 1 can be adopted. The reference frame is determined according to the frame type of this frame, prediction processing is performed based on the reference frame to obtain a prediction unit, and then the residual value between the prediction unit and the unit to be encoded is determined. Based on the prediction unit and the residual value, the reconstructed frame corresponding to this video frame can be determined (for example, the prediction unit and the residual value are added and then filtered to obtain the reconstructed frame). This reconstructed frame can then enter the reference frame queue to provide reference data for the subsequent encoding of video frames. It should be understood that a prediction unit can be obtained by performing prediction processing on the encoding unit of a video frame, the residual value can be determined based on the prediction unit, and subsequently, a reconstructed image (or called the reconstructed frame before filtering) can be determined based on the prediction unit and the residual value. By filtering the reconstructed image, the filtered reconstructed frame corresponding to this video frame can be obtained.
[0114] It can be understood that after the sample video data is input into the video encoding application, the video encoding application can encode and reconstruct each video frame to obtain a reconstructed frame corresponding to each video frame. In this application, during the encoding and reconstruction process, the reconstructed frame that has not undergone filtering processing (which can be called the sample reconstructed frame to be filtered) can be obtained. Each sample original video frame and its corresponding sample reconstructed frame to be filtered can be formed into a set of training data pairs, so that multiple sets of training data pairs can be obtained. These multiple sets of training data pairs can form a training data set, and this training data set can be used to train the filtering model in the video encoding application (that is, the model used to filter the reconstructed frame after inverse transformation), so that the filtering model of the video encoding application can have higher filtering performance.
[0115] In this application, when the filtering model in the video encoding application is trained and in a model convergence state, the trained filtering model can be deployed and updated (for example, replaced) to the filtering model deployed in the video encoding application. And a process of deploying and updating the filtering model deployed in the video encoding application can be understood as a deployment update of the video encoding application. This application can continuously generate new training data through the updated video encoding application, then use the new training data to repeatedly train the filtering model, and then perform another deployment update on the video encoding application based on the retrained filtering model until the filtering performance of the video encoding application meets the filtering quality requirement conditions. For example, taking the video encoding application that has not been deployed and updated as an example, the initial filtering model that has not been trained is included in this video encoding application that has not been deployed and updated. The sample video data can be input into the video encoding application, and the training data corresponding to the sample video data can be generated through this video encoding application. The initial filtering model can be trained and adjusted through this training data to obtain a trained and adjusted filtering model; subsequently, the trained and adjusted filtering model can be updated and deployed into the video encoding application to obtain the video application after the first deployment update; subsequently, new training data can be generated again through the video application after the first deployment update, and the filtering model can be retrained and adjusted based on the new training data to obtain a new trained filtering model. Then, the trained and adjusted filtering model can be updated and deployed into the video encoding application after the first deployment update to obtain the video encoding application after the second deployment update until the filtering performance of the video encoding application meets the requirement conditions.
[0116] That is, taking the current video encoding application as the video encoding application updated by the k-th deployment (deployed with the first filtering processing model), and the video encoding application updated by the k-th deployment has not yet met the filtering quality requirement conditions as an example. At this time, the sample video data can be input into the video encoding application updated by the k-th deployment, and the training data (which can be called the first training data) can be generated through the video encoding application updated by the k-th deployment. The first training data includes the sample original video frame as the training label, and the sample reconstructed frame to be filtered corresponding to the sample original video frame (which can be called the first sample reconstructed frame to be filtered). The first training data can be used to train and adjust the first filtering model. Among them, as can be seen from the above, the process of reconstructing the sample original video frame can include obtaining a prediction unit through prediction processing, determining a residual value based on the prediction unit, determining a reconstructed image (or called a reconstructed frame) based on the prediction unit and the residual value, and filtering the reconstructed image. Then, the sample reconstructed frame to be filtered can be understood as the reconstructed frame that has not been filtered during the process of reconstructing the sample original video frame. A set of training data pairs can be formed according to a sample original video frame and its corresponding sample reconstructed frame to be filtered, and the first training data can be formed according to multiple sets of training data pairs.
[0117] Step S102: Based on the sample original video frame and the first sample reconstructed frame to be filtered, train the filtering model to be trained in the video encoding application updated by the k-th deployment to obtain a second filtering model in a training convergence state, and update and deploy the second filtering model in the video encoding application updated by the k-th deployment to obtain the video encoding application updated by the (k + 1)-th deployment.
[0118] In this application, as can be seen from the above, when encoding a video frame, the frame type of each video frame can be determined first, and then encoding processing such as prediction processing can be performed on the encoding unit to be encoded based on the frame type. The filtering model in the video encoding application of this application can be used to filter the reconstructed frame belonging to the intra-frame prediction type (for example, the reconstructed frame corresponding to the I frame), and can also be used to filter the reconstructed frame belonging to the inter-frame prediction type (the reconstructed frame corresponding to the B frame). That is to say, the above first filtering model can refer to the filtering model commonly corresponding to the I frame and the B frame. In this case, the filtering model to be trained can refer to the first filtering model, and the first filtering model can be trained based on the first training data to obtain a second filtering model in a training convergence state. Subsequently, the first filtering model in the video encoding application updated by the k-th deployment can be replaced with the second filtering model, so that the video encoding application updated by the (k + 1)-th deployment can be obtained.
[0119] Among them, for training the filter model to be trained in the video encoding application updated in the k-th deployment based on the original video frame of the sample and the first sample frame to be filtered and reconstructed, the specific implementation of obtaining the second filter model in the training convergence state can be as follows: The first sample frame to be filtered and reconstructed can be input into the first filter model, and the sample filter reconstruction frame corresponding to the first sample frame to be filtered and reconstructed can be output through the first filter model. Subsequently, the error value between the sample filter reconstruction frame and the original video frame of the sample can be determined, and the model parameters of the first filter model can be adjusted through the error value to obtain the first filter model with the adjusted model parameters. When the first filter model with the adjusted model parameters meets the model convergence condition, the first filter model with the adjusted model parameters can be determined as the second filter model in the training convergence state. It should be noted that the model convergence condition here can refer to that the number of training adjustments reaches the preset number of adjustments, or it can refer to that the image quality of the filtered image output by the adjusted first filter model meets the quality requirement condition. For example, taking the model convergence condition as that the number of training adjustments reaches the preset number of adjustments as an example, assuming that the preset number of adjustments is 10 times, after adjusting the model parameters once based on the error value generated for the first time, the first sample frame to be filtered and reconstructed can be input into the first filter model adjusted once again, the sample filter reconstruction frame can be output again, a new error value can be generated, and the model parameters can be adjusted again until the model parameters are adjusted 10 times. At this time, the first filter model adjusted 10 times can be determined to meet the model convergence condition (that is, in the training convergence state).
[0120] Among them, the specific implementation of determining the error value between the sample filter reconstruction frame and the original video frame of the sample can be as follows: The loss function corresponding to the video encoding application updated in the k-th deployment can be obtained. Subsequently, the original image quality corresponding to the original video frame of the sample can be obtained based on the loss function, and the original image quality can be used as the image quality label. Subsequently, the filtered image quality corresponding to the sample filter reconstruction frame can be obtained based on the loss function, and the error value between the sample filtered video frame and the original video frame of the sample can be determined through the loss function, the image quality label, and the filtered image quality.
[0121] It can be understood that in this application, the iterative training of the filtering model is performed independently each time. That is to say, after generating the training data, the loss function for training the filtering model based on the training data can be different independent loss functions. For example, when training the filtering model based on the training data generated for the first time, the absolute value loss function L1 can be used to determine the error value, and the model parameters can be adjusted based on the error value; when training the filtering model based on the training data generated for the second time, the squared error loss function L2 can be used to determine the error value, and the model parameters can be adjusted based on the error value; when training the filtering model based on the training data generated for the third time, an error value can be determined first using the loss function L1, and then an error value can be determined using the loss function L2. The model parameters can be adjusted first based on the error value of L1, and then further adjusted based on the error value of L2 on the basis of the adjustment of L1. Of course, when training the filtering model with the subsequently generated training data, the error values of L1 and L2 can also be added together, and the model parameters can be adjusted based on the total error value obtained by the addition. The loss function for each iteration is not limited here.
[0122] Among them, for each loss function of this application, in addition to including the absolute value loss function L1 and the squared error loss function L2, it can also be any other loss function that can determine the error between the label and the output data. For example, the loss function can include the cross-entropy loss function, the mean squared error loss function, the logarithmic loss function, the exponential loss function, and so on. Taking the loss function corresponding to the video coding application updated in the (k + 1)-th deployment including the absolute value loss function and the squared error loss function as an example, the specific implementation method for determining the error value between the sample filtered reconstructed frame and the sample original video frame through the loss function, the image quality label, and the filtered image quality can be as follows: The first error value between the sample filtered reconstructed frame and the sample original video frame can be determined through the squared error loss function, the image quality label, and the filtered image quality; subsequently, the second error value between the sample filtered reconstructed frame and the sample original video frame can be determined through the absolute value loss function, the image quality label, and the filtered image quality; subsequently, the first error value and the second error value can be subjected to arithmetic processing (for example, addition arithmetic processing), and the result obtained by the arithmetic (for example, the addition result) can be determined as the error value between the sample filtered reconstructed frame and the sample original video frame.
[0123] Step S103: When the video coding application updated in the (k + 1)-th deployment meets the filtering quality requirement condition, determine the video coding application updated in the (k + 1)-th deployment as the target video coding application for video coding processing of video data.
[0124] In this application, after each deployment update of the video encoding application, the video encoding application can be detected to determine whether it meets the filtering quality requirement conditions. When it is determined that it meets the filtering quality requirement conditions, the video encoding application does not need to be iteratively updated and can be put into use (for example, as the target video encoding application to perform video encoding processing on subsequent video data). Taking the video encoding application with the (k + 1)-th deployment update as an example, the specific implementation for detecting the video encoding application can be as follows: The sample video data can be input into the video encoding application with the (k + 1)-th deployment update, and the filtered video data to be detected can be output through the video encoding application with the (k + 1)-th deployment update. Subsequently, the filtering video quality corresponding to the filtered video data to be detected and the original video quality corresponding to the sample video data can be obtained. The video encoding application with the (k + 1)-th deployment update can be detected based on the filtering video quality and the original video quality.
[0125] Among them, the specific implementation for detecting the video encoding application with the (k + 1)-th deployment update based on the filtering video quality and the original video quality can be as follows: The difference video quality between the filtering video quality and the original video quality can be determined. If the difference video quality is less than the difference quality threshold, it can be determined that the second video encoding application meets the filtering quality requirement conditions. If the difference video quality is greater than the difference quality threshold, it can be determined that the second video encoding application does not meet the filtering quality requirement conditions.
[0126] It can be understood that each filtered reconstruction frame (the reconstruction frame after being filtered by the second filtering model) output by the video encoding application with the (k + 1)-th deployment update can be obtained, and the filtering image quality of the filtered reconstruction frame can be compared with the original image quality of the corresponding original video frame. If the difference quality between the two is less than the preset threshold, it can be determined that after the second filtering model is deployed to the video encoding application, the video frames processed by the model meet the filtering requirements.
[0127] It can be understood that when the video encoding application with the (k + 1)-th deployment update does not meet the filtering quality requirement conditions, new training data can be generated again based on the video encoding application with the (k + 1)-th deployment update, and the second filtering model can be trained again based on the new training data to obtain a new filtering model (such as the third filtering model). Subsequently, the second filtering model deployed in the video encoding application with the (k + 1)-th deployment update can be updated to the third filtering model to obtain the video encoding application with the (k + 2)-th deployment update. At this time, the video encoding application with the (k + 3)-th deployment update can be detected again, and the application can be detected again until the video encoding application meets the filtering quality requirement conditions, so that the target video encoding application can be obtained.
[0128] To facilitate understanding of the specific processes of model training and application iterative update in this application, please also refer to Figure 5 , Figure 5 which is an architecture diagram of an iterative training system provided by an embodiment of this application. As Figure 5 shown, the system architecture may include a dataset generation module, a model training module, and an application integration module. For ease of understanding, each module will be described as follows:
[0129] Dataset generation module: mainly used to generate a training dataset (which can also be referred to as training data). It should be understood that through video coding applications, video data can be encoded and reconstructed. During the encoding and reconstruction process, the dataset generation module can use the frame to be filtered and reconstructed (i.e., the reconstructed frame that has not undergone filtering processing) and its corresponding original video frame as a pair of training data. Thus, a training data including multiple pairs of training data can be obtained.
[0130] Model training module: mainly used to train the filtering model. It should be understood that the training data generated by the above dataset generation module can be transmitted to the model training module. In the model training module, the filtering model can be trained based on the training data so that the filtering model can meet the training objectives.
[0131] Application integration module: mainly used to integrally deploy the trained filtering model into a video encoding application. It should be understood that the application integration module can deploy and integrate the above-mentioned trained filtering model into the video encoding application. At this time, video data can be input into the video encoding application integrated with the trained filtering model to detect whether the video encoding application can meet the filtering requirement conditions in actual application (that is, to detect whether the filtering performance of the application meets the preset requirement conditions). If it meets the requirements, the video encoding application can be put into use. If it does not meet the requirements, the dataset generation module can be used again to regenerate the training dataset based on the new video encoding application, and then retrain the filtering model based on the new training dataset, and then perform application integration and actual filtering processing until the video encoding application meets the filtering requirement conditions. That is to say, when the trained filtering model meets the training objective (that is, meets the filtering requirement conditions), after it is integrally deployed into the video encoding application, due to different reference relationships of the coded frames of the inter-frame prediction type (inter-frame prediction coded frames), the reconstructed frame before filtering corresponding to the inter-frame prediction coded frame will be inconsistent with the reconstructed frame before filtering in the training process, and the results of training and testing will be inconsistent. The filtering performance of the filtering model in the video encoding application will not be the expected effect. However, through the iterative training process of regenerating the training dataset with each newly deployed and updated video encoding application, retraining the filtering model, and then integrating it again into the video encoding application for actual filtering processing, the consistency between training and testing can be continuously improved until the filtering model in the video encoding application meets the filtering requirement conditions, thereby improving the encoding efficiency.
[0132] It should be noted that when the dataset generation module generates the training dataset, for the training dataset, software integrated with a filter with comparable performance can also be used to generate it. That is, for the current filtering model, software integrated with a filter with comparable performance to the current model can be used to generate the training data, and then the current filtering model can be trained based on the training data. For example, before the first iterative update of the application, a filter with comparable performance to the initial untrained filtering model (a traditional filter, such as a filter whose filtering parameters are determined based on manual experience) can be used to generate the training dataset, and this training dataset can be used to complete the model training process in the first iterative update (that is, use this training dataset to train the initial untrained filtering model).
[0133] In an embodiment of the present application, sample video data can be input into a video encoding application updated by the k-th deployment (such as the 1st, 2nd, 3rd,...), a video encoding application deployed with a first filtering model. Through the video encoding application updated by the k-th deployment, the first training data corresponding to the sample video data can be output; based on the first training data, the first filtering model can be retrained to obtain a second filtering model, and the second filtering model can be used to update the video encoding application updated by the k-th deployment again, so as to obtain a video encoding application updated by the (k + 1)-th deployment; at this time, the video encoding application updated by the (k + 1)-th deployment can be detected. When the video encoding application updated by the (k + 1)-th deployment meets the filtering quality requirement conditions, the video encoding application updated by the (k + 1)-th deployment can be determined as the target video encoding application for video encoding processing of video data. It should be understood that deploying a filtering model (such as the first filtering model) to a video encoding application is a deployment update of the video encoding application. After generating training data based on the video encoding application, the filtering model can be trained based on machine learning, and then the trained filtering model can be deployed and updated to the video encoding application to obtain an updated video encoding application; subsequently, the present application can continue to generate new training data based on the updated video encoding application, retrain the filtering model based on the new training data, and then deploy and update the trained filtering model to the video encoding application again to obtain an updated video encoding application again until the video encoding meets the filtering quality requirement conditions. That is to say, after deploying the filtering model to the video encoding application, the present application can iteratively update the video encoding application, repeat the training process of the filtering model by updating the training data, and continuously deploy and update the video encoding application, which can improve the consistency between the training effect and the test effect of the filtering model, improve the encoding efficiency, and can also improve the filtering performance of the filtering model deployed in the video encoding application without relying on manual experience, and reduce the distortion degree of the encoded video. In summary, the present application can improve the filtering performance, reduce the image distortion degree of the encoded video, and improve the encoding efficiency.
[0134] As can be seen from the above, when encoding a video frame, the frame type to which each video frame belongs can be determined first, and then encoding processing such as prediction processing of the encoding unit to be encoded can be performed based on the frame type. Since the prediction mode of an inter-frame prediction encoded frame (during the encoding process, each video frame can be called an encoded frame, and an inter-frame prediction encoded frame can be understood as an encoded frame whose frame type belongs to the inter-frame prediction type, such as a B frame) is an inter-frame prediction mode, it can refer to other reconstructed frames for prediction, and usually has a higher prediction accuracy; while the prediction mode of an intra-frame prediction encoded frame (which can be understood as an encoded frame whose frame type belongs to the intra-frame prediction type, such as an I frame) is an intra-frame prediction mode, and it refers to other regions of this frame, and usually the prediction accuracy is lower than that of an inter-frame prediction encoded frame. That is to say, for the filtering model, the image features before filtering of the intra-frame prediction encoded frame and the inter-frame prediction encoded frame (which can be understood as the features of the reconstructed frame (which can be called the reconstructed frame to be filtered) before the filtering process in the Figure 1 corresponding process) will be different, with obvious differences. Then, if the same filtering model is used for unified filtering processing, it will have a certain impact on the image quality after filtering. Then, in order to make the filtering effects of the intra-frame prediction encoded frame and the inter-frame prediction encoded frame better, that is, to improve the image quality after filtering, the present application can train different filtering models for the intra-frame prediction encoded frame and the inter-frame prediction encoded frame, and use different filtering models to perform filtering processing on the reconstructed frame to be filtered corresponding to the intra-frame prediction encoded frame and the reconstructed frame to be filtered corresponding to the inter-frame prediction encoded frame. For example, use an intra-frame filtering model to perform filtering processing on the reconstructed frame to be filtered corresponding to the intra-frame prediction encoded frame, and use an inter-frame filtering model to perform filtering processing on the reconstructed frame to be filtered corresponding to the inter-frame prediction encoded frame.
[0135] It can be understood that when different filtering models are used to filter the intra-predicted coded frame and the inter-predicted coded frame, in the video coding application of the present application, it may include an intra-filtering model for filtering the reconstructed frame belonging to the intra-prediction type (the reconstructed frame to be filtered corresponding to the intra-predicted coded frame, for example, the reconstructed frame corresponding to the I frame), and an inter-filtering model for filtering the reconstructed frame belonging to the inter-prediction type (the reconstructed frame to be filtered corresponding to the inter-predicted coded frame, for example, the reconstructed frame corresponding to the B frame). The present application can determine the model to be trained between the intra-filtering model and the inter-filtering model, and then train the model to be trained based on the training data. For example, in the first filtering model deployed in the video coding application updated in the k-th deployment, it includes a first intra-filtering model for filtering the reconstructed frame to be filtered belonging to the intra-prediction type (the first intra-filtering model is trained based on the initial intra-filtering model in the video coding application updated in the (k - 1)-th deployment), and the video coding application updated in the k-th deployment further includes an untrained initial inter-filtering model (the initial inter-filtering model is used to filter the reconstructed frame to be filtered belonging to the inter-prediction type). Taking this as an example, the present application can determine the model to be trained in the video coding application updated in the k-th deployment as the initial inter-filtering model and the first intra-filtering model, or can also determine the model to be trained only as the initial inter-filtering model.
[0136] For ease of understanding, please also refer to Figure 6 , Figure 6 which is a schematic flowchart of a process for training a filtering model provided by an embodiment of the present application. This process takes the model to be trained in the video coding application updated in the k-th deployment including the initial inter-filtering model and the first intra-filtering model as an example, and details the specific process of training the model to be trained in the video coding application updated in the k-th deployment based on the sample original video frame and the first sample reconstructed frame to be filtered, so as to obtain a second filtering model in a training convergence state. As Figure 6 shown, this process may at least include the following steps S501 - S505:
[0137] Step S501, obtain the reconstructed inter-frame to be filtered belonging to the inter-prediction type in the first sample reconstructed frame to be filtered, and use the sample original video frame corresponding to the reconstructed inter-frame to be filtered in the sample original video frame as the first video frame label.
[0138] Specifically, as can be seen from the above, the training data generated by the video coding application may include each sample original video frame and its corresponding sample filtered and reconstructed frame. During the process of encoding and reconstructing each sample original video frame, it is necessary to determine the frame type to which each sample original video frame belongs, and then perform prediction processing based on the frame type. Therefore, the sample filtered and reconstructed frames obtained also correspond to different frame types. For example, during the process of encoding and reconstructing the sample original video frame a, if the frame type of the encoded frame of the sample original video frame a is determined to be the intra-prediction type, then the frame type of the reconstructed frame corresponding to the sample original video frame a can also be understood as the intra-prediction type. In this application, for the intra-filtering model, video frames with the intra-prediction type frame type can be used for training, and for the inter-filtering model, video frames with the inter-prediction type frame type can be used for training. Then, when the filtering model to be trained in the video coding application updated in the k-th deployment includes the initial inter-filtering model and the first intra-filtering model, for the initial inter-filtering model, in the first training data, a filtered and reconstructed frame belonging to the inter-prediction type (which can be called the filtered inter-frame reconstructed frame) can be obtained, and the sample original video frame corresponding to the filtered inter-frame reconstructed frame in the sample original video frame is used as the first video frame label.
[0139] Step S502: Obtain a first filtered intra-frame reconstructed frame belonging to the intra-prediction type from the first sample filtered and reconstructed frame, and use the sample original video frame corresponding to the first filtered intra-frame reconstructed frame in the sample original video frame as the second video frame label.
[0140] Specifically, similarly, for the first intra-filtering model, in the first training data, a reconstructed frame belonging to the intra-prediction type (which can be called the first filtered intra-frame reconstructed frame) can also be obtained, and the sample original video frame corresponding to the first filtered intra-frame reconstructed frame in the sample original video frame is used as the second video frame label.
[0141] Step S503: Train the initial inter-filtering model based on the filtered inter-frame reconstructed frame and the first video frame label to obtain a first inter-filtering model in a training convergence state.
[0142] Specifically, through the filtered inter-frame reconstructed frame and its corresponding first video frame label, the initial inter-filtering model can be trained and adjusted to obtain an inter-filtering model in a training convergence state (which can be called the first inter-filtering model).
[0143] Step S504: Train the first intra-filtering model based on the filtered intra-frame reconstructed frame and the second video frame label to obtain a second intra-filtering model in a training convergence state.
[0144] Specifically, through the first in-loop reconstructed frame within the first frame to be filtered and its corresponding second video frame label, the in-loop filtering model for the first frame can be trained and adjusted to obtain an in-loop filtering model in a training convergence state (which can be called the second in-loop filtering model).
[0145] Step S505: Determine the first inter-filtering model and the second in-loop filtering model as the second filtering model.
[0146] Specifically, the second in-loop filtering model and the first inter-filtering model can be used as the second filtering model. For the specific training process of each filtering model, reference can be made to the description of the training process for the filtering model in the corresponding embodiment above Figure 3 and will not be elaborated here.
[0147] It should be noted that in the corresponding embodiment above Figure 2 when the first filtering model is the filtering model commonly corresponding to the inter-frame prediction coded frame and the in-frame prediction coded frame, for the k-th deployment update of the video coding application to update and deploy the second filtering model to obtain the (k + 1)-th deployment update of the video coding application, the specific process can be to simply replace the first filtering model with the second filtering model. Then in this embodiment, when the filtering model to be trained in the k-th deployment update of the video coding application includes the initial inter-filtering model and the first in-loop filtering model, for updating and deploying the second filtering model in the k-th deployment update of the video coding application to obtain the (k + 1)-th deployment update of the video coding application, the specific method can be: replacing and updating the first in-loop filtering model in the k-th deployment update of the video coding application with the second in-loop filtering model; replacing and updating the initial inter-filtering model in the k-th deployment update of the video coding application with the first inter-filtering model. That is to say, the two trained models are updated and replaced separately.
[0148] Optionally, it can be understood that when the initial inter-frame filtering model and the first intra-frame filtering model are included in the video encoding application updated in the k-th deployment, and the initial inter-frame filtering model is an untrained filtering model, the model to be trained in the video encoding application updated in the k-th deployment can be determined as only the initial inter-frame filtering model. That is to say, the first intra-frame filtering model does not need to be retrained first. After training and adjusting the initial inter-frame filtering model through the filtered inter-frame reconstructed frame and the first video frame label in the above first training data to obtain the first inter-frame filtering model, the first inter-frame filtering model can be determined as the second filtering model. Subsequently, the initial inter-frame filtering model in the video encoding application updated in the k-th deployment can be replaced and updated with the first inter-frame filtering model to obtain the video encoding application updated in the (k + 1)-th deployment. In the subsequent iterative update process of the video encoding application, new training data can be generated based on the video encoding application updated in the (k + 1)-th deployment, and then the first intra-frame filtering model and the first inter-frame filtering model can be trained together based on the new training data, and then the video encoding application updated in the (k + 1)-th deployment can be deployed and updated based on the trained intra-frame or inter-frame filtering model.
[0149] It can be understood that for the case where an intra-frame prediction coded frame (i.e., a coded frame or video frame belonging to the intra-frame prediction type, such as an I frame) has a type of model (i.e., an intra-frame filtering model, such as an I frame model), and an inter-frame prediction coded frame (i.e., a coded frame or video frame belonging to the inter-frame prediction type, such as a B frame) has a type of model, for the training of the filtering model, the Figure 6 shown process can be adopted. For ease of understanding, please refer to Figure 7 , Figure 7 which is a schematic diagram of a model training scenario provided by an embodiment of the present application. As Figure 7 shown, taking the intra-frame filtering model including an I frame model and the inter-frame filtering model including a B frame model as an example, in this case, the process of iterative training of the model can include the following steps (1) - step (5):
[0150] Step (1): When the initial untrained I frame model and B frame model are included in the video encoding application, the sample video data can be input into the video encoding application, and the training data can be output through the video encoding application.
[0151] Step (2): Use the initial I-frame model in the video encoding application as the model to be trained, and train the I-frame model based on the training data (train the I-frame model based on the original I-frame video frames and the I-frame frames to be filtered and reconstructed in the training data), to obtain an I-frame model in a training convergence state (i.e., the trained I-frame model). Subsequently, the initial I-frame model in the video encoding application can be replaced with the trained I-frame model, while keeping the initial B-frame model unchanged, thus obtaining the video encoding application with the first deployment update.
[0152] Step (3): Input the sample video data into the video encoding application with the first deployment update, and output new training data through this video encoding application with the first deployment update.
[0153] Step (4): Use the initial B-frame model in the video encoding application with the first deployment update as the model to be trained, and train the B-frame model based on the new training data (train the B-frame model based on the original B-frame video frames and the B-frame frames to be filtered and reconstructed in the training data), to obtain a B-frame model in a training convergence state (i.e., the trained B-frame model). The initial B-frame model in the video encoding application with the first deployment update can be replaced with the trained B-frame model, obtaining the video encoding application with the second deployment update. It should be noted that when training the B-frame model based on the new training data here, the trained I-frame model can also be trained simultaneously, to obtain the trained B-frame model and the re-trained I-frame model, and then update and replace the I-frame model and the B-frame model in the video encoding application respectively.
[0154] Step (5): Input the sample video data into the video encoding application with the second deployment update, and the video encoding application with the second deployment update can output new training data again. At this time, the I-frame model and the B-frame model in the video encoding application with the first deployment update can be re-trained based on the new training data, to obtain the trained I-frame model and B-frame model, and then replace and update the models in the video encoding application with the second deployment update respectively with the trained I-frame model and B-frame model, obtaining the video encoding application with the third deployment update.
[0155] It should be noted that after each deployment update of the video encoding application, its application can be detected to determine whether it meets the filtering requirement conditions. When the filtering requirement conditions are not met, a new training data set is generated, and model training is carried out again based on the new training data set, and the video encoding application is deployed and updated again based on the trained model until the video encoding application meets the filtering requirement conditions.
[0156] Optionally, in a feasible embodiment, the filter model to be trained in the video coding application updated in the k-th deployment may include a first inter-frame filter model and a first intra-frame filter model. The first inter-frame filter model is trained based on the initial inter-frame filter model in the video coding application updated in the (k - 1)-th deployment. That is to say, in this application, the initial inter-frame filter model and the initial intra-frame filter model can be trained together. After obtaining the first inter-frame filter model and the first intra-frame filter model, the initial inter-frame filter model and the initial intra-frame filter model in the video coding application updated in the (k - 1)-th deployment are respectively updated and replaced to obtain the video coding application updated in the k-th deployment, which includes the first inter-frame filter model and the first intra-frame filter model. For ease of understanding, please refer to Figure 8 , Figure 8 which is a schematic diagram of a model training scenario provided by an embodiment of this application. As Figure 8 shown, taking the intra-frame filter model including the I-frame model and the inter-frame filter model including the B-frame model as an example, in this case, the process of iterative training of the model may include the following steps (81)-(84):
[0157] Step (81): When the video coding application includes an initial untrained I-frame model and B-frame model, the sample video data can be input into the video coding application, and the training data is output through the video coding application.
[0158] Step (82): Taking the initial I-frame model in the video coding application as the model to be trained, the I-frame model is trained based on the training data (the I-frame original video frame and the I-frame to-be-filtered reconstructed frame in the training data are used to train the I-frame model), and at the same time, the B-frame model is trained based on the training data (the B-frame original video frame and the B-frame to-be-filtered reconstructed frame in the training data are used to train the B-frame model) to obtain a B-frame model in a training convergence state (i.e., the trained B-frame model). Subsequently, the initial I-frame model in the video coding application can be replaced with the trained I-frame model, and the initial B-frame model can be replaced with the trained B-frame model. Thus, the video coding application updated in the first deployment can be obtained.
[0159] Step (83): Input the sample video data into the video coding application updated in the first deployment, and new training data is output through the video coding application updated in the first deployment.
[0160] Step (84): Use the I-frame model and the B-frame model in the video encoding application updated in the first deployment as the models to be trained (or only use the B-frame model as the model to be trained). Based on the new training data, train the I-frame model and the B-frame model to obtain the I-frame model and the B-frame model in a training convergence state. The models in the video encoding application updated in the first deployment can be replaced and updated accordingly to obtain the video encoding application updated in the second deployment.
[0161] After each deployment update of the video encoding application, its application can be detected to determine whether it meets the filtering requirement conditions. When the filtering requirement conditions are not met, a new training data set is generated, and model training is performed again based on the new training data set. Then, the video encoding application is deployed and updated again based on the trained model until the video encoding application meets the filtering requirement conditions.
[0162] It should be understood that as described above, when encoding a video frame, the frame type of each video frame can be determined first, and then encoding processing such as prediction processing of the unit to be encoded can be performed based on the frame type. For an inter-frame prediction encoded frame (such as a B frame), since it needs other reconstructed frames as references for prediction encoding, the coupling relationship of the reference relationship of the inter-frame prediction encoded frame affects the consistency between training and actual testing. That is, even if the trained filtering model is in a model convergence state and reaches the training goal, after it is deployed to the video encoding application, due to the need for forward reference or bidirectional reference of the inter-frame prediction encoded frame, during the actual testing process, the filtered reconstructed frame output by the video encoding application is not the expected filtering effect. That is, the results of training and testing are not equivalent. For example, for a certain B frame, its corresponding reference frame is the filtered reconstructed frame of the previous I frame. After deploying the trained filtering model of the I frame and the filtering model of the B frame to the video encoding application, the quality of the filtered reconstructed frame of the I frame will be improved. Then, due to the improvement in the quality of the filtered reconstructed frame of the I frame, the prediction result of this B frame will also be improved, that is, the features before filtering of the B frame will be improved. Then, after being processed by the filtering model of the B frame, the filtering effect of the reconstructed frame of this B frame is not the filtering effect during the training process. That is, the training process and the testing process are not equivalent. Through the iterative training method provided by the embodiments of the present application, by updating the training data set to repeatedly train the B-frame model or the I-frame model, the consistency between the training and testing of the B frame can be improved, thereby improving the filtering performance of the video encoding application while improving the encoding efficiency and reducing the distortion degree of the encoded video.
[0163] In the embodiments of the present application, by updating the training data set to repeatedly train the inter-frame filtering model, the consistency in the training and testing of the inter-frame prediction coded frames can be improved, thereby improving the coding efficiency, enhancing the filtering performance of the video coding application, and reducing the distortion degree of the coded video.
[0164] Furthermore, it can be understood that the above Figure 6 The corresponding embodiments can be understood as the model training process where the inter-frame prediction coded frames only correspond to one type of filtering model (i.e., the inter-frame filtering model). In fact, some of the inter-frame prediction coded frames can be filtered by one type of filtering model (for example, taking B-1, B-3, B-5 frames as an example, they can be filtered by the B-1 model), and some frames can be filtered by another type of filtering model (for example, taking B-2, B-4, B-6 frames as an example, they can be filtered by the B-2 model). That is to say, the inter-frame filtering model can include two or more types of inter-frame filtering models (each inter-frame filtering model can refer to a model with different network structures and model parameters; of course, each inter-frame filtering model can also refer to a model with the same network structure but different model parameters). Here, for the convenience of distinction, the inter-frame prediction type (for example, the B-frame type) can be subdivided into the first inter-frame prediction type (such as the B-1, B-3, B-5 frame types) and the second inter-frame prediction type (such as the B-2, B-4, B-6 frame types), and a certain type of inter-frame filtering model is called a type inter-frame filtering model. For example, the inter-frame filtering model corresponding to the first inter-frame prediction type can be called the first type inter-frame filtering model, and the inter-frame filtering model corresponding to the second inter-frame prediction type can be called the second type inter-frame filtering model.
[0165] It can be understood that when different types of inter-frame filtering models are used to filter different inter-frame prediction coded frames, in the video coding application of the present application, it may include an intra-frame filtering model for filtering a reconstructed frame belonging to the intra-frame prediction type (the reconstructed frame to be filtered corresponding to the intra-frame prediction coded frame, for example, the reconstructed frame corresponding to the I frame), a first type of inter-frame filtering model for filtering a reconstructed frame belonging to the first inter-frame prediction type, and a second type of inter-frame filtering model for filtering a reconstructed frame belonging to the second inter-frame prediction type. The present application can determine a model to be trained between the intra-frame filtering model and different types of inter-frame filtering models, and then train the model to be trained based on training data. For example, taking the first filtering model including a first intra-frame filtering model for filtering a reconstructed frame to be filtered belonging to the intra-frame prediction type (the first intra-frame filtering model is trained based on the initial intra-frame filtering model in the video coding application updated at the (k - 1)th deployment) as an example, the video coding application updated at the kth deployment further includes an untrained first type of initial inter-frame filtering model and a second type of initial inter-frame filtering model (the first type of initial inter-frame filtering model is used to filter a reconstructed frame to be filtered belonging to the first inter-frame prediction type, and the second type of initial inter-frame filtering model is used to filter a reconstructed frame to be filtered belonging to the second inter-frame prediction type). The present application can first determine the model to be trained in the video coding application updated at the kth deployment as the first intra-frame filtering model and the untrained first type of initial inter-frame filtering model, or can also determine the model to be trained only as the untrained first type of initial inter-frame filtering model.
[0166] For ease of understanding, please also refer to Figure 9 , Figure 9 which is a schematic flowchart of a process for training a filtering model provided by an embodiment of the present application. As Figure 9 shown, the process may at least include the following steps S801 - step S806:
[0167] Step S801, determine the first intra-frame filtering model and the first type of initial inter-frame filtering model in the video coding application updated at the kth deployment as the model to be trained in the video coding application updated at the kth deployment.
[0168] Specifically, each time when determining the model to be trained, if there is a certain filtering model (including the intra-frame filtering model and the inter-frame filtering model) that has not been trained yet, one of the filtering models that have not been trained can be preferentially selected as the model to be trained until all the filtering models have been trained. At this time, when all the filtering models have been trained, when determining the model to be trained, all the trained filtering models can be used as the models to be trained, and all the models can be simultaneously trained and adjusted based on the new training data. Of course, each time when determining the model to be trained, if there is a certain filtering model (including the intra-frame filtering model and the inter-frame filtering model) that has not been trained yet, one of the filtering models that have not been trained can be preferentially selected and used together with the trained models as the models to be trained. Optionally, each time when determining the model to be trained, if there is a certain filtering model (including the intra-frame filtering model and the inter-frame filtering model) that has not been trained yet, all the filtering models that have not been trained can also be used together as the models to be trained. In the embodiment of the present application, the first intra-frame filtering model in the video coding application updated in the kth deployment is a trained model, and the first type of initial inter-frame filtering model and the second type of initial inter-frame filtering model are both models that have not been trained yet. In the embodiment of the present application, the first intra-frame filtering model and the first type of initial inter-frame filtering model (or the second type of initial inter-frame filtering model) can be used together as the models to be trained. Of course, only the first type of initial inter-frame filtering model (or the second type of initial inter-frame filtering model) can also be used as the model to be trained; the first type of initial inter-frame filtering model and the second type of initial inter-frame filtering model can also be used together as the models to be trained.
[0169] Step S802: Obtain a first type of inter-frame reconstructed frame to be filtered belonging to the first inter-frame prediction type in the first sample frame to be filtered and reconstructed, and use the sample original video frame corresponding to the first type of inter-frame reconstructed frame to be filtered in the sample original video frame as the third video frame label.
[0170] Specifically, for the intra-frame filtering model, a video frame with the frame type being the intra-frame prediction type can be used for training, and for the inter-frame filtering model, a video frame with the frame type being the inter-frame prediction type can be used for training. Then, when the filtering model to be trained in the video coding application updated in the kth deployment includes the first type of initial inter-frame filtering model, a reconstructed frame to be filtered belonging to the first inter-frame prediction type (which can be called the first type of inter-frame reconstructed frame to be filtered) can be obtained from the first training data, and the sample original video frame corresponding to the first type of inter-frame reconstructed frame to be filtered in the sample original video frame is used as its corresponding video frame label (which can be called the third video frame label).
[0171] Step S803: Obtain the second intra-frame reconstructed frame belonging to the intra-frame prediction type in the first sample frame to be filtered and reconstructed, and use the corresponding sample original video frame of the second intra-frame reconstructed frame in the sample original video frame as the fourth video frame label.
[0172] Specifically, similarly, when the filter model to be trained in the video encoding application updated in the k-th deployment includes the first intra-frame filter model, the frame to be filtered and reconstructed belonging to the intra-frame prediction type (which can be called the second intra-frame reconstructed frame to be filtered) can be obtained in the first training data, and the corresponding sample original video frame of the second intra-frame reconstructed frame in the sample original video frame is used as its corresponding video frame label (which can be called the fourth video frame label).
[0173] Step S804: Train the first type of initial inter-frame filter model based on the first type of inter-frame reconstructed frame to be filtered and the third video frame label to obtain the first type of inter-frame filter model in a training convergence state.
[0174] Specifically, through the first type of inter-frame reconstructed frame to be filtered and the third video frame label, the first type of initial inter-frame filter model can be trained and adjusted to obtain the first type of inter-frame filter model in a training convergence state.
[0175] Step S805: Train the first intra-frame filter model based on the second intra-frame reconstructed frame to be filtered and the fourth video frame label to obtain the second intra-frame filter model in a training convergence state.
[0176] Specifically, through the second intra-frame reconstructed frame to be filtered and the fourth video frame label, the first intra-frame filter model can be trained to obtain the second intra-frame filter model in a training convergence state.
[0177] Step S806: Determine the first type of inter-frame filter model and the second intra-frame filter model as the second filter model.
[0178] Specifically, the second intra-frame filter model and the first type of inter-frame filter model can be used as the second filter model. For the specific training process of each filter model, reference can be made to the description of the training process for training the filter model in the corresponding embodiments above. Figure 3 The description will not be repeated here.
[0179] It should be noted that when the model to be trained in the video coding application updated and deployed for the k-th time includes the first type of initial inter-frame filtering model and the first intra-frame filtering model, the specific method for updating and deploying the second filtering model in the video coding application updated and deployed for the k-th time to obtain the video coding application updated and deployed for the (k + 1)-th time can be as follows: the first intra-frame filtering model in the video coding application updated and deployed for the k-th time can be replaced and updated with the second intra-frame filtering model; the first type of initial inter-frame filtering model in the video coding application updated and deployed for the k-th time can be replaced and updated with the first type of inter-frame filtering model. That is to say, the two trained models are updated and replaced separately.
[0180] Optionally, it can be understood that in the case where the model to be trained in the video coding application updated and deployed for the k-th time includes the first type of initial inter-frame filtering model and the first intra-frame filtering model, when it is detected that the video coding application updated and deployed for the (k + 1)-th time does not meet the filtering quality requirement condition, the second training data can be generated by using the video coding application updated and deployed for the (k + 1)-th time and the sample video data; wherein, the second training data includes the sample original video frames as training labels and the second sample video frames to be filtered and reconstructed corresponding to the sample original video frames, and the second sample video frames to be filtered and reconstructed refer to the video frames that are not filtered by the second filtering model during the process of encoding and reconstructing the sample original video frames by using the video coding application updated and deployed for the (k + 1)-th time; subsequently, based on the sample original video frames and the second sample video frames to be filtered and reconstructed, the model to be trained for filtering in the video coding application updated and deployed for the (k + 1)-th time (for example, the second type of initial inter-frame filtering model) can be trained to obtain the third filtering model in a training convergence state; subsequently, the third filtering model can be updated and deployed in the video coding application updated and deployed for the (k + 1)-th time to obtain the video coding application updated and deployed for the (k + 2)-th time; when the video coding application updated and deployed for the (k + 2)-th time meets the filtering quality requirement condition, the video coding application updated and deployed for the (k + 2)-th time can be determined as the target video coding application.
[0181] Taking the untrained filtering model to be trained in the video encoding application updated in the (k + 1)-th deployment as the second type of initial inter-frame filtering model as an example, for the sample original video frame and the second sample reconstructed frame to be filtered, the specific implementation of training the untrained filtering model in the video encoding application updated in the (k + 1)-th deployment to obtain the third filtering model in a training convergence state can be as follows: Second-type reconstructed frames to be filtered belonging to the second inter-frame prediction type can be obtained in the second sample reconstructed frame to be filtered, and the sample original video frame corresponding to the second-type reconstructed frames to be filtered in the sample original video frame can be used as the fifth video frame label; The second-type initial inter-frame filtering model can be trained based on the second-type reconstructed frames to be filtered and the fifth video frame label to obtain the second-type inter-frame filtering model in a training convergence state; The second-type inter-frame filtering model can be determined as the third filtering model.
[0182] It can be understood that for the case where an intra-frame prediction coded frame (e.g., I frame) has one type of model (i.e., an intra-frame filtering model, such as an I frame model) and an inter-frame prediction coded frame (e.g., B frame) has two or more types of models, for the training of the filtering model, the Figure 9 shown process can be adopted. For ease of understanding, please also refer to Figure 10 , Figure 10 which is a schematic diagram of another model training scenario provided by the embodiments of the present application. As Figure 10 shown, taking the intra-frame filtering model including an I frame model and the inter-frame filtering model including B-1 and B-2 models as an example, in this case, combined with the scenario shown in Figure 10 , the process of iterative training of the model can include the following steps 1 to 6:
[0183] Step 1: When the video encoding application includes an initial untrained I frame model, B-1 model, and B-2 model, the sample video data can be input into the video encoding application, and training data can be output through the video encoding application.
[0184] Step 2: Taking the initial I frame model in the video encoding application as the model to be trained, training the I frame model based on the training data (training the I frame model based on the I frame original video frame and the I frame reconstructed frame to be filtered in the training data) to obtain the I frame model in a training convergence state (i.e., the trained I frame model), and the initial I frame model in the video encoding application can be replaced with the trained I frame model, keeping the initial B-1 and B-2 models unchanged, to obtain the video encoding application updated in the 1st deployment.
[0185] Step 3: Inputting the sample video data into the video encoding application updated in the 1st deployment, and outputting new training data through the video encoding application updated in the 1st deployment.
[0186] Step 4: Use the initial B-1 model in the video encoding application updated in the first deployment as the model to be trained, and train the initial B-1 model based on the new training data (train the B-1 model based on the original video frames and the frames to be filtered and reconstructed corresponding to the B-1 model in the training data), to obtain a B-1 model in a training convergence state (the trained B-1 model). The initial B-1 model in the video encoding application updated in the first deployment can be replaced with the trained B-1 model to obtain a video encoding application updated in the second deployment.
[0187] Step 5: Input the sample video data into the video encoding application updated in the second deployment, and output new training data through the video encoding application updated in the second deployment.
[0188] Step 6: Use the initial B-2 model in the video encoding application updated in the second deployment as the model to be trained, and train the B-2 model based on the new training data (train the B-2 model based on the original video frames and the frames to be filtered and reconstructed corresponding to the B-2 model in the training data), to obtain a B-2 model in a training convergence state. The initial B-2 model in the video encoding application updated in the second deployment can be replaced with the trained B-2 model to obtain a video encoding application updated in the third deployment. Of course, subsequently, the sample video data can be input into the video encoding application updated in the third deployment again, and new training data is output again based on the video encoding application updated in the third deployment. At this time, each filtering model in the video encoding application updated in the third deployment can be trained based on the new training data to obtain trained models, and then the video encoding application updated in the third deployment is updated and deployed based on the trained models to obtain a video encoding application updated in the fourth deployment.
[0189] Of course, it can be understood that in Step 4 above, the initial B-1 model and the initial B-2 model can also be trained based on the training data. After obtaining the trained B-1 model and B-2 model, the initial B-1 model and B-2 model in the video encoding application updated in the first deployment are updated and replaced respectively to obtain a video encoding application updated in the second deployment. After training the I-frame model, B-1 model, and B-2 model, new training data can be generated again, and the I-frame model, B-1 model, and B-2 model are trained together based on the new training data.
[0190] It should be noted that after each deployment update of the video encoding application, it can be detected to determine whether it meets the filtering requirement conditions. When the filtering requirement conditions are not met, a new training data set is generated, and the model is trained again based on the new training data set. Then, the video encoding application is deployed and updated again based on the trained model until the video encoding application meets the filtering requirement conditions.
[0191] In the embodiment of the present application, by updating the training data set to repeatedly train the inter-frame filtering model, the consistency in the training and testing of the inter-frame predicted coding frames can be improved, thereby improving the coding efficiency, enhancing the filtering quality of the video encoding application, and reducing the distortion degree of the encoded video.
[0192] It should be noted that the model iterative training process proposed in the present application can also be applied to the iterative training of other models. For example, it is also applicable to the iterative training of the inter-frame prediction model and the intra-frame prediction model. That is to say, the model training method proposed in the present application, which regenerates the training data based on the application after model deployment and repeats training the model, and then updates and deploys the application until the application meets the performance requirement conditions, is not limited to the iterative training of the filtering model.
[0193] It can be understood that after determining the target video encoding application through the model iterative training method proposed in the embodiment of the present application, the target video encoding application can be put into use, that is, the target video encoding application can be used to perform video encoding processing on video data. For example, in a video call scenario, when two users are having a video call, the terminal device corresponding to user a can perform video encoding processing on the video data associated with user a based on the target video encoding application. After obtaining the video compression bitstream, the video compression bitstream is transmitted to the terminal device corresponding to user b (the user having a video call with user a) so that the terminal device corresponding to user b can decode and output the video data associated with user a on the display interface. For ease of understanding, please also refer to Figure 11 , Figure 11 is a schematic flowchart of a data processing method provided by an embodiment of the present application, and this process can be the application process of the target video encoding application. As Figure 11 shown, this process can at least include the following steps S201 - step S202:
[0194] Step S201, inputting video data into a target video coding application, performing video coding processing on the video data through the target video coding application, and obtaining a video compression code stream corresponding to the video data; the target video coding application refers to a video coding application updated for the k+1th deployment that meets the filtering quality requirement; the video coding application updated for the k+1th deployment includes a second filtering model in a training convergence state; the second filtering model is obtained by training the filtering model to be trained in the video coding application updated for the kth deployment that includes the first filtering model based on the sample original video frame as a training label in the first training data and the first sample to-be-filtered reconstructed frame corresponding to the sample original video frame; the first training data is generated by the video coding application updated for the kth deployment and the sample video data; the first sample to-be-filtered reconstructed frame refers to a reconstructed frame that has not been filtered by the first filtering model during the process of reconstructing the sample original video frame through the video coding application updated for the kth deployment; the sample original video frame is a video frame in the sample video data; k is a positive integer.
[0195] For details on the specific process of determining the target video encoding application, please refer to the previous article Figure 2 The description in the corresponding embodiment will not be repeated here.
[0196] Step S202: sending the video compression code stream to a receiving device, so that the receiving device decodes the video compression code stream.
[0197] Specifically, the computer device may send the video compression code stream to a receiving device (eg, a terminal device that receives the video compression code stream), and the receiving device may decode the video compression code stream.
[0198] In an embodiment of the present application, by repeatedly training the inter-frame filtering model by updating the training data set, the consistency in training and testing of the inter-frame prediction coding frames can be improved, thereby improving the coding efficiency, improving the filtering quality of the video coding application, and reducing the distortion of the encoded video.
[0199] For further information, see Figure 12 , Figure 12 1 is a schematic diagram of the structure of a data processing device provided in an embodiment of the present application. The data processing device may be a computer program (including program code) running in a computer device, for example, the data processing device is an application software; the data processing device may be used to execute Figure 3 As shown in the method. Figure 12 As shown, the data processing device 1 may include: a training data generating module 11, a model training module 12, an application updating module 13 and a target application determining module 14.
[0200] A training data generation module 11 is configured to input sample video data into a video encoding application including a first filtering model updated in the k-th deployment;
[0201] The training data generation module 11 is further configured to generate first training data through the video encoding application updated in the k-th deployment and the sample video data; the first training data includes sample original video frames as training labels and first sample reconstructed frames to be filtered corresponding to the sample original video frames; the first sample reconstructed frames to be filtered refer to the reconstructed frames that are not filtered by the first filtering model during the process of reconstructing the sample original video frames through the video encoding application updated in the k-th deployment; the sample original video frames are video frames in the sample video data; k is a positive integer;
[0202] A model training module 12 is configured to train a filtering model to be trained in the video encoding application updated in the k-th deployment based on the sample original video frames and the first sample reconstructed frames to be filtered, so as to obtain a second filtering model in a training convergence state;
[0203] An application update module 13 is configured to update and deploy the second filtering model in the video encoding application updated in the k-th deployment to obtain a video encoding application updated in the (k + 1)-th deployment;
[0204] A target application determination module 14 is configured to, when the video encoding application updated in the (k + 1)-th deployment meets the filtering quality requirement condition, determine the video encoding application updated in the (k + 1)-th deployment as a target video encoding application for performing video encoding processing on video data.
[0205] Wherein, for the specific implementation manners of the training data generation module 11, the model training module 12, the application update module 13, and the target application determination module 14, reference may be made to the descriptions of steps S101 - S103 in the corresponding embodiments above, which will not be elaborated here. Figure 2 The descriptions of steps S101 - S103 in the corresponding embodiments above will not be elaborated here.
[0206] In one embodiment, the first filtering model includes a first intra-frame filtering model for filtering reconstructed frames to be filtered belonging to the intra-frame prediction type; the first intra-frame filtering model is trained based on an initial intra-frame filtering model in the video encoding application updated in the (k - 1)-th deployment, and the video encoding application updated in the k-th deployment further includes an untrained initial inter-frame filtering model; the initial inter-frame filtering model is used for filtering reconstructed frames to be filtered belonging to the inter-frame prediction type, and the filtering model to be trained in the video encoding application updated in the k-th deployment includes the initial inter-frame filtering model and the first intra-frame filtering model;
[0207] The model training module 12 may include: a video label acquisition unit 121, a model training unit 122, and a model determination unit 123.
[0208] The video label acquisition unit 121 is configured to obtain, from a first sample to-be-filtered and reconstructed frame, a to-be-filtered and inter-frame reconstructed frame belonging to an inter-frame prediction type, and use the sample original video frame corresponding to the to-be-filtered and inter-frame reconstructed frame in the sample original video frame as a first video frame label.
[0209] The video label acquisition unit 121 is further configured to obtain, from the first sample to-be-filtered and reconstructed frame, a first to-be-filtered and intra-frame reconstructed frame belonging to an intra-frame prediction type, and use the sample original video frame corresponding to the first to-be-filtered and intra-frame reconstructed frame in the sample original video frame as a second video frame label.
[0210] The model training unit 122 is configured to train an initial inter-frame filtering model based on the to-be-filtered and inter-frame reconstructed frame and the first video frame label to obtain a first inter-frame filtering model in a training convergence state.
[0211] The model training unit 122 is further configured to train a first intra-frame filtering model based on the to-be-filtered and intra-frame reconstructed frame and the second video frame label to obtain a second intra-frame filtering model in a training convergence state.
[0212] The model determination unit 123 is configured to determine the first inter-frame filtering model and the second intra-frame filtering model as a second filtering model.
[0213] Wherein, for the specific implementation manners of the video label acquisition unit 121, the model training unit 122, and the model determination unit 123, reference may be made to the descriptions of steps S501 - S505 in the corresponding embodiment above, which will not be elaborated here. Figure 6 The descriptions of steps S501 - S505 in the corresponding embodiment above, which will not be elaborated here.
[0214] In one embodiment, the application update module 13 may include: an intra-frame model replacement unit 131 and an inter-frame model replacement unit 132.
[0215] The intra-frame model replacement unit 131 is configured to replace and update the first intra-frame filtering model in the video coding application updated in the k-th deployment with the second intra-frame filtering model.
[0216] The inter-frame model replacement unit 132 is configured to replace and update the initial inter-frame filtering model in the video coding application updated in the k-th deployment with the first inter-frame filtering model.
[0217] Wherein, for the specific implementation manners of the intra-frame model replacement unit 131 and the inter-frame model replacement unit 132, reference may be made to the description of step S505 in the corresponding embodiment above, which will not be elaborated here. Figure 6 The descriptions of step S505 in the corresponding embodiment above, which will not be elaborated here.
[0218] In one embodiment, the first filtering model includes a first intra-frame filtering model for filtering a reconstructed frame to be filtered belonging to an intra-frame prediction type; the first intra-frame filtering model is trained based on an initial intra-frame filtering model in a video coding application updated in the (k - 1)-th deployment, and the video coding application updated in the k-th deployment further includes an untrained first-type initial inter-frame filtering model and a second-type initial inter-frame filtering model; the first-type initial inter-frame filtering model is used for filtering a reconstructed frame to be filtered belonging to a first inter-frame prediction type, and the second-type initial inter-frame filtering model is used for filtering a reconstructed frame to be filtered belonging to a second inter-frame prediction type;
[0219] The model training module 12 may include: a model to be trained determination unit 124, a video frame label determination unit 125, an inter-frame model training unit 126, an intra-frame model training unit 127, and a filtering model determination unit 128.
[0220] The model to be trained determination unit 124 is configured to determine the first intra-frame filtering model and the first-type initial inter-frame filtering model in the video coding application updated in the k-th deployment as the filtering model to be trained in the video coding application updated in the k-th deployment;
[0221] The video frame label determination unit 125 is configured to obtain a first-type reconstructed frame to be filtered belonging to a first inter-frame prediction type in a first sample reconstructed frame to be filtered, and use the sample original video frame corresponding to the first-type reconstructed frame to be filtered in the sample original video frame as a third video frame label;
[0222] The video frame label determination unit 125 is further configured to obtain a second reconstructed frame to be filtered belonging to an intra-frame prediction type in the first sample reconstructed frame to be filtered, and use the sample original video frame corresponding to the second reconstructed frame to be filtered in the sample original video frame as a fourth video frame label;
[0223] The inter-frame model training unit 126 is configured to train the first-type initial inter-frame filtering model based on the first-type reconstructed frame to be filtered and the third video frame label to obtain a first-type inter-frame filtering model in a training convergence state;
[0224] The intra-frame model training unit 127 is configured to train the first intra-frame filtering model based on the second reconstructed frame to be filtered and the fourth video frame label to obtain a second intra-frame filtering model in a training convergence state;
[0225] The filtering model determination unit 128 is configured to determine the first-type inter-frame filtering model and the second intra-frame filtering model as a second filtering model.
[0226] For the specific implementation manner, reference may be made to the above Figure 6The descriptions of steps S501 - S505 in the corresponding embodiments will not be elaborated here.
[0227] Among them, for the specific implementation manners of the to - be - trained model determination unit 124, the video frame label determination unit 125, the inter - frame model training unit 126, the intra - frame model training unit 127, and the filtering model determination unit 128, reference can be made to the Figure 9 The descriptions of steps S801 - S806 in the corresponding embodiments will not be elaborated here.
[0228] In one embodiment, the data processing device 1 may further include: a data generation module 15, a filtering model training module 16, a deployed model module 17, and a target application determination module 18.
[0229] The data generation module 15 is configured to, when the video encoding application updated in the (k + 1)-th deployment does not meet the filtering quality requirement condition, generate second training data through the video encoding application updated in the (k + 1)-th deployment and the sample video data; the second training data includes the sample original video frames as training labels, and the second sample video frames to be filtered and reconstructed corresponding to the sample original video frames, where the second sample video frames to be filtered and reconstructed refer to the video frames that are not filtered by the second filtering model during the process of encoding and reconstructing the sample original video frames through the video encoding application updated in the (k + 1)-th deployment;
[0230] The filtering model training module 16 is configured to train the to - be - trained filtering model in the video encoding application updated in the (k + 1)-th deployment based on the sample original video frames and the second sample video frames to be filtered and reconstructed, to obtain a third filtering model in a training convergence state;
[0231] The deployed model module 17 is configured to update and deploy the third filtering model in the video encoding application updated in the (k + 1)-th deployment to obtain the video encoding application updated in the (k + 2)-th deployment;
[0232] The target application determination module 18 is configured to, when the video encoding application updated in the (k + 2)-th deployment meets the filtering quality requirement condition, determine the video encoding application updated in the (k + 2)-th deployment as the target video encoding application.
[0233] Among them, for the specific implementation manners of the data generation module 15, the filtering model training module 16, the deployed model module 17, and the target application determination module 18, reference can be made to the Figure 9 The descriptions of steps S806 in the corresponding embodiments will not be elaborated here.
[0234] In one embodiment, the to - be - trained filtering model in the video encoding application updated in the (k + 1)-th deployment includes a second - type initial inter - frame filtering model;
[0235] The filtering model training module 16 may include: a label determination unit 161, a type inter-frame model training unit 162, and a filtering model determination unit 163.
[0236] The label determination unit 161 is configured to obtain, from the second sample to-be-filtered and reconstructed frame, a second type of to-be-filtered and inter-frame reconstructed frame belonging to the second inter-frame prediction type, and use the sample original video frame corresponding to the second type of to-be-filtered and inter-frame reconstructed frame in the sample original video frame as the fifth video frame label.
[0237] The type inter-frame model training unit 162 is configured to train the second type of initial inter-frame filtering model based on the second type of to-be-filtered and inter-frame reconstructed frame and the fifth video frame label, to obtain a second type of inter-frame filtering model in a training convergence state.
[0238] The filtering model determination unit 163 is configured to determine the second type of inter-frame filtering model as the third filtering model.
[0239] Wherein, for the specific implementation manners of the label determination unit 161, the type inter-frame model training unit 162, and the filtering model determination unit 163, reference may be made to the description of step S806 in the corresponding embodiment above, and details will not be elaborated herein. Figure 9 For the specific implementation manners of the label determination unit 161, the type inter-frame model training unit 162, and the filtering model determination unit 163, reference may be made to the description of step S806 in the corresponding embodiment above, and details will not be elaborated herein.
[0240] In one embodiment, the model training module 12 may include: a filtered frame output unit 129 and a parameter adjustment unit 120.
[0241] The filtered frame output unit 129 is configured to input the first sample to-be-filtered and reconstructed frame into the first filtering model, and output, through the first filtering model, a sample filtered and reconstructed frame corresponding to the first sample to-be-filtered and reconstructed frame.
[0242] The parameter adjustment unit 120 is configured to determine an error value between the sample filtered and reconstructed frame and the sample original video frame.
[0243] The parameter adjustment unit 120 is further configured to adjust the model parameters of the first filtering model through the error value, to obtain a first filtering model with adjusted model parameters.
[0244] The parameter adjustment unit 120 is further configured to, when the first filtering model with adjusted model parameters meets the model convergence condition, determine the first filtering model with adjusted model parameters as a second filtering model in a training convergence state.
[0245] Wherein, for the specific implementation manners of the filtered frame output unit 129, the parameter adjustment unit 120, and the model determination unit 123, reference may be made to the description of step S102 in the corresponding embodiment above, and details will not be elaborated herein. Figure 3 For the specific implementation manners of the filtered frame output unit 129, the parameter adjustment unit 120, and the model determination unit 123, reference may be made to the description of step S102 in the corresponding embodiment above, and details will not be elaborated herein.
[0246] In one embodiment, the parameter adjustment unit 120 includes: a function acquisition subunit 1201 and an error value determination subunit 1202.
[0247] The function acquisition subunit 1201 is configured to acquire the loss function corresponding to the video encoding application updated by the k-th deployment.
[0248] The error value determination subunit 1202 is configured to obtain the original image quality corresponding to the sample original video frame based on the loss function, and use the original image quality as the image quality label.
[0249] The error value determination subunit 1202 is further configured to obtain the filtered image quality corresponding to the sample filtered reconstruction frame based on the loss function, and determine the error value between the sample filtered video frame and the sample original video frame through the loss function, the image quality label, and the filtered image quality.
[0250] For the specific implementation manners of the function acquisition subunit 1201 and the error value determination subunit 1202, reference may be made to the description of step S102 in the corresponding embodiment above, which will not be elaborated here. Figure 2 For the specific implementation manners of the function acquisition subunit 1201 and the error value determination subunit 1202, reference may be made to the description of step S102 in the corresponding embodiment above, which will not be elaborated here.
[0251] In one embodiment, the loss function includes an absolute value loss function and a squared error loss function.
[0252] The error value determination subunit 1202 is further specifically configured to determine a first error value between the sample filtered reconstruction frame and the sample original video frame through the squared error loss function, the image quality label, and the filtered image quality.
[0253] The error value determination subunit 1202 is further specifically configured to determine a second error value between the sample filtered reconstruction frame and the sample original video frame through the absolute value loss function, the image quality label, and the filtered image quality.
[0254] The error value determination subunit 1202 is further specifically configured to perform an arithmetic operation on the first error value and the second error value, and determine the result of the operation as the error value between the sample filtered reconstruction frame and the sample original video frame.
[0255] In one embodiment, the data processing device 1 may further include: a filtered video output module 19, a video quality acquisition module 20, and an application detection module 21.
[0256] The filtered video output module 19 is configured to input the sample video data into the video encoding application updated by the (k + 1)-th deployment, and output the filtered video data to be detected through the video encoding application updated by the (k + 1)-th deployment.
[0257] A video quality acquisition module 20, configured to acquire the filtered video quality corresponding to the filtered video data to be detected, and the original video quality corresponding to the sample video data;
[0258] An application detection module 21, configured to detect the video encoding application updated and deployed for the (k + 1)-th time according to the filtered video quality and the original video quality.
[0259] Among them, for the specific implementation manners of the filtered video output module 19, the video quality acquisition module 20, and the application detection module 21, reference may be made to the description of step S103 in the corresponding embodiment above, which will not be elaborated here. Figure 3 The description corresponding to step S103 in the embodiment will not be elaborated here.
[0260] In one embodiment, the application detection module 21 may include: a difference quality determination unit 211 and a detection result determination unit 212.
[0261] The difference quality determination unit 211 is configured to determine the difference video quality between the filtered video quality and the original video quality;
[0262] The detection result determination unit 212 is configured to determine that the second video encoding application meets the filtered quality requirement condition if the difference video quality is less than the difference quality threshold;
[0263] The detection result determination unit 212 is further configured to determine that the second video encoding application does not meet the filtered quality requirement condition if the difference video quality is greater than the difference quality threshold.
[0264] Among them, for the specific implementation manners of the difference quality determination unit 211 and the detection result determination unit 212, reference may be made to the description of step S103 in the corresponding embodiment above, which will not be elaborated here. Figure 2 The description corresponding to step S103 in the embodiment will not be elaborated here.
[0265] In the embodiments of the present application, by updating the training data set to retrain the inter-frame filtering model, the consistency of the inter-frame predicted coding frames in training and testing can be improved, thereby improving the coding efficiency, enhancing the filtering performance of the video encoding application, and reducing the distortion degree of the encoded video.
[0266] Further, please refer to Figure 13 , Figure 13 which is a schematic structural diagram of a data processing device provided in the embodiments of the present application. The data processing device may be a computer program (including program code) running in a computer device. For example, the data processing device is an application software; the data processing device may be used to execute Figure 11 the method shown. As Figure 13 shown, the data processing device 2 may include: a bitstream generation module 31 and a bitstream sending module 32.
[0267] A bitstream generation module 31 is configured to input video data into a target video encoding application, and perform video encoding processing on the video data through the target video encoding application to obtain a video compression bitstream corresponding to the video data. The target video encoding application refers to the (k + 1)-th deployed and updated video encoding application that meets the filtering quality requirement conditions. The (k + 1)-th deployed and updated video encoding application includes a second filtering model in a training convergence state. The second filtering model is obtained by training a filtering model to be trained in the video encoding application including the first filtering model in the k-th deployed and updated video encoding application based on the sample original video frames as training labels in the first training data and the first sample frames to be filtered and reconstructed corresponding to the sample original video frames. The first training data is generated by the k-th deployed and updated video encoding application and sample video data. The first sample frames to be filtered and reconstructed refer to the reconstructed frames that are not filtered by the first filtering model during the process of reconstructing the sample original video frames through the k-th deployed and updated video encoding application. The sample original video frames are video frames in the sample video data. k is a positive integer;
[0268] A bitstream sending module 32 is configured to send the video compression bitstream to a receiving device so that the receiving device performs decoding processing on the video compression bitstream.
[0269] Among them, for the specific implementation manners of the bitstream generation module 31 and the bitstream sending module 32, reference may be made to the descriptions in the corresponding embodiments above Figure 11 and will not be elaborated here.
[0270] Further, please refer to Figure 14 , Figure 14 which is a schematic structural diagram of a computer device provided in an embodiment of the present application. As Figure 14 shown, the device 1 in the corresponding embodiment above or Figure 12 Figure 13The device 2 in the corresponding embodiment can be applied to the above computer device 8000. The computer device 8000 may include: a processor 8001, a network interface 8004, and a memory 8005. In addition, the computer device 8000 further includes: a user interface 8003 and at least one communication bus 8002. Among them, the communication bus 8002 is used to realize the connection and communication between these components. Among them, the user interface 8003 may include a display screen (Display) and a keyboard (Keyboard). Optionally, the user interface 8003 may further include a standard wired interface and a wireless interface. The network interface 8004 may optionally include a standard wired interface and a wireless interface (such as a WI-FI interface). The memory 8005 may be a high-speed RAM memory or a non-volatile memory, such as at least one disk memory. Optionally, the memory 8005 may further be at least one storage device located far from the aforementioned processor 8001. As Figure 14 shown, the memory 8005, as a computer-readable storage medium, may include an operating system, a network communication module, a user interface module, and a device control application program.
[0271] In Figure 14 the computer device 8000 shown, the network interface 8004 can provide network communication functions; while the user interface 8003 is mainly used to provide an input interface for users; and the processor 8001 can be used to call the device control application program stored in the memory 8005 to achieve:
[0272] Inputting the sample video data into the video coding application including the first filtering model updated in the k-th deployment, and generating first training data through the video coding application updated in the k-th deployment and the sample video data; the first training data includes the sample original video frames as training labels and the first sample reconstruction frames to be filtered corresponding to the sample original video frames; the first sample reconstruction frames to be filtered refer to the reconstruction frames that have not been filtered by the first filtering model during the process of reconstructing the sample original video frames through the video coding application updated in the k-th deployment; the sample original video frames are the video frames in the sample video data; k is a positive integer;
[0273] Training the filtering model to be trained in the video coding application updated in the k-th deployment based on the sample original video frames and the first sample reconstruction frames to be filtered, obtaining a second filtering model in a training convergence state, and updating and deploying the second filtering model in the video coding application updated in the k-th deployment to obtain the video coding application updated in the (k + 1)-th deployment;
[0274] When the video encoding application updated and deployed for the k+1th time meets the filtering quality requirement condition, the video encoding application updated and deployed for the k+1th time is determined as a target video encoding application for performing video encoding processing on the video data.
[0275] Or implement:
[0276] The video data is input into the target video coding application, and the video data is subjected to video coding processing by the target video coding application, and the video data corresponds to a video compression code stream; the target video coding application refers to the video coding application updated for the k+1th deployment that meets the filtering quality requirement; the video coding application updated for the k+1th deployment includes a second filtering model in a training convergence state; the second filtering model is obtained by training the filtering model to be trained in the video coding application updated for the kth deployment that includes the first filtering model based on the sample original video frame as a training label in the first training data and the first sample to-be-filtered reconstructed frame corresponding to the sample original video frame; the first training data is generated by the video coding application updated for the kth deployment and the sample video data; the first sample to-be-filtered reconstructed frame refers to a reconstructed frame that has not been filtered by the first filtering model during the process of reconstructing the sample original video frame by the video coding application updated for the kth deployment; the sample original video frame is a video frame in the sample video data; k is a positive integer;
[0277] The video compression code stream is sent to a receiving device so that the receiving device decodes the video compression code stream.
[0278] It should be understood that the computer device 8000 described in the embodiment of the present application can execute the above Figure 2 or Figure 11 The description of the data processing method in the corresponding embodiment can also be performed as described above. Figure 12 In the corresponding embodiment, the data processing device 1 or Figure 13 The description of the data processing device 2 in the corresponding embodiment will not be repeated here. In addition, the description of the beneficial effects of the same method will not be repeated here either.
[0279] In addition, it should be pointed out here that: the embodiment of the present application also provides a computer-readable storage medium, and the computer-readable storage medium stores a computer program executed by the data processing computer device 8000 mentioned above, and the computer program includes program instructions. When the processor executes the program instructions, the computer program can execute the above-mentioned data processing computer device 8000. Figure 3 or Figure 11The description of the above data processing method in the corresponding embodiments will not be repeated here. In addition, the description of the beneficial effects of using the same method will not be repeated either. For the technical details not disclosed in the embodiments of the computer-readable storage medium involved in this application, please refer to the description of the method embodiments of this application.
[0280] The above computer-readable storage medium may be the data processing device provided in any of the foregoing embodiments or the internal storage unit of the above computer device, such as the hard disk or memory of the computer device. The computer-readable storage medium may also be an external storage device of the computer device, such as a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. equipped on the computer device. Further, the computer-readable storage medium may also include both the internal storage unit and the external storage device of the computer device. The computer-readable storage medium is used to store the computer program and other programs and data required by the computer device. The computer-readable storage medium may also be used to temporarily store the data that has been output or is to be output.
[0281] In one aspect of this application, a computer program product or a computer program is provided. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the method provided in one aspect of the embodiments of this application.
[0282] The terms "first", "second", etc. in the description, claims and drawings of the embodiments of this application are used to distinguish different objects, rather than to describe a specific order. In addition, the term "comprising" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, device, product or equipment that includes a series of steps or units is not limited to the listed steps or modules, but may optionally further include steps or modules not listed, or may optionally further include other step units inherent to these processes, methods, devices, products or equipment.
[0283] Those of ordinary skill in the art will realize that the units and algorithm steps of the examples described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the components and steps of the examples have been generally described in terms of function in the above description. Whether these functions are executed in hardware or software depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.
[0284] The methods and related devices provided by the embodiments of this application are described with reference to the method flowcharts and / or structural schematic diagrams provided by the embodiments of this application. Specifically, each process and / or block of the method flowchart and / or structural schematic diagram, as well as the combination of the processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate a device for implementing the functions specified in Figure 1 one process or multiple processes and / or structural schematic Figure 1 one block or multiple blocks. These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured article including an instruction device, and the instruction device implements the functions specified in Figure 1 one process or multiple processes and / or structural schematic Figure 1 one block or multiple blocks. These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in Figure 1 one process or multiple processes and / or structural schematic one block or multiple blocks.
[0285] The above-disclosed are only the preferred embodiments of this application. Of course, the scope of the rights of this application cannot be limited thereby. Therefore, equivalent changes made according to the claims of this application still fall within the scope covered by this application.
Claims
1. A data processing method, characterized in that, comprising: obtaining a video encoding application whose filtering performance does not meet the filtering quality requirement conditions; a first filtering model is deployed in the video encoding application; encoding and processing the sample video data through the video encoding application, and training the first filtering model based on the unfiltered data generated by the encoding process and the sample video data; deploying and updating the trained first filtering model to the video encoding application to obtain an updated video encoding application; when the filtering performance of the updated video encoding application meets the filtering quality requirement conditions, it is used to perform video encoding processing on video data.
2. The method according to claim 1, characterized in that, the obtaining of the video encoding application whose filtering performance does not meet the filtering quality requirement conditions includes: obtaining the video encoding application updated in the kth deployment; k is a positive integer; detecting the filtering performance of the video encoding application updated in the kth deployment; if the filtering performance of the video encoding application updated in the kth deployment does not meet the filtering quality requirement conditions, then determining the video encoding application updated in the kth deployment as the video encoding application whose filtering performance does not meet the filtering quality requirement conditions.
3. The method according to claim 1, characterized in that, the video encoding application whose filtering performance does not meet the filtering quality requirement conditions is the video encoding application updated in the kth deployment; the sample video data includes at least one sample original video frame, and the encoding process performed by the video encoding application updated in the kth deployment on the sample video data includes encoding each of the sample original video frames, and the unfiltered data generated by the encoding process includes a first sample to-be-filtered reconstruction frame corresponding to each of the sample original video frames; the first sample to-be-filtered reconstruction frame corresponding to each of the sample original video frames refers to a reconstruction frame generated during the encoding process that has not been filtered by the first filtering model.
4. The method according to claim 3, characterized in that, the training of the first filtering model based on the sample video data and the unfiltered data generated by the encoding process includes: forming a set of training data pairs by each of the sample original video frames and the corresponding first sample to-be-filtered reconstruction frame; determining the set composed of the training data pairs corresponding to at least one sample original video frame as the first training data of the first filtering model; training the first filtering model through the first training data.
5. The method according to claim 4, characterized in that, the training of the first filtering model through the first training data includes: inputting the first training data into the first filtering model; filtering the first sample to-be-filtered reconstruction frame in the first training data through the first filtering model, and outputting a sample filtered reconstruction frame corresponding to the first sample to-be-filtered reconstruction frame; Adjust the model parameters of the first filtering model based on the error between the sample filtered reconstruction frame corresponding to the first sample to-be-filtered reconstruction frame and the corresponding sample original video frame, to obtain the trained first filtering model.
6. The method according to any one of claims 1 to 5, wherein, the adjusting the model parameters of the first filtering model based on the error between the sample filtered reconstruction frame corresponding to the first sample to-be-filtered reconstruction frame and the corresponding sample original video frame, to obtain the trained first filtering model, includes: Determine the error value between the sample filtered reconstruction frame and the corresponding sample original video frame, and adjust the model parameters of the first filtering model through the error value to obtain a first filtering model with adjusted model parameters; When the first filtering model with adjusted model parameters meets the model convergence condition, determine the first filtering model with adjusted model parameters as the trained first filtering model.
7. The method according to claim 6, wherein, the determining the error value between the sample filtered reconstruction frame and the corresponding sample original video frame includes: Obtain the loss function corresponding to the video coding application updated in the k-th deployment; Based on the loss function, obtain the original image quality corresponding to the sample original video frame, and use the original image quality as the image quality label; Based on the loss function, obtain the filtered image quality corresponding to the sample filtered reconstruction frame, and determine the error value between the sample filtered video frame and the sample original video frame through the loss function, the image quality label, and the filtered image quality.
8. The method according to claim 1, wherein, the video coding application whose filtering performance does not meet the filtering quality requirement condition is the video coding application updated in the k-th deployment, and the video coding application after the deployment update is the video coding application updated in the (k + 1)-th deployment; the trained first filtering model is the second filtering model; the method further includes: Detect the filtering performance of the video coding application updated in the (k + 1)-th deployment; When the filtering performance of the video coding application updated in the (k + 1)-th deployment does not meet the filtering quality requirement condition, generate second training data through the video coding application updated in the (k + 1)-th deployment and the sample video data; Train the second filtering model in the video coding application updated in the (k + 1)-th deployment based on the second training data; Deploy and update the trained second filtering model to the video coding application updated in the (k + 1)-th deployment to obtain the video coding application updated in the (k + 2)-th deployment; When the filtering performance of the video coding application updated in the (k + 2)-th deployment meets the filtering quality requirement condition, determine the video coding application updated in the (k + 2)-th deployment as the target video coding application.
9. The method according to claim 1, wherein, the detecting the filtering performance of the video coding application updated in the (k + 1)-th deployment includes: Input the sample video data into the video encoding application updated in the (k + 1)-th deployment, and output the video data to be detected for filtering through the video encoding application updated in the (k + 1)-th deployment; Obtain the filtered video quality corresponding to the video data to be detected for filtering, and the original video quality corresponding to the sample video data; Detect the video encoding application updated in the (k + 1)-th deployment according to the filtered video quality and the original video quality.
10. The method according to claim 9, wherein, the detecting the filtering quality in the video encoding application updated in the (k + 1)-th deployment according to the filtered video quality and the original video quality includes: Determine the differential video quality between the filtered video quality and the original video quality; If the differential video quality is less than the differential quality threshold, determine that the second video encoding application meets the filtering quality requirement condition; If the differential video quality is greater than the differential quality threshold, determine that the second video encoding application does not meet the filtering quality requirement condition.
11. A data processing device, wherein, it includes: A model training module, configured to obtain a video encoding application whose filtering performance does not meet the filtering quality requirement condition; a first filtering model is deployed in the video encoding application; A training data generation module, further configured to perform encoding processing on sample video data through the video encoding application, and train the first filtering model based on the data without filtering processing generated by the encoding processing and the sample video data; An application update module, configured to deploy and update the trained first filtering model into the video encoding application to obtain a video encoding application after deployment and update; When the filtering performance of the video encoding application after deployment and update meets the filtering quality requirement condition, it is configured to perform video encoding processing on video data.
12. A computer device, wherein, it includes: A processor, a memory, and a network interface; The processor is connected to the memory and the network interface, wherein the network interface is used to provide network communication functions, the memory is used to store program codes, and the processor is used to call the program codes so that the computer device executes the method according to any one of claims 1-10.
13. A computer-readable storage medium, wherein, a computer program is stored in the computer-readable storage medium, and the computer program is suitable for being loaded and executed by a processor to execute the method according to any one of claims 1-10.