Video coding method and apparatus

By extracting features and determining target encoding parameters from multiple videos to be encoded, the problem of low encoding efficiency in existing technologies is solved, achieving the effect of reducing bandwidth costs while ensuring video quality.

WO2026021082A1PCT designated stage Publication Date: 2026-01-29TAOBAO CHINA SOFTWARE
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2025/102686
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-07-26
Filing Date
2025-06-23
Publication Date
2026-01-29

AI Technical Summary

Technical Problem

Existing video coding technologies require multiple encoding processes to determine encoding parameters, resulting in low video coding efficiency and difficulty in maintaining video quality while reducing bandwidth costs.

Method used

By extracting features from multiple videos to be encoded, the target encoding parameters are determined based on the total bitrate and video features. The encoding process is then carried out using the target encoding parameters to ensure that the bitrate and quality evaluation indicators of the encoded video meet the requirements.

Benefits of technology

It improves video encoding efficiency, ensuring that the sum of the bitrates of the encoded videos is less than the total bitrate, and the sum of the quality assessment index values ​​is greater than the standards of the video demand side, thereby improving the quality of video encoding.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2025102686_29012026_PF_FP_ABST
    Figure CN2025102686_29012026_PF_FP_ABST
Patent Text Reader

Abstract

Embodiments of the present disclosure provide a video coding method and apparatus. The method comprises: upon extracting video features of videos to be coded among a plurality of videos to be coded, on the basis of a total bitrate corresponding to the plurality of videos to be coded and the video features of the videos to be coded, determining target coding parameters of the videos to be coded; and then using the target coding parameters to perform coding processing on the videos to be coded, so as to obtain coded videos. In the process, the target coding parameters of the videos to be coded can be determined directly by means of the total bitrate corresponding to the plurality of videos to be coded and the video features of the videos to be coded, without the need to perform pre-coding processing on the videos to be coded to obtain coding features of videos having undergone pre-processing. Therefore, the video coding quality is ensured while improving the efficiency of determining the target coding parameters of the videos to be coded.
Need to check novelty before this filing date? Find Prior Art

Description

Video encoding method and device

[0001] The present disclosure claims priority to Chinese Patent Application No. 202411023034.2, filed on July 26, 2024, with the Chinese Patent Office, entitled "Video encoding method and device", the entire contents of which are incorporated herein by reference. TECHNICAL FIELD

[0002] The present disclosure relates to the technical field of computer, in particular to a video encoding method and device. BACKGROUND

[0003] With the development of Internet technology, high-definition video applications (on-demand and live) are becoming more and more popular, but the problem is higher bandwidth cost and poorer playback fluency. Therefore, if the video transmission bandwidth cost is reduced while the video quality is guaranteed, it is a problem that needs to be concerned.

[0004] In the prior art, video encoding is usually performed based on different encoding parameter values, and the video is encoded multiple times to determine the target encoding parameter value, and then formal video encoding is performed, so as to reduce the bandwidth cost while guaranteeing the video quality.

[0005] However, the existing technology determines the encoding parameters after multiple encoding processes, which reduces the video encoding efficiency. SUMMARY

[0006] The embodiments of the present disclosure provide a video encoding method to improve the video encoding efficiency while guaranteeing the video encoding quality. The embodiments of the present disclosure also provide a video encoding device, an electronic device and a computer storage medium.

[0007] The embodiments of the present disclosure provide a video encoding method, comprising: performing feature extraction processing on a plurality of to-be-encoded videos to obtain video features of a to-be-encoded video in the plurality of to-be-encoded videos; determining a target encoding parameter of the to-be-encoded video based on a total code rate corresponding to the plurality of to-be-encoded videos and the video features of the to-be-encoded video in the plurality of to-be-encoded videos; and performing encoding processing on the to-be-encoded video according to the target encoding parameter of the to-be-encoded video; wherein when the encoding parameter of the to-be-encoded video in the plurality of to-be-encoded videos is the target encoding parameter corresponding thereto, the code rate of the encoded video after the encoding processing is a target code rate, the quality evaluation index value of the encoded video after the encoding processing is a target quality evaluation index value, the sum of the target code rates corresponding to the encoded videos respectively is less than the total code rate, and the sum of the target quality evaluation index values corresponding to the encoded videos respectively is greater than a video quality evaluation index standard value required by a video demand end for the plurality of encoded videos.

[0008] Optionally, the determining the target encoding parameter of the to-be-encoded video based on the total code rate corresponding to the plurality of to-be-encoded videos and the video feature of the to-be-encoded video comprises: determining a video content type to which the to-be-encoded video belongs according to the video feature of the to-be-encoded video; classifying the to-be-encoded video according to the video content type to which the to-be-encoded video belongs, to obtain at least one to-be-encoded video group, the video content types of the to-be-encoded videos in the to-be-encoded video group belong to the same video content type; determining the target encoding parameter of the to-be-encoded video group based on the total code rate corresponding to the plurality of to-be-encoded videos and the video content type of the to-be-encoded video group, as the target encoding parameter of the to-be-encoded video in the to-be-encoded video group; and correspondingly, the encoding processing of the to-be-encoded video according to the target encoding parameter of the to-be-encoded video comprises: jointly encoding processing the to-be-encoded video contained in the to-be-encoded video group according to the target encoding parameter of the to-be-encoded video group.

[0009] Optionally, the determining the target encoding parameter of the to-be-encoded video based on the total code rate corresponding to the plurality of to-be-encoded videos and the video feature of the to-be-encoded video comprises: determining an encoding parameter prediction model corresponding to the plurality of to-be-encoded videos based on the total code rate corresponding to the plurality of to-be-encoded videos; inputting the video feature of the to-be-encoded video into the encoding parameter prediction model respectively to obtain the target encoding parameter of the to-be-encoded video output by the encoding parameter prediction model; wherein the encoding parameter prediction model is used to determine the target encoding parameter of the to-be-encoded video according to the video feature of the to-be-encoded video, and the target encoding parameter satisfies the following conditions: after the to-be-encoded video is encoded by using the target encoding parameter, a target code rate and a target quality evaluation index value of an encoded video are obtained; and a sum of the target code rates corresponding to the plurality of encoded videos is less than the total code rate, and a sum of the target quality evaluation index values corresponding to the plurality of encoded videos is greater than a video quality evaluation index standard value required by a video demand end for the plurality of encoded videos.

[0010] Optionally, the encoding parameter prediction model is trained by: obtaining total code rates corresponding to a plurality of sample videos, and sample video features and sample encoding parameters of sample videos in the plurality of sample videos; inputting the total code rates corresponding to the plurality of sample videos and the sample video features of the sample videos into an initial encoding parameter prediction model to obtain initial encoding parameters of the sample videos output by the initial encoding parameter prediction model; obtaining loss values between the initial encoding parameters of the sample videos and the sample encoding parameters; summing the loss values corresponding to the sample videos in the plurality of sample videos to obtain total loss values corresponding to the plurality of sample videos; and adjusting parameters of the initial encoding parameter prediction model based on the total loss values to obtain the trained encoding parameter prediction model.

[0011] Optionally, the sample encoding parameter of the sample video is obtained by: obtaining a plurality of candidate rate-distortion point groups of a sample video in a plurality of candidate encoding parameters; selecting a predetermined rate-distortion point in the candidate rate-distortion point group corresponding to the sample video, the predetermined rate-distortion point including at least the following attribute parameters: a predetermined encoding parameter, a predetermined code rate, and a predetermined quality evaluation index value; encoding the sample video using the predetermined encoding parameter, so that the code rate of the encoded video is the predetermined code rate, and the quality evaluation index value of the encoded video is the predetermined quality evaluation index value; and if the above attribute parameters of the predetermined rate-distortion points corresponding to the sample videos in the plurality of sample videos satisfy the following conditions, the predetermined encoding parameter in the predetermined rate-distortion point of the sample video is taken as the sample encoding parameter of the sample video: the sum of the predetermined code rates corresponding to the encoded videos in the plurality of sample videos is less than the total code rate, and the sum of the predetermined quality evaluation index values corresponding to all encoded videos is greater than a video quality evaluation index standard value required by a video demand end for a plurality of encoded videos.

[0012] Optionally, the obtaining of the plurality of candidate rate-distortion point groups of a sample video in a plurality of candidate encoding parameters includes: encoding the sample video using a candidate encoding parameter to obtain a code rate and a quality evaluation index value of an encoded video after the encoding processing under the candidate encoding parameter; obtaining the code rate and the quality evaluation index value corresponding to the encoding processing of the sample video by a plurality of candidate encoding parameters; taking the candidate encoding parameter and the code rate and the quality evaluation index value corresponding to the candidate encoding parameter as a candidate rate-distortion point of the encoded video after the encoding processing; and taking the candidate rate-distortion points corresponding to the plurality of candidate encoding parameters as the plurality of candidate rate-distortion point groups of the sample video in the plurality of candidate encoding parameters.

[0013] Optionally, the method further comprises: obtaining a video quality evaluation index standard value required by the video demand end for the plurality of encoded videos; and determining the target encoding parameter of the to-be-encoded video based on the total code rate corresponding to the plurality of to-be-encoded videos, the video features of the to-be-encoded video in the plurality of to-be-encoded videos, and the video quality evaluation index standard value, comprises: determining the target encoding parameter of the to-be-encoded video based on the total code rate corresponding to the plurality of to-be-encoded videos, the video features of the to-be-encoded video in the plurality of to-be-encoded videos, and the video quality evaluation index standard value.

[0014] Optionally, the method further comprises: obtaining a video quality evaluation index standard value required by the video demand end for the plurality of encoded videos; and determining the target encoding parameter of the to-be-encoded video based on the total code rate corresponding to the plurality of to-be-encoded videos, the video features of the to-be-encoded video in the plurality of to-be-encoded videos, and the video quality evaluation index standard value, comprises: predicting a predicted bit rate required by the to-be-encoded video according to the video features of the to-be-encoded video; determining a predicted encoding parameter required by the to-be-encoded video according to the predicted bit rate required by the to-be-encoded video; obtaining a predicted code rate and a predicted quality evaluation index value of the video of the to-be-encoded video after encoding by the predicted encoding parameter; and if the sum of the predicted code rates corresponding to the plurality of encoded videos is less than the total code rate, and the sum of the predicted quality evaluation index values corresponding to the plurality of encoded videos is greater than the video quality evaluation index standard value, then the predicted encoding parameter is taken as the target encoding parameter of the to-be-encoded video.

[0015] Optionally, the predicted encoding parameter comprises a plurality of candidate encoding parameters; and the method further comprises: obtaining a plurality of predicted code rates and a plurality of predicted quality evaluation index values corresponding to the to-be-encoded video in the plurality of to-be-encoded videos by respectively encoding the to-be-encoded video using the plurality of candidate encoding parameters; and if the predicted code rates and the predicted quality evaluation index values corresponding to the plurality of candidate encoding parameters satisfy the following conditions: the sum of the predicted code rates corresponding to the plurality of encoded videos is less than the total code rate, and the predicted quality evaluation index values corresponding to the plurality of encoded videos are less than the video quality evaluation index standard value, then a target encoding parameter is selected from the plurality of candidate encoding parameters.

[0016] Optionally, the selecting the target coding parameter from the plurality of candidate coding parameters comprises: encoding the to-be-encoded video by using the plurality of candidate coding parameters respectively to obtain a plurality of predicted quality evaluation index values of the to-be-encoded video after encoding; comparing the plurality of predicted quality evaluation index values of the to-be-encoded video after encoding to obtain a target predicted quality evaluation index value, the target predicted quality evaluation index value being greater than other predicted quality evaluation index values except the target predicted quality evaluation index value in the plurality of predicted quality evaluation index values; and taking a candidate coding parameter corresponding to the target predicted quality evaluation index value as the target coding parameter of the to-be-encoded video.

[0017] The embodiments of the present disclosure further provide an electronic device, comprising: a processor; and a memory for storing a computer program, wherein the electronic device is powered on and executes the computer program by using the processor to perform the above method.

[0018] The embodiments of the present disclosure further provide a storage medium, wherein the storage medium stores a computer program, and the computer program is executed by a processor to perform the above method.

[0019] The embodiments of the present disclosure provide a computer program product, comprising: a computer program, when the computer program is executed by a processor of an electronic device, the processor executes the above method.

[0020] Compared with the prior art, the embodiments of the present disclosure have the following advantages:

[0021] The video coding method provided by the embodiments of the present disclosure can determine the target coding parameter of the to-be-coded video according to the total code rate corresponding to the plurality of to-be-coded videos and the video features of the to-be-coded video after extracting the video features of the plurality of to-be-coded videos. Then, the to-be-coded video is coded by using the target coding parameter to obtain the coded video. In this process, the target coding parameter of the to-be-coded video can be determined by using the total code rate corresponding to the plurality of to-be-coded videos and the video features of the to-be-coded video, without the need to obtain the coding features of the pre-processed video after pre-coding the to-be-coded video. Therefore, the efficiency of determining the target coding parameter of the to-be-coded video is improved. Moreover, the method provided by the present disclosure can determine the target coding parameter of the to-be-coded video to meet the following conditions: the code rate of the coded video obtained by coding the to-be-coded video by using the target coding parameter is the target code rate, and the quality evaluation index value of the coded video is the target quality evaluation index value; the sum of the target code rates corresponding to the plurality of coded videos in the plurality of to-be-coded videos is less than the total code rate, and the sum of the target quality evaluation index values corresponding to the plurality of coded videos is greater than the quality evaluation index standard value required by the video demand end for the coded video. Therefore, the video coding method provided by the present disclosure can improve the video coding efficiency while ensuring the video coding quality. BRIEF DESCRIPTION OF DRAWINGS DETAILED DESCRIPTION OF THE INVENTION BRIEF DESCRIPTION OF DRAWINGS

[0022] FIG. 1 is a flowchart of a video coding method provided by a first embodiment of the present disclosure.

[0023] FIG. 2 is a schematic diagram of a pixel pair of Pd(i, j) in a gray level co-occurrence matrix provided by the prior art.

[0024] FIG. 3 is a schematic diagram of two filter operators in spatial domain information.

[0025] FIG. 4 is a schematic diagram of a process in which an encoding parameter prediction model determines the target coding parameters of a plurality of to-be-coded videos and codes the to-be-coded videos according to the target coding parameters.

[0026] FIG. 5 is a schematic diagram of a video coding device provided by a second embodiment of the present disclosure.

[0027] FIG. 6 is a schematic diagram of an electronic device provided by a third embodiment of the present disclosure. DETAILED DESCRIPTION

[0028] In the following description, a large number of specific details are set forth in order to facilitate a thorough understanding of the present disclosure. However, the present disclosure can be implemented in many different ways than those described herein, and those skilled in the art can make similar extensions without departing from the spirit of the present disclosure, so the present disclosure is not limited to the specific implementations disclosed below.

[0029] The terminology used in the present disclosure is for the purpose of describing particular embodiments only and is not intended to be limiting of the present disclosure. As used in this present disclosure and the appended claims, the description will be described, for example, "an", "one", and "the" are not intended to limit the number or modify an ordinal position of the preceding words, but are used to distinguish a particular entity from one or more other entities.

[0030] Embodiments of the present disclosure provide a video encoding method and device. In the following embodiments, they are described one by one.

[0031] For ease of understanding, first give some related concepts and terms involved in the embodiments of the present disclosure.

[0032] For a batch of videos in the transmission process, the batch of video files will be subject to the same bandwidth limitation, where the bandwidth is the maximum rate of video transmission, that is, the amount of data that can be transmitted per unit time. The total bandwidth of the batch of video files in the transmission process can also be determined under the condition of the total code rate. How to determine the target encoding parameters of each video file to ensure that the sum of the code rates corresponding to each video in the batch of videos is less than the total code rate of the batch of videos. Under the condition of the total code rate, it can also make the sum of the target quality evaluation index values corresponding to each video greater than the sum of the candidate quality evaluation index values under other encoding parameter conditions.

[0033] The following explains the terms involved in the embodiments of the present disclosure respectively:

[0034] Video encoding: the process of converting video signals into digital codes to effectively use bandwidth and storage space in storage and transmission. It involves decomposing video signals into spatial and temporal encodable units and applying compression algorithms to reduce data volume.

[0035] Encoding efficiency: the data compression ratio that a video encoding algorithm can achieve under the same picture quality. High encoding efficiency means that it can encode with more code rate, thereby reducing the demand for storage space and transmission bandwidth.

[0036] Support vector regression (SVR): a machine learning algorithm used to establish a regression relationship between input variables and continuous target variables. In video encoding, SVR can be used to establish a rate control model to dynamically adjust encoding parameters according to video content and quality requirements, thereby achieving higher encoding efficiency and higher video quality.

[0037] An I-frame (Intra-coded frame), also known as a keyframe or intra-frame, is an independent frame in a video sequence. An I-frame does not rely on information from other frames for decoding; it contains complete image information. In a video sequence, an I-frame is typically inserted at regular intervals (e.g., per second). An I-frame can be decoded into a complete image, making it very useful for video editing and random access.

[0038] P-frame (Predictive-coded frame): Also known as a predictive frame, it is a frame type relative to I-frames. P-frames are encoded using information from decoded preceding reference frames (usually previous I-frames and P-frames). P-frames only contain the difference information between the current frame and the reference frame, not complete image information. This difference information can be represented using motion compensation techniques to reduce data volume and improve coding efficiency.

[0039] Quantization Parameter (QP): Used to control the quantization process in video coding. Quantization is the process of converting a continuous video signal into a discrete numerical representation. By controlling the QP value, the quantization precision can be adjusted, thereby affecting coding efficiency and video quality.

[0040] Constant Rate Factor (CRF): A bitrate control mode used to adjust the encoding bitrate while maintaining constant video quality. By adjusting the CRF value, different encoding efficiencies and file sizes can be achieved while maintaining relatively consistent video quality.

[0041] Rate-distortion curve: Describes the change in compression quality at different bitrates. During data compression, to reduce data size, some image or audio quality is usually sacrificed. The rate-distortion curve describes the degree of distortion between the compressed data and the original data at different compression ratios (bitrates). A typical rate-distortion curve plots bitrate (compression ratio) on the x-axis and distortion (quality loss) on the y-axis. Generally, as the compression ratio increases, distortion also increases, meaning the quality of the compressed data decreases. This is because higher compression ratios lead to more data loss or information loss.

[0042] Group knapsack problem: an optimization problem solving algorithm based on dynamic programming, used to solve a variant of the knapsack problem. In the traditional knapsack problem, there is a set of items and a knapsack, each item has a weight and a value, the goal is to select some items to put into the knapsack, so that their total weight does not exceed the capacity of the knapsack, while the total value is maximized. While in the group knapsack problem, the items are divided into several groups, each group can only choose one item in the group to put into the knapsack, the goal is to select one item from each group, so that their total weight does not exceed the capacity of the knapsack, while the total value is maximized.

[0043] It should be noted that the above disclosed information is only used to help understand the present disclosure, and does not mean to constitute the prior art known to those skilled in the art.

[0044] The video encoding method provided by the first embodiment of the present disclosure is described below in combination with FIG. 1 to FIG. 4. FIG. 1 is a flowchart of a video encoding method provided by the first embodiment of the present disclosure. The video encoding method shown in FIG. 1 comprises:

[0045] Step S101: performing feature extraction processing on a plurality of to-be-encoded videos to obtain video features of a to-be-encoded video in the plurality of to-be-encoded videos;

[0046] Step S102: determining a target encoding parameter of the to-be-encoded video based on a total code rate corresponding to the plurality of to-be-encoded videos and the video features of the to-be-encoded video in the plurality of to-be-encoded videos;

[0047] Step S103: encoding processing the to-be-encoded video according to the target encoding parameter of the to-be-encoded video; wherein when the encoding parameter of the to-be-encoded video in the plurality of to-be-encoded videos is the target encoding parameter corresponding thereto, the code rate of the encoded video after encoding processing is the target code rate, the quality evaluation index value of the encoded video after encoding processing is the target quality evaluation index value, and the sum of the target code rates corresponding to the encoded videos after encoding processing is less than the total code rate, and the sum of the target quality evaluation index values corresponding to the encoded videos after encoding processing is greater than the video quality evaluation index standard value required by the video demand end for the plurality of encoded videos.

[0048] The method provided by the embodiments of the present disclosure can be applied to the process that a video publisher uploads a batch of videos to a video platform. The client sends the batch of videos to the server. The server determines the target encoding parameter of each video in the batch of videos based on the method provided by the present disclosure. After encoding processing each video based on the target encoding parameter, the video is transmitted to the video platform. The method can also be applied to the process that the video platform selects a batch of videos for a target user. Before showing the batch of videos to the target user, the client of the video platform provides the batch of videos to the server. The server determines the target encoding parameter of each video in the batch of videos based on the method provided by the present disclosure. After encoding processing each video based on the target encoding parameter, the video is transmitted to the client to show to the target user.

[0049] The method provided by the embodiments of the present disclosure determines the target encoding parameter of the to-be-encoded video in the plurality of to-be-encoded videos based on the total code rate corresponding to the plurality of to-be-encoded videos and the video feature of the to-be-encoded video. The method does not need to determine the encoding feature data by pre-encoding the to-be-encoded video, and obtains the target encoding parameter by multiple encoding, thereby improving the efficiency of determining the target encoding parameter of the plurality of to-be-encoded videos. Moreover, after the target encoding parameter obtained by the video encoding method provided by the embodiments of the present disclosure is used to encode the to-be-encoded video, the encoded video can satisfy the following conditions: the sum of the code rates corresponding to the plurality of encoded videos in the plurality of to-be-encoded videos is less than the total code rate corresponding to the plurality of to-be-encoded videos, and the sum of the quality evaluation index values of the encoded videos is greater than the quality evaluation index standard value required by the video demand end for the plurality of encoded videos. Therefore, the video encoding method provided by the embodiments of the present disclosure improves the efficiency of determining the target encoding parameter of the plurality of to-be-encoded videos, while ensuring that the sum of the code rates of the encoded videos is less than the total code rate and the video quality of the encoded videos, that is, improving the encoding efficiency while ensuring the video encoding quality.

[0050] In the step S101, the video feature of the to-be-encoded video is determined by performing feature extraction processing on the plurality of to-be-encoded videos. Specifically, the video feature of each to-be-encoded video is extracted first, and then the target encoding parameter of the to-be-encoded video is determined. The plurality of to-be-encoded videos can be a plurality of to-be-encoded videos having a correlation relationship, for example, the plurality of to-be-encoded videos are all to-be-encoded videos about food introduction. The plurality of to-be-encoded videos can also be a plurality of to-be-encoded segment videos of a complete video, and the video features contained in each segment video are different. The plurality of to-be-encoded videos includes at least two to-be-encoded videos, and more specifically, each to-be-encoded video can further include a plurality of to-be-encoded segment videos.

[0051] The plurality of to-be-encoded videos can also be a plurality of to-be-encoded videos that are not associated with each other, for example, the plurality of to-be-encoded videos include three to-be-encoded videos, the first to-be-encoded video is a to-be-encoded video with a static picture frame ratio greater than a preset threshold, the second to-be-encoded video is a to-be-encoded video with a dynamic video frame ratio greater than a preset threshold, and the third to-be-encoded video is a to-be-encoded video with a content richness in a video frame greater than a preset threshold. Therefore, for the above three types of to-be-encoded videos, the corresponding video features are different, and accordingly, the target encoding parameters required for the to-be-encoded video in the subsequent steps are also different.

[0052] First, the process of extracting the video features of the to-be-encoded video is specifically described as follows:

[0053] Specifically, the to-be-encoded video is respectively extracted from the spatial domain and the time domain to obtain the spatial domain features and the derived features of the spatial domain features of the to-be-encoded video, and the time domain features and the derived features of the time domain features of the to-be-encoded video.

[0054] The spatial domain features of the to-be-encoded video include at least one of the following features: a grey level co-occurrence matrix GLCM for representing the texture features of an image; spatial information SI for representing the amount of spatial details in the image; and colorfulness CF for representing the saturation or intensity of colors in the image. The features are described one by one as follows:

[0055] First, the first aspect feature in the spatial domain features, the grey level co-occurrence matrix GLCM, is described as follows: the grey level co-occurrence matrix GLCM is used to describe the texture features of an image, specifically, the image is converted into an image with a grey level of L, then a grey level co-occurrence matrix is generated according to the relationship between the grey values of the pixel points, and a secondary statistic is obtained according to the grey level co-occurrence matrix to describe the texture features of the image.

[0056] The size of the grey level co-occurrence matrix is LxL, and each value (i, j) of the grey level co-occurrence matrix corresponds to the frequency of occurrence of the grey level j of the point reached after a fixed distance d=(D x , D y ) from the point with a grey level of i in the same direction. Wherein, P d (i, j) is represented by formula 1. P d (i, j) is represented by formula 1. P d (i, j) = P d (i, j | d, θ) (i, j = 0, 1, 2, …, L-1) Formula 1

[0057] Wherein, i, j respectively represent the gray value of the pixel, d represents the distance relationship between two points, θ represents the angle relationship between two points, and the value is usually 0°, 45°, 90°, 135°, which can be referred to as Fig. 2, which is a pixel pair of the GLCM P d (i, j) in the prior art.

[0058] Based on formula 1, when the distance relationship between two points is d and the angle relationship between two points is θ, the GLCM is Wherein, It is represented by formula 2.

[0059] In order to facilitate the analysis of each element in the GLCM, each element in the GLCM is normalized to obtain a normalized GLCM. Wherein, the specific way of normalizing each element is: dividing each element P d (i, j) in the GLCM by the sum of all elements in the GLCM to obtain the average value of each element Then, the second order statistics of the normalized GLCM are extracted, as follows:

[0060] Wherein, the second order statistics of the GLCM GLCM include at least one of the following:

[0061] (1) Entropy GLCM of the GLCM ent (Wherein, ent refers to the abbreviation of entropy), which is used to represent the complexity or non-uniformity of the image texture feature. The greater the entropy of the GLCM, the more complex the texture information in the image; on the contrary, the simpler the texture information in the image. Wherein, the entropy GLCM of the GLCM ent It is calculated by formula 3:

[0062] (2) Correlation GLCM of the GLCM cor (Wherein, cor refers to the abbreviation of correlation), which is used to represent the similarity of the elements of the GLCM in the column or row direction. When the element values in the GLCM are uniform and equal, the greater the value of the element correlation; on the contrary, when the elements in the matrix are not uniform, the smaller the value of the element correlation. Wherein, the correlation GLCM of the GLCM cor It is calculated by formula 4:

[0063] Wherein,

[0064] ​(3) Contrast of GLCM con (Wherein, con refers to the abbreviation of Contrast), used to measure how the element values of the matrix are distributed and the degree of local change in the image, representing the clarity of the image and the depth of the texture groove. The greater the contrast value, the deeper the texture groove, the greater the contrast, the clearer the effect; on the contrary, the smaller the contrast value, the shallower the texture groove, the more blurred the effect. Among them, the contrast of GLCM con It is calculated by formula 5:

[0065] (4) Homogeneity of GLCM hom (Wherein, hom refers to the abbreviation of Homogeneity), used to represent the degree of local change of image texture. The higher the homogeneity, the more uniform the local image, and the lower the homogeneity, the more uneven the local image. Among them, the homogeneity of GLCM hom It is calculated by formula 6:

[0066] (5) Energy of GLCM ene (Wherein, ene refers to the abbreviation of Energy), used to represent the uniformity of image gray scale distribution and the thickness of texture. Among them, the energy of GLCM ene It is calculated by formula 7:

[0067] The following explains the spatial and temporal features in combination with Table 1:

[0068] Table 1

[0069] Among them, as shown in Table 1, (1) the entropy of GLCM ent including the average value of the entropy of GLCM (F9, GLCM ent _mean, mean represents the average value): used to represent the average value between the values of the complexity or non-uniformity of the image texture feature; the standard deviation of the entropy of GLCM (F10, GLCM ent _std, std is the abbreviation of Standard Deviation): used to represent the standard deviation value between the values of the complexity or non-uniformity of the image texture feature.

[0070] (2) the correlation of GLCM cor including the average value of the correlation of GLCM (F1, GLCM corMean): the average value between the values used to characterize the degree of similarity of the elements of the gray level co-occurrence matrix in the column or row direction; the correlation standard deviation of the gray level co-occurrence matrix (F2, GLCM cor Std): the standard deviation value between the values used to characterize the degree of similarity of the elements of the gray level co-occurrence matrix in the column or row direction.

[0071] (3) Contrast of the gray level co-occurrence matrix GLCM con including the average value of the contrast of the gray level co-occurrence matrix (F3, GLCM con Mean): the average value between the values used to characterize the degree of clarity and the depth of the texture of the image; the contrast standard deviation of the gray level co-occurrence matrix (F4, GLCM con Std): the standard deviation value between the values used to characterize the degree of clarity and the depth of the texture of the image.

[0072] (4) Homogeneity of the gray level co-occurrence matrix GLCM hom including the average value of the homogeneity of the gray level co-occurrence matrix (F7, GLCM hom Mean): the average value between the values used to characterize the degree of local variation of the texture of the image; the homogeneity standard deviation of the gray level co-occurrence matrix (F8, GLCM hom Std): the standard deviation value between the values used to characterize the degree of local variation of the texture of the image.

[0073] (5) Energy of the gray level co-occurrence matrix GLCM ene including the average value of the energy of the gray level co-occurrence matrix (F5, GLCM ene Mean): the average value between the values used to characterize the uniformity of the gray level distribution of the image and the thickness of the texture; and the energy standard deviation of the gray level co-occurrence matrix (F6, GLCM ene Std): the standard deviation value between the values used to characterize the uniformity of the gray level distribution of the image and the thickness of the texture.

[0074] The above is the feature description information contained in the quadratic statistics of the gray level co-occurrence matrix.

[0075] The following describes the second aspect feature in the spatial domain feature: the spatial information SI representing the amount of spatial details in a frame of image, wherein the spatial domain refers to the distribution of the image or signal in the spatial position, corresponding to the time domain, which refers to the change of the signal with time. The more complex the scene is in space, the higher the SI value is.

[0076] wherein the calculation method of SI is as formula 8: SI = σ{G(x, y)} Formula 8

[0077] wherein σ denotes the variance of G(x, y);

[0078] G(x, y) is represented by formula 9:

[0079] wherein (x, y) denotes a pixel point in the image, G 1 , G 2 are images obtained by applying two filter operators to the image, wherein the two filter operators are represented by FIG. 3, which is a schematic diagram of two filter operators in spatial information. In FIG. 3, 301 denotes G 1 , and 302 denotes G 2 .

[0080] If the SI value is large, it indicates that the image contains a lot of edge information.

[0081] wherein, as shown in Table 1, the derived features of the spatial information SI include the following representations:

[0082] (1) Mean value of spatial information (F26, SI_mean): Spatial information (SI) refers to the gray value of each pixel point in the image or the intensity value of the signal in the spatial position. These values carry basic information of the image content or signal characteristics, such as brightness, color distribution, texture features, etc.

[0083] The mean value of spatial information (SI_mean) is calculated by performing an arithmetic mean on the gray values of all pixel points in the image to obtain a single numerical value to represent the average brightness or signal intensity level of the entire image.

[0084] (2) Standard deviation value of spatial information (F27, SI_std): used to represent the degree of change or dispersion of the gray values of all pixel points in the image relative to the mean value. That is, it represents an index of the uniformity or variability of the data distribution of all pixel points in the image to determine the spatial heterogeneity of all pixel points data or the clarity, noise level of the image.

[0085] (3) Mean value between mean values of spatial information (F28, mean_SI_mean): used to represent the similarity of spatial information of different regions in the image after calculating the mean value of spatial information of each region in the image.

[0086] or to represent the similarity of the image at different time points after calculating the mean value of spatial information of the same region at different time points.

[0087] (4) The standard deviation value between the average values of the spatial information (F29, mean SI std) is used to represent the fluctuation or deviation between the spatial information of different regions in the image. After the average value of the spatial information is calculated for the image, the standard deviation of the average values of different regions in the image is calculated to determine the fluctuation or deviation between the spatial information of different regions.

[0088] Or, the average value of the spatial information is calculated for the same region at different time points, and then the standard deviation of the corresponding average values at different time points is calculated again to determine the fluctuation or deviation between the spatial information of the image at different time points.

[0089] The following describes the third aspect feature in the spatial feature: colorfulness CF. The colorfulness feature mainly describes the saturation or intensity of the color in the image. By analyzing the colorfulness, the richness, intensity and color distribution of the image color can be understood. The colorfulness is calculated by formula 10: CF = σ rgyb + 0.3 * μ rgyb Formula 10

[0090] Wherein, rg(i,j) = |R(i,j) - B(i,j)|, yb(i,j) = |0.5*R(i,j) + G(i,j) - B(i,j)|

[0091] Wherein, μ rgyb represents the average color factor, σ rgyb represents the variance color factor.

[0092] R(i,j), G(i,j), B(i,j) represent the values of the (i,j)th pixel point in the RGB color space R, G, B of an image. μ rg represents the mean value of rg(i,j), μ yb represents the mean value of yb(i,j). σ rg represents the standard deviation of rg(i,j), σ yb represents the standard deviation of yb(i,j).

[0093] Wherein, as shown in Table 1, the derived features of colorfulness CF include the following several representations:

[0094] (1) The average value of colorfulness (F22, CF_mean): the colorfulness values of all pixel points in the image are summed, and then divided by the total number of pixels to obtain the average value of colorfulness (CF_mean), to represent the overall average colorfulness level of the image.

[0095] (2) Standard deviation of colorfulness (F23, CF_std): used to represent the difference between the colorfulness value of all pixels in the image and the corresponding average value.

[0096] The above is the description process of the three aspects of spatial features and their derived features.

[0097] The following describes the temporal features of the video to be encoded:

[0098] The temporal features of the video to be encoded include at least one of the following features: temporal information TI (Temporal Information), used to represent the amount of temporal change of the video sequence; normalized cross correlation NCC (Normalized Cross Correlation), used to represent the correlation between two adjacent video frames; temporal coherence TC (Temporal Coherence), used to represent the correlation relationship of the time domain signals of two adjacent video frames.

[0099] The following first describes the first aspect of the temporal features: temporal information TI, used to represent the amount of temporal change of the video sequence. The higher the TI value of the video sequence with higher motion degree, wherein the temporal information TI is calculated by taking the frame difference of the image brightness component of the nth frame and the (n-1)th frame. Specifically, it is calculated by formula 11:

[0100] wherein Luma n (x,y) represents the brightness component of the (x, y) point position of the nth image; Luma n-1 (x,y) represents the brightness component of the (x, y) point position of the (n-1)th image.

[0101] Wherein, as shown in Table 1, the derived features of the temporal information TI include the following representations:

[0102] (1) Mean of temporal information (F24, TI_mean), (2) Standard deviation of temporal information (F25, TI_std).

[0103] The following describes the second aspect of the temporal features: normalized cross correlation NCC, an algorithm for calculating the correlation between two groups of sample data based on statistics, with a value range of [-1, 1]. When the value of NCC is 1, it indicates a high correlation, and when the value of NCC is -1, it indicates no correlation. In the embodiments of the present disclosure, the image frame is first converted into a grayscale image, and then the correlation between two adjacent image frames is calculated. Wherein, the normalized cross correlation NCC is calculated by formula 12:

[0104] wherein f n(i,j) represents the value of the gray scale map of the nth video frame at the (i,j) coordinate value; H, W represent the height and width of the video frame, respectively.

[0105] Among them, as shown in Table 1, the derived features of the normalized cross-correlation NCC include the following representations:

[0106] (1) the average value of the normalized cross-correlation (F20, NCC_mean), (2) the standard deviation of the normalized cross-correlation (F21, NCC_std).

[0107] The following describes a third aspect feature of the time domain feature: time domain coherence TC, which is used to measure the correlation between two signals, and its value range is [0, 1], 0 represents no correlation. Among them, the calculation of the normalized cross-correlation is realized by the following way: the cross power spectrum density between the two signals and the power spectrum density of each other are calculated, and then the cross power spectrum density between the two signals and the power spectrum density of each other are combined, that is, the value of the normalized cross-correlation can be obtained. Specifically, the normalized cross-correlation can be calculated by the following formula:

[0108] Formula 13 is realized:

[0109] Among them, |P xy (μ,v)| 2 represents the square of the cross power spectrum modulus between the two signals; P xx (μ,v), P yy (μ,v) respectively represent the self power spectrum, P xy (μ,v) represents the cross power spectrum.

[0110] Among them, the self power spectrum and the cross power spectrum are calculated by the formula group provided in formula 14:

[0111] Among them, F x (μ,v), F y (μ,v) represent the Fourier transform of the two signals, respectively represent the conjugate function of the Fourier transform of the two signals.

[0112] Among them, as shown in Table 1, the derived features of the time domain coherence TC include the following representations:

[0113] (1) the average value between the average values of the time domain coherence (F11, mean_TC_mean): used to represent the average level of a group of time domain coherence average value data as a whole. It is calculated by the following way:

[0114] Step 1, calculate the average value mean_TC of a group of time domain coherence data:

[0115] Time-domain coherence (TC) is used to measure the correlation between two signals. A group contains 10 signal pairs, each of which corresponds to a time-domain coherence data. The average value mean_TC of the 10 time-domain coherence data is calculated to represent the overall time-domain coherence level of the group.

[0116] Step 2, calculate the average value mean_TC_mean between the average values of multiple groups of time-domain coherence data:

[0117] Obtain multiple groups of signal pairs, for example, 3 groups of signal pairs, each group containing 10 signal pairs. Based on the above method, obtain the average value of the time-domain coherence data of each group of signal pairs, then the three groups of signal pairs can obtain 3 time-domain coherence data average values. Take the 3 time-domain coherence data average values as a group of time-domain coherence average values. Divide the group number (3) by the group of time-domain coherence average values to obtain the average value between the time-domain coherence average values (mean_TC_mean). Through mean_TC_mean, the overall level of the time-domain coherence data of the 3 groups of signal pairs is reflected.

[0118] (2) The standard deviation between the average values of time-domain coherence (F12, mean_TC_std): a statistical quantity used to represent the dispersion degree of a group of time-domain coherence average data, that is, to reflect the volatility or change range of the time-domain coherence average data. It is calculated as follows:

[0119] Step 1, calculate the average value mean_TC of a group of time-domain coherence data:

[0120] Time-domain coherence (TC) is used to measure the correlation between two signals. A group contains 10 signal pairs, each of which corresponds to a time-domain coherence data. The average value mean_TC of the 10 time-domain coherence data is calculated to represent the overall time-domain coherence level of the group.

[0121] Step 2, calculate the standard deviation mean_TC_mean between the average values of multiple groups of time-domain coherence data:

[0122] Obtain multiple groups of signal pairs, for example, 3 groups of signal pairs, each group containing 10 signal pairs. Based on the above method, obtain the average value of the time-domain coherence data of each group of signal pairs, then the three groups of signal pairs can obtain 3 time-domain coherence data average values. Take the 3 time-domain coherence data average values as a group of time-domain coherence average values. Divide the group number (3) by the group of time-domain coherence average values to obtain the average value between the time-domain coherence average values (mean_TC_mean).

[0123] To measure the dispersion (i.e., the volatility or range of variation among the time-domain coherence data averages) between each time-domain coherence data average and its corresponding mean (mean_TC_mean), the standard deviation of the time-domain coherence data averages is calculated. The specific calculation is as follows: For each time-domain coherence data average, the square of the difference between it and the mean (mean_TC_mean) is calculated, obtaining the squares of the differences for each of the three time-domain coherence data averages. The squares of these three differences are then summed and divided by the number of time-domain coherence data averages (3) to obtain the variance. Finally, the square root of the variance is taken to obtain the standard deviation (mean_TC_std) among the time-domain coherence data averages.

[0124] (3) Average of the standard deviations of time-domain coherence (F13, std_TC_mean): This represents the average of the standard deviations of multiple sets of time-domain coherence data. Each set of time-domain coherence data has its own standard deviation, which reflects the dispersion or volatility of each dataset. std_TC_mean is the average of these standard deviations to obtain an indicator representing the overall volatility. It is calculated as follows:

[0125] Step 1, calculate the standard deviation (std_TC) among a set of time-domain coherence data:

[0126] Temporal coherence (TC) measures the correlation between two signals. A set contains 10 signal pairs, each corresponding to a set of time-domain coherence data. TC measures the correlation between each time-domain coherence data point and the average coherence data point of its corresponding pair. The standard deviation of this set of time-domain coherence data is calculated based on the degree of dispersion (i.e., the volatility or range of variation among these time-domain coherence data). The specific calculation method is as follows: Calculate the average of this set of time-domain coherence data (10 TC values). Calculate the TC of each time-domain coherence data separately. i (i = 1, 2, 3, ..., 10) and their corresponding average values The squares of the differences between the data points are used to obtain the squares of the differences for each of the 10 time-domain coherence data points. The sum of these squares is then divided by the number of time-domain coherence data points (10) to obtain the variance. Finally, the square root of the variance is taken to obtain the standard deviation (std_TC) of this set of time-domain coherence data.

[0127] Step 2, calculate the average of the standard deviations among multiple sets of time-domain coherence data, std_TC_mean:

[0128] For multiple sets of time-domain coherence data, in order to obtain the average value between the standard deviations of multiple sets of time-domain coherence data, for example, 3 sets of time-domain coherence data, each set of time-domain coherence data contains 10 time-domain coherence data, and each set of time-domain coherence data has a corresponding standard deviation (std_TC), the average value of the 3 corresponding standard deviations is taken, and the average value between the standard deviations of the 3 sets of time-domain coherence data (std_TC_mean) is obtained.

[0129] (4) Standard deviation between standard deviations of time-domain coherence (F14, std_TC_std): used to represent the standard deviation between the standard deviations of multiple sets of time-domain coherence data. Each set of time-domain coherence data has its own standard deviation (std_TC), which reflects the dispersion or fluctuation of the respective data set. std_TC_std is the standard deviation of the standard deviations of multiple sets of time-domain coherence data, to obtain the dispersion or fluctuation between the standard deviations of multiple sets of time-domain coherence data. It is calculated as follows:

[0130] Step 1, calculate the standard deviation between a set of time-domain coherence data (std_TC):

[0131] Time-domain coherence TC is used to measure the correlation between two signals. A set contains 10 signal pairs, and each signal pair corresponds to a time-domain coherence data. In order to measure the dispersion between each time-domain coherence data in the set of time-domain coherence data and its corresponding average value of coherence data , that is, the fluctuation or change range between the set of time-domain coherence data, the standard deviation of the set of time-domain coherence data is calculated. The specific calculation method is as follows: calculate the average value of the set of time-domain coherence data (10 TC values) , calculate the difference between each time-domain coherence data TC i (i = 1, 2, 3, …, 10) and its corresponding average value , obtain the square of the difference corresponding to each time-domain coherence data. Add the square of the difference corresponding to each time-domain coherence data in the set, and divide by the number of time-domain coherence data (10) to obtain the variance. Finally, take the square root of the variance to obtain the standard deviation between the set of time-domain coherence data (std_TC).

[0132] Step 2, calculate the standard deviation between the standard deviations of multiple sets of time-domain coherence data std_TC_std:

[0133] For multiple sets of time domain coherence data, in order to obtain the fluctuation and dispersion between the standard deviation values of multiple sets of time domain coherence data, for example, 3 sets of time domain coherence data, each set of time domain coherence data contains 10 time domain coherence data, each set of time domain coherence data has a corresponding standard deviation (std_TC), the 3 corresponding standard deviations are calculated again, specifically:

[0134] The average value of the 3 corresponding standard deviations (std_TC_mean) is calculated, then the square of the difference between each standard deviation (std_TC) and its corresponding average value (std_TC_mean) is calculated, and the square of the difference between the 3 standard deviations (std_TC) and their corresponding average values (std_TC_mean) is obtained. After adding the square of the difference between the 3 standard deviations (std_TC) and their corresponding average values (std_TC_mean), divide by 3, obtain the variance. Finally, take the square root of the variance to obtain the standard deviation between the standard deviations of the 3 sets of time domain coherence data (std_TC_std).

[0135] (5) Skewness of the mean value of time domain coherence (F15, mean_TC_skew, where skew is the abbreviation of skewness) refers to the statistical quantity obtained by skewness analysis of the mean value of a set of time domain coherence data, which is used to characterize the skew direction and degree of the mean value of the set of time domain coherence data. If the mean value distribution of the set of time domain coherence data is symmetric, the skewness value is 0; if the mean value distribution of the set of time domain coherence data is skewed to the right, the skewness value is positive; if the mean value distribution of the set of time domain coherence data is skewed to the left, the skewness value is negative. Through mean_TC_skew, it is determined whether the time series coherence data is concentrated on one side of the mean value, or whether there are outliers, etc.

[0136] Wherein, the skewness of the mean value of time domain coherence (mean_TC_skew) is calculated by the following steps:

[0137] Step 1: Calculate the mean value (mean_TC), calculate the mean value of a set of time domain coherence data (TC). Step 2: Calculate the skewness, calculate the skewness of the mean value calculated in step 1. The calculation formula of skewness is usually based on the third moment and the fourth moment of sample data, specifically as formula 15:

[0138] Wherein, n represents the number of sample data, TC i represents the time domain coherence value of each sample data, mean_TC represents the mean value of the time domain coherence data of n sample data, s is the standard deviation value of the time domain coherence data of n sample data.

[0139] (6) Skewness of standard deviation of time-domain coherence (F16, std_TC_skew): refers to a statistical quantity obtained by skewness analysis on the standard deviation values of a group of time-domain coherence data, used to represent the skew direction and degree of the standard deviation values of the group of time-domain coherence data. If the distribution of the standard deviation values of the group of time-domain coherence data is symmetric, the skewness value is 0; if the distribution of the standard deviation values of the group of time-domain coherence data is skewed to the right, the skewness value is positive; if the distribution of the standard deviation values of the group of time-domain coherence data is skewed to the left, the skewness value is negative. Through std_TC_skew, it is determined whether the standard deviation values of the time-domain coherence data are concentrated on one side, or whether there are outliers, etc.

[0140] wherein the skewness of the standard deviation of time-domain coherence (std_TC_skew) is calculated by the following steps:

[0141] Step 1: Calculate the standard deviation (std_TC), calculate the standard deviation of a group of time-domain coherence data (TC) to measure the dispersion of these data. Step 2: Calculate the skewness (skewness), calculate the skewness of the standard deviation value calculated in step 1. The calculation formula of skewness is usually based on the third moment and the fourth moment of sample data. The specific calculation formula is as formula 16:

[0142] wherein x is each value of std_TC, μ is the average value of std_TC, n is the number of std_TC, and std_TC is the standard deviation of time-domain coherence data.

[0143] (7) Dispersion of mean of time-domain coherence (F17, mean_TC_kur): used to represent the sharpness or flatness of the distribution of the mean of time-domain coherence data. If the value of mean_TC_kur is greater than 3 (the kurtosis of normal distribution is 3), it means that the distribution of mean_TC is more sharp than the normal distribution, i.e. the data is more concentrated around the mean, and the tail is shorter; if the value of mean_TC_kur is less than 3, it means that the distribution of mean_TC is more flat than the normal distribution, i.e. the data distribution is wider, and the tail is longer.

[0144] (8) Entropy of the mean of time-domain coherence (F18, mean_TC_entropy): used to measure the uncertainty or information content of a series of time-domain coherence data. If mean_TC_entropy is low, it indicates that the coherence distribution of the signal is relatively concentrated, that is, the coherence level of the signal at most time points or frequency bands is similar, which indicates that the signal behavior is regular or predictable. If mean_TC_entropy is high, it indicates that the coherence of the signal changes greatly over time or frequency, and there is high complexity or irregularity. The signal behavior is significantly different at different time periods or frequency bands, and is difficult to predict.

[0145] (9) Entropy of the standard deviation of time-domain coherence (F19, std_TC_entropy): used to represent the uncertainty and complexity of the change distribution of time-domain coherence. If std_TC_entropy is low, it indicates that the change of time-domain coherence is relatively small and regular, and the information uncertainty contained in the standard deviation is low, that is, the fluctuation pattern of time-domain coherence is single and predictable. If std_TC_entropy is high, it indicates that the standard deviation of time-domain coherence itself has high uncertainty or complexity, that is, the change pattern of time-domain coherence over time is very diverse and difficult to accurately predict, which indicates that the behavior of the signal under different time points or conditions is significantly different.

[0146] The above is the information of the derived features of the spatial features and the time-domain features included in Table 1, which includes 29 derived features. In addition, on the basis of the 29 features, other features can also be added.

[0147] By obtaining the above video features of the video image, the target encoding parameter for encoding processing of the to-be-encoded video is analyzed and determined according to the video features, and the efficiency and accuracy of obtaining the target encoding parameter are improved.

[0148] After obtaining the video features of the to-be-encoded video of the plurality of to-be-encoded videos, the target encoding parameter of the to-be-encoded video is determined according to the video features of the to-be-encoded video. When determining the target encoding parameter of the to-be-encoded video, as described in step S102, the target encoding parameter of the to-be-encoded video is determined based on the total code rate corresponding to the plurality of to-be-encoded videos and the video features of the to-be-encoded video in the plurality of to-be-encoded videos.

[0149] Among them, the total code rate corresponding to the plurality of to-be-encoded videos is specifically the total bandwidth condition of the plurality of to-be-encoded videos in the video transmission process, which is the data transmission rate that the network can support, that is, the amount of data that can be transmitted through the network per unit time.

[0150] In order to satisfy the sum of the code rates of all the encoded videos after the plurality of to-be-encoded videos are encoded is less than the total code rate, the encoding parameters of each to-be-encoded video in the plurality of to-be-encoded videos need to be adjusted. The finally determined target encoding parameters of each to-be-encoded video in the plurality of to-be-encoded videos can satisfy the following conditions: the sum of the code rates of the encoded videos is less than the total code rate, and the sum of the quality evaluation index values of the encoded videos is greater than the video quality evaluation index standard value required by the video demand end for the plurality of to-be-encoded videos.

[0151] The encoding parameters can include at least one of the following parameters: a bit rate, a resolution, a frame rate, a quantization parameter, or a constant quality rate.

[0152] The bit rate is a key factor directly determining the code rate of the encoded video, affecting the file size and picture quality of the encoded video. The bit rate refers to the amount of information transmitted per unit time, and the unit of the bit rate is usually bits per second or multiples thereof, for example, kilobits per second, megabits per second, and the like. Setting the target bit rate of the to-be-encoded video in the video encoding software means that the rate at which the encoder compresses the to-be-encoded video is specified, so that the encoded video reaches the target file size or meets the target network transmission requirement.

[0153] The resolution is the pixel size of the video, for example, 1080p, 720p, and the like. Reducing the resolution can significantly reduce the required bit rate for video encoding, but at the same time, the picture details of the video frame will also be reduced.

[0154] The frame rate is the number of frames displayed per second of the video. A high frame rate provides smoother dynamic effects, but increases the demand for bit rate. Appropriately reducing the frame rate can reduce the demand for bit rate, thereby reducing the bandwidth requirement for the encoded video.

[0155] The quantization parameter or the constant quality rate, both of which are used in the encoder to control the compression degree, the greater the value, the greater the compression degree, the smaller the code rate of the encoded video, and the lower the quality evaluation index value of the encoded video.

[0156] Based on the above analysis, the encoding parameters are particularly important for determining the code rate and the quality evaluation index value of the encoded video. Therefore, in order to make the target encoding parameters of the to-be-encoded video in the plurality of to-be-encoded videos satisfy the above conditions, the target encoding parameters can be determined in the following ways, which are described below.

[0157] The first target coding parameter determination manner includes: determining a video content type to which a to-be-encoded video in the plurality of to-be-encoded videos belongs according to a video feature of the to-be-encoded video; classifying the to-be-encoded video in the plurality of to-be-encoded videos according to the video content type to which the to-be-encoded video belongs, to obtain at least one to-be-encoded video group, the video content types of to-be-encoded videos in the to-be-encoded video group belong to the same video content type; and determining a target coding parameter of the to-be-encoded video group based on a total code rate corresponding to the plurality of to-be-encoded videos and in combination with the video content type of the to-be-encoded video group, as the target coding parameter of the to-be-encoded video in the to-be-encoded video group.

[0158] After the video feature of the to-be-encoded video is obtained, the video content type to which the to-be-encoded video belongs is determined. For example, a first video content type of the to-be-encoded video is a real-time transmission content type, for example, a video stream of a main speaker in a conference or a live video stream. A second video content type of the to-be-encoded video is a static transmission content type, for example, a recorded course video or a recorded commodity introduction video. Compared with the second video content type, the first video content type has a higher requirement on video quality, and therefore, a bit rate allocated to the video of the first video content type is higher than a bit rate allocated to the video of the second video content type.

[0159] For another example, a third video content type of the to-be-encoded video is a static display content type, for example, the to-be-encoded video is a video stream formed by a plurality of pictures similar in content, which belongs to a static display picture. A fourth video content type of the to-be-encoded video is a dynamic content type, for example, the to-be-encoded video contains a video stream formed by a plurality of different actions, which belongs to a dynamic picture. Compared with the third video content type, the fourth video content type has a higher requirement on video quality, and therefore, a bit rate allocated to the video of the fourth video content type is higher than a bit rate allocated to the video of the third video content type.

[0160] For another example, a fifth video content type of the to-be-encoded video is a type with rich picture detail information, for example, a plurality of video frames of the to-be-encoded video respectively contain detail information content with a proportion greater than a preset proportion threshold in the video frame. A sixth video content type of the to-be-encoded video is a type with poor picture detail information, for example, a plurality of video frames of the to-be-encoded video respectively contain detail information content with a proportion less than or equal to a preset proportion threshold in the video frame. Compared with the sixth video content type, the fifth video content type has a higher requirement on video quality, and therefore, a bit rate allocated to the video of the fifth video content type is higher than a bit rate allocated to the video of the sixth video content type.

[0161] The plurality of to-be-encoded videos can be to-be-encoded videos of the same video content type or to-be-encoded videos of different video content types. For the latter, the plurality of to-be-encoded videos are classified according to their video content types, and the target encoding parameters of the videos of the same video content type are determined according to the difference in the requirements of the video quality of different video content types.

[0162] Because the total code rate of the plurality of to-be-encoded videos corresponds to the total bandwidth condition of the network currently used, the target encoding parameters of the videos of different video content types are determined under the condition of meeting the total bandwidth condition, so that the sum of the code rates of all the to-be-encoded videos in the plurality of to-be-encoded videos is less than the total code rate. In this process, the videos of the same video content type in the plurality of to-be-encoded videos are classified, and a target encoding parameter is determined for the videos of the same video content type, thereby avoiding determining the target encoding parameter for each to-be-encoded video, so as to reduce the calculation amount in the process of determining the target encoding parameter.

[0163] For example, if the total bandwidth condition of a network is 5 Mbps (megabits per second), and the plurality of to-be-encoded videos includes 5 to-be-encoded videos. In this case, it is necessary to ensure that the sum of the code rates of the 5 to-be-encoded videos is less than or equal to 5 Mbps. Therefore, the target encoding parameters of the 5 to-be-encoded videos need to be allocated by trade-off. If the video content types of the first two to-be-encoded videos are the first type of video content type described above, and the video content types of the last three to-be-encoded videos are the second type of video content type described above, then the bit rate allocated to the videos of the first type of video content type is higher than the bit rate allocated to the videos of the second type of video content type. In order to reduce the calculation amount of the target encoding parameters of the to-be-encoded videos, the two to-be-encoded videos of the first type of video content type can be determined as one target encoding parameter, and the last three to-be-encoded videos of the second type of video content type can be determined as one target encoding parameter.

[0164] Correspondingly, the encoding processing of the to-be-encoded videos according to the target encoding parameters of the to-be-encoded videos includes joint encoding processing of the to-be-encoded videos included in the to-be-encoded video group according to the target encoding parameters of the to-be-encoded video group.

[0165] The target encoding parameters of the to-be-encoded videos of the same type of video content type are the same, and therefore the plurality of to-be-encoded videos of the same type of video content type in the plurality of to-be-encoded videos can be jointly encoded, thereby reducing the calculation amount in the process of determining the target encoding parameter and improving the encoding efficiency.

[0166] The second target coding parameter determination manner is: determining a coding parameter prediction model corresponding to the plurality of to-be-encoded videos based on a total code rate corresponding to the plurality of to-be-encoded videos; inputting video features of a to-be-encoded video in the plurality of to-be-encoded videos into the coding parameter prediction model respectively, to obtain a target coding parameter of the to-be-encoded video output by the coding parameter prediction model; wherein the coding parameter prediction model is configured to determine the target coding parameter of the to-be-encoded video according to the video features of the to-be-encoded video, and the target coding parameter satisfies the following conditions: after the to-be-encoded video is encoded by using the target coding parameter, a target code rate and a target quality evaluation index value of an encoded video are obtained; and a sum of target code rates of the plurality of encoded videos corresponding to the plurality of to-be-encoded videos is less than the total code rate, and a sum of target quality evaluation index values of the plurality of encoded videos is greater than a video quality evaluation index standard value required by a video demand end for the plurality of encoded videos.

[0167] In order to make the sum of code rates of the plurality of to-be-encoded videos after encoding less than the total code rate (total bandwidth condition), the embodiment of the present disclosure can determine the target coding parameter of the to-be-encoded video of the plurality of to-be-encoded videos by using the pre-trained coding parameter prediction model. One total code rate corresponds to one coding parameter prediction model.

[0168] First, the coding parameter prediction model corresponding to the plurality of to-be-encoded videos is determined according to the total code rate corresponding to the plurality of to-be-encoded videos. Then, the coding parameter prediction model determines the target coding parameter of the to-be-encoded video according to the video features of the to-be-encoded video. In this process, the coding parameter prediction model determines the target coding parameter according to the video features of the to-be-encoded video, rather than pre-encoding the to-be-encoded parameter to obtain the encoding feature data according to the pre-encoding result, and then determining the target coding parameter, which reduces the calculation amount in the process of determining the target coding parameter and improves the determination efficiency of the target coding parameter.

[0169] Moreover, the target coding parameter obtained by the coding parameter prediction model can satisfy the following conditions: after the to-be-encoded video is encoded by using the target coding parameter, the target code rate and the target quality evaluation index value of the encoded video are obtained. And the sum of the target code rates of the plurality of encoded videos is less than the total code rate, and the sum of the target quality evaluation index values is less than the video quality evaluation index standard value required by the video demand end for the plurality of encoded videos. Therefore, the process of determining the target coding parameter by using the pre-trained coding parameter prediction model not only improves the determination rate of the target coding parameter, but also ensures that the video quality of the encoded video meets the requirements of the video demand end, that is, the coding efficiency is improved while the video quality of the encoded video is ensured.

[0170] The coding parameter prediction model can be trained in the following manner:

[0171] obtaining a total code rate corresponding to a plurality of sample videos, and a sample video feature and a sample encoding parameter of a sample video in the plurality of sample videos; inputting the total code rate corresponding to the plurality of sample videos and the sample video feature of the sample video into an initial encoding parameter prediction model to obtain an initial encoding parameter for the sample video output by the initial encoding parameter prediction model; obtaining a loss value between the initial encoding parameter of the sample video and the sample encoding parameter; summing the loss values corresponding to the sample videos in the plurality of sample videos to obtain a total loss value corresponding to the plurality of sample videos; and adjusting parameters of the initial encoding parameter prediction model based on the total loss value to obtain a trained encoding parameter prediction model.

[0172] The plurality of sample videos can be a plurality of historical videos obtained in a historical processing process, and a target plurality of historical videos are selected from the plurality of historical videos as the plurality of sample videos. The sample encoding parameter of the sample video in the plurality of sample videos can satisfy the following conditions: a sample code rate obtained after the sample video is encoded by the sample encoding parameter is less than a total code rate corresponding to the plurality of sample videos, and a sample quality evaluation index value after the encoding is greater than a quality evaluation index standard value required by a video demand end after the plurality of sample videos are encoded. The initial encoding parameter prediction model is trained by the total code rate corresponding to the plurality of sample videos and the sample video feature of the sample video in the plurality of sample videos, until the sum of the loss values between the initial encoding parameter obtained by the initial encoding parameter prediction model and the sample encoding parameter satisfies a preset condition, which indicates that the encoding parameter prediction model is trained.

[0173] The encoding parameter prediction model used in the embodiments of the present disclosure can be a target model obtained after training a support vector regression (SVR) model. This can be explained in combination with FIG. 4, which is a process schematic diagram of determining target encoding parameters of multiple to-be-encoded videos by using an encoding parameter prediction model and encoding the to-be-encoded videos according to the target encoding parameters. For example, the multiple to-be-encoded videos include three to-be-encoded videos, namely, a first to-be-encoded video A, a second to-be-encoded video B, and a third to-be-encoded video C. After the total code rate of the multiple to-be-encoded videos is determined, the corresponding encoding parameter prediction model under the total code rate is obtained. Then, the encoding parameter prediction model determines the target encoding parameter Ab of the video A according to the video feature data of the video A, determines the target encoding parameter Bb of the video B according to the video feature data of the video B, and determines the target encoding parameter Cb of the video C according to the video feature data of the video C. On this basis, the video A is encoded by using the target encoding parameter Ab to obtain the target code rate and the target quality evaluation index value of the encoded video A. The video B is encoded by using the target encoding parameter Bb to obtain the target code rate and the target quality evaluation index value of the encoded video B. The video C is encoded by using the target encoding parameter Cb to obtain the target code rate and the target quality evaluation index value of the encoded video C. The sum of the target code rate of the encoded video A, the target code rate of the encoded video B, and the target code rate of the encoded video C is less than the total code rate corresponding to the multiple to-be-encoded videos, and the sum of the target quality evaluation index value of the encoded video A, the target quality evaluation index value of the encoded video B, and the target quality evaluation index value of the encoded video C is greater than the video quality evaluation index standard value required by the video demand end for the multiple encoded videos.

[0174] In addition, other machine learning models can also be used, for example, Gaussian regression model, limit tree model, random forest, or multi-layer perception model, which are not limited here.

[0175] The sample encoding parameters of the sample videos are obtained in the following manner:

[0176] The candidate rate-distortion point group obtained by a sample video in a plurality of sample videos under a plurality of candidate encoding parameter conditions is acquired; a predetermined rate-distortion point is selected from the candidate rate-distortion point group corresponding to the sample video, and the predetermined rate-distortion point at least includes the following attribute parameters: a predetermined encoding parameter, a predetermined code rate, and a predetermined quality evaluation index value; after the sample video is encoded by using the predetermined encoding parameter, the code rate of the encoded video is the predetermined code rate, and the quality evaluation index value of the encoded video is the predetermined quality evaluation index value; if the above attribute parameters of the predetermined rate-distortion points corresponding to the sample videos in the plurality of sample videos satisfy the following conditions, the predetermined encoding parameter in the predetermined rate-distortion point of the sample video is taken as the sample encoding parameter of the sample video: the sum of the predetermined code rates corresponding to the encoded videos in the plurality of sample videos is less than the total code rate, and the sum of the predetermined quality evaluation index values corresponding to all the encoded videos is greater than the video quality evaluation index standard value required by a video demand end for the plurality of encoded videos.

[0177] The determination process of the sample encoding parameter of the sample video can be determined in the above manner.

[0178] The candidate rate-distortion point group obtained by a sample video in a plurality of sample videos under a plurality of candidate encoding parameter conditions is acquired, specifically, the sample video is encoded by using a candidate encoding parameter, to obtain the code rate and the quality evaluation index value of the encoded video after the encoding under the candidate encoding parameter condition; the code rate and the quality evaluation index value corresponding to the sample video after the encoding by using a plurality of candidate encoding parameters are obtained; the candidate encoding parameter and the code rate and the quality evaluation index value corresponding to the candidate encoding parameter are taken as a candidate rate-distortion point of the encoded video after the encoding; and the candidate rate-distortion points corresponding to the plurality of candidate encoding parameters are taken as the candidate rate-distortion point group obtained by the sample video under the plurality of candidate encoding parameter conditions.

[0179] The plurality of sample videos include a plurality of sample videos, and for each sample video, a plurality of candidate encoding parameters are used for encoding, to obtain the candidate code rate and the candidate quality evaluation index value of the encoded video, wherein a candidate rate-distortion point at least includes the following parameters: a candidate encoding parameter, a candidate code rate, and a candidate quality evaluation index value. The plurality of candidate encoding parameters are used for encoding the sample video, to obtain the candidate code rate and the candidate quality evaluation index value corresponding to the plurality of candidate encoding parameters, that is, a plurality of candidate rate-distortion points. The plurality of candidate rate-distortion points corresponding to a sample video are taken as the candidate rate-distortion point group of the sample video.

[0180] After determining the candidate rate-distortion point groups corresponding to all sample videos respectively in the plurality of sample videos, the pre-determined rate-distortion point corresponding to each sample video is determined. Then, the pre-determined rate-distortion point is judged as follows. If the judgment result satisfies the following conditions, the pre-determined rate-distortion point of the sample video is determined as the sample rate-distortion point of the sample video, and the pre-determined encoding parameter contained in the pre-determined rate-distortion point of the sample video is the sample encoding parameter of the sample video:

[0181] The sum of the pre-determined code rates corresponding to the plurality of encoded videos is less than the total code rate corresponding to the plurality of sample videos, and the sum of the pre-determined quality evaluation index values corresponding to the plurality of encoded videos is greater than the video quality evaluation index standard value required by the video demand end for the encoded videos.

[0182] After determining the sample encoding parameters of the plurality of sample videos, the initial encoding parameter prediction model is trained with the sample encoding parameters as labels and the sample video features of the sample encoding videos, so as to obtain the encoding parameter prediction model corresponding to different total code rates.

[0183] The above is the way of determining the target encoding parameter of the to-be-encoded video through the encoding parameter prediction model.

[0184] The conditions satisfied by the target encoding parameters obtained by the first and second target encoding parameter determination manners are that the sum of the target code rates of the plurality of encoded videos is less than the total code rate, and the sum of the target quality evaluation index values is greater than the video quality evaluation index standard value required by the video demand end for the plurality of encoded videos. The video quality evaluation index standard value required by the video demand end for the plurality of encoded videos is different according to the specific video type or the requirements of the video demand end.

[0185] Therefore, the third way of determining the target encoding parameter of the to-be-encoded video is also provided in the embodiments of the present disclosure:

[0186] The video quality evaluation index standard value required by the video demand end for the plurality of encoded videos is obtained. The target encoding parameter of the to-be-encoded video is determined based on the total code rate corresponding to the plurality of to-be-encoded videos and the video features of the to-be-encoded video in the plurality of to-be-encoded videos. The target encoding parameter of the to-be-encoded video can be determined based on the total code rate corresponding to the plurality of to-be-encoded videos, the video features of the to-be-encoded video in the plurality of to-be-encoded videos, and the video quality evaluation index standard value.

[0187] The total code rate of the plurality of to-be-encoded videos defines the sum of code rates that can be used by the encoding process in all to-be-encoded videos in the plurality of to-be-encoded videos, and the sum of code rates in the encoding process of the to-be-encoded videos cannot exceed the total code rate. On this basis, in order to ensure that the quality evaluation index values of all to-be-encoded videos after encoding meet the preset requirements, it is necessary to allocate corresponding target encoding parameters to each to-be-encoded video. The target encoding parameter corresponding to each to-be-encoded parameter can be determined according to the video features of the to-be-encoded video and the video quality evaluation index standard value required by the video demand end.

[0188] If the video quality evaluation index standard value required by the video demand end is greater than the first preset video quality evaluation index standard threshold, the first preset video quality evaluation index standard threshold is that the video picture quality belongs to standard definition quality. In this case, the bit rate requirement of the target encoding parameter determined for the to-be-encoded video in the plurality of to-be-encoded videos is the bit rate required to obtain standard definition video picture quality.

[0189] If the video quality evaluation index standard value required by the video demand end is greater than the second preset video quality evaluation index standard threshold, the second preset video quality evaluation index standard threshold is that the video picture quality belongs to super definition quality. The code rate corresponding to the super definition quality is greater than the code rate corresponding to the standard definition quality. The bit rate in the target encoding parameter of the to-be-encoded video directly determines the code rate of the encoded video. In this case, the bit rate requirement of the target encoding parameter determined for the to-be-encoded video in the plurality of to-be-encoded videos is the bit rate required to obtain super definition video picture quality.

[0190] Based on the different video quality evaluation index standard values required by the video demand end for the plurality of encoded videos, the disclosure can specifically determine the target encoding parameter of the to-be-encoded video in the following manner:

[0191] Specifically, the target encoding parameter of the to-be-encoded video is determined based on the total code rate corresponding to the plurality of to-be-encoded videos, the video features of the to-be-encoded video in the plurality of to-be-encoded videos, and the video quality evaluation index standard value, which includes: predicting the predicted bit rate required by the to-be-encoded video according to the video features of the to-be-encoded video; determining the predicted encoding parameter required by the to-be-encoded video according to the predicted bit rate required by the to-be-encoded video; obtaining the predicted code rate and the predicted quality evaluation index value of the video of the to-be-encoded video after encoding by the predicted encoding parameter; if the sum of the predicted code rates corresponding to the encoded videos in the plurality of encoded videos corresponding to the plurality of to-be-encoded videos is less than the total code rate, and the sum of the predicted quality evaluation index values corresponding to the encoded videos is greater than the video quality evaluation index standard value, then the predicted encoding parameter is taken as the target encoding parameter of the to-be-encoded video.

[0192] The video features of the to-be-encoded videos in the plurality of to-be-encoded videos are different, and the bit rates required in the to-be-encoded video encoding processes are different. For a to-be-encoded video with a high video feature complexity of the video feature of the to-be-encoded video, the bit rate required is higher than that of a to-be-encoded video with a low video feature complexity of the video feature of the to-be-encoded video. For example, a to-be-encoded video is a to-be-encoded video with an independent video frame ratio greater than a preset ratio threshold. For example, a first to-be-encoded video includes 100 video frames, of which 70 are independent frames. An independent frame needs to display the entire content of the frame, while a predicted frame is determined according to the difference between the current predicted frame and the adjacent previous independent frame or predicted frame. A second to-be-encoded video includes 100 video frames, of which 30 are independent frames, and the rest are predicted frames. In this case, the video feature complexity of the first video frame is higher than that of the second to-be-encoded video. Correspondingly, the bit rate required for encoding the first to-be-encoded video is higher than that required for encoding the second to-be-encoded video.

[0193] According to the bit rate required for the to-be-encoded video, the prediction encoding parameters corresponding to the to-be-encoded video are determined, and the prediction code rate and the prediction quality evaluation index value of the to-be-encoded video after encoding by the prediction encoding parameters are obtained.

[0194] The sum of the prediction code rates obtained by the to-be-encoded videos in the plurality of to-be-encoded videos is added, and the sum of the prediction code rates is compared with the total code rate. The sum of the prediction quality evaluation index values obtained by the to-be-encoded videos is added, and the sum of the prediction quality evaluation index values is compared with the video quality evaluation index standard value.

[0195] If the sum of the prediction code rates corresponding to the plurality of encoded videos in the plurality of to-be-encoded videos is less than the total code rate, and the sum of the prediction quality evaluation index values corresponding to the encoded videos is greater than the video quality evaluation index standard value, it is indicated that the prediction encoding parameters predicted for the to-be-encoded videos in the plurality of to-be-encoded videos can meet the requirements of the video demand end for the quality of the encoded videos, and the prediction encoding parameters are used as the target encoding parameters of the to-be-encoded videos.

[0196] The above determination of the prediction encoding parameters for each to-be-encoded video can include a plurality of candidate encoding parameters. In this case, for the to-be-encoded videos in the plurality of to-be-encoded videos, the plurality of candidate encoding parameters are used to encode the to-be-encoded videos respectively to obtain a plurality of prediction code rates and a plurality of prediction quality evaluation index values corresponding to the to-be-encoded videos. If the prediction code rates and the prediction quality evaluation index values corresponding to the plurality of candidate encoding parameters satisfy the following conditions, a target encoding parameter is selected from the plurality of candidate encoding parameters: the sum of the prediction code rates corresponding to the encoded videos is less than the total code rate, and the sum of the prediction quality evaluation index values corresponding to the encoded videos is less than the video quality evaluation index standard value.

[0197] For example, the multiple videos to be encoded include two videos to be encoded, each of which determines two candidate encoding parameters, and the two videos to be encoded are encoded by using the two candidate encoding parameters to obtain two predicted code rates and two predicted quality evaluation index values corresponding to the two videos to be encoded respectively.

[0198] The predicted code rates of the encoded videos corresponding to all the videos to be encoded of the multiple videos to be encoded are added to obtain a total value of the predicted code rates of each encoded video corresponding to the multiple predicted code rates and the predicted code rates of other encoded videos. Similarly, the predicted quality evaluation index values of the encoded videos corresponding to all the videos to be encoded of the multiple videos to be encoded are added to obtain a total value of the predicted quality evaluation index values of each encoded video corresponding to the multiple predicted quality evaluation index values and the predicted quality evaluation index values of other encoded videos.

[0199] Continuing with the above example: the total code rate of the multiple videos to be encoded is 5 Mbps (megabits per second).

[0200] The multiple videos to be encoded include two videos to be encoded, a first video to be encoded A and a second video to be encoded B. The first video to be encoded A includes two candidate encoding parameters Aa1 and Aa2. Aa1 corresponds to a predicted code rate Aa11 and a predicted quality evaluation index value Aa12. Aa2 corresponds to a predicted code rate Aa21 and a predicted quality evaluation index value Aa22.

[0201] The second video to be encoded B includes two candidate encoding parameters Bb1 and Bb2. Bb1 corresponds to a predicted code rate Bb11 and a predicted quality evaluation index value Bb12. Bb2 corresponds to a predicted code rate Bb21 and a predicted quality evaluation index value Bb22.

[0202] Among them, the sum of the predicted code rates corresponding to all the videos to be encoded of the multiple videos to be encoded includes the following types: the sum of Aa11 and Bb11, the sum of Aa11 and Bb21, the sum of Aa21 and Bb11, and the sum of Aa21 and Bb21.

[0203] The sum of the predicted quality evaluation index values corresponding to all the videos to be encoded of the multiple videos to be encoded includes the following types: the sum of Aa12 and Bb12, the sum of Aa12 and Bb22, the sum of Aa22 and Bb12, and the sum of Aa22 and Bb22.

[0204] The sum of any one of the predicted code rates is less than the total code rate 5 Mbps, but the sum of any one of the predicted quality evaluation index values is less than the video demand end corresponding video quality evaluation index standard value. For example, the video demand end corresponding video quality evaluation index standard value can be greater than the second preset video quality evaluation index standard threshold, and the second preset video quality evaluation index standard threshold is that the video picture quality belongs to super clear quality. The code rate corresponding to the super clear quality is greater than the code rate corresponding to the standard clear quality.

[0205] Therefore, in this case, the target coding parameter can be selected from the multiple candidate coding parameters corresponding to each to-be-encoded video, specifically as follows:

[0206] The target coding parameter is selected from the multiple candidate coding parameters, including:

[0207] The to-be-encoded video is encoded by the multiple candidate coding parameters respectively to obtain multiple predicted quality evaluation index values of the to-be-encoded video after encoding; the multiple predicted quality evaluation index values of the to-be-encoded video after encoding are compared to obtain a target predicted quality evaluation index value, the target predicted quality evaluation index value is greater than other predicted quality evaluation index values except the target predicted quality evaluation index value in the multiple predicted quality evaluation index values; and the candidate coding parameter corresponding to the target predicted quality evaluation index value is taken as the target coding parameter of the to-be-encoded video.

[0208] Continuing to combine the above example, wherein Aa11 is greater than Aa21, and Aa12 is greater than Aa22. Wherein Bb11 is greater than Bb21, and Bb12 is greater than Bb22. Therefore, the target coding parameter corresponding to the first to-be-encoded video is the candidate coding parameter Aa1, and the target coding parameter corresponding to the second to-be-encoded video is the candidate coding parameter Bb1.

[0209] The above is a third determination manner of determining the target coding parameter.

[0210] The three determination manners of determining the target coding parameter, the target coding parameter obtained by the three determination manners satisfies the following conditions after the to-be-encoded video is encoded: the sum of the target code rates is less than the total code rate, and the sum of the target quality evaluation index values is greater than the video quality evaluation index standard value required by the video demand end for the encoded video.

[0211] In addition, in the preferred embodiment, the target coding parameter of the to-be-encoded video determined can satisfy the following conditions: the sum of the code rates of the encoded video is less than the total code rate, and the sum of the quality evaluation index values of the encoded video is greater than the sum of the subsequent quality evaluation index values corresponding to the encoded video under the condition of other candidate coding parameters.

[0212] When the encoding parameter of the to-be-encoded video is the target encoding parameter, the code rate of the encoded video is the target code rate, and the quality evaluation index value of the encoded video is the target quality evaluation index value. When the encoding parameter of the to-be-encoded video is the other candidate encoding parameter, the code rate of the encoded video is the candidate code rate, and the quality evaluation index value of the encoded video is the candidate quality evaluation index value.

[0213] The sum of the target code rates corresponding to the plurality of encoded videos is less than the total code rate, and the sum of the target quality evaluation index values corresponding to the plurality of encoded videos is greater than the sum of the candidate quality evaluation index values of the plurality of encoded videos under the condition of the other candidate encoding parameters.

[0214] In combination with the example of FIG. 4, the sum of the target code rate of the encoded A video, the target code rate of the encoded B video, and the target code rate of the encoded C video is less than the total code rate corresponding to the plurality of to-be-encoded videos, and the sum of the target quality evaluation index value of the encoded A video, the target quality evaluation index value of the encoded B video, and the target quality evaluation index value of the encoded C video is greater than the sum of the candidate quality evaluation index values of the encoded videos under the condition of the other candidate encoding parameters (the total value obtained by adding the candidate quality evaluation index values corresponding to the above three encoded videos).

[0215] Corresponding to the first embodiment, the second embodiment of the present disclosure provides a video encoding device, and the related parts can be understood by referring to the description of the corresponding method embodiment. Please refer to FIG. 5, which is a schematic diagram of a video encoding device provided by the second embodiment of the present disclosure. The video encoding device shown in the figure comprises: a video feature obtaining unit 501, configured to perform feature extraction processing on a plurality of to-be-encoded videos, and obtain the video feature of a to-be-encoded video in the plurality of to-be-encoded videos; a target encoding parameter determining unit 502, configured to determine the target encoding parameter of the to-be-encoded video based on the total code rate corresponding to the plurality of to-be-encoded videos and the video feature of the to-be-encoded video; and an encoding processing unit 503, configured to perform encoding processing on the to-be-encoded video according to the target encoding parameter of the to-be-encoded video. When the encoding parameter of the to-be-encoded video in the plurality of to-be-encoded videos is the target encoding parameter corresponding thereto, the code rate of the encoded video after the encoding processing is the target code rate, and the quality evaluation index value of the encoded video is the target quality evaluation index value. The sum of the target code rates corresponding to the plurality of encoded videos is less than the total code rate, and the sum of the target quality evaluation index values corresponding to the plurality of encoded videos is greater than the video quality evaluation index standard value required by a video demand end for the plurality of encoded videos.

[0216] Based on the above-mentioned embodiments, a third embodiment of the present disclosure provides an electronic device, and the related parts can be understood by referring to the corresponding description of the above-mentioned embodiments. Please refer to FIG. 6, which is a schematic diagram of an electronic device according to the third embodiment of the present disclosure. The electronic device shown in the figure comprises a memory and a processor; the memory is used to store a computer program, and the computer program is executed by the processor to perform the method provided by the embodiments of the present disclosure.

[0217] Based on the above-mentioned embodiments, a fourth embodiment of the present disclosure provides a computer storage medium, and the related parts can be understood by referring to the corresponding description of the above-mentioned embodiments. The schematic diagram of the computer storage medium is similar to FIG. 6, and the memory in the figure can be understood as the storage medium. The computer storage medium stores a computer program, and the computer program is executed by the processor to implement the method provided by the embodiments of the present disclosure.

[0218] Based on the above-mentioned embodiments, a fifth embodiment of the present disclosure provides a computer program product, which comprises a computer program. When the computer program is executed by the processor of the electronic device, the processor can at least implement the method provided in the above-mentioned embodiments.

[0219] In a typical configuration, a computing device includes one or more processors (CPUs), input / output interfaces, network interfaces, and memory. The memory can include non-persistent memory in the computer readable medium, random access memory (RAM), and / or non-volatile memory, such as read-only memory (ROM) or flash memory (flash RAM). The memory is an example of computer readable media.

[0220] 1. Computer readable media includes permanent and non-permanent, removable and non-removable media can be implemented by any method or technology to store information. Information can be computer readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassette, magnetic tape, magnetic disk storage or other magnetic storage device, or any other non-transmission medium that can be used to store information that can be accessed by a computing device. According to the definition in this paper, computer readable media does not include non-transitory computer readable media (transitory media), such as modulated data signals and carriers.

[0221] 2. As those skilled in the art will appreciate, the embodiments of the present disclosure can be provided as methods, systems or computer program products. Thus, the present disclosure can take the form of an entirely hardware embodiment, an entirely software embodiment or an embodiment combining software and hardware aspects. Furthermore, the present disclosure can take the form of a computer program product on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROMs, optical storage devices, etc.) embodying computer-readable program code.

[0222] The present disclosure is disclosed above with the preferred embodiments, but it is not intended to limit the present disclosure, and any person skilled in the art can make possible changes and modifications without departing from the spirit and scope of the present disclosure, so the protection scope of the present disclosure should be subject to the scope defined by the claims of the present disclosure.

[0223] It should be noted that the embodiments of the present disclosure can involve the use of user data. In actual applications, user-specific personal data can be used in the schemes described herein in a manner that complies with applicable laws and regulations of the country (for example, with the explicit consent of the user, with the actual notification of the user, etc.), within the scope permitted by applicable laws and regulations. It should be noted that the user information (including but not limited to user equipment information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in the present disclosure are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data need to comply with relevant laws, regulations and standards of relevant countries and regions, and provide corresponding operation portal for user to choose authorization or refusal.

Claims

1. A method of video encoding, wherein, The method comprises: feature extraction processing is performed on a plurality of to-be-encoded videos to obtain video features of a to-be-encoded video in the plurality of to-be-encoded videos; based on total code rates corresponding to the plurality of to-be-encoded videos and the video features of the to-be-encoded video, target encoding parameters of the to-be-encoded video are determined; encoding processing is performed on the to-be-encoded video according to the target encoding parameters of the to-be-encoded video; wherein when the encoding parameters of the to-be-encoded video in the plurality of to-be-encoded videos are the target encoding parameters corresponding thereto, the code rates of the encoded videos after the encoding processing are target code rates, the quality evaluation index values of the encoded videos after the encoding processing are target quality evaluation index values, the sum of the target code rates corresponding to the encoded videos respectively is less than the total code rates, and the sum of the target quality evaluation index values corresponding to the encoded videos respectively is greater than a video quality evaluation index standard value required by a video demand end for the plurality of encoded videos.

2. The method of claim 1, wherein, The method comprises: determining, according to the video features of the to-be-encoded video in the plurality of to-be-encoded videos, a video content type to which the to-be-encoded video belongs; classifying the to-be-encoded video in the plurality of to-be-encoded videos according to the video content type to which the to-be-encoded video belongs to obtain at least one to-be-encoded video group, the video content types of the to-be-encoded videos in the to-be-encoded video group belonging to the same video content type; based on the total code rates corresponding to the plurality of to-be-encoded videos, combining the video content types of the to-be-encoded video groups, and determining target encoding parameters of the to-be-encoded video groups as the target encoding parameters of the to-be-encoded videos in the to-be-encoded video groups; correspondingly, the encoding processing performed on the to-be-encoded video according to the target encoding parameters of the to-be-encoded video comprises joint encoding processing performed on the to-be-encoded videos contained in the to-be-encoded video group according to the target encoding parameters of the to-be-encoded video group.

3. The method of claim 1, wherein, The method comprises: based on the total code rates corresponding to the plurality of to-be-encoded videos, determining an encoding parameter prediction model corresponding to the plurality of to-be-encoded videos; inputting the video features of the to-be-encoded video in the plurality of to-be-encoded videos into the encoding parameter prediction model respectively to obtain target encoding parameters of the to-be-encoded video output by the encoding parameter prediction model; wherein the encoding parameter prediction model is used to determine the target encoding parameters of the to-be-encoded video according to the video features of the to-be-encoded video, and the target encoding parameters satisfy the following conditions: The target code rate and the target quality evaluation index value of the encoded video are obtained after the to-be-encoded video is encoded by using the target encoding parameter; the sum of the target code rates of the encoded videos corresponding to the plurality of to-be-encoded videos is less than the total code rate, and the sum of the target quality evaluation index values of the encoded videos is greater than the video quality evaluation index standard value required by the video demand end for the plurality of encoded videos.

4. The method of claim 3, wherein, The encoding parameter prediction model is trained in the following manner: obtaining the total code rate corresponding to a plurality of sample videos, and sample video features and sample encoding parameters of sample videos in the plurality of sample videos; inputting the total code rate corresponding to the plurality of sample videos and the sample video features of the sample videos into an initial encoding parameter prediction model to obtain initial encoding parameters of the sample videos output by the initial encoding parameter prediction model; obtaining a loss value between the initial encoding parameters of the sample videos and the sample encoding parameters; summing the loss values of the sample videos in the plurality of sample videos to obtain a total loss value corresponding to the plurality of sample videos; adjusting parameters of the initial encoding parameter prediction model based on the total loss value to obtain a trained encoding parameter prediction model.

5. The method of claim 4, wherein, The sample encoding parameters of the sample videos are obtained in the following manner: obtaining a plurality of candidate rate-distortion point groups of a sample video in a plurality of candidate encoding parameters; selecting a predetermined rate-distortion point from the candidate rate-distortion point groups corresponding to the sample video, the predetermined rate-distortion point comprising at least the following attribute parameters: a predetermined encoding parameter, a predetermined code rate, and a predetermined quality evaluation index value; after the sample video is encoded by using the predetermined encoding parameter, the code rate of the encoded video is the predetermined code rate, and the quality evaluation index value of the encoded video is the predetermined quality evaluation index value; if the attribute parameters of the predetermined rate-distortion points of the sample videos in the plurality of sample videos satisfy the following conditions, the predetermined encoding parameter in the predetermined rate-distortion point of the sample video is taken as the sample encoding parameter of the sample video: the sum of the predetermined code rates of the encoded videos corresponding to the plurality of sample videos is less than the total code rate, and the sum of the predetermined quality evaluation index values of all the encoded videos is greater than the video quality evaluation index standard value required by the video demand end for the plurality of encoded videos.

6. The method of claim 5, wherein, The candidate rate-distortion point groups of a sample video in a plurality of candidate encoding parameters are obtained in the following manner: encoding the sample video by using a candidate encoding parameter to obtain the code rate and the quality evaluation index value of the encoded video after the encoding processing under the condition of the candidate encoding parameter; obtaining the code rate and the quality evaluation index value corresponding to the encoding processing of the sample video by using a plurality of candidate encoding parameters; taking the candidate encoding parameter and the code rate and the quality evaluation index value corresponding to the candidate encoding parameter as a candidate rate-distortion point of the encoded video after the encoding processing; The multiple candidate rate-distortion points corresponding to the multiple candidate encoding parameters are used as a candidate rate-distortion point group obtained by the sample video under the condition of the multiple candidate encoding parameters.

7. The method of claim 1, wherein, Further comprising: Obtaining a video quality evaluation index standard value required by a video demand end for the multiple encoded videos; The target encoding parameter of the to-be-encoded video is determined based on the total code rate corresponding to the multiple to-be-encoded videos, the video features of the to-be-encoded video in the multiple to-be-encoded videos, and the video quality evaluation index standard value. The target encoding parameter of the to-be-encoded video is determined based on the total code rate corresponding to the multiple to-be-encoded videos, the video features of the to-be-encoded video in the multiple to-be-encoded videos, and the video quality evaluation index standard value.

8. The method of claim 7, wherein, The target encoding parameter of the to-be-encoded video is determined based on the total code rate corresponding to the multiple to-be-encoded videos, the video features of the to-be-encoded video in the multiple to-be-encoded videos, and the video quality evaluation index standard value. The predicted bit rate required by the to-be-encoded video is predicted according to the video features of the to-be-encoded video; The predicted encoding parameter required by the to-be-encoded video is determined according to the predicted bit rate required by the to-be-encoded video; The predicted code rate and the predicted quality evaluation index value of the video of the to-be-encoded video after encoding by the predicted encoding parameter are obtained; If the sum of the predicted code rates corresponding to the multiple encoded videos is less than the total code rate, and the sum of the predicted quality evaluation index values corresponding to the multiple encoded videos is greater than the video quality evaluation index standard value, the predicted encoding parameter is used as the target encoding parameter of the to-be-encoded video.

9. The method of claim 8, wherein, The predicted encoding parameter includes multiple candidate encoding parameters. The method further comprises: The multiple to-be-encoded videos are respectively encoded by using the multiple candidate encoding parameters to obtain corresponding multiple predicted code rates and multiple predicted quality evaluation index values; If the predicted code rate and the predicted quality evaluation index value corresponding to the multiple candidate encoding parameters satisfy the following conditions, a target encoding parameter is selected from the multiple candidate encoding parameters: The sum of the predicted code rates corresponding to the multiple encoded videos is less than the total code rate, and the predicted quality evaluation index value corresponding to the multiple encoded videos is less than the video quality evaluation index standard value.

10. The method of claim 9, wherein, The target encoding parameter is selected from the multiple candidate encoding parameters, comprising: The multiple predicted quality evaluation index values of the to-be-encoded video after encoding by the multiple candidate encoding parameters are obtained; The target predicted quality evaluation index value is obtained by comparing the multiple predicted quality evaluation index values of the to-be-encoded video after encoding by the multiple candidate encoding parameters, the target predicted quality evaluation index value being greater than other predicted quality evaluation index values except the target predicted quality evaluation index value in the multiple predicted quality evaluation index values; The candidate encoding parameter corresponding to the target predicted quality evaluation index value is used as the target encoding parameter of the to-be-encoded video.

11. An electronic device, comprising: Comprise: A processor; And A memory for storing a computer program, wherein the electronic device, after being powered on and running the computer program by the processor, executes the method of any one of claims 1-10.

12. A storage medium, wherein, The storage medium stores a computer program, wherein the computer program is run by the processor to execute the method of any one of claims 1-10.

13. A computer program product, wherein, Comprise: A computer program, when executed by the processor of an electronic device, causes the processor to execute the method of any one of claims 1-10.

Citation Information

Patent Citations

  • Video coding method and device and electronic equipment

    CN113938682A

  • Multimedia data coding method and device, electronic equipment and storage medium

    CN117478886A

  • Video coding method and device, electronic equipment and computer storage medium

    CN118158414A

  • Video coding method and device

    CN119182935A

  • Constant Quality Video Encoding with Encoding Parameter Fine-Tuning

    US20190253763A1