Video encoding methods and related products

By converting the video to a second color space and determining the target bitrate control parameters based on the baseline bitrate control parameters and the deviation, the video frames are segmented and encoded, which solves the problem of slow video encoding speed and achieves the saving of hardware resources and the improvement of encoding speed.

CN119110072BActive Publication Date: 2025-10-28XIAOHONGSHU TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411173004.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-08-26
Publication Date
2025-10-28
Estimated Expiration
2044-08-26

AI Technical Summary

Technical Problem

Existing video encoding technologies are slow when processing video frames from different scenarios, resulting in high hardware resource consumption and high costs.

Method used

By converting the video to be encoded to a second color space, the encoder determines the target bitrate control parameters based on the reference bitrate control parameters and the encoding deviation, and then segments and encodes the video frames, reducing the amount of data processing and improving the encoding speed.

Benefits of technology

By reducing data processing volume, hardware resource requirements are lowered, hardware costs are reduced, and video encoding speed is increased.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119110072B_ABST
    Figure CN119110072B_ABST
Patent Text Reader

Abstract

This application discloses a video encoding method and related products. The method includes: inputting the video to be encoded, n segmented frame numbers, a reference bitrate control parameter, and (n+1) encoding deviations into an encoder. The segmented frame numbers are the frame numbers of the video frames in the video to be encoded, and the n segmented frame numbers are used to determine (n+1) segments of video to be encoded from the video to be encoded, where n is a positive integer; determining the (n+1) segments of video to be encoded from the video to be encoded using the encoder; determining (n+1) target bitrate control parameters based on the reference bitrate control parameter and the (n+1) encoding deviations using the encoder; and encoding the (n+1) segments of video to be encoded using the encoder based on the (n+1) target bitrate control parameters to obtain the video encoding result of the video to be encoded output by the encoder. This method can improve the speed of video encoding.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of video coding technology, and in particular to a video coding method and related products. Background Technology

[0002] Encoding video allows for compression, reducing its data size and facilitating storage and transmission. Therefore, understanding video encoding is of paramount importance. Summary of the Invention

[0003] This application provides a video encoding method and related products to improve video encoding speed. The related products include video encoding devices, electronic devices, computer-readable storage media, and computer program products.

[0004] Firstly, a video encoding method is provided, the method comprising:

[0005] The video to be encoded, n segmented frame numbers, a baseline bitrate control parameter, and (n+1) encoding deviations are input to the encoder. The segmented frame numbers are the frame numbers of the video frames in the video to be encoded. The n segmented frame numbers are used to determine (n+1) segments of video to be encoded from the video to be encoded. The n is a positive integer. The encoding deviations correspond one-to-one with the segments of video to be encoded. The encoding deviations are the deviations of the bitrate control parameters of the segments of video to be encoded from the baseline bitrate control parameters.

[0006] The encoder determines (n+1) segments of video to be encoded from the video to be encoded, which are determined by the n segmentation frame numbers;

[0007] The encoder determines (n+1) target bitrate control parameters based on the baseline bitrate control parameters and the (n+1) to-be-encoded deviations, and the video segments to be encoded correspond one-to-one with the target bitrate control parameters;

[0008] The encoder encodes the (n+1) video segments to be encoded based on the (n+1) target bitrate control parameters, thereby obtaining the video encoding result of the video to be encoded output by the encoder.

[0009] In any embodiment of this application, the video to be encoded is a video in a first color space, and the encoder is used to process the video in a second color space, wherein the first color space is different from the second color space;

[0010] Before inputting the video to be encoded, n segmented frame numbers, baseline bitrate control parameters, and (n+1) encoding offsets to the encoder, the method further includes:

[0011] The video to be encoded is converted from the first color space to the second color space to obtain the converted video;

[0012] The step of inputting the video to be encoded, n segmented frame numbers, a reference bitrate control parameter, and (n+1) encoding deviations to the encoder includes:

[0013] The converted video, the n segmented frame numbers, the baseline bitrate control parameters, and the (n+1) unencoded deviations are input to the encoder;

[0014] The step of determining the (n+1) segments of video to be encoded from the video to be encoded by the encoder, which are determined by the n segmentation frame numbers, includes:

[0015] The encoder determines (n+1) segments of converted video corresponding to the (n+1) segments of video to be encoded from the converted video based on the n segmented frame numbers;

[0016] The step of determining (n+1) target bit rate control parameters by the encoder based on the reference bit rate control parameters and the (n+1) to-be-encoded deviations includes:

[0017] The encoder determines (n+1) target bitrate control parameters for the (n+1) segments of converted video based on the reference bitrate control parameters and the (n+1) unencoded deviations, and the converted video segments correspond one-to-one with the target bitrate control parameters;

[0018] The step of encoding the (n+1) video segments to be encoded using the encoder based on the (n+1) target bitrate control parameters to obtain the video encoding result of the video to be encoded output by the encoder includes:

[0019] The encoder encodes the (n+1) video segments to be encoded based on the (n+1) target bitrate control parameters to obtain the video encoding result.

[0020] In conjunction with any embodiment of this application, the step of encoding the (n+1) segments of the converted video based on the (n+1) keyframes, the reference bitrate control parameters, and the (n+1) encoding deviations to obtain the video encoding result includes:

[0021] The encoder determines the sum of the baseline bitrate control parameter and each of the (n+1) encoding deviations, thereby obtaining (n+1) target bitrate control parameters for the (n+1) converted video segments. The target bitrate control parameters correspond one-to-one with the converted video segments.

[0022] The encoder encodes the (n+1) converted video segments based on the (n+1) keyframes and the (n+1) target bitrate control parameters to obtain the video encoding result.

[0023] In any embodiment of this application, the (n+1) segments of video to be encoded include a target video segment, and the (n+1) offsets to be encoded include a target offset corresponding to the target video segment;

[0024] Before inputting the converted video, the n segmented frame numbers, the baseline bitrate control parameters, and the (n+1) to-be-encoded offsets to the encoder, the method further includes:

[0025] Feature extraction is performed on the target video segment to obtain at least one feature of the target video segment;

[0026] The at least one feature is input into the first model to obtain the first relation output by the first model;

[0027] The first model is used to predict the relationship between the bitrate control parameters of the video and the bitrate of the reference encoding result based on the video's features. The reference encoding result is obtained by encoding the video using the video's bitrate control parameters. The first relationship is the relationship between the target bitrate control parameters of the target video segment and the bitrate of the target encoding result. The target encoding result is obtained by encoding the target video segment using the target bitrate control parameters.

[0028] Based on the first relationship, the target bitrate control parameters of the target video segment are determined;

[0029] Alternatively, the at least one feature can be input into a second model to obtain a second relationship output by the second model. The first model is used to predict the relationship between the bitrate control parameters of the video and the video quality of the reference coding result based on the features of the video. The second relationship is the relationship between the target bitrate control parameters of the target video segment and the video quality of the target coding result.

[0030] Based on the second relationship, the target bitrate control parameters of the target video segment are determined;

[0031] The difference between the target bitrate control parameter and the baseline bitrate control parameter is determined to obtain the target deviation.

[0032] In conjunction with any embodiment of this application, before determining the target bitrate control parameter of the target video segment based on the first relationship, the method further includes:

[0033] The at least one feature is input into the second model to obtain the second relationship output by the second model;

[0034] The step of determining the target bitrate control parameters of the target video segment based on the first relationship includes:

[0035] Based on the first relationship and the second relationship, a bitrate control parameter is determined to make the target encoding result meet the preset requirements, which is used as the target bitrate control parameter. The preset requirements include the bitrate requirements of the target encoding result and the video quality requirements of the target encoding result.

[0036] In any embodiment of this application, in the video to be encoded, video frames located on both sides of the segmentation frame number have different scenes, and video frames located between two adjacent segmentation frame numbers have the same scene.

[0037] In any embodiment of this application, the n segmentation frame numbers are determined through a scene detection process, which includes:

[0038] When the scene of consecutive i video frames in the video to be encoded changes from the first scene to the second scene, the reference number of reference video frames with the second scene in the i video frames is determined, where i is an integer greater than 1.

[0039] If the number of references is greater than or equal to the first threshold, the smallest frame number of the reference video frame is determined as the segmentation frame number.

[0040] In any embodiment of this application, after the scene of the consecutive i-frame video frames switches from the first scene to the second scene, and then switches back to the first scene from the second scene, the scene detection process further includes:

[0041] If the number of references is greater than or equal to a first threshold, before determining the minimum frame number of the reference video frame as the segmentation frame number, the first threshold is used.

[0042] Determine the initial number of initial video frames containing the first scene in the i-frame video frame;

[0043] When the number of references is greater than or equal to a first threshold, determining the minimum frame number of the reference video frames as the segmentation frame number includes:

[0044] If the number of references is greater than or equal to a first threshold and the number of initial references is less than or equal to the number of references, the minimum frame number of the reference video frame is determined as the segmentation frame number.

[0045] In conjunction with any embodiment of this application, the scene detection process further includes: after determining the initial number of initial video frames containing the first scene in the i-frame video frames,

[0046] If the initial number is greater than the reference number, the frame number after the maximum frame number is determined as the segmentation frame number, and the maximum frame number is the maximum frame number of the initial video frame.

[0047] In any embodiment of this application, the scene of the consecutive i-frame video frames switches from the first scene to the second scene, and then switches from the second scene to the third scene;

[0048] The scene detection process further includes: before determining the minimum frame number of the reference video frame as the segmentation frame number when the number of references is greater than or equal to a first threshold.

[0049] Determine the initial number of initial video frames containing the first scene in the i-frame video frame;

[0050] When the number of references is greater than or equal to a first threshold, determining the frame number smaller than the first threshold among the frame numbers of the reference video frames as the segmentation frame number includes:

[0051] If the number of references is greater than or equal to the first threshold and the number of initial references is less than or equal to the second threshold, the minimum frame number of the reference video frame is determined as the segmentation frame number.

[0052] In conjunction with any embodiment of this application, the scene detection process further includes: after determining the initial number of initial video frames containing the first scene in the i-frame video frames,

[0053] If the initial number is greater than the second threshold and the reference number is less than the third threshold, determine the number of unconfirmed video frames in the i-frame video frame that have the third scene.

[0054] If the number of frames to be confirmed is greater than or equal to the third threshold, the smallest frame number of the video frames to be confirmed is determined as the segmentation frame number, and the third threshold is less than the second threshold.

[0055] Secondly, a video encoding apparatus is provided, the apparatus comprising:

[0056] The input unit is used to input the video to be encoded, n segmented frame numbers, a reference bitrate control parameter, and (n+1) encoding deviations to the encoder. The segmented frame numbers are the frame numbers of the video frames in the video to be encoded. The n segmented frame numbers are used to determine (n+1) segments of video to be encoded from the video to be encoded. The n is a positive integer. The encoding deviations correspond one-to-one with the segments of video to be encoded. The encoding deviations are the deviations of the bitrate control parameter of the segment to be encoded from the reference bitrate control parameter.

[0057] An encoding unit is configured to determine, via the encoder, (n+1) segments of video to be encoded from the video to be encoded, which are determined by the n segmentation frame numbers;

[0058] The encoding unit is used to determine (n+1) target bitrate control parameters by the encoder based on the reference bitrate control parameters and the (n+1) to-be-encoded deviations, wherein the video segments to be encoded correspond one-to-one with the target bitrate control parameters;

[0059] The encoding unit is used to encode the (n+1) segments of video to be encoded by the encoder based on the (n+1) target bitrate control parameters, so as to obtain the video encoding result of the video to be encoded output by the encoder.

[0060] In any embodiment of this application, the video to be encoded is a video in a first color space, and the encoder is used to process the video in a second color space, wherein the first color space is different from the second color space;

[0061] The video encoding device further includes: a conversion unit, used to convert the video to be encoded from the first color space to the second color space to obtain a converted video;

[0062] The encoding unit is specifically used for:

[0063] The converted video, the n segmented frame numbers, the baseline bitrate control parameters, and the (n+1) unencoded deviations are input to the encoder;

[0064] The encoder determines (n+1) segments of converted video corresponding to the (n+1) segments of video to be encoded from the converted video based on the n segmented frame numbers;

[0065] The encoder determines (n+1) target bitrate control parameters for the (n+1) segments of converted video based on the reference bitrate control parameters and the (n+1) unencoded deviations, and the converted video segments correspond one-to-one with the target bitrate control parameters;

[0066] The encoder encodes the (n+1) video segments to be encoded based on the (n+1) target bitrate control parameters to obtain the video encoding result.

[0067] In conjunction with any embodiment of this application, the encoding unit is specifically used for:

[0068] The encoder takes the first frame of the converted video and the video frames in the converted video that correspond to the n segmentation frame numbers as (n+1) key frames of the (n+1) segments of the converted video, and the key frames correspond one-to-one with the converted video segments.

[0069] The encoder encodes the (n+1) segments of the converted video based on the (n+1) keyframes, the reference bitrate control parameters, and the (n+1) encoding deviations to obtain the video encoding result.

[0070] In conjunction with any embodiment of this application, the encoding unit is specifically used for:

[0071] The encoder determines the sum of the baseline bitrate control parameter and each of the (n+1) encoding deviations, thereby obtaining (n+1) target bitrate control parameters for the (n+1) converted video segments. The target bitrate control parameters correspond one-to-one with the converted video segments.

[0072] The encoder encodes the (n+1) converted video segments based on the (n+1) keyframes and the (n+1) target bitrate control parameters to obtain the video encoding result.

[0073] In any embodiment of this application, the (n+1) segments of video to be encoded include a target video segment, and the (n+1) offsets to be encoded include a target offset corresponding to the target video segment;

[0074] The video encoding apparatus further includes: a determining unit configured to:

[0075] Feature extraction is performed on the target video segment to obtain at least one feature of the target video segment;

[0076] The at least one feature is input into the first model to obtain the first relation output by the first model;

[0077] The first model is used to predict the relationship between the bitrate control parameters of the video and the bitrate of the reference encoding result based on the video's features. The reference encoding result is obtained by encoding the video using the video's bitrate control parameters. The first relationship is the relationship between the target bitrate control parameters of the target video segment and the bitrate of the target encoding result. The target encoding result is obtained by encoding the target video segment using the target bitrate control parameters.

[0078] Based on the first relationship, the target bitrate control parameters of the target video segment are determined;

[0079] Alternatively, the at least one feature can be input into a second model to obtain a second relationship output by the second model. The first model is used to predict the relationship between the bitrate control parameters of the video and the video quality of the reference coding result based on the features of the video. The second relationship is the relationship between the target bitrate control parameters of the target video segment and the video quality of the target coding result.

[0080] Based on the second relationship, the target bitrate control parameters of the target video segment are determined;

[0081] The difference between the target bitrate control parameter and the baseline bitrate control parameter is determined to obtain the target deviation.

[0082] In conjunction with any embodiment of this application, the determining unit is further configured to:

[0083] The at least one feature is input into the second model to obtain the second relationship output by the second model;

[0084] Based on the first relationship and the second relationship, a bitrate control parameter is determined to make the target encoding result meet the preset requirements, which is used as the target bitrate control parameter. The preset requirements include the bitrate requirements of the target encoding result and the video quality requirements of the target encoding result.

[0085] In any embodiment of this application, in the video to be encoded, video frames located on both sides of the segmentation frame number have different scenes, and video frames located between two adjacent segmentation frame numbers have the same scene.

[0086] In any embodiment of this application, the n segmentation frame numbers are determined through a scene detection process, which includes:

[0087] When the scene of consecutive i video frames in the video to be encoded changes from the first scene to the second scene, the reference number of reference video frames with the second scene in the i video frames is determined, where i is an integer greater than 1.

[0088] If the number of references is greater than or equal to the first threshold, the smallest frame number of the reference video frame is determined as the segmentation frame number.

[0089] In any embodiment of this application, after the scene of the consecutive i-frame video frames switches from the first scene to the second scene, and then switches back to the first scene from the second scene, the scene detection process further includes:

[0090] Determine the initial number of initial video frames containing the first scene in the i-frame video frame;

[0091] When the number of references is greater than or equal to a first threshold, determining the minimum frame number of the reference video frames as the segmentation frame number includes:

[0092] If the number of references is greater than or equal to a first threshold and the number of initial references is less than or equal to the number of references, the minimum frame number of the reference video frame is determined as the segmentation frame number.

[0093] In conjunction with any embodiment of this application, the scene detection process further includes: after determining the initial number of initial video frames containing the first scene in the i-frame video frames,

[0094] If the initial number is greater than the reference number, the frame number after the maximum frame number is determined as the segmentation frame number, and the maximum frame number is the maximum frame number of the initial video frame.

[0095] In any embodiment of this application, the scene of the consecutive i-frame video frames switches from the first scene to the second scene, and then switches from the second scene to the third scene;

[0096] The scene detection process further includes: before determining the minimum frame number of the reference video frame as the segmentation frame number when the number of references is greater than or equal to a first threshold.

[0097] Determine the initial number of initial video frames containing the first scene in the i-frame video frame;

[0098] When the number of references is greater than or equal to a first threshold, determining the frame number smaller than the first threshold among the frame numbers of the reference video frames as the segmentation frame number includes:

[0099] If the number of references is greater than or equal to the first threshold and the number of initial references is less than or equal to the second threshold, the minimum frame number of the reference video frame is determined as the segmentation frame number.

[0100] In conjunction with any embodiment of this application, the scene detection process further includes: after determining the initial number of initial video frames containing the first scene in the i-frame video frames, if the initial number is greater than the second threshold and the reference number is less than the third threshold, determining the number of video frames to be confirmed containing the third scene in the i-frame video frames.

[0101] If the number of frames to be confirmed is greater than or equal to the third threshold, the smallest frame number of the video frames to be confirmed is determined as the segmentation frame number, and the third threshold is less than the second threshold.

[0102] Thirdly, an electronic device is provided, comprising: a processor and a memory, the memory being used to store computer program code, the computer program code including computer instructions, wherein, when the processor executes the computer instructions, the electronic device performs as described in the first aspect and any of its embodiments.

[0103] Fourthly, another electronic device is provided, comprising: a processor, a transmitting device, an input device, an output device, and a memory, the memory being used to store computer program code, the computer program code including computer instructions, wherein, when the processor executes the computer instructions, the electronic device performs as described in the first aspect and any of its embodiments.

[0104] Fifthly, a computer-readable storage medium is provided, wherein a computer program is stored therein, the computer program including program instructions that, when executed by a processor, cause the processor to perform the first aspect and any of its embodiments described above.

[0105] In a sixth aspect, a computer program product is provided, the computer program product comprising a computer program or instructions that, when the computer program or instructions are executed on a computer, cause the computer to perform the first aspect described above and any of its embodiments.

[0106] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and are not intended to limit this application.

[0107] In this application, the video to be encoded and n segmented frame numbers are used. Each segmented frame number is a frame number of a video frame in the video to be encoded. The n segmented frame numbers are used to determine (n+1) segments of video to be encoded from the video to be encoded. The encoding deviation is the deviation of the bitrate control parameter of the video segment to be encoded from the reference bitrate control parameter. After inputting the video to be encoded, the n segmented frame numbers, the reference bitrate control parameter, and the (n+1) encoding deviations into the encoder, the video encoding device can determine the (n+1) segments of video to be encoded from the video to be encoded using the n segmented frame numbers. Furthermore, the encoder can determine (n+1) target bitrate control parameters based on the reference bitrate control parameter and the (n+1) encoding deviations. Finally, the encoder can encode the (n+1) segments of video to be encoded based on the (n+1) target bitrate control parameters to obtain the video encoding result of the video to be encoded output by the encoder. Therefore, the encoder can encode (n+1) segments of video to be encoded by encoding the video in one operation, thereby reducing the amount of data processing required to encode different segments of video to be encoded using different bitrate control parameters, and improving the encoding speed. Furthermore, since the amount of data processing can be reduced, it means that this encoding method requires less hardware resources, thereby reducing hardware costs. Attached Figure Description

[0108] To more clearly illustrate the technical solutions in the embodiments of this application or the background art, the accompanying drawings used in the embodiments of this application or the background art will be described below.

[0109] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the specification, serve to explain the technical solutions of this application.

[0110] Figure 1 A flowchart illustrating a video encoding method provided in an embodiment of this application;

[0111] Figure 2 This is a schematic diagram of the structure of a deep learning model provided in an embodiment of this application;

[0112] Figure 3 This application provides a schematic diagram of the architecture of a video-on-demand system.

[0113] Figure 4 This is a schematic diagram of the structure of a video encoding device provided in an embodiment of this application;

[0114] Figure 5 This is a schematic diagram of the hardware structure of an electronic device provided in an embodiment of this application. Detailed Implementation

[0115] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present application, and not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of the present application.

[0116] The terms "first," "second," etc., in the specification, claims, and accompanying drawings of this application are used to distinguish different objects, not to describe a specific order. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or apparatus that includes a series of steps or units is not limited to the listed steps or units, but may optionally include steps or units not listed, or may optionally include other steps or units inherent to these processes, methods, products, or apparatuses.

[0117] In this document, the term "embodiment" means that a particular feature, structure, or characteristic described in connection with an embodiment may be included in at least one embodiment of this application. The appearance of this phrase in various places throughout the specification does not necessarily refer to the same embodiment, nor is it a separate or alternative embodiment mutually exclusive with other embodiments. It will be explicitly and implicitly understood by those skilled in the art that the embodiments described herein can be combined with other embodiments.

[0118] The execution subject of this application embodiment is a video encoding device, which can be any electronic device capable of executing the technical solutions disclosed in the method embodiments of this application. Optionally, the video encoding device can be one of the following: a computer or a server.

[0119] It should be understood that the method embodiments of this application can also be implemented by a processor executing computer program code. The embodiments of this application are described below with reference to the accompanying drawings. Please refer to... Figure 1 , Figure 1 This is a flowchart illustrating a video encoding method provided in an embodiment of this application.

[0120] 101. Input the video to be encoded, n segmented frame numbers, reference bit rate control parameters, and (n+1) encoding deviations into the encoder.

[0121] In this embodiment of the application, the video to be encoded can be any video. For example, the video to be encoded is a video on a video-on-demand platform, or, for example, a video captured by a mobile phone.

[0122] In this embodiment of the application, the segmentation frame number is the frame number of the video frame in the video to be encoded. The frame number of the video frame in the video to be encoded indicates the display order of the video frames in the video to be encoded. For example, the video frames to be encoded include video frame a and video frame b, where the frame number of video frame a is 1 and the frame number of video frame b is 2. Then, when displaying the video frames to be encoded, video frame a is displayed first, and then video frame b is displayed.

[0123] The n segmentation frame numbers are used to determine (n+1) segments of video to be encoded from the video to be encoded, where n is a positive integer. For example, if n = 1, the n segmentation frame numbers represent 101 frames. If the video to be encoded has a total of 180 frames, then based on the n segmentation frame numbers, the following two segments can be determined: frames 1 to 100 of the video to be encoded constitute one segment, and frames 101 to 180 constitute another segment. For another example, if n = 2, the n segmentation frame numbers include 88 frames and 150 frames. If the video to be encoded has 180 frames, then based on n segmentation frame numbers, the following three segments can be identified from the video to be encoded: frames 1 to 87, frames 88 to 149, and frames 150 to 180. It should be understood that identifying (n+1) segments based on n segmentation frame numbers means defining (n+1) segments from the entire video to be encoded, without actually dividing the video into (n+1) segments.

[0124] Optionally, in the video to be encoded, video frames located on either side of a segmentation frame number have different scenes, while video frames located between two adjacent segmentation frame numbers have the same scene. Since video frames located on either side of a segmentation frame number belong to different video segments to be encoded, and video frames located between two adjacent segmentation frame numbers belong to the same video segment to be encoded, different video segments to be encoded have different scenes, while video frames within the same video segment to be encoded have the same scene.

[0125] Optionally, scene detection can be performed on the video to be encoded using an encoder to determine the scenes in the video. Then, based on the scenes in the video to be encoded, segmentation frame numbers can be determined so that video frames located on either side of a segmentation frame number have different scenes, and video frames located between two adjacent segmentation frame numbers have the same scene. Optionally, the encoder used for scene detection on the video to be encoded can be a different encoder than the encoder used in step 101.

[0126] Optionally, the encoder can extract image information from each video frame in the video to be encoded by encoding it. Based on this image information, it can determine whether different video frames share the same scene, and thus determine the segmentation frame number. The image information includes color. In one possible implementation, the encoder generates a histogram for each video frame in the video to be encoded, where the histogram represents the color distribution of pixels in the video frame. For two adjacent video frames in the video to be encoded, the difference between their histograms is determined. If the difference is greater than or equal to a difference threshold, it indicates a large difference in the content of the two adjacent video frames, thus determining that they have different scenes. If the difference is less than the difference threshold, it indicates a small difference in the content of the two adjacent video frames, thus determining that they share the same scene. When it is determined that two adjacent video frames have different scenes, the larger frame number between the two adjacent video frames is used as the segmentation frame number.

[0127] In one possible implementation, the similarity of video frames within the same video segment to be encoded is greater than a similarity threshold, while the similarity of video frames within different video segments to be encoded is less than or equal to the similarity threshold. That is, video frames with the same scene have a similarity greater than the similarity threshold, while video frames with different scenes have a similarity less than or equal to the similarity threshold. In another possible implementation, the scene refers to the background. Video frames within the same video segment to be encoded contain the same background, while video frames within different video segments to be encoded contain different backgrounds. In yet another possible implementation, the scene includes both regions containing and not containing regions of human eye interest. Video frames containing and not containing regions of human eye interest belong to different video segments to be encoded. The region of human eye interest is the area in the video frame that is of interest to the human eye; optionally, the region of human eye interest is the region of interest (ROI). In another possible implementation, the scenarios include high and low image complexity. Video frames with image complexity exceeding a certain threshold and those with image complexity below the threshold belong to different video segments to be encoded. Higher image complexity indicates a higher bitrate required for encoding the video frame. Optionally, the image complexity of a video frame can be determined based on the magnitude of its gradient; specifically, the magnitude is positively correlated with image complexity.

[0128] In this embodiment, the bitrate control parameter is a parameter used to control the bitrate of the encoded video. Optionally, the bitrate control parameter is one of the following: constant rate factor (CRF), constant bit rate (CBR), variable bit rate (VBR), or constant quantization parameter (CQP).

[0129] In this embodiment, the baseline bitrate control parameter is a preset bitrate control parameter. Optionally, the baseline bitrate control parameter is determined based on historical experience. For example, the bitrate control parameter is CRF. Based on historical experience, it is known that when CRF is 32, the bitrate of the encoded video obtained by encoding using CRF meets expectations. Therefore, the baseline bitrate control parameter can be determined to be 32. As another example, the bitrate control parameter is CRF. Based on historical experience, it is known that when CRF is 33, the quality of the encoded video obtained by encoding using CRF and then decoding the encoded video meets expectations. Therefore, the baseline bitrate control parameter can be determined to be 33.

[0130] In this embodiment, the encoding deviation corresponds one-to-one with the video segment to be encoded, meaning each video segment has an encoding deviation. The encoding deviation is the deviation of the bitrate control parameter of the video segment to be encoded from the baseline bitrate control parameter. Based on the encoding deviation and the baseline bitrate control parameter, the bitrate control parameter of the video segment to be encoded can be determined. The encoding deviation can be positive or negative. For example, if the bitrate control parameter is CRF, the baseline bitrate control parameter is 32, and the encoding deviation of the video segment to be encoded is 3, then the bitrate control parameter of the video segment to be encoded is 32 + 3 = 35. As another example, if the bitrate control parameter is CRF, the baseline bitrate control parameter is 32, and the encoding deviation of the video segment to be encoded is -1, then the bitrate control parameter of the video segment to be encoded is 32 - 1 = 31.

[0131] Optionally, given different scenarios in different video segments to be encoded, the required bitrate for encoding these segments varies. Therefore, the bitrate control parameters for different video segments are different, resulting in different encoding deviations. In one possible implementation, video frames containing areas of human visual interest and those not containing such areas belong to different video segments to be encoded. For ease of explanation, the video segments to which video frames containing areas of human visual interest belong will be referred to as "eye-interested video segments," and those not containing such areas will be referred to as "eye-uninterested video segments." The bitrate required to encode eye-interested video segments is higher than that required to encode eye-uninterested video segments, leading to different encoding deviations for these segments. In another possible implementation, video frames with image complexity exceeding a complexity threshold and those with image complexity below a complexity threshold belong to different video segments to be encoded. For ease of explanation, the video segments to be encoded belonging to video frames with image complexity exceeding the complexity threshold will be referred to as complex video segments, and the video segments to be encoded belonging to video frames with image complexity exceeding the complexity threshold will be referred to as simple video segments. The bit rate required to encode complex video segments is higher than that required to encode simple video segments, which in turn leads to different encoding deviations for complex video segments and simple video segments.

[0132] In one possible implementation, after determining the bitrate control parameters of the video segment to be encoded, the video encoding device determines the difference between the bitrate control parameters of the video segment to be encoded and the reference bitrate control parameters, thus obtaining the encoding deviation.

[0133] In this embodiment, the encoder can be any encoder. Optionally, the encoder can be one of the following: an encoder built based on the H.265 standard, an encoder built based on the H.264 standard, an encoder built based on the H.266 standard, or an encoder built based on the Open Multimedia Consortium (AOMedia Video 1, AV1) standard. By inputting bitrate control parameters and video into the encoder, the encoder can encode the video using the bitrate control parameters.

[0134] By inputting a bitrate control parameter to the encoder, the encoder can complete one encoding operation based on that bitrate control parameter. Since different video segments to be encoded have different bitrate control parameters, in order to represent the bitrate control parameters of different video segments to be encoded using a single bitrate control parameter, the video encoding parameters input a reference bitrate control parameter and (n+1) encoding offsets to the encoder. In this way, the encoder can determine the bitrate control parameters of (n+1) video segments to be encoded during one encoding operation, based on the reference bitrate control parameter and the (n+1) encoding offsets.

[0135] Since encoding a single segment of video is performed by the encoder, and the video to be encoded comprises (n+1) segments, the video encoding device inputs the video to be encoded and n segmentation frame numbers into the encoder to enable the encoder to distinguish between different segments. Thus, the encoder can determine the (n+1) segments to be encoded from the video to be encoded based on the n segmentation frame numbers. Optionally, after obtaining the n segmentation frame numbers, the video encoder stores them as a comma-separated values ​​(CSV) file and inputs this CSV file into the encoder. Optionally, the order of the n segmentation frame numbers in the CSV file is determined based on the presentation time stamp (PTS) of the video frame corresponding to each segmentation frame number.

[0136] Therefore, the video encoding device inputs the video to be encoded, n segmented frame numbers, a reference bitrate control parameter, and (n+1) encoding deviations into the encoder. The encoder can then determine (n+1) segments of the video to be encoded from the video based on the n segmented frame numbers, and determine the bitrate control parameters for these (n+1) segments based on the reference bitrate control parameter and the (n+1) encoding deviations. Finally, the bitrate control parameters of these (n+1) segments are used to encode them, completing the encoding of the video to be encoded and obtaining the video encoding result, which is the encoding result of the video to be encoded.

[0137] 102. Using the encoder described above, determine the (n+1) segments of video to be encoded from the video to be encoded, which are determined by the above n segmentation frame numbers.

[0138] The encoder takes the first frame of the video to be encoded and the video frame corresponding to the segmentation frame number as the starting video frame of a segment to be encoded, and can determine (n+1) segments to be encoded from the video to be encoded, which are determined by n segmentation frame numbers.

[0139] 103. Based on the above-mentioned reference bit rate control parameters and the above-mentioned (n+1) encoding deviations, determine (n+1) target bit rate control parameters using the encoder described above.

[0140] In this embodiment, there is a one-to-one correspondence between the video segments to be encoded and the target bitrate control parameters. Since the encoding deviation represents the deviation of the bitrate control parameters of the video segment to be encoded from the reference bitrate control parameters, the video encoding device can determine (n+1) target bitrate control parameters based on the reference bitrate control parameters and (n+1) encoding deviations. In one possible implementation, the encoder determines the sum of the reference bitrate control parameters and each of the (n+1) encoding deviations to obtain the (n+1) target bitrate control parameters for the (n+1) converted video segments.

[0141] 104. The encoder encodes the (n+1) video segments to be encoded based on the (n+1) target bit rate control parameters to obtain the video encoding result of the video to be encoded output by the encoder.

[0142] The video encoding device encodes any one of the (n+1) segments of video to be encoded using the corresponding target bitrate control parameters. After encoding the (n+1) segments, the encoding of the video to be encoded is complete, yielding the video encoding result. In this process, the encoder uses different target bitrate control parameters to encode different parts of the video to be encoded, rather than dividing the video to be encoded into (n+1) segments and encoding each segment separately.

[0143] In one possible implementation, the encoder takes the first frame of the video to be encoded and the video frames corresponding to the n segmentation frame numbers as (n+1) keyframes (intra-coded frames) of the (n+1) segments of the video to be encoded. Based on the baseline bitrate control parameters and the (n+1) encoding offsets, the encoder determines the bitrate control parameters for the (n+1) segments of the video to be encoded. Based on the bitrate control parameters and the (n+1) keyframes of the (n+1) segments of the video to be encoded, the encoder encodes the (n+1) segments of the video to be encoded, obtaining the video encoding result.

[0144] In this embodiment, the video to be encoded and n segmented frame numbers are used. Each segmented frame number is a frame number of a video frame in the video to be encoded. The n segmented frame numbers are used to determine (n+1) segments of video to be encoded from the video to be encoded. The encoding deviation is the deviation of the bitrate control parameter of the video segment to be encoded from the reference bitrate control parameter. After inputting the video to be encoded, the n segmented frame numbers, the reference bitrate control parameter, and the (n+1) encoding deviations into the encoder, the video encoding device can determine the (n+1) segments of video to be encoded from the video to be encoded, determined by the n segmented frame numbers. Furthermore, the encoder can determine (n+1) target bitrate control parameters based on the reference bitrate control parameter and the (n+1) encoding deviations. Finally, the encoder can encode the (n+1) segments of video to be encoded based on the (n+1) target bitrate control parameters to obtain the video encoding result of the video to be encoded output by the encoder. Therefore, the encoder can encode (n+1) segments of video to be encoded by encoding the video in one operation, thereby reducing the amount of data processing required to encode different segments of video to be encoded using different bitrate control parameters, and improving the encoding speed. Furthermore, since the amount of data processing can be reduced, it means that this encoding method requires less hardware resources, thereby reducing hardware costs.

[0145] Specifically, traditional video coding methods typically involve first determining n segmentation frame numbers from the video to be encoded, and then dividing the video into (n+1) segments based on these n frame numbers. For each of these (n+1) segments, the segment to be encoded and its bitrate control parameters are input to the encoder. The encoder then encodes the segment based on these bitrate control parameters, resulting in a sub-encoded result. By encoding the (n+1) segments, (n+1) sub-encoded results are obtained. Finally, these (n+1) sub-encoded results are concatenated to obtain the encoded video.

[0146] However, if the segmentation frame number is not a keyframe, then there is no keyframe in the video segment to be encoded, and the encoder cannot encode the video segment. Therefore, when the segmentation frame number is not a keyframe, traditional video encoding methods need to encode the video to be encoded by setting the video frame corresponding to the segmentation frame number as a keyframe, and then dividing the video to be encoded into (n+1) segments based on n segmentation frame numbers, and then using the encoder to encode the (n+1) segments.

[0147] This application, by encoding the video to be encoded in one step, utilizes the bitrate control parameters of (n+1) segments of the video to be encoded to encode the (n+1) segments. Therefore, compared to traditional video encoding methods, it eliminates the overhead of setting the video frames corresponding to the segmented frame numbers as keyframes during encoding, thereby reducing data processing volume and increasing encoding speed. The reduced data processing volume also lowers hardware costs. For example, encoding a set of videos to be encoded using traditional methods requires 100,000 central processing unit (CPU) cores. Using the method provided in this application, the data processing volume is reduced by 20% compared to traditional methods, requiring only 80,000 CPU cores. Thus, encoding the same set of videos to be encoded saves 20,000 CPU cores, resulting in a cost saving of 20,000 CPU cores.

[0148] Furthermore, traditional encoding methods concatenate (n+1) sub-encoding results based on the timestamp of the video to be encoded to obtain the encoded video. Therefore, if the timestamp of the video to be encoded is incorrect, the display order of the video frames in the concatenated encoded result will differ from the display order of the video frames in the original video. For example, in the video to be encoded, the timestamp of frame 10 is t1, and the timestamp of frame 11 is t2. If t2 is less than t1, then in the concatenated encoded result, frame 11 will be displayed before frame 10. However, in the original video to be encoded, frame 10 will be displayed before frame 11. If the display order of the video frames in the concatenated encoded result differs from the display order of the video frames in the original video, then after decoding the encoded result to obtain the decoded video, playing the decoded video and its corresponding audio will result in a mismatch between audio and video frames, i.e., audio-visual desynchronization. Furthermore, if the timestamp of the video to be encoded is incorrect, the encoded result obtained by splicing may also have frame drops compared to the video to be encoded. However, since the embodiment of this application does not segment the video to be encoded, the video encoding result of the video to be encoded can be directly obtained through the encoder, thus eliminating the need to perform the step of splicing the sub-encoded results, thereby avoiding the problems of audio-visual asynchrony and frame drops caused by splicing the sub-encoded results.

[0149] Furthermore, since the embodiments of this application determine the (n+1) video segments to be encoded from the video during the encoding process using the bitrate control parameters of the (n+1) video segments to be encoded, those skilled in the art can integrate the determination of the (n+1) video segments to be encoded from the video and the encoding of the (n+1) video segments to be encoded into the same program code when developing program code to implement the technical solution of the embodiments of this application. This saves development manpower costs and facilitates subsequent maintenance of the program code, compared to having different personnel develop program code separately for determining the (n+1) video segments to be encoded from the video and encoding the (n+1) video segments to be encoded.

[0150] In one possible scenario, the video encoding device is a server running a video-on-demand platform. After receiving a video to be encoded uploaded by a user from the video-on-demand platform, the video encoding device can encode (n+1) video segments to be encoded using the video encoding method provided in this application embodiment to obtain a video encoding result. This allows for the use of different bitrate control parameters to encode different video segments, enabling the selection of different bitrate control parameters based on the video content. This improves the video quality of the encoded result and reduces its bitrate. After obtaining the video encoding result, it can be published to the video-on-demand platform.

[0151] As an optional implementation, the video to be encoded is a video in a first color space, and the encoder is used to process the video in a second color space, wherein the first color space and the second color space are different. Optionally, the first color space is the RGB color space and the second color space is the YUV color space. In this case, the video to be encoded is an RGB format video, and the encoder is used to process the YUV format video, where R represents red, G represents green, B represents blue, Y represents luminance, U represents the red chroma component, and V represents the blue chroma component.

[0152] In this embodiment, the video encoding device performs the following steps before performing step 101:

[0153] 201. Convert the video to be encoded from the first color space to the second color space to obtain the converted video.

[0154] Since the encoder is used to process video in the second color space, the video encoding device can convert the video to be encoded into a converted video that can be processed by the encoder by performing step 201. In one possible implementation, the first color space is the RGB color space and the second color space is the YUV color space. The video encoding device obtains the converted video by converting the video to be encoded from the RGB color space to the YUV color space.

[0155] After obtaining the converted video, the video encoding device performs the following steps during step 101:

[0156] 202. Input the above-mentioned converted video, the above-mentioned n segmented frame numbers, the above-mentioned baseline bit rate control parameters, and the above-mentioned (n+1) to-be-encoded deviations into the above-mentioned encoder.

[0157] After obtaining the converted video, the video encoding device performs the following steps during step 102:

[0158] 203. Using the encoder described above, based on the above n segmented frame numbers, determine the (n+1) segments of converted video that correspond to the above (n+1) segments of video to be encoded from the above converted video.

[0159] In this embodiment of the application, the (n+1) segments of the converted video are determined based on n segmentation frame numbers. The implementation method of determining the (n+1) segments of the converted video based on n segmentation frame numbers can be found in step 101, which describes the implementation method of determining the (n+1) segments of the video to be encoded based on n segmentation frame numbers. It will not be repeated here.

[0160] Since the converted video is a video that can be processed by the encoder, after the video encoding device inputs the converted video, n segmented frame numbers, reference bit rate control parameters and (n+1) offsets to be encoded into the encoder, the encoder can determine (n+1) segments of the converted video from the converted video based on the n segmented frame numbers.

[0161] After obtaining the converted video, the video encoding device performs the following steps during step 103:

[0162] 204. Based on the above-mentioned reference bitrate control parameters and the above-mentioned (n+1) encoding deviations, the encoder determines the (n+1) target bitrate control parameters of the above-mentioned (n+1) converted video segments.

[0163] Since there is a one-to-one correspondence between the video segments to be encoded and the target bitrate control parameters, a one-to-one correspondence between the video segments to be encoded and the converted video segments, and a one-to-one correspondence between the converted video segments and the target bitrate control parameters, the encoder can determine the (n+1) target bitrate control parameters for the (n+1) converted video segments based on the reference bitrate control parameters and the (n+1) encoding deviations.

[0164] After obtaining the converted video, the video encoding device performs the following steps during step 104:

[0165] 205. Using the encoder described above, the above (n+1) target bitrate control parameters are used to encode the above (n+1) video segments to be encoded, and the above video encoding result is obtained.

[0166] The video encoding device encodes any one of the (n+1) converted video segments using the corresponding target bitrate control parameters. After encoding the (n+1) converted video segments, the encoding of the converted video is complete. Since the converted video is obtained by converting the video to be encoded, the result of encoding the converted video is the same as the result of encoding the video to be encoded. Therefore, by encoding the (n+1) video segments to be encoded based on (n+1) target bitrate control parameters, the video encoding result can be obtained.

[0167] In this implementation, since the encoder is used to process video in the second color space, and the video to be encoded is not in the second color space, the video to be encoded needs to be converted to the second color space to obtain a converted video before the encoder is used to encode the video. Then, the converted video, n segmented frame numbers, reference bitrate control parameters, and (n+1) encoding deviations are input to the encoder, and the video encoding result can be obtained through the encoder's encoding.

[0168] In this implementation, the encoder only needs to perform encoding once, so the video encoding device only needs to perform a color space conversion on the video to be encoded once. This color space conversion refers to converting the video to be encoded from a first color space to a second color space. Traditional video encoding methods require first encoding the video to be encoded, setting the video frames corresponding to the segmented frame numbers as keyframes, and then encoding the (n+1) segments of the video to be encoded. That is, traditional video encoding methods require two encoding operations, and therefore two color space conversions. Therefore, compared to traditional video encoding methods, this implementation saves the overhead of one color space conversion, thereby reducing data processing volume, increasing encoding speed, and lowering hardware costs.

[0169] As an optional implementation, the video encoding device performs the following steps during step 205:

[0170] 301. Using the encoder described above, the first frame of the converted video and the video frames in the converted video that correspond to the above n segmentation frame numbers are taken as the (n+1) key frames of the above (n+1) segments of the converted video.

[0171] In this embodiment, the first frame of the converted video and the video frames corresponding to the n segmentation frame numbers are all keyframes. Each keyframe corresponds one-to-one with a converted video segment; that is, each converted video segment has one keyframe. Optionally, the keyframe is an instantaneous decoding refresh (IDR) frame.

[0172] 302. The encoder encodes the (n+1) segments of the converted video based on the (n+1) keyframes, the reference bit rate control parameters, and the (n+1) encoding deviations to obtain the video encoding result.

[0173] The encoder can determine the bitrate control parameters of (n+1) converted video segments based on the reference bitrate control parameters and (n+1) encoding deviations. Then, based on the bitrate control parameters of (n+1) converted video segments and (n+1) keyframes, it can encode the (n+1) converted video segments to obtain the video encoding result.

[0174] In one possible implementation, the encoder determines the sum of the baseline bitrate control parameters and each of the (n+1) encoding deviations to be encoded, thus obtaining (n+1) target bitrate control parameters for the (n+1) converted video segments. Each target bitrate control corresponds one-to-one with a converted video segment. The encoder then encodes the (n+1) converted video segments based on the (n+1) keyframes and the (n+1) target bitrate control parameters to obtain the video encoding result.

[0175] As an optional implementation, the (n+1) video segments to be encoded include a target video segment, and the (n+1) offsets to be encoded include target offsets corresponding to the target video segment, wherein the target video segment can be any one of the (n+1) video segments to be encoded.

[0176] In this embodiment, the video encoding device obtains the target deviation by performing the following steps:

[0177] 401. Perform feature extraction on the above target video segment to obtain at least one feature of the above target video segment.

[0178] In this embodiment of the application, when the number of features in the target video segment is greater than 1, different features include information of different dimensions. In one possible implementation, the features of the target video segment include: video content features, encoding result features, and precoding features.

[0179] Video content features include one or more of the following: texture information, noise information, Blind Video Quality Assessment (BVQA) information, blockiness information, blurriness information, and spatiotemporal information. Texture information refers to the texture information within the target video segment. Noise information can be used to assess the intensity and distribution of noise in the video; noise affects the video's clarity and visual appeal. Blind quality assessment information refers to the quality score predicted from the target video segment based on statistical analysis or a machine learning model of the video signal; this prediction does not depend on a reference video. Blockiness information can represent image blockiness caused by compression or other processing, which typically manifests as obvious boundaries in the image (i.e., video frames of the target video segment), affecting visual continuity. Blurriness information measures the degree of blurriness in the video, usually determined by analyzing the sharpness or edge information of video frames. Spatiotemporal information includes the temporal and spatial information of the target video segment; for example, temporal information can be the correlation between different video frames in the target video segment, and spatial information can be the scene complexity within video frames in the target video segment.

[0180] Optionally, the texture information is the video spatial texture feature. The video encoding device can extract the video spatial texture features of the target video segment based on the gray-level co-occurrence matrix (GLCM). The gray-level co-occurrence matrix is ​​a commonly used image processing technique for analyzing the texture features of an image. The co-occurrence matrix reflects the co-occurrence probability of different gray levels in space, thereby capturing the image's texture information, such as roughness and regularity. When applied to video, each color video frame can first be converted to a grayscale image, and then the co-occurrence matrix of the pixel gray values ​​in each frame can be calculated. By analyzing these matrices, indicators reflecting the texture characteristics of the video frame can be extracted, such as contrast, correlation, and entropy.

[0181] Optionally, the spatiotemporal information includes video spatiotemporal correlation features, which are extracted from the target video segment based on a preset spatiotemporal domain algorithm. The spatiotemporal domain algorithm may include motion estimation, optical flow calculation, or time-series analysis techniques to capture information such as the dynamic behavior, motion trajectory, and scene transitions of objects in the video. Through these analyses, features reflecting the changing patterns of video content, the continuity of actions, or the complexity of the scene can be extracted in both time and space dimensions.

[0182] The encoding result features are the characteristics of the encoded result obtained by encoding the target video segment using a preset bitrate control parameter, where the preset bitrate control parameter is determined based on historical experience. For example, the bitrate control parameter is CRF. Based on historical experience, it is known that when the CRF is 32, the bitrate of the encoded result obtained by encoding the video using CRF meets expectations, therefore, the preset bitrate control parameter can be determined to be 32. As another example, the bitrate control parameter is CRF. Based on historical experience, it is known that when the CRF is 33, the quality of the encoded video obtained by encoding the video using CRF and then decoding the encoded result meets expectations, therefore, the preset bitrate control parameter can be determined to be 33.

[0183] The encoding result features include one or more of the following: the preset bitrate control parameters and the multi-method assessment fusion (VMAF) metric for the encoding result obtained by encoding the target video segment using the preset bitrate control parameters, the preset bitrate control parameters and the bitrate of the encoding result obtained by encoding the target video segment using the preset bitrate control parameters.

[0184] Precoding features include one or more of the following: intra-frame prediction coding information, inter-frame prediction coding information, motion vector information, quantization step size information, peak signal-to-noise ratio, prediction mode (or coding mode) information, block partitioning information, video bitrate of the precoded result, VMAF value of the precoded result, video width and height of the target video segment, and video frame rate of the target video segment.

[0185] Intra-frame prediction refers to predicting the content of the current block within the same frame using the encoded information around the pixel block. It is suitable for I-frames (keyframes) and helps with random access and error recovery in video. Inter-frame prediction uses the similarity between consecutive frames to predict the content of the current frame. It is typically used for P-frames (forward prediction frames) and B-frames (bidirectional prediction frames) and can significantly improve compression efficiency. Motion vectors, in inter-frame prediction, represent the displacement of the current block relative to the corresponding block in the reference frame. They are used to compensate for motion and are crucial for reducing spatial redundancy in video compression. The quantization step size during encoding affects the degree of compression and video quality. Block partitioning information and prediction mode information refer to the macroblock (or sub-macroblock) partitioning of video frames and the selection of the optimal prediction mode during encoding. Different block sizes and shapes (e.g., 16x16, 8x8, 4x4) and various prediction modes (e.g., DC, planar, various angle orientation predictions) are allowed during encoding to achieve optimal compression results.

[0186] 402. Input at least one of the above features into the first model to obtain the first relation output by the first model.

[0187] In this embodiment, the first model can predict the relationship between the video's bitrate control parameters and the bitrate of a reference encoding result based on the video's features. The reference encoding result is obtained by encoding the video using the video's bitrate control parameters. In this relationship, there is a one-to-one correspondence between the bitrate control parameters and the bitrate of the reference encoding result. Based on this relationship, it can be determined what bitrate the encoded result will be obtained by encoding the video using a specific bitrate control parameter. It can also be determined what bitrate control parameter, for a specific bitrate, will result in an encoded result with that specific bitrate.

[0188] Optionally, the first model is trained through the following process: Training features from the training video are input into the first model to obtain a first training relation predicted by the first model, where the first training relation is the bitrate control parameter of the training video and the bitrate of the encoding result obtained by encoding the training video using the bitrate control parameter. Based on a first difference between the first training relation and the first ground truth (GT), a first loss of the first model is obtained, where the first difference is positively correlated with the first loss. The parameters of the first model are updated based on the first loss until the first loss converges, completing the training of the first model.

[0189] When the first model is used to predict the relationship between the bitrate control parameters of a video and the bitrate of a reference coding result based on video features, the video coding device inputs at least one feature into the first model. The first model can then predict the relationship between the target bitrate control parameters of the target video segment and the bitrate of the target coding result based on at least one feature of the target video segment. The target coding result is obtained by encoding the target video segment using the target bitrate control parameters. Therefore, the first relationship is the relationship between the target bitrate control parameters of the target video segment and the bitrate of the target coding result. Optionally, the first relationship is a curve representing the relationship between the target bitrate control parameters of the target video segment and the bitrate of the target coding result.

[0190] 403. Based on the first relationship mentioned above, determine the target bitrate control parameters for the target video segment.

[0191] Because the first relation determines the optimal bitrate control parameter for encoding a video at a specific bitrate, resulting in a specific bitrate, the video encoding device can determine the target bitrate control parameter based on this first relation, thereby ensuring the target bitrate of the encoded result meets the requirements. For example, if the target bitrate requirement is to not exceed a bitrate threshold, then the first relation can be used to determine the bitrate control parameter that ensures the target bitrate does not exceed the threshold, which can then be used as the target bitrate control parameter.

[0192] 404. Determine the difference between the above target bitrate control parameter and the above baseline bitrate control parameter to obtain the above target deviation.

[0193] The video encoding device determines the target bitrate control parameter minus the reference bitrate control parameter as the target deviation.

[0194] In this embodiment, after extracting features from a target video segment to obtain at least one feature, the video encoding device inputs the at least one feature into a first model to obtain a first relationship output by the first model. This first relationship represents the relationship between the target bitrate control parameter of the target video segment and the bitrate of the target encoding result. Based on this first relationship, the video encoding device determines the target bitrate control parameter of the target video segment, ensuring that the bitrate of the target encoding result obtained by encoding the target video segment using the target bitrate control parameter meets the requirements. After obtaining the target bitrate control parameter, the difference between the target bitrate control parameter and the reference bitrate control parameter is determined, thus obtaining the target deviation.

[0195] In this embodiment, the video encoding device can also obtain the target deviation by performing the following steps:

[0196] 501. Perform feature extraction on the above target video segment to obtain at least one feature of the above target video segment.

[0197] The implementation of this step can be found in step 401, and will not be repeated here.

[0198] 502. Input at least one of the above features into the second model to obtain the second relationship output by the second model.

[0199] In this embodiment, the second model can predict the relationship between the video's bitrate control parameters and the video quality of the reference encoding result based on the video's features. It should be understood that the video quality of the encoded result is the quality of the video obtained by decoding the encoded result; optionally, this quality is represented by VMAF. In this relationship, there is a one-to-one correspondence between the bitrate control parameters and the video quality of the reference encoding result. Based on this relationship, it can be determined what the video quality of the encoded result is for a specific bitrate control parameter. Based on this relationship, it can also be determined what bitrate control parameter is needed to encode the video to achieve a specific video quality.

[0200] Optionally, the second model is trained through the following process: The training features of the training video are input into the second model to obtain the second training relation predicted by the second model. The second training relation represents the bitrate control parameter of the training video and the video quality of the encoded result obtained by encoding the training video using the bitrate control parameter. Based on the second difference between the first training relation and the second ground truth (GT), a second loss of the second model is obtained, where the second difference is positively correlated with the second loss. The parameters of the second model are updated based on the second loss until the second loss converges, completing the training of the second model.

[0201] In the case where the second model is used to predict the relationship between the bitrate control parameters of a video and the video quality of a reference coding result based on video features, after the video coding device inputs at least one feature into the second model, the second model can predict the relationship between the target bitrate control parameters of the target video segment and the video quality of the target coding result based on at least one feature of the target video segment. Therefore, the second relationship is the relationship between the target bitrate control parameters of the target video segment and the video quality of the target coding result. Optionally, the second relationship is a curve representing the relationship between the target bitrate control parameters of the target video segment and the video quality of the target coding result.

[0202] 503. Based on the second relationship mentioned above, determine the target bitrate control parameters for the target video segment.

[0203] Because the second relationship determines the optimal bitrate control parameter for encoding a specific video quality, ensuring the encoded video meets that quality standard, the video encoding device can determine the target bitrate control parameter based on this relationship, thereby ensuring the target encoded video quality meets the requirements. For example, if the target bitrate requirement is no lower than a video quality threshold, the second relationship can be used to determine the bitrate control parameter that ensures the target bitrate is not lower than the video quality threshold, and this parameter can then be used as the target bitrate control parameter.

[0204] 504. Determine the difference between the above target bitrate control parameters and the above baseline bitrate control parameters to obtain the above target deviation.

[0205] The video encoding device determines the target bitrate control parameter minus the reference bitrate control parameter as the target deviation.

[0206] In this embodiment, after extracting features from a target video segment to obtain at least one feature, the video encoding device inputs that feature into a second model to obtain a second relationship output by the second model. This second relationship represents the relationship between the target bitrate control parameter of the target video segment and the video quality of the target encoding result. Based on this second relationship, the video encoding device determines the target bitrate control parameter of the target video segment, ensuring that the video quality of the target encoding result obtained by encoding the target video segment using the target bitrate control parameter meets the requirements. After obtaining the target bitrate control parameter, the difference between the target bitrate control parameter and the reference bitrate control parameter is determined, thus obtaining the target deviation.

[0207] As an optional implementation, before performing step 403, the video encoding device also performs the following steps: inputting at least one feature into the second model to obtain the second relation output by the second model. The implementation of this step can be found in step 502, and will not be described again here.

[0208] After obtaining the second relationship, the video encoding device performs the following steps during step 403: based on the first and second relationships, determines the bitrate control parameters that make the target encoding result meet the preset requirements, and uses them as the target bitrate control parameters.

[0209] In this embodiment, the preset requirements include requirements for the bitrate of the target encoded result and requirements for the video quality of the target encoded result. Since the first relationship is the relationship between the target bitrate control parameter of the target video segment and the bitrate of the target encoded result, and the second relationship is the relationship between the target bitrate control parameter of the target video segment and the video quality of the target encoded result, the video encoding device determines the target bitrate control parameter of the target video segment based on the first and second relationships, so that both the bitrate and video quality of the target encoded result meet the requirements.

[0210] It should be understood that the target video segment and target deviation in the embodiments of this application are descriptive objects selected for the purpose of concisely describing the technical solution, and should not be construed as the video encoding device determining one of the (n+1) deviations to be encoded solely by obtaining the target deviation. In practical applications, the video encoding device can determine any one of the (n+1) deviations to be encoded by obtaining the target deviation.

[0211] Optionally, both the first and second models described above can be deep learning models. Please refer to [link / reference]. Figure 2 , Figure 2 This is a schematic diagram of the structure of a deep learning model provided in an embodiment of this application. Figure 2 As shown, this deep learning model can interface with the feature extraction module used to implement the method flow in step 401 above.

[0212] like Figure 2 As shown, the feature extraction module can include video content feature extraction, online complete encoding feature extraction, and fast pre-encoding feature extraction. Video content feature extraction is used to extract video content features, online complete encoding feature extraction is used to extract encoding result features, and fast pre-encoding feature extraction is used to extract pre-encoded features. Figure 2 As shown, video content features include texture information, parameterless video quality assessment, and spatiotemporal information. Encoding result features include `anchor_crf`, `anchor_vmaf`, and `anchor_bitrate`. When extracting features from the video segment to be encoded, the feature extraction module performs a complete encoding based on `anchor_crf`, obtaining the corresponding `anchor_vmaf` and `anchor_bitrate`. `anchor_crf` is a preset bitrate control parameter; `anchor_vmaf` represents the video quality of the target encoding result when the video segment is encoded using `anchor_crf`; and `anchor_bitrate` represents the bitrate of the target encoding result when the video segment is encoded using `anchor_crf`. For example, the bitrate control parameter is CRF. Based on historical experience, with a CRF of 32, the bitrate of the encoded video obtained using CRF meets expectations; therefore, `anchor_crf` can be determined to be 32. Precoding features include prediction mode, coded frame type features, and video width / height / frame rate features. The coded frame type features include intra-frame prediction coding information or inter-frame prediction coding information. By using a feature extraction module to extract features from the video to be encoded, at least one feature of any given video segment can be obtained.

[0213] like Figure 2 As shown, a deep learning model may include a first base-net, a dangling operator, and a second base-net connected in sequence.

[0214] In one possible implementation, at least one feature of the video segment to be encoded output by the feature extraction module can be input into a first base network, which then predicts a first relation or a second relation based on the at least one feature. Figure 2The curves shown are pred 1pass vmaf / bitrate, where pred 1pass bitrate represents the first relationship and pred 1pass vmaf represents the second relationship.

[0215] In one possible implementation, the feature extraction module outputs a video segment to be encoded with more than one feature. The server can first concatenate at least one feature output by the feature extraction module to obtain the corresponding target feature. This target feature is then input into a first base network, which predicts either a first relation or a second relation based on the target feature. It should be understood that feature concatenation integrates features from different sources or of different types to obtain a more comprehensive feature set. Simultaneously, it allows the network to learn the interactions and combination effects between features, thereby improving the model's predictive performance.

[0216] like Figure 2 As shown, the first base network can include a batch normalization layer, a self-attention layer, a residual feature layer, and a fully connected layer connected in sequence. Below, we will combine... Figure 2 The internal structure of the first basic network is shown, and the process of the first basic network predicting the first relation or the second relation is described.

[0217] First, the target feature, obtained by concatenating at least one feature, is input into a batch normalization layer. This layer performs batch normalization on at least one feature contained in the target feature and outputs at least one normalized feature. It should be understood that normalization can transform features at different scales to the same scale, for example, adjusting the values ​​of at least one feature to the range of [0, 1], which is beneficial for comparison and combination between different features.

[0218] Then, at least one normalized feature is input into a self-attention mechanism layer, which extracts the correlation information between the at least one feature and integrates the at least one feature into a deep feature based on the correlation information. In one illustrated embodiment, the correlation information between the at least one feature may include the correlation coefficient between the at least one feature.

[0219] Then, the depth features output from the self-attention mechanism layer are input to the residual feature layer, which performs residual correction on the depth features to obtain the residual features. In one illustrated implementation, as... Figure 2As shown, the residual feature layer may include two residual blocks (res-blocks). It should be understood that setting residual blocks can help the network learn the residual, i.e., the difference between the output and the input, more easily, promote deeper feature learning, and effectively solve the gradient vanishing problem common in deep networks, etc. This application does not make specific limitations in this regard.

[0220] Finally, the residual features output by the residual feature layer are input into the fully connected layer, which then predicts and outputs the first or second relation based on these residual features.

[0221] Thus, by using at least one feature combination as input to the first base network, and employing reasonable normalization methods and self-attention mechanisms in the first base network, this application can maximize network performance and improve the prediction accuracy of the first or second relationship.

[0222] Optionally, when extracting features from the video segment to be encoded, the feature extraction module performs a complete encoding of the video segment based on anchor_crf, obtaining the corresponding anchor_vmaf and anchor_bitrate. The true first relation corresponding to the video segment to be encoded should pass through the first anchor point (anchor_crf, anchor_bitrate) composed of anchor_crf and anchor_bitrate, and the true second relation corresponding to the video segment to be encoded should pass through the second anchor point (anchor_crf, anchor_vmaf) composed of anchor_crf and anchor_vmaf.

[0223] Based on this, such as Figure 2 As shown, this application further includes a dangling operator connected to the aforementioned first base network in the deep learning model. This dangling operator can further dangling the first or second relation predicted by the first base network based on the aforementioned video features (anchor_crf, anchor_vmaf, and anchor_bitrate), resulting in the dangling-processed first or second relation. Figure 2 The pred vmaf / bitrate dangling curves are shown, where pred bitrate represents the first relationship after dangling and pred vmaf represents the second relationship after dangling.

[0224] The first relation after dangling passes through the first anchor point (anchor_crf, anchor_bitrate), and the second relation after dangling passes through the second anchor point (anchor_crf, anchor_vmaf), thereby improving the prediction accuracy of the deep learning model and making the predicted relation closer to the true relation. Figure 2 The label shows the vmaf / bitrate curve.

[0225] In one illustrated embodiment, the dangling process may specifically include: adding a first bias to all bitrates of the predicted first relation, wherein the first bias may be equal to anchor_bitrate minus the bitrate corresponding to anchor_crf in the first relation; and adding a second bias to all video qualities of the predicted second relation, wherein the second bias may be equal to anchor_vmaf minus the video quality corresponding to anchor_crf in the crf-vmaf curve.

[0226] Thus, after the dangling process, the first relationship passes through the first anchor point (anchor_crf, anchor_bitrate), and the second relationship passes through the second anchor point (anchor_crf, anchor_vmaf).

[0227] In one illustrated implementation, to further improve the prediction accuracy of the deep learning model, such as Figure 2 As shown, this application also sets up a second basic network connected to the aforementioned dangling operator in the deep learning model. This second basic network can perform residual correction on the first or second relation after dangling processing to obtain the residual-corrected first or second relation, i.e. Figure 2 The pred 2pass vmaf / bitrate curves shown are: pred 2pass bitrate represents the first relation after residual correction, and pred 2pass vmaf represents the second relation after residual correction.

[0228] It should be noted that this application does not impose any particular limitation on the specific structure of the second basic network. In one illustrated embodiment, the second basic network may perform residual learning based on skip connections to enable the model to have better predictive performance, etc., and this application does not impose any specific limitation in this regard.

[0229] Optionally, deep learning models can be trained through supervised training. Specifically, such as... Figure 2As shown, the supervised training process may include: using the L1 loss function to supervise the difference between each point in the first or second relation output by the first base network and the label point pair in the first or second relation output by the second base network in the deep learning model, where label is the label. Simultaneously, the L1 loss function is used to supervise the gradient difference (i.e.,... Figure 2 In the gradloss (in CRF), the gradient can be defined as the difference between the vmaf_diff value and the bitrate_diff value of adjacent CRF points. The vmaf_diff value is the difference between the two video qualities corresponding to adjacent CRF points, and the bitrate_diff value is the difference between the two bitrates corresponding to adjacent CRF points.

[0230] As an optional implementation, the n segmentation frame numbers are determined through a scene detection process. The scene detection process can be performed by the aforementioned video encoding device, or by any electronic device other than the video encoding device. The scene detection process includes:

[0231] 601. When the scene of consecutive i-frame video frames in the video to be encoded changes from the first scene to the second scene, determine the number of reference video frames in the i-frame video frames that have the second scene.

[0232] In this embodiment, i is an integer greater than 1. Both the first scene and the second scene can be arbitrary scenes, and the first scene is different from the second scene. In step 601, when the video encoding device determines that a scene change occurs in i consecutive frames, it determines the number of reference video frames containing the changed scene to obtain a reference number.

[0233] 602. When the number of references is greater than or equal to the first threshold, the smallest frame number of the reference video frame is determined as the segmentation frame number.

[0234] If the number of reference frames is greater than or equal to the first threshold, it indicates that there are a large number of video frames with the second scene. This suggests that the appearance of the second scene in i consecutive video frames is not due to misidentification of the scene. Therefore, the video encoding device determines the smallest frame number of the reference video frames as the segmentation frame number, and then the segmentation of the first and second scenes can be completed using the segmentation frame number. For example, let A represent the video frame with the first scene, and let B represent the reference video frame. The i consecutive video frames are AAAB, and the first threshold is 1. Since the number of B frames is greater than or equal to 1, the segmentation frame number is the frame number of B.

[0235] As an optional implementation, the scene of consecutive i-frame video frames switches from a first scene to a second scene, and then switches back to the first scene. In this implementation, the scene detection process further includes: determining the initial number of initial video frames containing the first scene in the i-frame video frames before executing step 602.

[0236] After obtaining the initial quantity, step 602 in the scene detection process specifically includes the following steps: when the reference quantity is greater than or equal to the first threshold and the initial quantity is less than or equal to the reference quantity, the minimum frame number of the reference video frame is determined as the segmentation frame number.

[0237] If the initial number is less than or equal to the reference number, it means that the number of video frames with the first scene does not exceed the number of video frames with the second scene. This indicates that the second scene did not appear between the first scenes due to misidentification of the scene. Therefore, the video encoding device determines the smallest frame number of the reference video frame as the segmentation frame number, which can be used to segment the first and second scenes. For example, let A represent the initial video frame and B represent the reference video frame. The i consecutive video frames are AAABBBBA, and the first threshold is 1. Since the number of A's is equal to the number of B's, and the number of B's ​​is greater than or equal to 1, the segmentation frame number is the frame number of the first B's.

[0238] Optionally, the scene detection process further includes: after determining the initial number of initial video frames containing the first scene in the i-frame video frames, if the initial number is greater than the reference number, determining the frame number after the maximum frame number as the segmentation frame number, where the maximum frame number is the maximum frame number of the initial video frames. The initial number being greater than the reference number indicates that the number of video frames containing the first scene is greater than the number of video frames containing the second scene, meaning the second scene appears between the first scenes due to misidentification. Therefore, the video encoding device determines the maximum frame number as the segmentation frame number to avoid errors in the segmentation frame number caused by misidentification of scenes. For example, let A represent the initial video frame and B represent the reference video frame. The consecutive i-frame video frames are AAAABBAAAA. Since the number of A's equals the number of B's, the segmentation frame number is the frame number of the video frame after the last A. For example, if the frame number of the last A is 100, then the segmentation frame number is 101.

[0239] It should be understood that when the scene of consecutive i-frame video frames switches from the first scene to the second scene and then back to the first scene, if the initial number is greater than the reference number and the reference number is greater than the first threshold, the video encoding device determines the maximum frame number as the segmentation frame number.

[0240] As an optional implementation, the scene of consecutive i-frame video frames switches from a first scene to a second scene, and then from the second scene to a third scene. The first scene, second scene, and third scene can all be arbitrary scenes, and any two of the first scene, second scene, and third scene are different. The scene detection process further includes: before step 602, determining the initial number of initial video frames containing the first scene in the i-frame video frames.

[0241] After obtaining the initial quantity, step 602 in the scene detection process specifically includes the following steps: when the reference quantity is greater than or equal to the first threshold and the initial quantity is less than or equal to the second threshold, the minimum frame number of the reference video frame is determined as the segmentation frame number.

[0242] If the initial number is less than or equal to the second threshold, it indicates that the number of video frames with the first scene is small. This means that the scene switching from the first scene to the second scene in i consecutive video frames is not due to misidentification of the scene. Therefore, the video encoding device determines the minimum frame number of the reference video frame as the segmentation frame number, and then the segmentation of the first scene and the second scene can be completed through the segmentation frame number.

[0243] As an optional implementation, after determining the initial number, the scene detection step further includes the following steps: if the initial number is greater than a second threshold and the reference number is less than a third threshold, determine the number of video frames in the i-frame that have the third scene to be confirmed. If the number of frames to be confirmed is greater than or equal to the third threshold, determine the minimum frame number of the first video frame as the segmentation frame number.

[0244] If the initial number is greater than the second threshold, it indicates that there are a large number of video frames with the first scene. This suggests that the scene switching from the first scene to the second scene in consecutive i-frames may be due to misidentification. Therefore, the video encoding device determines whether the second scene is due to misidentification based on the reference number. A high reference number indicates that it is not due to misidentification, while a low reference number indicates that it is due to misidentification. In this embodiment, the video encoding device uses a third threshold to determine whether the reference number is high or low, where the third threshold is lower than the second threshold. Specifically, a reference number less than the third threshold indicates a low reference number, and a reference number greater than or equal to the third threshold indicates a high reference number. Therefore, if the initial number is greater than the second threshold and the reference number is less than the third threshold, the video encoding device determines that it is not necessary to separate the first scene and the second scene. It then continues to determine the number of video frames with the third scene to be confirmed, in order to further determine whether the third scene needs to be separated from other scenes (the first scene and the second scene). If the number of frames to be confirmed is greater than or equal to the third threshold, it indicates a large number of frames to be confirmed, thus determining that the appearance of the third scene is not due to misidentification. Therefore, the video encoding device determines the smallest frame number of the video frames to be confirmed as the segmentation frame number. This segmentation frame number is then used to segment the second and third scenes, separating the third scene from other scenes (the first and second scenes). For example, let A represent the initial video frame, B represent the reference video frame, and C represent the video frames to be confirmed. The i consecutive video frames are AAABBCCCC, the second threshold is 2, and the third threshold is 3. Since the number of A frames is greater than 2 and the number of B frames is less than 3, A and B are not segmented. However, since the number of C frames is greater than 3, the video encoding device segments B and C, and the segmentation frame number is the frame number of the first C frame.

[0245] It should be understood that the first, second, and third scenarios in this embodiment are the descriptive objects selected for the concise description of the technical solution. In practical applications, the number of switching scenarios in consecutive i-frame video frames can be greater than 3. If the first scene satisfies the conditions required for the first scenario, the last scene satisfies the conditions required for the third scenario, and all intermediate scenes satisfy the conditions required for the second scenario, then the video encoding device uses the smallest frame number in the last scene as the segmentation frame number. For example, let A, B, C, D, E, and F represent video frames with different scenarios. The consecutive i-frame video frames are AAABBCCDDEEFFFF, the second threshold is 2, and the third threshold is 3. Since the number of A is greater than 2, and the number of B, C, D, and E are all less than 3, A and B are not segmented, nor are B and C, C and D, or D and E. However, since the number of F is greater than 3, the video encoding device segments E and F, and the segmentation frame number is the frame number of the first F.

[0246] Based on the technical solutions provided in the embodiments of this application, the embodiments of this application also provide a possible application scenario. Please refer to... Figure 3 , Figure 3 This application provides a schematic diagram of the architecture of a video-on-demand system. Figure 3 As shown, the video-on-demand system 1 includes a client 11, a client 12, and a video encoding device 13. Both client 11 and client 12 have a communication connection with the video encoding device 13. Both client 11 and client 12 can upload videos to the video encoding device 13 through this communication connection, and can also obtain videos from the video-on-demand platform running on the video encoding device 13 through this communication connection. Optionally, client 11 and client 12 can be one of the following: a mobile phone, a computer, a tablet computer, or a wearable smart device. For example, client 11 is a mobile phone and client 12 is a computer; or both client 11 and client 12 are tablet computers. Optionally, the video encoding device 13 is a server.

[0247] In this video-on-demand scenario, the video encoding device 13 serves as the server running the video-on-demand platform. After receiving the video to be encoded uploaded by a user from the video-on-demand platform, the video encoding device 13 can determine n segmentation frame numbers from the video to be encoded using the method for obtaining segmentation frame numbers provided in this embodiment, and then store these n segmentation frame numbers as a CSV file. Based on the method for determining the target deviation provided in this embodiment, the video encoding device 13 can determine the bitrate control parameters for (n+1) video segments to be encoded, and then determine (n+1) encoding deviations based on the bitrate control parameters of the (n+1) video segments to be encoded and the baseline bitrate control parameters. The video encoding device 13 then inputs the video to be encoded, the CSV file, the baseline bitrate control parameters, and the (n+1) encoding deviations to the encoder, so that the encoder encodes the (n+1) video segments to be encoded based on the baseline bitrate control parameters and the (n+1) encoding deviations, obtaining the video encoding result of the video to be encoded output by the encoder. Finally, the video encoding result is published to the video-on-demand platform.

[0248] Those skilled in the art will understand that, in the above-described method of the specific implementation, the order in which each step is written does not imply a strict execution order and does not constitute any limitation on the implementation process. The specific execution order of each step should be determined by its function and possible internal logic.

[0249] The methods of the embodiments of this application have been described in detail above, and the apparatus of the embodiments of this application is provided below.

[0250] Please see Figure 4 , Figure 4This is a schematic diagram of a video encoding device provided in an embodiment of this application. The video encoding device 2 includes: an input unit 21 and an encoding unit 22. Optionally, the video encoding device 2 further includes: a conversion unit 23 and a determination unit 24. Specifically:

[0251] Input unit 21 is used to input the video to be encoded, n segmented frame numbers, a reference bitrate control parameter, and (n+1) encoding deviations to the encoder. The segmented frame numbers are the frame numbers of the video frames in the video to be encoded. The n segmented frame numbers are used to determine (n+1) segments of video to be encoded from the video to be encoded. The n is a positive integer. The encoding deviations correspond one-to-one with the segments of video to be encoded. The encoding deviations are the deviations of the bitrate control parameter of the segment to be encoded from the reference bitrate control parameter.

[0252] Encoding unit 22 is used to determine (n+1) segments of video to be encoded from the video to be encoded by the encoder, which are determined by the n segmentation frame numbers;

[0253] The encoding unit 22 is used to determine (n+1) target bitrate control parameters by the encoder based on the reference bitrate control parameters and the (n+1) to-be-encoded deviations, wherein the video segments to be encoded correspond one-to-one with the target bitrate control parameters;

[0254] The encoding unit 22 is used to encode the (n+1) segments of video to be encoded by the encoder based on the (n+1) target bitrate control parameters, so as to obtain the video encoding result of the video to be encoded output by the encoder.

[0255] In any embodiment of this application, the video to be encoded is a video in a first color space, and the encoder is used to process the video in a second color space, wherein the first color space is different from the second color space;

[0256] The video encoding device 2 further includes: a conversion unit 23, used to convert the video to be encoded from the first color space to the second color space to obtain a converted video;

[0257] The encoding unit 22 is specifically used for:

[0258] The converted video, the n segmented frame numbers, the baseline bitrate control parameters, and the (n+1) unencoded deviations are input to the encoder;

[0259] The encoder determines (n+1) segments of converted video corresponding to the (n+1) segments of video to be encoded from the converted video based on the n segmented frame numbers;

[0260] The encoder determines (n+1) target bitrate control parameters for the (n+1) segments of converted video based on the reference bitrate control parameters and the (n+1) unencoded deviations, and the converted video segments correspond one-to-one with the target bitrate control parameters;

[0261] The encoder encodes the (n+1) video segments to be encoded based on the (n+1) target bitrate control parameters to obtain the video encoding result.

[0262] In conjunction with any embodiment of this application, the encoding unit 22 is specifically used for:

[0263] The encoder takes the first frame of the converted video and the video frames in the converted video that correspond to the n segmentation frame numbers as (n+1) key frames of the (n+1) segments of the converted video, and the key frames correspond one-to-one with the converted video segments.

[0264] The encoder encodes the (n+1) segments of the converted video based on the (n+1) keyframes, the reference bitrate control parameters, and the (n+1) encoding deviations to obtain the video encoding result.

[0265] In conjunction with any embodiment of this application, the encoding unit 22 is specifically used for:

[0266] The encoder determines the sum of the baseline bitrate control parameter and each of the (n+1) encoding deviations, thereby obtaining (n+1) target bitrate control parameters for the (n+1) converted video segments. The target bitrate control parameters correspond one-to-one with the converted video segments.

[0267] The encoder encodes the (n+1) converted video segments based on the (n+1) keyframes and the (n+1) target bitrate control parameters to obtain the video encoding result.

[0268] In any embodiment of this application, the (n+1) segments of video to be encoded include a target video segment, and the (n+1) offsets to be encoded include a target offset corresponding to the target video segment;

[0269] The video encoding device 2 further includes: a determining unit 24 configured to:

[0270] Feature extraction is performed on the target video segment to obtain at least one feature of the target video segment;

[0271] The at least one feature is input into the first model to obtain the first relation output by the first model;

[0272] The first model is used to predict the relationship between the bitrate control parameters of the video and the bitrate of the reference encoding result based on the video's features. The reference encoding result is obtained by encoding the video using the video's bitrate control parameters. The first relationship is the relationship between the target bitrate control parameters of the target video segment and the bitrate of the target encoding result. The target encoding result is obtained by encoding the target video segment using the target bitrate control parameters.

[0273] Based on the first relationship, the target bitrate control parameters of the target video segment are determined;

[0274] Alternatively, the at least one feature can be input into a second model to obtain a second relationship output by the second model. The first model is used to predict the relationship between the bitrate control parameters of the video and the video quality of the reference coding result based on the features of the video. The second relationship is the relationship between the target bitrate control parameters of the target video segment and the video quality of the target coding result.

[0275] Based on the second relationship, the target bitrate control parameters of the target video segment are determined;

[0276] The difference between the target bitrate control parameter and the baseline bitrate control parameter is determined to obtain the target deviation.

[0277] In conjunction with any embodiment of this application, the determining unit 24 is further configured to:

[0278] The at least one feature is input into the second model to obtain the second relationship output by the second model;

[0279] Based on the first relationship and the second relationship, a bitrate control parameter is determined to make the target encoding result meet the preset requirements, which is used as the target bitrate control parameter. The preset requirements include the bitrate requirements of the target encoding result and the video quality requirements of the target encoding result.

[0280] In any embodiment of this application, in the video to be encoded, video frames located on both sides of the segmentation frame number have different scenes, and video frames located between two adjacent segmentation frame numbers have the same scene.

[0281] In any embodiment of this application, the n segmentation frame numbers are determined through a scene detection process, which includes:

[0282] When the scene of consecutive i video frames in the video to be encoded changes from the first scene to the second scene, the reference number of reference video frames with the second scene in the i video frames is determined, where i is an integer greater than 1.

[0283] If the number of references is greater than or equal to the first threshold, the smallest frame number of the reference video frame is determined as the segmentation frame number.

[0284] In any embodiment of this application, after the scene of the consecutive i-frame video frames switches from the first scene to the second scene, and then switches back to the first scene from the second scene, the scene detection process further includes:

[0285] Determine the initial number of initial video frames containing the first scene in the i-frame video frame;

[0286] When the number of references is greater than or equal to a first threshold, determining the minimum frame number of the reference video frames as the segmentation frame number includes:

[0287] If the number of references is greater than or equal to a first threshold and the number of initial references is less than or equal to the number of references, the minimum frame number of the reference video frame is determined as the segmentation frame number.

[0288] In conjunction with any embodiment of this application, the scene detection process further includes: after determining the initial number of initial video frames containing the first scene in the i-frame video frames,

[0289] If the initial number is greater than the reference number, the frame number after the maximum frame number is determined as the segmentation frame number, and the maximum frame number is the maximum frame number of the initial video frame.

[0290] In any embodiment of this application, the scene of the consecutive i-frame video frames switches from the first scene to the second scene, and then switches from the second scene to the third scene;

[0291] The scene detection process further includes: before determining the minimum frame number of the reference video frame as the segmentation frame number when the number of references is greater than or equal to a first threshold.

[0292] Determine the initial number of initial video frames containing the first scene in the i-frame video frame;

[0293] When the number of references is greater than or equal to a first threshold, determining the frame number smaller than the first threshold among the frame numbers of the reference video frames as the segmentation frame number includes:

[0294] If the number of references is greater than or equal to the first threshold and the number of initial references is less than or equal to the second threshold, the minimum frame number of the reference video frame is determined as the segmentation frame number.

[0295] In conjunction with any embodiment of this application, the scene detection process further includes: after determining the initial number of initial video frames containing the first scene in the i-frame video frames,

[0296] If the initial number is greater than the second threshold and the reference number is less than the third threshold, determine the number of unconfirmed video frames in the i-frame video frame that have the third scene.

[0297] If the number of frames to be confirmed is greater than or equal to the third threshold, the smallest frame number of the video frames to be confirmed is determined as the segmentation frame number, and the third threshold is less than the second threshold.

[0298] In this embodiment, the video to be encoded and n segmented frame numbers are used. Each segmented frame number is a frame number of a video frame in the video to be encoded. The n segmented frame numbers are used to determine (n+1) segments of video to be encoded from the video to be encoded. The encoding deviation is the deviation of the bitrate control parameter of the video segment to be encoded from the reference bitrate control parameter. After inputting the video to be encoded, the n segmented frame numbers, the reference bitrate control parameter, and the (n+1) encoding deviations into the encoder, the video encoding device can determine the (n+1) segments of video to be encoded from the video to be encoded, determined by the n segmented frame numbers. Furthermore, the encoder can determine (n+1) target bitrate control parameters based on the reference bitrate control parameter and the (n+1) encoding deviations. Finally, the encoder can encode the (n+1) segments of video to be encoded based on the (n+1) target bitrate control parameters to obtain the video encoding result of the video to be encoded output by the encoder. Therefore, the encoder can encode (n+1) segments of video to be encoded by encoding the video in one operation, thereby reducing the amount of data processing required to encode different segments of video to be encoded using different bitrate control parameters, and improving the encoding speed. Furthermore, since the amount of data processing can be reduced, it means that this encoding method requires less hardware resources, thereby reducing hardware costs.

[0299] In some embodiments, the functions or modules of the apparatus provided in this application can be used to perform the methods described in the above method embodiments. The specific implementation can be referred to the description of the above method embodiments, and for the sake of brevity, it will not be repeated here.

[0300] Figure 5This is a schematic diagram of the hardware structure of an electronic device provided in an embodiment of this application. The electronic device 3 includes a processor 31 and a memory 32. Optionally, the electronic device 3 also includes an input device 33 and an output device 34. The processor 31, memory 32, input device 33, and output device 34 are coupled together via connectors, which include various interfaces, transmission lines, or buses, etc., and are not limited in this embodiment. It should be understood that in the various embodiments of this application, coupling refers to mutual connection in a specific way, including direct connection or indirect connection through other devices, such as through various interfaces, transmission lines, buses, etc.

[0301] Processor 31 may include one or more processors, such as one or more CPUs. If the processor is a CPU, it may be a single-core CPU or a multi-core CPU. Optionally, processor 31 may be a processor group consisting of multiple CPUs, with the multiple processors coupled to each other via one or more buses. Optionally, the processor may also be other types of processors, etc., which are not limited in this embodiment.

[0302] The memory 32 can be used to store computer program instructions, as well as various types of computer program code, including program code for executing the scheme of this application. Optionally, the memory includes, but is not limited to, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM), or compact disc read-only memory (CD-ROM), which is used for related instructions and data.

[0303] Input device 33 is used to input data and / or signals, and output device 34 is used to output data and / or signals. Input device 33 and output device 34 can be independent devices or an integrated device.

[0304] It is understood that in this embodiment of the application, the memory 32 can be used not only to store related instructions, but also to store related data. This embodiment of the application does not limit the specific data stored in the memory.

[0305] Understandable, Figure 5 This is merely a simplified design of an electronic device. In practical applications, the electronic device may also include other necessary components, including, but not limited to, any number of input / output devices, processors, memories, etc., and all electronic devices that can implement the embodiments of this application are within the protection scope of this application.

[0306] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0307] Those skilled in the art will readily understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here. Those skilled in the art will also readily understand that the various embodiments of this application have different focuses, and for the sake of convenience and brevity, the same or similar parts may not be repeated in different embodiments. Therefore, parts not described or not described in detail in one embodiment can be referred to the descriptions in other embodiments.

[0308] In the several embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.

[0309] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0310] In addition, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.

[0311] In the above embodiments, implementation can be achieved entirely or partially through software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented entirely or partially as a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of this application are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted through the computer-readable storage medium. The computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that integrates one or more available media. The available media may be magnetic media (e.g., floppy disks, hard disks, magnetic tapes), optical media (e.g., digital versatile discs (DVDs)), or semiconductor media (e.g., solid-state disks (SSDs)).

[0312] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. This program can be stored in a computer-readable storage medium, and when executed, it can include the processes described in the above method embodiments. The aforementioned storage medium includes various media capable of storing program code, such as read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

Claims

1. A video encoding method, characterized in that, The method includes: The video to be encoded, n segmented frame numbers, a reference bitrate control parameter, and n+1 encoding deviations are input to the encoder. The segmented frame number is the frame number of the video frame in the video to be encoded. The n segmented frame numbers are used to determine n+1 segments to be encoded from the video to be encoded. n is a positive integer. The encoding deviations correspond one-to-one with the segments to be encoded. The encoding deviations are the deviations of the bitrate control parameter of the segment to be encoded from the reference bitrate control parameter. The encoder determines n+1 segments of video to be encoded from the video to be encoded, which are determined by the n segmentation frame numbers; The encoder determines n+1 target bitrate control parameters based on the baseline bitrate control parameters and the n+1 to-be-encoded deviations, and the video segments to be encoded correspond one-to-one with the target bitrate control parameters. The encoder encodes the n+1 video segments to be encoded based on the n+1 target bitrate control parameters, and obtains the video encoding result of the video to be encoded output by the encoder.

2. The method according to claim 1, characterized in that, The video to be encoded is a video in a first color space, and the encoder is used to process the video in a second color space. The first color space is different from the second color space. Before inputting the video to be encoded, n segmented frame numbers, baseline bitrate control parameters, and n+1 encoding deviations to the encoder, the method further includes: The video to be encoded is converted from the first color space to the second color space to obtain the converted video; The step of inputting the video to be encoded, n segmented frame numbers, a reference bitrate control parameter, and n+1 encoding deviations into the encoder includes: The converted video, the n segmented frame numbers, the baseline bitrate control parameters, and the n+1 offsets to be encoded are input to the encoder. The step of determining the n+1 video segments to be encoded from the video to be encoded by the encoder, which are determined by the n segmentation frame numbers, includes: The encoder determines, based on the n segmented frame numbers, the n+1 segments of converted video corresponding to the n+1 segments of video to be encoded from the converted video; The step of determining n+1 target bitrate control parameters by the encoder based on the baseline bitrate control parameters and the n+1 to-be-encoded deviations includes: The encoder determines n+1 target bitrate control parameters for the n+1 converted video segments based on the baseline bitrate control parameters and the n+1 to be encoded deviations, and the converted video segments correspond one-to-one with the target bitrate control parameters; The process of encoding the n+1 video segments to be encoded using the encoder based on the n+1 target bitrate control parameters to obtain the video encoding result of the video to be encoded output by the encoder includes: The encoder encodes the n+1 video segments to be encoded based on the n+1 target bitrate control parameters to obtain the video encoding result.

3. The method according to claim 2, characterized in that, The process of encoding the n+1 video segments to be encoded using the encoder based on the n+1 target bitrate control parameters to obtain the video encoding result includes: The encoder uses the first frame of the converted video and the video frames in the converted video that correspond to the n segmentation frame numbers as the n+1 key frames of the n+1 converted video segments, and the key frames correspond one-to-one with the converted video segments. The encoder encodes the n+1 segments of the converted video based on the n+1 keyframes, the baseline bitrate control parameters, and the n+1 encoding deviations to obtain the video encoding result.

4. The method according to claim 3, characterized in that, The process involves encoding the n+1 segments of the converted video based on the n+1 keyframes, the baseline bitrate control parameters, and the n+1 encoding deviations to obtain the video encoding result, including: The encoder determines the sum of the baseline bitrate control parameter and each of the n+1 encoding deviations, thereby obtaining n+1 target bitrate control parameters for the n+1 converted video segments. The target bitrate control parameters correspond one-to-one with the converted video segments. The encoder encodes the n+1 converted video segments based on the n+1 keyframes and the n+1 target bitrate control parameters to obtain the video encoding result.

5. The method according to any one of claims 2 to 4, characterized in that, The n+1 video segments to be encoded include the target video segment, and the n+1 offsets to be encoded include the target offsets corresponding to the target video segment; Before inputting the converted video, the n segmented frame numbers, the baseline bitrate control parameters, and the n+1 to-be-encoded deviations into the encoder, the method further includes: Feature extraction is performed on the target video segment to obtain at least one feature of the target video segment; The at least one feature is input into the first model to obtain the first relation output by the first model; The first model is used to predict the relationship between the bitrate control parameters of the video and the bitrate of the reference encoding result based on the video's features. The reference encoding result is obtained by encoding the video using the video's bitrate control parameters. The first relationship is the relationship between the target bitrate control parameters of the target video segment and the bitrate of the target encoding result. The target encoding result is obtained by encoding the target video segment using the target bitrate control parameters. Based on the first relationship, the target bitrate control parameters of the target video segment are determined; Alternatively, the at least one feature can be input into a second model to obtain a second relationship output by the second model. The first model is used to predict the relationship between the bitrate control parameters of the video and the video quality of the reference coding result based on the features of the video. The second relationship is the relationship between the target bitrate control parameters of the target video segment and the video quality of the target coding result. Based on the second relationship, the target bitrate control parameters of the target video segment are determined; The difference between the target bitrate control parameter and the baseline bitrate control parameter is determined to obtain the target deviation.

6. The method according to claim 5, characterized in that, Before determining the target bitrate control parameter of the target video segment based on the first relationship, the method further includes: The at least one feature is input into the second model to obtain the second relationship output by the second model; The step of determining the target bitrate control parameters of the target video segment based on the first relationship includes: Based on the first relationship and the second relationship, a bitrate control parameter is determined to make the target encoding result meet the preset requirements, which is used as the target bitrate control parameter. The preset requirements include the bitrate requirements of the target encoding result and the video quality requirements of the target encoding result.

7. The method according to any one of claims 1 to 4, characterized in that, In the video to be encoded, video frames located on both sides of the segmentation frame number have different scenes, while video frames located between two adjacent segmentation frame numbers have the same scene.

8. The method according to claim 7, characterized in that, The n segmentation frame numbers are determined through a scene detection process, which includes: When the scene of i consecutive video frames in the video to be encoded changes from the first scene to the second scene, the reference number of reference video frames with the second scene in the i consecutive video frames is determined, where i is an integer greater than 1. If the number of references is greater than or equal to the first threshold, the smallest frame number of the reference video frame is determined as the segmentation frame number.

9. The method according to claim 8, characterized in that, After the scene of the consecutive i-frame video frames switches from the first scene to the second scene, and then switches back to the first scene from the second scene, the scene detection process further includes: when the reference number is greater than or equal to the first threshold, before determining the minimum frame number of the reference video frame as the segmentation frame number, determining the initial number of the initial video frames with the first scene in the i-frame video frames. When the number of references is greater than or equal to a first threshold, determining the minimum frame number of the reference video frames as the segmentation frame number includes: If the number of references is greater than or equal to a first threshold and the number of initial references is less than or equal to the number of references, the minimum frame number of the reference video frame is determined as the segmentation frame number.

10. The method according to claim 9, characterized in that, The scene detection process further includes: after determining the initial number of initial video frames containing the first scene in the i-frame video frames, if the initial number is greater than the reference number, determining a frame number after the maximum frame number as the segmentation frame number, wherein the maximum frame number is the maximum frame number of the initial video frame.

11. The method according to claim 8, characterized in that, After the scene of the consecutive i-frame video frames switches from the first scene to the second scene, and then switches from the second scene to the third scene, the scene detection process further includes: determining the initial number of the initial video frames containing the first scene in the i-frame video frames before determining the minimum frame number of the reference video frame as the segmentation frame number when the reference number is greater than or equal to the first threshold. When the number of references is greater than or equal to a first threshold, determining the frame number smaller than the first threshold among the frame numbers of the reference video frames as the segmentation frame number includes: If the number of references is greater than or equal to the first threshold and the number of initial references is less than or equal to the second threshold, the minimum frame number of the reference video frame is determined as the segmentation frame number.

12. The method according to claim 11, characterized in that, The scene detection process further includes: after determining the initial number of initial video frames containing the first scene in the i-frame video frames, if the initial number is greater than the second threshold and the reference number is less than the third threshold, determining the number of video frames to be confirmed containing the third scene in the i-frame video frames. If the number of frames to be confirmed is greater than or equal to the third threshold, the smallest frame number of the video frames to be confirmed is determined as the segmentation frame number, and the third threshold is less than the second threshold.

13. A video encoding device, characterized in that, The device includes: The acquisition unit is used to acquire the video to be encoded and n segmentation frame numbers, wherein the segmentation frame number is the frame number of the video frame in the video to be encoded, and the n segmentation frame numbers are used to divide the video to be encoded into n+1 segments to be encoded, where n is a positive integer; The acquisition unit is used to acquire the baseline bitrate control parameter and n+1 encoding deviations of n+1 video segments to be encoded. The encoding deviations correspond one-to-one with the video segments to be encoded, and the encoding deviations are the deviations of the bitrate control parameters of the video segments to be encoded from the baseline bitrate control parameter. The encoding unit is used to input the video to be encoded, the n segmented frame numbers, the reference bitrate control parameters, and the n+1 encoding deviations to the encoder, so that the encoder encodes the n+1 segments of the video to be encoded based on the reference bitrate control parameters and the n+1 encoding deviations, and obtains the video encoding result of the video to be encoded output by the encoder.

14. An electronic device, characterized in that, include: A processor and a memory, the memory being used to store computer program code, the computer program code including computer instructions, wherein, when the processor executes the computer instructions, the electronic device performs the method as described in any one of claims 1 to 12.

15. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, the computer program including program instructions that, when executed by a processor, cause the processor to perform the method according to any one of claims 1 to 12.

16. A computer program product, characterized in that, The computer program product includes a computer program or instructions; when the computer program or instructions are executed on a computer, the computer causes the computer to perform the method according to any one of claims 1 to 12.

Citation Information

Patent Citations

  • Video coding method and device, computer equipment and medium

    CN118055240A

  • Video processing method and device, and storage medium

    WO2021254139A1