Video Encoding Method and Apparatus, Storage Medium, and Electronic Device

By acquiring the quality weight of the video frame, the problem of poor visual perception of video encoding in the prior art is solved, and the user experience of video playback is improved.

CN114760475BActive Publication Date: 2025-07-25BEIJING KINGSOFT CLOUD NETWORK TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202110026092.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-01-08
Publication Date
2025-07-25
Estimated Expiration
2041-01-08

AI Technical Summary

Technical Problem

The prior art fails to effectively consider the impact of the subjective quality of a single-frame video on the overall video quality in video encoding, resulting in poor viewing.

Method used

By obtaining the quality weight of the video frame to be encoded, adjusting the fixed code rate coefficient CRF value, encoding the video frame according to the quality weight, allocating a larger code rate value to frames with a greater impact on the overall quality, and a smaller code rate value to frames with a less impact on the subjective quality.

Benefits of technology

On the premise of ensuring the overall quality of the video, the code rate of keyframes is improved, and the user's visual effect and visual experience during video playback is enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114760475B_ABST
    Figure CN114760475B_ABST
Patent Text Reader

Abstract

The present invention discloses a video encoding method and device, a storage medium, and an electronic device, and belongs to the field of video encoding technology. The method includes: obtaining the quality weight of the source video frame to be encoded, wherein the quality weight is used to characterize the degree of influence of the source video frame on the source video quality; adjusting the original fixed code rate coefficient CRF value of the source video frame based on the quality weight to obtain a target CRF value; encoding the source video frame according to the target CRF value. Through the present invention, the technical problem that the related technology adopts fixed CRF value encoding to cause poor video perception is solved, and the code rate of the key frame is improved while ensuring the overall quality of the video, thereby enhancing the user's viewing effect and viewing experience during video playback.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of video coding, and in particular, to a video coding method, an apparatus, a storage medium, and an electronic device. Background Art

[0002] When performing video coding in related technologies, it refers to a method of converting a file in the original video format into another video format file through compression technology. Common codec standards in video stream transmission include H.264, H.265, AVS, AV1, etc.

[0003] There are various bitrate control methods in related technologies during video coding. There is the CBR (Constant BitRate, fixed bitrate) mode with a constant bitrate, and the CRF (Constant Rate Factor, fixed bitrate coefficient) mode with a constant visual quality. Each bitrate control method ultimately achieves the effect of bitrate control by changing the QP value of each frame. Among them, the CRF mode uses the input CRF value as the reference QP (Quant parameter, quantization parameter). The way to save bitrate and maintain picture quality in video coding is achieved by optimizing the encoder or by optimizing objective metrics, and objective metrics are all used to guide coding, which will increase the compression amount of the video or increase the size of the video file.

[0004] The CRF method in related technologies usually only considers objective metrics and does not consider the impact of the subjective quality of a single video frame on the overall subjective quality of the video. However, the applicant has found that at different bitrates, different frames have different subjective feelings for users. For example, users are more concerned about the foreground picture frames, less concerned about the transition frames between scenes, more concerned about the close-up pictures, and not concerned about the distant view pictures, etc.

[0005] In view of the above problems existing in related technologies, no effective solution has been found yet. Summary of the Invention

[0006] Embodiments of the present invention provide a video coding method, an apparatus, a storage medium, and an electronic device.

[0007] According to one aspect of the embodiments of the present application, a video coding method is provided, including: obtaining a quality weight of a source video frame to be coded, where the quality weight is used to characterize the influence degree of the source video frame on the source video quality; adjusting an original fixed bitrate coefficient CRF value of the source video frame based on the quality weight to obtain a target CRF value; and coding the source video frame according to the target CRF value.

[0008] Further, adjusting the original CRF value of the source video frame based on the quality weight to obtain a target CRF value includes: comparing the quality weight with a preset threshold; if the quality weight is greater than the preset threshold, reducing the original CRF value of the source video frame to obtain a first target CRF value; if the quality weight is less than the preset threshold, increasing the original CRF value of the source video frame to obtain a second target CRF value; if the quality weight is equal to the preset threshold, maintaining the original CRF value of the source video frame to obtain a third target CRF value.

[0009] Further, encoding the source video frame according to the target CRF value includes: calculating a quantization parameter QP value based on the target CRF value; compressing the source video frame using the QP value to generate a target video frame, where the bitrate value of the target video frame is positively correlated with the quality weight.

[0010] Further, obtaining the quality weight of the source video frame to be encoded includes: calculating a first quality value of the source video frame and calculating a second quality value of the source video; using the following algorithm model to calculate the quality weight: V mos = F mos W; where V mos is the second quality value, F mos is a one-dimensional matrix composed of the first quality values of multiple source video frames, and W is the quality weight, and the source video includes the multiple source video frames.

[0011] Further, before calculating the quality weight using the algorithm model, the method further includes: collecting sample video data; determining the video quality of the sample video in the sample video data and the video frame quality of each frame of the sample video; using the video frame quality as input information and the video quality as output information to train a preset original model to obtain V mos = F mos W.

[0012] Further, determining the video quality of the sample video in the sample video data and the video frame quality of each frame of the sample video includes one of the following: calculating the video quality of the sample video in the sample video data using an image quality evaluation algorithm and calculating the video frame quality of each frame of the sample video using an image quality evaluation algorithm; receiving first label information of the sample video and second label information of the video frame quality of each frame of the sample video; respectively determining the first label information and the second label information as the video quality and the video frame quality.

[0013] Further, obtaining the quality weight of the source video frame to be encoded includes one of the following: obtaining the quality weights of multiple consecutive source video frames to be encoded in the source video in units of a preset number of frames, where the sum of the quality weights of the multiple consecutive source video frames is 1; obtaining the quality weights of all source video frames to be encoded in the source video frame by frame, where the sum of the quality weights of all source video frames is 1; sampling in the source video to obtain the quality weights of multiple discrete source video frames to be encoded, where the sum of the quality weights of the multiple discrete source video frames is 1.

[0014] According to another aspect of the embodiments of the present application, there is also provided a video encoding apparatus, including: an obtaining module, configured to obtain the quality weight of a source video frame to be encoded, where the quality weight is used to characterize the influence degree of the source video frame on the source video quality; an adjustment module, configured to adjust the original fixed code rate coefficient CRF value of the source video frame based on the quality weight to obtain a target CRF value; and an encoding module, configured to encode the source video frame according to the target CRF value.

[0015] Further, the adjustment module includes: a comparison unit, configured to compare the quality weight with a preset threshold; an adjustment unit, configured to, if the quality weight is greater than the preset threshold, reduce the original CRF value of the source video frame to obtain a first target CRF value; if the quality weight is less than the preset threshold, increase the original CRF value of the source video frame to obtain a second target CRF value; and if the quality weight is equal to the preset threshold, maintain the original CRF value of the source video frame to obtain a third target CRF value.

[0016] Further, the encoding module includes: a calculation unit, configured to calculate a quantization parameter QP value based on the target CRF value; and a compression unit, configured to compress the source video frame by using the QP value to generate a target video frame, where the code rate value of the target video frame is positively correlated with the quality weight.

[0017] Further, the obtaining module includes: a first calculation unit, configured to calculate a first quality value of the source video frame and calculate a second quality value of the source video; and a second calculation unit, configured to calculate the quality weight by using the following algorithm model: V mos = F mos W; where V mos is the second quality value, F mos is a one-dimensional matrix composed of the first quality values of multiple source video frames, and W is the quality weight, and the source video includes the multiple source video frames.

[0018] Further, the device further includes: an acquisition module, configured to acquire sample video data before the acquisition module calculates the quality weight by using an algorithm model; a determination module, configured to determine the video quality of the sample video in the sample video data, and the video frame quality of each frame of the sample video; a training module, configured to use the video frame quality as input information and the video quality as output information to train a preset original model to obtain V mos = F mos W.

[0019] Further, the determination module includes one of the following: a calculation unit, configured to calculate the video quality of the sample video in the sample video data by using an image quality evaluation algorithm, and calculate the video frame quality of each frame of the sample video by using an image quality evaluation algorithm; a determination unit, configured to receive first label information of the sample video and second label information of the video frame quality of each frame of the sample video; and determine the first label information and the second label information as the video quality and the video frame quality, respectively.

[0020] Further, the acquisition module includes one of the following: a first acquisition unit, configured to acquire the quality weights of multiple consecutive source video frames to be encoded in the source video in units of a preset number of frames, where the sum of the quality weights of the multiple consecutive source video frames is 1; a second acquisition unit, configured to acquire the quality weights of all source video frames to be encoded in the source video frame by frame, where the sum of the quality weights of all source video frames is 1; a third acquisition unit, configured to sample and acquire the quality weights of multiple discrete source video frames to be encoded in the source video, where the sum of the quality weights of the multiple discrete source video frames is 1.

[0021] According to another aspect of the embodiments of the present application, there is also provided a storage medium, including a stored program, where the program, when running, executes the above steps.

[0022] According to another aspect of the embodiments of the present application, there is also provided an electronic device, including a processor, a communication interface, a memory, and a communication bus, where the processor, the communication interface, and the memory communicate with each other through the communication bus; where: the memory is used to store a computer program; the processor is used to execute the steps in the above method by running the program stored on the memory.

[0023] The embodiments of the present application also provide a computer program product including instructions, when running on a computer, enabling the computer to execute the steps in the above method.

[0024] Through the present invention, the quality weight of the source video frame to be encoded is obtained, and then the original fixed Constant Rate Factor (CRF) value of the source video frame is adjusted based on the quality weight to obtain the target CRF value. Finally, the source video frame is encoded according to the target CRF value. By adjusting the CRF value of a single frame based on the influence degree of the single-frame image on the overall video quality, the method of subjective quality evaluation is applied to video encoding, thereby guiding the encoder to encode. For frames with a large quality weight, their influence on the overall quality is more important, and a larger bitrate value is allocated to them. For frames with a smaller weight, their influence on the subjective quality is smaller, and a relatively smaller bitrate value can be allocated accordingly. This solves the technical problem that the related art uses a fixed CRF value for encoding, resulting in poor video viewing experience. While ensuring the overall video quality, the bitrate of key frames is increased, enhancing the user viewing effect and viewing experience during video playback. BRIEF DESCRIPTION OF THE DRAWINGS

[0025] The drawings described herein are used to provide a further understanding of the present invention and constitute a part of this application. The schematic embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation to the present invention. In the drawings:

[0026] Figure 1 is a hardware structure block diagram of a server according to an embodiment of the present invention;

[0027] Figure 2 is a flowchart of a video encoding method according to an embodiment of the present invention;

[0028] Figure 3 is a schematic diagram of consecutive source video frames according to an embodiment of the present invention;

[0029] Figure 4 is a structure block diagram of a video encoding device according to an embodiment of the present invention;

[0030] Figure 5 is a structure block diagram of an electronic device for implementing an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0031] In order to enable those skilled in the art of the present technology to better understand the solutions of this application, the technical solutions in the embodiments of this application will be clearly and completely described below in conjunction with the drawings in the embodiments of this application. Obviously, the described embodiments are only a part of the embodiments of this application, rather than all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the scope of protection of this application. It should be noted that, without conflict, the embodiments in this application and the features in the embodiments can be combined with each other.

[0032] It should be noted that the terms "first", "second", etc. in the description, claims and the above-mentioned drawings of this application are used to distinguish similar objects, and do not necessarily describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments of this application described here can be implemented in an order different from those illustrated or described here. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, product or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0033] Embodiment 1

[0034] The method embodiment provided by the first embodiment of this application can be executed on a server, computer, imaging device, mobile phone, tablet or similar computing device. Taking running on a server as an example, Figure 1 is a hardware structure block diagram of a server according to an embodiment of the present invention. As Figure 1 shown, the server may include one or more ( Figure 1 only one is shown in the figure) processors 102 (the processor 102 may include, but is not limited to, processing devices such as a microprocessor MCU or a programmable logic device FPGA) and a memory 104 for storing data. Optionally, the above-mentioned server may further include a transmission device 106 for communication functions and an input / output device 108. Those of ordinary skill in the art can understand that Figure 1 the structure shown is only schematic and does not limit the structure of the above-mentioned server. For example, the server may further include more or fewer components than those Figure 1 shown in the figure, or have a different configuration from that Figure 1 shown in the figure.

[0035] The memory 104 can be used to store server programs. For example, software programs and modules of application software, such as the server program corresponding to a video encoding method in an embodiment of the present invention. The processor 102 executes various functional applications and data processing by running the server program stored in the memory 104, that is, implements the above-mentioned method. The memory 104 may include high-speed random access memory, and may also include non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memories. In some instances, the memory 104 may further include a memory remotely disposed relative to the processor 102, and these remote memories may be connected to the server through a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.

[0036] The transmission device 106 is used to receive or send data via a network. Specific examples of the above-mentioned network may include a wireless network provided by the communication provider of the server. In one example, the transmission device 106 includes a network adapter (Network Interface Controller, abbreviated as NIC), which can be connected to other network devices through a base station so as to communicate with the Internet. In one example, the transmission device 106 can be a Radio Frequency (RF) module, which is used to communicate with the Internet wirelessly.

[0037] In this embodiment, a video encoding method is provided. Figure 2 It is a flowchart of a video encoding method according to an embodiment of the present invention, as Figure 2 shown. The process includes the following steps:

[0038] Step S202, obtain the quality weight of the source video frame to be encoded, where the quality weight is used to characterize the influence degree of the source video frame on the source video quality;

[0039] The quality weight in this embodiment is a pre-allocated weight coefficient for the source video frame. The greater the influence degree of a single frame on the overall quality of the source video, the greater the quality weight. Conversely, the smaller the quality weight. At the same time, in order to meet the constant visual quality of the video, the quality weights of multiple source video frames meet the normalization constraint conditions.

[0040] Step S204, adjust the original fixed bitrate coefficient CRF value of the source video frame based on the quality weight to obtain the target CRF value;

[0041] In the encoding mode of constant visual quality (CRF), the CRF value is related to the video bitrate after encoding. The bitrate is the number of data bits transmitted per unit time during data transmission, the bitrate per second of the encoded video. The bitrate is also called the sampling rate. The larger the sampling rate per unit time, the higher the precision, and the closer the processed file is to the original file. The bitrate affects the distortion degree of the video. The higher the bitrate, the clearer the video. Conversely, the picture is rough and has many mosaics.

[0042] Step S206, encode the source video frame according to the target CRF value.

[0043] In the CRF mode (i.e., the constant visual quality encoding mode) of this embodiment, the specified original CRF value is replaced with an adjusted target CRF as an encoding parameter for encoding, so as to stabilize the output visual quality. For video frames with a small quality weight, their distortion is relatively difficult to be detected by the naked eye. Therefore, the CRF value can be increased for encoding to save the video bitrate; while for videos with a large quality weight, even a small distortion is more likely to be detected by people. By reducing the CRF value for encoding and increasing the bitrate, the distortion of the video is reduced.

[0044] Through the above steps, the quality weight of the source video frame to be encoded is obtained, and then the original fixed code rate coefficient CRF value of the source video frame is adjusted based on the quality weight to obtain the target CRF value. Finally, the source video frame is encoded according to the target CRF value. By adjusting the CRF value of a single frame based on the degree of influence of the single-frame image on the overall video quality, the method of subjective quality evaluation is applied to video encoding, so as to guide the encoder to encode. For frames with a large quality weight, their influence on the overall quality is more important, and a larger bitrate value is allocated to them, while for frames with a smaller weight, their influence on subjective quality is smaller, and a relatively smaller bitrate value can be allocated accordingly. This solves the technical problem that the video looks poor due to encoding with a fixed CRF value in the related art. While ensuring the overall quality of the video, the bitrate of key frames is increased, and the user's viewing experience and viewing effect during video playback are enhanced.

[0045] In an implementation manner of this embodiment, adjusting the original CRF value of the source video frame based on the quality weight to obtain the target CRF value includes:

[0046] S11, comparing the quality weight with a preset threshold;

[0047] Optionally, the preset threshold in this embodiment is a middle value of the quality weight, such as 0, 0.5, etc. If the quality weight is greater than the preset threshold, the influence degree of the source video frame on the source video quality is larger; conversely, if the quality weight is less than the preset threshold, the influence degree of the source video frame on the source video quality is smaller.

[0048] S12, if the quality weight is greater than the preset threshold, reducing the original CRF value of the source video frame to obtain the first target CRF value; if the quality weight is less than the preset threshold, increasing the original CRF value of the source video frame to obtain the second target CRF value; if the quality weight is equal to the preset threshold, maintaining the original CRF value of the source video frame to obtain the third target CRF value.

[0049] In an implementation manner of this embodiment, encoding the source video frame according to the target CRF value includes: calculating the QP (Quant parameter) value based on the target CRF value; compressing the source video frame using the QP value to generate the target video frame, where the bitrate value of the target video frame is positively correlated with the quality weight.

[0050] The greater the quality weight, the greater the true bitrate of the video after encoding; the smaller the quality weight, the smaller the true bitrate of the video after encoding. When playing the target video containing the target video frame, the distortion of the key frame that the user cares about is smaller, and the distortion of the non-key frame that the user cares about is larger, and it is not easily noticed by the user, improving the overall visual quality of the video.

[0051] By obtaining the relationship between each frame of the video and the overall quality of the video, and based on the impact of the video frame on the overall quality of the video, the importance degree of each frame is obtained, so as to guide the encoder to encode. For frames with large weights, their impact on the overall quality is more important, so a larger bitrate value is allocated to them; while for frames with smaller weights, their impact on the subjective is smaller, and a smaller bitrate value can be allocated accordingly to save bitrate.

[0052] In an implementation manner of this embodiment, the quality weight of a single frame is calculated through an algorithm model. Obtaining the quality weight of the source video frame to be encoded includes:

[0053] S21, calculating the first quality value of the source video frame and calculating the second quality value of the source video;

[0054] S22, calculating the quality weight using the following algorithm model: V mos = F mos W;

[0055] where, V mos is the second quality value, F mos is a one-dimensional matrix composed of the first quality values of multiple source video frames, W is the quality weight, and the source video includes multiple source video frames.

[0056] In an example based on the above algorithm model, the source video includes three source video frames, frame 1, frame 2, and frame 3, and their quality weights are 1.7, 0.4, and -1.1 respectively, F mos = [2, 1, 3], V mos = 0.5.

[0057] In this embodiment, the algorithm model can be pre-trained based on sample data and a basic model, or obtained from a third party. In an example, before calculating the quality weight using the algorithm model, it further includes: collecting sample video data; determining the video quality of the sample video in the sample video data and the video frame quality of each frame of the sample video; using the video frame quality as input information and the video quality as output information to train a preset original model to obtain the algorithm model, that is, V mos = F mos W.

[0058] The video frame quality and the video quality are the input label data and the output label data of a preset original model respectively. Through training, an algorithm model is obtained that can output the quality weight of the relative overall video based on the video frame quality. Since there are a vast number of source video frames in the source video, a vast number of quality weights will also be generated corresponding to the source video frames. The quality of the source video frames can be obtained through quantitative calculation and automatically output by the algorithm model without manual tagging, which improves the automation degree of video encoding, increases the encoding speed, and reduces the influence of artificially assigned quality weights on the encoding result.

[0059] In this embodiment, the quality of the sample can be determined by using an image quality evaluation algorithm analysis or by using the method of manual marking, and the video quality of the sample video in the sample video data and the video frame quality of each frame of the sample video are determined, including one of the following: calculating the video quality of the sample video in the sample video data by using an Image Quality Assessment (IQA) algorithm, and calculating the video frame quality of each frame of the sample video by using an image quality evaluation algorithm; receiving the first label information of the sample video and the second label information of the video frame quality of each frame of the sample video; and respectively determining the first label information and the second label information as the video quality and the video frame quality.

[0060] In this embodiment, obtaining the quality weight of the source video frame to be encoded includes one of the following:

[0061] Example 1: Obtain the quality weights of multiple consecutive source video frames to be encoded in the source video in units of a preset number of frames, where the sum of the quality weights of the multiple consecutive source video frames is 1;

[0062] Optionally, the preset number of frames can be 16, 24, etc. Figure 3 This is a schematic diagram of consecutive source video frames in an embodiment of the present invention. Taking 16 consecutive frames of images as one encoding unit, the source video includes several encoding units, and the sum of the quality weights of 16 consecutive source video frames is 1;

[0063] Example 2: Obtain the quality weights of all source video frames to be encoded in the source video frame by frame, where the sum of the quality weights of all source video frames is 1;

[0064] Example 3: Sample in the source video to obtain the quality weights of multiple discrete source video frames to be encoded, where the sum of the quality weights of the multiple discrete source video frames is 1.

[0065] During the adoption process of this example, sampling can be performed at a predetermined frame interval or randomly.

[0066] In an implementation scenario of this embodiment, a method is proposed to guide the encoding process of an encoder by considering the impact of the subjective quality of each video frame on the overall subjective quality of the video, and to apply subjective quality evaluation to video encoding. The implementation process includes:

[0067] (1) Dataset collection and annotation: Crawl video data from the network and use manual annotation to score each frame of the video and the overall quality of the video; in the present invention, the quality of each frame of the video and the overall quality of the video are manually annotated. It is also possible to use the method of image quality evaluation to obtain the relative frame quality. Finally, it is also feasible to obtain the relationship between the relative frame quality and the overall quality of the video.

[0068] (2) Use a deep learning model to learn the relationship between each frame of the video and the overall quality of the video. The input of the model is the quality of each frame of the video, denoted as a one-dimensional matrix Fmos (i.e., a vector), and the output is the overall quality of the video Vmos. The relationship between Fmos and Vmos is expressed as follows:

[0069] V mos = F mos W; where W is the influence weight of each frame of the video on the overall video.

[0070] (3) Map the weight W of each frame to Δcrf. On the basis of the originally specified crf in the video encoding process, add Δcrf to reassign a corresponding crf to each frame. For frames with a large weight, their impact on the overall quality is more important. Therefore, after mapping, the Δcrf value is negative, and the original crf value decreases, and a larger bitrate value is allocated to it; while for frames with a smaller weight, their impact on the subjective quality is smaller. After mapping, the Δcrf value is positive, and the original crf value increases, and a smaller bitrate value can be allocated accordingly to save bitrate.

[0071] By obtaining the relationship between each frame of the video and the overall quality of the video, and according to the impact of each video frame on the overall quality of the video, the importance of each frame is obtained, and the method of subjective quality evaluation is applied to encoding, so as to guide the encoder to encode. For frames with a large quality weight, their impact on the overall quality is more important, and a larger bitrate value is allocated to them; while for frames with a smaller weight, their impact on the subjective quality is smaller, and a smaller bitrate value can be allocated accordingly to save bitrate.

[0072] Through the description of the above embodiments, those skilled in the art can clearly understand that the method according to the above embodiments can be implemented by means of software plus a necessary general hardware platform. Of course, it can also be implemented by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art can be embodied in the form of a software product. The computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes several instructions for causing a terminal device (which can be a mobile phone, computer, server, or network device, etc.) to execute the methods described in various embodiments of the present invention.

[0073] Embodiment 2

[0074] In this embodiment, a video encoding device is further provided to implement the above embodiments and preferred implementation manners. Those that have been described will not be repeated. As used below, the term "module" can be a combination of software and / or hardware that can implement a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, implementation in hardware, or a combination of software and hardware is also possible and contemplated.

[0075] Figure 4 is a structural block diagram of a video encoding device according to an embodiment of the present invention. As Figure 4 shown, the device includes: an acquisition module 40, an adjustment module 42, and an encoding module 44, where

[0076] The acquisition module 40 is used to acquire the quality weight of the source video frame to be encoded, where the quality weight is used to characterize the influence degree of the source video frame on the source video quality;

[0077] The adjustment module 42 is used to adjust the original fixed code rate factor CRF value of the source video frame based on the quality weight to obtain a target CRF value;

[0078] The encoding module 44 is used to encode the source video frame according to the target CRF value.

[0079] Optionally, the adjustment module includes: a comparison unit for comparing the quality weight with a preset threshold; an adjustment unit for, if the quality weight is greater than the preset threshold, reducing the original CRF value of the source video frame to obtain a first target CRF value; if the quality weight is less than the preset threshold, increasing the original CRF value of the source video frame to obtain a second target CRF value; if the quality weight is equal to the preset threshold, maintaining the original CRF value of the source video frame to obtain a third target CRF value.

[0080] Optionally, the encoding module includes: a calculation unit configured to calculate a quantization parameter QP value based on the target CRF value; a compression unit configured to compress the source video frame using the QP value to generate a target video frame, wherein a bitrate value of the target video frame is positively correlated with the quality weight.

[0081] Optionally, the obtaining module includes: a first calculation unit configured to calculate a first quality value of the source video frame and a second quality value of the source video; a second calculation unit configured to calculate the quality weight using the following algorithm model: V mos = F mos W; where V mos is the second quality value, F mos is a one-dimensional matrix composed of first quality values of multiple source video frames, and W is the quality weight, and the source video includes the multiple source video frames.

[0082] Optionally, the apparatus further includes: a collection module configured to collect sample video data before the obtaining module calculates the quality weight using the algorithm model; a determination module configured to determine a video quality of a sample video in the sample video data and a video frame quality of each frame of the sample video; a training module configured to train a preset original model with the video frame quality as input information and the video quality as output information to obtain V mos = F mos W.

[0083] Optionally, the determination module includes one of the following: a calculation unit configured to calculate the video quality of the sample video in the sample video data using an image quality evaluation algorithm and calculate the video frame quality of each frame of the sample video using an image quality evaluation algorithm; a determination unit configured to receive first label information of the sample video and second label information of the video frame quality of each frame of the sample video; and determine the first label information and the second label information as the video quality and the video frame quality, respectively.

[0084] Optionally, the obtaining module includes one of the following: a first obtaining unit configured to obtain quality weights of multiple consecutive source video frames to be encoded in the source video in units of a preset number of frames, wherein a sum of the quality weights of the multiple consecutive source video frames is 1; a second obtaining unit configured to obtain quality weights of all source video frames to be encoded in the source video frame by frame, wherein a sum of the quality weights of all source video frames is 1; a third obtaining unit configured to sample and obtain quality weights of multiple discrete source video frames to be encoded in the source video, wherein a sum of the quality weights of the multiple discrete source video frames is 1.

[0085] It should be noted that the above-mentioned various modules can be implemented by software or hardware. For the latter, it can be implemented in the following ways, but not limited to: the above-mentioned modules are all located in the same processor; or, the above-mentioned various modules are respectively located in different processors in any combination form.

[0086] Embodiment 3

[0087] An embodiment of the present invention also provides a storage medium, in which a computer program is stored, wherein the computer program is set to execute the steps in any one of the above method embodiments when running.

[0088] Optionally, in this embodiment, the above storage medium can be set to store a computer program for executing the following steps:

[0089] S1. Obtain the quality weight of the source video frame to be encoded, where the quality weight is used to characterize the influence degree of the source video frame on the source video quality;

[0090] S2. Adjust the original fixed code rate factor CRF value of the source video frame based on the quality weight to obtain a target CRF value;

[0091] S3. Encode the source video frame according to the target CRF value.

[0092] Optionally, in this embodiment, the above storage medium may include but is not limited to: various media such as USB flash drives, read-only memories (ROM for short), random access memories (RAM for short), mobile hard disks, magnetic disks or optical discs that can store computer programs.

[0093] An embodiment of the present invention also provides an electronic device, including a memory and a processor, where a computer program is stored in the memory, and the processor is set to run the computer program to execute the steps in any one of the above method embodiments.

[0094] Optionally, the above electronic device may further include a transmission device and an input / output device, where the transmission device is connected to the above processor, and the input / output device is connected to the above processor.

[0095] Optionally, in this embodiment, the above processor can be set to execute the following steps through a computer program:

[0096] S1. Obtain the quality weight of the source video frame to be encoded, where the quality weight is used to characterize the influence degree of the source video frame on the source video quality;

[0097] S2. Adjust the original constant rate factor (CRF) value of the source video frame based on the quality weight to obtain the target CRF value;

[0098] S3. Encode the source video frame according to the target CRF value.

[0099] Optionally, specific examples in this embodiment may refer to the examples described in the above embodiments and optional implementation manners, and will not be elaborated herein.

[0100] Figure 5 It is a structural block diagram of an electronic device for implementing an embodiment of the present invention. As Figure 5 shown, it includes a processor 51 and a memory 52 for storing data, which are connected through a communication bus 54. It further includes a communication interface 53 connected to the communication bus 54 for adaptively connecting to other components or external devices.

[0101] The serial numbers of the above embodiments of the present application are only for description and do not represent the superiority or inferiority of the embodiments.

[0102] In the above embodiments of the present application, the descriptions of the respective embodiments have their own emphases. For parts not detailed in a certain embodiment, reference may be made to the relevant descriptions of other embodiments.

[0103] In several embodiments provided by the present application, it should be understood that the disclosed technical content can be implemented in other ways. Among them, the device embodiments described above are only illustrative. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection between each other can be through some interfaces. The indirect coupling or communication connection of units or modules can be in an electrical or other form.

[0104] The units described as separate components may or may not be physically separated. The components shown as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0105] In addition, the functional units in each embodiment of the present application can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above integrated units can be implemented in the form of hardware or in the form of software functional units.

[0106] When the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of this application. The aforementioned storage medium includes: various media that can store program codes, such as USB flash drives, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), mobile hard disks, magnetic disks, or optical discs.

[0107] The above are only the preferred embodiments of this application. It should be noted that for those of ordinary skill in the art, without departing from the principle of this application, several improvements and refinements can still be made, and these improvements and refinements should also be regarded as the protection scope of this application.

Claims

1. A video encoding method, characterized in that, Including: Obtain the quality weight of the source video frame to be encoded, where the quality weight is used to characterize the influence degree of the source video frame on the source video quality; Adjust the original constant rate factor CRF value of the source video frame based on the quality weight to obtain the target CRF value; Encode the source video frame according to the target CRF value; Among them, obtaining the quality weight of the source video frame to be encoded includes: calculating a first quality value of the source video frame and calculating a second quality value of the source video; calculating the quality weight using the following algorithm model: ; where is the second quality value, is a one-dimensional matrix composed of the first quality values of multiple source video frames, is the quality weight, and the source video includes the multiple source video frames; among them, before calculating the quality weight using the algorithm model, it further includes: collecting sample video data; determining the video quality of the sample video in the sample video data and the video frame quality of each frame of the sample video; using the video frame quality as input information and the video quality as output information to train a preset original model to obtain ; Among them, obtaining the quality weight of the source video frame to be encoded includes one of the following: obtaining the quality weights of multiple consecutive source video frames to be encoded in the source video in units of a preset number of frames, where the sum of the quality weights of multiple consecutive source video frames is 1; obtaining the quality weights of all source video frames to be encoded in the source video frame by frame, where the sum of the quality weights of all source video frames is 1; sampling in the source video to obtain the quality weights of multiple discrete source video frames to be encoded, where the sum of the quality weights of multiple discrete source video frames is 1.

2. The method according to claim 1, characterized in that, Adjusting the original CRF value of the source video frame based on the quality weight to obtain the target CRF value includes: Compare the quality weight with a preset threshold; If the quality weight is greater than the preset threshold, decrease the original CRF value of the source video frame to obtain the first target CRF value; if the quality weight is less than the preset threshold, increase the original CRF value of the source video frame to obtain the second target CRF value; if the quality weight is equal to the preset threshold, maintain the original CRF value of the source video frame to obtain the third target CRF value.

3. The method according to claim 1, characterized in that, Encoding the source video frame according to the target CRF value includes: Calculate the quantization parameter QP value based on the target CRF value; Compress the source video frame using the QP value to generate a target video frame, where the bitrate value of the target video frame is positively correlated with the quality weight.

4. The method according to claim 1, characterized in that, Determine the video quality of the sample video in the sample video data, and the video frame quality of each frame of the sample video, including one of the following: Calculate the video quality of the sample video in the sample video data using an image quality evaluation algorithm, and calculate the video frame quality of each frame of the sample video using an image quality evaluation algorithm; Receive the first label information of the sample video and the second label information of the video frame quality of each frame of the sample video; determine the first label information and the second label information as the video quality and the video frame quality respectively.

5. A video encoding device, characterized in that, Including: An acquisition module for obtaining the quality weight of the source video frame to be encoded, where the quality weight is used to characterize the influence degree of the source video frame on the source video quality; An adjustment module for adjusting the original constant rate factor CRF value of the source video frame based on the quality weight to obtain the target CRF value; An encoding module for encoding the source video frame according to the target CRF value; Among them, the obtaining module includes: a first calculation unit, configured to calculate a first quality value of the source video frame and a second quality value of the source video; and a second calculation unit, configured to calculate the quality weight by using the following algorithm model: ; where is the second quality value, is a one-dimensional matrix composed of first quality values of multiple source video frames, is the quality weight, and the source video includes the multiple source video frames; Among them, the device further includes: an acquisition module, configured to acquire sample video data before the acquisition module calculates the quality weight by using an algorithm model; a determination module, configured to determine the video quality of the sample video in the sample video data, and the video frame quality of each frame of the sample video; a training module, configured to use the video frame quality as input information and the video quality as output information to train a preset original model to obtain ; Among them, the obtaining module includes one of the following: a first obtaining unit, configured to obtain quality weights of multiple consecutive source video frames to be encoded in the source video in units of a preset number of frames, where the sum of the quality weights of the multiple consecutive source video frames is 1; a second obtaining unit, configured to obtain quality weights of all source video frames to be encoded in the source video frame by frame, where the sum of the quality weights of all source video frames is 1; a third obtaining unit, configured to sample and obtain quality weights of multiple discrete source video frames to be encoded in the source video, where the sum of the quality weights of the multiple discrete source video frames is 1.

6. A storage medium, characterized in that, The storage medium includes a stored program, where the program, when running, executes the method steps described in any one of claims 1 to 4 above.

7. An electronic device, comprising a processor, a communication interface, a memory, and a communication bus, wherein, The processor, communication interface, and memory complete mutual communication through a communication bus; among them: The memory is used to store a computer program; The processor is configured to execute the method steps described in any one of claims 1 to 4 by running the program stored on the memory.

Citation Information

Patent Citations

  • Video processing method and device

    CN109286825A