Video encoding method and apparatus, and electronic device

By building a buffer area and dynamically adjusting the video attribute information, the problem of quality and bandwidth balance of video encoding in a dynamic network environment is solved, and efficient resource utilization and network stability are achieved.

CN120075447BActive Publication Date: 2025-07-01SHENZHEN SHANGMI NETWORK TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510535619.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-27
Publication Date
2025-07-01
Estimated Expiration
2045-04-27

AI Technical Summary

Technical Problem

Existing video encoding technology cannot effectively balance video quality with transmission bandwidth in dynamic network environments, resulting in buffer underflow or bandwidth waste when the network environment changes.

Method used

By building a buffer area, the video attribute information is dynamically adjusted based on the packet loss rate, round trip time and BBR congestion control algorithm, such as resolution, frame rate and quantization step size, and multiple encoding optimizations are performed to adapt to network changes to ensure that the buffered data meets preset requirements.

Benefits of technology

It realizes effective balance of video quality and transmission bandwidth in a dynamic network environment, reduces data volume, improves resource usage efficiency, and avoids network congestion and playback jitter.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120075447B_ABST
    Figure CN120075447B_ABST
Patent Text Reader

Abstract

The present invention provides a video encoding method, apparatus, and electronic device, relating to the technical field of video processing. The method includes: constructing a buffer area, obtaining the current bitrate of a target video, and determining initial buffer data based on the current state of the buffer area and the current bitrate of the target video. If the initial buffer data does not meet the preset buffer requirements, adjust the attribute information of the target video based on the current state of the buffer area to obtain a first intermediate video corresponding to the target video, obtain the bitrate of the first intermediate video, and determine intermediate buffer data based on the current state of the buffer area and the bitrate of the first intermediate video. If the intermediate buffer data does not meet the preset buffer requirements, perform encoding optimization based on the scene type of the first intermediate video to obtain a second intermediate video, and transmit the second intermediate video based on the size of the second intermediate video and the buffer area, thereby improving the resource utilization efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of video processing technology, and in particular to a video encoding method and device, and electronic equipment. Background Art

[0002] Video encoding refers to the process of converting raw video data into a format more suitable for storage and transmission through a specific compression algorithm. The amount of raw video data may be very large. For example, an uncompressed high-definition movie may require hundreds of GB or even more storage space. Through video encoding, various compression technologies can be used, such as removing redundant information in the video and using efficient encoding algorithms, to significantly reduce the amount of data, making it easier to store in storage devices such as hard disks and optical disks, and also reducing storage costs.

[0003] In the field of video encoding, rate control is the core technology for balancing video quality and transmission bandwidth. In existing technologies, constant bit rate (CBR) is often used, but the fixed bit rate forces a constant output bit rate and cannot adapt to dynamic network environments. When the network bandwidth decreases, it is easy to cause buffer underflow and cause freezes; when the bandwidth is sufficient, the image quality cannot be improved due to the bit rate upper limit, resulting in bandwidth waste. Another example is the use of variable bit rate (VBR). Although the bit rate fluctuates, its adjustment is usually based on offline content analysis (such as two-pass encoding), which cannot respond to network changes in real time, and the bit rate fluctuation range lacks precise constraints, which can easily cause network congestion or playback jitter. Summary of the invention

[0004] In view of the above technical problems, the technical solution adopted by the present invention is:

[0005] According to a first aspect of the present invention, a video encoding method is provided, the method comprising the following steps:

[0006] A buffer area is constructed on the client, where the client is used to send the target video, and the size of the buffer area is determined based on the packet loss rate q of the transmission channel of the target video between the client and the server, the round-trip time RTT, and the BBR congestion control algorithm;

[0007] Based on the current bit rate of the acquired target video and the current state of the buffer area, initial buffer data is determined, where the initial buffer data includes: initial buffer time, which is the cache time of the target video in the buffer area at the current bit rate;

[0008] If the initial buffer data does not meet the preset buffer requirement, based on the current state of the buffer area, adjusting the attribute information of the target video to obtain a first intermediate video corresponding to the target video, the attribute information including resolution, frame rate, and quantization step size;

[0009] Based on the intermediate bitrate of the acquired first intermediate video and the current state of the buffer area, determine intermediate buffer data, where the intermediate buffer data includes: an intermediate buffer time, which is the buffer time of the first intermediate video in the buffer area at the intermediate bitrate;

[0010] If the intermediate buffer data does not meet the preset buffer requirements, perform encoding optimization based on the scene type of the first intermediate video, obtain a second intermediate video, and transmit the second intermediate video.

[0011] According to a second aspect of the present invention, there is provided a video encoding device, the device comprising:

[0012] A buffer area construction module for constructing a buffer area, wherein the buffer area is constructed on the client, where the client is used to send a target video, and the size of the buffer area is determined based on the packet loss rate q, round-trip time RTT, and BBR congestion control algorithm of the transmission channel between the client and the server;

[0013] An initial buffer data module for determining initial buffer data based on the current bitrate of the acquired target video and the current state of the buffer area, where the initial buffer data includes: an initial buffer time, which is the caching time of the target video in the buffer area at the current bitrate;

[0014] An adjustment module for, if the initial buffer data does not meet the preset buffer requirements, adjusting the attribute information of the target video based on the current state of the buffer area to obtain a first intermediate video corresponding to the target video, where the attribute information includes resolution, frame rate, and quantization step size;

[0015] An intermediate buffer data module for determining intermediate buffer data based on the intermediate bitrate of the acquired first intermediate video and the current state of the buffer area, where the intermediate buffer data includes: an intermediate buffer time, which is the buffer time of the first intermediate video in the buffer area at the intermediate bitrate;

[0016] An optimization module for, if the intermediate buffer data does not meet the preset buffer requirements, performing encoding optimization based on the scene type of the first intermediate video, obtaining a second intermediate video, and transmitting the second intermediate video.

[0017] According to a third aspect of the present invention, there is provided an electronic device, comprising: a processor, a memory, and a computer program stored on the memory and executable on the processor, where the processor implements the foregoing method when executing the computer program.

[0018] The present invention has at least the following beneficial effects:

[0019] In summary, a buffer area is constructed, the current bitrate of the target video is obtained, and based on the current state of the buffer area and the current bitrate of the target video, initial buffer data is determined. If the initial buffer data does not meet the preset buffer requirements, based on the current state of the buffer area, the attribute information of the target video is adjusted to obtain a first intermediate video corresponding to the target video. The bitrate of the first intermediate video is obtained, and based on the current state of the buffer area and the bitrate of the first intermediate video, intermediate buffer data is determined. If the intermediate buffer data does not meet the preset buffer requirements, encoding optimization is performed based on the scene type of the first intermediate video to obtain a second intermediate video. Based on the size of the second intermediate video and the buffer area, the second intermediate video is transmitted. The present invention determines whether to perform two video encodings on the target video based on the buffer area, improving the resource utilization efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention, and those of ordinary skill in the art can obtain other drawings based on these drawings without creative efforts.

[0021] Figure 1 It is a flowchart of a video encoding method provided in Embodiment 1 of the present invention;

[0022] Figure 2 It is a flowchart of S100 provided in Embodiment 1 of the present invention;

[0023] Figure 3 It is a flowchart of S300 provided in Embodiment 1 of the present invention;

[0024] Figure 4 It is a flowchart of S320 provided in Embodiment 1 of the present invention;

[0025] Figure 5 It is a flowchart of an embodiment of S500 provided in Embodiment 1 of the present invention;

[0026] Figure 6 It is a flowchart of another embodiment of S500 provided in Embodiment 1 of the present invention;

[0027] Figure 7 It is a schematic structural diagram of a video encoding device provided in Embodiment 2 of the present invention;

[0028] Figure 8 It is a schematic structural diagram of a buffer area construction module provided in Embodiment 2 of the present invention;

[0029] Figure 9 It is a schematic structural diagram of an adjustment module provided in Embodiment 2 of the present invention;

[0030] Figure 10 It is a schematic structural diagram of the new attribute information acquisition sub-module provided in the second embodiment of the present invention;

[0031] Figure 11 It is a schematic structural diagram of an embodiment of the optimization module provided in the second embodiment of the present invention;

[0032] Figure 12 It is a schematic structural diagram of another embodiment of the optimization module provided in the second embodiment of the present invention. Detailed implementation manners

[0033] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative efforts belong to the scope of protection of the present invention.

[0034] It should be noted that the terms "first", "second", etc. in the specification and claims of the present invention and the above-mentioned drawings are used to distinguish similar tasks, and do not have to be used to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present invention described here can be implemented in an order other than those illustrated or described here. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or server including a series of steps or units does not have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0035] Embodiment 1

[0036] The first embodiment of the present invention provides a video encoding method, and the method includes the following steps, as Figure 1 shown:

[0037] S100, build a buffer area on the client, where the client is used to send the target video, and the size of the buffer area is determined based on the packet loss rate q, round-trip time RTT, and BBR congestion control algorithm of the transmission channel between the client and the server. It can be understood that the buffer area is the buffer area of the client for the target video, and is used to temporarily store the data that has been encoded but has not been sent out through the transmission channel and the data that has been sent to the server but has not received a receipt.

[0038] S200, determine the initial buffer data based on the currently obtained bit rate of the target video and the current state of the buffer area.

[0039] Specifically, the initial buffer data at least includes an initial buffer time, which is the caching time of the target video in the buffer area at the current bitrate. The initial buffer data may further include: the current data volume in the buffer area, the buffer change trend, and the maximum data volume threshold of the buffer area. The buffer change trend is an increasing rate or a decreasing rate.

[0040] S300. If the initial buffer data does not meet the preset buffer requirements, based on the current state of the buffer area, adjust the attribute information of the target video to obtain a first intermediate video corresponding to the target video. The attribute information includes resolution, frame rate, and quantization step size.

[0041] In an embodiment of the present invention, if the initial buffer data does not meet the preset buffer requirements, it includes: the initial buffer time > the preset buffer time threshold. In another embodiment of the present invention, if the initial buffer data does not meet the preset buffer requirements, it includes: the initial buffer time > the preset buffer time threshold, and the buffer change trend is an increasing rate.

[0042] Specifically, if the initial buffer data meets the preset buffer requirements, directly encode and transmit the target video. Those skilled in the art know that any method of encoding and transmitting the target video in the prior art belongs to the protection scope of the present invention and will not be elaborated here.

[0043] Specifically, based on the adjusted attribute information, use an encoder to optimize the encoding of the target video to obtain a first intermediate video.

[0044] S400. Based on the intermediate bitrate of the obtained first intermediate video and the current state of the buffer area, determine the intermediate buffer data. The intermediate buffer data includes: an intermediate buffer time, which is the buffering time of the first intermediate video in the buffer area at the intermediate bitrate. It can be understood that after adjusting the attribute information of the target video, make a judgment again.

[0045] S500. If the intermediate buffer data does not meet the preset buffer requirements, perform encoding optimization based on the scene type of the first intermediate video, obtain a second intermediate video, and transmit the second intermediate video. That is, perform encoding again through the scene type to further reduce the data volume of the target video.

[0046] Specifically, if the intermediate buffer data meets the preset buffer requirements, transmit the first intermediate video.

[0047] In summary, a buffer area is constructed, the current bitrate of the target video is obtained, and based on the current state of the buffer area and the current bitrate of the target video, initial buffer data is determined. If the initial buffer data does not meet the preset buffer requirements, based on the current state of the buffer area, the attribute information of the target video is adjusted, the first intermediate video corresponding to the target video is obtained, the bitrate of the first intermediate video is obtained, and based on the current state of the buffer area and the bitrate of the first intermediate video, intermediate buffer data is determined. If the intermediate buffer data does not meet the preset buffer requirements, encoding optimization is performed based on the scene type of the first intermediate video, the second intermediate video is obtained, and based on the size of the second intermediate video and the buffer area, the second intermediate video is transmitted. The present invention determines whether to perform two video encodings on the target video based on the buffer area to reduce the data volume of the target video, dynamically determines the encoding method based on the buffer area, balances the target video bitrate and the buffer, and improves the resource utilization efficiency.

[0048] Specifically, as Figure 2 shown, in S100, the buffer area size is determined based on the packet loss rate q, the round-trip time RTT, and the BBR congestion control algorithm of the transmission channel between the client and the server for the target video, specifically including:

[0049] S110, obtain the round-trip time RTT and the packet loss rate q at which the client of the target video sends target data packets to the server at a preset period.

[0050] Specifically, RTprop = min (several RTTs of the preset period). RTprop is the minimum round-trip time, and min() is the function to take the minimum value.

[0051] S120, obtain the preset gain coefficient corresponding to the packet loss rate q. Specifically, the corresponding relationship between the packet loss rate and the preset gain coefficient is set in advance, so that after obtaining the packet loss rate q in the current network state, the preset gain coefficient corresponding to q can be found. Further, it also includes: determining the corresponding relationship between the packet loss rate and the preset gain coefficient through the packet loss rate and the current stage of BBR.

[0052] Specifically, q and the gain coefficient are inversely proportional. When q increases, the gain coefficient is gradually reduced.

[0053] S130, based on the preset gain coefficient corresponding to the packet loss rate q, use the BBR congestion control algorithm to obtain the congestion window size, and determine the buffer area size based on the congestion window size.

[0054] Specifically, the congestion window cwnd = BtlBW × RTprop × G(q), where G(q) is the gain coefficient corresponding to q, and BtlBW is the minimum bottleneck bandwidth.

[0055] In an embodiment of the present invention, the buffer area D = cwnd + cwnd × q / (1 - q).

[0056] In summary, by dynamically associating the packet loss rate q and combining with the bandwidth, the buffer area is dynamically determined for more effective transmission.

[0057] Specifically, as Figure 3 shown, based on the current state of the buffer area in S300, adjusting the attribute information of the target video further includes:

[0058] S310, determining a specified bitrate based on the initial buffer time and the current bitrate. The bitrate ratio is directly proportional to the time ratio. The bitrate ratio is the ratio of the specified bitrate to the current bitrate, and the time ratio is the ratio of the preset buffer time threshold to the initial buffer time.

[0059] Specifically, the specified bitrate Rtar = α × Rcur × (1 - a), where Rcur is the current bitrate, a is the bitrate adjustment factor, and α is the weight adjustment factor. In an embodiment of the present invention, a = (initial buffer time - preset buffer time threshold) / initial buffer time. When the difference between the initial buffer time and the preset buffer time threshold is too large, a will increase for emergency downgrading.

[0060] S320, determining a new quantization step size, a new frame rate, and a new resolution based on the specified bitrate.

[0061] In an embodiment of the present invention, the new quantization step size is the sum of the initial quantization step size and the adjusted quantization step size. The adjusted quantization step size is the product of the adjustment coefficient and the preset maximum increment. The adjustment coefficient is the quotient of the bitrate difference and the current bitrate. The bitrate difference is the difference between the current bitrate and the specified bitrate. The new quantization step size QP1 = QP0 + △QP × β QP , where QP0 is the initial quantization step size, △QP = (Rcur - Rtar) / Rcur × QPmax, QPmax is the preset maximum allowable increment, and β QP is the weight coefficient of the quantization step size, where β QP + β F + β R = 1. The new frame rate F1 = F0 × (1 - (Rcur - Rtar) / Rcur × β F ), β F is the frame rate weight coefficient, and F0 is the initial frame rate; the new resolution width W1 = W0 × (1 - (Rcur - Rtar) / Rcur × β R ), 1 / 2 the new resolution length H1 = H0 × (1 - (Rcur - Rtar) / Rcur × β R ), 1 / 2 , where W0 is the initial resolution width, H0 is the initial resolution length, and βR is the resolution weight coefficient. Among them, β is determined based on the scene type of the first intermediate video QP , β F , β R .

[0062] In another embodiment of the present invention, as Figure 4 shown, S320 determines a new quantization step size, a new frame rate, and a new resolution based on a specified bit rate, and further includes:

[0063] S321, constructing a relationship model for determining the quantization step size, frame rate, and resolution by the bit rate. Among them, in the relationship model, the bit rate and the quantization step size are in an inverse relationship, the bit rate and the frame rate are in a direct relationship, and the bit rate and the resolution are in a direct relationship.

[0064] In one embodiment of the present invention, the function of the relationship model is: video bit rate R = k × W × H × F / QP c , where W is the number of horizontal pixels of the resolution, H is the number of vertical pixels of the resolution, F is the frame rate, QP is the quantization step size, k is the encoding coefficient, and c is the quantization step size sensitivity.

[0065] S322, obtaining a training data list, where the training data list includes several training data, and the training data includes an actual training bit rate, a training quantization step size, a training bit rate, and a training resolution.

[0066] S323, based on the training data list, using the gradient descent method to update the parameters of the relationship model to obtain an updated relationship model.

[0067] S324, based on the specified bit rate and the updated relationship model, determining a new quantization step size, a new frame rate, and a new resolution.

[0068] In summary, constructing a relationship model for determining the quantization step size, frame rate, and resolution by the bit rate, obtaining a training data list, where the training data list includes several training data, based on the training data list, using the gradient descent method to update the parameters of the relationship model to obtain an updated relationship model, based on the specified bit rate and the updated relationship model, determining a new quantization step size, a new frame rate, and a new resolution, and realizing automatic parameter tuning by constructing a relationship model.

[0069] In one embodiment of the present invention, as Figure 5 shown, S500 performs encoding optimization based on the scene type of the first intermediate video to obtain a second intermediate video, and further includes:

[0070] S510, extracting key frames from the video content of the first intermediate video to obtain several key frames. Specifically, those skilled in the art know that any method for extracting key frames in the prior art belongs to the protection scope of the present invention and will not be elaborated here.

[0071] S520, construct a scene binary classification model, input the key frames into the scene binary classification model, and determine the classification result of each key frame. In an embodiment of the present invention, a scene binary classification model is constructed, the key frames are input into the scene binary classification model, and it is determined whether the key frames are static scenes or dynamic scenes.

[0072] S530, based on the classification result of each key frame, determine to use the intra-frame prediction method for encoding optimization to obtain a second intermediate video.

[0073] Specifically, based on the classification result of the key frames, determine to use the intra-frame prediction method for encoding optimization, or determine to use the inter-frame prediction method for encoding optimization; further, for the video frames of the first intermediate video other than the key frames, methods such as time interpolation and motion propagation are used to determine the parameters of encoding optimization; furthermore, when using the intra-frame encoding method for encoding optimization of the key frames, for the video frames of the first intermediate video other than the key frames, the time difference method is used to determine the parameters of encoding optimization; when using the inter-frame encoding method for encoding optimization of the key frames, for the video frames of the first intermediate video other than the key frames, the motion propagation method is used to determine the parameters of encoding optimization.

[0074] In summary, the present invention realizes the re-compression of the first intermediate video through the scene binary classification model and a simple decision of motion-static binary classification, and avoids causing resource usage pressure on the client.

[0075] In another embodiment of the present invention, as Figure 6 shown, S500 performs encoding optimization based on the scene type of the first intermediate video to obtain a second intermediate video, and further includes:

[0076] S501, extract x initial key frames from the first intermediate video, where x is less than a preset frame threshold.

[0077] S502, input the initial key frames into the constructed scene multi-classification model to obtain initial scene labels.

[0078] S503, based on the initial scene labels, determine a key frame extraction method, and use the key frame extraction method to extract key frames from the first intermediate video again to obtain intermediate key frames, and the number of intermediate key frames is greater than the initial key frames.

[0079] Specifically, first determine the scene labels, and based on the scene labels, determine the key frame extraction method, so as to extract key frames more accurately; further, it also includes determining the extraction quantity of key frames based on the scene labels.

[0080] In an embodiment of the present invention, the above steps are implemented through the corresponding relationship between preset scene labels and key frame extraction methods.

[0081] S504, input the intermediate key frame into the constructed scene multi-classification model to obtain the intermediate scene label.

[0082] S505, input the intermediate key frame into the constructed motion detection model to obtain the motion intensity score.

[0083] S506, input the intermediate key frame into the constructed texture feature extraction model to obtain the texture score.

[0084] In an embodiment of the present invention, the constructed texture feature extraction model is Local Binary Pattern (LBP). LBP determines the texture score by comparing the gray value relationship between each pixel in the image and its neighboring pixels, and uses a lightweight model to assist in determination.

[0085] S507, generate encoding parameters for encoding optimization based on the intermediate scene label, motion intensity score, and texture score to obtain the second intermediate video.

[0086] In summary, extract x initial key frames from the first intermediate video, input the initial key frames into the constructed scene multi-classification model to obtain the initial scene label, determine the key frame extraction method based on the initial scene label, and use the key frame extraction method to extract key frames from the first intermediate video again to obtain the intermediate key frames. Input the intermediate key frames into the constructed scene multi-classification model to obtain the intermediate scene label, input the intermediate key frames into the constructed motion detection model to obtain the motion intensity score, input the intermediate key frames into the constructed texture feature extraction model to obtain the texture score, generate encoding parameters for encoding optimization based on the intermediate scene label, motion intensity score, and texture score to obtain the second intermediate video; by jointly using multiple models, the encoding efficiency is improved.

[0087] Embodiment 2

[0088] Embodiment 2 of the present invention provides a video encoding device, as Figure 7 shown, the device includes:

[0089] A buffer area construction module for constructing a buffer area. Among them, a buffer area is constructed on the client, where the client is used to send the target video, and the size of the buffer area is determined based on the packet loss rate q, round-trip time RTT, and BBR congestion control algorithm of the transmission channel between the client and the server;

[0090] An initial buffer data module for determining the initial buffer data based on the current bit rate of the obtained target video and the current state of the buffer area. The initial buffer data includes: the initial buffer time, which is the caching time of the target video in the buffer area at the current bit rate.

[0091] An adjustment module, configured to, if the initial buffer data does not meet the preset buffer requirements, adjust the attribute information of the target video based on the current state of the buffer area to obtain a first intermediate video corresponding to the target video, where the attribute information includes resolution, frame rate, and quantization step size;

[0092] An intermediate buffer data module, configured to determine intermediate buffer data based on the intermediate bit rate of the obtained first intermediate video and the current state of the buffer area, where the intermediate buffer data includes: an intermediate buffer time, and the intermediate buffer time is the buffer time of the first intermediate video in the buffer area at the intermediate bit rate;

[0093] An optimization module, configured to, if the intermediate buffer data does not meet the preset buffer requirements, perform encoding optimization based on the scene type of the first intermediate video, obtain a second intermediate video, and transmit the second intermediate video.

[0094] Among them, as Figure 8 shown, the buffer area construction module further includes:

[0095] An RTT acquisition sub-module, configured to acquire the round-trip time RTT and packet loss rate q of the client of the target video sending target data packets to the server at a preset period.

[0096] A gain coefficient acquisition sub-module, configured to acquire a preset gain coefficient corresponding to the packet loss rate q.

[0097] A buffer area size acquisition sub-module, configured to, based on the preset gain coefficient corresponding to the packet loss rate q, use the BBR congestion control algorithm to acquire the congestion window size, and determine the buffer area size based on the congestion window size.

[0098] Among them, as Figure 9 shown, the adjustment module further includes:

[0099] A specified bit rate acquisition sub-module, configured to determine a specified bit rate based on the initial buffer time and the current bit rate, where the bit rate ratio is directly proportional to the time ratio, the bit rate ratio is the ratio of the specified bit rate to the current bit rate, and the time ratio is the ratio of the preset buffer time threshold to the initial buffer time.

[0100] A new attribute information acquisition sub-module, configured to determine a new quantization step size, a new frame rate, and a new resolution based on the specified bit rate.

[0101] Among them, as Figure 10 shown, the new attribute information acquisition sub-module further includes:

[0102] A relationship model acquisition unit, configured to construct a relationship model for the bit rate to determine the quantization step size, the frame rate, and the resolution. Among them, in the relationship model, the bit rate and the quantization step size are in an inverse relationship, the bit rate and the frame rate are in a direct relationship, and the bit rate and the resolution are in a direct relationship.

[0103] A training data acquisition unit for acquiring a training data list, where the training data list includes a number of training data, and the training data includes an actual training bitrate, a training quantization step size, a training bitrate, and a training resolution.

[0104] A parameter update unit for updating the parameters of the relationship model using the gradient descent method based on the training data list to obtain an updated relationship model.

[0105] A new attribute information acquisition unit for determining a new quantization step size, a new frame rate, and a new resolution based on a specified bitrate and the updated relationship model.

[0106] Among them, as Figure 11 shown, the optimization module further includes:

[0107] A key frame acquisition sub-module for extracting key frames from the video content of the first intermediate video to obtain a number of key frames.

[0108] A scene binary classification model acquisition sub-module for constructing a scene binary classification model and inputting the key frames into the scene binary classification model to determine the classification result of each key frame.

[0109] A second intermediate video acquisition sub-module for determining to perform encoding optimization using the intra-frame prediction method based on the classification result of each key frame to obtain a second intermediate video.

[0110] Among them, as Figure 12 shown, the optimization module further includes:

[0111] An initial key frame acquisition sub-module for extracting x initial key frames from the first intermediate video, where x is less than a preset frame threshold.

[0112] A scene multi-classification model sub-module for inputting the initial key frames into the constructed scene multi-classification model to obtain initial scene labels.

[0113] An intermediate key frame acquisition sub-module for determining a key frame extraction method based on the initial scene labels and using the key frame extraction method to extract key frames from the first intermediate video again to obtain intermediate key frames, and the number of intermediate key frames is greater than the initial key frames.

[0114] An intermediate scene label acquisition sub-module for inputting the intermediate key frames into the constructed scene multi-classification model to obtain intermediate scene labels.

[0115] A motion intensity score acquisition sub-module for inputting the intermediate key frames into the constructed motion detection model to obtain motion intensity scores.

[0116] A texture score acquisition sub-module for inputting the intermediate key frames into the constructed texture feature extraction model to obtain texture scores.

[0117] A final sub-module, configured to generate encoding parameters for encoding optimization based on intermediate scene tags, exercise intensity scores, and texture scores, and obtain a second intermediate video.

[0118] Embodiment III

[0119] An embodiment of the present invention further provides an electronic device, including: a processor, a memory, and a computer program stored on the memory and executable on the processor, where the processor implements the method provided in the above embodiment when executing the computer program.

[0120] Although some specific embodiments of the present invention have been described in detail by way of examples, those skilled in the art should understand that the above examples are for illustrative purposes only and not for limiting the scope of the present invention. Those skilled in the art should also understand that various modifications can be made to the embodiments without departing from the scope and spirit of the present invention.

Claims

1. A video encoding method, characterized in that: The method comprises the following steps: A buffer area is constructed on the client, where the client is used to send the target video, and the size of the buffer area is determined based on the packet loss rate q of the transmission channel of the target video between the client and the server, the round-trip time RTT, and the BBR congestion control algorithm; Based on the current bit rate of the acquired target video and the current state of the buffer area, initial buffer data is determined, where the initial buffer data includes: initial buffer time, which is the cache time of the target video in the buffer area at the current bit rate; If the initial buffer data does not meet the preset buffer requirement, based on the current state of the buffer area, adjusting the attribute information of the target video to obtain a first intermediate video corresponding to the target video, the attribute information including resolution, frame rate, and quantization step size; Based on the obtained intermediate bit rate of the first intermediate video and the current state of the buffer area, determine the intermediate buffer data, the intermediate buffer data including: intermediate buffer time, the intermediate buffer time is the buffer time of the first intermediate video in the buffer area at the intermediate bit rate; If the intermediate buffer data does not meet the preset buffer requirement, encoding optimization is performed based on the scene type of the first intermediate video, a second intermediate video is obtained and the second intermediate video is transmitted; Performing encoding optimization based on the scene type of the first intermediate video to obtain the second intermediate video also includes: Extract key frames from the video content of the first intermediate video to obtain a plurality of key frames; Construct a scene binary classification model, input the key frames into the scene binary classification model, and determine the classification result of each key frame; Based on the classification result of each key frame, it is determined to use an intra-frame prediction method for encoding optimization to obtain a second intermediate video.

2. The video encoding method according to claim 1, characterized in that The size of the buffer area is determined based on the packet loss rate q of the target video transmission channel between the client and the server, the round-trip time RTT, and the BBR congestion control algorithm, including: The client that obtains the target video sends the target data packet to the server according to the preset period. The round-trip time RTT and packet loss rate q; Obtain a preset gain coefficient corresponding to the packet loss rate q; Based on the preset gain coefficient corresponding to the packet loss rate q, the BBR congestion control algorithm is used to obtain the congestion window size, and the buffer area size is determined based on the congestion window size.

3. The video encoding method according to claim 1, characterized in that: Based on the current state of the buffer area, the attribute information of the target video is adjusted, including: Based on the initial buffer time and the current bit rate, the specified bit rate is determined, the bit rate ratio and the time ratio are in direct proportion, the bit rate ratio is the ratio of the specified bit rate to the current bit rate, and the time ratio is the ratio of the preset buffer time threshold to the initial buffer time; Determine a new quantization step size, a new frame rate, and a new resolution based on the specified bit rate.

4. The video encoding method according to claim 3, characterized in that: Determining a new quantization step size, a new frame rate, and a new resolution based on a specified bit rate also includes: Constructing a relationship model in which the bit rate determines the quantization step size, the frame rate, and the resolution, wherein in the relationship model, the bit rate and the quantization step size are inversely proportional, the bit rate and the frame rate are in direct proportion, and the bit rate and the resolution are in direct proportion; Obtain a training data list, wherein the training data list includes a plurality of training data, and the training data includes an actual training bit rate, a training quantization step size, a training bit rate, and a training resolution; Based on the training data list, the parameters of the relational model are updated using the gradient descent method to obtain an updated relational model; Based on the specified bit rate and the updated relationship model, a new quantization step size, a new frame rate, and a new resolution are determined.

5. The video encoding method according to claim 3, characterized in that: The new quantization step size is the sum of the initial quantization step size and the adjusted quantization step size. The adjusted quantization step size is the product of the adjustment coefficient and the preset maximum increment. The adjustment coefficient is the quotient of the bit rate difference and the current bit rate. The bit rate difference is the difference between the current bit rate and the specified bit rate.

6. The video encoding method according to claim 1, characterized in that: Performing encoding optimization based on the scene type of the first intermediate video to obtain the second intermediate video also includes: Extracting x initial key frames from the first intermediate video, where x is less than a preset frame threshold; Input the initial key frame into the constructed scene multi-classification model to obtain the initial scene label; Based on the initial scene label, determine a key frame extraction method, and use the key frame extraction method to extract key frames from the first intermediate video again to obtain intermediate key frames, where the number of intermediate key frames is greater than that of initial key frames; Input the intermediate key frames into the constructed scene multi-classification model to obtain the intermediate scene labels; The intermediate key frames are input into the constructed motion detection model to obtain the motion intensity score; Input the intermediate key frame into the constructed texture feature extraction model to obtain the texture score; Based on the intermediate scene label, the motion intensity score and the texture score, encoding parameters are generated to perform encoding optimization and obtain a second intermediate video.

7. The video encoding method according to claim 6, characterized in that: The constructed texture feature extraction model is the local binary pattern LBP.

8. A video encoding device, characterized in that: The device comprises: A buffer area construction module is used to construct a buffer area, wherein the buffer area is constructed on a client, wherein the client is used to send a target video, and the size of the buffer area is determined based on a packet loss rate q of a transmission channel of the target video between the client and the server, a round-trip time RTT, and a BBR congestion control algorithm; An initial buffering data module is used to determine initial buffering data based on the current bit rate of the acquired target video and the current state of the buffering area, wherein the initial buffering data includes: an initial buffering time, which is the cache time of the target video in the buffering area at the current bit rate; An adjustment module, configured to adjust the attribute information of the target video based on the current state of the buffer area to obtain a first intermediate video corresponding to the target video if the initial buffer data does not meet the preset buffer requirement, wherein the attribute information includes resolution, frame rate, and quantization step size; An intermediate buffering data module is used to determine intermediate buffering data based on the intermediate bit rate of the first intermediate video and the current state of the buffer area, wherein the intermediate buffering data includes: intermediate buffering time, which is the buffering time of the first intermediate video in the buffer area at the intermediate bit rate; An optimization module, configured to perform encoding optimization based on the scene type of the first intermediate video, obtain the second intermediate video and transmit the second intermediate video if the intermediate buffer data does not meet the preset buffer requirement; The optimization module also includes: A key frame acquisition submodule is used to extract key frames from the video content of the first intermediate video to obtain a plurality of key frames; The scene binary classification model acquisition submodule is used to construct a scene binary classification model, input key frames into the scene binary classification model, and determine the classification result of each key frame; The second intermediate video acquisition submodule is used to determine to use the intra-frame prediction method for encoding optimization based on the classification result of each key frame to acquire the second intermediate video.

9. An electronic device, comprising: A processor, a memory, and a computer program stored in the memory and executable on the processor, wherein the processor implements the video encoding method according to any one of claims 1 to 7 when executing the computer program.

Citation Information

Patent Citations

  • Short video code rate adaptive transmission method based on multi-agent reinforcement learning

    CN116506626A

  • Streaming media video coding method and device, equipment, storage medium and product

    CN119545111A