Video encoding method, image sending device, and storage medium
By dynamically adjusting the insertion time interval and encoding strategy of long-term reference frames, the video quality decline caused by the fixed long-term reference frame strategy when the network is poor is solved, and video quality assurance in different network environments is achieved.
Patent Information
- Application Number
- CN202111586374.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-22
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2041-12-22
AI Technical Summary
In the prior art, fixed long-term reference frame strategy is prone to cause the receiver to lose LTR frames when the network is poor, and it is impossible to effectively ensure the transmission and image quality of video.
By obtaining the average packet loss rate and round trip delay of the video image, dynamically adjust the insertion time interval and encoding strategy of long-term reference frames, insert LTR frames in time, and adjust the insertion time interval and encoding strategy of reference frames in real time according to network conditions to ensure video quality.
In the case of poor network, reduce screen discoloration, improve the smoothness of the video picture, and ensure the quality of the video.
Smart Images

Figure CN114302142B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image transmission, and particularly to a video encoding method, an image sending device, and a storage medium. Background Art
[0002] In the related art, a fixed long-term reference frame strategy is mainly adopted, that is, a long-term reference frame LTR (Long Term Reference) is inserted at a fixed period. Or a fixed number of frames after the LTR frame all refer to the LTR frame.
[0003] However, when the network is poor, there is a situation where the receiving end loses the LTR frame, thus unable to effectively ensure video transmission and image quality. Summary of the Invention
[0004] The main object of the present invention is to provide a video encoding method, an image sending device, and a storage medium, aiming to solve the problem that when the receiving end loses the LTR frame in the prior art, the video transmission and image quality cannot be effectively guaranteed.
[0005] To achieve the above object, in a first aspect, the present invention provides a video encoding method, and the method includes:
[0006] Obtain the average packet loss rate of the video image within the current detection time period;
[0007] Determine a long-term reference frame according to the average packet loss rate, and obtain the real-time average packet loss rate at each moment and the current round-trip delay of the current round-trip delay period at each moment;
[0008] Adjust the reference frame insertion time interval according to the real-time average packet loss rate;
[0009] Determine a target long-term reference frame according to the current round-trip delay;
[0010] Encode each P frame of the video image by using the target long-term reference frame and the reference frame insertion time interval.
[0011] In an embodiment, the determining a target long-term reference frame according to the current round-trip delay includes:
[0012] Judge whether the current round-trip delay is greater than a preset delay threshold;
[0013] If the current round-trip delay is greater than the preset delay threshold, determine the long-term reference frame as the target long-term reference frame;
[0014] If the current round-trip delay is less than the preset delay threshold, determine the target long-term reference frame according to the feedback information sent by the receiving end.
[0015] In one embodiment, if the current round-trip delay is less than the preset delay threshold, determining a target long-term reference frame according to the feedback information sent by the receiving end includes:
[0016] If the current round-trip delay is less than the preset delay threshold, determine whether feedback information indicating that the receiving end has received the historical long-term reference frame is received;
[0017] If the feedback information is received, determine the historical long-term reference frame as the target long-term reference frame;
[0018] If the feedback information is not received, filter out the target long-term reference frame from each P frame according to the reference frame insertion time interval.
[0019] In one embodiment, before obtaining the average packet loss rate of the video image in the current detection time period, the method further includes:
[0020] Obtain the packet loss rate at the previous moment of the current moment and the duration of the previous detection time period;
[0021] Determine the duration of the current detection time period according to the packet loss rate and the duration of the previous detection time period.
[0022] In one embodiment, determining the duration of the current detection time period according to the packet loss rate and the duration of the previous detection time period includes:
[0023] Determine the ratio of the packet loss rate to the preset packet loss rate threshold;
[0024] If the ratio is greater than or equal to the preset reference value, determine the duration of the current detection time period according to the duration of the previous detection time period and the first preset formula: The first preset formula is:
[0025] m i = m i-1 - adjustmentFact 2 ;
[0026] If the ratio is less than the preset reference value, determine the duration of the current detection time period according to the duration of the previous detection time period and the second preset formula: The second preset formula is:
[0027] m i = m i-1 + 1;
[0028] Wherein, m i is the duration of the current detection time period, m i-1 is the duration of the previous detection time period, and adjustmentFact is the ratio.
[0029] In one embodiment, determining the duration of the current detection period according to the packet loss rate and the duration of the previous detection period includes:
[0030] Obtaining a duration to be determined according to the packet loss rate and the duration of the previous detection period;
[0031] If the duration to be determined is greater than or equal to a preset maximum detection duration, then using the preset maximum detection duration as the duration of the current detection period;
[0032] If the duration to be determined is less than the preset maximum detection duration and greater than the preset minimum detection duration, then using the duration to be determined as the duration of the current detection period;
[0033] If the duration to be determined is less than or equal to the preset minimum detection duration, then using the preset minimum detection duration as the duration of the current detection period.
[0034] In one embodiment, adjusting the reference frame insertion time interval according to the real-time average packet loss rate includes:
[0035] Obtaining a reference frame insertion time interval according to the real-time average packet loss rate, the preset packet loss rate threshold, and a first preset formula; the first preset formula is:
[0036]
[0037] wherein, LTRFramePEriod is the reference frame insertion time interval, averageLossRate is the real-time average packet loss rate, n is the preset packet loss rate threshold, and t is a natural number.
[0038] In one embodiment, after adjusting the reference frame insertion time interval according to the real-time average packet loss rate, the method further includes:
[0039] If the reference frame insertion time interval is greater than a preset maximum insertion time interval, then using the preset maximum insertion time interval as the reference frame insertion time interval.
[0040] In a second aspect, the present invention further provides an image sending device, including: a memory, a processor, and a video encoding program stored on the memory and executable on the processor, where the video encoding program is configured to implement the steps of the video encoding method as described above.
[0041] In a third aspect, the present invention further provides a computer-readable storage medium, on which a video encoding program is stored, and when the video encoding program is executed by a processor, it implements the steps of the video encoding method as described above.
[0042] An embodiment of the present invention provides a video encoding method, an image sending device, and a storage medium. The video encoding method determines a long-term reference frame for encoding based on the average packet loss rate within the current detection time period to instantaneously insert the long-term reference frame, avoiding the situation of reduced image quality such as screen freeze caused by the loss of the long-term reference frame. Moreover, the target long-term reference frame is determined according to the current round-trip delay RTT, and the reference frame insertion time interval is adjusted in real time according to the real-time average packet loss rate at subsequent moments and the current round-trip delay within the current round-trip delay period, so that the insertion time interval of subsequent target long-term reference frames is more in line with the actual situation of network transmission, and different encoding strategies are adopted under different RTTs, thereby ensuring the video quality, reducing the screen freeze time, and making the video picture smoother in the case of a poor network. BRIEF DESCRIPTION OF THE DRAWINGS
[0043] Figure 1 It is a schematic diagram of the modules of the image sending device of the present invention;
[0044] Figure 2 It is a schematic flowchart of the first embodiment of the video encoding method of the present invention;
[0045] Figure 3 It is a schematic flowchart of the second embodiment of the video encoding method of the present invention;
[0046] Figure 4 It is a schematic flowchart of the third embodiment of the video encoding method of the present invention;
[0047] Figure 5 It is a schematic flowchart of the fourth embodiment of the video encoding method of the present invention.
[0048] The realization, functional features, and advantages of the object of the present invention will be further described in conjunction with the embodiments with reference to the drawings. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0049] It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0050] In the existing related technologies, a fixed long-term reference frame strategy is mainly adopted, that is, a long-term reference frame LTR (Long Term Reference) is inserted at a fixed period. Or the fixed n frames after the LTR frame all refer to the LTR frame. However, when the network is poor, there is a situation where the receiving end will lose the LTR frame, thus unable to effectively guarantee the transmission of the video and the image quality.
[0051] To this end, the present application provides a video encoding method. When the network is poor, an LTR is inserted immediately, and in the subsequent process, the insertion time interval of the long-term reference frame is dynamically adjusted according to the packet loss rate. At the same time, the encoding strategy of the encoded frame is determined according to the round-trip delay, overcoming the deficiency of inserting the long-term reference frame at a fixed period, enabling the video encoding to better adapt to the network transmission in strong network and weak network environments, and providing better video quality.
[0052] Refer to Figure 1 , Figure 1 which is a schematic structural diagram of an image sending device for the hardware operating environment involved in the solution of the embodiment of the present application.
[0053] As Figure 1 shown, the image sending device may include: a processor 1001, such as a central processing unit (CPU), a communication bus 1002, a user interface 1003, a network interface 1004, and a memory 1005. Among them, the communication bus 1002 is used to realize the connection and communication between these components. The user interface 1003 may include a display screen (Display) and an input unit such as a keyboard (Keyboard). Optionally, the user interface 1003 may further include a standard wired interface and a wireless interface. The network interface 1004 may optionally include a standard wired interface and a wireless interface (such as a wireless fidelity (WI-FI) interface). The memory 1005 may be a high-speed random access memory (Random Access Memory, RAM) or a stable non-volatile memory (Non-Volatile Memory, NVM), such as a disk memory. Optionally, the memory 1005 may also be a storage device independent of the foregoing processor 1001.
[0054] Those skilled in the art can understand that Figure 1 the structure shown in
[0055] As Figure 1 shown, the memory 1005, as a storage medium, may include an operating system, a data storage module, a Bluetooth communication module, a user interface module, and a video encoding program.
[0056] In Figure 1In the image sending device shown, the network interface 1004 is mainly used for data communication with a network server; the user interface 1003 is mainly used for data interaction with a user; the processor 1001 and the memory 1005 in the image sending device of the present invention can be arranged in the image sending device. The image sending device calls the video encoding program stored in the memory 1005 through the processor 1001 and executes the video encoding method provided by the embodiments of the present application.
[0057] Based on the above hardware devices but not limited to the above hardware devices, a first embodiment of a video encoding method of the present application is proposed. Refer to Figure 2 , Figure 2 which is a schematic flowchart of the first embodiment of the video encoding method of the present application.
[0058] In this embodiment, the method includes:
[0059] Step S101, obtain the average packet loss rate of the video image during the current detection time period;
[0060] In this embodiment, the execution subject of the video encoding method is an image sending device, which is used to encode the video image data in the buffer and send it to the receiving end, and the receiving end decodes it after receiving it. It can be understood that the image sending device can be a video server, a cloud game server, or a live broadcast server, etc., and the receiving end can be a mobile terminal such as a mobile phone or a tablet, or a terminal device such as a computer.
[0061] During the encoding and transmission process of the image sending device, the average packet loss rate during the current detection time period can be obtained. Among them, the current detection time period is also a time period formed by taking the moment with a duration from the current moment as the start moment and the current moment as the end moment. The average packet loss rate during the current detection time period is used to evaluate the network transmission quality of the time period where the current moment is located. Among them, the duration of the current detection time period can be expressed as m i , in seconds, so that the network transmission situation at the current moment can be evaluated according to the average packet loss rate within the previous m i seconds. If the average packet loss rate is high, the network at the current moment is poor. If the average packet loss rate is low, the network at the current moment is good.
[0062] Step S102, determine a long-term reference frame according to the average packet loss rate, and obtain the real-time average packet loss rate at each moment and the current round-trip delay of each moment in the current round-trip delay period.
[0063] After obtaining the average packet loss rate, the current actual network condition can be determined according to the average packet loss rate. For example, after obtaining the average packet loss rate, the average packet loss rate can be compared with a preset packet loss rate threshold. If the average packet loss rate is greater than or equal to the preset packet loss rate threshold, then the network condition suddenly fluctuates at this time, and the network condition is poor, and the situation of losing long-term reference frames may occur. If the average packet loss rate is lower than the preset packet loss rate threshold, then the network condition is good at this time.
[0064] Specifically, when the average packet loss rate is greater than or equal to the preset packet loss rate threshold, a long-term reference frame LTR can be immediately inserted to avoid the situation of reduced image quality such as screen freeze caused by the loss of long-term reference frames. When the average packet loss rate is lower than the preset packet loss rate threshold, encoding can continue according to the preset LTR insertion strategy without change.
[0065] In order to avoid the algorithm being relatively complex due to frequent insertion of long-term reference frames LTR when the network is poor, make the bitrate allocation more reasonable, suitable for network transmission, and ensure image quality. In a specific embodiment, if the average packet loss rate is greater than or equal to the preset packet loss rate threshold for the first time, the current frame can be used as the long-term reference frame, and the current moment can be used as the cycle start moment to enter the long-term reference frame insertion cycle.
[0066] After obtaining the average packet loss rate, the average packet loss rate can be compared with the preset packet loss rate threshold. If the average packet loss rate is greater than or equal to the preset packet loss rate threshold for the first time, then the network condition suddenly fluctuates at this time, and the network condition is poor, and the situation of losing long-term reference frames may occur. A long-term reference frame LTR can be immediately inserted to avoid the situation of reduced image quality such as screen freeze caused by the loss of long-term reference frames.
[0067] It can be understood that when the sending end starts to encode and send the video image data stream, the sending end inserts LTR regularly or dynamically based on certain preset rules. At this time, if the network condition at the current moment is detected to be poor, the current moment can be used as the cycle start moment to enter a long-term reference frame insertion cycle. That is to say, after entering the long-term reference frame insertion cycle, a new reference frame insertion time interval or insertion strategy can be adopted to adapt to the network transmission in the weak network environment at this time.
[0068] During the long-term reference frame insertion cycle, the real-time average packet loss rate at each moment and the current round-trip delay of the current round-trip delay period at each moment are obtained.
[0069] Specifically, during the encoding and transmission process, the image sending device obtains the current round-trip delay of the current round-trip delay period at each moment.
[0070] Specifically, the sender constructs and sends network probe packets to the receiver. The network probe packets carry the sending timestamp begin_ntp_time. After the receiver receives the latest network probe packet, it records the current timestamp last_recv_ntp_time. The receiver b constructs a network probe packet reply packet and sets the probe packet reply packet timestamp processTime = the current ntp timestamp - last_recv_ntp_time. After the sender receives the network probe packet reply packet from the receiver, it records the arrival time cur_ntp_time of the network probe packet reply packet.
[0071] Thus, the calculation method of the current round-trip delay RTT is as follows:
[0072] RTT = cur_ntp_time - begin_ntp_time - processTime.
[0073] For the current moment, the specific value of the current round-trip delay RTT in the current round-trip delay period can be queried.
[0074] Step S103: Adjust the reference frame insertion time interval according to the real-time average packet loss rate;
[0075] That is, in this step, within the long-term reference frame insertion period, the reference frame insertion time interval is dynamically and adaptively adjusted according to the real-time average packet loss rate, overcoming the deficiency of the fixed long-term reference frame, enabling video coding to better adapt to network transmission in strong network and weak network environments, and providing better video quality.
[0076] Among them, the reference frame insertion time interval is adaptively obtained according to the average packet loss rate and the preset packet loss rate threshold, and is adjusted according to the real-time average packet loss rate in the subsequent process, so as to ensure video quality, reduce the time of frozen frames, and make the video picture smoother.
[0077] Step S104: Determine the target long-term reference frame according to the current round-trip delay.
[0078] Specifically, within the long-term reference frame insertion period, the target long-term reference frame is adjusted according to the current round-trip delay and the preset delay threshold, so as to determine a suitable coding frame strategy according to the specific current round-trip delay situation, set an appropriate target long-term reference frame, and avoid the decoding process being aggravated by frozen frames or even increasing the possibility of network congestion caused by inappropriate target long-term reference frames.
[0079] Step S105: Encode each P frame of the video image by using the target long-term reference frame and the reference frame insertion time interval.
[0080] After determining the target long-term reference frame and the reference frame insertion time interval according to the network transmission situation, that is, after determining the encoding strategy at any moment within the long-term reference frame insertion period, the P frames of the video image can be encoded by using the target long-term reference frame and the reference frame insertion time interval to obtain the code stream data to be sent. The sending end sends the sent code stream data to the receiving end. After receiving the code stream data to be sent, the receiving end decodes it to obtain the video image data.
[0081] And it can be understood that during the video transmission process, the sending end always monitors the network transmission situation in real time. At each moment, it continuously obtains the real-time average packet loss rate at the current moment and the current round-trip delay in the current round-trip delay period where the current moment is located. In this embodiment, the reference frame insertion time interval and the encoding strategy are continuously adjusted according to the real-time average packet loss rate and the current round-trip delay, that is, according to the actual network state, so as to balance the image quality and network transmission in the case of poor network, ensure the video quality, reduce the time of screen freeze, and make the video picture smoother.
[0082] In addition, in this embodiment, when it is detected that a preset condition is met, the preset long-term reference frame insertion period can be ended, that is, return to the normal preset rule to insert LTR regularly or dynamically. Specifically, the end condition of the preset long-term reference frame insertion period can be the completion of the video image transmission, or the packet loss rate of each frame within several frames is greater than the preset packet loss rate threshold, that is, the network transmission quality becomes better. For example, the packet loss rate of each frame within the current detection time period is greater than the preset packet loss rate threshold. This embodiment does not limit this.
[0083] In this embodiment, the video encoding method determines the network state based on the average packet loss rate, thereby immediately inserting the long-term reference frame to avoid the situation of reduced image quality such as screen freeze caused by the loss of the long-term reference frame, and adjusts the reference frame insertion time interval according to the average packet loss rate at each subsequent moment in real time, so that the insertion time interval of the long-term reference frame more conforms to the actual situation of network transmission, and also determines the target long-term reference frame according to the current round-trip delay RTT, so as to adopt different encoding strategies under different RTTs, thereby ensuring the video quality in the case of poor network, reducing the time of screen freeze, and making the video picture smoother.
[0084] Based on the above embodiment, the second embodiment of the video encoding method of the present application is proposed. Refer to Figure 3 , Figure 3 It is the flowchart of the second embodiment of the video encoding method of the present application.
[0085] In this embodiment, step S104 includes:
[0086] Step A10, determine whether the current round-trip delay is greater than the preset delay threshold;
[0087] Step A20: If the current round-trip delay is greater than the preset delay threshold, then determine the current long-term reference frame as the target long-term reference frame;
[0088] Step A30: If the current round-trip delay is less than the preset delay threshold, then determine the target long-term reference frame according to the feedback information sent by the receiving end.
[0089] Specifically, if the network RTT delay is greater than the preset delay threshold thresholdRtt. Among them, thresholdRtt is determined according to empirical values, such as 500ms, then determine the long-term reference frame determined in step S102 as the target long-term reference frame, and force the ordinary p-frames within the LTRFramePeriod cycle to be encoded with reference to this LTR frame. Because when the RTT is relatively long, the feedback period is also relatively long, and the feedback packets may also be lost, which causes the step size of the LTR to become very long, resulting in a sharp deterioration of the image quality. Therefore, the LTR inserted at the current moment can be used as the target long-term reference frame for reference encoding.
[0090] If the network RTT delay is less than the preset delay threshold thresholdRtt, further determine the target long-term reference frame according to the feedback information sent by the receiving end. That is, when the network is in a good condition at this time, the currently determined long-term reference frame is not used, and a simpler encoding strategy can be used or return to the normal preset rule to insert LTRs regularly or dynamically to ensure the video quality, reduce the time of screen freeze, and make the video picture smoother.
[0091] Specifically, refer to Figure 4 , step A30, includes:
[0092] Step A301: If the current round-trip delay is less than the preset delay threshold, then determine whether the feedback information indicating that the decoding end has received the historical long-term reference frame is received.
[0093] Step A302: If the feedback information is received, then execute the step of determining the historical long-term reference frame as the target long-term reference frame.
[0094] Step A303: If the feedback information is not received, then screen out the target long-term reference frame from each P-frame according to the reference frame insertion time interval.
[0095] When the sender encodes the current frame, it uses the historical LTR frames confirmed by the receiver as a reference for encoding, which can ensure better video frame smoothness. Here, the historical LTR frames confirmed by the receiver refer to the sender receiving the confirmation message sent by the receiver, and the above confirmation message indicates that the above LTR frames can be normally decoded by the receiver. The advantage of this reference relationship is that when the receiver receives video frames, they are all encoded with the confirmed historical LTR frames as reference frames. As long as the received video frames are complete, they can be decoded and displayed. Therefore, the sender will use the previously confirmed historical LTR as a reference to encode non-LTR P frames. That is, when the network is in good condition at this time, instead of using the currently determined long-term reference frame, the previous historical LTR is still used to ensure video quality, reduce the time of screen freezing, and make the video frames smoother.
[0096] It can be understood that if the feedback information is received, the historical LTR is normally received by the receiver, and at this time, encoding can be normally performed with the historical LTR as a reference.
[0097] If the feedback information is not received, an unexpected situation of historical LTR loss may occur. At this time, normal I-P-P-P frame encoding can be performed according to the reference frame insertion time interval, but some of the P frames will be selected and marked as LTR frames, that is, the target long-term reference frames.
[0098] Based on the above embodiments, the third embodiment of the video encoding method of the present application is proposed. Refer to Figure 5 , Figure 5 which is the flowchart of the third embodiment of the video encoding method of the present application.
[0099] In this embodiment, the method includes:
[0100] Step S201, obtain the packet loss rate at the previous moment of the current moment and the duration of the previous detection time period.
[0101] In this embodiment, the packet loss rate is the packet loss rate within the previous moment. For example, when the previous moment is the previous 1 s, the packet loss rate is the packet loss rate within the previous 1 s. The previous detection time period is the detection time period corresponding to the previous moment, that is, a time period with the moment at a duration of the previous detection time period before the previous moment as the start moment and the previous moment as the end moment. At the previous moment, the sender obtains the average packet loss rate of the video image within the duration of the previous detection time period before the previous moment.
[0102] In this embodiment, the sender re-determines a new current detection time period at each moment. The specific value of the duration of the current detection time period changes with the change of the packet loss rate. For example, the sender updates the previous duration to the duration of the previous detection time period every 1 s and dynamically determines the duration of the new current detection time period.
[0103] Step S202: Determine the duration of the current detection time period according to the packet loss rate and the duration of the previous detection time period.
[0104] Specifically, step S202 includes:
[0105] Step B10: Determine the ratio of the packet loss rate to the preset packet loss rate threshold;
[0106] Step B20: If the ratio is greater than or equal to a preset reference value, determine the duration of the current detection time period according to the duration of the previous detection time period and the first preset formula: The first preset formula is:
[0107] m i = m i-1 - adjustmentFact 2 ;
[0108] Step B30: If the ratio is less than the preset reference value, determine the duration of the current detection time period according to the duration of the previous detection time period and the second preset formula: The second preset formula is:
[0109] m i = m i-1 + 1;
[0110] Where, m i is the duration of the current detection time period, m i-1 is the duration of the previous detection time period, and adjustmentFact is the ratio.
[0111] Specifically,
[0112] Where, adjustmentFact is the ratio, curSecondLossRate is the packet loss rate, and n is the preset packet loss rate threshold. The initial value of m i can be taken as 10.
[0113] If the ratio is greater than or equal to the preset reference value, that is, the packet loss rate at the previous moment of the current moment is relatively high, then quickly converge the duration m i of the current detection time period. On the contrary, when the packet loss rate at the previous moment of the current moment is relatively low, then increase the duration m i of the current detection time period.
[0114] After obtaining the duration m i , the time period obtained by pushing forward m i duration with the current moment as the end moment is the current detection time period of the current moment.
[0115] Therefore, in this embodiment, the duration value of the current detection time period corresponding to each moment is dynamically adjusted according to the packet loss rate to better determine the network transmission quality at the current moment.
[0116] As an embodiment, step S202 includes:
[0117] (1) Obtain the duration to be determined according to the packet loss rate and the duration of the previous detection time period;
[0118] (2) If the duration to be determined is greater than or equal to the preset maximum detection duration, then use the preset maximum detection duration as the duration of the current detection time period;
[0119] (3) If the duration to be determined is less than the preset maximum detection duration and greater than the preset minimum detection duration, then use the duration to be determined as the duration of the current detection time period;
[0120] (4) If the duration to be determined is less than or equal to the preset minimum detection duration, then use the preset minimum detection duration as the duration of the current detection time period.
[0121] Specifically, in this embodiment,
[0122] Therefore, when the duration to be determined is greater than the preset maximum detection duration, the maximum detection duration is taken as the duration of the current detection time period. When the duration to be determined is less than the preset minimum duration, the preset minimum detection duration is taken as the duration of the current detection time period. When the duration to be determined is between the preset maximum detection duration and the preset minimum detection duration, then the duration to be determined is determined as the duration m of the current detection time period i .
[0123] In this embodiment, the upper limit value and the lower limit value of the duration of the current detection time period are limited by the preset maximum detection duration and the preset minimum detection duration, so as to avoid the duration of the current detection time period taking values outside the preset maximum detection duration and the preset minimum detection duration from affecting the accuracy of the detection result of the network quality.
[0124] Step S203, obtain the average packet loss rate of the video image within the current detection time period.
[0125] Step S204, determine the long-term reference frame according to the average packet loss rate, and obtain the real-time average packet loss rate at each moment and the current round-trip delay of the current round-trip delay period at each moment.
[0126] Step S205, adjust the reference frame insertion time interval according to the real-time average packet loss rate.
[0127] In this embodiment, this step specifically includes:
[0128] Based on the real-time average packet loss rate, the preset packet loss rate threshold, and the first preset formula, obtain the reference frame insertion time interval of the long-term reference frame insertion period; the first preset formula is:
[0129]
[0130] where LTRFramePEriod is the reference frame insertion time interval, averageLossRate is the real-time average packet loss rate, and t is a natural number. Optionally, t = 3.
[0131] Thus, when the network is good, that is, when the real-time average packet loss rate is low, the reference frame insertion time interval is adjusted to be shorter, and when the network is poor, the reference frame insertion time interval is adjusted to be longer, so as to balance image quality and network transmission.
[0132] Step S206: If the reference frame insertion time interval is greater than the preset maximum insertion time interval, then use the preset maximum insertion time interval as the reference frame insertion time interval.
[0133] Specifically,
[0134]
[0135] where the preset maximum insertion time interval maxRefFrameNum is optionally 16. If LTRFramePeriod is less than the preset maximum number of reference frames, then LTRFramePeriod remains unchanged, otherwise LTRFramePeriod is equal to 16.
[0136] In this step, the step value of inserting the LTR is limited by the preset maximum insertion time interval to avoid an overly large insertion time interval affecting the normal coding quality.
[0137] Step S207: Determine the target long-term reference frame according to the current round-trip delay.
[0138] Step S208: Encode each P frame of the video image by using the target long-term reference frame and the reference frame insertion time interval.
[0139] For ease of understanding, a specific implementation is shown below:
[0140] The live server sends video stream data to the mobile terminal. During this process, the live server obtains the current round-trip delay RTT of the current round-trip delay period at the current moment and the average packet loss rate within the previous 10 s before the current moment. When the average packet loss rate is greater than or equal to the preset packet loss rate threshold for the first time at the moment of 03:00, the live server takes the current frame as the long-term reference frame LTR and enters the long-term reference frame insertion period at this time. And according to the ratio of the average packet loss rate to the preset packet loss rate threshold, the reference frame insertion time interval of 16 s is calculated. At this time, if RTT>500 ms, then within the first reference frame insertion time interval, encoding and sending are performed with this LTR as the reference.
[0141] At each moment after the long-term reference frame insertion period, the live server obtains the current round-trip delay RTT of the current round-trip delay period at each moment and the real-time average packet loss rate within the current detection time period corresponding to each moment, and continuously updates the reference frame insertion time interval according to the real-time average packet loss rate, so as to make the LTR step length adjustment shorter when the network is good and longer when the network is poor, so as to balance the image quality and network transmission.
[0142] When the live server detects at 30:10 that the packet loss rate per second within 10 consecutive seconds is less than the preset packet loss rate threshold, the long-term reference frame insertion period is ended, and the live server inserts the LTR according to the conventional LTR insertion algorithm for video encoding and transmission.
[0143] In addition, an embodiment of the present invention further provides a computer storage medium, on which a video encoding program is stored. When the video encoding program is executed by a processor, the steps of the video encoding method as described above are implemented. Therefore, it will not be elaborated here. In addition, the description of the beneficial effects of adopting the same method will not be elaborated either. For the technical details not disclosed in the embodiment of the computer-readable storage medium involved in this application, please refer to the description of the method embodiment of this application. By way of example, the program instructions can be deployed to be executed on one computing device, or on multiple computing devices located at one location, or on multiple computing devices distributed at multiple locations and interconnected by a communication network.
[0144] Those of ordinary skill in the art can understand that all or part of the processes of implementing the methods in the above embodiments can be completed by instructing relevant hardware through a computer program. The above program can be stored in a computer-readable storage medium. When the program is executed, it can include the processes of the embodiments of the above methods. Among them, the above storage medium can be a magnetic disk, an optical disk, a read-only memory (ROM), or a random access memory (RAM), etc.
[0145] In addition, it should be noted that the device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. In addition, in the attached drawings of the device embodiments provided by the present invention, the connection relationships between the modules indicate that they have communication connections, which can be specifically implemented as one or more communication buses or signal lines. Those of ordinary skill in the art can understand and implement without creative efforts.
[0146] Through the description of the above embodiments, those skilled in the art can clearly understand that the present invention can be implemented by means of software plus necessary general hardware, and of course, it can also be implemented by dedicated hardware including application-specific integrated circuits, dedicated CPUs, dedicated memories, dedicated components, etc. Generally, functions completed by computer programs can be easily implemented by corresponding hardware, and the specific hardware structures for implementing the same function can also be various, such as analog circuits, digital circuits or dedicated circuits. However, in more cases, software program implementation is a better implementation manner for the present invention. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a readable storage medium, such as a floppy disk, a USB flash drive, a mobile hard disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a magnetic disk or an optical disc of a computer, etc., and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods of various embodiments of the present invention.
[0147] The above are only the preferred embodiments of the present invention, and do not limit the patent scope of the present invention accordingly. Any equivalent structural or equivalent process transformations made by using the description and drawings of the present invention, or directly or indirectly applied to other related technical fields, are equally included in the patent protection scope of the present invention.
Claims
1. A video encoding method, characterized in that, The method includes: Obtaining the average packet loss rate of the video image within the current detection time period; Determining a long-term reference frame according to the average packet loss rate, and obtaining the real-time average packet loss rate at each moment and the current round-trip delay of the current round-trip delay period at each moment; Adjusting the reference frame insertion time interval according to the real-time average packet loss rate; Determining a target long-term reference frame according to the current round-trip delay; Encoding each P frame of the video image by using the target long-term reference frame and the reference frame insertion time interval; The determining the long-term reference frame according to the average packet loss rate includes: When the average packet loss rate is greater than a preset packet loss rate threshold, immediately inserting a long-term reference frame; or, if the average packet loss rate is greater than or equal to the preset packet loss rate threshold for the first time, taking the current frame as the long-term reference frame and taking the current moment as the start moment of the period, and entering the long-term reference frame insertion period; The determining the target long-term reference frame according to the current round-trip delay includes: Judging whether the current round-trip delay is greater than a preset delay threshold; If the current round-trip delay is greater than the preset delay threshold, determining the long-term reference frame as the target long-term reference frame; If the current round-trip delay is less than the preset delay threshold, determining the target long-term reference frame according to the feedback information sent by the receiving end.
2. The video encoding method according to claim 1, wherein The if the current round-trip delay is less than the preset delay threshold, determining the target long-term reference frame according to the feedback information sent by the receiving end includes: If the current round-trip delay is less than the preset delay threshold, judging whether the feedback information that the receiving end confirms to have received the historical long-term reference frame is received; If the feedback information is received, determining the historical long-term reference frame as the target long-term reference frame; If the feedback information is not received, screening out the target long-term reference frame from each P frame according to the reference frame insertion time interval.
3. The video encoding method according to claim 1, wherein Before obtaining the average packet loss rate of the video image within the current detection time period, the method further includes: Obtaining the packet loss rate at the previous moment of the current moment and the duration of the previous detection time period; Determining the duration of the current detection time period according to the packet loss rate and the duration of the previous detection time period.
4. The video encoding method according to claim 3, wherein The determining the duration of the current detection time period according to the packet loss rate and the duration of the previous detection time period, including: Determining the ratio of the packet loss rate to the preset packet loss rate threshold; If the ratio is greater than or equal to a preset reference value, determining the duration of the current detection time period according to the duration of the previous detection time period and a first preset formula: The first preset formula is: m i = m i-1 - adjustmentFact 2 ; If the ratio is less than the preset reference value, determining the duration of the current detection time period according to the duration of the previous detection time period and a second preset formula: The second preset formula is: m i = m i-1 + 1; where m i is the duration of the current detection time period, and m i-1 is the duration of the previous detection time period, and adjustmentFact is the ratio.
5. The video encoding method according to claim 3, wherein The determining the duration of the current detection time period according to the packet loss rate and the duration of the previous detection time period includes: Obtaining a duration to be determined according to the packet loss rate and the duration of the previous detection time period; If the to-be-determined duration is greater than or equal to the preset maximum detection duration, then use the preset maximum detection duration as the duration of the current detection time period; If the to-be-determined duration is less than the preset maximum detection duration and greater than the preset minimum detection duration, then use the to-be-determined duration as the duration of the current detection time period; If the to-be-determined duration is less than or equal to the preset minimum detection duration, then use the preset minimum detection duration as the duration of the current detection time period.
6. The video encoding method according to claim 4, wherein The adjusting the reference frame insertion time interval according to the real-time average packet loss rate includes: Obtaining the reference frame insertion time interval according to the real-time average packet loss rate, the preset packet loss rate threshold, and a first preset formula; the first preset formula is: Wherein, LTRFramePEriod is the reference frame insertion time interval, averageLossRate is the real-time average packet loss rate, n is the preset packet loss rate threshold, and t is a natural number.
7. The video encoding method according to any one of claims 1 to 6, characterized in that, After adjusting the reference frame insertion time interval according to the real-time average packet loss rate, the method further includes: If the reference frame insertion time interval is greater than the preset maximum insertion time interval, then use the preset maximum insertion time interval as the reference frame insertion time interval.
8. An image sending device, characterized in that, Including: A memory, a processor, and a video encoding program stored on the memory and executable on the processor, the video encoding program being configured to implement the steps of the video encoding method according to any one of claims 1 to 7.
9. A computer-readable storage medium, characterized in that, A video encoding program is stored on the computer-readable storage medium, and when the video encoding program is executed by the processor, the steps of the video encoding method according to any one of claims 1 to 7 are implemented.
Citation Information
Patent Citations
Video encoding method using long-term reference frame, electronic equipment, and system
CN106817585A
Video image transmission method, video image sending equipment, video call method and video call equipment
CN112532908A