Method for transmitting and receiving a video stream
By using SVC and FEC technologies, the redundancy of the enhancement layer is reduced first and then the redundancy of the base layer is reduced layer by layer, which solves the problem of low video stream success rate in streaming media transmission and achieves efficient video stream transmission under limited network bandwidth.
Patent Information
- Application Number
- CN202310360126.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-03-31
- Publication Date
- 2025-11-28
- Estimated Expiration
- 2043-03-31
AI Technical Summary
In existing streaming media transmission technologies, the success rate of video streams needs to be improved, especially when network bandwidth is limited, making it impossible to guarantee successful transmission of video streams.
The video stream is divided into a base layer and an enhancement layer using Scalable Video Coding (SVC) technology. Redundancy information is added through forward error correction coding (FEC), and when network bandwidth is insufficient, the redundancy of the enhancement layer is reduced first, and then layer by layer, the redundancy is reduced to that of the base layer until the video stream can be successfully transmitted.
While ensuring the quality of the base layer transmission, the success rate of video stream transmission is improved and the probability of video stream transmission failure is reduced.
Smart Images

Figure CN116506658B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of streaming media transmission, in particular to a video stream sending method and a receiving method. BACKGROUND
[0002] Streaming media is also called streaming media, which is a media that is transmitted and played simultaneously. The user of the client can continuously receive and watch or listen to the transmitted media while the media provider is transmitting the media on the network. Therefore, in order to ensure that the user of the client can successfully watch or listen to the transmitted media, the successful transmission of the streaming media must be ensured. However, the success rate of the streaming media transmission needs to be further improved. SUMMARY
[0003] The present application provides a video stream sending method and a receiving method, which can improve the success rate of video stream sending.
[0004] The first aspect of the embodiment of the present application provides a video stream sending method, which comprises the following steps: determining a target network bandwidth; judging whether the current corresponding encoding data of a first picture group in an SVC video stream can be transmitted under the target network bandwidth, wherein the first picture group comprises a base layer and at least one enhancement layer, the base layer and the enhancement layer each comprise a plurality of video frames, and the current corresponding encoding data of the first picture group is obtained after the video frames in the first picture group are encoded according to their respective current redundancy; if the current corresponding encoding data of the first picture group cannot be transmitted, reducing the current redundancy of the video frames in the at least one enhancement layer; after the current redundancy of the video frames in the at least one enhancement layer is reduced, in response to the current corresponding encoding data of the first picture group still being unable to be transmitted under the target network bandwidth, reducing the current redundancy of the video frames in the base layer; and sending each video frame in the first picture group to a receiving end in sequence after the video frame is redundantly encoded according to its current redundancy.
[0005] The second aspect of the embodiment of the present application provides a video stream receiving method, which comprises the following steps: playing the encoding data sent by the sending end after the encoding data is decoded; wherein the encoding data is sent by the sending end by using any one of the above methods.
[0006] The third aspect of the embodiment of the present application provides an electronic device, which comprises a processor, a memory and a communication circuit. The processor is coupled to the memory and the communication circuit, respectively. The memory stores program data. The processor realizes the steps in the above method by executing the program data in the memory.
[0007] The fourth aspect of the embodiment of the present application provides a computer readable storage medium, which stores a computer program, and the computer program can be executed by a processor to implement the steps in the above method.
[0008] The beneficial effect is that, when it is determined that the current encoding data corresponding to the first picture group cannot be transmitted under the target network bandwidth, the redundancy of the current video frame in the enhancement layer is reduced preferentially, and then only when the current encoding data corresponding to the first picture group still cannot be transmitted under the target network bandwidth, the redundancy of the current video frame in the base layer is reduced, that is, the layered redundancy reduction processing is performed, which can improve the success rate of video stream transmission under the premise of ensuring the transmission quality of the base layer. BRIEF DESCRIPTION OF DRAWINGS
[0009] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings needed in the embodiment description will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor on the basis of these drawings.
[0010] Figure 1 is a flowchart of an embodiment of the video stream sending method of the present application;
[0011] Figure 2 is a schematic diagram of the principle of FEC encoding technology;
[0012] Figure 3 is a flowchart of step S13 in Figure 1 ;
[0013] Figure 4 is a flowchart of step S132 in Figure 3 ;
[0014] Figure 5 is a schematic diagram of reducing the redundancy of the current video frame in the enhancement layer first according to the present application;
[0015] Figure 6 is a schematic diagram of discarding the video frame in the enhancement layer according to the present application;
[0016] Figure 7 is a schematic diagram of reducing the redundancy of the current video frame in the base layer according to the present application;
[0017] Figure 8 is a flowchart of step S15 in Figure 1 ;
[0018] Figure 9 is a flowchart of step S15 in Figure 8 ;
[0019] Figure 10 is a flowchart of an embodiment of a receiving method of a video stream of the present application;
[0020] Figure 11 is Figure 10 is a flowchart of step S21 in
[0021] Figure 12 is a structural diagram of a sending end and a receiving end of the present application;
[0022] Figure 13 is a structural diagram of an embodiment of an electronic device of the present application;
[0023] Figure 14 is a structural diagram of an embodiment of a computer readable storage medium of the present application. DETAILED DESCRIPTION
[0024] The technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only some of the embodiments of the present application, but not all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative work fall within the scope of protection of the present application.
[0025] It should be noted that the terms “first” and “second” in the present application are only used for the purpose of description, and cannot be understood as indicating or implying relative importance or implicitly indicating the number of the indicated technical features. Therefore, the features limited by “first” and “second” can explicitly or implicitly include at least one of the features. In the description of the present application, the meaning of “multiple” is at least two, for example, two, three, etc., unless otherwise specifically limited. In addition, the terms “include” and “have” and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device including a series of steps or units is not limited to the listed steps or units, but can optionally include steps or units not listed, or can optionally include other steps or units inherent to the process, method, product or device.
[0026] Referring to Figure 1 In an embodiment of the present application, the sending method of the video stream includes:
[0027] S11: determining a current target network bandwidth.
[0028] The sending method in the present embodiment is executed by a sending end, which can be a camera or other device, without limitation.
[0029] The network bandwidth at the current time is defined as the target network bandwidth, and the method for determining the target network bandwidth can be referred to the following description.
[0030] S12: judging whether the current encoding data corresponding to the first picture group in the SVC video stream can be transmitted under the target network bandwidth.
[0031] The first picture group comprises one base layer and at least one enhancement layer, and the base layer and the enhancement layer each comprise a plurality of video frames. The current encoding data corresponding to the first picture group is obtained after the video frames in the first picture group are encoded according to respective current redundancy of the video frames.
[0032] When the judgment result is that the current encoding data can be transmitted, step S15 is executed; and when the judgment result is that the current encoding data cannot be transmitted, step S13 is executed.
[0033] Specifically, the SVC (Scalable Video Coding) technology is used to process the video stream to be transmitted to obtain the SVC video stream. The SVC technology is a technology capable of dividing the video stream into multiple resolution, quality and frame rate layers, and is an extension of the H.264 video coding standard, and is referred to as H.264-SVC. The SVC technology divides the video stream into one base layer and at least one enhancement layer (the number of the enhancement layers can be one or multiple). The data in the base layer is data capable of providing the most basic video quality, frame rate and resolution, and the data in the enhancement layer is data capable of improving the video quality, frame rate and resolution of the data in the base layer. If the data in the base layer is damaged, the influence on the video stream is great, and if the data in the enhancement layer is damaged, the influence on the video stream is relatively small.
[0034] The GOP (Group of Pictures) to be transmitted by the sending end to the receiving end is defined as the first picture group, and steps S120 and subsequent steps are executed for the first picture group. The first picture group also comprises one base layer and at least one enhancement layer because the first picture group is in the SVC video stream. The number of the enhancement layers can be one or multiple. The base layer and the enhancement layer each comprise a plurality of video frames, that is, the video frames in the first picture group are either the video frames in the base layer or the video frames in the enhancement layer.
[0035] Meanwhile, in combination with Figure 2 , the forward error correction technology (FEC technology) is introduced as follows:
[0036] The sending end adds some redundant information in the data packet when transmitting the data. The receiving end can recover the complete valid data according to the valid data and the redundant information even if some data packets are lost.
[0037] If the video frame has redundancy, the video frame is forward error correction coded according to the redundancy; and if the video frame has no redundancy, the video frame is not forward error correction coded.
[0038] However, when step S12 is performed, some of the video frames in the first picture group currently have corresponding redundancy, and some of the video frames in the first picture group currently can have no corresponding redundancy. The current redundancy of the video frames in the first picture group can be preset, for example, the maximum redundancy corresponding to each of the video frames in the first picture group is preset, and then the maximum redundancy corresponding to each of the video frames in the first picture group is determined as the current redundancy of the video frames when step S12 is performed.
[0039] The current encoding data corresponding to the first picture group is defined as the encoding data obtained after each of the video frames in the first picture group is encoded according to the current redundancy of the video frame. It can be understood that the current redundancy of the video frames can change in the subsequent processing process, and thus the current encoding data corresponding to the first picture group also changes.
[0040] If the current encoding data corresponding to the first picture group can be transmitted under the target network bandwidth, it means that the encoding data can be successfully sent to the receiving end after each of the video frames in the first picture group is encoded according to the current redundancy of the video frame. If the current encoding data corresponding to the first picture group cannot be transmitted, it means that there can be a failure when the encoding data is sent to the receiving end. Therefore, in order to improve the probability of successful transmission, the redundancy of the video frames in the first picture group needs to be reduced, that is, the redundancy of the video frames in the first picture group is reduced.
[0041] S13: Reduce the current redundancy of the video frames in at least one enhancement layer.
[0042] Specifically, from the above introduction of the base layer and the enhancement layer, it can be known that the base layer stores data that can provide the most basic video quality, frame rate and resolution, and the enhancement layer stores data that can improve the video quality, frame rate and resolution of the base layer data. Therefore, the importance of the base layer is higher than that of the enhancement layer, and thus in the application, the current redundancy of the video frames in the enhancement layer is reduced first.
[0043] Referring to Figure 3 In the embodiment, step S13 includes:
[0044] S131: Determine a first enhancement layer in the first picture group.
[0045] The first enhancement layer is the highest enhancement layer in the first picture group.
[0046] S132: Reduce the current redundancy of the video frames in the first enhancement layer.
[0047] wherein, after reducing the current redundancy of the video frames in the first enhancement layer, the current redundancy of each video frame in the first enhancement layer is not lower than the corresponding minimum redundancy of the first enhancement layer.
[0048] S133: determining whether the current encoding data corresponding to the first picture group can be transmitted under the target network bandwidth.
[0049] If the result of the determination is that the current encoding data corresponding to the first picture group can be transmitted under the target network bandwidth, step S134 is executed; if the result of the determination is that the current encoding data corresponding to the first picture group cannot be transmitted under the target network bandwidth, step S135 is executed.
[0050] S134: stopping reducing the current redundancy of the video frames in any enhancement layer.
[0051] S135: discarding the video frames in the first enhancement layer, and determining a second enhancement layer, which is located before the first enhancement layer and adjacent to the first enhancement layer in the first picture group, as the first enhancement layer.
[0052] After step S135, step S132 is executed.
[0053] wherein, if the video frames in the first enhancement layer are determined to be discarded in step S135, the video frames in the first enhancement layer will not be transmitted to the receiving end by the subsequent sending end.
[0054] Specifically, when reducing the current redundancy of the video frames in at least one enhancement layer, the current redundancy of the video frames in the highest enhancement layer is reduced first; when the current redundancy of the video frames in the highest enhancement layer is reduced, if the current encoding data corresponding to the first picture group still cannot be transmitted under the target network bandwidth, the video frames in the highest enhancement layer are discarded, and the current redundancy of the video frames in the next highest enhancement layer is continuously reduced; when the current redundancy of the video frames in the next highest enhancement layer is reduced, if the current encoding data corresponding to the first picture group still cannot be transmitted under the target network bandwidth, the video frames in the next highest enhancement layer are discarded, and the current redundancy of the video frames in the enhancement layer, which is adjacent to the next highest enhancement layer and lower than the next highest enhancement layer, is continuously reduced; then the above process is repeatedly performed until the current encoding data corresponding to the first picture group can be transmitted under the target network bandwidth or the video frames in all enhancement layers are discarded. Meanwhile, in the above process, if the current encoding data corresponding to the first picture group can be transmitted under the target network bandwidth, the reduction of the current redundancy of the video frames in the enhancement layer is stopped, and the current redundancy of the video frames in the base layer will not be reduced subsequently.
[0055] For the convenience of understanding, the following examples are provided:
[0056] If there are three enhancement layers, from low to high, they are the first enhancement layer, the second enhancement layer and the third enhancement layer, the redundancy of the current video frame in the third enhancement layer is reduced first, and after the redundancy is reduced, if the encoding data corresponding to the current first picture group can be transmitted under the target network bandwidth, the reduction of the redundancy of the current video frame in the enhancement layer is stopped, and the redundancy of the current video frame in the base layer will not be reduced subsequently. However, if the encoding data corresponding to the current first picture group cannot be transmitted under the target network bandwidth, the video frame in the third enhancement layer is discarded, the redundancy of the current video frame in the second enhancement layer is reduced, and after the redundancy is reduced, if the encoding data corresponding to the current first picture group can be transmitted under the target network bandwidth, the reduction of the redundancy of the current video frame in the enhancement layer is stopped, and the redundancy of the current video frame in the base layer will not be reduced subsequently. However, if the encoding data corresponding to the current first picture group cannot be transmitted under the target network bandwidth, the video frame in the second enhancement layer is discarded, the redundancy of the current video frame in the first enhancement layer is reduced, and after the redundancy is reduced, if the encoding data corresponding to the current first picture group can be transmitted under the target network bandwidth, the reduction of the redundancy of the current video frame in the enhancement layer is stopped, and the redundancy of the current video frame in the base layer will not be reduced subsequently. However, if the encoding data corresponding to the current first picture group cannot be transmitted under the target network bandwidth, the video frame in the first enhancement layer is discarded, and the redundancy of the current video frame in the base layer is reduced subsequently.
[0057] Each enhancement layer corresponds to a minimum redundancy, and after the redundancy of the current video frame in the enhancement layer is reduced, the redundancy of any current video frame in the enhancement layer is not lower than the minimum redundancy corresponding to the enhancement layer.
[0058] In the embodiment, referring to Figure 4 , the step S132 includes:
[0059] S1321: determining a first video frame in the first enhancement layer.
[0060] The first video frame is arranged at the last in the first enhancement layer.
[0061] S1322: reducing the redundancy of the current first video frame.
[0062] S1323: judging whether the encoding data corresponding to the current first picture group can be transmitted under the target network bandwidth.
[0063] If the result of the judgment is that the encoding data can be transmitted, the step S1324 is executed, otherwise the step S1325 is executed.
[0064] S1324: stopping the reduction of the redundancy of the current video frame in any enhancement layer.
[0065] S1325: determining a second video frame located before the first video frame and adjacent to the first video frame in the first enhancement layer as the first video frame.
[0066] After step S1325, return to step S1322.
[0067] Specifically, when reducing the current redundancy of the video frames in the first enhancement layer, the redundancy of the tail video frame (the last video frame in the first enhancement layer) is reduced first, and if the transmission under the target network bandwidth is possible, the reduction of the redundancy of the video frames in the first enhancement layer is stopped, but if the transmission under the target network bandwidth is still not possible, the redundancy of the second last video frame is reduced, and so on, until the current redundancy of the first video frame in the first enhancement layer is reduced.
[0068] S14: after reducing the current redundancy of the video frames in at least one enhancement layer, in response to the current encoding data of the first picture group still being unable to be transmitted under the target network bandwidth, reducing the current redundancy of the video frames in the base layer.
[0069] Specifically, if the current encoding data of the first picture group is still unable to be transmitted under the target network bandwidth after reducing the redundancy of the video frames in the enhancement layer, the redundancy of the video frames in the base layer is reduced.
[0070] In the embodiment, the process of reducing the current redundancy of the video frames in the base layer is basically similar to the process of reducing the current redundancy of the video frames in the first enhancement layer, and the process of reducing the current redundancy of the video frames in the base layer comprises:
[0071] determining a first video frame in the base layer; reducing the current redundancy of the first video frame, wherein after reducing the redundancy of the first video frame, the current redundancy of the first video frame is not lower than the minimum redundancy corresponding to the base layer; if after reducing the current redundancy of the first video frame, in response to the current encoding data of the first picture group still being unable to be transmitted under the target network bandwidth, determining a second video frame located before the first video frame and adjacent to the first video frame in the base layer as the first video frame, and then returning to the step of reducing the current redundancy of the first video frame.
[0072] Specifically, when reducing the current redundancy of the video frames in the base layer, the redundancy of the tail video frame (the last video frame in the base layer) is reduced first, and if the transmission under the target network bandwidth is possible, the reduction of the redundancy of the video frames in the base layer is stopped, but if the transmission under the target network bandwidth is still not possible, the redundancy of the second last video frame is reduced, and so on, until the transmission under the current target network bandwidth is possible or the current redundancy of the first video frame in the base layer is reduced.
[0073] Meanwhile, in the embodiment, when the redundancy of the first video frame in the base layer is reduced, if the encoding data corresponding to the current first picture group still cannot be transmitted under the target network bandwidth, it is determined that the video frames in the base layer will not be processed in redundancy in the future, i.e., the redundancy of the video frames in the base layer will not be set.
[0074] In order to better understand the process of reducing the redundancy, the following further describes the process of reducing the redundancy by taking an example of Figure 5 to Figure 7 , which first assumes that in the example, the first picture group includes a base layer and an enhancement layer, and in Figure 5 to Figure 7 , the graphics filled with patterns represent the video frames in the base layer, and the graphics not filled with patterns represent the video frames in the enhancement layer.
[0075] When it is determined that the encoding data corresponding to the current first picture group cannot be transmitted under the target network bandwidth, the redundancy of the video frames in the enhancement layer is first reduced, and when the redundancy is reduced, the redundancy of the video frames in the enhancement layer is sequentially reduced in the order from back to front. Among them, even after the redundancy of the video frames is reduced, the redundancy of each video frame in the enhancement layer is not lower than the minimum redundancy corresponding to the enhancement layer.
[0076] After the redundancy of all the video frames in the enhancement layer is reduced, if the encoding data corresponding to the current first picture group still cannot be transmitted under the target network bandwidth, all the video frames in the enhancement layer are discarded.
[0077] If after all the video frames in the enhancement layer are discarded, the encoding data corresponding to the current first picture group still cannot be transmitted under the target network bandwidth, the redundancy of the video frames in the base layer is reduced, and when the redundancy is reduced, the redundancy of the video frames in the base layer is sequentially reduced in the order from back to front. Among them, even after the redundancy of the video frames is reduced, the redundancy of each video frame in the base layer is not lower than the minimum redundancy corresponding to the base layer.
[0078] After the redundancy of all the video frames in the base layer is reduced, if the encoding data corresponding to the current first picture group still cannot be transmitted under the target network bandwidth, it is determined that the base layer will not be processed in redundancy in the future, i.e., the redundancy of the video frames in the base layer will not be set.
[0079] It should be noted that in other embodiments, the specific process of reducing the redundancy of the video frames in the enhancement layer is not limited, and the specific process of reducing the redundancy of the video frames in the base layer is not limited, as long as the redundancy of the video frames in the enhancement layer is preferentially reduced.
[0080] For example, in the process of reducing the current redundancy of the video frames in the enhancement layer, the reduction can be performed in the order from front to back (i.e. the current redundancy of the head video frame in the enhancement layer is reduced first), and after the current redundancy of the video frames in the enhancement layer is reduced, the video frames in the enhancement layer can not be discarded even if they cannot be transmitted at the target network bandwidth.
[0081] In the process of reducing the current redundancy of the video frames in the base layer, the reduction can also be performed in the order from front to back (i.e. the current redundancy of the head video frame in the base layer is reduced first), and after the current redundancy of the video frames in the base layer is reduced, the video frames in the base layer can not be abandoned for redundancy processing even if they cannot be transmitted at the target network bandwidth.
[0082] S15: sequentially encode and send each video frame in the first picture group to the receiving end according to the current redundancy thereof.
[0083] Specifically, after the current redundancy of the video frames in the first picture group is finally determined, each video frame in the first picture group is sequentially encoded and sent to the receiving end, wherein when each video frame in the first picture group is encoded, the corresponding encoded data of the video frame is sent to the receiving end.
[0084] As can be seen from the above, when it is determined that the current encoded data corresponding to the first picture group cannot be transmitted at the target network bandwidth, the current redundancy of the video frames in the enhancement layer is reduced first, and then only when the current encoded data corresponding to the first picture group still cannot be transmitted at the target network bandwidth, the current redundancy of the video frames in the base layer is reduced, i.e. the redundancy is reduced in layers, which can improve the success rate of video stream transmission on the premise of ensuring the transmission quality of the base layer.
[0085] Firstly, as can be seen from the above, in the first picture group after the redundancy is reduced, as long as there is a video frame in the enhancement layer, there must be corresponding redundancy for the video frame, and for the video frames in the base layer, there can be corresponding redundancy or there can be no corresponding redundancy.
[0086] The specific process of step S15 will be described below in conjunction with Figure 8 .
[0087] S101: Obtain the next first to-be-sent video frame to be sent.
[0088] S102: Determine whether the first to-be-sent video frame belongs to the video frames in the enhancement layer.
[0089] If the determination result is that it belongs, S107 is executed, and if the determination result is that it does not belong, S103 is executed.
[0090] S103: Determine whether the target network bandwidth can meet the minimum redundancy transmission requirements of the base layer.
[0091] Specifically, step S103 means: determining whether the encoded data corresponding to the first frame group can be transmitted under the target network bandwidth after setting the current redundancy of each video frame in the base layer to the minimum redundancy corresponding to the base layer.
[0092] If the judgment result is that the condition can be met, then proceed to step S104; otherwise, proceed to step S106.
[0093] S104: Obtain the redundancy of the first video frame to be sent.
[0094] S105: Perform FEC encoding on the first video frame to be sent according to the redundancy.
[0095] Specifically, step S105 means: performing forward error correction coding on the first video frame to be sent according to the current redundancy of the first video frame to be sent, to obtain multiple first data packets corresponding to the first video frame to be sent.
[0096] After step S105, step S112 is executed.
[0097] S106: Segment the first video frame to be sent.
[0098] In step S106, the first video frame to be sent is not subjected to redundancy processing. Instead, the first video frame to be sent is directly segmented to obtain multiple first data packets.
[0099] After step S106, step S113 is executed.
[0100] S107: Determine whether there are any dropped video frames in the enhancement layer where the first video frame to be sent is located.
[0101] If it exists, proceed to step S108; otherwise, proceed to step S109.
[0102] S108: Discard the first video frame to be sent.
[0103] After S108, return to step S101.
[0104] S109: Determine whether the target network bandwidth can meet the minimum redundancy transmission of the target enhancement layer.
[0105] Specifically, step S109 means: determining whether the encoded data corresponding to the first frame group can be transmitted under the target network bandwidth if the current redundancy of each video frame in the target enhancement layer is set to the minimum redundancy corresponding to the target enhancement layer. Here, the target enhancement layer is the enhancement layer where the first video frame to be sent is located.
[0106] If the result is not satisfied, step S108 is executed, otherwise step S110 is executed.
[0107] S110: Obtain the redundancy of the first to-be-sent video frame.
[0108] S111: Perform FEC encoding on the first to-be-sent video frame according to the redundancy.
[0109] Specifically, after FEC encoding, a plurality of first data packets can be obtained.
[0110] S112: Obtain the plurality of first data packets after FEC encoding.
[0111] S113: Package the plurality of first data packets to obtain a target data packet group, and send the target data packet group.
[0112] Specifically, the plurality of first data packets corresponding to the first to-be-sent video frame are packaged to obtain a target data packet group corresponding to the first to-be-sent video frame, the target data packet group includes a plurality of target data packets, and then the target data packet group is sent to the receiving end.
[0113] Specifically, the plurality of first data packets are specifically RTP packaged to obtain a target data packet group including a plurality of RTP data packets, and the target data packet group is subsequently sent to the receiving end through the RTP protocol.
[0114] In step S150, for any video frame to be sent in the first picture group, the redundancy of the video frame can be determined according to the following formula:
[0115] When the video frame k is a video frame in the enhancement layer, the redundancy b thereof is calculated according to the following formula:
[0116]
[0117] When the video frame l is a video frame in the base layer, the redundancy a thereof is calculated according to the following formula:
[0118]
[0119] Meanwhile, the calculation of the redundancy needs to satisfy the following constraint condition:
[0120]
[0121] wherein B l is the total size of all data packets obtained after FEC encoding of the video frame l in the base layer, P i is the size of the i-th original data packet of the video frame l, and R jMj is the size of the jth redundancy packet added after FEC encoding of the video frame I, and when j is 0, no redundancy packet is added. k Q is the total size of all data packets obtained after FEC encoding of the video frame k in the enhancement layer. i H is the size of the ith original data packet of the video frame k. j Mj is the size of the jth redundancy packet added after FEC encoding of the video frame k, and when j is 0, no redundancy packet is added. t g is the data amount of the video frames to be sent in the current GOP (the current GOP is the first picture group), g is the number of video frames in the base layer, and u is the number of video frames in the enhancement layer. Meanwhile, the redundancy a and b calculated above need to be less than the packet loss rate corresponding to the first picture group, and the process of obtaining the packet loss rate corresponding to the first picture group can be seen below.
[0122] In the above embodiment, the purpose of performing step S103 and step S109 is to further improve the success rate of transmission.
[0123] Specifically, when performing step S103, if the current redundancy of each video frame in the base layer is set to the minimum redundancy corresponding to the base layer, and the current encoding data corresponding to the first picture group cannot be transmitted under the target network bandwidth, it means that the current video frames in the base layer cannot be transmitted under the target network bandwidth only by redundancy processing, and in order to ensure the success rate of transmission, the first video frame to be sent is not processed by redundancy.
[0124] Similarly, when performing step S109, if the current redundancy of each video frame in the target enhancement layer is set to the minimum redundancy corresponding to the target enhancement layer, and the current encoding data corresponding to the first picture group cannot be transmitted under the target network bandwidth, it means that the current video frames in the target enhancement layer cannot be transmitted under the target network bandwidth only by redundancy processing, and in order to ensure the success rate of transmission, the video frames in the target enhancement layer are directly discarded, that is, the first video frame to be sent is lost.
[0125] If there are discarded video frames in the enhancement layer in which the first video frame to be sent is located, the video frames in the enhancement layer are discarded, that is, the video frames in the enhancement layer will not be sent to the receiving end in the future.
[0126] It should be noted that in other embodiments, steps S103 and / or S109 can not be performed.
[0127] Or, in other embodiments, step S107 can also not be performed, that is, even if there is a lost video frame in the target enhancement layer where the first to-be-sent video frame is located, the first to-be-sent video frame is sent to the receiving end as long as the current encoding data of the first picture group can be transmitted under the target network bandwidth.
[0128] In combination Figure 8 And Figure 9 After step S113, further comprising:
[0129] S114: determining whether a target identifier fed back by the receiving end is received.
[0130] If received, step S115 is performed, otherwise step S118 is performed.
[0131] Wherein, the receiving end sends the target identifier corresponding to the target data packet to the sending end when detecting that the target data packet is not received.
[0132] S115: determining whether the target data packet corresponding to the target identifier corresponds to a video frame in the base layer.
[0133] If the determination result is yes, step S116 is performed, otherwise step S117 is performed.
[0134] S116: sending the target data packets corresponding to the M target identifiers to the receiving end.
[0135] S117: sending the target data packets corresponding to the N target identifiers to the receiving end.
[0136] Wherein, M is greater than N.
[0137] S118: determining whether the transmission is ended.
[0138] If the determination result is yes, step S119 is performed, otherwise step S101 is returned.
[0139] Specifically, in the process of receiving the target data packet, if the receiving end detects a packet loss, it will send the target identifier corresponding to the target data packet to the receiving end.
[0140] When the sending end receives the target identifier, if it is determined that the target data packet corresponding to the target identifier is the target data packet corresponding to the video frame in the base layer, the target data packet is retransmitted M times, but if it is determined that the target data packet corresponding to the target identifier is the target data packet corresponding to the video frame in the enhancement layer, the target data packet is retransmitted N times, wherein considering that the importance of the base layer is higher than that of the enhancement layer, M is set to be greater than N, for example, M is equal to 2 times and N is equal to 1 time.
[0141] It should be noted that in other embodiments, the target data packet corresponding to the target identifier is retransmitted T times regardless of whether the target data packet corresponding to the target identifier is a target data packet corresponding to a video frame in the base layer or in the enhancement layer, i.e., the number of retransmissions is equal.
[0142] In the embodiment, after successful transmission of one GOP is completed, the receiving end receives the network bandwidth, packet loss rate, and the like fed back by the sending end, and then the receiving end takes the network bandwidth, packet loss rate, and the like as the target network bandwidth and the corresponding packet loss rate of the next GOP.
[0143] Referring to Figure 10 , Figure 10 is a flowchart of an embodiment of a receiving method of the video stream of the application. The receiving method comprises the following steps.
[0144] S20: playing the decoded data sent by the sending end.
[0145] The encoded data is sent by the sending end using the sending method in any of the above embodiments. For details, refer to the above description, which will not be repeated here.
[0146] Referring to Figure 11 , in the embodiment, step S20 comprises the following steps.
[0147] S201: receiving the target data packet of the target video frame sent by the sending end.
[0148] Specifically, the video frame currently received by the receiving end is defined as the target video frame.
[0149] S202: when none of the target data packets corresponding to the target video frame is currently received, determining whether the time from the second target time point reaches a second preset time.
[0150] If the determination result is yes, step S203 is performed, otherwise step S204 is performed.
[0151] Specifically, the second target time point can be the time point at which the target data packet group corresponding to the previous video frame is completed, and the second preset time can be determined according to the target time of the target video frame. The target time will be described below.
[0152] In an application scenario, the second preset time is equal to the target time of the target video frame.
[0153] When the time from the second target time point reaches the second preset time, if any target data packet corresponding to the target video frame has not been received, the sending end is directly notified to discard the target video frame.
[0154] It should be noted that in other embodiments, step S202 can also not be performed, i.e., step S202 is not a necessary step.
[0155] S203: notify the sending end to discard the sending of the target video frame.
[0156] S204: determine whether the target data packet corresponding to the target video frame that has been received can be decoded and processed.
[0157] If the determination result is yes, step S205 is performed, otherwise step S206 is performed.
[0158] S205: decode and play the target data packet corresponding to the target video frame that has been received.
[0159] Specifically, if the decoding and processing can be performed, it means that the receiving end has completed the reception of the target video frame, and then the decoding and playing is directly performed.
[0160] S206: determine whether there is a missing target data packet at present.
[0161] If the determination result is yes, step S207 is performed, otherwise step S201 is performed.
[0162] Specifically, if there is no missing target data packet at present, the target data packet corresponding to the target video frame is continuously received. If there is, step S207 is performed.
[0163] S207: determine whether the time from the first target time point is the first preset time.
[0164] If the determination result is yes, step S208 is performed, otherwise step S209 is performed.
[0165] Specifically, if the time from the first target time point is not the first preset time, it is determined that the missing target data packet is only delayed sending, and is not a real loss, so as to avoid the repeated sending of the sending end. Therefore, only when the time from the first target time point is the first preset time, the sending end is notified to resend the missing target data packet.
[0166] The first target time point and the second target time point described above can be the same time point or can not be the same time point.
[0167] In an application scenario, the first preset time is equal to one half of the target time.
[0168] S208: notify the sending end of the identifier of the missing target data packet.
[0169] S209: determining whether the target video frame belongs to the video frame in the base layer.
[0170] If the result of the determination is yes, step S210 is performed, otherwise step S211 is performed.
[0171] S210: determining whether the number of target data packets accumulated lost by the target video frame reaches a first number.
[0172] If the result of the determination is yes, step S208 is performed, otherwise the process returns to step S201.
[0173] S211: determining whether the number of target data packets accumulated lost by the target video frame reaches a second number.
[0174] If the result of the determination is yes, step S208 is performed, otherwise the process returns to step S201.
[0175] The calculation process of the first number comprises: determining a first product of the redundancy corresponding to the target video frame and a first proportion; determining a sum of one and the first product to obtain a first sum value; and determining a product of the total amount of data packets included in the target data packet group and the first sum value to obtain the first number.
[0176] The calculation process of the second number comprises: determining a second product of the redundancy corresponding to the target video frame and a second proportion; determining a sum of one and the second product to obtain a second sum value; and determining a product of the total amount of data packets included in the target data packet group and the second sum value to obtain the second number.
[0177] Specifically, if the target video frame is a video frame in the base layer, the first number will be used, wherein the first number = the total amount of data packets included in the target data packet group * (1 + the redundancy of the target video frame * the first proportion), wherein the total amount of data packets included in the target data packet group is the total amount of data packets obtained after the target video frame is redundantly encoded.
[0178] If the target video frame is a video frame in the enhancement layer, the second number will be used, wherein the second number = the total amount of data packets included in the target data packet group * (1 + the redundancy of the target video frame * the second proportion).
[0179] In an application scenario, the first proportion is greater than the second proportion, so as to preferentially guarantee the demand of network transmission of the base layer.
[0180] It should be noted that in other embodiments, as long as the target data packet is detected to be lost, the identifier corresponding to the target data packet is sent to the sending end, so that the sending end retransmits the lost target data packet. The process of retransmitting the lost target data packet by the sending end has been introduced above, and details can be found in the above related content, which will not be repeated here.
[0181] Wherein, after step S205, the receiving end will first determine a time length t according to the current target video frame receiving time length, decoding time, frame rate and decoding existence time, and determine the time length t as the target time corresponding to the next video frame received. Meanwhile, the second preset time can be equal to the target time, and the second preset time can be one half, one third or two thirds of the target time, etc., which is not limited here.
[0182] Meanwhile, after step S205, it can also be judged whether the receiving end has completed the reception of the GOP where the target video frame is located. If so, the current network bandwidth and packet loss rate can be determined according to the parameters related to the GOP, and the network bandwidth and packet loss rate can be fed back to the sending end, so that the sending end takes the network bandwidth as the target network bandwidth of the next GOP and takes the packet loss rate as the current packet loss rate used by the sending end when sending the next GOP.
[0183] From the above, it can be seen that in the above scheme, the retransmission time is limited and the packet loss retransmission is used for reliable transmission, which can reduce the frame loss caused by video transmission.
[0184] The structure of the sending end and the receiving end in the present application will be introduced as follows: Figure 12
[0185] The sending end 110 includes a network perception unit 111, a redundancy calculation unit 112, an error correction coding unit 113 and a network sending unit 114; the receiving end 120 includes a network receiving unit 121, an error correction decoding unit 122 and a network feedback unit 123.
[0186] In the sending end 110, the network perception unit 111 is used to obtain the packet loss information (specifically, the packet loss rate), network delay, network bandwidth and other information fed back by the receiving end 120; the redundancy calculation unit 112 is used to calculate the redundancy of each video frame in the current GOP (i.e. the first picture group) to be sent; the error correction coding unit 113 performs error correction coding on the corresponding video frame according to the current calculated redundancy and generates block data; the network sending unit 114 sends the block data generated by the error correction coding unit 113 to the receiving end 120 through the network transmission channel.
[0187] In the receiving end 120, the network receiving unit 121 receives data through a network transmission channel, and sends the block data to the error correction decoding unit 122 for data recombination. If the data recombination fails, the network feedback unit 123 notifies the sending end 110; if the data recombination succeeds, the reconstructed data is sent to the user, and at the same time, the network feedback unit 123 of the receiving end measures the data packet loss rate, network delay information, network bandwidth and other information of the sending end 110 in real time and feeds back to the sending end 110.
[0188] From the above, in an application scenario, the following scenarios can exist:
[0189] When the sending end sends the same video stream (including the base layer, the enhancement layer 1 and the enhancement layer 2) to different receiving ends, because the network bandwidths of different receiving ends can be different, the sending end can perform different redundancy reduction processing on the video stream corresponding to different receiving ends, so that the video stream received by the receiving end 1 can only include the base layer and the enhancement layer 1 (i.e., the sending end discards the enhancement layer 2), and the video stream received by the receiving end 2 can only include the base layer (i.e., the sending end discards the enhancement layer 1 and the enhancement layer 2).
[0190] Referring to Figure 13 , Figure 13 is a structural schematic diagram of an embodiment of an electronic device of the present application. The electronic device 200 includes a processor 210, a memory 220 and a communication circuit 230, the processor 210 is coupled to the memory 220 and the communication circuit 230 respectively, the memory 220 stores program data, and the processor 210 implements the steps in the method of any one of the above embodiments by executing the program data in the memory 220, wherein the detailed steps can be referred to the above embodiments, which will not be repeated here.
[0191] The electronic device 200 can be a device for executing the video stream sending method in any one of the above embodiments, or a device for executing the video stream receiving method in any one of the above embodiments, which is not limited here.
[0192] For example, when the electronic device 200 is a device for executing the video stream sending method in any one of the above embodiments, the electronic device 200 can be a camera.
[0193] When the electronic device 200 is a device for executing the video stream receiving method in any one of the above embodiments, the electronic device 200 can be a mobile phone, a computer or any other device.
[0194] Referring to Figure 14 , Figure 14is a structural schematic diagram of an embodiment of the computer readable storage medium of the present application. The computer readable storage medium 400 stores a computer program 410, which can be executed by a processor to implement the steps in any of the above methods.
[0195] The computer readable storage medium 400 can be a U disk, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, or any device that can store the computer program 410, or a server that stores the computer program 410, which can send the stored computer program 410 to other devices for running, or can run the stored computer program 410.
[0196] The above is only an embodiment of the present application, and does not limit the patent scope of the present application. Any equivalent structure or equivalent process transformation using the content of the specification and drawings, or direct or indirect application in other related technical fields, is also included in the patent protection scope of the present application.
Claims
1. A method for sending a video stream, characterized in that, The method includes: Determine the current target network bandwidth; Determine whether the encoded data corresponding to the first frame group in the SVC video stream can be transmitted under the target network bandwidth. The first frame group includes a base layer and at least one enhancement layer. Both the base layer and the enhancement layer include multiple video frames. After encoding the video frames in the first frame group according to the current redundancy of each video frame, the encoded data corresponding to the first frame group is obtained. If transmission is not possible, reduce the current redundancy of the video frames in the at least one enhancement layer; After reducing the current redundancy of video frames in at least one enhancement layer, in response to the fact that the encoded data currently corresponding to the first picture group still cannot be transmitted under the target network bandwidth, the current redundancy of video frames in the base layer is reduced. Each video frame in the first frame group is sequentially encoded according to its current redundancy level and then sent to the receiving end. The video frames in the enhancement layer all have corresponding redundancy; the step of sequentially encoding each video frame in the first frame group according to its current redundancy and then sending it to the receiving end includes: Obtain the first video frame to be sent from the first screen group; Determine whether the first video frame to be sent is a video frame in the base layer or the enhancement layer; In response to the fact that the first video frame to be sent is a video frame in the enhancement layer, the current redundancy of the first video frame to be sent is obtained; Based on the current redundancy of the first video frame to be sent, forward error correction encoding is performed on the first video frame to be sent to obtain multiple first data packets corresponding to the first video frame to be sent. Encapsulate multiple first data packets corresponding to the first video frame to be sent to obtain a target data packet group corresponding to the first video frame to be sent, wherein the target data packet group includes multiple target data packets; Send the target data packet group to the receiving end; The step of obtaining the current redundancy of the first video frame to be sent in response to the first video frame to be sent being a video frame in the enhancement layer includes: in response to the first video frame to be sent being a video frame in the enhancement layer, determining whether there are any discarded video frames in the enhancement layer where the first video frame to be sent is located; if so, discarding the first video frame to be sent and identifying another video frame to be sent after the first video frame to be sent as the first video frame to be sent, and then returning to execute the step of determining whether the first video frame to be sent is a video frame in the base layer or the enhancement layer; if not, obtaining the redundancy of the first video frame to be sent.
2. The method according to claim 1, characterized in that, The step of reducing the current redundancy of video frames in the at least one enhancement layer includes: In the first frame group, a first enhancement layer is determined, wherein the first enhancement layer is the highest enhancement layer in the first frame group; Reduce the current redundancy of video frames in the first enhancement layer, wherein, after reducing the current redundancy of video frames in the first enhancement layer, the current redundancy of each video frame in the first enhancement layer is not lower than the minimum redundancy corresponding to the first enhancement layer. If, after reducing the current redundancy of video frames in the first enhancement layer, the encoded data corresponding to the first picture group still cannot be transmitted under the target network bandwidth, the video frames in the first enhancement layer are discarded, and the second enhancement layer in the first picture group that is located before the first enhancement layer and adjacent to the first enhancement layer is determined as the first enhancement layer. Then, the process returns to the step of reducing the current redundancy of video frames in the first enhancement layer.
3. The method according to claim 2, characterized in that, The step of reducing the current redundancy of video frames in the first enhancement layer includes: A first video frame is determined in the first enhancement layer, wherein the first video frame is arranged last in the first enhancement layer; Reduce the current redundancy of the first video frame; If, after reducing the current redundancy of the first video frame, the encoded data corresponding to the first picture group still cannot be transmitted under the target network bandwidth, the second video frame in the first enhancement layer that is located before the first video frame and adjacent to the first video frame is determined as the first video frame, and then the process of reducing the current redundancy of the first video frame is returned.
4. The method according to claim 1, characterized in that, The step of reducing the current redundancy of video frames in the base layer includes: The first video frame is determined in the base layer; Reduce the current redundancy of the first video frame, wherein, after reducing the redundancy of the first video frame, the current redundancy of the first video frame is not lower than the minimum redundancy corresponding to the base layer. If, after reducing the current redundancy of the first video frame, the encoded data corresponding to the first picture group still cannot be transmitted under the target network bandwidth, the second video frame in the base layer that is located before the first video frame and adjacent to the first video frame is determined as the first video frame, and then the step of reducing the current redundancy of the first video frame is returned to be executed.
5. The method according to claim 4, characterized in that, The step of reducing the current redundancy of video frames in the base layer further includes: When the first video frame is the first video frame in the base layer, in response to the fact that the encoded data corresponding to the first picture group cannot be transmitted under the target network bandwidth, it is determined that the redundancy of each video frame in the base layer is not set.
6. The method according to claim 1, characterized in that, The method further includes: In response to the fact that the first video frame to be sent is a video frame in the base layer, the redundancy of the first video frame to be sent is obtained; If the first video frame to be sent has no redundancy, then the first video frame to be sent is fragmented to obtain multiple first data packets corresponding to the first video frame to be sent; otherwise, forward error correction coding is performed on the first video frame to be sent according to the current redundancy of the first video frame to be sent to obtain multiple first data packets corresponding to the first video frame to be sent.
7. The method according to claim 6, characterized in that, The step of obtaining the redundancy of the first video frame to be sent in response to the fact that the first video frame to be sent is a video frame in the base layer includes: In response to the fact that the first video frame to be sent is a video frame in the base layer, it is determined whether the encoded data corresponding to the first picture group can be transmitted under the target network bandwidth if the current redundancy of each video frame in the base layer is set to the minimum redundancy corresponding to the base layer. If transmission is not possible, the redundancy of the first video frame to be sent is not set; If transmission is possible, the redundancy of the first video frame to be sent is obtained.
8. The method according to claim 1, characterized in that, The step of obtaining the redundancy of the first video frame to be sent if it does not exist includes: If it does not exist, then determine whether the encoded data corresponding to the first picture group can be transmitted in the target network bandwidth after setting the current redundancy of each video frame in the target enhancement layer to the minimum redundancy corresponding to the target enhancement layer. Here, the target enhancement layer is the enhancement layer where the first video frame to be sent is located. If transmission is possible, the redundancy of the first video frame to be sent is obtained; If transmission fails, the first video frame to be sent is discarded, and another video frame to be sent after the first video frame to be sent is identified as the first video frame to be sent. Then, the process returns to the step of determining whether the first video frame to be sent is a video frame in the base layer or the enhancement layer.
9. The method according to claim 1, characterized in that, The method further includes: Upon receiving the target identifier sent by the receiving end, it is determined whether the target data packet corresponding to the target identifier corresponds to the video frame in the base layer or the video frame in the enhancement layer; If it corresponds to the video frame in the base layer, then send M target data packets corresponding to the target identifier to the receiving end; If it corresponds to the video frame in the enhancement layer, then N target data packets corresponding to the target identifier are sent to the receiving end, where M is greater than N; When the receiving end detects that it has not received the target data packet, it sends the target identifier corresponding to the target data packet to the sending end.
10. A method for receiving a video stream, characterized in that, The method includes: The encoded data sent by the sending end is decoded and then played back; wherein the encoded data is sent by the sending end using the method described in any one of claims 1 to 9.
11. The method according to claim 10, characterized in that, The step of decoding and playing the encoded data sent by the sending end includes: During the process of receiving the target data packet group corresponding to the target video frame, it is determined whether there are any lost target data packets; If it exists, then determine whether the target video frame is a video frame in the base layer or a video frame in the enhancement layer; If the target video frame is a video frame in the base layer, then when the number of target data packets lost in the cumulative number of target video frames reaches a first number, the sender is notified of the identifier of the lost target data packets; otherwise, the sender continues to receive the target data packets in the target data packet group. If the target video frame is a video frame in the enhancement layer, then when the number of target data packets lost in the cumulative number of target video frames reaches a second number, the sender is notified of the identifier of the lost target data packets; otherwise, the sender continues to receive the target data packets in the target data packet group. The calculation process for the first quantity includes: determining the first product of the redundancy corresponding to the target video frame and the first ratio; determining the sum of the first product to obtain a first sum value; and determining the product of the total number of data packets included in the target data packet group and the first sum value to obtain the first quantity. The calculation process for the second quantity includes: determining the second product of the redundancy corresponding to the target video frame and the second ratio; determining the sum of the first product and the second product to obtain a second sum value; and determining the product of the total number of data packets included in the target data packet group and the second sum value to obtain the second quantity.
12. The method according to claim 11, characterized in that, Before determining whether the target video frame is a video frame in the base layer or a video frame in the enhancement layer, the method further includes: Determine whether the current time remaining from the first target time point has reached the first preset time; If the target data packet is lost, the sender is notified of its identifier. If the condition is not met, then the step of determining whether the target video frame is a video frame in the base layer or a video frame in the enhancement layer is performed.
13. The method according to claim 11, characterized in that, The step of determining whether there are any lost target data packets during the process of receiving the target data packet group corresponding to the target video frame includes: During the process of receiving the target data packet corresponding to the target video frame, in response to the fact that the target data packet corresponding to the received target video frame has been decoded, the target data packet corresponding to the received target video frame is decoded and played. Otherwise, determine whether there is a lost target data packet.
14. The method according to claim 11, characterized in that, The method further includes: If no target data packet corresponding to the target video frame is received and the time remaining until the second target time point reaches the second preset time, then the sending end is notified to discard the transmission of the target video frame.
15. An electronic device, characterized in that, The electronic device includes a processor, a memory, and a communication circuit. The processor is coupled to the memory and the communication circuit. The memory stores program data. The processor executes the program data in the memory to implement the steps of the method as described in any one of claims 1-14.
16. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that can be executed by a processor to implement the steps of the method as described in any one of claims 1-14.