Streaming media transmission method, electronic equipment and storage medium
By compressing voice frames in a narrowband communication network and combining a broadband transmission protocol to generate data packets containing current and target frames, the problem of large voice delay in the prior art is solved, and the reliability and efficiency improvement of real-time voice transmission is achieved.
Patent Information
- Application Number
- CN202510377692.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-26
- Publication Date
- 2025-07-25
AI Technical Summary
In the prior art, the voice delay based on the loss retransmission strategy is relatively large and cannot meet the needs of real-time voice transmission.
Voice frames are obtained and compressed through a narrowband communication network, and data packets containing the current voice frame and target frame are generated. The broadband network transmission protocol is used to send them to the receiver. The receiver parses and recovers the lost voice frames, and adopts a redundant design to maintain voice continuity and reduce delay.
It effectively solves the continuous packet loss problem caused by the shadow effect, improves the reliability and transmission efficiency of narrowband real-time voice in mobile communication networks, and reduces voice delay.
Smart Images

Figure CN120378412A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of communications, and in particular, to a streaming media transmission method, an electronic device, and a storage medium. Background Art
[0002] When narrowband walkie-talkie voice is transmitted in a mobile communication network, the voice often freezes. Later, it was found through testing that this is caused by packet loss in the 4G network. In related technologies, a loss retransmission strategy is adopted when packets are lost, that is, the sending end marks the characteristics of voice packets, and after the receiving end determines packet loss through the characteristics of voice packets, it requests the sending end to retransmit.
[0003] However, in the prior art, the loss retransmission strategy will generate a relatively large delay when the delay in the mobile communication network is large, which is unacceptable for real-time voice. Summary of the Invention
[0004] Embodiments of this application provide a streaming media transmission method, an electronic device, and a storage medium, so as to at least solve the problem of large voice delay in related technologies based on the loss retransmission strategy.
[0005] According to one aspect of the embodiments of this application, a streaming media transmission method is provided, which is applied to a sending end. The method includes: obtaining a plurality of consecutive voice frames through a narrowband communication network, where each voice frame in the plurality of voice frames is obtained by compressing voice data in the current voice cycle according to a narrowband transmission protocol, each voice frame is the current voice frame, and a set of at least one consecutive voice frame before the current voice frame is the target frame; generating a plurality of data packets according to a broadband network transmission protocol, where each data packet in the plurality of data packets includes the current voice frame and the target frame; and sending the plurality of data packets to a receiving end through a broadband communication network, so that the receiving end parses the received data packets and obtains the voice frames included in each data packet.
[0006] According to another aspect of the embodiments of this application, a streaming media transmission method is further provided, which is applied to a receiving end. The method includes: receiving, through a broadband communication network, a plurality of data packets periodically sent by a sending end, where each data packet in the plurality of data packets includes a current voice frame and a target frame, the current voice frame is a voice frame obtained by compressing voice data in the current voice cycle according to a narrowband transmission protocol, and the target frame is a set of at least one consecutive voice frame before the current voice frame; parsing and obtaining the voice frames included in the plurality of data packets; and sending the obtained voice frames through a narrowband communication network according to a narrowband transmission protocol.
[0007] According to another aspect of the embodiments of the present application, there is also provided a computer-readable storage medium storing a computer program, wherein the computer program is configured to execute the steps in any one of the above method embodiments when running.
[0008] According to another aspect of the embodiments of the present application, there is provided a computer program product or a computer program. The computer program product or the computer program includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the steps in any one of the above method embodiments.
[0009] According to another aspect of the embodiments of the present application, there is also provided an electronic device including a memory and a processor. A computer program is stored in the memory, and the processor is configured to execute the steps in any one of the above method embodiments through the computer program.
[0010] In the embodiments of the present application, each packet sent not only includes the current voice frame to be sent this time, but also includes a specific number of target frames before the current voice frame at least. That is, the target frame is a redundant design of the packet. On the one hand, it is convenient for the receiving end to use the target frame to replace the lost voice frame in the case of random or continuous loss of packets, so as to recover the randomly or continuously lost data, keep the continuity of real-time media data, and improve the reliability of transmitting narrowband real-time voice using a mobile communication network, and can effectively solve the problem of continuous packet loss caused by the shadow effect. On the other hand, when the packet carries the target frame and the packet is lost at the receiving end, there is no need to retransmit the packet, reducing the voice delay in the communication process. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation to the present application. In the drawings:
[0012] Figure 1 is a schematic diagram of an application environment of a streaming media transmission method according to an embodiment of the present application;
[0013] Figure 2 is a schematic flowchart of a streaming media transmission method according to an embodiment of the present application;
[0014] Figure 3 is a schematic diagram of a sending end sending packets according to a voice period according to an embodiment of the present application;
[0015] Figure 4Schematic diagram of another streaming media transmission method according to an embodiment of the present application;
[0016] Figure 5 Schematic diagram of a receiving end receiving data packets according to a voice period according to an embodiment of the present application;
[0017] Figure 6 Schematic diagram of another receiving end receiving data packets according to a voice period according to an embodiment of the present application;
[0018] Figure 7 Schematic diagram of the structure of a streaming media transmission device according to an embodiment of the present application;
[0019] Figure 8 Schematic diagram of the structure of another streaming media transmission device according to an embodiment of the present application. Detailed implementation manners
[0020] In order to enable those skilled in the art to better understand the solution of the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.
[0021] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and do not have to be used to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present application described here can be implemented in an order other than those illustrated or described here. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device including a series of steps or units does not have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0022] According to one aspect of the embodiments of the present application, a streaming media transmission method is provided. Optionally, as an alternative implementation manner, the above-mentioned streaming media transmission method can be but is not limited to being applied to, for example, Figure 1In the application environment shown, the sending end 102 is remotely interconnected with the receiving end 104 via a 4G network. When the walkie-talkie on the sending end 102 speaks, after the sending end 102 receives the voice data of the walkie-talkie, it sends it to the remote receiving end 104 via the 4G network. After receiving the call, the remote receiving end 104 broadcasts it via radio, and the walkie-talkies within its radio coverage can receive the voice call. The sending end 102 and the receiving end 104 can be repeaters adopting PDT (Professional Digital Trunked Radio) / DMR (Digital Mobile Radio Standard).
[0023] The streaming media transmission method of the embodiment of the present application can be executed by the sending end 102. Figure 2 It is a schematic flowchart of an optional streaming media transmission method according to an embodiment of the present application, as Figure 2 shown, the process of this method can include steps S202 to S204.
[0024] Step S202, obtain a plurality of consecutive voice frames through a narrowband communication network. Each voice frame in the plurality of voice frames is obtained by compressing the voice data within the current voice cycle according to the narrowband transmission protocol. Each voice frame is the current voice frame, and the set of at least one consecutive voice frame before the current voice frame is the target frame.
[0025] Among them, a narrowband communication network refers to a network that uses a relatively small bandwidth for communication. Compared with a broadband network, the data transmission rate of a narrowband network is lower, but they can provide a longer transmission distance and lower power consumption, and are especially suitable for remote communication and low-power devices. In a narrowband network, the frequency band resources are more finely divided and managed to support multiple low-rate communication links. The narrowband transmission protocol is a communication protocol designed for efficient data transmission in a narrowband network. For example, commonly used narrowband transmission protocols include PDT (Professional Digital Trunked Radio) and DMR (Digital Mobile Radio), which are used for professional radio communication to ensure clear and reliable voice communication under limited bandwidth.
[0026] The voice cycle refers to the time unit in which the sending end divides the continuous voice data stream at fixed time intervals and compresses the voice data within each time interval into one frame for transmission. For example, under the PDT standard, a duration of 60 ms can be set as a cycle, and the voice within each fixed cycle is compressed into a voice frame. The current voice cycle refers to the time cycle in which the sending end is currently sending voice data. As the sending end sends data packets in sequence, the current voice cycle also changes accordingly.
[0027] A voice frame refers to an independent data unit formed by dividing a continuous voice data stream at time intervals of the voice period and processing the voice data within this time period through compression, encoding, etc. For example, in a TDMA (Time Division Multiple Access) system, a TDMA frame contains at least two time slots, and communication is transmitted on different time slots for different calls. For example, in a PDT system, there are two time slots, a TDMA frame is 60 ms, and a time slot is 30 ms; in the embodiments of the present application, the length of a voice frame can be 30 ms (including more than 30 ms of compressed voice data), and it is sent once every 60 ms, and 6 voice frames form a voice superframe (A - F). The embodiments of the present application do not limit specific narrow - band system types (in addition to the above - mentioned TDMA system, it can also be used in frequency - division multiple access (FDMA) and other wireless communications), but only illustrate the essential transmission method of voice frames. In short, a voice frame is a compressed representation of voice data within a voice period in a voice data stream, and is the basic unit for audio encoding, transmission, and storage. Each voice frame contains compressed data of voice data within a voice period and possible metadata (such as time stamps, encoding information, etc.), and is used for transmission in a communication system or storage on a storage medium. Each voice frame corresponds to an RTP (Real - Time Transport Protocol) sequence number, a time stamp, and a sub - frame number, which are convenient for assisting the receiving end to sort voice frames. For example, in PDT / DMR, the sub - frame numbers E, F, A, B, C, D, or 1, 2, 3, 4, 5, 6 can be directly used to mark continuous voice frames.
[0028] A target frame refers to a specific number of consecutive voice frames arranged in chronological order in a voice data stream and located before the currently processed current voice frame. It can be understood that the target frame must be located before the current voice frame in time, and according to the normal data - stream processing logic, the target frame should have been generated, processed, and sent (or in some cases, may be recorded or stored).
[0029] It can be understood that the more target frames contained in the data packet sent by the sending end, the more voice frames the receiving end can recover. For example, if the data packet only contains 2 target frames (i.e., the first 2 voice frames of the current voice frame), then the receiving end can only recover the lost voice frames when the number of continuously lost data packets is within 2. When the number of lost data packets exceeds 2, the receiving end can only recover some voice frames and cannot recover all the lost voice frames. However, the number of target frames contained in the data packet sent by the sending end is not the more the better. The number of target frames should be associated with the number of consecutive lost packets. Too many target frames will result in a large number of redundant target frames at the receiving end, wasting resources. For example, assume that the number of target frames is set to 5 and the number of consecutive lost packets in actual application is 2. At this time, only 2 target frames need to be included in the data packet to enable the receiving end to recover the lost voice frames. Obviously, setting the number of target frames to 5 will result in a large number of redundant target frames at the receiving end. Generally, the number of target frames is set to 2.
[0030] Specifically, the sending end receives the voice data stream sent by the walkie-talkie, divides the continuous voice data stream according to a fixed voice period, and for the current voice period, compresses the voice data in the current voice period into the current voice frame, and obtains multiple previously sent voice frames before the current voice frame in the buffer as target frames.
[0031] Step S204, generate multiple data packets according to the broadband network transmission protocol, and each data packet among the multiple data packets contains the current voice frame and the target frame.
[0032] Among them, the broadband network transmission protocol refers to a communication protocol used for efficiently and reliably transmitting a large amount of data in a broadband communication network. Common broadband network transmission protocols include TCP / IP (Transmission Control Protocol / Internet Protocol), UDP (User Datagram Protocol), HTTP (Hypertext Transfer Protocol), FTP (File Transfer Protocol), etc. The compression technology in the narrowband transmission protocol is usually optimized to adapt to the low-bandwidth environment and support the operation of low-power devices, such as walkie-talkies. Therefore, in the embodiments of this application, using the narrowband transmission protocol to compress the voice frames can reduce the amount of data required for transmission, thereby saving bandwidth resources, reducing communication costs, and ensuring voice quality at the same time. However, when the compressed voice frames need to be transmitted in a broadband network, using a broadband network transmission protocol (such as TCP / IP, UDP, RTP, etc.) can better utilize the high-bandwidth advantage of the broadband network, improve data transmission efficiency, and ensure the quality of service (QoS). Therefore, in the embodiments of this application, the voice frames are first compressed by the narrowband transmission protocol to ensure efficient encoding and decoding on low-bandwidth devices; then they are transmitted in the broadband network using the broadband transmission protocol, which can make full use of the performance of the high-bandwidth network to ensure high-speed and stable data transmission. This combination takes advantage of the strengths of both protocols and meets the multi-level communication requirements from walkie-talkies to public network transmission.
[0033] A data packet is the basic data unit transmitted in network communication. It consists of a series of bytes and contains the current voice frame to be transmitted within the current voice cycle and multiple consecutive voice frames (i.e., target frames) before the current voice frame.
[0034] Specifically, the sending end compresses each voice frame (i.e., the current voice frame) and the corresponding target frame according to the broadband network transmission protocol to obtain multiple data packets.
[0035] Step S206: Send the multiple data packets to the receiving end through the broadband communication network so that the receiving end can parse the received data packets and obtain the voice frames contained in each data packet.
[0036] After the receiving end receives the data packets sent by the sending end, it parses the received data packets to obtain the voice frames in the data packets. Each voice frame corresponds to an RTP (Real-Time Transport Protocol) sequence number, a timestamp, and a sub-frame number. Therefore, when the RTP (Real-Time Transport Protocol) sequence number, timestamp, or sub-frame number received by the receiving end is not continuous, it is determined that a data packet loss has occurred. At this time, the receiving end can recover the lost voice frames based on the voice frames carried in the received data packets, reduce the consumption of the voice buffer pool caused by transmission packet loss, and ensure voice continuity.
[0037] Specifically, the sending end sends multiple data packets to the receiving end through a broadband communication network. The receiving end periodically receives the data packets sent by the sending end. If the sequence numbers or timestamps of the received data packets are not continuous, it determines that a data packet loss has occurred, determines the lost data packets, obtains the data packets before and / or after the lost data packets, parses the data packets before and / or after the lost data packets to obtain multiple speech frames, and based on the data packets before the lost data packet and the target frames carried in the data packets after the lost data packet, the lost speech frames can be restored.
[0038] For example, Figure 3 FIG. is a schematic diagram of the sending end sending data packets according to a voice period in an embodiment. As Figure 3 shown, for four consecutive data packets, combined with the characteristics of narrowband private network voice transmission (1. The interval between speech frames is long, and once lost, it has a great impact on the voice quality; 2. The voice rate is low and the occupied bandwidth is extremely small). Each voice packet carries the speech of the first two sub-frames and marks the sub-frame information. The sub-frame numbers E, F, A, B, C, D represent consecutive speech frames. Each data packet includes, in addition to the speech frames that should be sent in the current voice period, the speech frames of the first 2 frames before the current speech frame. For example, the last data packet includes three speech frames with sub-frame numbers E, F, A, where speech frame E is the current speech frame, and speech frames F and A are the target frames. When the receiving end finds that there are data packets lost (such as the second and third data packets are lost) and directly receives the data of the last data packet, it directly stores the speech data of the first two sub-frames carried by the last data packet into the speech buffer, eliminates the impact caused by the packet loss of the public mobile communication network, reduces the consumption of the second and third data packets lost on the speech buffer pool, and ensures voice continuity.
[0039] In the embodiment of the present application, based on the current speech frame and the target frame, data packets for the current voice period are generated. When there is a phenomenon of consecutive frame loss, the lost speech frames can still be restored based on the target frames in the data packets before and after the packet loss. That is, the embodiment of the present application adopts a unique redundant transmission method, which can improve the transmission reliability. For example, as Figure 3 shown, even if the second and third data packets are lost, the lost speech frames can still be replaced based on the target frames carried in the first and last data packets, thereby reducing the consumption of the transmission packet loss on the speech buffer pool and ensuring voice continuity.
[0040] In some embodiments, due to the low rate of the private network vocoder, multi-frame transmission has little impact on the 4G / 5G network. Taking PDT as an example, a 60ms data frame only requires 9 bytes of voice data, and even if it is repeated twice, it only requires 27 bytes, which is still less than the protocol overhead (20 bytes for IP header, 8 bytes for UDP header, and at least 12 bytes for RTP header). Therefore, the embodiments of the present application can disperse the transmission payload and reduce network bandwidth consumption.
[0041] In some embodiments, two self-organizing network devices are connected to the switching center via 4G / 5G, and the use effect is verified in static and dynamic scenarios respectively.
[0042] The use effect under the condition of random packet loss on the air interface was verified in a static scenario. The test results showed that the random packet loss rate of the 4G / 5G base station is required to be less than 1%. According to the performance requirements of the vocoder, it has obviously affected the voice quality. After adopting the embodiment of the present application, at least 3 packets of data must be lost continuously to cause real voice loss. The probability is one in a million, which has no obvious impact on the voice.
[0043] The use effect under the condition of continuous air interface packet loss was verified in a dynamic scenario. The test results were as follows: during the movement of the 4G / 5G terminal, the signal strength often changes drastically. In actual measurements, continuous packet loss often occurs (once every 3-5 sentences). After adopting the embodiment of the present application, it was observed through logs that one packet was often lost, and in some cases, two packets were lost continuously, but the voice remained continuous.
[0044] Through the above steps, each data packet sent includes at least a specific number of target frames before the current voice frame in addition to the current voice frame to be sent this time, that is, the target frame is designed as a redundant data packet. On the one hand, it is convenient for the receiving end to use the target frame to replace the lost voice frame when the data packet is randomly or continuously lost, thereby recovering the randomly or continuously lost data, maintaining the continuity of real-time media data, and improving the reliability of transmitting narrowband real-time voice using mobile communication networks. It can effectively solve the problem of continuous packet loss caused by shadow effect. On the other hand, the data packet carries the target frame. When the data packet is lost at the receiving end, there is no need to resend the data packet, which reduces the voice delay in the communication process.
[0045] In an exemplary embodiment, before acquiring a plurality of continuous voice frames through a narrowband communication network, the method further includes:
[0046] Determine a first number; and determine a set of the first number of continuous speech frames before the current speech frame as a target frame.
[0047] Among them, the first quantity refers to a specific, pre-set number used to determine how many consecutive speech frames need to be selected as target frames before the current speech frame. If the system requires high speech recovery ability and fault tolerance, then the "first quantity" may be set larger so that when packets are lost, the receiving end can use more redundant information to recover the lost speech frames. On the contrary, if the system has high requirements for real-time performance, and the network conditions are good and the packet loss rate is low, then the "first quantity" may be set smaller to reduce the redundancy of data packets and transmission delay.
[0048] Specifically, the sending end determines the first quantity according to a specific selection method. During the current speech period, it obtains the first quantity of consecutive speech frames before the current speech frame from the buffer, and determines the first quantity of consecutive speech frames before the current speech frame as the target frames.
[0049] In some embodiments, the sending end can determine the first quantity according to experience or industry standards. It can also determine the first quantity according to experimental test methods. For example, through experimental statistical analysis, the mapping relationship between the packet loss rate and the number of target frames is obtained. In actual application, the sending end first only sends the current speech frame to the receiving end, and the receiving end statistically analyzes the packet loss rate and determines the first quantity according to this mapping relationship.
[0050] In some embodiments, the sending end can also determine the first quantity according to system requirements. For example, if the system has high requirements for real-time performance, it may be necessary to reduce the "first quantity" to reduce processing delay; if the system has high requirements for fault tolerance, it may be necessary to increase the "first quantity" to provide more redundant information to cope with packet loss.
[0051] In this embodiment, by pre-determining the first quantity of target frames, it is convenient to maintain the balance between the packet loss rate at the receiving end and the target frames set by the sending end, and ensure the speech continuity at the receiving end.
[0052] In an exemplary embodiment, determining the first quantity includes:
[0053] Randomly determine a value from a preset range as the first quantity; the upper limit value of the preset range is determined based on the maximum speech delay, and the lower limit value of the preset range is determined based on the minimum number of lost frames.
[0054] Among them, the preset range refers to a numerical interval that defines the set of values that may be used as the number of target frames. The preset range is jointly defined by the upper limit value and the lower limit value, and includes all values from the lower limit value to the upper limit value (including both). The preset range is used to randomly select a value, which is then used as the "first quantity" to determine the number of consecutive speech frames that need to be examined before the current speech frame.
[0055] The upper limit value is the maximum value in the preset range, which limits the maximum value that the "first quantity" can take. The upper limit value is determined based on the maximum voice delay, that is, the maximum voice delay that the receiving end can tolerate is used as a reference to set the upper limit of the "first quantity". This is done to ensure that even in the case of high network latency, the receiving end still has sufficient redundant information to cope with possible voice frame loss or errors.
[0056] In some embodiments, the method for determining the upper limit value can be: determine the maximum voice delay that the receiving end can tolerate, determine the maximum number of data packets sent under the maximum voice delay according to the proportional relationship between the maximum voice delay and the voice period, and use the maximum number as the upper limit of the preset range. Further, after determining the maximum number, the maximum number can be corrected according to the link delay and the intercom's own delay, and the corrected maximum number is used as the upper limit of the preset range.
[0057] For example, in the field of intercoms, the voice delay requirement is not more than 1 second, preferably not more than 500 milliseconds, and the voice period is 60 ms. Then the corresponding number of voice data packets is 16 packets and 8 packets. Therefore, for a delay of 1 second, the corresponding upper limit value is 16; for a delay of 500 milliseconds, the corresponding upper limit value is 8. Further, if some corrections are made according to the link delay and the intercom's own delay, usually three more packets are reduced, that is, 13 packets and 5 packets. In other words, if the maximum voice delay that the receiving end can tolerate is 1 second or 500 milliseconds, then for a delay of 1 second, the corresponding upper limit value is 13; for a delay of 500 milliseconds, the corresponding upper limit value is 5.
[0058] The lower limit value is the minimum value in the preset range, which limits the minimum value that the "first quantity" can take. The lower limit value is determined based on the minimum number of lost frames, that is, the minimum number of lost frames expected by the receiving end is used as a reference to set the lower limit of the "first quantity". The purpose of setting the lower limit value may be to minimize unnecessary data redundancy while ensuring a certain degree of fault tolerance, so as to improve the efficiency and real-time performance of the system. Generally, the minimum number of lost frames is 1. Therefore, the lower limit value is generally 1. Thus, combining the above content, the preset range can preferably be [1, 5].
[0059] Specifically, the sending end determines the maximum number of data packets sent under the maximum voice delay according to the proportional relationship between the maximum voice delay and the voice period, and uses the maximum number as the upper limit of the preset range. Further, after determining the maximum number, the maximum number can be corrected according to the link delay and the intercom's own delay, and the corrected maximum number is used as the upper limit of the preset range. The sending end uses the minimum number of lost frames as the lower limit of the preset range, and determines the preset range according to the upper limit value and the lower limit value. The sending end randomly determines a value from the preset range as the first quantity.
[0060] In this embodiment, the upper limit value of the preset range is determined based on the maximum voice delay and voice period, the lower limit value of the preset range is determined based on the minimum number of lost packets, and a value is randomly determined from the preset range as the first quantity. By considering the maximum voice delay to determine the upper limit value, the voice transmission delay can be reduced, and by ensuring the minimum number of lost packets to determine the lower limit value, the voice quality of the voice transmission can be maintained.
[0061] In an exemplary embodiment, the first quantity is determined randomly in the above embodiment. This method has the problem that the determined first quantity cannot ensure that the packet loss rate at the receiving end meets the requirements. For example, assume that the first quantity is determined to be 1, that is, in addition to the current voice frame in the data packet, it also includes the previous voice frame (i.e., the target frame) of the current voice frame. If the receiving end continuously loses 3 data packets, then the receiving end can only recover one frame and cannot fully recover the lost voice frames, resulting in the packet loss rate at the receiving end still not meeting the requirements. Therefore, to solve the above problem, in this embodiment, determining the first quantity includes:
[0062] 1. Obtain multiple packet loss events within the historical time period; different packet loss events in the multiple packet loss events correspond to different numbers of lost packets.
[0063] Among them, a packet loss event refers to an event in which a data packet fails to be successfully transmitted from the sending end to the receiving end due to various reasons during network transmission and is thus discarded or lost.
[0064] The number of lost packets refers to the number of data packets actually lost in a certain packet loss event. Different packet loss events may correspond to different numbers of lost packets. For example, at a certain time point, due to network congestion, only a small number of data packets may be discarded; while at another time point, due to equipment failure, a large number of data packets may be lost.
[0065] Specifically, the sending end collects or queries all packet loss events that occurred in the past period of time, statistically summarizes the extracted packet loss events, and calculates the number of lost packets for each event.
[0066] 2. Determine the occurrence probability of each packet loss event among the multiple packet loss events.
[0067] Specifically, the sending end counts the number of times each number of lost packets appears and divides it by the total number of events to calculate the occurrence probability of each number of lost packets (i.e., each packet loss situation).
[0068] 3. Determine the number of lost packets corresponding to the packet loss event with the highest occurrence probability among the multiple packet loss events as the first quantity.
[0069] Specifically, after the sending end determines the occurrence probabilities of various numbers of lost packets, it selects the number of lost packets with the highest occurrence probability as the "first quantity".
[0070] In this embodiment, the number of lost packets corresponding to the lost packet event with the highest occurrence probability is used as the first quantity, that is, the first quantity represents the lost packet situation that occurs most frequently during this time period. The first quantity determined by this method is more in line with the actual situation. Therefore, the determined first quantity can ensure that the packet loss rate at the receiving end meets the requirements and significantly reduce the packet loss rate in the network.
[0071] In an exemplary embodiment, after determining the first quantity, determining the target frame based on the first quantity, and generating data packets based on the target frame, the above method further includes:
[0072] In the case of receiving an adjustment notice fed back by the receiving end, according to the adjustment notice, adjust the first quantity to obtain a second quantity, and determine the second quantity of consecutive speech frames before the current speech frame as the target frame.
[0073] Among them, the adjustment notice indicates that the packet loss rate at the receiving end does not meet the requirements. For example, the packet loss rate is too high, or the packet loss rate is very small but there are too many redundant target frames. The adjustment notice is used to instruct the sending end to adjust the first quantity, and the adjustment strategy can be to increase the first quantity or to decrease the first quantity. It can be understood that the adjustment notice is sent after the receiving end receives the data packet and calculates the packet loss rate.
[0074] The second quantity is the quantity obtained after adding / subtracting based on the first quantity. It should be noted that the second quantity needs to meet the requirement of being within the preset range in the above embodiment. As can be seen from the above embodiment, the preset range is a set of values that may be used as the number of target frames. Therefore, the adjusted second quantity should also be within the preset range.
[0075] Specifically, after the receiving end receives the data packet, it calculates the packet loss rate. When the packet loss rate meets the requirements, it does not perform any processing; when the packet loss rate does not meet the requirements, it sends an adjustment notice to the sending end. After the sending end receives the adjustment notice fed back by the receiving end, it adjusts the first quantity according to the adjustment notice to obtain a second quantity. The sending end takes the next speech cycle as the current speech cycle, determines the second quantity of consecutive speech frames before the current speech frame as the target frame, generates the data packet for the current speech cycle with the current speech frame and the re-determined target frame, and sends the data packet to the receiving end. The receiving end repeats the step of calculating the packet loss rate.
[0076] In this embodiment, the receiving end monitors the packet loss rate in real time and immediately sends an adjustment notice to the sending end when the requirements are not met. This immediate feedback mechanism allows the sending end to quickly respond to changes in network conditions to match the current network situation, and by dynamically adjusting the number of target frames, it can reduce packet loss and delay caused by network problems to a certain extent, thereby improving the fluency and clarity of voice communication.
[0077] In an exemplary embodiment, adjusting the first quantity according to an adjustment notification to obtain a second quantity includes:
[0078] When the adjustment notification indicates that the packet loss rate is greater than a first preset threshold, increasing the first quantity to obtain the second quantity.
[0079] Wherein, the first preset threshold refers to the upper limit value of the packet loss rate. The first preset threshold can be determined according to actual requirements or data statistics. For example, the first preset threshold can be set to one-thousandth. The packet loss rate being greater than the first preset threshold means that the number of target frames sent by the sending end is too small, resulting in the receiving end being unable to recover all the lost voice frames based on the target frames sent by the sending end, thereby causing the packet loss rate to be greater than the first preset threshold.
[0080] Specifically, after the sending end receives the adjustment notification, when the adjustment notification indicates that the packet loss rate is greater than the first preset threshold, increasing the first quantity according to a fixed step value or other methods to obtain the second quantity. The sending end takes the next voice cycle as the current voice cycle and determines the second quantity of consecutive voice frames before the current voice frame as the target frames.
[0081] In this embodiment, when the adjustment notification indicates that the packet loss rate is greater than the first preset threshold, increasing the first quantity to obtain the second quantity can increase the number of target frames in the data packets sent by the sending end, facilitating the receiving end to recover all the lost voice frames based on the target frames sent by the sending end and reducing the packet loss rate.
[0082] In an exemplary embodiment, the above method further includes:
[0083] When the adjustment notification indicates that the packet loss rate is less than a second preset threshold, decreasing the first quantity to obtain the second quantity.
[0084] Wherein, the second preset threshold refers to the lower limit value of the packet loss rate. The second preset threshold can be determined according to actual requirements or data statistics. For example, the second preset threshold can be set to one-ten-thousandth. It can be understood that when the packet loss rate is between the first preset threshold and the second preset threshold, it is determined that the packet loss rate meets the requirements and the receiving end does not perform any processing.
[0085] The packet loss rate being less than the second preset threshold means that the number of target frames sent by the sending end is too large, resulting in a large number of redundant target frames at the receiving end.
[0086] Specifically, after the sending end receives the adjustment notification, when the adjustment notification indicates that the packet loss rate is less than the second preset threshold, decreasing the first quantity according to a fixed step value or other methods to obtain the second quantity. The sending end takes the next voice cycle as the current voice cycle and determines the second quantity of consecutive voice frames before the current voice frame as the target frames.
[0087] In this embodiment, when the adjusted notification representation of the packet loss rate is less than the second preset threshold, the first quantity is reduced to obtain the second quantity, which can reduce the number of target frames in the data packets sent by the sending end, thereby reducing the number of redundant target frames received by the receiving end and reducing resource waste.
[0088] According to another aspect of the embodiments of the present application, there is also provided a streaming media transmission method, which is applied to Figure 1 the receiving end in Figure 4 FIG. is a schematic flowchart of an optional streaming media transmission method according to an embodiment of the present application. As Figure 4 shown, the process of this method may include steps S402 to S406.
[0089] Step S402, receiving, through a broadband communication network, multiple data packets periodically sent by a sending end, where each of the multiple data packets contains a current voice frame and a target frame. The current voice frame is a voice frame obtained by compressing voice data within the current voice cycle according to a narrowband transmission protocol, and the target frame is a set of at least one consecutive voice frame before the current voice frame.
[0090] Among them, the broadband communication network, the current voice cycle, the data packet, the current voice frame, the target frame, and the narrowband transmission protocol have all been explained in the above embodiments, and therefore will not be elaborated here.
[0091] Step S404, parsing and obtaining the voice frames included in the multiple data packets.
[0092] Specifically, if the sequence numbers or timestamps of the data packets received by the receiving end are not continuous, it is determined that a data packet loss has occurred, and the lost data packets are determined. The data packets before and / or after the lost data packets are obtained, and the data packets before and / or after the lost data packets are parsed to obtain multiple voice frames. According to the target frames carried in the data packets before and / or after the lost data packets, the lost voice frames can be restored.
[0093] For example, the sending end sends data packets according to the voice cycle as shown in Figure 3 shown. As shown in Figure 3 shown, the second data packet and the third data packet are interfered and lost. Figure 5 FIG. is a schematic diagram of the receiving end receiving data packets according to the voice cycle in an embodiment. As Figure 5As shown in the figure, after the second and third data packets sent by the sending end are interfered and lost, the receiving end can only receive the first data packet (including three voice frames with sub-frame numbers B, C, and D) and the last data packet (including three voice frames with sub-frame numbers E, F, and A), and perform parsing and caching. Six voice frames with sub-frame numbers E, F, A, B, C, and D can be obtained. However, the sub-frame numbers E, F, A, B, C, and D represent consecutive voice frames. Therefore, after sorting the six parsed voice frames, complete and consecutive voice frames can be obtained. At this time, the receiving end can restore the complete voice frames according to the target frames in the first and last data packets, ensuring voice continuity. In summary, if consecutive packet losses occur during 4G data transmission, the processing process is as follows Figure 5 As shown in the figure, by using interleaved retransmission and sub-frame sequence numbers, when two consecutive data packets are lost, the complete voice data can still be successfully restored using this solution.
[0094] Step S406: Send the obtained voice frames through the narrowband communication network according to the narrowband transmission protocol.
[0095] Specifically, the receiving end sorts the parsed voice frames according to the frame numbers of the voice frames and sends the sorted voice frames according to the narrowband transmission protocol.
[0096] Through the above steps, each data packet sent each time contains at least a specific number of target frames before the current voice frame in addition to the current voice frame to be sent this time. That is, the target frame is used as the redundancy design of the data packet. On the one hand, it is convenient for the receiving end to use the target frame to replace the lost voice frame in the case of random or consecutive loss of data packets, so as to restore the randomly or continuously lost data, maintain the continuity of real-time media data, and improve the reliability of transmitting narrowband real-time voice using the mobile communication network, and can effectively solve the problem of consecutive packet losses caused by the shadow effect. On the other hand, when the data packet carries the target frame and the data packet is lost at the receiving end, there is no need to retransmit the data packet, reducing the voice delay in the communication process.
[0097] In an exemplary embodiment, parsing and obtaining the voice frames included in multiple data packets includes:
[0098] First, each of the multiple data packets is used as the current data packet. In the case where there are data packets lost before or after the current data packet, the current data packet is parsed to obtain multiple voice frames; each voice frame in the multiple voice frames is assigned an identifier; the identifier is used to represent the voice period corresponding to the voice frame.
[0099] Among them, an identifier is a kind of label or code used to uniquely identify or distinguish different speech frames. For example, in PDT / DMR, the sub-frame numbers E, F, A, B, C, D, or 1, 2, 3, 4, 5, 6 can be directly used to label consecutive speech frames.
[0100] Specifically, the receiving end sorts the received multiple data packets according to the time sequence of the speech cycle indicated by the identifier of the data packet, obtains the sorted multiple data packets, takes each data packet in the sorted multiple data packets as the current data packet, determines whether the sequence numbers or timestamps of the data packets before and after the current data packet are continuous. If not, it determines that a data packet loss occurs before or after the current data packet, determines the lost data packet, and parses the current data packet to obtain multiple speech frames.
[0101] Second, eliminate the speech frames with duplicate identifiers among the multiple speech frames, sort the eliminated speech frames according to the speech cycle represented by the identifier, and based on the sorted speech frames, recover the lost speech frames.
[0102] Specifically, after the receiving end parses the data packets before and after the loss to obtain multiple speech frames, it eliminates the speech frames with duplicate identifiers, sorts the eliminated speech frames in the order of the speech cycle represented by the identifier, and based on the sorted speech frames, recovers the lost speech frames.
[0103] For example, Figure 6 is a schematic diagram of the receiving end receiving data packets according to the speech cycle in another embodiment, as Figure 6As shown, for four consecutive data packets, the sub-frame numbers E, F, A, B, C, D, H represent consecutive speech frames. Each data packet contains, in addition to the speech frames to be sent in the current speech cycle, the speech frames of the first 3 frames of the current speech frame. For example, the last data packet contains four speech frames with sub-frame numbers E, F, A, B, where speech frame E is the current speech frame, and speech frames F, A, and B are target frames. When interference packet loss occurs in the second and third data packets sent by the sending end, the receiving end can only receive the first data packet (containing four speech frames with sub-frame numbers B, C, D, H) and the last data packet (containing four speech frames with sub-frame numbers E, F, A, B). If the current data packet is the first data packet (containing four speech frames with sub-frame numbers B, C, D, H) and the packet loss occurs after it, then the three speech frames B, C, D in the second data packet can be recovered based on the first data packet. If the current data packet is the last data packet (containing four speech frames with sub-frame numbers E, F, A, B) and the packet loss occurs before it, then the three speech frames F, A, B in the third data packet can be recovered based on the last data packet. To ensure complete recovery of the lost speech frames, in this embodiment, the receiving end determines the lost data packets based on the sequence number or timestamp of the received data packets, and uses the data packet before the packet loss (the first data packet) and the data packet after the packet loss (the last data packet) for parsing and caching, and eight speech frames with sub-frame numbers E, F, A, B, C, D, H can be obtained, where there are two speech frames B. However, the sub-frame numbers E, F, A, B, C, D, H represent consecutive speech frames. Therefore, after sorting the eight parsed speech frames and removing one redundant speech frame B, a complete and consecutive set of speech frames can be obtained. At this time, the receiving end can recover the complete speech frames based on the target frames in the first and last data packets, ensuring speech continuity.
[0104] It should be noted that: if there are no speech frames with duplicate identifiers, there is no need to remove them.
[0105] In this embodiment, by assigning identifiers to each speech frame, after removing the speech frames with duplicate identifiers among multiple speech frames, and sorting according to the speech cycles represented by these identifiers, it can be ensured that even in the case of data packet loss, the playback order of the speech frames can remain correct, thereby enhancing the continuity of the speech.
[0106] In a detailed embodiment, a streaming media transmission method includes the following steps:
[0107] First, the sending end randomly determines a value as the first quantity from a preset range; the upper limit value of the preset range is determined based on the maximum speech delay, and the lower limit value of the preset range is determined based on the minimum number of lost packets;
[0108] Alternatively, the sending end obtains multiple packet loss events within a historical time period; different packet loss events among the multiple packet loss events correspond to different packet loss quantities; determines the occurrence probability of each packet loss event among the multiple packet loss events; and determines the packet loss quantity corresponding to the packet loss event with the highest occurrence probability among the multiple packet loss events as the first quantity.
[0109] Second, the sending end determines the first quantity of consecutive speech frames before the current speech frame as the target frames.
[0110] Third, the sending end obtains multiple consecutive speech frames through a narrowband communication network. Each speech frame among the multiple speech frames is obtained by compressing the speech data within the current speech cycle according to the narrowband transmission protocol. Each speech frame is the current speech frame, and the set of at least one consecutive speech frame before the current speech frame is the target frame.
[0111] Fourth, the sending end generates multiple data packets according to the broadband network transmission protocol. Each data packet among the multiple data packets contains the current speech frame and the target frames; and sends the multiple data packets to the receiving end through the broadband communication network so that the receiving end parses the received data packets and obtains the speech frames contained in each data packet.
[0112] Fifth, the receiving end receives multiple data packets periodically sent by the sending end through the broadband communication network. Among them, each data packet among the multiple data packets contains the current speech frame and the target frames. The current speech frame is a speech frame obtained by compressing the speech data within the current speech cycle according to the narrowband transmission protocol, and the target frames are the set of at least one consecutive speech frame before the current speech frame.
[0113] Sixth, the receiving end takes each data packet among the multiple data packets as the current data packet. In the case of packet loss of the data packet before or after the current data packet, parses the current data packet to obtain multiple speech frames; each speech frame among the multiple speech frames is assigned an identifier; and the identifier is used to represent the speech cycle corresponding to the speech frame.
[0114] Seventh, the receiving end eliminates the speech frames with duplicate identifiers among the multiple speech frames, sorts the eliminated speech frames according to the speech cycles represented by the identifiers, and restores the lost speech frames based on the sorted speech frames, and sends the obtained speech frames through the narrowband communication network according to the narrowband transmission protocol.
[0115] Eighth, the receiving end calculates the packet loss rate. When the packet loss rate does not meet the requirements, it feeds back an adjustment notice to the sending end.
[0116] 9. When the sending end receives the adjustment notice fed back by the receiving end, if the adjustment notice indicates that the packet loss rate is greater than the first preset threshold, increase the first quantity to obtain the second quantity; if the adjustment notice indicates that the packet loss rate is less than the second preset threshold, decrease the first quantity to obtain the second quantity.
[0117] In this embodiment, each data packet sent each time includes, in addition to the current voice frame to be sent this time, at least a specific number of target frames before the current voice frame. That is, the target frame is a redundant design of the data packet. On the one hand, it is convenient for the receiving end to use the target frame to replace the lost voice frame in the case of random or continuous loss of data packets, so as to restore the randomly or continuously lost data, keep the continuity of real-time media data, and improve the reliability of transmitting narrowband real-time voice using the mobile communication network, and can effectively solve the problem of continuous packet loss caused by the shadow effect. On the other hand, when the data packet carries the target frame, in the case of data packet loss at the receiving end, there is no need to retransmit the data packet, reducing the voice delay in the communication process.
[0118] According to another aspect of the embodiments of the present application, a streaming media transmission device is further provided. This device can be used to implement the streaming media transmission method provided in the above embodiments, and those that have been described will not be repeated. As used below, the term "module" can be a combination of software and / or hardware that can achieve a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, implementation in hardware, or a combination of software and hardware is also possible and contemplated.
[0119] Figure 7 is a structural block diagram of an optional streaming media transmission device according to the embodiments of the present application. As Figure 7 shown in, the device includes:
[0120] An acquisition module 701, configured to acquire a plurality of consecutive voice frames through a narrowband communication network. Each voice frame in the plurality of voice frames is obtained by compressing the voice data within the current voice cycle according to the narrowband transmission protocol. Each voice frame is the current voice frame, and a set of at least one consecutive voice frame before the current voice frame is the target frame;
[0121] A sending module 702, configured to generate a plurality of data packets according to the broadband network transmission protocol. Each data packet in the plurality of data packets includes the current voice frame and the target frame; and send the plurality of data packets to the receiving end through the broadband communication network, so that the receiving end can parse the received data packets and obtain the voice frames included in each data packet.
[0122] In an exemplary embodiment, the acquisition module 701 is further configured to determine a first quantity before acquiring a plurality of consecutive voice frames through the narrowband communication network; and determine a set of the first quantity of consecutive voice frames before the current voice frame as the target frame.
[0123] In an exemplary embodiment, the obtaining module 701 is further configured to randomly determine a value within a preset range as the first quantity; the upper limit value of the preset range is determined based on the maximum voice delay, and the lower limit value of the preset range is determined based on the minimum number of lost frames.
[0124] In an exemplary embodiment, the obtaining module 701 is further configured to obtain a plurality of packet loss events within a historical time period; different packet loss events among the plurality of packet loss events correspond to different numbers of lost packets; determine the occurrence probability of each packet loss event among the plurality of packet loss events; and determine the number of lost packets corresponding to the packet loss event with the highest occurrence probability among the plurality of packet loss events as the first quantity.
[0125] In an exemplary embodiment, the obtaining module 701 is further configured to, when receiving an adjustment notification fed back by the receiving end, adjust the first quantity according to the adjustment notification to obtain a second quantity, and determine the continuous voice frames of the second quantity before the current voice frame as the target frames.
[0126] In an exemplary embodiment, the obtaining module 701 is further configured to, when the adjustment notification indicates that the packet loss rate is greater than a first preset threshold, increase the first quantity to obtain a second quantity.
[0127] In an exemplary embodiment, the obtaining module 701 is further configured to, when the adjustment notification indicates that the packet loss rate is less than a second preset threshold, decrease the first quantity to obtain a second quantity.
[0128] Figure 8 is a structural block diagram of another optional streaming media transmission device according to an embodiment of the present application, as Figure 8 shown in, the device includes:
[0129] A receiving module 801, configured to receive a plurality of data packets periodically sent by a sending end through a broadband communication network, where each data packet among the plurality of data packets includes a current voice frame and target frames, the current voice frame is a voice frame obtained by compressing voice data within the current voice period according to a narrowband transmission protocol, and the target frames are a set of at least one continuous voice frame before the current voice frame;
[0130] A recovery module 802, configured to parse and obtain the voice frames included in the plurality of data packets; and send the obtained voice frames through a narrowband communication network according to the narrowband transmission protocol.
[0131] In an exemplary embodiment, the recovery module 802 is further configured to use each data packet among a plurality of data packets as a current data packet, and in the case of loss of data packets before or after the current data packet, parse the current data packet to obtain a plurality of speech frames; each speech frame among the plurality of speech frames is assigned an identifier; the identifier is used to represent the speech period corresponding to the speech frame; discard the speech frames with duplicate identifiers among the plurality of speech frames, sort the discarded speech frames according to the speech period represented by the identifier, and recover the lost speech frames based on the sorted speech frames.
[0132] The serial numbers of the embodiments of the present application above are only for description and do not represent the advantages or disadvantages of the embodiments.
[0133] If the integrated unit in the above embodiments is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in the above computer-readable storage medium. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in the storage medium and includes several instructions for causing one or more computer devices (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods in the various embodiments of the present application.
[0134] In the above embodiments of the present application, the descriptions of the various embodiments have their own emphases. For the parts not detailed in a certain embodiment, reference may be made to the relevant descriptions of other embodiments.
[0135] In several embodiments provided by the present application, it should be understood that the disclosed client can be implemented in other ways. Among them, the device embodiments described above are only illustrative. For example, the division of units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces. The indirect couplings or communication connections of units or modules can be in an electrical or other form.
[0136] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0137] In addition, in each embodiment of the present application, each functional unit may be integrated into one processing unit, may exist physically alone for each unit, or two or more units may be integrated into one unit. The above integrated unit may be implemented in the form of hardware or in the form of a software functional unit.
[0138] The above are only the preferred embodiments of the present application. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present application, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present application.
Claims
1. A streaming media transmission method, characterized in that, Applied to the sending end, the method includes: Obtain a plurality of consecutive speech frames through a narrowband communication network. Each speech frame in the plurality of speech frames is obtained by compressing the speech data within the current speech period according to the narrowband transmission protocol. Each of the speech frames is the current speech frame, and the set of at least one consecutive speech frame before the current speech frame is the target frame; Generate a plurality of data packets according to the broadband network transmission protocol. Each data packet in the plurality of data packets contains the current speech frame and the target frame; Send the plurality of data packets to the receiving end through a broadband communication network, so that the receiving end parses the received data packets and obtains the speech frames contained in each data packet.
2. The method according to claim 1, characterized in that Before obtaining the plurality of consecutive speech frames through the narrowband communication network, the method further includes: Determine a first quantity; Determine the set of the first quantity of consecutive speech frames before the current speech frame as the target frame.
3. The method according to claim 2, characterized in that, The determining the first quantity includes: Randomly determine a value from a preset range as the first quantity; the upper limit value of the preset range is determined based on the maximum speech delay, and the lower limit value of the preset range is determined based on the minimum number of lost frames.
4. The method according to claim 2, characterized in that The determining the first quantity includes: Obtain a plurality of packet loss events within a historical time period; different packet loss events in the plurality of packet loss events correspond to different numbers of lost packets; Determine the occurrence probability of each packet loss event in the plurality of packet loss events; Determine the number of lost packets corresponding to the packet loss event with the highest occurrence probability among the plurality of packet loss events as the first quantity.
5. The method according to claim 2, characterized in that, The method further includes: In the case of receiving an adjustment notification feedback from the receiving end, according to the adjustment notification, adjust the first quantity to obtain a second quantity, and determine the set of the second quantity of consecutive speech frames before the current speech frame as the target frame.
6. The method according to claim 5, wherein The adjusting the first quantity according to the adjustment notification to obtain a second quantity includes: In the case where the adjustment notification indicates that the packet loss rate is greater than a first preset threshold, increase the first quantity to obtain a second quantity.
7. The method according to claim 6, characterized in that, The method further includes: In the case where the adjustment notification indicates that the packet loss rate is less than a second preset threshold, decrease the first quantity to obtain a second quantity.
8. A streaming media transmission method, characterized in that, Applied to the receiving end, the method includes: Receive a plurality of data packets periodically sent by the sending end through a broadband communication network. Among them, each data packet in the plurality of data packets contains a current speech frame and a target frame. The current speech frame is a speech frame obtained by compressing the speech data within the current speech period according to the narrowband transmission protocol, and the target frame is the set of at least one consecutive speech frame before the current speech frame; Parse and obtain the speech frames contained in the plurality of data packets; Send the obtained speech frames through the narrowband communication network according to the narrowband transmission protocol.
9. The method according to claim 8, wherein The method further includes: Take each of the multiple data packets as the current data packet. In the case of packet loss before or after the current data packet, parse the current data packet to obtain multiple speech frames; each speech frame in the multiple speech frames is assigned an identifier; the identifier is used to represent the speech period corresponding to the speech frame. Eliminate the speech frames with duplicate identifiers in the multiple speech frames, sort the eliminated speech frames according to the speech period represented by the identifiers, and recover the lost speech frames based on the sorted speech frames.
10. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the method described in any one of claims 1 to 7, or implements the steps of the method described in any one of claims 8 to 9.
11. A computer-readable storage medium, characterized in that, A computer program is stored in the computer-readable storage medium, wherein when the computer program is executed by a processor, it implements the steps of the method described in any one of claims 1 to 7, or implements the steps of the method described in any one of claims 8 to 9.