Data transmission method, transmitting end and receiving end

By layering audio and video frames, the basic layer frames are transmitted using low-latency, low-bandwidth paths to ensure video continuity and low latency, while the enhancement layer frames are transmitted using high-bandwidth paths. This solves the problem of video quality degradation caused by network anomalies and improves the overall quality of audio and video data transmission and user experience.

CN120915952AActive Publication Date: 2025-11-07SHENZHEN HDCVT TECH CO LTD

Patent Information

Application Number
CN202511447099.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-11
Publication Date
2025-11-07
Estimated Expiration
2045-10-11

AI Technical Summary

Technical Problem

Existing AVOIP technology is prone to keyframe loss during sudden network anomalies, resulting in pixelation, blurry images, or audio stuttering, which affects the quality of audio and video data transmission and user experience.

Method used

By performing layered processing on video frames, the video frames are divided into a base layer and an enhancement layer. The base layer frames are transmitted using the first path, which has low latency and low bandwidth, to ensure video continuity and low latency. The enhancement layer frames are transmitted using the second path, which has high bandwidth, to improve image quality and detail and dynamically adapt to network conditions.

Benefits of technology

It achieves the goal of ensuring video continuity and basic quality even during network anomalies, while improving image details, maximizing the use of network resources, avoiding overall quality degradation caused by single-path congestion, and providing a stable, high-quality audio and video experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120915952A_ABST
    Figure CN120915952A_ABST
Patent Text Reader

Abstract

The invention discloses a data transmission method, a sending end and a receiving end, and relates to the field of audio and video signal transmission, and the method comprises the steps: carrying out the coding of original data, and obtaining a plurality of video frames; according to the video feature of each video frame, determining a data hierarchy to which each video frame belongs, the data hierarchy comprising a base layer and an enhancement layer; sending a video frame corresponding to the base layer to a receiving end through the first path; the video frame corresponding to the enhancement layer is sent to a receiving end through a second path, the time delay of the first path is lower than that of the second path, and the bandwidth of the first path is lower than that of the second path. According to the invention, through layered transmission, low-delay and high-priority delivery of the base layer video is guaranteed preferentially, and the quality of data transmission is improved on the whole.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of audio and video signal transmission, in particular to a data transmission method, a sending end and a receiving end. BACKGROUND

[0002] The current AVoIP (Audio-Visual over Internet Protocol) technology mainly dynamically selects the optimal path for audio and video data transmission according to the real-time detection of network bandwidth and time delay, so as to realize smooth and low-delay transmission.

[0003] However, when a network anomaly occurs, key frame loss and other situations are likely to occur, which may cause subsequent prediction frames to be unable to be correctly restored and decoded, resulting in mosaic, blurred pictures or audio lag in the finally restored video picture, and thus causing the quality of audio and video data transmission to decrease and the user experience to decrease. SUMMARY

[0004] The main purpose of the present application is to provide a data transmission method, a sending end and a receiving end, aiming to solve the technical problem of how to improve the transmission quality of data transmission.

[0005] To achieve the above-mentioned purpose, the present application provides a data transmission method, which is applied to a sending end, and the method comprises: encoding original data to obtain a plurality of video frames; determining the data level to which each of the video frames belongs according to the video features of each of the video frames, wherein the data level comprises a basic layer and an enhanced layer; sending the video frames corresponding to the basic layer to a receiving end through a first path; sending the video frames corresponding to the enhanced layer to the receiving end through a second path, wherein the time delay of the first path is lower than that of the second path, and the bandwidth of the first path is lower than that of the second path.

[0006] In addition, to achieve the above-mentioned purpose, the present application also provides a data transmission method, which is applied to a receiving end, and the method comprises: receiving a plurality of video frames sent by a client through a preset first path and a preset second path, wherein the video frames are obtained by encoding original data by the client, the first path is used for sending video frames with a basic layer as the data level, the second path is used for sending video frames with an enhanced layer as the data level, the data level of each of the video frames is determined according to the video features of each of the video frames, the time delay of the first path is lower than that of the second path, and the bandwidth of the first path is lower than that of the second path; Reorganize and decode each of the video frames to obtain a restored video.

[0007] In addition, to achieve the above object, the application further provides a data transmission device, which is applied to a sending end and comprises: a coding module, configured to code original data to obtain a plurality of video frames; a layering module, configured to determine a data layer level to which each of the video frames belongs according to a video feature of each of the video frames, wherein the data layer level comprises a base layer and an enhancement layer; a first sending module, configured to send the video frame corresponding to the base layer to a receiving end through a first path; a second sending module, configured to send the video frame corresponding to the enhancement layer to the receiving end through a second path, wherein the time delay of the first path is lower than that of the second path, and the bandwidth of the first path is lower than that of the second path.

[0008] In addition, to achieve the above object, the application further provides a data transmission device, which is applied to a receiving end and comprises: a receiving module, configured to receive a plurality of video frames sent by a client through a preset first path and a preset second path, wherein the video frames are obtained by coding original data by the client, the first path is used to send a video frame with a data layer level of a base layer, the second path is used to send a video frame with a data layer level of an enhancement layer, the data layer level of each of the video frames is determined according to a video feature of each of the video frames, the time delay of the first path is lower than that of the second path, and the bandwidth of the first path is lower than that of the second path; a decoding module, configured to reorganize and decode each of the video frames to obtain a restored video.

[0009] In addition, to achieve the above object, the application further provides a sending end, which comprises a memory, a processor and a computer program stored in the memory and executable on the processor, and the computer program is configured to implement the steps of the data transmission method as described above.

[0010] In addition, to achieve the above object, the application further provides a receiving end, which comprises a memory, a processor and a computer program stored in the memory and executable on the processor, and the computer program is configured to implement the steps of the data transmission method as described above.

[0011] In addition, to achieve the above object, the application further provides a storage medium, which is a computer readable storage medium, and the storage medium stores a computer program, and the computer program is executed by a processor to implement the steps of the data transmission method as described above.

[0012] In addition, to achieve the above object, the application further provides a computer program product, which comprises a computer program, and the computer program, when executed by a processor, implements the steps of the data transmission method as described above.

[0013] The one or more technical solutions provided by the application have at least the following technical effects: first, the original data is encoded by the sending end to convert the continuous media stream into a discrete frame sequence that is convenient for network transmission, so as to be processed in layers; then, the data level to which each video frame belongs is determined according to the video features of each video frame, so as to realize the importance layering of the video content, and to realize the differentiated transmission of the video frames of different data levels; then, the video frames corresponding to the base layer are sent to the receiving end through the first path with lower delay and lower bandwidth, so as to preferentially ensure that the receiving end can quickly and stably receive the base layer data necessary to maintain the continuous playing and basic quality of the video, thereby effectively reducing the initial loading time and the risk of freezing; at the same time, the video frames corresponding to the enhancement layer are sent to the receiving end through the second path with higher bandwidth, so as to fully utilize the high-bandwidth channel to transmit the enhancement layer data for improving the quality, without affecting the basic fluency, and finally combine the base layer data to restore a higher-quality video picture. The application intelligently layers the original data before transmission, and according to the real-time state of the network, preferentially guarantees the low-delay transmission of the base layer to maintain the video continuity, and at the same time optimizes the detail transmission by utilizing the high-bandwidth demand characteristics of the enhancement layer, so as to dynamically adapt to the network condition and maximize the use of available network resources; and the independent transmission of the first path and the second path realizes the load sharing and isolation of the video frames on different paths, avoids the overall quality degradation caused by the congestion of a single path, especially guarantees the low delay and high priority delivery of the base layer video, thereby improving the quality of data transmission as a whole, and providing stable and coherent high-quality audio and video experience for users. BRIEF DESCRIPTION OF DRAWINGS

[0014] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the application and serve to explain the principles of the application together with the specification.

[0015] In order to more clearly illustrate the technical solutions in the embodiments of the application or the prior art, the accompanying drawings needed to be used in the embodiments or prior art description will be briefly introduced as follows. Obviously, for those skilled in the art, other drawings can also be obtained based on these drawings without creative labor.

[0016] Figure 1 The flowchart provided by the data transmission method embodiment one of the application; Figure 2The schematic diagram of the overall process of the data transmission method provided for Embodiment Three of the present application is shown in the figure below. Figure 3 The schematic diagram of the module structure of the data transmission device of Embodiment One of the present application is shown in the figure below. Figure 4 The schematic diagram of the module structure of another data transmission device of the present application is shown in the figure below. Figure 5 The schematic diagram of the device structure of the hardware running environment of the sending end involved in the data transmission method of the present application is shown in the figure below. Figure 6 The schematic diagram of the device structure of the hardware running environment of the receiving end involved in the data transmission method of the present application is shown in the figure below.

[0017] The object implementation, functional features and advantages of the present application will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. DETAILED DESCRIPTION

[0018] It should be understood that the specific embodiments described herein are merely intended to explain the technical solutions of the present application, and are not intended to limit the present application.

[0019] In order to better understand the technical solutions of the present application, the specific embodiments will be described in detail below in conjunction with the drawings and the specific embodiments.

[0020] When audio and video are transmitted through AVoIP technology, key frame loss and other situations may occur when network anomalies occur, which may cause subsequent prediction frames to be unable to be correctly restored and decoded, resulting in mosaic, blurred pictures or audio lag in the finally restored video picture, and thus causing the quality of audio and video transmission to decrease and the user experience to decrease.

[0021] The present application provides a solution, according to the video features of each video frame, the video frame is divided into a base layer and an enhancement layer, the importance of the video content is layered, so as to realize the differentiated transmission of different content video frames; further, according to the delay and bandwidth and other parameters of each network path, the transmission paths corresponding to the base layer and the enhancement layer are determined, wherein the base layer is preferentially allocated to a first path with low delay to ensure the real-time and continuity of the video, and the enhancement layer is allocated to a high-throughput path (second path) with high bandwidth weight to carry large traffic; further, according to the real-time state of the network, the low-delay transmission of the base layer is preferentially guaranteed to maintain the continuity of the video, and the high-bandwidth demand characteristics of the enhancement layer are utilized to optimize the detail transmission, so as to dynamically adapt to the network condition, maximize the use of available network resources, and ensure the low delay and high priority delivery of the base layer video, thereby improving the quality of data transmission.

[0022] It should be noted that the execution entities in this embodiment involve a sending end and a receiving end. The sending end or receiving end can be an electronic device with data processing, network communication, and program execution functions, such as a tablet computer, personal computer, or mobile phone. Specifically, the sending end refers to the electronic device that sends raw data, used to acquire, encode, and package raw audio and video signals into a network-transmittable format before pushing them to network nodes; while the receiving end refers to the electronic device that receives data, used to receive bitstreams or data packets from the network, perform unpacking, deduplication, sorting, decoding, synchronization, and other operations, and render them to the screen and / or speakers.

[0023] The following description uses the sending end as the execution subject to illustrate this embodiment and the subsequent embodiments. Based on this, the embodiments of this application provide a data transmission method, referring to... Figure 1 , Figure 1 This is a flowchart illustrating the first embodiment of the data transmission method of this application.

[0024] In this embodiment, the data transmission method includes steps S10 to S40: Step S10: Encode the raw data to obtain multiple video frames; Raw data refers to data streams that have not undergone compression or encoding processing. It can include audio and video data, such as analog audio signals captured by a microphone and analog video signals captured by a camera. In computer systems, video data is usually represented as a series of continuous bitmap image arrays, such as YUV or RGB format, while audio data is represented as a PCM (Pulse Code Modulation) sample sequence.

[0025] A video frame is the basic unit of encoded video data. It typically includes independent keyframes (I-frames) that can represent a complete image, prediction frames (P-frames) that record the difference information of the previous frame, and bidirectional prediction frames (B-frames) that record the difference between consecutive frames.

[0026] For example, the video data in the original data can be processed by a video encoder to remove redundant information from the original video data and encode the video stream in the original video data into multiple video frames.

[0027] Step S20: Determine the data level to which each video frame belongs based on the video features of each video frame, wherein the data level includes a base layer and an enhancement layer; Video features refer to various attribute information in a video frame that can reflect its content, structure, motion, etc. They can be extracted and analyzed through image processing and computer vision algorithms, such as color histograms, texture features, motion vectors, etc.

[0028] Data hierarchy refers to logical labels used to identify data importance and transmission priority for hierarchical management, which can be mainly divided into base layer and enhancement layer.

[0029] The base layer refers to the minimum data combination containing key information to ensure basic understanding and availability of data, while the enhancement layer refers to a data set providing additional information to improve data quality or enrich data content based on the base layer.

[0030] For example, after dividing the video frames, the base layer contains the minimum data set absolutely necessary for the receiving end to reconstruct recognizable and coherent video sequences, usually including I frames and a few key P frames; while the enhancement layer contains additional data for enhancing the video quality of the base layer, usually including most P frames, B frames and possibly high-frequency detail information.

[0031] For example, the video encoder can extract and analyze the features of each video frame, and divide the video frames with rich video features to the base layer. For example, the motion vector of a video frame with large motion of a person or dynamic change of a background is high, and the video frame with rich video features can be determined as a video frame of the base layer.

[0032] It can be understood that by dividing the video frames into base layer and enhancement layer based on video features, the transmission strategy can be dynamically adjusted according to the network condition, differential transmission is realized, and the network resource utilization is improved.

[0033] In a possible implementation, the video features include motion vector, histogram difference and gradient variance mean, and step S20 includes: Step S21, in the case that the motion vector of the video frame is greater than a first threshold or the histogram difference is greater than a second threshold, the video frame is determined as a video frame belonging to the base layer. The motion vector refers to a parameter used to describe the motion direction and displacement size of an image block (macro block) from the previous frame to the current frame, which is usually output by the encoder in the P / B frame prediction stage, and is used to quantify the motion activity of the frame.

[0034] The histogram is a statistical chart of the intensity distribution of pixels in an image; and the histogram difference is a scalar value calculated by comparing the brightness (and / or color) histograms of the current frame and the previous frame, which is used to measure the overall brightness and color distribution change between the two frames, so as to determine whether scene switching occurs between the two frames.

[0035] The first threshold and the second threshold are preset or dynamically self-adaptively adjusted critical parameters, wherein the first threshold is used for comparison and judgment with the motion vector of the video frame; and the second threshold is used for comparison with the histogram difference of the video frame to judge whether the video frame has scene switching. Both of the two values can be preset by the user or self-adjusted according to the network state, for example, when the network delay is high, the threshold can be appropriately increased to reduce the number of video frames of the base layer and preferentially ensure the normal transmission of the base layer video frame.

[0036] Exemplarily, in the case that the video frame is a P frame or a B frame, the sending end can directly parse the motion vector of each macroblock in the video frame from the output data of the encoder and statistically aggregate the motion vector of the video frame at the frame level; and then compare the motion vector with the first threshold, if the motion vector is greater than the first threshold, it indicates that the video frame is a high motion video frame, which is divided into the base layer to ensure the picture continuity of the restored video and avoid the sudden change of the video object motion. Alternatively, in the case that the video frame is an I frame, the "hist_diff" parameter output by the encoder can also be directly read and determined as the histogram difference of the video frame; if the histogram difference is greater than the second threshold, it indicates that the scene switching occurs between the video frame and the previous video frame, and the video frame is determined as the video frame of the base layer to ensure the picture continuity of the restored video and avoid the scene mutation.

[0037] In step S22, in the case that the motion vector of the video frame is less than or equal to the first threshold, the histogram difference is less than or equal to the second threshold, and the gradient variance mean is less than or equal to the third threshold, the video frame is determined as the video frame belonging to the enhancement layer.

[0038] The gradient variance mean is a value obtained by calculating the variance of the gradient values of each local region in the video frame and then taking the average, which is used to reflect the richness and change of the image details and edges in the video frame. The higher the value, the more complex the scene.

[0039] The third threshold is also a dynamically self-adaptively adjusted critical parameter, which is used to judge the scene complexity of the current video frame.

[0040] Exemplarily, the sending end can directly use SSIM (Structural Similarity Index) or BRISQUE (Blind / Reference-less Image Spatial Quality Evaluator) to determine the scene complexity of the video frame, but the resource consumption of the two is large; and when the resource is limited, the gradient variance value of the video frame in the encoder can be directly read, the mean value thereof is calculated, and the mean value is compared with the third threshold value; if the mean value of the gradient variance of the video frame is greater than the third threshold value, it is indicated that the current scene complexity is high, and if the mean value of the gradient variance of the video frame is less than or equal to the third threshold value, it is indicated that the scene complexity is low.

[0041] Optionally, in the case that the motion vector of the video frame is less than or equal to the first threshold value, the histogram difference is less than or equal to the second threshold value, and the mean value of the gradient variance is greater than the third threshold value, the video frame is determined as the video frame belonging to the base layer.

[0042] Exemplarily, for any video frame, in the case that the motion vector of the video frame is less than or equal to the first threshold value and the histogram difference is less than or equal to the second threshold value, it can be determined that the video frame neither appears high motion change nor scene switching; further, whether the video frame belongs to the base layer or the enhancement layer can be determined through the mean value of the gradient variance of the video frame; if the mean value of the gradient variance is greater than the third threshold value, it is indicated that the scene complexity of the video frame is low, even if the network is poor and the frame is lost, it will not affect the picture continuity of the restored video, and therefore the video frame can be determined as the video frame of the enhancement layer.

[0043] In the embodiment, the importance of the video frame to maintaining the picture continuity is determined according to the video features such as the motion vector, the histogram difference and the mean value of the gradient variance of the video frame, so that the accuracy of the video frame layering can be ensured, and the high-quality layered transmission is realized.

[0044] In step S30, the video frame corresponding to the base layer is sent to the receiving end through the first path; In step S40, the video frame corresponding to the enhancement layer is sent to the receiving end through the second path, wherein the time delay of the first path is lower than that of the second path, and the bandwidth of the first path is lower than that of the second path.

[0045] The first path represents a network path for transmitting the base layer data, and the second path represents a network path for transmitting the enhancement layer; it can be the same network path or completely different two independent paths, and the embodiment does not make specific limitation thereto.

[0046] Exemplarily, the probe data packets can be continuously or periodically sent to all available network paths, the current latency and available bandwidth of each network path are measured, and then the network path with lower latency is determined as the first path and the network path with higher available bandwidth is determined as the second path.

[0047] It can be understood that, based on the completely different transmission requirements (such as low latency, high bandwidth, etc.) of the base layer data and the enhancement layer data, the base layer data with high priority and low latency requirement is guided to the most stable path to ensure the continuity and the lowest acceptable quality of the video stream, and the enhancement layer data with high bandwidth requirement is guided to the path with the largest capacity to provide the highest video definition and details as much as possible when the network condition allows, thereby realizing the maximum utilization of network resources and improving the overall quality of video data transmission and user perception.

[0048] Optionally, between sending the video frames, the video frames can be packaged to obtain video packets corresponding to the video frames, and the video packets corresponding to the video frames of the base layer are sent to the receiving end through the first path, and the video packets corresponding to the video frames of the enhancement layer are sent to the receiving end through the second path.

[0049] Packaging refers to the process of encapsulating data (such as video frame data, audio frame data) of a specific format according to certain protocols and rules, adding necessary header information (such as sequence number, timestamp, etc.) to make it a data packet that meets the transmission or storage requirements. Common network protocols include: RTP (Real-time Transport Protocol, Real-time Transport Protocol), SRTP (Secure RTP, Secure Real-time Transport Protocol), etc. The video packet refers to a data unit formed by packaging the video frame, which contains video frame data and its corresponding header information. The receiving end can analyze and process the video frame data according to the header information in the video packet.

[0050] Exemplarily, in the process of packaging the video frame, first, a packet header is constructed according to a predetermined network protocol, including steps of assigning a sequence number to the video frame, stamping a timestamp according to the packaging time, and writing a base layer or enhancement layer identifier in the extension field of the packet header; and then the packet header and the video frame are encapsulated into a video packet.

[0051] In one possible implementation, step S10 includes: Step A10, encoding the original data to obtain a plurality of audio frames; The audio frame refers to the basic unit after audio encoding, including compressed samples of all audio channels in a very short time (such as tens of milliseconds).

[0052] Exemplarily, the audio data in the original data can be processed by an audio encoder to be divided into individual audio frames.

[0053] The data transmission method further comprises: Step A20, performing voice activity detection on each audio frame to obtain a detection result of each audio frame; Voice activity detection (VAD) can be used to detect whether valid speech exists in an audio signal. In this embodiment, a voice activity detection method is used to detect whether human voice exists in an audio frame. The detection result can be a result indicating whether human voice exists in an audio frame.

[0054] Optionally, for any audio frame to be processed, the short-time energy and the zero-crossing rate of the audio signal in the frame are calculated by a VAD algorithm; in the case where both the short-time energy and the zero-crossing rate of the frame are greater than a preset threshold, it is determined that the detection result of the frame is that human voice exists.

[0055] For example, the short-time energy of a frame of audio signal can be determined by calculating the sum of squares of the amplitudes of sample points in the frame; the zero-crossing rate of a frame of audio signal can be determined by calculating the number of sign changes of sample points in the frame.

[0056] Step A30, in the case where the detection result of the audio frame is that human voice exists, determining that the data level to which the audio frame belongs is the base layer; Step A40, in the case where the detection result of the audio frame is that human voice does not exist, determining that the data level to which the audio frame belongs is the enhancement layer; For example, since the base layer contains the minimum data set absolutely necessary for the receiving end to reconstruct an identifiable and coherent audio sequence, and the enhancement layer contains additional data for enhancing audio quality, the audio frame containing human voice detected is divided into the high-priority base layer, and the audio frame without human voice (silence or pure noise) is divided into the enhancement layer, thereby achieving differentiation of the importance of audio data, so as to achieve differentiated transmission of audio data.

[0057] Optionally, after obtaining the detection result of whether human voice exists in the audio frame, the audio frame with the detection result that human voice exists can be further determined as a candidate frame, and keyword detection is performed on each candidate frame, and the candidate frame containing a preset keyword is determined as a target frame; then, each target frame can be determined as an audio frame belonging to the base layer, and other audio frames other than the target frames can be determined as audio frames belonging to the enhancement layer.

[0058] Exemplarily, the process of the keyword detection includes: for any candidate frame, Mel Frequency Cepstral Coefficents (MFCC) or Log-Mel Spectrogram is extracted from the candidate frame as a feature vector representing the human auditory characteristics; then, the feature vector is input into a preset lightweight deep neural network or convolutional neural network, and a confidence score of the feature vector compared with a preset keyword is output; then, the candidate frame with a confidence score greater than a preset confidence threshold is determined as a target frame.

[0059] In step A50, the audio frame corresponding to the base layer is sent to the receiving end through the first path, and the audio frame corresponding to the enhancement layer is sent to the receiving end through the second path.

[0060] It can be understood that the audio frame corresponding to the base layer is transmitted through the first path with low delay and stable transmission, which ensures the accurate transmission of the key audio information; at the same time, the audio frame corresponding to the enhancement layer is transmitted through the second path with high bandwidth, which can provide as much audio details as possible when the network condition allows, and the two together improve the quality of audio data transmission.

[0061] Optionally, the audio frame can also be packaged before being sent, and the specific processing process is the same as the packaging and sending process of the video frame in step S40, and thus will not be described again.

[0062] The embodiment provides a data transmission method, and a sending end divides video frames and audio frames into a base layer and an enhancement layer according to content characteristics of the video frames and the audio frames, so as to realize differentiated data transmission; then, the video frames / audio frames corresponding to the base layer are sent to a receiving end through a first path with low delay, so as to ensure stable transmission of the base layer data and ensure that the receiving end can decode a coherent and usable video / audio stream, effectively avoiding problems such as video / audio lag and discontinuity, thereby improving the quality of audio / video data transmission; at the same time, the video frames / audio frames corresponding to the enhancement layer are sent to the receiving end through a second path with high bandwidth, so that the overall scene details of the audio / video can be enriched when the network condition is good, and the overall quality of the audio / video data transmission and the user's viewing experience are improved.

[0063] Based on the first embodiment, a second embodiment of the data transmission method of the application is provided, and in the embodiment, before step S30, the following steps are further included: In step S301, the health degree score of each network path is determined according to the network state parameters of each network path and the corresponding parameter weights, wherein the network state parameters include at least one of a delay, a jitter, a packet loss rate, an available bandwidth and a stability, the delay weight of the base layer is higher than the delay weight of the enhancement layer, the jitter weight of the base layer is higher than the jitter weight of the enhancement layer, the packet loss rate weight of the base layer is higher than the packet loss rate weight of the enhancement layer, the stability weight of the base layer is lower than the stability weight of the enhancement layer, and the bandwidth weight of the base layer is lower than the bandwidth weight of the enhancement layer. A network path refers to a logical or physical communication channel that a data packet can pass through from a sending end to a receiving end, which can be an IP (Internet Protocol) route formed by different routing protocols, or different physical network interfaces (such as a 5G cellular network, Wi-Fi, a wired Ethernet network) or a logical interface after binding.

[0064] A network state parameter refers to a set of indexes for quantifying the transmission quality of a path, which are obtained by active or passive detection of a network, and mainly include a delay, a jitter, a packet loss rate, an available bandwidth and a stability. The delay refers to a one-way or round-trip time of a data packet from a sending end to a receiving end, and the jitter refers to a variation of the delay, i.e., a variance of time intervals of consecutive data packets. The packet loss rate represents a proportion of lost data packets in total sent packets in a transmission process. The available bandwidth refers to a maximum data transmission efficiency that a network path can provide. The stability refers to an ability of a network path to maintain a normal transmission state within a period of time, which is determined based on a variation degree of other network state parameters (such as a bandwidth, a delay, a packet loss rate, etc.).

[0065] A parameter weight refers to a coefficient assigned to each network state parameter in a health degree score calculation formula. The size of the weight determines the relative importance of the parameter in the health degree score of each network path.

[0066] A health degree score refers to a comprehensive measurement value for quantifying a current transmission quality of a network path, which can be obtained by weighted calculation of network state parameters.

[0067] Optionally, the health degree score of each network path can be obtained by multiplying and adding the scores of each network state parameter and the corresponding parameter weights of each network state parameter, wherein the network state parameters include at least one of a delay, a jitter, a packet loss rate, an available bandwidth and a stability.

[0068] Exemplarily, a calculation formula of the health degree score score is as follows:

[0069] wherein, , , , , respectively represent the delay weight, the jitter weight, the packet loss rate weight, the bandwidth weight and the availability weight; f(D) represents a processing function for delay, which is used to convert the delay into a numerical value quantitatively evaluating the high and low of the delay; f(J) and f(P) respectively represent processing functions for jitter and packet loss rate, which are respectively used to evaluate the influence of jitter and packet loss on the quality of data transmission; f(AvailBW) represents a processing function for available bandwidth, which is used to evaluate the bandwidth condition of the network path; f(Stability) represents a processing function for stability, which is used to evaluate the stability and consistency of the network path.

[0070] wherein, the health score of the base layer and the enhancement layer can be provided with two different sets of parameter weights, the delay weight of the base layer is higher than that of the enhancement layer, the jitter weight of the base layer is higher than that of the enhancement layer, the packet loss rate weight of the base layer is higher than that of the enhancement layer, the bandwidth weight of the base layer is lower than that of the enhancement layer, and the stability weight of the base layer is lower than that of the enhancement layer.

[0071] In step S302, the first path and the second path are determined from the network paths according to the health scores of the network paths.

[0072] Exemplarily, the probe data packets can be continuously or periodically sent to all available network paths, and the current delay and available bandwidth of each network path are measured; for each network path, the health scores of the network paths corresponding to the base layer and the enhancement layer are respectively calculated using two sets of weight configuration data, wherein the delay weight of the base layer is higher than that of the enhancement layer, and the bandwidth weight of the base layer is lower than that of the enhancement layer; and then the network path with the highest health score corresponding to the weight configuration data of the base layer is determined as the first path, and the network path with the highest health score corresponding to the weight configuration data of the enhancement layer is determined as the second path.

[0073] It can be understood that, by setting different parameter weights according to the characteristics of the base layer and the enhancement layer, the adaptation degree of each network path to the video frames of different levels can be more accurately evaluated, the base layer data is guided to a high-reliability path (the first path) with low delay, low jitter and low packet loss rate, and the enhancement layer data is guided to a large-capacity path (the second path) with high bandwidth and high stability, thereby realizing the optimal pairing between the network path capability and the data transmission demand, and improving the quality of audio and video data transmission.

[0074] In a feasible implementation, the first path includes a main path and a backup path, and step S302 includes: Step S3021, in the case that the video frame belongs to the base layer, a main path and a secondary path are determined according to the health scores of the network paths, wherein the main path is used to transmit all the video packets corresponding to the video frame in full amount, and the secondary path is used to transmit the redundant data of the video frame according to a preset redundancy replication ratio; In an embodiment, the sending end can select two transmission paths for the same base layer video frame according to the health score ranking, determine the network path with a higher score as the main path and the network path with a lower score as the secondary path, and then transmit all the data packets corresponding to the video frame through the main path, determine the redundant packets in proportion to the redundancy replication ratio from all the data packets corresponding to the video frame, and transmit the redundant packets through the secondary path. When the main path has a problem, the secondary path can provide certain data backup and supplement, thereby improving the reliability and fault tolerance of video transmission.

[0075] The redundancy replication ratio is a preset parameter used to determine how many percentages of the original data packets are to be replicated and transmitted through the secondary path. By setting different redundancy replication ratios, the reliability and bandwidth utilization of transmission can be balanced.

[0076] For example, for a base layer video frame, the sending end can query the health scores of all the network paths, select the one with the highest score as the main path and the one with the second highest score as the secondary path, and then put all the video packets corresponding to the video frame into the sending queue of the main path, and according to the preset redundancy replication ratio, replicate the corresponding proportion of the packets from all the video packets corresponding to the video frame according to a rule (such as selecting one out of every several packets or randomly selecting) and put them into the sending queue of the secondary path. Then, the original all video packets are transmitted in full amount to the receiving end through the main path, and the redundant video packets replicated are transmitted to the receiving end through the secondary path. After receiving the video packets transmitted by the main path and the secondary path, the receiving end can perform deduplication according to the sequence numbers of the video packets.

[0077] It can be understood that by transmitting the base layer data through the dual paths, even if the main path has packet loss or a short interruption, the receiving end still has a high probability of recovering the complete data from the redundant packets transmitted by the secondary path, thereby reducing the possibility of loss of base layer video frames caused by network burst anomalies, reducing the occurrence of video freezing, frame loss and other phenomena, and improving the quality and reliability of audio / video transmission. At the same time, by presetting the redundancy replication ratio, the secondary path does not transmit the video packets in full amount, but reasonably utilizes the bandwidth resources on the premise of ensuring a certain fault tolerance, thereby improving the reliability of transmission and avoiding excessive occupation of bandwidth, and improving the overall bandwidth utilization efficiency.

[0078] Step S3022, in the case that the video frame belongs to the enhancement layer, the video frame is disassembled into multiple network abstraction layer units, and the second path corresponding to each network abstraction layer unit is determined through the weighted round-robin scheduling algorithm and the health score of each network path.

[0079] The network abstraction layer unit (NALU) is a syntax structure defined in the video coding standards such as H.264, H.265, and represents the basic encapsulation unit of the coded video data.

[0080] The weighted round-robin (WRR) scheduling algorithm refers to an algorithm for task allocation among multiple resources (such as network paths), which allocates tasks to each resource in a certain order according to the weight value of each resource, and the higher the weight value of the resource, the greater the probability of being allocated to the task. In this embodiment, the resource refers to the network path, and the weight value is determined according to the health score of each network path.

[0081] Exemplarily, for the video frame of the enhancement layer, the sending end can read its code stream through the parser, and according to the start code of the NALU, it is divided into multiple independent NALUs; then, the health scores of all network paths of the enhancement layer are obtained, and the health scores of the network paths are used as their weights; then, through the WRR algorithm, each network path is sequentially and circularly pointed to, and each NALU is allocated to a path according to its weight, and the higher the weight of the path, the greater the probability of being selected and the more frequent the selection.

[0082] Optionally, the FEC (Forward Error Correction) can be used to add redundant codes to each NALU, so that even if there is packet loss in the transmission process of each NALU, other NALUs can also recover the complete video frame according to the redundant codes without retransmission, improving the quality and efficiency of the enhancement layer video transmission.

[0083] It can be understood that for the enhancement layer video frame with low importance, it is disassembled into smaller and independent transmission units (network abstraction layer units), and then the weighted round-robin scheduling algorithm is used to distribute these small units to multiple paths for transmission according to the health scores of the paths, which can effectively aggregate the bandwidth of multiple paths, thereby efficiently and quickly transmitting large-capacity enhancement layer data.

[0084] In a feasible implementation, the data transmission method further comprises: Step E10: Using a pre-defined second-order exponential weighted moving average model, determine the short-term fluctuation values ​​and long-term trend values ​​of each network state parameter in the main path and the second path, respectively. The Exponentially Weighted Moving Average (EWMA) model is a statistical method for calculating the weighted average of time series data. Its core idea is to assign weights to historical data that decay exponentially over time, with data points closer to the present having higher weights and those further away having lower weights. The formula can be seen as follows: ,in It is the EWMA value at time t. This refers to the observed value of a certain network state parameter (such as latency, packet loss rate, etc.) at the current moment. It is the smoothing coefficient (0.05 < The value <0.3 determines the rate of weight decay.

[0085] The second-order EWMA model calculates the weighted average by assigning different weights to data at different times. It not only considers the relationship between the current data point and the data point at the previous time, but also takes into account the second-order dynamic characteristics of the data points as they change over time, thus more effectively capturing the changing trends and fluctuations of the data.

[0086] Short-term fluctuations refer to the EWMA values ​​with relatively large smoothing coefficients for the current observations, reflecting the rapid fluctuations of the network; while long-term trend values ​​refer to the EWMA values ​​with relatively small smoothing coefficients for the current observations, reflecting the overall trend of the network.

[0087] For example, the second-order EWMA model calculation formula for any network state parameter is as follows:

[0088]

[0089] in, This represents the observed value of a certain network state parameter at the current moment; and These are the smoothing coefficients calculated for short-term fluctuations and long-term trend values, respectively. ; and These represent the short-term fluctuation value and long-term trend value of a certain network state parameter, respectively.

[0090] Step E20: If the difference between the short-term fluctuation value and the long-term trend value of at least one network state parameter in the main path is greater than the preset fluctuation threshold, increase the redundancy replication ratio of the secondary path in the base layer. The fluctuation threshold refers to a preset threshold value for determining whether the short-term fluctuation is severe enough to trigger the emergency mechanism.

[0091] Exemplarily, the sending end can monitor the change of each network status parameter of the main path in real time; then, according to the observation value of each network status parameter and the second-order EWMA model, the short-term fluctuation value and the long-term trend value of each network status parameter are calculated, and the difference between them is further calculated; then, the difference is compared with the preset fluctuation threshold value, and if the difference of at least one network status parameter is greater than the fluctuation threshold value, it indicates that the network status of the current main path is deteriorating, and the emergency mechanism is automatically triggered, and the redundant replication ratio of the secondary path is appropriately increased, for example, the redundant replication ratio of the secondary path of the base layer is adjusted to 100%, that is, a complete copy is created for each base layer data packet and sent through the secondary path, so that even in the case of data loss or damage due to high load of the main path, the redundant data of the secondary path can be used for recovery, ensuring the integrity of the key data.

[0092] Step E30, in the case that the difference between the short-term fluctuation value and the long-term trend value of at least one network status parameter of the second path is greater than the fluctuation threshold value, the second path is adjusted to a preset backup path.

[0093] The backup path refers to a predetermined backup network path; when the currently used second path fails, its performance decreases or meets a certain trigger condition, the sending end will automatically switch to this backup path to ensure the continuity and stability of data transmission. In the present embodiment, the backup path can be determined as a network path whose health score determined based on the weight configuration of the enhancement layer is lower than that of the original second path, or the health scores of each network path can be re-determined based on the current network status parameters of each network path, and the backup path can be determined according to the re-determined health scores. The present embodiment does not make specific limitations on this.

[0094] Exemplarily, the sending end can monitor the change of each network status parameter of the second path in real time; the triggering process of the emergency mechanism is the same as that in step E20 described above, and therefore will not be described again; then, after triggering the emergency mechanism, the second path is adjusted to the backup path with the suboptimal health score, thereby ensuring the continuity of the enhancement layer data transmission.

[0095] Optionally, in the case that the difference between the short-term fluctuation value and the long-term trend value of at least one network status parameter of the second path is greater than the fluctuation threshold value, in addition to switching to the backup path, the redundant ratio of the FEC of the second path can also be increased; and the change of each network status parameter is continuously detected, and in the case that the time length during which the difference of at least one network status parameter is greater than the fluctuation threshold value is greater than the preset time length, the backup path is switched to.

[0096] It can be understood that by analyzing the short-term fluctuations and long-term trends of the parameters, abnormal signs of the path can be perceived in advance before serious packet loss or delay occurs, so as to quickly avoid unstable network paths and effectively avoid transmission interruption and other situations caused by sudden abnormalities of a single path, thereby improving the stability of audio and video transmission.

[0097] In a possible implementation, the data transmission method further includes: By a preset change point detection algorithm, it is determined whether there is a change point in each network state parameter of the main path and the second path; in the case that there is at least one network state parameter with a change point in the main path, the redundancy duplication ratio of the secondary path of the base layer is adjusted to 100%; in the case that there is at least one network state parameter with a change point in the second path, the second path is adjusted to a preset backup path.

[0098] The change point detection algorithm refers to an algorithm for identifying points at which statistical characteristics (such as mean, variance) of time series data suddenly change. It can analyze the data sequence to determine whether the distribution of data after a certain point has deviated significantly and continuously compared with before; this point can be referred to as a change point, i.e., a data point at which the data distribution deviates. Common change point detection algorithms include CUSUM (Cumulative Sum) detection algorithm, Page-Hinkley detection algorithm, etc.

[0099] Exemplarily, after detecting the change point, an emergency mechanism is triggered immediately. The specific implementation of the emergency mechanism can refer to the specific implementation modes of steps E20 and E30, and thus will not be described here.

[0100] In this embodiment, by using the second-order EWMA model and / or the change point detection algorithm, the fluctuation of the network state parameter is predicted at the sending end, so that the network state can be judged and actively scheduled in advance, thereby improving the stability and continuity of audio and video data transmission. By using the second-order EWMA model, trend prediction can be smoothed, and congestion tendency can be perceived in advance by hundreds of milliseconds; by using the change point detection algorithm, sudden abnormalities can be captured, and fast response can be achieved; and by combining the second-order EWMA model and the change point detection algorithm, the end-to-end delay can be stabilized within 50 ms in a metropolitan area network / campus network environment, and the audio and video can be kept from obvious lagging under 2% sudden packet loss.

[0101] Based on the first and / or second embodiments described above, a third embodiment of the data transmission method of the present application is proposed. In this embodiment, the data transmission method is applied to a receiving end, and the data transmission method includes: Step A10, receiving a plurality of video frames sent by the preset first path and the preset second path, wherein the video frames are obtained by encoding the original data by the client, the first path is used to send the video frames of the data level of the base layer, the second path is used to send the video frames of the data level of the enhancement layer, the data level of each video frame is determined according to the video characteristics of each video frame, the time delay of the first path is lower than that of the second path, and the bandwidth of the first path is lower than that of the second path. Optionally, when receiving the data packets (including video packets and audio packets) of the first path and the second path, different reordering windows can be set respectively, wherein the reordering window for receiving the base layer data from the first path is smaller than the reordering window for receiving the enhancement layer data from the second path.

[0102] Illustratively, the reordering window of the base layer data can be set to 12-24 ms, and the reordering window of the enhancement layer data can be set to 40-80 ms, allowing a larger delay. Within the reordering window, the data packets within the same video frame / audio frame can be sorted according to the sequence numbers of the data packets. In addition, in the process of sorting, further de-duplication can be performed according to the sequence numbers of the data packets.

[0103] Illustratively, in order to avoid congestion, a deadline for data packet decodability can also be set in the reordering window of the base layer; for example, for a video of 30 frames per second (fps), the deadline can be set to 33 milliseconds; for a video of 60 fps, the deadline can be set to 16.7 milliseconds. If a data packet cannot complete reordering and de-duplication and other processes before the deadline, it cannot be guaranteed to be decoded on time, so the data packet will be discarded to avoid affecting subsequent decoding and playing.

[0104] Step A20, recombining and decoding each video frame to obtain a restored video.

[0105] The restored video refers to a sequence of video frames consistent with the video data of the original data of the sending end after recombination and decoding processing by the receiving end.

[0106] Illustratively, the sending end sends the video frames of the base layer to the preset receiving end through the low-delay first path and sends the video frames of the enhancement layer to the receiving end through the high-bandwidth second path; then, after receiving the video frames, the receiving end first reorders (recombines) each audio frame according to the sequence number of each audio frame and de-duplicates according to the sequence number; then the receiving end can decode the recombined video frames to obtain a restored video.

[0107] In a feasible implementation, the data transmission method further comprises: Step A30, receiving the multiple audio frames sent by the client through the first path and the second path; Step A40, recombining and decoding each audio frame to obtain the restored audio; The restored audio refers to the audio signal consistent with the audio data in the original data of the sending end after recombination and decoding processing of the receiving end.

[0108] Exemplarily, after receiving the multiple audio frames sent by the client, the receiving end can recombine each audio frame into a continuous audio stream according to the header information (such as the time sequence number, the timestamp, etc.) corresponding to each audio frame in the original order relationship; and then, decodes the recombined audio stream through an audio decoder to obtain the restored audio.

[0109] The video frame and the audio frame both include a presentation timestamp and a sending timestamp, and the data transmission method further includes: Step A50, determining the audio frame with the same presentation timestamp as the presentation timestamp of the video frame as the reference audio frame; The presentation timestamp (PTS) refers to the information used to identify the time point at which the media data (such as the video frame and the audio frame) should be presented when playing, which is usually expressed in time units (such as milliseconds, microseconds, etc.).

[0110] It can be understood that the stability and accuracy of the video and picture synchronization are ensured by using the stably transmitted audio frame as the reference for alignment.

[0111] Step A60, determining the local presentation time of the corresponding restored video frame in the restored video according to the presentation timestamp of the video frame and the offset time, and determining the local presentation time of the corresponding restored audio frame in the restored audio according to the presentation timestamp of the reference audio frame and the offset time, wherein the offset time is determined according to the sending timestamp and the local receiving time; The sending timestamp refers to the time marked when the data is sent out from the sending end, which is usually based on the system clock of the sending end and is used to record the time information of data sending. The local receiving time refers to the receiving time recorded by the receiving end when the data arrives at the receiving end, which is based on the system clock of the receiving end.

[0112] The offset time refers to the time difference between the sending end and the receiving end caused by factors such as network transmission delay and device processing time difference; it can be calculated according to the difference between the sending timestamp of the last sent data packet and the local receiving time of the video frame / audio frame, and is used to measure the time delay experienced by the video frame / audio frame from sending to receiving.

[0113] The local presentation time refers to a time point at which the restored video frame and the audio frame should be actually presented on the receiving end device, which can be calculated according to the presentation timestamp, the offset time and other factors, and is used to ensure that the audio and video can be played in the correct time sequence and rhythm at the receiving end.

[0114] Exemplarily, for a video frame / audio frame, its corresponding local presentation time wherein, pts represents the presentation timestamp, represents the offset time.

[0115] Exemplarily, before calculating the local presentation time, the offset time can also be processed by Kalman filtering or moving average to avoid introducing jittery "drag / accelerate", for example, the offset time is wherein, represents the sending timestamp, represents the local receiving time, and E represents the Kalman filtering processing.

[0116] Step A33, determining the difference between the offset time of the video frame and the offset time of the reference audio frame as the time deviation of the restored video frame. The time deviation is used to measure the degree of inconsistency between the video and the audio in time; by calculating the time deviation, the synchronization between the video and the audio can be understood, which provides a basis for subsequent adjustment.

[0117] Exemplarily, the calculation formula of the time deviation is as follows:

[0118] wherein, and respectively represent the presentation timestamp of the video frame and the reference audio frame, and respectively represent the local presentation time of the video frame and the reference audio frame.

[0119] Step A34, adjusting the local presentation time of the restored video frame according to the time deviation until the difference between the local presentation time of the restored video frame and the local presentation time of the restored audio frame is less than or equal to a preset deviation threshold, so as to synchronize the restored video and the restored audio.

[0120] The deviation threshold refers to a maximum range value of the allowed difference between the local presentation times of the audio and the video; when the difference between the local presentation times of the audio and the video is less than or equal to this threshold, it can be considered that the audio and the video have reached a synchronization state; when the difference is greater than this threshold, adjustment needs to be continued. For example, in a normal network state, the threshold can be set to 20 ms; or in a network congestion state, the threshold can be set to 50 ms.

[0121] Exemplarily, the time deviation can be directly determined as an adjustment range of the local presentation time of the restored video frame, ensuring the audio-visual synchronization of the current video frame and audio frame, but the operation for other video frames is time-consuming and complex.

[0122] Exemplarily, in order to achieve overall audio-visual synchronization, a slight resampling correction drift (±100~300ppm) is made to the restored audio, or a ≤1ms fine-tuning of frame presentation is made to the video, to avoid audio-visual asynchronization. For example, according to the time deviation drift, the fine-tuned sampling rate of the restored audio can be determined as:

[0123] wherein kp is a gain coefficient for controlling the correction amplitude of the audio sampling rate, the clamp function is used to ensure that the correction does not exceed ±300ppm, and base_rate is the nominal audio sampling rate when the audio acquisition device acquires audio. Further, the fine-tuned sampling rate can be sent to an audio resampler or clock fine-tuning to slightly offset the audio stream and video stream at a very small rate, and the video end only makes a small presentation advance / delay of less than 1ms, and the main alignment is completed by the audio end, thereby continuously suppressing the lip difference without multiple adjustments for each video frame.

[0124] Exemplarily, after the synchronized restored video and restored audio are combined, the restored audio-video can be obtained, which is the complete audio-video content finally presented to the user.

[0125] In this embodiment, after receiving the video packet and audio packet sent by the sending end, the video packet and audio packet are recombined and decoded, and the restored video and restored audio obtained by decoding are synchronized to improve the audio-video synchronization experience of the user. During the synchronization process, the local presentation time of the audio and video is calculated respectively, and the playback scheduling is performed based on the local clock of the receiving end, to accurately compensate for the transmission delay difference caused by different network paths, solve the problem of audio-visual asynchronization, and ensure the immersive experience of audio-visual synchronization.

[0126] Exemplarily, in order to assist in understanding the implementation process of the data transmission method obtained after combining the above-mentioned embodiment one, please refer to Figure 2 , Figure 2 A general flowchart of a data transmission method is provided, specifically: First, the sending end performs B101 to obtain the original audio and video, and performs B102 to encode and content feature analyze the original audio and video, to obtain each video frame and audio frame corresponding to the original audio and video, and to determine the video features of each video frame, such as motion vector, histogram difference, gradient variance mean, etc., and voice activity detection can also be performed on the audio frame; then, B103 is performed to perform stream layering according to the content feature analysis result, and each video frame and audio frame is divided into a base layer or an enhancement layer, for example, the video frame with motion vector and histogram difference both greater than the corresponding threshold is divided into the base layer, and the audio frame with voice activity detection result as existing voice is divided into the base layer; at the same time, B104 can be performed to perform path health score and path scheduling, and the health score of each network path can be determined according to the weighted sum of the network state parameters (including delay and bandwidth, etc.) on each network path and the corresponding parameter weight, since the base layer and the enhancement layer are provided with two different sets of parameter weights, the delay weight of the base layer is higher than that of the enhancement layer, and the bandwidth weight of the base layer is lower than that of the enhancement layer, so the transmission path of the base layer is determined as a low-delay path (corresponding to the first path), and the transmission path of the enhancement layer is determined as a high-bandwidth path (corresponding to the second path); then, the video frame and the audio frame corresponding to the base layer are sent to the receiving end through the low-delay path, and the video frame and the audio frame corresponding to the enhancement layer are sent to the receiving end through the high-bandwidth path. Then, the receiving end performs B105 to recombine and decode the received video frame and audio frame, respectively, the receiving end can recombine and de-duplicate according to the sequence number of each data packet, decode the data packet corresponding to the recombined audio and the data packet corresponding to the video, respectively, to obtain the restored video and the restored audio, and synchronize the restored audio and the restored video to obtain the restored audio and video; finally, B106 is performed to output the restored audio and video to a preset interactive interface for the user to watch.

[0127] For example, the following gives the key parameters and timing budget based on the above data transmission method: For the data of the base layer, the packaging delay of the sending side is 1-3 ms, the copy processing needs 0.2-0.8 ms, and the one-way round-trip delay under the city area or park is less than or equal to 6 ms; the receiving rearrangement window is usually set to 12-24 ms, no FEC redundancy is set or 0-10% FEC redundancy is set, and the deviation threshold is usually set to 20 ms.

[0128] For the data of the enhancement layer, the packaging delay of the sending side is 3-8 ms, the copy processing needs 1-3 ms, and the one-way round-trip delay under the city area or park is less than or equal to 12 ms; the receiving rearrangement window is usually set to 40-80 ms, 10-25% FEC redundancy is set, and the deviation threshold is usually set to 50 ms.

[0129] Exemplarily, the sending end or the receiving end can be implemented based on a SoC (System-on-Chip) or an FPGA (Field-Programmable Gate Array) to accelerate processing efficiency; the following improvements can also be made: For the sending end, a NIC (Network Interface Card) can be used in combination with a DMA (Direct Memory Access) technology to directly transmit data from the memory to the network interface card at the transmission end, bypassing the traditional CPU (Central Processing Unit) processing, to provide more efficient data transmission; a RaptorQ acceleration can also be implemented in the FPGA IP core to improve real-time error correction capability and reduce the impact of packet loss on transmission.

[0130] For the receiving end, a hardware queue can be used to de-duplicate and reorder data packets (including video packets and audio packets) to improve the recombination efficiency of the data packets; a hardware acceleration module can also be set to accelerate data recovery of the FEC technology; a digital signal processor (DSP) unit can also be used to resample the audio to achieve more accurate audio-visual synchronization.

[0131] By using the above data transmission method, the base layer data required by the 1080p60 (1920x1080 resolution + 60 frames per second) requirement under the metropolitan area network can be transmitted to the receiving end with a low latency of less than 35ms; even in the case of 2% packet loss and 5ms jitter in a single path, the restored video display has no noticeable lag; even in the case of 3 congestion within a 10-second window, the time deviation of audio-visual synchronization can be controlled within 20ms; and in the case of 1 path being disconnected and being restored within 100ms, the complex strategy of the base layer can ensure the normal display of the picture. It can be seen that the data transmission method of the present application has the advantages of stable low latency, anti-packet loss, and anti-interruption.

[0132] It should be noted that the above examples are only used to understand the present application and do not limit the data transmission method of the present application, and more forms of simple changes based on this technical concept are within the protection scope of the present application.

[0133] The present application also provides a data transmission device, please refer to Figure 3 , the data transmission device is applied to a sending end, and the device comprises: An encoding module 10 is configured to encode original data to obtain a plurality of video frames. The hierarchical module 20 is configured to determine a data level to which each video frame belongs according to a video feature of each video frame, wherein the data level includes a base layer and an enhancement layer. The first sending module 30 is configured to send the video frame corresponding to the base layer to the receiving end through the first path. The second sending module 40 is configured to send the video frame corresponding to the enhancement layer to the receiving end through the second path, wherein the time delay of the first path is lower than that of the second path, and the bandwidth of the first path is lower than that of the second path.

[0134] The data transmission device provided by the embodiment of the present application adopts the data transmission method in the above embodiment, and can solve the technical problem of how to improve the transmission quality of data transmission. Compared with the prior art, the data transmission device provided by the present application has the same beneficial effects as the data transmission method provided by the above embodiment, and other technical features in the data transmission device are the same as the features disclosed in the above embodiment method, which will not be repeated here.

[0135] The embodiment of the present application also provides another data transmission device, please refer to Figure 4 The data transmission device is applied to a receiving end, and the device comprises: The receiving module 50 is configured to receive a plurality of video frames sent by a client through a preset first path and a preset second path, wherein the video frames are obtained by encoding original data by the client, the first path is used for sending video frames with a base layer as a data level, the second path is used for sending video frames with an enhancement layer as a data level, the data level of each video frame is determined according to a video feature of each video frame, the time delay of the first path is lower than that of the second path, and the bandwidth of the first path is lower than that of the second path. The decoding module 60 is configured to recombine and decode each video frame to obtain a restored video.

[0136] The data transmission device provided by the embodiment of the present application adopts the data transmission method in the above embodiment, and can solve the technical problem of how to improve the transmission quality of data transmission. Compared with the prior art, the data transmission device provided by the present application has the same beneficial effects as the data transmission method provided by the above embodiment, and other technical features in the data transmission device are the same as the features disclosed in the above embodiment method, which will not be repeated here.

[0137] The embodiment of the present application provides a sending end, which comprises at least one processor and a memory connected with the at least one processor, wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the data transmission method in the above embodiment one.

[0138] The following refers to Figure 5The diagram illustrates a structural schematic suitable for implementing the transmitting end in the embodiments of this application. The transmitting end in the embodiments of this application may include, but is not limited to, mobile terminals such as mobile phones, laptops, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Portable Application Descriptions), PMPs (Portable Media Players), in-vehicle terminals (e.g., in-vehicle navigation terminals), and fixed terminals such as digital TVs and desktop computers. Figure 5 The sending end shown is merely an example and should not impose any limitations on the functionality and scope of use of the embodiments of this application.

[0139] like Figure 5 As shown, the transmitting end may include a processing device 1001 (e.g., a central processing unit, a graphics processing unit, etc.), which can perform various appropriate actions and processes according to a program stored in the read-only memory 1002 or a program loaded from the storage device 1003 into the random access memory 1004. The random access memory 1004 also stores various programs and data required for the operation of the transmitting end. The processing device 1001, the read-only memory 1002, and the random access memory 1004 are interconnected via a bus 1005. An input / output interface 1006 is also connected to the bus. Typically, the following systems can be connected to the input / output interface 1006: input devices 1007 including, for example, a touch screen, touchpad, keyboard, mouse, image sensor, microphone, accelerometer, gyroscope, etc.; output devices 1008 including, for example, a liquid crystal display (LCD), speaker, vibrator, etc.; storage devices 1003 including, for example, magnetic tape, hard disk, etc.; and communication devices 1009. The communication device 1009 allows the transmitting end to exchange data via wireless or wired communication with other devices. Although the diagram shows transmitters with various systems, it should be understood that it is not required to implement or have all of the systems shown. More or fewer systems may be implemented alternatively.

[0140] In particular, according to the embodiments of the present application, the process described above with reference to the flowchart can be implemented as a computer software program. For example, the embodiments of the present application include a computer program product comprising a computer program carried on a computer readable medium, the computer program containing program code for executing the method shown in the flowchart. In such embodiments, the computer program can be downloaded and installed from a network by a communication device, or installed from a storage device 1003, or installed from a read-only memory 1002. When the computer program is executed by the processing device 1001, the above-mentioned functions defined in the method of the embodiments of the present application are executed.

[0141] The sending end provided by the embodiments of the present application adopts the data transmission method in the above embodiments, and can solve the technical problem of how to improve the transmission quality of data transmission. Compared with the prior art, the beneficial effects of the sending end provided by the present application are the same as those of the data transmission method provided by the above embodiments, and other technical features in the sending end are the same as those disclosed in the previous embodiment method, which will not be repeated here.

[0142] It should be understood that parts of the present application can be realized by hardware, software, firmware or a combination thereof. In the description of the above embodiments, specific features, structures, materials or characteristics can be combined in any one or more embodiments or examples in a suitable manner.

[0143] The above describes only the specific implementation of the present application, but the protection scope of the present application is not limited thereto, any person skilled in the art can easily think of changes or replacements within the technical scope disclosed by the present application, which should be covered within the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

[0144] The embodiments of the present application provide a receiving end, which comprises at least one processor, and a memory connected with the at least one processor in communication; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the data transmission method in the above embodiment one.

[0145] The following will be described with reference to the accompanying drawings Figure 6The diagram illustrates a suitable structural schematic for implementing the receiver in the embodiments of this application. The receiver in the embodiments of this application may include, but is not limited to, mobile terminals such as mobile phones, laptops, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Portable Application Descriptions), PMPs (Portable Media Players), in-vehicle terminals (e.g., in-vehicle navigation terminals), and fixed terminals such as digital TVs and desktop computers. Figure 6 The receiver shown is merely an example and should not impose any limitations on the functionality and scope of use of the embodiments of this application.

[0146] like Figure 6 As shown, the receiving end may include a processing device 2001 (e.g., a central processing unit, a graphics processing unit, etc.), which can perform various appropriate actions and processes according to a program stored in the read-only memory 2002 or a program loaded from the storage device 2003 into the random access memory 2004. The random access memory 2004 also stores various programs and data required for the operation of the receiving end. The processing device 2001, the read-only memory 2002, and the random access memory 2004 are interconnected via a bus 2005. An input / output interface 2006 is also connected to the bus. Typically, the following systems can be connected to the input / output interface 2006: input devices 2007 including, for example, touchscreens, touchpads, keyboards, mice, image sensors, microphones, accelerometers, gyroscopes, etc.; output devices 2008 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 2003 including, for example, magnetic tapes, hard disks, etc.; and communication devices 2009. The communication device 2009 allows the receiving end to exchange data via wireless or wired communication with other devices. Although the diagram shows receivers with various systems, it should be understood that it is not required to implement or have all of the systems shown. More or fewer systems may be implemented alternatively.

[0147] In particular, according to the embodiments disclosed in the present application, the process described above with reference to the flowchart can be implemented as a computer software program. For example, the embodiments disclosed in the present application include a computer program product comprising a computer program carried on a computer readable medium, the computer program containing program codes for executing the method shown in the flowchart. In such embodiments, the computer program can be downloaded and installed from a network through a communication device, or installed from the storage device 2003, or installed from the read-only memory 2002. When the computer program is executed by the processing device 2001, the above-mentioned functions defined in the method of the embodiments disclosed in the present application are executed.

[0148] The receiving end provided by the embodiments of the present application adopts the data transmission method in the above embodiments, and can solve the technical problem of how to improve the transmission quality of data transmission. Compared with the prior art, the receiving end provided by the present application has the same beneficial effects as the data transmission method provided by the above embodiments, and other technical features in the receiving end are the same as the features disclosed in the previous embodiment method, which will not be repeated here.

[0149] It should be understood that parts of the present application can be realized by hardware, software, firmware or a combination thereof. In the description of the above embodiments, specific features, structures, materials or characteristics can be combined in any one or more embodiments or examples in a suitable manner.

[0150] The above describes only the specific implementation of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art can easily think of changes or replacements within the technical scope disclosed in the present application, which should be covered within the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

[0151] The embodiments of the present application provide a computer readable storage medium having computer readable program instructions (i.e. computer program) stored thereon, the computer readable program instructions being used to execute the data transmission method in the above embodiments.

[0152] The computer readable storage medium provided by the embodiments of the present application may, for example, be a U disk, but is not limited to an electric, magnetic, optical, electromagnetic, infrared, or semiconductor system or device, or any combination thereof. More specific examples of the computer readable storage medium may include, but are not limited to, an electric connection with one or more conductive wires, a portable computer disk, a hard disk, a random access memory (RAM), a read only memory (ROM), an erasable programmable read only memory (EPROM or flash memory), an optical fiber, a portable compact disk read only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof. In the embodiments, the computer readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system or device. The program code contained on the computer readable storage medium may be transmitted by any suitable medium, including but not limited to an electric wire, an optical cable, an RF (Radio Frequency), and the like, or any suitable combination thereof.

[0153] The computer readable storage medium described above may be contained in a sending end / receiving end, or may exist separately without being assembled into the sending end / receiving end.

[0154] The computer readable storage medium described above carries one or more programs, which, when executed by the sending end, cause the sending end to: encode original data to obtain a plurality of video frames; determine a data level to which each video frame belongs according to a video feature of each video frame, wherein the data level includes a base layer and an enhancement layer; send the video frame corresponding to the base layer to the receiving end through a first path; and send the video frame corresponding to the enhancement layer to the receiving end through a second path, wherein the time delay of the first path is lower than that of the second path, and the bandwidth of the first path is lower than that of the second path.

[0155] The computer readable storage medium described above carries one or more programs, which, when executed by the receiving end, cause the receiving end to: receive a plurality of video frames sent by a client through a preset first path and a preset second path, wherein the video frames are obtained by encoding original data by the client, the first path is used to send a video frame of a base layer, the second path is used to send a video frame of an enhancement layer, the data level of each video frame is determined according to a video feature of each video frame, the time delay of the first path is lower than that of the second path, and the bandwidth of the first path is lower than that of the second path; recombine and decode each video frame to obtain a restored video.

[0156] Computer program code for carrying out operations of the present application can be written in any combination of one or more programming languages, including an object oriented programming language such as Java, Smalltalk, C++ or the like and conventional procedural programming languages, such as the "C" programming language or similar programming languages. The program code can execute entirely on the user's computer, partly on the user's computer, as a stand-alone software package, partly on the user's computer and partly on a remote computer or entirely on the remote computer or server. In the latter scenario, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection can be made to an external computer (for example, through the Internet using an Internet Service Provider).

[0157] The flow diagrams and the block diagrams in the drawings are meant as illustrative representations of the architectures, functions, and operations of possible implementations of systems, methods and computer program products according to the present application. It should be noted that each block in the flow diagrams and the block diagrams, and combinations of blocks in the flow diagrams and the block diagrams, can be implemented by either hardware, software, or combinations thereof. The flow diagrams of FIGS. 6-8 and the block diagrams of FIGS. 1-5 illustrate the architecture, functionality, and operations of possible implementations of systems, methods and computer program products according to the present application. In this regard, each block in the flow diagrams or block diagrams can represent a module, segment, or portion of code, which comprises one or more executable instructions for implementing the specified logical functions. It should also be noted that in some alternative implementations, the functions noted in the blocks can occur out of the order noted in the figures. For example, two blocks shown in succession may, in fact, be executed substantially concurrently or the blocks may

[0158] The modules involved in the embodiments of the present application can be implemented in the form of software or in the form of hardware. In some cases, the name of the module does not constitute a limitation on the module itself.

[0159] The readable storage medium provided by the embodiments of the present application is a computer readable storage medium, which stores computer readable program instructions (i.e. computer programs) for executing the above data transmission method, and can solve the technical problem of improving the transmission quality of audio and video transmission. Compared with the prior art, the computer readable storage medium provided by the present application has the same beneficial effects as the data transmission method provided by the above embodiments, and will not be described here.

[0160] The embodiment of the present application further provides a computer program product comprising a computer program, which, when executed by a processor, implements the steps of the data transmission method as described above.

[0161] The computer program product provided by the embodiment of the present application can solve the technical problem of improving the transmission quality of audio and video transmission. Compared with the prior art, the beneficial effects of the computer program product provided by the embodiment of the present application are the same as those of the data transmission method provided by the above-mentioned embodiment, and are not described here.

[0162] The above only describes some embodiments of the present application, and does not limit the patent scope of the present application. Any equivalent structural transformation made by using the content of the present application specification and drawings, or direct / indirect application in other related technical fields is included in the patent protection scope of the present application.

Claims

1. A data transmission method, characterized by, The data transmission method is applied to a sending end, and the method comprises: encoding original data to obtain a plurality of video frames; determining data levels to which the video frames belong according to video features of the video frames, wherein the data levels comprise a basic layer and an enhanced layer; sending video frames corresponding to the basic layer to a receiving end through a first path; sending video frames corresponding to the enhanced layer to the receiving end through a second path, wherein a time delay of the first path is lower than that of the second path, and a bandwidth of the first path is lower than that of the second path.

2. The data transmission method of claim 1, wherein, The video features comprise a motion vector, a histogram difference and a gradient variance mean, and the step of determining the data levels to which the video frames belong according to the video features of the video frames comprises: in a case where the motion vector of the video frame is greater than a first threshold value or the histogram difference is greater than a second threshold value, determining that the data level to which the video frame belongs is the basic layer; in a case where the motion vector of the video frame is less than or equal to the first threshold value, the histogram difference is less than or equal to the second threshold value, and the gradient variance mean is less than or equal to a third threshold value, determining that the data level to which the video frame belongs is the enhanced layer.

3. The data transmission method as claimed in claim 1, characterized in that, Before the step of sending the video frames corresponding to the basic layer to the receiving end through the first path, the method further comprises: determining health scores of network paths according to network state parameters of the network paths and corresponding parameter weights, wherein the network state parameters comprise at least one of a time delay, a jitter, a packet loss rate, an available bandwidth and stability, a time delay weight of the basic layer is higher than that of the enhanced layer, a jitter weight of the basic layer is higher than that of the enhanced layer, a packet loss rate weight of the basic layer is higher than that of the enhanced layer, a stability weight of the basic layer is lower than that of the enhanced layer, and a bandwidth weight of the basic layer is lower than that of the enhanced layer; determining the first path and the second path from the network paths according to the health scores of the network paths.

4. The data transmission method of claim 3, wherein, The first path comprises a main path and a secondary path, and the step of determining the first path and the second path from the network paths according to the health scores of the network paths comprises: in a case where the video frame belongs to the basic layer, determining the main path and the secondary path according to the health scores of the network paths, wherein the main path is used to send the video frame in full amount, and the secondary path is used to send copy data of the video frame according to a preset redundancy replication ratio; in a case where the video frame belongs to the enhanced layer, decomposing the video frame into a plurality of network abstraction layer units, and determining second paths corresponding to the network abstraction layer units according to a weighted round-robin scheduling algorithm and the health scores of the network paths.

5. The data transmission method of claim 4, wherein, The data transmission method further comprises: respectively determining short-term fluctuation values and long-term trend values of the network state parameters in the main path and the second path through a preset second-order exponentially weighted moving average model; In a case where a difference between a short-term fluctuation value and a long-term trend value of at least one network state parameter of the main path is greater than a preset fluctuation threshold, increasing a redundant replication ratio of the secondary path of the base layer; In a case where a difference between a short-term fluctuation value and a long-term trend value of at least one network state parameter of the second path is greater than the fluctuation threshold, adjusting the second path to a preset backup path.

6. The data transmission method of claim 1, wherein, The step of encoding the original data comprises: encoding the original data to obtain a plurality of audio frames; The data transmission method further comprises: performing voice activity detection on each of the audio frames to obtain a detection result of each of the audio frames; in a case where the detection result of the audio frame is that there is human voice, determining that the data level to which the audio frame belongs is a base layer; in a case where the detection result of the audio frame is that there is no human voice, determining that the data level to which the audio frame belongs is an enhancement layer; sending the audio frame corresponding to the base layer to the receiving end through the first path, and sending the audio frame corresponding to the enhancement layer to the receiving end through the second path.

7. A data transmission method, characterized by, The data transmission method is applied to a receiving end, and the method comprises: receiving a plurality of video frames sent by a client through a preset first path and a preset second path, wherein the video frames are obtained by encoding original data by the client, the first path is used to send video frames of a base layer, the second path is used to send video frames of an enhancement layer, the data level of each of the video frames is determined according to the video characteristics of each of the video frames, the time delay of the first path is lower than that of the second path, and the bandwidth of the first path is lower than that of the second path; recombining and decoding each of the video frames to obtain a restored video.

8. The data transmission method of claim 7 wherein, The data transmission method further comprises: receiving a plurality of audio frames sent by the client through the first path and the second path; recombining and decoding each of the audio frames to obtain a restored audio; The video frames and the audio frames each comprise a presentation timestamp and a sending timestamp, and the data transmission method further comprises: determining an audio frame with the same presentation timestamp as the presentation timestamp of the video frame as a reference audio frame; determining a local presentation time of a corresponding restored video frame in the restored video according to the presentation timestamp of the video frame and an offset time, and determining a local presentation time of a corresponding restored audio frame in the restored audio according to the presentation timestamp of the reference audio frame and the offset time, wherein the offset time is determined according to a sending timestamp and a local receiving time; determining a difference between the offset time of the video frame and the offset time of the reference audio frame as a time deviation of the restored video frame; adjusting the local presentation time of the restored video frame according to the time deviation until a difference between the local presentation time of the restored video frame and the local presentation time of the restored audio frame is less than or equal to a preset deviation threshold, so as to synchronize the restored video and the restored audio.

9. A transmitting end, characterized by, The sending end comprises a memory, a processor, and a computer program stored in the memory and capable of running on the processor, and the computer program is configured to implement the steps of the data transmission method according to any one of claims 1 to 6.

10. A receiving end, characterized by The receiving end comprises a memory, a processor, and a computer program stored in the memory and capable of running on the processor, and the computer program is configured to implement the steps of the data transmission method according to claim 7 or 8.

Citation Information

Patent Citations

  • Method and system for pushing video streaming based on layered coding

    CN101909063A

  • Multi-path data stream drainage method and device

    CN117749694A

  • Autonomous controllable architecture software and hardware collaborative audio and video signal transmission method and system

    CN120751133A

  • System and method for uploading 3D video to video website by user

    US20150334437A1

Cited By

  • Data processing method and device

    CN121349738A

  • Transmission optimization method and device

    CN122068945A