Audio and video data transmission method, device, equipment, medium and program product

By determining the candidate transmission order for audio and video data and determining the loss parameters based on frame attributes, the problem of audio and video playback fluency is solved, and more efficient audio and video data transmission is achieved and user experience is improved.

CN117651169BActive Publication Date: 2025-05-09BEIJING BAIDU NETCOM SCI & TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202311753396.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-12-19
Publication Date
2025-05-09
Estimated Expiration
2043-12-19

AI Technical Summary

Technical Problem

In video network transmission scenarios, it is necessary to improve the smoothness of audio and video playback of clients, especially in the planning of transmission sequence of audio frames and video frames, which is difficult to effectively solve in the prior art.

Method used

By determining the candidate transmission order for the data frames in the audio and video data to be transmitted, the loss parameters are determined according to the frame attributes, the target transmission order is determined, and the data frames are transmitted to the client using the target transmission order.

Benefits of technology

It improves the smoothness of audio and video playback of the client. By optimizing the transmission order of audio frames and video frames, it reduces playback lag and redundant time difference, and improves the user's viewing experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117651169B_ABST
    Figure CN117651169B_ABST
Patent Text Reader

Abstract

The present disclosure provides a method, device, equipment and medium for transmitting audio and video data, and relates to the field of data processing, and in particular to the field of data transmission technology. The specific implementation scheme is: determining a candidate transmission order for data frames in the audio and video data to be transmitted; wherein the data frames include audio frames and video frames; determining the loss parameters of the candidate transmission order according to the frame attributes of the data frames; determining the target transmission order from the candidate transmission order according to the loss parameters, and transmitting the data frames to the client using the target transmission order. The scheme of the present disclosure improves the accuracy and rationality of the determination of the target transmission order through quantitative analysis, thereby improving the fluency of the video playback on the client.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of data processing, in particular to the field of data transmission technology, and specifically to a method, device, equipment and medium for transmitting audio and video data. Background Art

[0002] In a video network transmission scenario, the audio and video data cached in the server needs to be sent to the client, and the client downloads and plays the received audio and video data.

[0003] In order to improve the audio and video playback process of the client and thus improve the user's viewing experience, it is necessary to reasonably plan the transmission order of audio frames and video frames in the played video. Summary of the invention

[0004] The present disclosure provides a method, apparatus, device and medium for transmitting audio and video data.

[0005] According to one aspect of the present disclosure, a method for transmitting audio and video data is provided, comprising:

[0006] Determine a candidate transmission order for data frames in the audio and video data to be transmitted; wherein the data frames include audio frames and video frames;

[0007] Determining a loss parameter of the candidate transmission sequence according to a frame attribute of the data frame;

[0008] A target transmission order is determined from the candidate transmission orders according to the loss parameter, and the data frame is transmitted to the client using the target transmission order.

[0009] According to another aspect of the present disclosure, there is provided a transmission device for audio and video data, comprising:

[0010] A candidate sequence determination module, used to determine a candidate transmission sequence for data frames in the audio and video data to be transmitted; wherein the data frames include audio frames and video frames;

[0011] A loss parameter determination module, used to determine the loss parameter of the candidate transmission sequence according to the frame attribute of the data frame;

[0012] The target sequence determination module is used to determine a target transmission sequence from the candidate transmission sequences according to the loss parameter, and to transmit the data frame to the client using the target transmission sequence.

[0013] According to another aspect of the present disclosure, there is provided an electronic device, comprising:

[0014] at least one processor; and

[0015] a memory communicatively connected to the at least one processor; wherein,

[0016] The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the audio and video data transmission method described in any embodiment of the present disclosure.

[0017] According to another aspect of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to enable the computer to execute the method for transmitting audio and video data according to any embodiment of the present disclosure.

[0018] According to another aspect of the present disclosure, a computer program product is provided, including a computer program, wherein when the computer program is executed by a processor, the method for transmitting audio and video data according to any embodiment of the present disclosure is implemented.

[0019] It should be understood that the content described in this section is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it intended to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] The accompanying drawings are used to better understand the present solution and do not constitute a limitation of the present disclosure.

[0021] Figure 1 is a schematic diagram of a method for transmitting audio and video data according to an embodiment of the present disclosure;

[0022] Figure 2 is a schematic diagram of candidate transmission sequences of audio and video data to be transmitted according to an embodiment of the present disclosure;

[0023] Figure 3 is a schematic diagram of another method for transmitting audio and video data according to an embodiment of the present disclosure;

[0024] Figure 4 is a schematic diagram of another method for transmitting audio and video data according to an embodiment of the present disclosure;

[0025] Figure 5 is a structural schematic diagram of a transmission device for audio and video data according to an embodiment of the present disclosure;

[0026] Figure 6 The block diagram is a block diagram of an electronic device used to implement the method for transmitting audio and video data according to an embodiment of the present disclosure. DETAILED DESCRIPTION

[0027] The following is a description of exemplary embodiments of the present disclosure in conjunction with the accompanying drawings, including various details of the embodiments of the present disclosure to facilitate understanding, which should be considered as merely exemplary. Therefore, it should be recognized by those of ordinary skill in the art that various changes and modifications may be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.

[0028] Figure 1 This is a schematic diagram of a method for transmitting audio and video data according to an embodiment of the present disclosure. This embodiment is applicable to the case where a method for determining the transmission order of audio and video data is optimized. This method can be executed by a transmission device for audio and video data. The device can be implemented by software and / or hardware and integrated in an electronic device. The electronic device involved in this embodiment can be a server with computing and communication capabilities. Specifically, refer to Figure 1 , the method specifically includes the following:

[0029] S110. Determine a candidate transmission order for data frames in the audio and video data to be transmitted; wherein the data frames include audio frames and video frames.

[0030] The audio and video data to be transmitted refers to the audio and video data obtained from the server according to the playback requirements of the client, and the client performs the operation of downloading and playing the audio and video data transmitted by the server. For example, in a video on demand scenario or an initial cache delivery scenario of a live video, the client needs to obtain the audio and video data to be transmitted from the server, and after receiving the audio and video data, the client performs the operation of downloading and playing the audio and video data to be transmitted that is pre-cached in the server. In the network transmission scenario, in order to improve the playback experience of the client, it is necessary to determine a reasonable interleaved transmission order for the audio frames and video frames in the audio and video data to be transmitted, so as to ensure the smoothness of the client when downloading and playing the audio and video data according to the interleaved transmission order.

[0031] Specifically, the audio and video data to be transmitted is a video clip, including multiple data frames, such as at least two audio frames and at least two video frames, and the candidate transmission order is the staggered order of different combinations between the multiple data frames. Exemplarily, the audio and video data to be transmitted is divided into an audio frame sequence and a video frame sequence, and the order of audio frames in the audio frame sequence is the same as the order in the audio and video data to be transmitted. Similarly, the order of video frames in the video frame sequence is the same as the order in the audio and video data to be transmitted, so as to avoid the transmission causing the transmission order of audio or video to be disordered, which affects the playback order. The candidate transmission order is determined according to different interlaced combinations between the audio frame sequence and the video frame sequence, that is, the positional relationship between the audio frame and the video frame in different candidate transmission orders is different, but the order between the audio and video frames and the order between the video frames are the same.

[0032] like Figure 2 The figure shows a schematic diagram of candidate transmission order of audio and video data to be transmitted, wherein the audio and video data to be transmitted includes n a audio frames and n v video frames, and the audio frame sequence is The video frame sequence is x i represents the number of video frames before the i-th audio frame in the audio frame sequence, Indicates the nth in the audio frame sequence a The number of video frames after audio frames, x i and Can be equal to 0, with a length of n a +1 vector represents the video candidate transmission order of video frames, where Among them, X satisfies n a +1 continuous video frame set is The collection is The same. Then different combinations of values ​​in vector X represent different candidate transmission orders. Similarly, Figure 2 The audio frames and video frames in the video frame can also be exchanged. This embodiment is only used for illustrative purposes and does not affect the protection scope of the present disclosure.

[0033] S120. Determine loss parameters of candidate transmission sequences according to frame attributes of the data frames.

[0034] The frame attributes of the data frame include parameters that affect the data frame during transmission, such as frame attributes including at least one of the following: frame size, display timestamp, and network transmission speed. The frame size is used to indicate the size of the data frame. For example, based on the above example, the audio frame sequence The corresponding frame size is The video frame sequence is The corresponding frame size is The presentation time stamp (PTS) is used to indicate the playback time of the data frame in the client. The corresponding display timestamp is The video frame sequence is The corresponding frame size is The network transmission speed refers to the speed at which data frames are sent from the server to the client. The network transmission speed is determined based on factors such as bandwidth and the average speed within a preset time period.

[0035] The loss parameter refers to the impact parameter brought about when the audio frame sequence and the video frame sequence are transmitted according to different candidate transmission orders. The loss parameter can be considered from different impact angles, such as the impact angle brought about by different candidate transmission orders on the actual start time of the client, or the impact angle brought about by the length of unused time of the data frame on the client, the unused time length refers to the length between the time when the data frame is transmitted to the client and the time when the frame starts to be played, or the impact angle brought about by the uniformity of the interleaving between the audio frame and the video frame. Exemplarily, the loss parameter of the candidate transmission order can be considered from at least one of the following angles: the time required for each data frame to be sent to the client, the actual start time of the playback on the client, and the difference between the display timestamp of each data frame and the display playback timestamp of the first video frame that starts to be played. The loss parameters of the above angles are determined respectively according to the frame attribute information of the data frame to characterize the degree of playback impact corresponding to different candidate transmission data.

[0036] In another optional implementation of this embodiment, the frame attributes of the data frame include at least a frame size and a network transmission speed;

[0037] S120, including:

[0038] Determine the target frame corresponding to the candidate transmission sequence according to the preset client playback condition and the candidate transmission sequence; wherein the target frame includes the audio frame and / or video frame that has been received when the client playback condition is met;

[0039] A playback start time is determined for the candidate transmission sequence according to the frame size of the target frame and the network transmission speed, and the playback start time is used as a loss parameter.

[0040] Among them, the client playback condition refers to the trigger condition for the client to start playing the video. It can be that the client starts playing when it receives at least one audio frame and at least one video frame, or it can be that the client starts playing when it receives at least one audio frame or at least one video frame, or it can be that the client starts playing when it receives any data frame. The client playback condition is determined according to the playback mechanism of different clients and is not limited here. Different client playback conditions and candidate transmission sequences correspond to different target frames. For example, if the client playback condition is that the client starts playing when it receives at least one audio frame and at least one video frame, the corresponding target frame is at least one audio frame and at least one video frame; if the client playback condition is that the client starts playing when it receives at least one audio frame or at least one video frame, the corresponding target frame is at least one audio frame or at least one video frame. The number of audio frames and video frames included in the target frame is determined according to the arrangement order of the audio frames and video frames in the candidate transmission sequence.

[0041] Since the actual start time of playback on the client will affect the client's playback experience, the loss parameter is determined based on the actual start time of playback on the client. Specifically, the client starts playback after receiving the target frame, and the transmission time of the target frame is the playback start time. The transmission time of the target frame can be determined based on the frame size of the target frame and the network transmission speed, and then the playback start time is determined as the loss parameter.

[0042] The corresponding client playback start time is determined for the candidate transmission order through the client playback condition, and the client playback start time is used as the loss parameter. That is, the client playback start time corresponding to different candidate transmission orders is considered, and the target transmission order is determined according to the client playback start time, thereby improving the client's playback smoothness.

[0043] In another optional implementation of this embodiment, the client playback condition is that the client starts playing when receiving at least one audio frame and at least one video frame; the target frame includes at least one audio frame and at least one video frame;

[0044] Determine a playback start time for a candidate transmission sequence according to a frame size of a target frame and a network transmission speed, including:

[0045] Determining a first frame size according to the sum of a frame size of a target audio frame and a frame size of a target video frame in the target frame;

[0046] The playback start time is determined according to the quotient of the first frame size and the network transmission speed.

[0047] The sizes of all frames transmitted from the server to the client when the client starts playing are determined based on the sum of the frame sizes of all data frames included in the target frame, and then the quotient of all frame sizes and the network transmission speed is used as the start time of playback, that is, when the server transmits all target frames to the client, the client starts playing the received target frames.

[0048] In this embodiment, the client playback condition is that the client starts playback when it receives at least one audio frame and at least one video frame. Similarly, under other client playback conditions, the corresponding target frames are different, and the corresponding method for determining the playback start time is different. Informed by this embodiment, it is easy for those skilled in the art to think of the calculation method of the playback start time corresponding to other client playback conditions, which will not be repeated here.

[0049] The target frame is determined by the client playback conditions, and then the playback start time is determined according to the frame attributes of the target frame, which improves the accuracy of the loss parameter determination and the rationality of the target transmission sequence determination.

[0050] In another optional implementation of this embodiment, the client playback condition is that the client starts playing when receiving at least one audio frame and at least one video frame; the target frame includes at least one audio frame and at least one video frame; the playback start time L is determined according to the following formula: first :

[0051]

[0052] in, represents the frame size of the i-th audio frame in the candidate transmission order; represents the frame size of the jth video frame in the candidate transmission order; x m represents the number of video frames between the m-1th audio frame and the mth audio frame in the candidate transmission order; V represents the network transmission speed; m satisfies and That is, when m satisfies this condition, it satisfies the client playback condition of receiving at least one audio frame and at least one video frame.

[0053] S130. Determine a target transmission sequence from candidate transmission sequences according to the loss parameter, and transmit data frames to the client using the target transmission sequence.

[0054] The target transmission order is determined from the candidate transmission orders based on the comparison results of the loss parameters corresponding to different candidate transmission orders. Exemplarily, when the loss parameter is the playback start time, the playback start time measures the first frame playback time of the client caused by different arrangements. Therefore, the smaller the playback start time, the better the playback smoothness for the client. The candidate transmission orders are sorted in ascending order according to the playback start time, and the candidate transmission order with the smallest playback start time is selected as the target transmission order, and the audio and video data to be transmitted are transmitted to the client according to the target transmission order.

[0055] The solution of this embodiment determines the loss parameters of different candidate transmission orders through the frame attributes of the data frame, and determines the target transmission order from the candidate transmission orders based on the loss parameters. The accuracy and rationality of the target transmission order determination are improved through quantitative analysis, thereby improving the smoothness of the client's video playback.

[0056] Figure 3 is a schematic diagram of another method for transmitting audio and video data according to an embodiment of the present disclosure. This embodiment is a further refinement of the above technical solution, and the frame attribute of the data frame also includes a display timestamp. The technical solution in this embodiment can be combined with each optional solution in one or more of the above embodiments. Figure 3 As shown, the transmission method of audio and video data includes the following:

[0057] S210. Determine a candidate transmission order for data frames in the audio and video data to be transmitted; wherein the data frames include audio frames and video frames.

[0058] S220: Determine the redundant time difference between the play time and the transmission arrival time of the data frame according to the frame size, network transmission speed and display timestamp of the data frame and the candidate transmission sequence.

[0059] The play time of a data frame refers to the actual play time of the data frame determined according to the transmission order of the data frame in the candidate transmission order and the client playback conditions; the transmission arrival time of a data frame refers to the arrival time of the data frame from the server to the client determined according to the transmission order of the data frame in the candidate transmission order. The redundant time difference between the play time of a data frame and the transmission arrival time refers to the unused time of the data frame on the client, that is, it represents the time difference between the arrival of the data frame at the client and the actual play. This time difference can determine the rationality of the candidate transmission order. If the redundant time difference of the data frame is very small or negative, it will cause playback jams and lack of picture or sound signals.

[0060] Specifically, the redundant time difference between the play time and the transmission arrival time of each data frame corresponding to different candidate transmission sequences is determined according to the frame size of the data frame, the network transmission speed and the display timestamp. Exemplarily, the play time of each data frame is determined according to the position information of each data frame in each candidate transmission sequence and the client playback start time, and the transmission arrival time of each data frame is determined according to the position information of each data frame in each candidate transmission sequence, and the difference between the play time and the transmission arrival time of each data frame is the redundant time difference of the data frame.

[0061] In another optional implementation of this embodiment, S220 includes:

[0062] Determine a preamble frame preceding the data frame according to the candidate transmission sequence;

[0063] Determine the transmission arrival time of the data frame according to the frame size of the preceding frame, the frame size of the data frame and the network transmission speed;

[0064] Determine the playback time of the data frame according to the playback start time, the display timestamp of the data frame and the display timestamp of the start frame of the corresponding type of the data frame;

[0065] The redundant time difference of the data frame is determined according to the difference between the play time and the transmission arrival time.

[0066] Among them, the preceding frame corresponding to the data frame refers to all audio frames and video frames that are located before the data frame in the candidate transmission order. The playback start time refers to the time when the client starts playing, and the method for determining the playback time is: according to the preset client playback conditions and the candidate transmission order, determine the target frame corresponding to the candidate transmission order; wherein the target frame includes the audio frames and / or video frames that have been received when the client playback conditions are met; determine the playback start time for the candidate transmission order according to the frame size of the target frame and the network transmission speed. The specific determination process refers to Example 1 and will not be repeated here. The starting frame of the corresponding type of the data frame refers to the first frame in the frame sequence of the corresponding type of the data frame. For example, if the data frame is an audio frame, the starting frame of the corresponding type of the data frame is A1.

[0067] Specifically, according to the candidate transmission order and the frame size and network transmission speed in the frame attributes of the data frame, the total frame size required to be transmitted when transmitting to the data frame when transmitting according to the candidate transmission order is determined, and then the transmission arrival time corresponding to the data frame is determined according to the total frame size and the network transmission speed. Similarly, according to the candidate transmission order and the client playback conditions, the playback start time is determined, and then the playback time difference of the data frame is determined according to the display timestamp of the data frame and the display timestamp of the start frame of the corresponding type of the data frame, and finally the playback time of the data frame is determined according to the playback start time and the playback time difference. The redundant time difference of the data frame is determined according to the difference between the playback time and the transmission arrival time of the data frame.

[0068] By determining the playback time and transmission arrival time of each data frame, and using the redundant time difference between the playback time and transmission arrival time to determine the loss parameter, the loss parameter can characterize the overall playback redundancy time of data frames in different candidate transmission sequences, which is beneficial to reducing the client rendering jamming phenomenon when audio and video data are sent down according to the target transmission sequence, and improving the client playback smoothness.

[0069] In another optional implementation of this embodiment, determining the transmission arrival time of the data frame according to the frame size of the preceding frame, the frame size of the data frame, and the network transmission speed includes:

[0070] Determine a second frame size according to the sum of the frame size of the data frame and the frame size of the preceding frame;

[0071] The transmission arrival time of the data frame is determined according to the quotient of the second frame size and the network transmission speed.

[0072] The second frame size determined by the sum of the frame size of the data frame and the frame size of the preceding frame corresponding to the data frame represents the total frame size required to be transmitted when the data frame is transmitted according to the candidate transmission sequence. The transmission arrival time of the data frame to the client is obtained according to the quotient of the second frame size and the network transmission speed.

[0073] The transmission arrival time is determined by the frame size of the data frame and the network transmission speed, thereby improving the accuracy of determining the redundant time difference.

[0074] In another optional implementation of this embodiment, if the data frame is an audio frame, the transmission arrival time of the audio frame is determined according to the following formula:

[0075]

[0076] If the data frame is a video frame, the transmission arrival time of the video frame is determined according to the following formula:

[0077]

[0078] in, represents the transmission arrival time of the i-th audio frame in the candidate transmission sequence; represents the transmission arrival time of the kth video frame between the i-1th audio frame and the ith audio frame in the candidate transmission order; represents the frame size of the jth audio frame in the candidate transmission order; represents the frame size of the jth video frame in the candidate transmission order; x q represents the number of video frames between the q-1th audio frame and the qth audio frame in the candidate transmission order; V represents the network transmission speed.

[0079] Specifically, according to Figure 2 When the candidate transmission order shown is transmitted, the formula for determining the transmission arrival time of the corresponding audio frame and video frame is as shown above.

[0080] In another optional implementation of this embodiment, determining the playback time of the data frame according to the playback start time, the display timestamp of the data frame, and the display timestamp of the start frame of the corresponding type of the data frame includes:

[0081] Determine the display time difference according to the difference between the display time stamp of the data frame and the display time stamp of the start frame of the corresponding type of the data frame;

[0082] The playback time of the data frame is determined according to the sum of the display time difference and the playback start time.

[0083] Since the display timestamp indicates the time information of the data frame being played on the client, the difference between the display timestamp of the data frame and the display timestamp of the start frame of the corresponding type can be used to determine the play time difference between the data frame and the first frame in the frame sequence of the corresponding type. The play start time is determined according to the client play condition, and the play time of the data frame under the client play condition can be determined according to the sum of the play start time and the display time difference. This embodiment is described based on the play start condition that the client starts playing after receiving at least one audio frame and at least one video frame, and the play time determination method corresponding to other client play conditions is not repeated here.

[0084] The playback time is determined by the display timestamp of the data frame and the playback start time, thereby improving the accuracy of determining the redundant time difference.

[0085] In another optional implementation of this embodiment, if the data frame is an audio frame, the play time of the audio frame is determined according to the following formula:

[0086]

[0087] If the data frame is a video frame, the playback time of the video frame is determined according to the following formula:

[0088]

[0089] in, represents the playback time of the i-th audio frame in the candidate transmission order; represents the playback time of the kth video frame between the i-1th audio frame and the ith audio frame in the candidate transmission order; represents the presentation timestamp of the i-th audio frame in the candidate transmission order; Indicates the first Display timestamp of each video frame; Indicates the display timestamp of the first audio frame in the candidate transmission order, that is, the display timestamp of A1. Indicates the display timestamp of the first video frame in the candidate transmission order, that is, the display timestamp of V1.

[0090] S230, determining a redundant loss parameter for the candidate transmission sequence according to the redundant time difference of the data frame, and using the redundant loss parameter as a loss parameter.

[0091] The redundant loss parameter corresponding to the candidate transmission order is determined according to the redundant time difference of each data frame in the audio and video data to be transmitted, and the redundant loss parameter is used to characterize the overall level of the data frame playback and transmission time difference corresponding to the candidate transmission order. The larger the redundant loss parameter is, the larger the redundant time difference of most data frames is, and the better the playback effect of the client when transmitting according to the candidate transmission order.

[0092] In another optional implementation of this embodiment, S230 includes:

[0093] Determine the total redundant time according to the sum of the redundant time differences of the data frames;

[0094] A redundancy loss parameter is determined for the candidate transmission sequence based on the quotient of the total redundancy time and the total number of data frames.

[0095] The total redundant time corresponding to the audio and video data to be transmitted when it is transmitted according to the candidate transmission order is determined according to the sum of the redundant time differences of each data frame in the audio and video data to be transmitted, and the average redundant time value determined according to the quotient of the total redundant time and the total number of all data frames in the audio and video data to be transmitted is used as the redundant loss parameter corresponding to the candidate transmission order, that is, the redundant loss parameter characterizes the average redundant time difference of the data frames in the audio and video data to be transmitted.

[0096] The redundant loss parameter is determined according to the average value of the redundant time difference of the data frames in the audio and video data to be transmitted, which improves the accuracy of the redundant loss parameter in representing the overall redundancy level of the audio and video data to be transmitted, thereby improving the rationality of determining the target transmission sequence.

[0097] In another optional implementation of this embodiment, the redundancy loss parameter L is determined according to the following formula: redundancy :

[0098]

[0099] in, represents the redundant time difference of the i-th audio frame in the candidate transmission order, represents the redundant time difference of the kth video frame between the i-1th audio frame and the i-th audio frame in the candidate transmission order, x i Indicates the number of video frames between the i-1th audio frame and the i-th audio frame in the candidate transmission order; n a Indicates the total number of audio frames in the audio and video frame data sequence to be transmitted; n v Indicates the total number of video frames in the audio and video frame data sequence to be transmitted.

[0100]

[0101] S240. Determine a target transmission sequence from candidate transmission sequences according to the loss parameter, and transmit data frames to the client using the target transmission sequence.

[0102] When the loss parameter is a redundant loss parameter, the redundant loss parameter represents the overall level of the difference between the data frame playback time and the transmission arrival time corresponding to the candidate transmission sequence. Therefore, the larger the redundant loss parameter, the better the playback smoothness for the client. The candidate transmission sequences are sorted in descending order according to the redundant loss parameter, and the candidate transmission sequence with the largest redundant loss parameter is selected as the target transmission sequence, and the audio and video data to be transmitted are transmitted to the client according to the target transmission sequence.

[0103] The scheme of this embodiment determines the redundant loss parameters of different candidate transmission orders through the frame attributes of the data frames. The redundant loss parameters characterize the overall level of the difference between the data frame playback time and the transmission arrival time corresponding to the candidate transmission order. The target transmission order is determined from the candidate transmission orders based on the redundant loss parameters. The accuracy and rationality of the determination of the target transmission order are improved through the quantitative analysis of the redundant time difference, thereby improving the smoothness of the client's video playback.

[0104] Figure 4 is a schematic diagram of another method for transmitting audio and video data according to an embodiment of the present disclosure. This embodiment is a further refinement of the above technical solution. The frame attributes of the data frame at least include a display timestamp. The technical solution in this embodiment can be combined with various optional solutions in one or more of the above embodiments. Figure 4 As shown, the transmission method of audio and video data includes the following:

[0105] S310, determining a candidate transmission order for data frames in the audio and video data to be transmitted; wherein the data frames include audio frames and video frames.

[0106] S320. Determine a reference frame corresponding to the data frame according to position information of the data frame in the candidate transmission sequence; wherein the data frame and the reference frame corresponding to the data frame are of different types.

[0107] The reference frame corresponding to the data frame is used to represent other frames of different types adjacent to the data frame, that is, there are no other frames of the same type as the data frame between the reference frame and the data frame. For example, when the data frame is an audio frame, the reference frame corresponding to the audio frame is the target video frame adjacent to the audio frame before and after, and the target video frame satisfies the condition that there are no other audio frames between the audio frame; similarly, when the data frame is a video frame, the reference frame corresponding to the video frame is the target audio frame adjacent to the video frame before and after, and the target audio frame satisfies the condition that there are no other video frames between the video frame.

[0108] Specifically, according to the position information of the data frame in the candidate transmission sequence, different types of frames adjacent to and before the data frame are determined as reference frames.

[0109] In another optional implementation of this embodiment, S320 includes:

[0110] Determine a previous frame and a next frame of the same type as the data frame according to the type of the data frame and the candidate transmission sequence;

[0111] The frame between the previous frame and the data frame and the frame between the data frame and the next frame are determined as reference frames corresponding to the data frame according to the candidate transmission sequence.

[0112] The frames between the previous frame and the next frame of the same type as the data frame and the data frame are the target audio frame and the target video frame, ensuring that there are no other frames of the same type as the data frame between the target audio frame and the target video frame and the data frame.

[0113] The reference frame corresponding to the data frame is determined according to the previous frame and the next frame of the same type as the data frame, thereby improving the accuracy of determining the reference frame, thereby improving the accuracy of determining the interleaving loss parameter between the audio frame and the video frame.

[0114] S330 , determining an interlaced loss parameter according to the display timestamp of the reference frame, the display timestamp of the data frame, and the total number of frames of the same type corresponding to the data frame, and using the interlaced loss parameter as a loss parameter.

[0115] The reference frame corresponding to the data frame is the adjacent frame of the data frame. The interleaving time difference between the audio frame and the video frame can be determined based on the display timestamp of the data frame and the display timestamp of the reference frame. The average value of the interleaving time difference can be determined by combining the total number of the same type corresponding to the data frame, and the average value is used as the interleaving loss parameter.

[0116] In another optional implementation of this embodiment, S330 includes:

[0117] Determine the interleaving time difference according to the sum of squares of differences between the display time stamp of the data frame and the display time stamp of each reference frame;

[0118] The interleaving loss parameter is determined according to the quotient of the interleaving time difference and the total number of frames of the same type corresponding to the data frame.

[0119] Specifically, when determining the interleaving loss parameters of the candidate transmission order, one type of data frame is selected as the object for calculating the interleaving time difference. For example, an audio frame is selected to calculate the interleaving loss parameter, and the reference frame corresponding to each audio frame in the audio and video data to be transmitted is determined, and the interleaving time difference is determined based on the sum of the squares of the difference in display timestamps between each audio frame and the corresponding reference frame. The interleaving time difference reflects the display time distance information between the audio frame and the video frame. The interleaving loss parameter is determined by the quotient of the interleaving time difference and the total number of audio frames. The interleaving loss parameter reflects the average display time distance between the audio frame and the video, and can also characterize the interleaving uniformity between the audio frame and the video frame.

[0120] The interleaving time difference is determined by data frames of the same type and corresponding reference frames, and the interleaving loss parameter is then determined based on the total number of data frames of this type. This improves the accuracy of determining the interleaving loss parameter and the accuracy of characterizing the interleaving uniformity between audio frames and video frames through the interleaving loss parameter.

[0121] In another optional implementation of this embodiment, the data frame is an audio frame, and the reference frame is a video frame;

[0122] The interleaving loss parameter is determined according to the following formula:

[0123]

[0124]

[0125] Among them, n a Indicates the total number of audio frames in the audio and video frame data sequence to be transmitted; xq represents the number of video frames between the q-1th audio frame and the qth audio frame in the candidate transmission order; represents the presentation timestamp of the i-th audio frame in the candidate transmission order; Indicates the presentation timestamp of the j-th video frame in the candidate transmission order.

[0126] According to Figure 2 For the candidate transmission order shown in , for the i-th audio frame, the corresponding reference frames are two sets of video frames, namely and Corresponding to Video frames to video frames, Corresponding to Video frames to video frames. i Indicates the interleaving time difference of the i-th audio frame.

[0127] S340. Determine a target transmission sequence from candidate transmission sequences according to the loss parameter, and transmit data frames to the client using the target transmission sequence.

[0128] When the loss parameter is an interleaving loss parameter, the interleaving loss parameter characterizes the interleaving uniformity between the audio frames and video frames corresponding to the candidate transmission order. Therefore, the smaller the interleaving loss parameter, the more uniform the audio and video interleaving, and the better the playback smoothness for the client. The candidate transmission orders are sorted in ascending order according to the interleaving loss parameter, and the candidate transmission order with the smallest interleaving loss parameter is selected as the target transmission order, and the audio and video data to be transmitted are transmitted to the client according to the target transmission order.

[0129] The scheme of this embodiment determines the interleaving loss parameters of different candidate transmission orders through the frame attributes of the data frame. The interleaving loss parameters characterize the overall interleaving uniformity between the audio frames and the video frames corresponding to the candidate transmission order. The target transmission order is determined from the candidate transmission orders based on the interleaving loss parameters. The accuracy and rationality of the determination of the target transmission order are improved through quantitative analysis of the interleaving uniformity, thereby improving the smoothness of the video playback on the client.

[0130] In the embodiment of the present disclosure, the loss parameter includes at least one of the playback start time, the redundant loss parameter or the interleaved loss parameter, or may be any combination of two or three thereof. For example, the loss parameter includes the playback start time, the redundant loss parameter and the interleaved loss parameter.

[0131] A target transmission sequence is determined from candidate transmission sequences according to a playback start time, a redundancy loss parameter, and an interleaving loss parameter, and data frames are transmitted to a client using the target transmission sequence.

[0132] Specifically, a total loss parameter is determined according to the playback start time, the redundant loss parameter and the interleaving loss parameter, and a target transmission sequence is determined from candidate transmission sequences according to the total loss parameter.

[0133] Exemplarily, based on the above example, the expression of the total loss parameter is α1L first -α2L redundancy +α3L distance , α1, α2 and α3 are the weight information of the corresponding loss parameters, which can be set according to the actual scenario. Then, the target transmission order is determined from the candidate transmission order according to the total loss parameter, which is the min argx (α1L first -α2L redundancy +α3L distance ) is a problem of solving. Vector The video candidate transmission order representing the video frames. When the number of frames of audio and video data to be transmitted is small, the value of the vector X can be directly determined by enumeration. When the number of frames of audio and video data to be transmitted is large, it can be solved by a greedy algorithm, a Monte Carlo algorithm, or a gradient descent algorithm, and the solution algorithm is not limited here.

[0134] Figure 5 is a structural diagram of a transmission device for audio and video data according to an embodiment of the present disclosure, and the device can execute the transmission method for audio and video data involved in any embodiment of the present disclosure; Figure 5 The audio and video data transmission device 400 includes: a candidate sequence determination module 410, a loss parameter determination module 420 and a target sequence determination module 430.

[0135] The candidate sequence determination module 410 is used to determine a candidate transmission sequence for data frames in the audio and video data to be transmitted; wherein the data frames include audio frames and video frames;

[0136] A loss parameter determination module 420, configured to determine a loss parameter of the candidate transmission sequence according to a frame attribute of the data frame;

[0137] The target sequence determination module 430 is used to determine a target transmission sequence from the candidate transmission sequences according to the loss parameter, and transmit the data frame to the client using the target transmission sequence.

[0138] The solution of this embodiment determines the loss parameters of different candidate transmission orders through the frame attributes of the data frame, and determines the target transmission order from the candidate transmission orders based on the loss parameters. The accuracy and rationality of the target transmission order determination are improved through quantitative analysis, thereby improving the smoothness of the client's video playback.

[0139] In an optional implementation of this embodiment, the frame attributes of the data frame include at least a frame size and a network transmission speed;

[0140] The loss parameter determination module comprises:

[0141] A target frame determination unit, configured to determine a target frame corresponding to the candidate transmission sequence according to a preset client playback condition and the candidate transmission sequence; wherein the target frame includes an audio frame and / or a video frame that has been received when the client playback condition is met;

[0142] A playback start time determination unit is used to determine the playback start time for the candidate transmission sequence according to the frame size of the target frame and the network transmission speed, and use the playback start time as the loss parameter.

[0143] In an optional implementation of this embodiment, the client playback condition is that the client starts playing when receiving at least one audio frame and at least one video frame; the target frame includes at least one audio frame and at least one video frame;

[0144] The playback start time determination unit is specifically used to:

[0145] Determining a first frame size according to the sum of a frame size of a target audio frame and a frame size of a target video frame in the target frame;

[0146] The playback start time is determined according to a quotient of the first frame size and the network transmission speed.

[0147] In an optional implementation of this embodiment, the frame attributes of the data frame further include a display timestamp;

[0148] The loss parameter determination module further includes:

[0149] a redundant time difference determining unit, configured to determine a redundant time difference between a play time and a transmission arrival time of the data frame according to a frame size, a network transmission speed and a display timestamp of the data frame, and the candidate transmission sequence;

[0150] The redundant loss determining unit is used to determine a redundant loss parameter for the candidate transmission sequence according to the redundant time difference of the data frame, and use the redundant loss parameter as the loss parameter.

[0151] In an optional implementation of this embodiment, the redundant time difference determining unit includes:

[0152] A preamble frame determination subunit, configured to determine a preamble frame preceding the data frame according to the candidate transmission sequence;

[0153] a transmission arrival time determination subunit, configured to determine the transmission arrival time of the data frame according to the frame size of the preceding frame, the frame size of the data frame and the network transmission speed;

[0154] a play time determination subunit, configured to determine the play time of the data frame according to the play start time, the display timestamp of the data frame, and the display timestamp of the start frame of the type corresponding to the data frame;

[0155] The redundant time difference determining subunit is used to determine the redundant time difference of the data frame according to the difference between the playing time and the transmission arrival time.

[0156] In an optional implementation of this embodiment, the redundancy loss determining unit includes:

[0157] A total redundant time determination subunit, used to determine the total redundant time according to the sum of the redundant time differences of the data frames;

[0158] The redundancy loss determination subunit is used to determine a redundancy loss parameter for the candidate transmission sequence according to a quotient of the total redundancy time and the total number of the data frames.

[0159] In an optional implementation of this embodiment, the transmission arrival time determination subunit is specifically configured to:

[0160] Determine a second frame size according to the sum of the frame size of the data frame and the frame size of the preceding frame;

[0161] The transmission arrival time of the data frame is determined according to the quotient of the second frame size and the network transmission speed.

[0162] In an optional implementation of this embodiment, the play time determination subunit is specifically used to:

[0163] Determine a display time difference according to a difference between a display time stamp of the data frame and a display time stamp of a start frame of a type corresponding to the data frame;

[0164] The play time of the data frame is determined according to the sum of the display time difference and the play start time.

[0165] In an optional implementation of this embodiment, the frame attributes of the data frame include at least a display timestamp;

[0166] The loss parameter determination module comprises:

[0167] A reference frame determining unit, configured to determine a reference frame corresponding to the data frame according to position information of the data frame in the candidate transmission sequence; wherein the data frame and the reference frame corresponding to the data frame are of different types;

[0168] The interlacing loss determining unit is used to determine an interlacing loss parameter according to the display timestamp of the reference frame, the display timestamp of the data frame and the total number of frames of the same type corresponding to the data frame, and use the interlacing loss parameter as the loss parameter.

[0169] In an optional implementation of this embodiment, the reference frame determination unit is specifically configured to:

[0170] Determine, according to the type of the data frame and the candidate transmission order, a previous frame and a next frame of the same type as the data frame;

[0171] A frame between the previous frame and the data frame and a frame between the data frame and the next frame are determined as reference frames corresponding to the data frame according to the candidate transmission order.

[0172] In an optional implementation of this embodiment, the interleaving loss determination unit is specifically configured to:

[0173] Determine the interleaving time difference according to the sum of squares of differences between the display timestamp of the data frame and the display timestamps of each reference frame;

[0174] An interleaving loss parameter is determined according to a quotient of the interleaving time difference and a total number of frames of the same type corresponding to the data frame.

[0175] The above-mentioned audio and video data transmission device can execute the audio and video data transmission method provided by any embodiment of the present disclosure, and has the corresponding functional modules and beneficial effects of the execution method. For technical details not fully described in this embodiment, please refer to the audio and video data transmission method provided by any embodiment of the present disclosure.

[0176] In the technical solution of the present disclosure, the collection, storage, use, processing, transmission, provision and disclosure of user personal information involved are in compliance with the provisions of relevant laws and regulations and do not violate public order and good morals.

[0177] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium and a computer program product.

[0178] Figure 6A schematic block diagram of an example electronic device 500 that can be used to implement an embodiment of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processing, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present disclosure described and / or required herein.

[0179] like Figure 6 As shown, the device 500 includes a computing unit 501, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 502 or a computer program loaded from a storage unit 508 into a random access memory (RAM) 503. In the RAM 503, various programs and data required for the operation of the device 500 can also be stored. The computing unit 501, the ROM 502, and the RAM 503 are connected to each other via a bus 504. An input / output (I / O) interface 505 is also connected to the bus 504.

[0180] A number of components in the device 500 are connected to the I / O interface 505, including: an input unit 506, such as a keyboard, a mouse, etc.; an output unit 507, such as various types of displays, speakers, etc.; a storage unit 508, such as a disk, an optical disk, etc.; and a communication unit 509, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 509 allows the device 500 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.

[0181] The computing unit 501 may be a variety of general and / or special processing components with processing and computing capabilities. Some examples of the computing unit 501 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, digital signal processors (DSPs), and any appropriate processors, controllers, microcontrollers, etc. The computing unit 501 performs the various methods and processes described above, such as a method for transmitting audio and video data. For example, in some embodiments, the method for transmitting audio and video data may be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as a storage unit 508. In some embodiments, part or all of the computer program may be loaded and / or installed on the device 500 via ROM 502 and / or communication unit 509. When the computer program is loaded into RAM 503 and executed by the computing unit 501, one or more steps of the method for transmitting audio and video data described above may be performed. Alternatively, in other embodiments, the computing unit 501 may be configured to perform the transmission of method audio and video data by any other appropriate means (e.g., by means of firmware).

[0182] Various implementations of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chips (SOCs), load programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various implementations can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.

[0183] The program code for implementing the method of the present disclosure may be written in any combination of one or more programming languages. These program codes may be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device, so that the program code, when executed by the processor or controller, enables the functions / operations specified in the flow chart and / or block diagram to be implemented. The program code may be executed entirely on the machine, partially on the machine, partially on the machine and partially on a remote machine as a stand-alone software package, or entirely on a remote machine or server.

[0184] In the context of the present disclosure, a machine-readable medium may be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, device, or equipment. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium may include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0185] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).

[0186] The systems and techniques described herein may be implemented in a computing system that includes backend components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes frontend components (e.g., a user computer with a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such backend components, middleware components, or frontend components. The components of the system may be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), a blockchain network, and the Internet.

[0187] A computer system may include a client and a server. The client and the server are generally remote from each other and usually interact through a communication network. The relationship of client and server is generated by computer programs running on respective computers and having a client-server relationship with each other. The server may be a cloud server, a server of a distributed system, or a server combined with a blockchain.

[0188] It should be understood that the various forms of processes shown above can be used to reorder, add or delete steps. For example, the steps recorded in this disclosure can be executed in parallel, sequentially or in different orders, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved, and this document does not limit this.

[0189] The above specific implementations do not constitute a limitation on the protection scope of the present disclosure. It should be understood by those skilled in the art that various modifications, combinations, sub-combinations and substitutions can be made according to design requirements and other factors. Any modification, equivalent substitution and improvement made within the spirit and principle of the present disclosure shall be included in the protection scope of the present disclosure.

Claims

1. A method for transmitting audio and video data, comprising: Determine a candidate transmission order for data frames in the audio and video data to be transmitted; wherein the data frames include at least two audio frames and at least two video frames, and the candidate transmission order is an interleaved order of different combinations between multiple data frames; Determine the loss parameters of the candidate transmission sequence according to the frame attributes of the data frame; wherein the loss parameters include at least two of the following: a playback start time, a redundancy loss parameter, and an interleaving loss parameter, wherein the playback start time measures the first frame playback duration of the client caused by the arrangement of different candidate transmission sequences, the redundancy loss parameter represents the overall level of the data frame playback and transmission time difference corresponding to different candidate transmission sequences, and the interleaving loss parameter represents the average display time distance between audio frames and video frames in different candidate transmission sequences; A target transmission order is determined from the candidate transmission orders according to the loss parameter, and the data frame is transmitted to the client using the target transmission order.

2. The method according to claim 1, wherein: The frame attributes of the data frame include at least frame size and network transmission speed; The step of determining the loss parameter of the candidate transmission sequence according to the frame attribute of the data frame comprises: Determine a target frame corresponding to the candidate transmission sequence according to a preset client playback condition and the candidate transmission sequence; wherein the target frame includes an audio frame and / or a video frame that has been received when the client playback condition is met; A play start time is determined for the candidate transmission sequence according to the frame size of the target frame and the network transmission speed, and the play start time is used as the loss parameter.

3. The method according to claim 2, wherein: The client playback condition is that the client starts playing after receiving at least one audio frame and at least one video frame; the target frame includes at least one audio frame and at least one video frame; The step of determining the playback start time for the candidate transmission sequence according to the frame size of the target frame and the network transmission speed includes: Determining a first frame size according to the sum of a frame size of a target audio frame and a frame size of a target video frame in the target frame; The playback start time is determined according to a quotient of the first frame size and the network transmission speed.

4. The method according to claim 2 or 3, wherein: The frame attributes of the data frame also include a display timestamp; The step of determining the loss parameter of the candidate transmission sequence according to the frame attribute of the data frame further includes: Determine a redundant time difference between a play time and a transmission arrival time of the data frame according to a frame size, a network transmission speed and a display timestamp of the data frame, and the candidate transmission sequence; A redundant loss parameter is determined for the candidate transmission sequence according to the redundant time difference of the data frame, and the redundant loss parameter is used as the loss parameter.

5. The method according to claim 4, wherein: The determining, according to the frame size, network transmission speed and display timestamp of the data frame, and the candidate transmission sequence, of the redundant time difference between the play time and the transmission arrival time of the data frame comprises: Determine, according to the candidate transmission order, a preceding frame located before the data frame; Determining the transmission arrival time of the data frame according to the frame size of the preceding frame, the frame size of the data frame and the network transmission speed; Determine the playback time of the data frame according to the playback start time, the display timestamp of the data frame, and the display timestamp of the start frame of the type corresponding to the data frame; The redundant time difference of the data frame is determined according to the difference between the play time and the transmission arrival time.

6. The method according to claim 4, wherein: The determining of a redundancy loss parameter for the candidate transmission sequence according to the redundant time difference of the data frame comprises: Determine the total redundant time according to the sum of the redundant time differences of the data frames; A redundancy loss parameter is determined for the candidate transmission sequence according to a quotient of the total redundancy time and the total number of the data frames.

7. The method according to claim 5, wherein: The determining the transmission arrival time of the data frame according to the frame size of the preceding frame, the frame size of the data frame and the network transmission speed comprises: Determine a second frame size according to the sum of the frame size of the data frame and the frame size of the preceding frame; The transmission arrival time of the data frame is determined according to the quotient of the second frame size and the network transmission speed.

8. The method according to claim 5, wherein: The determining the playback time of the data frame according to the playback start time, the display timestamp of the data frame, and the display timestamp of the start frame of the type corresponding to the data frame comprises: Determine a display time difference according to a difference between a display time stamp of the data frame and a display time stamp of a start frame of a type corresponding to the data frame; The play time of the data frame is determined according to the sum of the display time difference and the play start time.

9. The method according to claim 1, wherein: The frame attributes of the data frame at least include a display timestamp; The step of determining the loss parameter of the candidate transmission sequence according to the frame attribute of the data frame comprises: Determine, according to the position information of the data frame in the candidate transmission sequence, a reference frame corresponding to the data frame; wherein the data frame and the reference frame corresponding to the data frame are of different types; An interlaced loss parameter is determined according to the display timestamp of the reference frame, the display timestamp of the data frame, and the total number of frames of the same type corresponding to the data frame, and the interlaced loss parameter is used as the loss parameter.

10. The method according to claim 9, wherein: The determining, according to the position information of the data frame in the candidate transmission sequence, a reference frame corresponding to the data frame comprises: Determine, according to the type of the data frame and the candidate transmission order, a previous frame and a next frame of the same type as the data frame; A frame between the previous frame and the data frame and a frame between the data frame and the next frame are determined as reference frames corresponding to the data frame according to the candidate transmission order.

11. The method according to claim 9 or 10, wherein: The determining of the interlaced loss parameter according to the display timestamp of the reference frame, the display timestamp of the data frame, and the total number of frames of the same type corresponding to the data frame includes: Determine the interleaving time difference according to the sum of squares of differences between the display timestamp of the data frame and the display timestamps of each reference frame; An interleaving loss parameter is determined according to a quotient of the interleaving time difference and a total number of frames of the same type corresponding to the data frame.

12. A transmission device for audio and video data, comprising: A candidate sequence determination module, used to determine a candidate transmission sequence for data frames in the audio and video data to be transmitted; wherein the data frames include at least two audio frames and at least two video frames, and the candidate transmission sequence is an interleaved sequence of different combinations between multiple data frames; A loss parameter determination module, used to determine the loss parameters of the candidate transmission sequence according to the frame attributes of the data frame; wherein the loss parameters include at least two of the following: a playback start time, a redundant loss parameter and an interleaving loss parameter, wherein the playback start time measures the first frame playback duration of the client caused by the arrangement of different candidate transmission sequences, the redundant loss parameter represents the overall level of the data frame playback and transmission time difference corresponding to different candidate transmission sequences, and the interleaving loss parameter represents the average display time distance between audio frames and video frames in different candidate transmission sequences; The target sequence determination module is used to determine a target transmission sequence from the candidate transmission sequences according to the loss parameter, and to transmit the data frame to the client using the target transmission sequence.

13. The device according to claim 12, wherein: The frame attributes of the data frame include at least frame size and network transmission speed; The loss parameter determination module comprises: A target frame determination unit, configured to determine a target frame corresponding to the candidate transmission sequence according to a preset client playback condition and the candidate transmission sequence; wherein the target frame includes an audio frame and / or a video frame that has been received when the client playback condition is met; A playback start time determination unit is used to determine the playback start time for the candidate transmission sequence according to the frame size of the target frame and the network transmission speed, and use the playback start time as the loss parameter.

14. The device according to claim 13, wherein: The client playback condition is that the client starts playing after receiving at least one audio frame and at least one video frame; the target frame includes at least one audio frame and at least one video frame; The playback start time determination unit is specifically used to: Determining a first frame size according to the sum of a frame size of a target audio frame and a frame size of a target video frame in the target frame; The playback start time is determined according to a quotient of the first frame size and the network transmission speed.

15. The device according to claim 13 or 14, wherein: The frame attributes of the data frame also include a display timestamp; The loss parameter determination module further includes: a redundant time difference determining unit, configured to determine a redundant time difference between a play time and a transmission arrival time of the data frame according to a frame size, a network transmission speed and a display timestamp of the data frame, and the candidate transmission sequence; The redundant loss determining unit is used to determine a redundant loss parameter for the candidate transmission sequence according to the redundant time difference of the data frame, and use the redundant loss parameter as the loss parameter.

16. The device according to claim 15, wherein: The redundant time difference determining unit comprises: A preamble frame determination subunit, configured to determine a preamble frame preceding the data frame according to the candidate transmission sequence; a transmission arrival time determination subunit, configured to determine the transmission arrival time of the data frame according to the frame size of the preceding frame, the frame size of the data frame and the network transmission speed; a play time determination subunit, configured to determine the play time of the data frame according to the play start time, the display timestamp of the data frame, and the display timestamp of the start frame of the type corresponding to the data frame; The redundant time difference determining subunit is used to determine the redundant time difference of the data frame according to the difference between the playing time and the transmission arrival time.

17. The device according to claim 15, wherein: The redundancy loss determination unit comprises: A total redundant time determination subunit, used to determine the total redundant time according to the sum of the redundant time differences of the data frames; The redundancy loss determination subunit is used to determine a redundancy loss parameter for the candidate transmission sequence according to a quotient of the total redundancy time and the total number of the data frames.

18. The device according to claim 16, wherein: The transmission arrival time determination subunit is specifically used for: Determine a second frame size according to the sum of the frame size of the data frame and the frame size of the preceding frame; The transmission arrival time of the data frame is determined according to the quotient of the second frame size and the network transmission speed.

19. The device according to claim 16, wherein: The play time determination subunit is specifically used for: Determine a display time difference according to a difference between a display time stamp of the data frame and a display time stamp of a start frame of a type corresponding to the data frame; The play time of the data frame is determined according to the sum of the display time difference and the play start time.

20. The device according to claim 12, wherein: The frame attributes of the data frame at least include a display timestamp; The loss parameter determination module comprises: A reference frame determining unit, configured to determine a reference frame corresponding to the data frame according to position information of the data frame in the candidate transmission sequence; wherein the data frame and the reference frame corresponding to the data frame are of different types; The interlacing loss determining unit is used to determine an interlacing loss parameter according to the display timestamp of the reference frame, the display timestamp of the data frame and the total number of frames of the same type corresponding to the data frame, and use the interlacing loss parameter as the loss parameter.

21. The device according to claim 20, wherein: The reference frame determination unit is specifically used to: Determine, according to the type of the data frame and the candidate transmission order, a previous frame and a next frame of the same type as the data frame; A frame between the previous frame and the data frame and a frame between the data frame and the next frame are determined as reference frames corresponding to the data frame according to the candidate transmission order.

22. The device according to claim 20 or 21, wherein: The interleaving loss determination unit is specifically configured to: Determine the interleaving time difference according to the sum of squares of differences between the display timestamp of the data frame and the display timestamps of each reference frame; An interleaving loss parameter is determined according to a quotient of the interleaving time difference and a total number of frames of the same type corresponding to the data frame.

23. An electronic device comprising: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 11.

24. A non-transitory computer-readable storage medium storing computer instructions, wherein: The computer instructions are used to cause the computer to execute the method according to any one of claims 1-11.

25. A computer program product comprising a computer program / instructions, characterized in that When the computer program / instructions are executed by a processor, the steps of the method according to any one of claims 1-11 are implemented.

Citation Information

Patent Citations

  • Data transmission method and device, electronic equipment and storage medium

    CN113038128A