Method and apparatus for streaming media data transmission, device, and medium
By transmitting audio and video data associated with the first frame in the same data block, the problems of long first frame time and data misalignment in the DASH protocol are solved, and faster first frame rendering and better user experience are achieved.
Patent Information
- Application Number
- PCT/CN2025/073311
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-01-31
- Filing Date
- 2025-01-20
- Publication Date
- 2025-08-07
AI Technical Summary
In the existing streaming media data transmission based on the DASH protocol, the first frame time is long and the audio and video data is easily misaligned, affecting the user's viewing experience.
Transfer audio data and video data associated with the first frame in the same data block, including media demonstration description MPD, reduce the number of HTTP requests and ensure data alignment through merge transmission.
Reduce the first frame time, improve user viewing experience, prevent audio and video data from being misaligned, and ensure correct rendering of the first frame.
Smart Images

Figure CN2025073311_07082025_PF_FP_ABST
Abstract
Description
Method, device, equipment and medium for streaming media data transmission
[0001] This application claims priority to the Chinese invention patent application entitled “Methods, devices, equipment and media for streaming media data transmission” and application number 202410134527.7, filed on January 31, 2024. The entire contents of that application are incorporated herein by reference. Technical Field
[0002] Embodiments of the present disclosure generally relate to the field of computers, and more particularly, to methods, devices, apparatuses, and media for streaming media data transmission. Background Art
[0003] In recent years, mobile terminals such as mobile phones and tablet computers have become important tools for people's daily lives, study and work. People can use mobile terminals to watch movies, live broadcasts, hold remote meetings, and so on. These functions can be achieved with the help of streaming data transmission. The Dynamic Adaptive Streaming over Hypertext Transfer Protocol (DASH) protocol is a common media data distribution protocol for streaming media. As the DASH protocol is increasingly widely used in streaming data transmission, how to improve the efficiency of streaming data transmission based on the DASH protocol and ensure the transmission quality has become an urgent problem to be solved. Summary of the Invention
[0004] In a first aspect of the present disclosure, a method for transmitting streaming media data is provided. In this method, a request for a first data block of streaming media data is received from a first device. The first data block is associated with a first frame of the streaming media data at the first device, and the first data block includes at least a media presentation description (MPD) associated with the streaming media data, audio data associated with the first frame, and video data associated with the first frame. Furthermore, the first data block is sent to the first device.
[0005] In a second aspect of the present disclosure, another method for transmitting streaming media data is provided. In this method, a first device sends a request for a first data block of streaming media data to a second device, where the first data block is associated with a first frame of the streaming media data at the first device and includes at least a media presentation description (MPD) associated with the streaming media data, audio data associated with the first frame, and video data associated with the first frame; and the first data block is received from the second device.
[0006] In a third aspect of the present disclosure, an apparatus for transmitting streaming media data is provided. The apparatus includes a request receiving module and a data block sending module. The request receiving module is configured to receive a request for a first data block of streaming media data from a first device, where the first data block is associated with a first frame of the streaming media data at the first device, and the first data block includes at least a media presentation description (MPD) associated with the streaming media data, audio data associated with the first frame, and video data associated with the first frame. The data block sending module is configured to send the first data block to the first device.
[0007] In a fourth aspect of the present disclosure, an apparatus for transmitting streaming media data is provided. The apparatus includes a request sending module and a data block receiving module. The request sending module is configured to: at a first device, send a request for a first data block of streaming media data to a second device, the first data block being associated with a first frame of the streaming media data at the first device, and the first data block including at least a media presentation description (MPD) associated with the streaming media data, audio data associated with the first frame, and video data associated with the first frame. The data block receiving module is configured to: receive the first data block from the second device.
[0008] In a fifth aspect of the present disclosure, an electronic device is provided. The electronic device includes: at least one processing unit; and at least one memory, the at least one memory being coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit, wherein the instructions, when executed by the at least one processing unit, cause the electronic device to perform the method according to the first or second aspect of the present disclosure.
[0009] In a sixth aspect of the present disclosure, a computer-readable storage medium is provided, on which instructions are stored. When the instructions are executed by a processor, the processor implements the method according to the first aspect or the second aspect of the present disclosure.
[0010] It should be understood that the content described in this summary section is not intended to limit the key features or important features of the embodiments of the present disclosure, nor is it intended to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] The above and other features, advantages and aspects of the embodiments of the present disclosure will become more apparent hereinafter with reference to the following detailed description in conjunction with the accompanying drawings. In the accompanying drawings, the same or similar reference numerals represent the same or similar elements, wherein:
[0012] FIG1 is a schematic diagram illustrating an example environment in which various embodiments of the present disclosure can be implemented;
[0013] FIG2 shows a signaling diagram for streaming media data transmission according to some embodiments of the present disclosure;
[0014] FIG3 is a schematic diagram showing a structure of a first data block according to some embodiments of the present disclosure;
[0015] FIG4 shows a flowchart of a method for streaming media data transmission according to some embodiments of the present disclosure;
[0016] FIG5 shows a flowchart of a method for streaming media data transmission according to some embodiments of the present disclosure;
[0017] FIG6 shows a block diagram of an example apparatus for streaming media data transmission according to some embodiments of the present disclosure;
[0018] FIG7 shows a block diagram of an example apparatus for streaming media data transmission according to some embodiments of the present disclosure; and
[0019] FIG8 illustrates a block diagram of a device in which one or more embodiments of the present disclosure may be implemented. DETAILED DESCRIPTION
[0020] The following describes embodiments of the present disclosure in more detail with reference to the accompanying drawings. Although certain embodiments of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the embodiments described herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are for illustrative purposes only and are not intended to limit the scope of protection of the present disclosure.
[0021] In the description of the embodiments of the present disclosure, the term "including" and similar terms should be understood as open inclusion, that is, "including but not limited to". The term "based on" should be understood as "based at least in part on". The term "one embodiment" or "the embodiment" should be understood as "at least one embodiment". The term "some embodiments" should be understood as "at least some embodiments". The following may also include other explicit and implicit definitions. As used herein, the term "model" can represent the association relationship between various data. For example, the above-mentioned association relationship can be obtained based on a variety of technical solutions currently known and / or to be developed in the future.
[0022] As used herein, the term "in response to" refers to a state in which a corresponding event occurs or a condition is satisfied. It will be understood that the timing of executing a subsequent action executed in response to the event or condition is not necessarily strongly correlated with the time when the event occurs or the condition is satisfied. For example, in some cases, the subsequent action may be executed immediately when the event occurs or the condition is satisfied; in other cases, the subsequent action may be executed some time after the event occurs or the condition is satisfied.
[0023] It is understandable that the data involved in this technical solution (including but not limited to the data itself, the acquisition or use of the data) must comply with the requirements of relevant laws, regulations and relevant provisions.
[0024] It is understandable that before using the technical solutions disclosed in the various embodiments of this disclosure, the type, scope of use, usage scenarios, etc. of the personal information involved in this disclosure should be informed to the user and the user's authorization should be obtained in an appropriate manner in accordance with relevant laws and regulations.
[0025] For example, in response to a user's active request, a prompt message is sent to the user to clearly inform the user that the operation requested will require the acquisition and use of the user's personal information. This allows the user to independently choose whether to provide personal information to the electronic device, application, server, storage medium, or other software or hardware that performs the operations of the disclosed technical solution based on the prompt message.
[0026] As an optional but non-limiting embodiment, in response to receiving a user's active request, the prompt information may be sent to the user in the form of a pop-up window, in which the prompt information may be presented in text form. Furthermore, the pop-up window may also include a selection control for the user to select "agree" or "disagree" to provide personal information to the electronic device.
[0027] It is understandable that the above notification and user authorization process are merely illustrative and do not limit the embodiments of the present disclosure. Other methods that meet relevant laws and regulations may also be applied to the embodiments of the present disclosure.
[0028] As briefly mentioned above, with the diversification of online entertainment, streaming media data transmission based on the DASH protocol is becoming increasingly widespread. A key metric for the streaming viewing experience is time to first frame, which is the time it takes for the first frame of the video to be displayed after the user clicks "Start Play."
[0029] In one existing solution, if a user initiates playback on the client, the client first needs to request a Media Presentation Description (MPD) file from the server. After receiving and parsing the MPD file, the client can request an initialization segment for audio and an initialization segment for video from the server, and further request a media segment for audio and a media segment for video. In this solution, it takes at least three Hypertext Transfer Protocol (HTTP) requests and responses from the time the user initiates viewing on the client to the time the first frame is presented on the client. This results in a longer first frame time and a poorer viewing experience for users.
[0030] In another existing solution, the MPD file, initialization segment, and media segments are combined into a tuning-in segment, thus reducing the time required for two requests. However, in practice, to better support features such as adaptive bit-rate (ABR) and multiple audio tracks, DASH often adopts a multiple adaptation set configuration, with audio and video content belonging to different adaptation sets. This results in the client needing to request the tuning-in segments for both audio and video simultaneously when starting playback.
[0031] The inventors discovered through research that: since the two HTTP requests for the audio start segment and the video start segment are independent of each other, there may be differences in the time they arrive at the server, the cache status of the corresponding files, etc., which will cause the content of the media segments contained in the audio start segment and the video start segment actually received by the client to be misaligned, such as inconsistent start times, etc. This will cause errors in the content of the first frame, seriously affecting the user's viewing experience. In order to avoid this problem, the client often needs to execute an additional HTTP request to achieve the alignment of the audio start segment and the video start segment, and then start normally. However, this will also make the first frame time longer and the user's viewing experience poor.
[0032] To this end, the various embodiments of the present disclosure propose a scheme for merging the audio data and video data for the first frame for transmission. Specifically, according to the embodiments of the present disclosure, a scheme for streaming media data transmission is proposed. In this scheme, a request for a first data block of streaming media data is received from a first device. The first data block is associated with the first frame of the streaming media data at the first device, and the first data block includes at least a media presentation description MPD associated with the streaming media data, audio data associated with the first frame, and video data associated with the first frame. Further, the first data block is sent to the first device. Correspondingly, according to the embodiments of the present disclosure, another scheme for streaming media data transmission is proposed. In this scheme, at the first device, a request for the first data block of streaming media data is sent to the second device, and the first data block is received from the second device.
[0033] It will be more clearly understood through the description below that according to the embodiments of the present disclosure, by transmitting the audio data and video data for the first frame in the same data block, on the one hand, the number of HTTP requests required to start the broadcast can be reduced, thereby reducing the first frame time and improving the user's viewing experience. On the other hand, it can avoid the misalignment between the transmitted audio data and video data, thereby ensuring the correctness of the first frame data and avoiding the need to request the start of the segment again, thereby ensuring that the first frame is rendered correctly. In other words, the solution according to the embodiments of the present disclosure can effectively prevent the misalignment of audio and video data while reducing the first frame time.
[0034] Various example implementations of the solution will be described in detail below with reference to the accompanying drawings. First, refer to Figure 1, which shows a schematic diagram of an example environment 100 in which various embodiments of the present disclosure can be implemented. The example environment 100 may generally include a client 110, a server 120, and a user 130. The client 110 may be communicatively coupled to the server 120. In some embodiments, the client 110 may communicate directly with the server 120. In other embodiments, a content delivery network (CDN) may be deployed between the client 110 and the server 120 so that the client 110 can obtain the required content nearby. The embodiments of the present disclosure are not limited in this respect.
[0035] In FIG1 , client 110 is shown as a mobile phone, but client 110 may also be any type of mobile terminal or portable terminal, including a laptop computer, a notebook computer, a netbook computer, a tablet computer, a media computer, a multimedia tablet, a personal communication system (PCS) device, a personal navigation device, a personal digital assistant (PDA), an audio / video player, a digital camera / camcorder, a positioning device, a television receiver, a radio receiver, an e-book device, a gaming device, or any combination thereof, including accessories and peripherals of these devices or any combination thereof. In some embodiments, client 110 may also support any type of interface for user 130 (such as a "wearable" circuit, etc.).
[0036] As shown in Figure 1, user 130 can, for example, initiate viewing of streaming content, such as live video, online movies, and the like, by operating an application on client 110. Client 110 can send a request for corresponding streaming data to server 120, and then server 120 can send start data for starting to play streaming content to client 110. Client 110 can then render and play the first frame of content based on the start data for user 130 to watch. This will be described in further detail below. In some embodiments, server 120 can be a device with computing and communication functions, such as a workstation, cloud server, etc. It should be understood that the structure and function of environment 100 are described for exemplary purposes only, and does not imply any limitation on the scope of the present disclosure.
[0037] Example signaling
[0038] FIG2 shows a signaling diagram 200 for streaming media data transmission according to some embodiments of the present disclosure. The signaling diagram 200 generally relates to a first device 201 and a second device 202. In some embodiments, the first device 201 can be implemented as the client 110 in FIG1 , and the second device 202 can be implemented as the server 120 in FIG1 . It should be understood that the first device 201 and / or the second device 202 can also be implemented as any other suitable device. For example, the first device 201 can also be implemented as another server. The scope of the present disclosure is not limited in this respect.
[0039] As shown in Figure 2, the first device 201 sends 205 a request for a first data block of streaming media data to the second device 202. By way of example and not limitation, the streaming media data includes live broadcast data, online video data, and the like. The first data block is associated with the first frame of the streaming media data at the first device 201. For example, the first data block may include some or all of the data required to generate the first frame at the first device 201. In some embodiments, the streaming media data may be transmitted according to the DASH protocol. Exemplarily, the request for the first data block may be an HTTP request, such as a GET request. The request may, for example, include a Uniform Resource Locator (URL) for the first data block.
[0040] The first data block includes at least a media presentation description MPD associated with the streaming media data, audio data associated with the first frame, and video data associated with the first frame. The MPD file (also referred to as MPD in this article) is used to describe information about the streaming media data, such as bit rate, resolution, characteristics of different media components (such as audio, video, text, etc.) included in the multimedia content, and so on. The audio data associated with the first frame may include the encoded code stream of the audio of the first frame, and the video data associated with the first frame may include the encoded code stream of the video of the first frame. In this way, the audio data and video data for the first frame can be requested at the same time, and the audio data and video data for the first frame can be sent in the same data block, thereby effectively preventing the misalignment of audio and video data.
[0041] In some embodiments, the data block may be a segment, and the first data block may also be referred to as a merged start segment. As used herein, a segment refers to the entity body of a response to an HTTP GET request or a portion of an HTTP GET request from a DASH client. It should be understood that the data block may also be implemented in any other suitable manner. The scope of the present disclosure is not limited in this respect.
[0042] In some embodiments, the first data block may further include initialization information associated with the decoded audio data and video data. Examples of initialization information include, but are not limited to, initialization setting information for the decoder at the first device 201, a picture parameter set (PPS), a sequence parameter set (SPS), and the like. In this way, after receiving the first data block, the first device 201 can generate and display the first frame based on the first data block, so that only one HTTP request is required to present the first frame, thereby further reducing the first frame time and improving the response speed to the start-up operation of the user 130.
[0043] Alternatively or additionally, the first data block can also include the start time of audio data and / or the start time of video data.Alternatively or additionally, the first data block can also include the sequence number of the audio data block corresponding to the audio data and / or the sequence number of the video data block corresponding to the audio data.Exemplarily and non-restrictively, the audio data block can be the media fragmentation (also referred to as audio fragmentation) for audio content, and the video data block can be the media fragmentation (also referred to as video fragmentation) for video content.By transmitting the start time and / or sequence number, the first equipment 201 can directly determine the start time and / or the sequence number of the data block of subsequent required request based on this information, thereby can more specifically initiate request.In this way, can more efficiently request data, and shorten the response time at the second equipment 202 places, thereby further reduce first frame time.
[0044] Correspondingly, the second device 202 may receive 210 a request for the first data block from the first device 201. In some embodiments, the second device 202 may generate the first data block in response to receiving the request. In other words, the second device 202 generates the first data block on demand according to the request of the user 130.
[0045] For example, the second device 202 may select a target video data block including video data associated with the first frame from at least one video data block encapsulated for the streaming media data. In some embodiments, the second device 202 may determine the latest currently available video data block as the target video data block. In this way, delays in playback on the client 110 may be minimized, allowing the user 130 to obtain the latest real-time video images in a timely manner.
[0046] Alternatively, the target video data block can be selected based on the first frame delay. This delay can be predetermined or sent by the first device 201 to the second device 202. For example, the first device 201 can send its desired first frame delay to the second device 202 along with the request for the first data block. It should be understood that the first frame delay can also be determined in any other suitable manner, and the scope of the present disclosure is not limited in this respect.
[0047] As an example, the second device 202 can determine the target time based on the current time and the delay time. Exemplarily, the second device 202 can directly subtract the delay time (e.g., 5 seconds) from the current time (e.g., 01:45:20) to obtain the target time (e.g., 01:45:15). Further, the second device 202 can select the target video data block from the packaged video data blocks based on the target time. For example, the second device 202 can determine the video data block including the video data of the target time as the target video data block. In this way, a buffer can be established between the content played by the client 110 and the real-time content to avoid problems such as playback jams and other problems caused by problems such as network congestion and network speed drop, thereby ensuring the viewing experience of the user 130.
[0048] Furthermore, the second device 202 can select a target audio data block including the audio data associated with the first frame from at least one audio data block encapsulated for the streaming media data. In some embodiments, the second device 202 can determine the latest audio data block currently available as the target audio data block. In this way, the delay existing during playback of the client 110 can be avoided as much as possible, so that the user 130 can obtain the latest audio content in real time.
[0049] Alternatively, the target audio data block can be selected based on the delay time of first frame. As an example, the second device 202 can determine the target time based on current time and delay time. Exemplarily, the second device 202 can directly deduct delay time (for example, 5 seconds) to obtain target time (for example, 01:45:15 seconds) from the current time (for example, 01:45:15). Further, the second device 202 can select the target audio data block from the audio data blocks that have been packaged based on this target time. For example, the audio data block of the audio data comprising the target time can be determined as the target audio data block by the second device 202. In this way, buffering can be set up between client 110 playback content and real-time content, avoid the problems such as playback jam, network speed drop that cause such problems as stuck, thereby ensure the viewing experience of user 130.
[0050] In some embodiments, the target video data block can be first selected according to the above-mentioned method, and the audio data block corresponding to the target video data block can be directly used as the target audio data block. For example, the audio data block with the same sequence number as the target video data block can be used as the target audio data block. In other embodiments, the target audio data block can be first selected according to the above-mentioned method, and the video data block corresponding to the target audio data block can be directly used as the target video data block. For example, the video data block with the same sequence number as the target audio data block can be used as the target video data block. In other embodiments, the selection operation of the target video data block and the target audio data block can also be performed independently of each other in parallel. It should be understood that the second device 202 can also select the target video data block and the target audio in any other suitable manner, and the scope of the present disclosure is not limited in this respect.
[0051] Furthermore, the second device 202 may encapsulate the content of the target video data block and the content of the target video data block in the first data block. For example, the content of the target video data block and the content of the target video data block may be encapsulated in a media data block (e.g., a media segment) in the first data block.
[0052] FIG3 shows a schematic diagram of the structure 300 of the first data block according to some embodiments of the present disclosure. As shown in FIG3 , the second device 202 can encapsulate the current MPD file into the first data block. Further, the sequence number and start time of the selected audio segment and video segment can be written. Then, the initialization data block (for example, the initialization segment) including the initialization information can be encapsulated into the first data block, and the content of the target video segment and the content of the target audio segment can be encapsulated into the media segment in the first data block. It should be understood that this structure 300 is merely exemplary and non-restrictive. The structure 300 may also include additional parts not shown and / or may omit a certain (or some) part shown. For example, the initialization segment can be omitted, and the initialization information can be directly encapsulated into the media segment. The scope of the present disclosure is not limited in this respect.
[0053] In some embodiments, audio data and video data can be interleaved and encapsulated into the first data block based on a preset sorting criterion. For example, after writing the video data of the first frame, the audio data of the first frame can be written. Then, the video data and audio data of another frame immediately following the first frame based on the same sorting criterion can be written, and so on. In one embodiment, the preset sorting criterion can be decoding timestamp order. In other words, the audio data and video data can be encapsulated into the first data block based on decoding timestamp order. For example, after writing the video data of the first frame, the audio data of the first frame can be written. Then, the video data and audio data of another frame immediately following the first frame based on the same sorting criterion can be written, and so on. This approach effectively avoids the problem in existing solutions of requiring the video content of all frames in the first data block to be decoded before decoding the audio content of the first frame. In this way, the media content in the first data block can be decoded more efficiently at client 110, further reducing the first frame time and improving the response speed to the start-up operation of user 130.
[0054] In some embodiments, the first data block may be encapsulated based on the MP4 format, thereby more efficiently encapsulating data and achieving compatibility with the DASH protocol. It should be understood that the first data block may also be encapsulated based on any other suitable file format, and the scope of the present disclosure is not limited in this respect.
[0055] The above describes an exemplary embodiment in which the second device 202 generates the first data block in response to receiving the request. In this way, the second device 202 can generate the first data block on demand, thereby saving computing power and storage space, so as to facilitate the application of the solution according to the embodiment of the present disclosure in resource-constrained scenarios.
[0056] In other embodiments, the second device 202 may generate data blocks for starting the broadcast in real time as the streaming media data is generated. For example, the second device 202 may generate a set of data blocks based on at least a portion of the streaming media data. For example, the second device 202 may encapsulate the encoded frames into corresponding data blocks as the live broadcast progresses.
[0057] Exemplarily, in response to the encoding of the current frame in the streaming media data being completed, it is determined whether the audio data and video data of the current frame belong to the current data block in a group of data blocks. For example, the current frame may refer to the frame currently being encoded, and the current data block may refer to the data block currently being generated. For ease of explanation, it is assumed that the current data block corresponds to the content from 01:01:00 to 01:01:02. If the current frame is 01:01:01, it can be determined that the audio data and video data of the current frame belong to the current data block. If the current frame is 01:01:03, it can be determined that the audio data and video data of the current frame do not belong to the current data block.
[0058] If it is determined that the audio data and video data of the current frame belong to the current data block, the second device 202 may encapsulate the audio data and video data of the current frame in the current data block. For example, the audio data and video data of the current frame may be interleaved and encapsulated in the current data block based on the decoding timestamp order.
[0059] If it is determined that the audio data and video data of the current frame do not belong to the current data block, the second device 202 can create a new data block and add it to the set of data blocks as another data block. Exemplarily, the second device 202 can generate another data block in the set of data blocks based on the audio data and video data of the current frame. For example, the second device 202 can encapsulate the current MPD in the other data block, and further encapsulate the audio data and video data of the current frame in the other data block. Additionally, the second device 202 can also write corresponding initialization information for decoding (e.g., as an initialization segment, or as part of a media segment), the start time of the audio data (e.g., the audio segment start time), the start time of the video data (e.g., the video segment start time), the sequence number of the audio data block corresponding to the audio data (e.g., the audio segment sequence number), the sequence number of the video data block corresponding to the video data (e.g., the video segment sequence number), etc. into the other data block. Furthermore, the second device 202 may use the newly created data block as the current data block and the next encoded frame as the current frame, thereby iteratively performing the above operations to generate one or more data blocks for starting broadcasting.
[0060] Furthermore, in response to receiving the request, the second device 202 may select a first data block from the generated set of data blocks. In some embodiments, the most recently available data block in the set of data blocks may be determined as the first data block. In this way, delays in playback by the client 110 can be minimized, allowing the user 130 to obtain the latest real-time media content in a timely manner.
[0061] Alternatively, the first data block may be selected based on the first frame delay. The delay may be predetermined or may be sent by the first device 201 to the second device 202. For example, the first device 201 may send its desired first frame delay to the second device 202 along with the request for the first data block. It should be understood that the first frame delay may also be determined in any other suitable manner, and the scope of the present disclosure is not limited in this respect.
[0062] As an example, the second device 202 can determine the target time based on the current time and the delay time. Exemplarily, the second device 202 can directly subtract the delay time (e.g., 8 seconds) from the current time (e.g., 01:05:20) to obtain the target time (e.g., 01:05:12). Further, the second device 202 can select the first data block from the generated group of data blocks based on the target time. For example, the second device 202 can determine the data block of the audio and video data of the target time in the group of data blocks as the first data block. In this way, a buffer can be established between the content played by the client 110 and the real-time content to avoid problems such as playback jams caused by problems such as network congestion and network speed reduction, thereby ensuring the viewing experience of the user 130.
[0063] The above describes an exemplary embodiment in which the second device 202 generates data blocks for starting broadcasting in real time as streaming media data is generated. In this way, after receiving a request for a first data block, the second device 202 can directly determine the first data block from the generated data blocks, thereby responding to the request more quickly and further shortening the first frame time.
[0064] Returning to reference Figure 2 , in response to the request for the first data block, the second device 202 can send 215 the first data block to the first device 201. By way of example and not limitation, the second device 202 can send the entire first data block in a single HTTP response. Accordingly, the first device 201 receives 220 the data block from the second device 202. In this manner, the audio data and video data for the first frame can be sent in the same data block, effectively preventing misalignment of the audio and video data.
[0065] Furthermore, the first device 201 can display the first frame based on the first data block. For example, the first device 201 can parse the first data block and perform decoding operations on the audio and video streams, and then render and present the first frame based on the obtained data. Furthermore, the first device 201 can continue to initiate requests for subsequent audio and video content to the second device 202 to continue playing subsequent content of the streaming media.
[0066] From the above description in combination with Figures 1 to 3, it can be seen that in the method for streaming media data transmission according to each embodiment of the present disclosure, by transmitting the audio data and video data for the first frame in the same data block, on the one hand, the number of HTTP requests required for start-up can be reduced, thereby reducing the first frame time and improving the user's viewing experience. On the other hand, it can avoid the misalignment of the transmitted audio data and video data, thereby ensuring the correctness of the first frame data and avoiding re-requesting the start-up segment, thereby ensuring that the first frame is rendered correctly. In other words, the method according to each embodiment of the present disclosure can effectively prevent the misalignment of audio and video data while reducing the first frame time.
[0067] Example Method
[0068] FIG4 shows a flow chart of a method 400 for transmitting streaming media data according to some embodiments of the present disclosure. In some embodiments, the method 400 may be executed at the server 120 shown in FIG1 . It should be understood that the method 400 may include additional blocks not shown and / or may omit one (or some) of the blocks shown, and the scope of the present disclosure is not limited in this respect.
[0069] In box 402, a request for a first data block of streaming media data is received from a first device, the first data block is associated with a first frame of the streaming media data at the first device, and the first data block includes at least a media presentation description MPD associated with the streaming media data, audio data associated with the first frame, and video data associated with the first frame.
[0070] At block 404, a first data block is sent to a first device.
[0071] In some embodiments, the audio data and the video data are encapsulated in the first data block in an order based on a decoding timestamp.
[0072] In some embodiments, the method 400 further includes: generating a first data block in response to receiving the request.
[0073] In some embodiments, generating the first data block includes: selecting a target video data block from at least one video data block encapsulated for streaming media data, the target video data block including video data associated with the first frame; selecting a target audio data block from at least one audio data block encapsulated for streaming media data, the target audio data block including audio data associated with the first frame; and encapsulating the content of the target video data block and the content of the target audio data block in a media data block in the first data block.
[0074] In some embodiments, the target video data block or the target audio data block is selected based on the delay time of the first frame.
[0075] In some embodiments, method 400 further includes: generating a set of data chunks based on at least a portion of the streaming media data; and selecting a first data chunk from the set of data chunks in response to receiving the request.
[0076] In some embodiments, the first data block is selected based on a delay time of the first frame.
[0077] In some embodiments, generating a set of data blocks includes: in response to encoding of a current frame in streaming media data being completed, determining whether the audio data and video data of the current frame belong to a current data block in a set of data blocks; in response to determining that the audio data and video data of the current frame belong to the current data block, encapsulating the audio data and video data of the current frame in the current data block; and in response to determining that the audio data and video data of the current frame do not belong to the current data block, generating another data block in a set of data blocks based on the audio data and video data of the current frame.
[0078] In some embodiments, generating another data block in the set of data blocks includes: encapsulating the current MPD in the another data block; and encapsulating the audio data and the video data of the current frame in the another data block.
[0079] In some embodiments, the first data block further includes at least one of the following: initialization information associated with decoded audio data and video data, a start time of the audio data, a start time of the video data, a sequence number of an audio data block corresponding to the audio data, or a sequence number of a video data block corresponding to the video data.
[0080] In some embodiments, the streaming media data includes live data.
[0081] In some embodiments, the streaming media data is transmitted according to a Dynamic Adaptive Streaming over Hypertext Transfer Protocol (DASH) protocol, and the first data block is a segment and is encapsulated based on the MP4 format.
[0082] In some embodiments, method 400 is implemented at a server, and the first device comprises a client.
[0083] FIG5 shows a flow chart of a method 500 for transmitting streaming media data according to some embodiments of the present disclosure. In some embodiments, the method 500 may be executed at the client 110 shown in FIG1 . It should be understood that the method 500 may include additional blocks not shown and / or may omit one (or some) of the blocks shown, and the scope of the present disclosure is not limited in this respect.
[0084] In box 502, at a first device, a request for a first data block of streaming media data is sent to a second device, the first data block is associated with a first frame of the streaming media data at the first device, and the first data block includes at least a media presentation description MPD associated with the streaming media data, audio data associated with the first frame, and video data associated with the first frame.
[0085] At block 504 , a first data block is received from a second device.
[0086] In some embodiments, the audio data and the video data are encapsulated in the first data block in an order based on a decoding timestamp.
[0087] In some embodiments, the method 500 further includes: displaying a first frame based on the first data block.
[0088] In some embodiments, the first data block further includes at least one of the following: initialization information associated with decoded audio data and video data, a start time of the audio data, a start time of the video data, a sequence number of an audio data block corresponding to the audio data, or a sequence number of a video data block corresponding to the video data.
[0089] In some embodiments, the streaming media data includes live data.
[0090] In some embodiments, the streaming media data is transmitted according to a Dynamic Adaptive Streaming over Hypertext Transfer Protocol (DASH) protocol, and the first data block is a segment and is encapsulated based on the MP4 format.
[0091] In some embodiments, the first device comprises a client and the second device comprises a server.
[0092] Example devices and equipment
[0093] Embodiments of the present disclosure also provide corresponding apparatuses and devices for implementing the above-described methods or processes. FIG6 shows a block diagram of an example apparatus 600 for transmitting streaming media data according to some embodiments of the present disclosure. The apparatus 600 can, for example, be used to implement the methods according to some embodiments of the present disclosure. In some embodiments, the apparatus 600 can be implemented at the server 120 shown in FIG1 .
[0094] As shown in FIG6 , apparatus 600 may include a request receiving module 602 and a data block sending module 604. Request receiving module 602 is configured to receive a request for a first data block of streaming media data from a first device, the first data block being associated with a first frame of the streaming media data at the first device, and the first data block including at least a media presentation description (MPD) associated with the streaming media data, audio data associated with the first frame, and video data associated with the first frame. Data block sending module 604 is configured to send the first data block to the first device.
[0095] In some embodiments, the audio data and the video data are encapsulated in the first data block in an order based on a decoding timestamp.
[0096] In some embodiments, the apparatus 600 further includes a first generating module configured to: generate a first data block in response to receiving the request.
[0097] In some embodiments, the first generation module includes a first selection module, a second selection module, and a first encapsulation module. The first selection module is configured to select a target video data block from at least one video data block encapsulated for streaming media data, the target video data block including video data associated with the first frame. The second selection module is configured to select a target audio data block from at least one audio data block encapsulated for streaming media data, the target audio data block including audio data associated with the first frame. The first encapsulation module is configured to encapsulate the contents of the target video data block and the contents of the target audio data block into a media data block within the first data block.
[0098] In some embodiments, the target video data block or the target audio data block is selected based on the delay time of the first frame.
[0099] In some embodiments, the apparatus 600 further includes a second generation module and a selection module. The second generation module is configured to generate a set of data blocks based on at least a portion of the streaming media data. The selection module is configured to select the first data block from the set of data blocks in response to receiving the request.
[0100] In some embodiments, the first data block is selected based on a delay time of the first frame.
[0101] In some embodiments, the second generation module includes a determination module, a second encapsulation module, and a third generation module. The determination module is configured to: in response to encoding of a current frame in the streaming media data being completed, determine whether the audio data and video data of the current frame belong to a current data block in a group of data blocks. The second encapsulation module is configured to: in response to determining that the audio data and video data of the current frame belong to the current data block, encapsulate the audio data and video data of the current frame in the current data block. The third generation module is configured to: in response to determining that the audio data and video data of the current frame do not belong to the current data block, generate another data block in the group of data blocks based on the audio data and video data of the current frame.
[0102] In some embodiments, the third generation module includes a third encapsulation module and a fourth encapsulation module. The third encapsulation module is configured to encapsulate the current MPD in another data block. The fourth encapsulation module is configured to encapsulate the audio data and video data of the current frame in another data block.
[0103] In some embodiments, the first data block further includes at least one of the following: initialization information associated with decoded audio data and video data, a start time of the audio data, a start time of the video data, a sequence number of an audio data block corresponding to the audio data, or a sequence number of a video data block corresponding to the video data.
[0104] In some embodiments, the streaming media data includes live data.
[0105] In some embodiments, the streaming media data is transmitted according to a Dynamic Adaptive Streaming over Hypertext Transfer Protocol (DASH) protocol, and the first data block is a segment and is encapsulated based on the MP4 format.
[0106] In some embodiments, the apparatus 600 is implemented at a server, and the first device comprises a client.
[0107] FIG7 shows a block diagram of an example apparatus 700 for streaming media data transmission according to some embodiments of the present disclosure. The apparatus 700 can be used to implement the method according to some embodiments of the present disclosure. In some embodiments, the apparatus 700 can be implemented at the client 110 shown in FIG1 .
[0108] As shown in FIG7 , apparatus 700 may include a request sending module 702 and a data block receiving module 704. Request sending module 702 is configured to, at a first device, send a request for a first data block of streaming media data to a second device. The first data block is associated with a first frame of the streaming media data at the first device, and the first data block includes at least a media presentation description (MPD) associated with the streaming media data, audio data associated with the first frame, and video data associated with the first frame. Data block receiving module 704 is configured to receive the first data block from the second device.
[0109] In some embodiments, the audio data and the video data are encapsulated in the first data block in an order based on a decoding timestamp.
[0110] In some embodiments, the apparatus 700 further includes a display module configured to display a first frame based on the first data block.
[0111] In some embodiments, the first data block further includes at least one of the following: initialization information associated with decoded audio data and video data, a start time of the audio data, a start time of the video data, a sequence number of an audio data block corresponding to the audio data, or a sequence number of a video data block corresponding to the video data.
[0112] In some embodiments, the streaming media data includes live data.
[0113] In some embodiments, the streaming media data is transmitted according to a Dynamic Adaptive Streaming over Hypertext Transfer Protocol (DASH) protocol, and the first data block is a segment and is encapsulated based on the MP4 format.
[0114] In some embodiments, the first device comprises a client and the second device comprises a server.
[0115] The modules and / or units included in the device 600 and / or the device 700 can be implemented in various ways, including software, hardware, firmware or any combination thereof. In some embodiments, one or more units can be implemented using software and / or firmware, such as machine executable instructions stored on a storage medium. In addition to or as an alternative to machine executable instructions, some or all of the units in the device 600 and / or the device 700 can be implemented at least in part by one or more hardware logic components. As an example and not limitation, exemplary types of hardware logic components that can be used include field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), system on chip (SOCs), complex programmable logic devices (CPLDs), and the like.
[0116] The modules and / or units shown in Figures 6 and / or 7 may be partially or entirely implemented as hardware modules, software modules, firmware modules, or any combination thereof. In particular, in some embodiments, the processes, methods, or procedures described above may be implemented by hardware in a storage system, a host corresponding to the storage system, or other computing devices independent of the storage system.
[0117] FIG8 shows a block diagram of a device 800 in which one or more embodiments of the present disclosure may be implemented. It should be understood that the electronic device 800 shown in FIG8 is merely exemplary and should not be construed as limiting the functionality and scope of the embodiments described herein. The electronic device 800 shown in FIG8 can be used to implement the client 110 and server 120 shown in FIG1 and / or the methods described above.
[0118] As shown in FIG8 , electronic device 800 is a general-purpose electronic device. Components of electronic device 800 may include, but are not limited to, one or more processors or processing units 810, memory 820, storage device 830, one or more communication units 840, one or more input devices 850, and one or more output devices 860. Processing unit 810 may be a real or virtual processor and is capable of performing various processes according to programs stored in memory 820. In a multi-processor system, multiple processing units execute computer-executable instructions in parallel to enhance the parallel processing capabilities of electronic device 800.
[0119] The electronic device 800 typically includes a plurality of computer storage media. Such media can be any available media accessible to the electronic device 800, including but not limited to volatile and non-volatile media, removable and non-removable media. The memory 820 can be a volatile memory (e.g., registers, cache, random access memory (RAM)), a non-volatile memory (e.g., read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory), or some combination thereof. The storage device 830 can be a removable or non-removable medium and can include a machine-readable medium, such as a flash drive, a disk, or any other medium that can be used to store information and / or data (e.g., training data for training) and can be accessed within the electronic device 800.
[0120] The electronic device 800 may further include additional removable / non-removable, volatile / non-volatile storage media. Although not shown in FIG8 , a disk drive for reading from or writing to a removable, non-volatile disk (e.g., a “floppy disk”) and an optical drive for reading from or writing to a removable, non-volatile optical disk may be provided. In these cases, each drive may be connected to a bus (not shown) by one or more data media interfaces. The memory 820 may include a computer program product 825 having one or more program modules configured to perform various methods or actions of various implementations of the present disclosure.
[0121] The communication unit 840 enables communication with other electronic devices via a communication medium. Additionally, the functions of the components of the electronic device 800 can be implemented in a single computing cluster or multiple computing machines that can communicate via a communication connection. Thus, the electronic device 800 can operate in a networked environment using a logical connection with one or more other servers, a network personal computer (PC), or another network node.
[0122] The input device 850 may be one or more input devices, such as a mouse, keyboard, or trackball. The output device 860 may be one or more output devices, such as a display, a speaker, or a printer. The electronic device 800 may also communicate with one or more external devices (not shown) via the communication unit 840 as needed, such as a storage device, a display device, or the like, with one or more devices that allow a user to interact with the electronic device 800, or with any device that allows the electronic device 800 to communicate with one or more other electronic devices (e.g., a network card, a modem, etc.). Such communication may be performed via an input / output (I / O) interface (not shown).
[0123] According to an exemplary implementation of the present disclosure, a computer-readable storage medium is provided, on which computer-executable instructions are stored, wherein the computer-executable instructions are executed by a processor to implement the method described above. According to an exemplary implementation of the present disclosure, a computer program product is also provided, which is tangibly stored on a non-transitory computer-readable medium and includes computer-executable instructions, and the computer-executable instructions are executed by a processor to implement the method described above. According to an exemplary implementation of the present disclosure, a computer program product is provided, on which a computer program is stored, and when the program is executed by a processor, the method described above is implemented.
[0124] Various aspects of the present disclosure are described herein with reference to flowcharts and / or block diagrams of methods, apparatuses, devices, and computer program products implemented according to the present disclosure. It should be understood that each block of the flowcharts and / or block diagrams, and combinations of blocks in the flowcharts and / or block diagrams, can be implemented by computer-readable program instructions.
[0125] These computer-readable program instructions can be provided to a processing unit of a general-purpose computer, a special-purpose computer, or other programmable data processing device, thereby producing a machine, such that when these instructions are executed by the processing unit of the computer or other programmable data processing device, a device is generated that implements the functions / actions specified in one or more blocks in the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium, where these instructions cause the computer, programmable data processing device, and / or other device to operate in a specific manner. Thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing various aspects of the functions / actions specified in one or more blocks in the flowchart and / or block diagram.
[0126] Computer-readable program instructions can be loaded onto a computer, other programmable data processing apparatus, or other device so that a series of operational steps are performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions executed on the computer, other programmable data processing apparatus, or other device to implement the functions / actions specified in one or more boxes in the flowchart and / or block diagram.
[0127] The flow charts and block diagrams in the accompanying drawings show the possible architecture, functions and operations of the systems, methods and computer program products according to multiple implementations of the present disclosure. In this regard, each box in the flow chart or block diagram can represent a part for a module, program segment or instruction, and a part for a module, program segment or instruction comprises one or more executable instructions for realizing the logical function of the specification. In some alternative implementations, the functions marked in the box can also occur in a sequence different from that marked in the accompanying drawings. For example, two continuous boxes can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flow chart, and the combination of the boxes in the block diagram and / or flow chart can be realized by a special hardware-based system that performs the function or action of the specification, or can be realized by a combination of special hardware and computer instructions.
[0128] While various implementations of the present disclosure have been described above, the foregoing description is intended to be illustrative, not exhaustive, and not limited to the disclosed implementations. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described implementations. The terminology used herein is selected to best explain the principles of the implementations, their practical applications, or improvements to existing technologies, or to enable others skilled in the art to understand the various implementations disclosed herein.
Claims
1. A method for transmitting streaming media data, comprising: Receiving a request for a first data block of streaming media data from a first device, the first data block being associated with a first frame of the streaming media data at the first device, and the first data block comprising at least a media presentation description (MPD) associated with the streaming media data, audio data associated with the first frame, and video data associated with the first frame; as well as Send the first data block to the first device. 2 . The method according to claim 1 , wherein the audio data and the video data are encapsulated in the first data block in a sequence based on a decoding timestamp.
3. The method according to claim 1, further comprising: In response to receiving the request, the first data block is generated.
4. The method according to claim 3, wherein generating the first data block comprises: Selecting a target video data block from at least one video data block encapsulated for the streaming media data, the target video data block including the video data associated with the first frame; Selecting a target audio data block from at least one audio data block encapsulated for the streaming media data, the target audio data block including the audio data associated with the first frame; as well as The content of the target video data block and the content of the target audio data block are encapsulated in a media data block in the first data block. The method according to claim 4 , wherein the target video data block or the target audio data block is selected based on a delay time of the first frame.
6. The method according to claim 1, further comprising: generating a set of data chunks based on at least a portion of the streaming media data; as well as In response to receiving the request, the first data block is selected from the set of data blocks. The method of claim 6 , wherein the first data block is selected based on a delay time of the first frame.
8. The method of claim 6, wherein generating a set of data blocks comprises: In response to encoding of a current frame in the streaming media data being completed, determining whether audio data and video data of the current frame belong to a current data block in the set of data blocks; In response to determining that the audio data and the video data of the current frame belong to a current data block, encapsulating the audio data and the video data of the current frame in the current data block; as well as In response to determining that the audio data and the video data of the current frame do not belong to a current data block, another data block in the set of data blocks is generated based on the audio data and the video data of the current frame.
9. The method of claim 8, wherein generating another data block in the set of data blocks comprises: Encapsulating the current MPD in the other data block; as well as The audio data and the video data of the current frame are encapsulated in the another data block.
10. The method according to claim 1, wherein the first data block further comprises at least one of the following: initialization information associated with decoding the audio data and the video data, The start time of the audio data, The start time of the video data, The sequence number of the audio data block corresponding to the audio data, or The sequence number of the video data block corresponding to the video data. The method according to claim 1 , wherein the streaming media data comprises live data. 12 . The method according to claim 1 , wherein the streaming media data is transmitted according to a Dynamic Adaptive Streaming over Hypertext Transfer Protocol (DASH) protocol, and the first data block is a segment and is encapsulated based on an MP4 format.
13. The method according to any one of claims 1 to 12, wherein the method is implemented at a server and the first device comprises a client.
14. A method for transmitting streaming media data, comprising: At a first device, sending a request for a first data block of streaming media data to a second device, where the first data block is associated with a first frame of the streaming media data at the first device, and the first data block includes at least a media presentation description (MPD) associated with the streaming media data, audio data associated with the first frame, and video data associated with the first frame; as well as The first data block is received from the second device. 15 . The method according to claim 14 , wherein the audio data and the video data are encapsulated in the first data block in a sequence based on a decoding timestamp.
16. The method according to claim 14, further comprising: The first frame is displayed based on the first data block.
17. The method according to claim 14, wherein the first data block further comprises at least one of the following: initialization information associated with decoding the audio data and the video data, The start time of the audio data, The start time of the video data, The sequence number of the audio data block corresponding to the audio data, or The sequence number of the video data block corresponding to the video data. The method of claim 14 , wherein the streaming media data comprises live data.
19. The method according to claim 16, wherein the streaming media data is transmitted according to a Dynamic Adaptive Streaming over Hypertext Transfer Protocol (DASH) protocol, and the first data block is a segment and is encapsulated based on an MP4 format.
20. The method of any one of claims 14 to 19, wherein the first device comprises a client and the second device comprises a server.
21. A device for transmitting streaming media data, comprising: a request receiving module configured to: receive a request for a first data block of streaming media data from a first device, where the first data block is associated with a first frame of the streaming media data at the first device, and the first data block includes at least a media presentation description (MPD) associated with the streaming media data, audio data associated with the first frame, and video data associated with the first frame; as well as The data block sending module is configured to: send the first data block to the first device.
22. A device for transmitting streaming media data, comprising: a request sending module configured to: send, at a first device, a request for a first data block of streaming media data to a second device, where the first data block is associated with a first frame of the streaming media data at the first device, and the first data block includes at least a media presentation description (MPD) associated with the streaming media data, audio data associated with the first frame, and video data associated with the first frame; as well as The data block receiving module is configured to receive the first data block from the second device.
23. An electronic device comprising: at least one processing unit; as well as At least one memory, the at least one memory being coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit, the instructions, when executed by the at least one processing unit, causing the electronic device to perform the method according to any one of claims 1 to 13 or the method according to any one of claims 14 to 20.
24. A computer-readable storage medium having instructions stored thereon, which, when executed by a processor, cause the processor to implement the method according to any one of claims 1 to 13 or the method according to any one of claims 14 to 20.
25. A computer program product tangibly stored in a computer storage medium and comprising computer executable instructions which, when executed by a device, cause the device to perform the method of any one of claims 1 to 13 or the method of any one of claims 14 to 20.
Citation Information
Patent Citations
Media resource push method, client side and server
CN107920108A
DASH media stream transmission method, electronic equipment and storage medium
CN113794898A
Live broadcast starting method, equipment and program product
CN114630157A
Reception device for receiving a plurality of real-time transfer streams, transmission device for transmitting same, and method for playing multimedia content
US20130293677A1