Data processing method and apparatus, protocol conversion method and apparatus, and device, medium and program product
By adding slice position identification to the live broadcast system, the problem of inconsistent media stream data on different protocol conversion devices is solved, and distributed consistent slicing is achieved, which improves the efficiency and playback effect of protocol conversion.
Patent Information
- Application Number
- PCT/CN2025/081218
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-01-08
- Filing Date
- 2025-03-07
- Publication Date
- 2025-07-17
AI Technical Summary
In a live broadcast system, when the same media stream performs protocol conversion on different protocol conversion devices, the generated data files are inconsistent, resulting in playback errors from the player, and the existing solutions are less flexible and usable.
By adding slice position identification to the original media stream data, it is ensured that the target data complies with the same first data protocol as the original data, and generate slice files that comply with different from the first data protocol, so as to achieve distributed consistency of slices.
Ensure that the protocol conversion results of different protocol conversion devices for the same media stream are consistent, avoid errors caused by inconsistent slicing results of the player, improve the efficiency and fault tolerance of protocol conversion, and improve the playback effect of media stream data.
Smart Images

Figure CN2025081218_17072025_PF_FP_ABST
Abstract
Description
Method, apparatus, device, medium and program product for data processing and protocol conversion
[0001] This application claims priority to the Chinese invention patent application entitled “Data processing method, protocol conversion method, device, equipment and storage medium” and application number 202410029489.9, filed on January 8, 2024. The entire contents of that application are incorporated herein by reference. Technical Field
[0002] The embodiments of the present disclosure generally relate to the field of computer technology, and more particularly, to methods, apparatuses, devices, media, and program products for data processing, and methods, apparatuses, devices, media, and program products for protocol conversion. Background Art
[0003] In recent years, with the increasing popularity of basic communications such as 4G / 5G, the live streaming industry has flourished both domestically and globally. Currently, in the internet live streaming technology system, the production or streaming end of live content generally uses protocols based on Hypertext Transfer Protocol (HTTP) or Real Time Messaging Protocol (RTMP) to push data streams to cloud servers. The streaming end of live content generally uses protocols such as Dynamic Adaptive Streaming over HTTP (DASH), HTTP-Flash Video (FLV), HTTP Live Streaming (HLS), and RTMP for streaming.
[0004] Segment-based streaming protocols (such as the DASH protocol, the HLS protocol, and so on) use a method of dividing the media stream into multiple media segments for transmission. This allows efficient and flexible transmission of audio and video content over the Internet, providing users with a high-quality viewing experience under different devices and network conditions. In the case of using a segment-based streaming protocol for pulling streams, since the push stream adopts the HTTP-FLV or RTMP protocol, protocol conversion is involved. However, when performing protocol conversion, the data files generated by the protocol conversion of the same media stream on different protocol conversion devices may be inconsistent, resulting in playback errors in the player on the user side. Therefore, how to ensure the consistency of the data files generated by the protocol conversion of the same media stream on different protocol conversion devices (i.e., the distributed consistency of the slices) has become a problem that needs to be solved urgently. Summary of the Invention
[0005] In a first aspect of the present disclosure, a method for data processing is provided. In this method, original data for a media stream is received and sliced into multiple data segments. Furthermore, at least one slice position identifier for each of the multiple data segments is added to the original data to obtain target data for the media stream. The target data and the original data conform to the same first data protocol.
[0006] In a second aspect of the present disclosure, a method for protocol conversion is provided. In this method, target data for a media stream is received. The target data conforms to a first data protocol and includes at least one slice location identifier for multiple data segments in the target data. Furthermore, based on the target data and the at least one slice location identifier, multiple slice files corresponding to the multiple data segments are generated. The multiple slice files conform to a second data protocol different from the first data protocol.
[0007] In a third aspect of the present disclosure, a device for data processing is provided. The device includes a receiving module, a slicing module, and an adding module. The receiving module is configured to receive original data for a media stream. The slicing module is configured to slice the original data into multiple data segments. The adding module is configured to add at least one slice position identifier for each of the multiple data segments to the original data to obtain target data for the media stream. The target data conforms to the same first data protocol as the original data.
[0008] In a fourth aspect of the present disclosure, a device for protocol conversion is provided. The device includes a receiving module and a generating module. The receiving module is configured to receive target data for a media stream. The target data conforms to a first data protocol and includes at least one slice location identifier for multiple data segments in the target data. The generating module is configured to generate, based on the target data and the at least one slice location identifier, multiple slice files corresponding to the multiple data segments. The multiple slice files conform to a second data protocol different from the first data protocol.
[0009] In a fifth aspect of the present disclosure, an electronic device is provided. The electronic device includes: at least one processor; and at least one memory, the at least one memory being coupled to the at least one processor and storing instructions for execution by the at least one processor, wherein the instructions, when executed by the at least one processor, cause the electronic device to perform the method according to the first aspect of the present disclosure or the method according to the second aspect of the present disclosure.
[0010] In a sixth aspect of the present disclosure, a computer-readable storage medium is provided, on which instructions are stored. When the instructions are executed by a processor, the processor implements the method according to the first aspect of the present disclosure or the method according to the second aspect of the present disclosure.
[0011] In a seventh aspect of the present disclosure, a computer program product is provided. The computer program product is tangibly stored in a computer storage medium and includes computer-executable instructions that, when executed by a device, cause the device to perform the method according to the first aspect of the present disclosure or the method according to the second aspect of the present disclosure.
[0012] It should be understood that the content described in this summary section is not intended to limit the key features or important features of the embodiments of the present disclosure, nor is it intended to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0013] The above and other features, advantages and aspects of the embodiments of the present disclosure will become more apparent hereinafter with reference to the following detailed description in conjunction with the accompanying drawings. In the accompanying drawings, the same or similar reference numerals represent the same or similar elements, wherein:
[0014] FIG1 is a schematic diagram illustrating an example environment in which various embodiments of the present disclosure can be implemented;
[0015] FIG2 shows a flow chart of a method for data processing according to some embodiments of the present disclosure;
[0016] FIG3 shows a schematic diagram of data processing according to some embodiments of the present disclosure;
[0017] FIG4 shows a flowchart of a method for protocol conversion according to some embodiments of the present disclosure;
[0018] FIG5 shows a schematic diagram of slicing processing for target media stream data according to some embodiments of the present disclosure;
[0019] FIG6 shows a schematic diagram of an example process of media stream distribution according to some embodiments of the present disclosure;
[0020] FIG7 shows a schematic diagram of another example process of media stream distribution according to some embodiments of the present disclosure;
[0021] FIG8 shows a block diagram of an example apparatus for data processing according to some embodiments of the present disclosure;
[0022] FIG9 shows a block diagram of an example apparatus for protocol conversion according to some embodiments of the present disclosure; and
[0023] FIG10 illustrates a block diagram of a device in which one or more embodiments of the present disclosure may be implemented. DETAILED DESCRIPTION
[0024] The following describes embodiments of the present disclosure in more detail with reference to the accompanying drawings. Although certain embodiments of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the embodiments described herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are for illustrative purposes only and are not intended to limit the scope of protection of the present disclosure.
[0025] It should be understood that the various steps described in the method embodiments of the present disclosure may be performed in different orders and / or in parallel. In addition, the method embodiments may include additional steps and / or omit the steps shown. The scope of the present disclosure is not limited in this respect.
[0026] In the description of the embodiments of the present disclosure, the term "including" and similar terms should be understood as open inclusion, that is, "including but not limited to". The term "based on" should be understood as "based at least in part on". The term "one embodiment" or "the embodiment" should be understood as "at least one embodiment". The term "some embodiments" should be understood as "at least some embodiments". The following may also include other explicit and implicit definitions. As used herein, the term "model" can represent the association relationship between various data. For example, the above-mentioned association relationship can be obtained based on a variety of technical solutions currently known and / or to be developed in the future.
[0027] As used herein, the term "in response to" refers to a state in which a corresponding event occurs or a condition is satisfied. It will be understood that the timing of executing a subsequent action executed in response to the event or condition is not necessarily strongly correlated with the time when the event occurs or the condition is satisfied. For example, in some cases, the subsequent action may be executed immediately when the event occurs or the condition is satisfied; in other cases, the subsequent action may be executed some time after the event occurs or the condition is satisfied.
[0028] It is understandable that the data involved in this technical solution (including but not limited to the data itself, the acquisition or use of the data) must comply with the requirements of relevant laws, regulations and relevant provisions.
[0029] It is understandable that before using the technical solutions disclosed in the various embodiments of this disclosure, the type, scope of use, usage scenarios, etc. of the personal information involved in this disclosure should be informed to the user and the user's authorization should be obtained in an appropriate manner in accordance with relevant laws and regulations.
[0030] For example, in response to a user's active request, a prompt message is sent to the user to clearly inform the user that the operation requested will require the acquisition and use of the user's personal information. This allows the user to independently choose whether to provide personal information to the electronic device, application, server, storage medium, or other software or hardware that performs the operations of the disclosed technical solution based on the prompt message.
[0031] As an optional but non-limiting embodiment, in response to receiving a user's active request, the prompt information may be sent to the user in the form of a pop-up window, in which the prompt information may be presented in text form. Furthermore, the pop-up window may also include a selection control for the user to select "agree" or "disagree" to provide personal information to the electronic device.
[0032] It is understandable that the above notification and user authorization process are merely illustrative and do not limit the embodiments of the present disclosure. Other methods that meet relevant laws and regulations may also be applied to the embodiments of the present disclosure.
[0033] As briefly mentioned above, in a live broadcast scenario, when different protocols are used in the push and pull processes, such as when FLV or RTMP protocols are used in the push process and DASH or HLS protocols are used in the pull process, a conversion from FLV or RTMP protocol to DASH or HLS protocol is involved. This protocol conversion is generally performed by a protocol conversion device. The protocol conversion device usually slices the media stream data according to certain requirements and converts it into small transport stream (TS) files or fragmented MPEG-4 (Fragmented MP4, FMP4) files, and updates the playlist file (manifest) in a timely manner, such as an m3u8 file or a media presentation description (MPD).
[0034] When the same media stream is converted on different protocol converters, the resulting slices may be inconsistent. Inconsistent slices can cause unexpected playback errors in players. Existing solutions often use a single protocol converter, meaning that the same media stream is always converted on the same device. This makes it impossible to switch protocol converters, resulting in poor flexibility and usability.
[0035] To this end, various embodiments of the present disclosure propose a solution for achieving distributed consistency of slices by adding slice position identifiers to the original media stream data. Specifically, according to some embodiments of the present disclosure, a solution for data processing is proposed, which includes: receiving original data for a media stream; slicing the original data into multiple data segments; and adding at least one slice position identifier for the multiple data segments to the original data to obtain target data for the media stream, the target data and the original data conforming to the same first data protocol. In addition, according to some other embodiments of the present disclosure, a solution for protocol conversion is also proposed, which includes: receiving target data for a media stream, the target data conforming to a first data protocol and including at least one slice position identifier for multiple data segments in the target data; and generating multiple slice files corresponding to the multiple data segments based on the target data and the at least one slice position identifier, the multiple slice files conforming to a second data protocol different from the first data protocol.
[0036] It will be more clearly understood through the description below that according to the embodiments of the present disclosure, a slice position identifier is added to the original media stream data to obtain the target media stream data, and the target media stream data is transparently transmitted in the transmission system for the media stream. At the device that performs the protocol conversion, the media stream data is subjected to protocol conversion based on the slice position identifier. In this way, it is possible to ensure that the protocol conversion results of different protocol conversion devices for the same media stream data are consistent, thereby achieving distributed consistency of slices. Therefore, this makes it possible to support protocol conversion of the same media stream on different protocol conversion devices, and ensures that when switching protocol conversion devices, the player will not have playback errors due to inconsistent slicing results, thereby improving the efficiency and fault tolerance of the protocol conversion and improving the playback effect of the media stream data.
[0037] Various example implementations of the solution will be described in detail below with reference to the accompanying drawings. First, refer to Figure 1, which shows a schematic diagram of an example environment 100 in which the various embodiments of the present disclosure can be implemented. The example environment 100 is shown for a live broadcast scenario, wherein the example environment 100 as a whole may include an anchor 110, a streaming device 120, a streaming system 130, streaming devices 140-1, 140-2, ..., 140-N, and viewers 150-1, 150-2, ..., 150-N, where N is any suitable positive integer, such as 1, 5, 70, 500, and so on. The streaming device 120 is communicatively coupled to the streaming system 130, and the streaming system 130 is communicatively coupled to the streaming device 140. The streaming system 130 can be implemented through a content delivery network (CDN) or other network. In the example shown in FIG1 , the streaming system 130 may include transmission devices 132-1, 132-2, ..., 132-M, where M is any suitable positive integer, such as 20, 100, 500, etc. For ease of description, the transmission devices 132-1, 132-2, ..., 132-M may also be individually or collectively referred to as transmission devices 132, the stream pulling devices 140-1, 140-2, ..., 140-N may also be individually or collectively referred to as stream pulling devices 140, and the viewers 150-1, 150-2, ..., 150-N may also be individually or collectively referred to as viewers 150.
[0038] As shown in Figure 1, the host 110 can, for example, initiate a live broadcast by operating the streaming device 120. The streaming device 120 can then upload local audio and video data to the streaming system 130 via the network according to a specific protocol (such as the RTMP protocol or the FLV protocol). The viewer 150 can, for example, initiate viewing of the host 110's live content by operating the streaming device 140. The streaming device 140 can then request the host 110's media stream from the streaming system 130, obtain the audio and video data from the streaming system 130, and decode and play it for the viewer 150 to watch and listen to.
[0039] In FIG1 , the streaming device 120 and the streaming device 140 are shown as mobile phones, but the streaming device 120 and / or the streaming device 140 can also be any type of mobile terminal or portable terminal, including a laptop computer, a notebook computer, a netbook computer, a tablet computer, a media computer, a multimedia tablet, a personal communication system (PCS) device, a personal navigation device, a personal digital assistant (PDA), an audio / video player, a digital camera / camcorder, a positioning device, a television receiver, a radio broadcast receiver, an e-book device, a gaming device, or any combination thereof, including accessories and peripherals of these devices or any combination thereof. In some embodiments, the streaming device 120 and / or the streaming device 140 can also support any type of interface for the user (such as a "wearable" circuit, etc.).
[0040] It should be understood that the structure and functionality of the environment 100 are described for exemplary purposes only and do not imply any limitation on the scope of the present disclosure. For example, the methods according to the embodiments of the present disclosure may also be applied to any other suitable scenarios besides live broadcast scenarios, such as on-demand scenarios, etc.
[0041] Figure 2 shows a flow chart of a method 200 for data processing according to some embodiments of the present disclosure. In some embodiments, method 200 may be performed at streaming device 120 or transmission device 132 as shown in Figure 1. It should be understood that method 200 may include additional blocks not shown and / or may omit one (or some) of the blocks shown, and the scope of the present disclosure is not limited in this respect.
[0042] In block 202, original data for a media stream is received. The original data of the media stream may conform to a first data protocol and be media stream data to be converted to a protocol, such as media stream data generated or uploaded by a host during a live broadcast. In the context of this disclosure, the original data of the media stream may also be referred to as original media stream data. Original media stream data may include video data and / or audio data. Data protocols may include data transmission protocols, data encapsulation protocols, and / or other related protocols. For example, the first data protocol may be the data transmission protocol used by the original media stream data, i.e., the original data transmission protocol to be converted to a protocol. In some embodiments, the first data protocol may include the FLV protocol or the RTMP protocol. For example, the original media stream data may be transmitted based on the FLV protocol or the RTMP protocol. It should be understood that the first data protocol may also be any other suitable data protocol, and the scope of this disclosure is not limited in this respect. In the example shown in FIG. 1 , the transmission device 132-1 may receive original media stream data uploaded by the streaming device 120 based on the first data protocol. In this case, the transmission device 132-1 may also be referred to as a streaming node.
[0043] In block 204, the original data is sliced into a plurality of data segments. In some embodiments, the original data may be sliced based on a preset data length of a second data protocol. The second data protocol is different from the first data protocol and may be a data transmission protocol to which the original media stream data needs to be converted, i.e., a target data transmission protocol for protocol conversion. Exemplarily, the second data protocol may be a segment-based streaming protocol, such as the DASH protocol, the HLS protocol, or the like. For example, the original media stream data may be converted from the first data protocol to the DASH protocol or the HLS protocol.
[0044] In some embodiments, the preset data length can be a predetermined duration of a data segment, that is, the slice length (i.e., slice duration) when converted to the second data protocol. For example, the expected slice length, minimum slice length, etc. when converting the original media stream data from the first data protocol to the second data protocol can be set in advance as needed. The original data can be divided into multiple data segments according to the preset data length. For example, the slice positions in the original media stream data can be determined in sequence according to the preset data length of the second data protocol, and the data length between two adjacent slice positions is greater than or equal to the preset data length. For example, the data length between two adjacent slice positions is not less than the preset data length and the difference with the preset data length is within a set range. Exemplarily, the data length is as close to the preset data length as possible while meeting the requirements.
[0045] In one example embodiment, the specific position of the slice position in the original media stream data may not be limited. For example, the relative positional relationship between the slice position and the frames included in the original media stream data may be disregarded, and the slice position in the original media stream data may be determined directly based on a preset data length. For example, a slice position may be determined at intervals of a preset data length in the original media stream data. In this case, the data length between two adjacent slice positions in the original media stream data may be equal to the preset data length.
[0046] In another example embodiment, the specific position of the slice position in the original media stream data can be defined. For example, the relative positional relationship between the slice position and the position of a random access point (e.g., an I-frame) in the media stream can be considered. In other words, the relative positional relationship between the slice position and at least a portion of the frames included in the original media stream data can be considered. Exemplarily, the slice position can be defined to meet a preset position condition. In this case, the original slice position of the original media stream data can be determined based on a preset data length of the second data protocol. Here, the original slice position can be understood as a candidate slice position preliminarily determined based on the preset data length for the second data protocol. The original slice position may or may not meet the preset position condition. If the determined original slice position meets the preset position condition, the original slice position is used as the actual slice position of the original media stream data. If the determined original slice position does not meet the preset position condition, the original slice position is adjusted based on the preset position condition, and the adjusted original slice position is used as the actual slice position of the original media stream data.
[0047] Exemplarily, the original slice position of the original media stream data can be determined based on a preset data length for the second data protocol. For example, the current original slice position is determined by separating the current original slice position from the previous slice position in the original media stream data by a preset data length. A determination is then made as to whether the original slice position meets a preset position condition. If the preset position condition is met, the original slice position can be determined as an actual slice position for the original media stream data. If the preset position condition is not met, the original slice position can be adjusted based on the preset position condition to a slice position that meets the preset position condition. In this manner, when the preset position condition exists, the actual slice position can be a slice position that meets the preset position condition, such as a slice position determined based on the preset data length of the second data protocol and the preset position condition.
[0048] In some embodiments, the preset position condition may be a condition for defining the relative positional relationship between the slice position and at least some frames in the original media stream data. For example, these at least some frames may include video key frames (I frames), video forward prediction frames (P frames), audio frames (A frames), and so on. The preset position condition may be set as needed. As an example, the preset position condition may be that frames located after and adjacent to the slice position must be video key frames. For example, when adjusting the original slice position based on the preset position condition, the original slice position may be adjusted to a position located before and adjacent to a certain video key frame. For example, the original slice position may be adjusted backward to a position located before and adjacent to a certain video key frame. The video key frame considered may also be set as needed. Exemplarily, the video key frame may be the video key frame in the original media stream data that is closest to the original slice position. This ensures that the data length between two adjacent slice positions is as close as possible to the preset data length of the second data protocol, while satisfying the preset position condition, thereby improving the effectiveness of subsequent protocol conversion. In this case, the original slice position can be adjusted backward to before the video key frame closest to the original slice position, that is, the original slice position can be adjusted backward to a position before and adjacent to the video key frame closest to the original slice position. In this way, it can be ensured that the first frame of the slice file (such as a video file) obtained by protocol conversion based on the slice position is a video key frame, thereby achieving random access and independent start-up, further improving the playback effect of the data file obtained by protocol conversion.
[0049] FIG3 shows a schematic diagram 300 of data processing according to some embodiments of the present disclosure. As shown in the upper half of FIG3 , the original data is sliced at slice position 321, slice position 322, and slice position 323, thereby obtaining data segments 311, data segment 312, and data segment 313. In FIG3 , V(I) represents the data of a video I frame, V(P) represents the data of a video P frame, and A represents the data of an audio frame. It can be seen that in the example of FIG3 , slice position 321, slice position 322, and slice position 323 are all before and adjacent to the I frame, thereby ensuring random access and independent start of play.
[0050] It should be noted that, in box 204, the original data is only sliced into multiple data segments, and multiple slice files that comply with the second data protocol are not generated. In other words, what is executed in box 204 is a simulated slicing process (i.e., a simulated protocol conversion process) for the purpose of determining the slice position, and no transpackaging operation is performed. In other words, no slice file for the second data protocol is generated in the simulated slicing process. It should be understood that the original data can also be sliced into multiple data segments in any other suitable manner, for example, by considering other factors that affect the slice position. The scope of the present disclosure is not limited in this respect.
[0051] Returning to reference FIG. 2 , at block 206, at least one slice position identifier for the plurality of data segments is added to the original data to obtain target data for the media stream. The target data and the original data conform to the same first data protocol. Each of the at least one slice position identifier may correspond to two adjacent data segments from the plurality of data segments and indicate the location of a boundary between the two adjacent data segments. For example, each of the at least one slice position identifier may be inserted between the corresponding two adjacent data segments.
[0052] In some embodiments, a slice position identifier can be added to the original data based on the slice position determined by the slicing operation in block 204 to obtain media stream data with the slice position identifier added, i.e., the target data of the media stream. In the context of this disclosure, the target data of the media stream may also be referred to as the target media stream data. Referring to Figure 3, the target data has slice position identifiers 331, 332, and 333 added to the original data. Slice position identifier 331 can correspond to data segment 311 and data segment 312 and be inserted between data segment 311 and data segment 312 to indicate the boundary between the two data segments. Similarly, slice position identifier 332 can correspond to data segment 312 and data segment 313 and be inserted between data segment 312 and data segment 313 to indicate the boundary between the two data segments. In this way, slice positions can be efficiently indicated, thereby improving the efficiency of subsequent protocol conversion. It should be understood that the slice position identifier may also indicate the position of the boundary between two adjacent data segments in any other suitable manner. For example, the slice position identifier may record the slice position based on a presentation time stamp (PTS) or a decoding time stamp (DTS) of a frame adjacent to the slice position. The scope of the present disclosure is not limited in this respect.
[0053] In some embodiments, the slice location identifier may include description information of multiple media segments corresponding to the multiple data segments of the media stream, and the description information is used for the second data protocol. In the context of the present disclosure, the code stream of the corresponding media segment can be obtained from a data segment by decapsulation. In some embodiments, the description information may include identification information, timestamp starting position, slice duration, Uniform Resource Locator (URL), and the like of the latest consecutive one or more video files and audio files. For example, the description information may be implemented in the form of a playlist file. Exemplarily, when the second data protocol includes the DASH protocol, the description information may include an MPD file for the DASH protocol of the multiple media segments. In the example shown in Figure 3, the slice location identifiers 331, 332, and 333 are script frames containing MPD information. Alternatively or additionally, when the second data protocol includes the HLS protocol, the description information may include index information for the HLS protocol of the multiple media segments, such as an m3u8 file. For the purpose of facilitating understanding, a non-limiting example of an m3u8 file is given below, wherein the text after “ / / ” is an explanation of the corresponding content added for the purpose of facilitating understanding.
[0054] It should be understood that the information mentioned in the above m3u8 file is only exemplary and non-restrictive. It may also include additional information not listed and / or may omit some (or some) of the listed information, and the scope of the present disclosure is not limited in this respect.
[0055] In some embodiments, the slice location identifier may further include an identifier for a method. For example, the identifier for a method may be implemented in the form of a method name. The method name of the method according to each embodiment of the present disclosure may be, for example, onSegmentSync or any other suitable string. It should be understood that the slice location identifier may also include any other suitable information, and the scope of the present disclosure is not limited in this respect.
[0056] In some embodiments, the information included in the slice location identifier can be encapsulated based on a target data structure. The target data structure can depend on the first data protocol. Exemplarily, when the first data protocol includes the FLV protocol, the target data structure includes a script data (SCRIPTDATA) structure, also known as a Script frame structure. The script data structure can include a script tag body (ScriptTagBody) encoded in an action message format (AMF). The script tag body is used to encapsulate method calls, which include: an item containing a method name and a corresponding set of parameters (arguments). The set of parameters can, for example, have the type of an ECMA array (ECMA Array). The array can include properties for achieving distributed consistency of slices. The availability of these properties can depend on the software used to create the FLV stream, the second data protocol, and whether multitrack streaming (Multitrack Streaming) via Enhanced RTMP (E-RTMP) is being used, etc. Table 1 below shows some exemplary properties.
[0057] Table 1 – Typical properties available in the onSegmentSync parameter object
[0058] For ease of understanding, a non-limiting example of the attribute trackidM3U8Map is given below.
[0059] It can be seen that the track of trackId=1 corresponds to the m3u8 playlist for high-definition video (expressed as video-hd), the track of trackId=2 corresponds to the m3u8 playlist for ultra-high-definition video (expressed as video-uhd), and the track of trackId=3 corresponds to the m3u8 playlist for audio (expressed as audio). In this way, the m3u8 playlist corresponding to the track can be determined based on the track index of each track of the media stream based on the above mapping relationship. It should be pointed out that, in the above example, " / n" represents a line feed. It should be understood that this example for the attribute trackidM3U8Map is shown only for illustrative purposes, and the scope of the present disclosure is not limited in this respect.
[0060] It should be understood that the attributes listed in Table 1 are exemplary and non-restrictive. Additional attributes not listed may also be included and / or one (or some) of the listed attributes may be omitted, and the scope of the present disclosure is not limited in this respect. For example, when multi-track streaming is not enabled, the trackidM3U8Map attribute may be omitted and the attribute M3U8 may be additionally included. The attribute M3U8 is of type string and carries the m3u8 playlist for the HLS protocol corresponding to the current media data segment.
[0061] In some embodiments, the target media stream data may further include one or more exemplary attributes as shown in Table 2 below.
[0062] Table 2 – Typical properties available in onSegmentSync or onMetaData parameter objects
[0063] Here, the end-to-end synchronization mechanism means that the attributes are transparently transmitted along with the media stream data in the streaming system 130 .
[0064] In some embodiments, the trackRepresentationMap attribute can adopt a key-value structure, where the key can be trackId, which serves as a unique identifier for a specific audio track or video track; the value can be a DASH representation ID corresponding to the track. It should be understood that the value can also be an International Organization for Standardization Base Media File Format (ISOBMFF) file track ID, etc. For ease of understanding, a non-limiting example of the trackRepresentationMap attribute is given below.
[0065] In the above example, 0, 1, 2, and 3 are keys, and "video_1080p", "video_720p", "audio_eng_128k", and "audio_eng_256k" are values. It can be seen that the track with trackId=0 corresponds to the DASH representation ID "video_1080p", which represents a video with a resolution of 1080p; the track with trackId=1 corresponds to the DASH representation ID "video_720p", which represents a video with a resolution of 720p; the track with trackId=2 corresponds to the DASH representation ID "audio_eng_128k", which represents English audio with a bit rate of 128kbps; and the track with trackId=3 corresponds to the DASH representation ID "audio_eng_256k", which represents English audio with a bit rate of 256kbps. In this way, the DASH representation ID of each track of the media stream can be determined according to the track index of the track based on the above mapping relationship. It should be understood that this example for the attribute trackRepresentationMap is shown for illustration purposes only and the scope of the present disclosure is not limited in this regard.
[0066] In some embodiments, the attribute trackHlsStreamMap may adopt a key-value pair structure, where the key may be trackId, which serves as a unique identifier for a specific audio track or video track; and the value may be the HLS stream m3u8 file name corresponding to the track. For ease of understanding, a non-limiting example of the attribute trackHlsStreamMap is given below.
[0067] In the above example, 0, 1, 2, and 3 are keys, and "video_1080p.m3u8", "video_720p.m3u8", "audio_eng.m3u8", and "audio_eng_alt.m3u8" are values. It can be seen that the track with trackId = 0 corresponds to the m3u8 file name "video_1080p.m3u8" for a video stream with a resolution of 1080p, the track with trackId = 1 corresponds to the m3u8 file name "video_720p.m3u8" for a video stream with a resolution of 720p, the track with trackId = 2 corresponds to the m3u8 file name "audio_eng.m3u8" for an English audio stream, and the track with trackId = 3 corresponds to the m3u8 file name "audio_eng_alt.m3u8" for another English audio stream. In this way, based on the above mapping relationship, the HLS stream corresponding to each track of the media stream can be determined according to the track index of the track. It should be understood that this example for the attribute trackHlsStreamMap is shown for illustration purposes only and the scope of the present disclosure is not limited in this regard.
[0068] In some embodiments, one or more of the above-mentioned attributes hlsMultivariantPlaylist, trackRepresentationMap and trackHlsStreamMap may be included in the slice location identifier, for example, encapsulated in a Script frame. In other words, each slice location identifier carries one or more of these attributes. Alternatively, since the information carried by the above-mentioned attributes hlsMultivariantPlaylist, trackRepresentationMap and trackHlsStreamMap is relatively fixed, one or more of these attributes may also be transmitted once in the session, that is, independent of the slice location identifier. In this case, the one or more attributes may be included in the metadata onMetaData. The metadata onMetaData is mainly used to describe the relevant parameters of the encoding format of audio and video, and may adopt the script data format described above. By transmitting the above-mentioned attributes hlsMultivariantPlaylist, trackRepresentationMap and / or trackHlsStreamMap in the metadata onMetaData, network bandwidth can be saved.
[0069] It should be understood that the names of the attributes mentioned above are merely exemplary and non-limiting, and the names of these attributes may also be any other suitable character strings. The scope of the present disclosure is not limited in this respect.
[0070] In some embodiments, each slice position identifier can carry one or more playlist files, which correspond to the current media data segment associated with the script data segment. In one example embodiment, the current media data segment associated with the script data segment can be composed of all media data that are located after the script data segment until the next script data segment or the end of the media stream (whichever is earlier). In another example embodiment, the current media data segment associated with the script data segment can be composed of all media data that are located before the script data segment until the previous script data segment or the beginning of the media stream (whichever is later).
[0071] In addition, in the case where the first data protocol includes the RTMP protocol, the target data structure may correspond to a metadata message (Metadata Message) or a data message (Data Message). For example, the target data structure may also include an item containing a method name and a corresponding set of parameters. The set of parameters may, for example, have the type of an ECMA array (ECMA Array), and the array may include properties for achieving distributed consistency of slices. This is similar to the properties described above with reference to the script tag body, and the present disclosure will not repeat them here. In this way, a predetermined data structure can be used to better encapsulate the parameters and information used for the methods according to the various embodiments of the present disclosure, so that it can be transmitted through the first data protocol.
[0072] In some embodiments, the description information includes description information corresponding to each track of at least one track of the media stream, and the target data includes a mapping relationship between each track and the corresponding description information. Exemplarily, when the second data protocol includes the HLS protocol, the slice location identifier includes a first indication (e.g., the above-mentioned attribute trackidM3U8Map) for indicating the mapping relationship between the track index and the index information of the multiple media segments for the HLS protocol, and the target data includes a second indication (e.g., the above-mentioned attribute hlsMultivariantPlaylist) for indicating the HLS multi-variant playlist, and a third indication (e.g., the above-mentioned attribute trackHlsStreamMap) for indicating the mapping between the track index and the HLS variant stream. When the second data protocol includes the DASH protocol, the slice location identifier includes the MPD for the DASH protocol of the multiple media segments (e.g., the above-mentioned attribute mpd), and the target data includes a fourth indication (e.g., the above-mentioned attribute trackRepresentationMap) for indicating the mapping between the track index and the DASH representation. In this way, the method according to each embodiment of the present disclosure can better support the application scenarios of multi-track streaming transmission, thereby expanding the usability and flexibility of the method. In an example embodiment, the second indication, the third indication and / or the fourth indication may be included in the slice location identifier. Alternatively, the second indication, the third indication, and / or the fourth indication may be included in other data parts of the target media stream data (e.g., metadata onMetaData). It should be understood that the above-mentioned first indication, second indication, third indication, and fourth indication may also be implemented in any other suitable manner, and the scope of the present disclosure is not limited in this respect. For example, the fourth indication may also be implemented as one or more flags, each flag indicating that a representation ID in the MPD corresponds to a single track index in the original RTMP / FLV multi-track stream.
[0073] In some embodiments, the slice position identifier is unencrypted. In other words, even if the original data itself is encrypted, the slice position identifier can also be unencrypted. In this way, consistent protocol conversion of media stream data by different protocol conversion devices can be better achieved.
[0074] In some embodiments, the slice position identifiers added to different slice positions may be the same in some aspects. Exemplarily, the slice position identifiers added to different slice positions may be slice position identifiers of the same type, for example, all are data frames (e.g., Script frames) or metadata messages. Alternatively or additionally, the slice position identifiers added to different slice positions may be different in some aspects. Exemplarily, the slice position identifiers added to different slice positions may contain different content, for example, the slice position identifier added at a specific slice position may contain the latest content corresponding to the slice position. The slice position identifier may contain current description information for the second data protocol. In this case, the current description information for the second data protocol of one or more media segments may be obtained, a slice position identifier containing the current description information may be generated, and the generated slice position identifier may be added to the corresponding slice position.
[0075] The description information can be generated by simulating slicing of the original media stream data (i.e., simulating protocol conversion). The simulated protocol conversion may be different from the actual protocol conversion. For example, when performing the simulated protocol conversion, only the slice position may be determined and / or the description information may be generated and updated without ultimately generating a data file that conforms to the second data protocol. The current description information may be the latest description information for the second data protocol of one or more media segments when generating a slice position identifier to be added for a specific slice position. The latest description information may be, for example, a playlist file corresponding to the second data protocol. In some embodiments, the slice position identifiers may be added one by one as the original media stream data is uploaded, and the description information is updated as the original media stream data is uploaded. In this case, the current description information contained in the slice position identifiers added at different slice positions in the original media stream data may be different, so that the description information can be restored from the slice position identifier when performing protocol conversion later.
[0076] For example, for the slice position to which the slice position identifier is currently to be added, information from a playlist file for the second data protocol can be obtained as current description information, and this current description information can be written into the slice position identifier, for example, in the form of a data frame or a data message. Furthermore, the slice position identifier can be added to the corresponding slice position. It should be understood that the description information and the slice position identifier can also be obtained in any other suitable manner. The scope of the present disclosure is not limited in this respect.
[0077] From the above description in combination with Figures 2 to 3, it can be seen that in the method for data processing according to each embodiment of the present disclosure, by slicing the original media stream data into multiple data segments and adding at least one slice position identifier for the multiple data segments to the original data, target media stream data carrying at least slice position information can be obtained. In this way, it is convenient to perform protocol conversion directly based on the slice position information carried in the target media stream data during the subsequent protocol conversion process, which can effectively ensure that the results of protocol conversion performed by different protocol conversion devices on the same media stream data remain consistent, thereby achieving consistent protocol conversion of media stream data by different protocol conversion devices. Therefore, it can be ensured that when switching protocol conversion devices, the player will not have playback errors due to inconsistent protocol conversion results, thereby effectively improving the consistency of protocol conversion and ensuring the playback effect of media stream data.
[0078] In some embodiments, the data processing method according to various embodiments of the present disclosure can be implemented at a first device within a streaming system 130 for a media stream. The first device can receive a request for a media stream from a second device. Further, based on the request, the first device can determine whether the second device is also located within the streaming system 130. For example, the first device can determine whether the second device is located within the streaming system 130 based on parameters or attributes of the request. Alternatively, the request from the second device itself can carry information indicating whether the second device is located within the streaming system 130.
[0079] If it is determined that the second device is within the streaming system 130, the first device can send the target data for the media stream to the second device. In this way, the target media stream data carrying the slice position information can be transparently transmitted in the streaming system. If it is determined that the second device is outside the streaming system 130, the first device can send the original data for the media stream to the second device. In an example embodiment, the original data can be the original data received by the first device. In another example implementation, the first device can delete all added slice position identifiers from the target media stream data, thereby restoring the original media stream data. In this way, errors can be avoided when devices outside the streaming system parse or play the target media stream data.
[0080] In some embodiments, when the streaming system 130 is implemented through a content distribution network, the first device may be a source station within the content distribution network. For example, the first device may be the first source station in the original media stream data upload link. In one example embodiment, the second device may also be a source station within the content distribution network. In another example embodiment, the second device may be an edge node within the content distribution network. Referring to Figure 1, the stream pulling device 140 may request media stream data from the transmission device 132-M nearby. In this case, the second device may correspond to the transmission device 132-M and may also be referred to as a stream pulling node. In another example embodiment, the second device may be a device outside the content distribution network, such as the stream pulling device 140 shown in Figure 1.
[0081] In other embodiments, the first device may be an edge node within a content distribution network. For example, the first device may be the first edge node in the original media stream data upload link, i.e., a push node. In an example embodiment, the second device may be a source station within a content distribution network. In this case, after generating the target media stream data, the push node may send the target media stream data to the source station of the content distribution network for storage. Subsequently, the protocol conversion device may obtain the target media stream data from the data source station of the content distribution network. In another example embodiment, the second device may be an edge node within the same content distribution network. In yet another example embodiment, the second device may be a device outside the content distribution network, such as the pull device 140 shown in FIG. 1 .
[0082] It can be seen that compared with the existing solutions of performing protocol conversion at the central source station and distributing the converted data files to the edge nodes, the method for data processing according to the various embodiments of the present disclosure can support the use of FLV or RTMP protocols for transparent transmission in the streaming system, and the edge nodes also use FLV or RTMP protocols to obtain data back to the source. On the one hand, this can make the encapsulation redundancy lower, and on the other hand, it can also omit the frequent interaction requests when the edge nodes return to the source, thereby saving network bandwidth. In addition, the method for data processing according to the various embodiments of the present disclosure can achieve distributed consistency slicing without relying on external network services, thereby saving costs.
[0083] FIG4 shows a flow chart of a method 400 for protocol conversion according to some embodiments of the present disclosure. In some embodiments, the method 400 may be performed at the transmission device 132 shown in FIG1 . It should be understood that the method 400 may also include additional blocks not shown and / or may omit one (or some) of the blocks shown, and the scope of the present disclosure is not limited in this respect.
[0084] At block 402, target data for a media stream is received. The target data conforms to a first data protocol and includes at least one slice location identifier for a plurality of data segments within the target data. The target data may be generated according to the data processing method described in conjunction with FIG. 2 and FIG. 3 , which has been described in detail above. Therefore, the concepts of target data, first data protocol, slice location identifiers, etc. are not further elaborated herein.
[0085] In block 404, a plurality of slice files corresponding to the plurality of data segments are generated based on the target data and at least one slice location identifier. These slice files conform to a second data protocol that is different from the first data protocol. The second data protocol may be a data transmission protocol to which the original media stream data needs to be converted, i.e., a target data transmission protocol for the protocol conversion. Exemplarily, the second data protocol may be a segment-based streaming protocol, such as the DASH protocol, the HLS protocol, and the like. For example, the original media stream data may be converted from the first data protocol to the DASH protocol or the HLS protocol.
[0086] In some embodiments, a slice position identifier added to the target media stream data can be identified. The slice position identifier at least indicates a slice position when converting from a first data protocol to a second data protocol. For example, after acquiring the target media stream data, the slice position identifier added to the target media stream data can be identified, for example, by identifying a data frame or data message included in the target media stream data that indicates a slice position.
[0087] Furthermore, the target media stream data can be subjected to protocol conversion based on the slice position indicated by the slice position identifier. In some embodiments, the slice position for slicing the target media stream data can be extracted from the slice position identifier. Alternatively, the position of the slice position identifier in the target media stream data can be determined as the slice position indicated by the slice position identifier. Then, the target media stream data can be sliced based on the determined slice position to obtain a data file that conforms to the second data protocol. For example, a plurality of data segments can be extracted from the target media stream data based on the slice position, and these data segments can be transpacked to obtain a plurality of slice files that conform to the second data protocol. Given that in the method according to each embodiment of the present disclosure, the slice position is a position for performing protocol conversion, and the slice position identifier can be used to implement protocol conversion, the slice position can also be referred to as a protocol conversion position, and the slice position identifier can also be referred to as a protocol conversion position.
[0088] In some embodiments, the slice file can be in TS format, FMP4 format, or Common Media Application Format (CMAF), etc. The slice file can include a video file and / or an audio file. Exemplarily, after determining the slice position, the target media stream data can be sliced, and a data file that conforms to the second data protocol is generated based on the media stream data segments obtained by slicing. For example, a video segment file that conforms to the second data protocol is generated based on the video frames (such as video key frames and video forward reference frames, etc.) contained in the media stream data segments obtained by slicing, and an audio segment file that conforms to the second data protocol is generated based on the audio frames contained in the media stream data segments obtained by slicing.
[0089] In some embodiments, the slice position identifier includes current description information for the second data protocol of multiple media segments. In this case, the current description information contained in the slice position identifier can also be obtained, and the description file for the second data protocol can be updated based on the current description information. Additionally, the target media stream data can be sliced based on the slice position and the updated description file. For example, the slice position identifier added at each slice position can include current description information for the second data protocol of one or more media segments before and after the slice position, such as the latest description information. Furthermore, the description file corresponding to the second data protocol when converting the target media stream data from the first data protocol to the second data protocol can be updated based on the current description information. Based on the slice position and the information carried by the updated description file (e.g., timestamp starting position, URL, etc.), the corresponding slice file can be restored from the data segment, thereby achieving protocol conversion of the target media stream data.
[0090] FIG5 shows a schematic diagram 500 of slicing processing for target media stream data according to some embodiments of the present disclosure. As shown in FIG5 , after receiving the target media stream data, the rules for inserting these Script frames are reversed. S(mpd) represents a Script frame containing MPD information, which indicates the boundary between data segments. That is, the media stream data segments between two adjacent Script frames can be transcapsulated into a slice file. Based on the media stream data segments, slice X (such as video slice x and audio slice x') can be generated according to the MPD information carried in the Script frame (for example, the slice file timestamp start position and URL, etc.), and slice X+1 (such as video slice x+1 and audio slice x+1') can be generated after a period of time, thereby achieving conversion from the FLV protocol to the DASH protocol. It should be understood that the Script frame containing MPD information shown in FIG5 is merely exemplary. In other embodiments, the Script frame may also include any other suitable information, such as m3u8 information, etc. The scope of the present disclosure is not limited in this respect.
[0091] Through the above description in combination with Figures 4 to 5, it can be seen that in the method for protocol conversion according to each embodiment of the present disclosure, a slice file is generated from the target data based on the slice position identifier carried in the target media stream data to achieve conversion from the first data protocol to the second data protocol. In this way, it is possible to achieve consistency in the protocol conversion results of different protocol conversion devices for the same media stream data, thereby ensuring the distributed consistency of the slices. Therefore, this makes it possible to support protocol conversion of the same media stream on different protocol conversion devices, and ensure that when switching protocol conversion devices, the player will not have playback errors due to inconsistent slicing results, thereby improving the efficiency and fault tolerance of protocol conversion and improving the playback effect of media stream data.
[0092] In some embodiments, the method for protocol conversion according to each embodiment of the present disclosure can be implemented at a third device within the streaming system 130 for the media stream. For example, the third device can be the same device as the second device described above. Alternatively, the third device can be different from the second device but communicatively coupled to the second device. The scope of the present disclosure is not limited in this respect.
[0093] The third device may receive a request for a media stream from a fourth device. Furthermore, the third device may determine, based on the request, whether the fourth device is also located within the streaming system 130. For example, the third device may determine whether the fourth device is located within the streaming system 130 based on parameters or attributes of the request. Alternatively, the request from the fourth device may itself carry information indicating whether the fourth device is located within the streaming system 130. Additionally, the third device may determine which protocol is used for data transmission between the third and fourth devices.
[0094] If it is determined that the fourth device is within the streaming system 130, the target data is sent to the fourth device. In this way, the target media stream data carrying the slice position information can be transparently transmitted in the streaming system. If it is determined that the fourth device is outside the streaming system 130 and the second data protocol is applied, the third device sends the slice file to the fourth device. Additionally, the third device can also send descriptive information for the second data protocol extracted from the slice position identifier to the fourth device, such as an MPD file for the DASH protocol or an m3u8 file for the HLS protocol. If it is determined that the fourth device is outside the streaming system 130 and the first data protocol is applied, the third device can obtain the original data for the media stream by deleting all slice position identifiers from the target data, and send the original data to the fourth device. In this way, on the one hand, errors can be avoided when devices outside the streaming system parse or play the target media stream data; on the other hand, when some services need to distribute media stream data that conforms to the first data protocol and media stream data that conforms to the second data protocol at the same time, the solution of the embodiment of the present disclosure only needs to return one path of the media stream data that conforms to the first data protocol to the source, and then generate media stream data that conforms to the second data protocol through protocol conversion, which can effectively save the cost of returning to the source bandwidth.
[0095] In some embodiments, when the streaming system 130 is implemented through a content distribution network, the third device may be an edge node within the content distribution network. For example, the third device may be an edge node within the content distribution network that serves the stream pulling device 140. In one example embodiment, the fourth device may be an edge node within the content distribution network. In another example embodiment, the fourth device may be a device outside the content distribution network, such as the stream pulling device 140 shown in FIG1 . In this case, compared with the existing solution in which protocol conversion is performed at the central source station and the edge node returns to the source to obtain data, the method according to the embodiment of the present disclosure performs protocol conversion at the edge node, and when the stream pulling device requests the slice file, the slice file already exists at the edge node and can respond directly. This greatly reduces the resource response time, thereby optimizing the playback response speed of the media streaming data.
[0096] Figure 6 shows a schematic diagram 600 of an example process of media stream distribution according to some embodiments of the present disclosure. In the example of Figure 6, the first data protocol is the FLV protocol, the second data protocol is the HLS protocol, and the protocol conversion device is source station 1. The live broadcast side uploads original media stream data that complies with the FLV protocol. The protocol conversion process is simulated at the first upstream source station (i.e., source station 1), but no actual file is generated. After the m3u8 file is updated, the m3u8 information (i.e., the current description information) is written into the FLV stream in the form of a Script frame as a slice position identifier, thereby obtaining an FLV stream containing a Script frame (i.e., target media stream data). The m3u8 information is transparently transmitted with the Script frame in the CDN system. The stream pulling node 1 obtains the FLV stream containing the Script frame from the source station 1 by returning to the source. The stream pulling node 1 restores the HLS file, such as the m3u8 file and the slice file, based on the m3u8 information in the Script frame. Further, the stream pulling node 1 sends the HLS file to the client of the viewer 1. Similarly, the stream pulling node 2 obtains the FLV stream containing the Script frame from the source station 1 by returning to the source. Pull Node 2 restores the HLS file based on the m3u8 information in the Script frame, such as the m3u8 file and the slice file. Furthermore, Pull Node 2 sends the HLS file to the client of Viewer 2. Here, Pull Node 1 and Pull Node 2 can generate consistent HLS files, thus achieving distributed consistency of slices.
[0097] In addition, when distributing FLV streams externally, the Script frame carrying the m3u8 information in the FLV stream can be deleted. For example, during internal cascading, the request parameters can be used to identify whether the other end is another streaming media server within the system or an external device. If it is an external device, the customized tag information (i.e., the Script frame) is deleted. In the example of Figure 6, after deleting the Script frame, the pull stream node 1 and / or the pull stream node 2 can send an FLV stream that does not contain a Script frame to the external device, which is not shown in Figure 6. In addition, the source station 1 can also send an FLV stream that does not contain a Script frame to other CDNs.
[0098] Figure 7 shows a schematic diagram 700 of another example process of media stream distribution according to some embodiments of the present disclosure. In the example of Figure 7, the first data protocol is the FLV protocol, the second data protocol is the DASH protocol, and the protocol conversion device is a streaming node. The live broadcast side uploads the original media stream data that complies with the FLV protocol. The first upstream streaming media server (i.e., the streaming node) simulates the protocol conversion process, but does not generate an actual file. After the MPD file is updated, the MPD information (i.e., the current description information) is written into the FLV stream in the form of a Script frame as a slice position identifier, thereby obtaining an FLV stream containing a Script frame (i.e., the target media stream data). The MPD information is transparently transmitted with the Script frame in the CDN system. The stream pulling node obtains the FLV stream containing the Script frame from the data source station 2 by returning to the source. The DASH file, such as the MPD file and the slice file, is restored according to the MPD information in the Script frame. Furthermore, the stream pulling node sends the DASH file to the viewer's client.
[0099] Furthermore, when distributing FLV streams externally, the Script frame carrying the MPD information can be deleted from the FLV stream. For example, during internal cascading, request parameters can be used to identify whether the peer is another streaming server within the system or an external device. If it is an external device, the customized tag information (i.e., the Script frame) can be deleted. In the example of Figure 7, Origin Site 1 can send an FLV stream without the Script frame to other CDNs.
[0100] In the example shown in FIG7 , the protocol conversion device (i.e., the device that performs protocol conversion in the CDN) may include a protocol conversion source station or edge node in the CDN. When central slicing is used for protocol conversion, the protocol conversion device may include a source station in the CDN, such as slice source station 1, slice source station 2, and so on. When edge slicing is used for protocol conversion, the protocol conversion device may include an edge node in the CDN, such as a streaming node, and so on.
[0101] In the examples of Figures 6 and 7, the cascade protocol (i.e., the transmission protocol between the stream pulling node and the upstream node or source station) is the FLV protocol. In this case, the slice position identifier can be implemented in the form of a custom Script frame, and the playlist information (e.g., MPD or m3u8 file) is encapsulated in the Script frame. In other embodiments, the RTMP protocol is adopted as the cascade protocol. In this case, the MPD information can be transmitted using the AMF0 package via Metadata Message. At this time, the protocol conversion identifier can be added to the original media stream data in the form of Metadata Message. It should be understood that the solutions according to the various embodiments of the present disclosure also support any other suitable cascade protocols, and the scope of the present disclosure is not limited in this respect.
[0102] As can be seen, by inserting slice-related information (such as protocol conversion identifiers) into the media stream data and passing the slice-related information to each streaming server in the system as the media stream data is transmitted, these streaming servers can generate consistent slice files based on the media stream data and the slice information it carries, thereby achieving the purpose of distributed consistent slicing. In this way, without relying on external services and only utilizing the stream transmission between streaming servers, distributed consistent slicing is efficiently achieved, thereby solving the problems of poor flexibility and availability of existing solutions that fix slicing tasks on a single slicing device.
[0103] Embodiments of the present disclosure also provide corresponding apparatuses and devices for implementing the above-described methods or processes. FIG8 shows a block diagram of an example apparatus 800 for data processing according to some embodiments of the present disclosure. The apparatus 800 can, for example, be used to implement the method for data processing according to some embodiments of the present disclosure. In some embodiments, the apparatus 800 can be implemented at the streaming device 120 or the transmission device 132 as shown in FIG1 .
[0104] As shown in FIG8 , apparatus 800 may include a receiving module 802, a slicing module 804, and an adding module 806. Receiving module 802 is configured to receive original data for a media stream. Slicing module 804 is configured to slice the original data into multiple data segments. Adding module 806 is configured to add at least one slice position identifier for the multiple data segments to the original data to obtain target data for the media stream. The target data conforms to the same first data protocol as the original data.
[0105] In some embodiments, each slice position identifier of the at least one slice position identifier corresponds to two adjacent data segments of the plurality of data segments and indicates a position of a boundary between the two adjacent data segments.
[0106] In some embodiments, each slice position identifier of the at least one slice position identifier is inserted between corresponding two adjacent data segments.
[0107] In some embodiments, the at least one slice location identifier includes description information of a plurality of media segments of the media stream corresponding to the plurality of data segments, and the description information is for a second data protocol different from the first data protocol.
[0108] In some embodiments, at least one slice location identifier also includes an identifier for a method.
[0109] In some embodiments, if it is determined that the second data protocol includes the Hypertext Transfer Protocol (HTTP) Live Streaming (HLS) protocol, the description information includes index information of the multiple media segments for the HLS protocol, or if it is determined that the second data protocol includes the Dynamic Adaptive Streaming over HTTP (DASH) protocol, the description information includes the Media Presentation Description (MPD) of the multiple media segments for the DASH protocol.
[0110] In some embodiments, the description information includes description information corresponding to each track of the at least one track of the media stream, and the target data includes a mapping relationship between each track and the corresponding description information.
[0111] In some embodiments, if it is determined that the second data protocol includes the HLS protocol, the slice location identifier includes a first indication for indicating a mapping relationship between a track index and index information for the HLS protocol of multiple media segments, and the target data includes a second indication for indicating an HLS multi-variant playlist, and a third indication for indicating a mapping between the track index and the HLS variant stream, or if it is determined that the second data protocol includes the DASH protocol, the slice location identifier includes an MPD for the DASH protocol of multiple media segments, and the target data includes a fourth indication for indicating a mapping between the track index and the DASH representation.
[0112] In some embodiments, information included in the at least one slice location identifier is encapsulated based on a target data structure, and the target data structure depends on the first data protocol.
[0113] In some embodiments, if it is determined that the first data protocol includes the Flash Video FLV protocol, the target data structure includes a script data SCRIPTDATA structure, or if it is determined that the first data protocol includes the Real Time Messaging Protocol RTMP, the target data structure corresponds to a metadata message.
[0114] In some embodiments, at least one slice location identifier is unencrypted.
[0115] In some embodiments, the original data is sliced based on at least one of: a predetermined duration of the data segment, or a location of a random access point in the media stream.
[0116] In some embodiments, apparatus 800 is implemented at a first device within a streaming system for a media stream. Apparatus 800 further includes a request receiving module and a sending module. The request receiving module is configured to receive a request for a media stream from a second device. The sending module is configured to: if the second device is determined to be within the streaming system, send target data for the media stream to the second device; or if the second device is determined to be outside the streaming system, send original data for the media stream to the second device.
[0117] In some embodiments, the streaming system is implemented via a content delivery network, and the first device comprises an origin station or an edge node within the content delivery network, and / or the second device comprises an origin station or an edge node within the content delivery network, or a device outside the content delivery network.
[0118] Figure 9 shows a block diagram of an example apparatus 900 for protocol conversion according to some embodiments of the present disclosure. The apparatus 900 can be used, for example, to implement a method for protocol conversion according to some embodiments of the present disclosure. In some embodiments, the apparatus 900 can be implemented at the transmission device 132 as shown in Figure 1.
[0119] As shown in FIG9 , apparatus 900 may include a receiving module 902 and a generating module 904. Receiving module 902 is configured to receive target data for a media stream. The target data conforms to a first data protocol and includes at least one slice location identifier for a plurality of data segments in the target data. Generating module 904 is configured to generate, based on the target data and the at least one slice location identifier, a plurality of slice files corresponding to the plurality of data segments. The plurality of slice files conforms to a second data protocol different from the first data protocol.
[0120] In some embodiments, each slice position identifier of the at least one slice position identifier corresponds to two adjacent data segments of the plurality of data segments and indicates a position of a boundary between the two adjacent data segments.
[0121] In some embodiments, each slice position identifier of the at least one slice position identifier is inserted between corresponding two adjacent data segments.
[0122] In some embodiments, the at least one slice location identifier includes description information of a plurality of media segments of the media stream corresponding to the plurality of data segments, and the description information is for a second data protocol.
[0123] In some embodiments, at least one slice location identifier also includes an identifier for a method.
[0124] In some embodiments, if it is determined that the second data protocol includes the HLS protocol, the description information includes index information for the HLS protocol of multiple media segments, or if it is determined that the second data protocol includes the DASH protocol, the description information includes MPD for the DASH protocol of multiple media segments.
[0125] In some embodiments, the description information includes description information corresponding to each track of the at least one track of the media stream, and the target data includes a mapping relationship between each track and the corresponding description information.
[0126] In some embodiments, if it is determined that the second data protocol includes the HLS protocol, the slice location identifier includes a first indication for indicating a mapping relationship between a track index and index information for the HLS protocol of multiple media segments, and the target data includes a second indication for indicating an HLS multi-variant playlist, and a third indication for indicating a mapping between the track index and the HLS variant stream, or if it is determined that the second data protocol includes the DASH protocol, the slice location identifier includes an MPD for the DASH protocol of multiple media segments, and the target data includes a fourth indication for indicating a mapping between the track index and the DASH representation.
[0127] In some embodiments, information included in the at least one slice location identifier is encapsulated based on a target data structure, and the target data structure depends on the first data protocol.
[0128] In some embodiments, if it is determined that the first data protocol includes the FLV protocol, the target data structure includes a SCRIPTDATA structure, or if it is determined that the first data protocol includes RTMP, the target data structure corresponds to a metadata message.
[0129] In some embodiments, at least one slice location identifier is unencrypted.
[0130] In some embodiments, the generation module 904 is further configured to: extract multiple data segments from the target data based on at least one slice location identifier; and transpack the multiple data segments to obtain multiple slice files.
[0131] In some embodiments, the apparatus 900 is implemented at a third device within a streaming system for a media stream. The apparatus 900 further includes a request receiving module and a sending module. The request receiving module is configured to receive a request for a media stream from a fourth device. The sending module is configured to: if it is determined that the fourth device is outside the streaming system and the second data protocol is applied, send a slice file to the fourth device, or if it is determined that the fourth device is outside the streaming system and the first data protocol is applied, obtain original data for the media stream by deleting all slice position identifiers from the target data, and send the original data to the fourth device, or if it is determined that the fourth device is within the streaming system, send target data to the fourth device.
[0132] In some embodiments, the streaming system is implemented via a content delivery network, and the third device comprises an edge node within the content delivery network, and / or the fourth device comprises an edge node within the content delivery network, or a device outside the content delivery network.
[0133] The modules and / or units included in the apparatus 800 and the apparatus 900 can be implemented in various ways, including software, hardware, firmware, or any combination thereof. In some embodiments, one or more units can be implemented using software and / or firmware, such as machine executable instructions stored on a storage medium. In addition to or as an alternative to machine executable instructions, some or all of the units in the apparatus 800 and the apparatus 900 can be implemented at least in part by one or more hardware logic components. As an example and not limitation, exemplary types of hardware logic components that can be used include field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chip (SOCs), complex programmable logic devices (CPLDs), and the like.
[0134] The modules and / or units shown in Figures 8 and 9 may be partially or entirely implemented as hardware modules, software modules, firmware modules, or any combination thereof. In particular, in some embodiments, the processes, methods, or procedures described above may be implemented by hardware in a storage system, a host corresponding to the storage system, or other computing devices independent of the storage system.
[0135] FIG10 illustrates a block diagram of a device 1000 in which one or more embodiments of the present disclosure may be implemented. It should be understood that the electronic device 1000 shown in FIG10 is merely exemplary and should not be construed as limiting the functionality and scope of the embodiments described herein. The electronic device 1000 shown in FIG10 may be used to implement the streaming device 120 and transmission device 132 shown in FIG1 and / or the methods described above.
[0136] As shown in FIG10 , electronic device 1000 is in the form of a general electronic device. Components of electronic device 1000 may include, but are not limited to, one or more processing units or processors 1010, memory 1020, storage device 1030, one or more communication units 1040, one or more input devices 1050, and one or more output devices 1060. Processor 1010 may be a real or virtual processor and is capable of performing various processes according to a program stored in memory 1020. In a multi-processor system, multiple processing units execute computer-executable instructions in parallel to increase the parallel processing capabilities of electronic device 1000.
[0137] The electronic device 1000 typically includes a plurality of computer storage media. Such media can be any available media accessible to the electronic device 1000, including but not limited to volatile and non-volatile media, removable and non-removable media. The memory 1020 can be a volatile memory (e.g., a register, a cache, a random access memory (RAM)), a non-volatile memory (e.g., a read-only memory (ROM), an electrically erasable programmable read-only memory (EEPROM), flash memory), or some combination thereof. The storage device 1030 can be a removable or non-removable medium and can include a machine-readable medium, such as a flash drive, a disk, or any other medium that can be used to store information and / or data (e.g., training data for training) and can be accessed within the electronic device 1000.
[0138] The electronic device 1000 may further include additional removable / non-removable, volatile / non-volatile storage media. Although not shown in FIG. 10 , a disk drive for reading from or writing to a removable, non-volatile disk (e.g., a “floppy disk”) and an optical drive for reading from or writing to a removable, non-volatile optical disk may be provided. In these cases, each drive may be connected to a bus (not shown) by one or more data media interfaces. The memory 1020 may include a computer program product 1025 having one or more program modules configured to perform various methods or actions of various implementations of the present disclosure.
[0139] The communication unit 1040 enables communication with other electronic devices via a communication medium. Additionally, the functions of the components of the electronic device 1000 can be implemented as a single computing cluster or multiple computing machines that can communicate via a communication connection. Thus, the electronic device 1000 can operate in a networked environment using logical connections to one or more other servers, network personal computers (PCs), or other network nodes.
[0140] Input device 1050 may be one or more input devices, such as a mouse, keyboard, or trackball. Output device 1060 may be one or more output devices, such as a display, a speaker, or a printer. Electronic device 1000 may also communicate with one or more external devices (not shown) via communication unit 1040 as needed, such as storage devices, display devices, or the like, with one or more devices that allow a user to interact with electronic device 1000, or with any device that allows electronic device 1000 to communicate with one or more other electronic devices (e.g., a network card, a modem, etc.). Such communication may be performed via an input / output (I / O) interface (not shown).
[0141] According to an exemplary implementation of the present disclosure, a computer-readable storage medium is provided, on which computer-executable instructions are stored, wherein the computer-executable instructions are executed by a processor to implement the method described above. According to an exemplary implementation of the present disclosure, a computer program product is also provided, which is tangibly stored on a non-transitory computer-readable medium and includes computer-executable instructions, and the computer-executable instructions are executed by a processor to implement the method described above. According to an exemplary implementation of the present disclosure, a computer program product is provided, on which a computer program is stored, and when the program is executed by a processor, the method described above is implemented.
[0142] Various aspects of the present disclosure are described herein with reference to flowcharts and / or block diagrams of methods, apparatuses, devices, and computer program products implemented according to the present disclosure. It should be understood that each block of the flowcharts and / or block diagrams, and combinations of blocks in the flowcharts and / or block diagrams, can be implemented by computer-readable program instructions.
[0143] These computer-readable program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, thereby producing a machine such that when these instructions are executed by the processor of the computer or other programmable data processing device, a device is generated that implements the functions / actions specified in one or more blocks in the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium, where these instructions cause the computer, programmable data processing device, and / or other device to operate in a specific manner. Thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing various aspects of the functions / actions specified in one or more blocks in the flowchart and / or block diagram.
[0144] Computer-readable program instructions can be loaded onto a computer, other programmable data processing apparatus, or other device so that a series of operational steps are performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions executed on the computer, other programmable data processing apparatus, or other device to implement the functions / actions specified in one or more boxes in the flowchart and / or block diagram.
[0145] The flow charts and block diagrams in the accompanying drawings show the possible architecture, functions and operations of the systems, methods and computer program products according to multiple implementations of the present disclosure. In this regard, each box in the flow chart or block diagram can represent a part for a module, program segment or instruction, and a part for a module, program segment or instruction comprises one or more executable instructions for realizing the logical function of the specification. In some alternative implementations, the functions marked in the box can also occur in a sequence different from that marked in the accompanying drawings. For example, two continuous boxes can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flow chart, and the combination of the boxes in the block diagram and / or flow chart can be realized by a special hardware-based system that performs the function or action of the specification, or can be realized by a combination of special hardware and computer instructions.
[0146] While various implementations of the present disclosure have been described above, the above descriptions are intended to be illustrative, non-exhaustive, and non-limiting to the disclosed implementations. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described implementations. The terminology used herein is selected to best explain the principles of the implementations, their practical applications, or improvements to existing technologies, or to enable others skilled in the art to understand the various implementations disclosed herein.
Claims
1. A method for data processing, comprising: Receiving raw data for a media stream; Slicing the raw data into a plurality of data segments; And Adding at least one slice position identifier for the plurality of data segments to the raw data to obtain target data for the media stream, the target data conforming to the same first data protocol as the raw data.
2. The method according to claim 1, wherein each slice position identifier in the at least one slice position identifier corresponds to two adjacent data segments among the plurality of data segments and indicates the position of the boundary between the two adjacent data segments.
3. The method according to claim 2, wherein each slice position identifier in the at least one slice position identifier is inserted between the corresponding two adjacent data segments.
4. The method according to any one of claims 1 to 3, wherein the at least one slice position identifier includes description information of a plurality of media segments corresponding to the plurality of data segments of the media stream, and the description information is for a second data protocol different from the first data protocol.
5. The method according to claim 4, wherein the at least one slice position identifier further includes an identifier for the method.
6. The method according to any one of claims 4 to 5, wherein if it is determined that the second data protocol includes the Hypertext Transfer Protocol HTTP Live Streaming (HLS) protocol, the description information includes index information of the plurality of media segments for the HLS protocol, or if it is determined that the second data protocol includes the HTTP-based Dynamic Adaptive Streaming over HTTP (DASH) protocol, the description information includes the Media Presentation Description (MPD) of the plurality of media segments for the DASH protocol.
7. The method according to any one of claims 4 to 6, wherein the description information includes description information corresponding to each track in at least one track of the media stream, and the target data includes a mapping relationship between each track and the corresponding description information.
8. The method according to claim 7, wherein if it is determined that the second data protocol includes the HLS protocol, the slice position identifier includes a first indication for indicating a mapping relationship between a track index and index information of the plurality of media segments for the HLS protocol, and the target data includes a second indication for indicating an HLS variant playlist and a third indication for indicating a mapping between a track index and an HLS variant stream, or if it is determined that the second data protocol includes the DASH protocol, the slice position identifier includes the MPD of the plurality of media segments for the DASH protocol, and the target data includes a fourth indication for indicating a mapping between a track index and a DASH representation.
9. The method according to any one of claims 4 to 8, wherein the information included in the at least one slice position identifier is encapsulated based on a target data structure, and the target data structure depends on the first data protocol.
10. The method according to claim 9, wherein if it is determined that the first data protocol includes the Flash video FLV protocol, the target data structure includes the script data SCRIPTDATA structure, or if it is determined that the first data protocol includes the Real-Time Messaging Protocol RTMP, the target data structure corresponds to a metadata message.
11. The method according to any one of claims 1 to 10, wherein the at least one slice position identifier is unencrypted.
12. The method according to any one of claims 1 to 11, wherein the original data is sliced based on at least one of the following: a predetermined duration of a data segment, or the position of a random access point in the media stream.
13. The method according to any one of claims 1 to 12, wherein the method is implemented at a first device within a streaming system for the media stream, and the method further includes: receiving a request for the media stream from a second device; and if it is determined that the second device is within the streaming system, sending the target data for the media stream to the second device, or if it is determined that the second device is outside the streaming system, sending the original data for the media stream to the second device.
14. The method according to claim 13, wherein the streaming system is implemented through a content delivery network, and the first device includes an origin station or an edge node within the content delivery network, and / or the second device includes an origin station or an edge node within the content delivery network, or a device outside the content delivery network.
15. A method for protocol conversion, comprising: receiving target data for a media stream, the target data conforming to a first data protocol and including at least one slice position identifier for a plurality of data segments in the target data; and generating, based on the target data and the at least one slice position identifier, a plurality of slice files corresponding to the plurality of data segments, the plurality of slice files conforming to a second data protocol different from the first data protocol.
16. The method according to claim 15, wherein each slice position identifier in the at least one slice position identifier corresponds to two adjacent data segments among the plurality of data segments, and indicates the position of the boundary between the two adjacent data segments.
17. The method according to claim 16, wherein each slice position identifier in the at least one slice position identifier is inserted between the corresponding two adjacent data segments.
18. The method according to any one of claims 15 to 17, wherein the at least one slice position identifier includes description information of a plurality of media segments of the media stream corresponding to the plurality of data segments, and the description information is for the second data protocol.
19. The method according to claim 18, wherein the at least one slice position identifier further includes an identifier for the method.
20. The method according to any one of claims 18 to 19, wherein if it is determined that the second data protocol includes the HLS protocol, the description information includes index information for the HLS protocol of the plurality of media segments, or if it is determined that the second data protocol includes the DASH protocol, the description information includes the MPD for the DASH protocol of the plurality of media segments.
21. The method according to any one of claims 18 to 20, wherein the description information includes description information corresponding to each track in at least one track of the media stream, and the target data includes a mapping relationship between each track and the corresponding description information.
22. The method according to claim 21, wherein if it is determined that the second data protocol includes the HLS protocol, the slice position identifier includes a first indication for indicating a mapping relationship between a track index and the index information for the HLS protocol of the plurality of media segments, and the target data includes a second indication for indicating an HLS multi-variant playlist and a third indication for indicating a mapping between the track index and the HLS variant stream, or if it is determined that the second data protocol includes the DASH protocol, the slice position identifier includes the MPD for the DASH protocol of the plurality of media segments, and the target data includes a fourth indication for indicating a mapping between the track index and the DASH representation.
23. The method according to any one of claims 18 to 22, wherein the information included in the at least one slice position identifier is encapsulated based on a target data structure, and the target data structure depends on the first data protocol.
24. The method according to claim 23, wherein if it is determined that the first data protocol includes the FLV protocol, the target data structure includes a SCRIPTDATA structure, or if it is determined that the first data protocol includes RTMP, the target data structure corresponds to a metadata message.
25. The method according to any one of claims 15 to 24, wherein the at least one slice position identifier is unencrypted.
26. The method according to any one of claims 15 to 25, wherein generating the plurality of slice files includes: extracting the plurality of data segments from the target data based on the at least one slice position identifier; and re-encapsulating the plurality of data segments to obtain the plurality of slice files.
27. The method according to any one of claims 15 to 26, wherein the method is implemented at a third device within a streaming system for the media stream, and the method further includes: receiving a request for the media stream from a fourth device; and if it is determined that the fourth device is outside the streaming system and the second data protocol is applied, sending the slice file to the fourth device, or If it is determined that the fourth device is outside the streaming system and the first data protocol is applied, the original data for the media stream is obtained by deleting all the slice position identifiers from the target data, and the original data is sent to the fourth device, or If it is determined that the fourth device is inside the streaming system, the target data is sent to the fourth device.
28. The method according to claim 27, wherein the streaming system is implemented through a content delivery network, and the third device includes an edge node within the content delivery network, and / or the fourth device includes an edge node within the content delivery network or a device outside the content delivery network.
29. A data processing apparatus, comprising: a receiving module configured to receive original data for a media stream; a slicing module configured to slice the original data into a plurality of data segments; and an adding module configured to add at least one slice position identifier for the plurality of data segments to the original data to obtain target data for the media stream, the target data conforming to the same first data protocol as the original data.
30. A protocol conversion apparatus, comprising: a receiving module configured to receive target data for a media stream, the target data conforming to a first data protocol and including at least one slice position identifier for a plurality of data segments in the target data; and a generating module configured to generate, based on the target data and the at least one slice position identifier, a plurality of slice files corresponding to the plurality of data segments, the plurality of slice files conforming to a second data protocol different from the first data protocol.
31. An electronic device, comprising: at least one processor; and at least one memory coupled to the at least one processor and storing instructions for execution by the at least one processor, the instructions, when executed by the at least one processor, causing the electronic device to perform the method according to any one of claims 1 to 14 or the method according to any one of claims 15 to 28.
32. A computer-readable storage medium having instructions stored thereon, the instructions, when executed by a processor, causing the processor to implement the method according to any one of claims 1 to 14 or the method according to any one of claims 15 to 28.
33. A computer program product tangibly stored in a computer storage medium and including computer-executable instructions, the computer-executable instructions, when executed by a device, causing the device to perform the method according to any one of claims 1 to 14 or the method according to any one of claims 15 to 28.
Citation Information
Patent Citations
Data slicing method and system
CN108600859A
Conversion method, device, system and computer readable medium for streaming media protocol
CN109495505A
Transfer stream file generation method, device and equipment and storage medium
CN112367527A
Data slicing method and device, electronic equipment and storage medium
CN115988286A
Media stream slicing method, device, system and equipment and storage medium
CN117201894A