Real-time interactive streaming media system
By mapping WebRTC session descriptions to streaming media description files and adding adaptive sets and user tag structures, the high latency and poor interoperability issues in multi-user scenarios in existing technologies are resolved, resulting in better live streaming performance and complex interactive scenarios, and supporting the transmission of various interactive information in the metaverse.
Patent Information
- Application Number
- CN202310411835.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-04-17
- Publication Date
- 2025-11-25
- Estimated Expiration
- 2043-04-17
AI Technical Summary
Existing real-time interactive streaming methods cannot meet the needs of the metaverse, especially in scenarios with multiple participants, they suffer from high latency, poor interoperability, inconsistent signaling, inability to implement subtitles, time metadata and advertising management, insufficient digital rights protection, and limitations of advanced encoding formats.
By mapping WebRTC session descriptions to streaming media description files, adding adaptive collections and user tag structures, it supports real-time interaction with multiple participants, transmits additional scalable metadata, including motion gesture information, digital assets, and virtual environment information, and leverages advanced features in streaming media description files such as digital rights protection and live ad insertion to achieve better live streaming performance and complex interactive scenarios.
It significantly reduces the latency of streaming interaction, improves the complexity of interactive scenarios and live streaming performance, supports real-time interaction with multiple participants, and implements digital rights protection and advanced audio-visual encoder functions.
Smart Images

Figure CN116366614B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] Embodiments of the present disclosure relate to the field of computer technology, and in particular to a real-time interactive streaming media system. BACKGROUND
[0002] The metaverse is a virtual world built by humans using digital technology, which is mapped from or beyond the real world, can interact with the real world, and has a new social system of digital living space. The metaverse itself is not a new technology, but an integration of a large number of existing technologies, including 5G, cloud computing, artificial intelligence, virtual reality, blockchain, digital currency, the Internet of Things, human-computer interaction, etc.
[0003] The existing real-time interactive streaming media method cannot meet the requirements of the metaverse. For example, the RTSP / RTP (Real Time Streaming Protocol / Real Time Transport Protocol) real-time media stream uses a push stream to solve part of the real-time streaming problem and can achieve low latency. However, because it is based on point-to-point transmission, it cannot be used for multi-person participation scenarios (i.e., watching streams and sending streams), so its use is limited.
[0004] The multi-person video call of webRTC (Web Real-Time Communications) and the conference technology (MESH) solve the problem of how to implement video conferencing and team online workflow based on browsers, and can achieve extremely low latency. However, there is no method and protocol for signaling, resulting in no interoperability between products. Based on the SDP (Session Description Protocol) description configuration, the protocol is relatively simple, and a large amount of configuration information is missing for advanced real-time stream pushing. Although custom extensions can be used, using custom extensions further leads to fragmentation of product implementation. An additional signaling server is required, and the incompatibility of the signaling server limits its use. There is no unified session establishment mechanism. There is no advantage of subtitles, time metadata, advertisement management, digital copyright protection, and high-level encoding format.
[0005] Improvements based on the SFU (select forward unit) of webRTC can partially solve the problem of multi-person participation by selecting the media stream that needs to be forwarded to generate a forwarding matrix to realize multi-person conferences. However, when the number of participants is large, the forwarding matrix will be very complex, and when the number of participants reaches a certain level, the overly complex forwarding matrix cannot be designed and implemented, which limits the number of participants.
[0006] Based on the webRTC MCU (multipoint control unit) improvement, the basic webRTC implementation of multi-person online conference is extended, although the problem of multi-person participation is solved, but due to the need for transcoding and mixing on the server side, although it can realize large-scale simultaneous online number in technology, the performance demand of the server side will increase a lot. SUMMARY
[0007] Embodiments of the present disclosure propose a real-time interactive streaming media system.
[0008] Embodiments of the present disclosure provide a real-time interactive streaming media system, comprising: an organizer terminal configured to send a registration session request including participant information to a server, establish a direct connection with a participant terminal and interact data; the server is configured to respond to receiving the registration session request, map the session description into a streaming media description file, store the participant information into the streaming media description file, respond to receiving a request for obtaining a streaming media description file sent by a participant terminal, and send the streaming media description file to the participant terminal; the participant terminal is configured to obtain the streaming media description file from the server, parse the connection information from the streaming media description file, and establish a direct connection with the organizer terminal according to the connection information and interact data.
[0009] In some embodiments, the organizer terminal is further configured to send at least one of the following event information to the server: control information, meta-universe information, media information; the server is further configured to update the streaming media description file according to the event information; the participant terminal is further configured to obtain the updated streaming media description file from the server at a regular time, and process events according to the updated streaming media description file.
[0010] In some embodiments, the mapping of the session description into the streaming media description file comprises: adding a new adaptive set on the basis of an existing adaptive set configuration of the streaming media description file, wherein the new adaptive set includes the description type of the attribute of the session description and the related parameter agreement under the description type; adding a user tag structure for describing participant information and participant device information in the new adaptive set.
[0011] In some embodiments, the server is further configured to add a pre-selection option in the streaming media description file, for recommending adaptive sets to participants and allowing participants to set adaptive sets to be accepted by themselves.
[0012] In some embodiments, the server is further configured to group all participants according to the user tag structure; and recommend adaptive sets suitable for each group of participants to the participants in the group.
[0013] In some embodiments, the server is further configured to add at least one of the following to the adaptive set on the basis of the existing adaptive set configuration of the streaming media description file: external reference media information, action gesture information, digital asset information, virtual environment information.
[0014] In some embodiments, the server is further configured to send the streaming media description file to the participant terminal through an HTTP channel; in response to receiving the media data, map the media data into segment data and send the segment data to the participant terminal through an RTC data channel.
[0015] In some embodiments, the participant terminal comprises: a streaming media description file parser configured to parse the streaming media description file to obtain an adaptive set and a pre-selection option; an adaptive bit rate selection controller configured to perform adaptive bit rate selection according to the pre-selection option; a terminal reporter configured to report error terminal information and session information; a segment parser configured to parse the segment data to find additional information at the encapsulation level; an event processor configured to insert events in the streaming media description file and in the media data; and a download controller configured to perform media stream download according to the pre-selection option.
[0016] In some embodiments, at least one participant terminal is further configured to send at least one piece of interaction information to the server; and the server is further configured to synthesize the at least one piece of interaction information to update the streaming media description file, which is acquired by the organizer terminal and all participant terminals at regular intervals.
[0017] In some embodiments, the server is further configured to write the participant's copyright information in the streaming media description file; and the participant terminal is further configured to decrypt the received media data according to the version information and render the decrypted media data.
[0018] Embodiments of the present disclosure provide a real-time interactive streaming media system, which maps a webRTC session stream to a streaming media description file protocol and packages and transmits additional extensible metadata information. The delay of streaming media-based streaming media interaction can be significantly reduced, the advantages of webRTC ultra-low delay are utilized, and advanced streaming media functions in the streaming media description file are used, such as digital copyright protection, live advertisement insertion, additional information transmission, advanced audio and video encoder functions, extensible server events, error terminal information collection and reporting, etc. Better live performance and more complex interactive scenarios can be achieved.
[0019] It is to be understood that the details set forth in this section are intended to be illustrative only and are not intended to limit the scope of the disclosure. Other features, objects, and advantages of the disclosure will be apparent to one of ordinary skill in the art from the following detailed description of non-limiting embodiments. BRIEF DESCRIPTION OF DRAWINGS
[0020] Other features, objects, and advantages of the disclosure will become more apparent from the following detailed description of non-limiting embodiments when read in connection with the following drawings:
[0021] Figure 1 is an exemplary system architecture diagram in which one embodiment of the present disclosure can be applied;
[0022] Figure 2 is a schematic diagram of information stored in a streaming media description file;
[0023] Figure 3 is a schematic diagram of the interaction between the metaverse additional information and the webRTC information and the existing functions in the streaming media description file;
[0024] Figure 4 is a schematic diagram of data transmission between the server and the terminal;
[0025] Figure 5 is a schematic diagram of the interaction process between the various modules of the real-time interactive streaming media system;
[0026] Figure 6 is a schematic diagram of the mixing process of the streaming media description file;
[0027] Figure 7 is a schematic diagram of the streaming media description file after the multi-party data mixing and injection;
[0028] Figure 8 is a schematic diagram of the streaming media description file after the insertion of advertisements in time sequence;
[0029] Figure 9 is a schematic diagram of the streaming media description file in which the metaverse additional information is stored;
[0030] Figure 10 is a schematic diagram of filtering the media stream that meets the user's own authority according to the pre-selected options set by the user;
[0031] Figure 11 is a schematic diagram of mixing the streaming media description file and filtering the mixed information. DETAILED DESCRIPTION
[0032] The present disclosure will be further described in details below with reference to the accompanying drawings and embodiments. It can be understood that the specific embodiments described herein are only used to explain the related application, but not to limit the application. In addition, it should be noted that, for the convenience of description, only the parts related to the application are shown in the drawings.
[0033] It should be noted that the embodiments in the present disclosure and the features in the embodiments can be combined with each other without conflict. The present disclosure will be described in detail below with reference to the accompanying drawings and embodiments.
[0034] Figure 1 An exemplary system architecture of an embodiment of the real-time interactive streaming media system to which the present disclosure can be applied is shown.
[0035] As shown in Figure 1 , the system architecture can include an organizer terminal, a server and a participant terminal. Among them, the number of participant terminals can be multiple.
[0036] The organizer terminal is configured to send a registration session request including participant information to the server, establish direct connection with the participant terminal and interact data;
[0037] The server is configured to respond to the registration session request, map the session description into a streaming media description file, store the participant information into the streaming media description file, and send the streaming media description file to the participant terminal in response to the request of the participant terminal for obtaining the streaming media description file;
[0038] The participant terminal is configured to obtain the streaming media description file from the server, parse the connection information from the streaming media description file, and establish direct connection with the organizer terminal according to the connection information and interact data.
[0039] WebRTC uses SDP as a one-to-one session description protocol, and WebRTC has made special definitions in individual fields within the standard specification of SDP, which need to be mapped to the related concepts of the media presentation description file (MPD, media presentation description). In MPEG-DASH, a set of media content with different encoding parameters and the corresponding description are defined as a media presentation. The media content here is composed of single or multiple media periods that are continuous in time, and the content of these media periods may be completely independent of each other, such as the inserted advertising content in the main film. Each media period contains one or more media content components, such as audio components in different languages, video components providing different perspectives for the same program, and subtitle components in different languages. Each media content component has a corresponding media content component type identifier, such as audio or video. Each media content component may have multiple encoded versions, referred to as media streams. Each media stream inherits the properties of media content, media period and media content component, and also has its own encoding parameter properties, such as bit rate, resolution, encoder type, etc. The MPD file is a metadata file used to describe the above-mentioned abstract concepts and their relationships.
[0040] As shown in the Exit section, the original MPD file contains the content. The webRTC mapping section is the newly added session description mapping to the media presentation description file content. The Metaverse info is the newly added meta-universe information mapping to the media presentation description file content. Figure 2
[0041] In some optional implementations of the embodiment, the mapping of the session description to the media presentation description file includes: adding a new adaptive set on the basis of the existing adaptive set configuration of the media presentation description file, wherein the new adaptive set includes a description type of the properties of the session description and related parameter conventions under the description type; and adding a user tag structure for describing participant information and participant device information in the new adaptive set.
[0042] As shown in the Exit section, the original MPD file contains the content. The webRTC mapping section is the newly added session description mapping to the media presentation description file content. The Metaverse info is the newly added meta-universe information mapping to the media presentation description file content. Figure 2 The mapping relationship is as follows: in the top description of the streaming media description file, a newly added profile identifier is added to indicate that the streaming media description file is with webRTC. At the same time, a webRTC session of a participant is mapped to an adaptation set in a period. Specifically, on the basis of the existing adaptation set configuration, the description type of an EssentialProperty (session description attribute) is newly added, and relevant parameter conventions under the description type are added to realize the description conversion from a webRTC session (SDP) to a streaming media description file. Multiple participants are mapped to multiple different adaptation sets under the same period. The UserLabel (user label) structure of the adaptation set is added to describe the information and state of the participant user and the user access device. The users can be grouped through the user label.
[0043] In some optional implementations of the embodiment, the server is further configured to add a pre-selection option in the streaming media description file, for recommending an adaptation set to a participant and allowing the participant to set the adaptation set to be accepted.
[0044] The simulcast and the webRTC multi-person transfer matrix (SFU) are each mapped to a preselection (pre-selection option) of the received stream information, and each preselection is a stream media signal that can be received by the participant. In this way, filtering of the content that can be received by different receivers can be realized on the server side.
[0045] In some optional implementations of the embodiment, the organizer terminal is further configured to send at least one of the following event information to the server: control information, meta-universe information, and media information; the server is further configured to update the streaming media description file according to the event information; and the participant terminal is further configured to obtain the updated streaming media description file from the server at a timing, and perform event processing according to the updated streaming media description file.
[0046] Additional scalable information used in the meta-universe, storage, compression, encapsulation, transmission, and decapsulation of the scalable information.
[0047] In the meta-universe, not only audio and video information needs to be transmitted, but also other rich information needs to be transmitted to realize a richer interactive experience. Therefore, the present application designs and considers the following rich information types that are missing in the existing streaming media.
[0048] As Figure 3As shown, it includes audio information (Video Info) and video information (Audio Info) and metaverse information. The media manifest lists some existing functions, such as DRM (Digital rights), SDP / DASH mapping, subtitle processing, metaverse info processing, AD (advertising), and interactive.
[0049] In some optional implementations of the embodiment, the server is further configured to add an adaptive set including at least one of the following to the existing adaptive set configuration of the media manifest file: external reference media information, action posture information, digital asset information, and virtual environment information.
[0050] 1. External reference media information (External Media)
[0051] In a scenario similar to a lecture or a briefing, in addition to live capture, there are already generated, processed, and no longer changed audio and video information. For example, a static picture form of a speech, an already edited audio and video file, etc.
[0052] 2. Action posture information (Posture Info)
[0053] In the interaction of xR in the metaverse, in order to realize further interaction between participants and the environment, and interaction between participants, real-time collected action posture information needs to be transmitted together to realize direct body action posture interaction in the virtual space.
[0054] Human body action and posture information is to describe the current static skeletal posture state and facial expression information.
[0055] Among them, the human skeletal node information is to describe the absolute position of the node (such as Rootbone, Fixbone, Loosebone, etc.) shown in FIG. 1 and the relative position angle information between nodes. Figure 2
[0056] 3. Digital asset information (Digital Asserts)
[0057] In order to display the digital assets of the participants in different virtual spaces in the metaverse, it is necessary to identify and authenticate the system type and the identification of the digital certificate of the authentication system in the streaming protocol.
[0058] 4. Virtual world info
[0059] In the metaverse, participants often act in multiple different virtual environments, each of which has different content, style, and purpose.
[0060] Therefore, in order to realize the coordinated display of characters and environments, the information needs to be labeled, so that the client can render different virtual environments for different participating users.
[0061] In some optional implementations of the embodiment, the server is further configured to: group all participants according to a user tag structure; and recommend an adaptive set adapted to each group of participants.
[0062] In some optional implementations of the embodiment, the server is further configured to: send the stream description file to the participant terminal through an HTTP channel; and in response to receiving the media data, map the media data into segment data and send the segment data to the participant terminal through an RTC data channel.
[0063] In some optional implementations of the embodiment, the participant terminal comprises: a stream description file parser configured to parse the stream description file to obtain an adaptive set and a pre-selection option; an adaptive bit rate selection controller configured to perform adaptive bit rate selection according to the pre-selection option; a terminal reporter configured to report error terminal information and session information; a segment parser configured to parse the segment data to find additional information at the encapsulation level; an event processor configured to insert events in the stream description file and in the media data; and a download controller configured to perform media stream download according to the pre-selection option.
[0064] Figure 4 A stream media server (server) and client (client) structure are provided to support webRTC. The stream media server is Figure 1 the server, and the client is the organizer terminal or the participant terminal.
[0065] The server needs to map the existing webRTC description file (SDP) to the stream description file. The mapping relationship has been described above.
[0066] Both the server and the client contain multiple network communication mechanisms, such as ordinary HTTP and end-to-end webSocket network protocols, etc. The server can choose to transmit data through ordinary HTTP (TCP) or RTCDataChannel (UDP).
[0067] In the client, the following parts are included: streaming description file parser, ABR (adaptive bit rate selection) controller, terminal reporter, segment parser, event processor, download controller. Their functions are as follows:
[0068] 1. Streaming description file parser
[0069] The adaptation set mapped by the SDP and the preselection of different user accesses controlled by the server can be parsed according to the mapping relationship above.
[0070] The server or session initiator can limit the access of each user. It groups different users according to the userGroupID and describes the information of the acceptable stream in each preselection, such as the number of acceptable streams, code rate, whether to carry additional information, etc.
[0071] 2. ABR controller and download controller
[0072] The existing webrtc control is the selection of the client by the server or the push stream end. In this application, a variety of ABR controllers already existing in the streaming description file ecology are used to select the code rate in the client. The code rate selection that meets the current device best can be realized.
[0073] 3. Terminal reporter
[0074] The information reporting used by the current webrtc is determined according to the implementation of each manufacturer, and there is no clear specification. The information reporting cannot realize intercommunication.
[0075] Here, the json description load conforming to the CMCD (common media client data) (CTA-5004) specification is used to report the load such as stream type, session id, maximum code rate, and current cache duration. The reporting information is reported to the server in the DVB message conforming manner.
[0076] 4. Segment parser
[0077] In order to obtain relevant information such as configuration information and event information before the media stream is decoded, the media stream needs to be parsed after the transmission of the stream is completed, and the additional information of the encapsulation level is found, such as the general encryption information specified in iso23001-7 and the information for expressing the insertion stream time node specified in scte-35.
[0078] 5. Event processor
[0079] Referring to the fragment parser, in the existing webRTC mechanism, the event control and media transmission are separated into two channels, which brings certain event synchronization control problems. In order to solve this problem, the existing event triggering and processing process in the streaming industry is used, 1) insert events in the manifest, wait for the streaming description file to update each time, and there is an opportunity to process and trigger events. 2) In the media stream, insert events conforming to ISO BMFF in the encapsulation level.
[0080] Figure 5 The flowchart of starting the link establishment and playing process and the multi-client synchronization mechanism is shown.
[0081] First, the organizer terminal (webRTC Raiser) initiates a registration session request (Register event) to the streaming description file relay (Manifest / RTCSession Relay Time sync, i.e. server), and the streaming description file relay stores the participant information of this session,
[0082] When other participant terminals (Streaming Client) join, they can access the media resource information (Request Manifest) on the streaming description file relay to obtain the streaming description file (Manifest) of this activity through the RTC stream (RTC stream).
[0083] The participant terminal processes the streaming description file (Manifest process), for example, chooses the preselection.
[0084] After the participant terminal receives the streaming description file, it parses the webRTC media description layer in it, uses the RTC library (RTC connection lib) for parsing, and obtains the RTC stream connection information (RTC Streams connection info). Then the participant terminal establishes a session SDP exchange with the organizer terminal to set up a direct link (Session SDP exchange, set-upp2p link).
[0085] When active, the organizer terminal sends events (e.g., control information (Control msg), meta-universe information (Mateverse Info), media information (Media info)), update the MPD streaming description file in the MPD relay, for example, introduce event information (Event relay&ingest), introduce meta-universe information (Mateverse info ingest), introduce media information (media info ingest).
[0086] When the content of the media changes, or new media elements are added, it is completed by updating the streaming description file in the streaming description file relay.
[0087] The participant terminal will periodically re-request the streaming description file (Manifest updating request) in the streaming description file relay, and the streaming description file on the MPD relay will be updated with the event (Manifest with event).
[0088] The participant terminal will respond to the updated streaming description file processing event (Manifest process event reaction).
[0089] Real-time data stream (RTP) is transmitted through the UDP data transmission channel (RTP / UDP data transfer). Meta-universe information is transmitted through the HTTP data transmission channel (Mateverse Info / HTTP data transfer).
[0090] The media data decrypted by the RTC connection library (media decrypt) is rendered. The meta-universe data parsed by the HTTP connection library (Mateverse data) is directly rendered as meta-universe information (Mateverse Info).
[0091] In some optional implementations of the embodiment, at least one participant terminal is further configured to send at least one piece of interaction information to the server; the server is further configured to synthesize the at least one piece of interaction information to update the streaming description file for the organizer terminal and all participant terminals to obtain at regular intervals.
[0092] In some optional implementations of the embodiment, the server is further configured to write the copyright information of the participant in the streaming media description file; the participant terminal is further configured to decrypt the received media data according to the version information, and render the decrypted media data. The original MPD contains copyright information, and the SDP maps the MPD while retaining the copyright information. The organizer will assign authorized users a key corresponding to the copyright information, and the authorized users will use the key to decrypt the encrypted data protected by the copyright to see the original data when receiving the encrypted data.
[0093] When there are different participants issuing interactive behaviors, the server synthesizes them. As shown in Figure 6 Fig. 1, the main publisher 1 provides a main stage view in the form of a video stream, a background of a virtual world, and a character 1 of a digital assert. The main publisher 2 provides a left 1 view in the form of a video stream and a right 1 view in the form of a video stream, and a character 2 of a digital assert. An AD publisher provides an AD and a digital store good. Other users provide interactive attendees and gestures. The above-mentioned contents are mixed into an RTC mapper by a manifest mixer, and a merged manifest is obtained after mixing. Then, a user prefer view is obtained through preselection based choices mapping.
[0094] The system provided by the present application can realize an interactive large-scale online real-time activity, such as an Astronomical concert. The scene has a small number of main performers who send a large amount of rich media information (concert performers, media publishers), a large number of viewers who mainly accept but have interactions, sponsors, advertisers, and recording and archiving storage, etc.
[0095] The application is further described by examples:
[0096] Step 1.1: Participants with main content of the actor type: register an interactive large-scale online real-time activity in the metaverse, publish the main content through the session management center, and the main content refers to the multi-angle, multi-channel and multi-stage rich media information in the concert. Reference Figure 7 , which can be regarded as a WebRTC initiator role. It completes the registration of the real-time activity and publishes it to the public participants. This type of participant sends main data to the MPD hybrid relay. Its data is scattered in multiple protocols and multiple ways.
[0097] The characteristics of this type of participant are to consume main bandwidth, occupy a large part of the total duration of the activity, produce multiple-angle audio-visual information, and often require high-level copyright protection. The generated information needs to be synchronized among multiple other participants, and is the main producer of real-time data.
[0098] Step 1.2: MPD relay and RTC session management, after receiving data from the participants, generate the main MPD information of the real-time activity, as shown in the figure, the MPD describes two user-selectable perspectives, cameras, their ids are mainstage1 and leftStage respectively, representing two different perspectives, and there is a sender-recommended configuration option id preselection.
[0099] Step 2.1: Participants with pre-made content, such as sponsors, advertisers, and digital art providers, send pre-made data to the MPD relay. This type of participant sends data that is incorporated into the main audio-visual data stream to the MPD hybrid relay.
[0100] The characteristics of this type of participant are that the inserted content is pre-made, the duration is short, the time point of insertion into the main content is not fixed, and the inserted content can vary with the type of recipient. It is fragmented data. For example Figure 8 , id = AD1 represents Advertiser 1, and the provided uri content is ad (advertisement).
[0101] Step 2.2: MPD relay and RTC session management, after receiving the pre-made data, mix the MPD according to the display requirements, such as at the main program before and after the time node, in the spatial position of the main program, perspective, i.e. spatiotemporal relativity. The mixed MPD is as shown in Figure 9 .
[0102] Step3.1: Participants as audience receive mixed MPD from MPD relay, filter the view angle according to preselection and personal interest, and interact with performers through event reporting mechanism. As shown in FIG. 3, preselection is used to set preselection, in which id=localSuggestion is the recommended option "mainstage1" of the server, and id=costume1 is the option "mainstage1leftstage" set by the user. The user not only accepts the recommended main stage (mainstage1), but also accepts the left stage (leftstage) option. The user filters out the right stage (rightstage). After setting, the user will only receive video data of the main stage and the left stage, and will not download data of the right stage. Figure 10
[0103] Step3.2: After receiving the interactive information from the audience, the MPD relay forwards the message to the target group. As shown in FIG. 4, the main publisher 1 and the main publisher 2 provide the main content pose information, the advertiser provides the advertisement, and the other users provide the interactive behavior information. These information are mixed into the mixed MPD by the MPD relay and injected into the RTC interrupt. Based on the preselected media / event / resource, the user will filter out the content that the user does not want to download, for example, the content provided by the main publisher 2 is filtered out in the figure. Figure 11
[0104] Compared with the prior art, the application can significantly reduce the delay of streaming media interaction based on streaming media, take advantage of the ultra-low delay of webRTC, and use advanced streaming media functions in the streaming media description file, such as digital copyright protection, live ad insertion, additional information transmission, advanced audio and video encoder functions, extensible server events, error terminal information collection and reporting, etc. Better live performance and more complex interactive scenarios can be achieved.
[0105] The above detailed description does not constitute a limitation on the protection scope of the present disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations and substitutions can be made according to design requirements and other factors. Any modification, equivalent replacement and improvement within the spirit and principles of the present disclosure should be included in the protection scope of the present disclosure.
Claims
1. A real-time interactive streaming system, comprising: an organizer terminal configured to send a registration session request including participant information to a server, establish a direct connection with a participant terminal and interact data; the server configured to, in response to receiving the registration session request, map a session description into an MPD streaming description file, store the participant information into the streaming description file, in response to receiving a request for acquiring the streaming description file sent by a participant terminal, send the streaming description file to the participant terminal; a participant terminal configured to acquire the streaming description file from the server, parse connection information from the streaming description file, and establish a direct connection with the organizer terminal and interact data according to the connection information; wherein the mapping of the session description into the MPD streaming description file comprises: adding a new adaptive set on the basis of an existing streaming description file with adaptive set configuration, wherein the new adaptive set includes a description type of attributes of the session description and related parameter agreements under the description type; adding a user tag structure for describing participant information and participant device information in the new adaptive set.
2. The system of claim 1, wherein: the organizer terminal is further configured to send at least one of the following event information to the server: control information, meta-universe information, media information; the server is further configured to update the streaming description file according to the event information; the participant terminal is further configured to acquire the updated streaming description file from the server at a regular time, and perform event processing according to the updated streaming description file.
3. The system of claim 1, wherein, the server is further configured to: add a pre-selection option in the streaming description file, for recommending adaptive sets to participants and allowing participants to set adaptive sets to be accepted by themselves.
4. The system of claim 3, wherein, the server is further configured to: group all participants according to the user tag structure; recommend adaptive sets suitable for each group of participants to the participants in the group.
5. The system of claim 1 or 3, wherein, the server is further configured to: add a new adaptive set including at least one of the following on the basis of an existing streaming description file with adaptive set configuration: external reference media information, action pose information, digital asset information, virtual environment information.
6. The system of claim 1, wherein, the server is further configured to: send the streaming description file to the participant terminal through an HTTP channel; in response to receiving media data, map the media data into segment data, and send the segment data to the participant terminal through an RTC data channel.
7. The system of claim 6, wherein, the participant terminal comprises: a streaming description file parser configured to parse the streaming description file to obtain adaptive sets and pre-selection options; an adaptive bit rate selection controller configured to perform adaptive bit rate selection according to the pre-selection options; a terminal reporter configured to report error terminal information and session information; a segment parser configured to parse the segment data and find additional information at an encapsulation level. an event handler configured to insert events in the streaming description file and insert events in the media data; a download controller configured to perform media stream downloading according to pre-selected options.
8. The system of claim 1, wherein, the at least one participant terminal is further configured to send at least one piece of interaction information to the server; the server is further configured to synthesize the at least one piece of interaction information to update the streaming description file for periodic acquisition by the organizer terminal and all participant terminals.
9. The system of claim 1, wherein, the server is further configured to write copyright information of the participant in the streaming description file; the participant terminal is further configured to decrypt the received media data according to the copyright information and render the decrypted media data.
Citation Information
Patent Citations
Media stream transmission control method and device, storage medium and electronic equipment
CN115865878A