Audio processing method and apparatus, device, and storage medium
By pushing data streams of target music and audio content during live interactive events and unbinding them across different terminals, the problem of one-way audio transmission in live interactive events is solved, improving user experience and interactive effects.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- BEIJING ZITIAO NETWORK TECH CO LTD
- Filing Date
- 2025-11-14
- Publication Date
- 2026-05-21
AI Technical Summary
In live streaming interactions, existing technology prevents the lead singer from hearing the backing vocals, and vice versa, thus reducing the user's interactive experience.
By pushing data streams including target music and audio content during live interactive events and unbinding them on different terminals, the lead singer and the backing singer can hear each other's audio content, achieving two-way audio transmission.
It enhances the interactive experience for users in the live streaming room, allowing participants to adjust the volume and sound effects of the accompaniment according to their own situation, thereby enhancing the real-time nature and sense of participation in the interaction.
Smart Images

Figure CN2025135130_21052026_PF_FP_ABST
Abstract
Description
Audio processing methods, apparatus, devices and storage media
[0001] This application claims priority to Chinese Patent Application No. 202411640735.0, filed on November 15, 2024, entitled "Audio Processing Method, Apparatus, Device and Storage Medium", the entire contents of which are incorporated herein by reference. Technical Field
[0002] The exemplary embodiments disclosed herein generally relate to the field of computers, and particularly to audio processing methods, apparatus, devices, and computer-readable storage media. Background Technology
[0003] With the rapid development of internet technology, live streaming has gradually become a new way of exchanging information. Users can interact with each other via electronic devices. How to improve the user's interactive experience in the live streaming room is a key concern. Summary of the Invention
[0004] In a first aspect of this disclosure, an audio processing method is provided, comprising: in response to a first participant joining a live interactive event associated with target music, pushing a first data stream at a first terminal corresponding to the first participant, the first data stream including target music content and first audio content collected by the first terminal; and in response to the first participant being a first role in the live interactive event, pushing a second data stream simultaneously with pushing the first data stream, the second data stream including the first audio content.
[0005] In a second aspect of this disclosure, an audio processing method is provided, comprising: in response to a second participant joining a live interactive event associated with target music, pushing a third data stream at a second terminal corresponding to the second participant, the third data stream including second audio content collected by the second terminal; and in response to the second participant being a second role in the live interactive event, receiving a second data stream pushed by a first terminal, the second data stream including first audio content collected by the first terminal, the first terminal corresponding to a first participant in the live interactive event.
[0006] In a third aspect of this disclosure, an audio processing method is provided, comprising: in response to a third participant joining a live interactive event associated with target music, receiving, at a third terminal corresponding to the third participant, a first data stream created by a first terminal and a third data stream created by a second terminal, wherein the first data stream includes the target music and first audio content collected by the first terminal, and the third data stream includes second audio content collected by the second terminal, wherein the first terminal corresponds to the first participant in the live interactive event, and the second terminal corresponds to the second participant in the live interactive event.
[0007] In a fourth aspect of this disclosure, an audio processing apparatus is provided, comprising: a first push module configured to, in response to a first participant joining a live interactive event associated with target music, push a first data stream at a first terminal corresponding to the first participant, the first data stream including target music content and first audio content captured by the first terminal; and a second push module configured to, in response to the first participant being a first role in the live interactive event, push a second data stream simultaneously with the first data stream, the second data stream including the first audio content.
[0008] In a fifth aspect of this disclosure, an audio processing apparatus is provided, comprising: a third push module configured to push a third data stream at a second terminal corresponding to the second participant in response to a second participant joining a live interactive event associated with target music, the third data stream including second audio content collected by the second terminal; and a first receiving module configured to receive a second data stream pushed by a first terminal in response to the second participant being a second role in the live interactive event, the second data stream including first audio content collected by the first terminal, the first terminal corresponding to a first participant in the live interactive event.
[0009] In a sixth aspect of this disclosure, an audio processing apparatus is provided, comprising: a second receiving module configured to, in response to a third participant joining a live interactive event associated with target music, receive, at a third terminal corresponding to the third participant, a first data stream created by a first terminal and a third data stream pushed by a second terminal, the first data stream including the target music and first audio content collected by the first terminal, the third data stream including second audio content collected by the second terminal, the first terminal corresponding to the first participant in the live interactive event, and the second terminal corresponding to the second participant in the live interactive event.
[0010] In a seventh aspect of this disclosure, an electronic device is provided, comprising: at least one processor; and at least one memory coupled to the at least one processor and storing instructions for execution by the at least one processor, the instructions, when executed by the at least one processor, causing the terminal device to perform the method of the first aspect, the method of the second aspect, or the method of the third aspect.
[0011] In an eighth aspect of this disclosure, a computer-readable storage medium is provided. The computer-readable storage medium stores computer-executable instructions that can be executed by a processor to implement the method of the first aspect.
[0012] In a ninth aspect of this disclosure, a computer program product is provided. The computer program product is tangibly stored in a computer storage medium and includes computer-executable instructions, which, when executed by a device, perform the method of the first aspect, the method of the second aspect, or the method of the third aspect.
[0013] It should be understood that the content described in this content section is not intended to limit the key or essential features of the embodiments of this disclosure, nor is it intended to restrict the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description
[0014] The above and other features, advantages, and aspects of the embodiments of this disclosure will become more apparent from the accompanying drawings and the following detailed description. In the drawings, the same or similar reference numerals denote the same or similar elements, wherein:
[0015] Figure 1 shows a schematic diagram of an example environment in which embodiments of the present disclosure can be implemented;
[0016] Figure 2 illustrates a schematic diagram of the interaction process involved in a live interactive event according to some embodiments of the present disclosure;
[0017] Figure 3 illustrates a schematic diagram of an example of an audio processing procedure according to some embodiments of the present disclosure;
[0018] Figure 4 illustrates another example of an audio processing procedure according to some embodiments of the present disclosure;
[0019] Figure 5 shows a schematic diagram of yet another example of an audio processing procedure according to some embodiments of the present disclosure;
[0020] Figure 6 shows a schematic diagram of an example of an audio processing apparatus according to some embodiments of the present disclosure;
[0021] Figure 7 shows a schematic diagram of another example of an audio processing apparatus according to some embodiments of the present disclosure;
[0022] Figure 8 shows a schematic diagram of yet another example of an audio processing apparatus according to some embodiments of the present disclosure;
[0023] Figure 9 illustrates a schematic diagram of an example of an audio processing system according to some embodiments of the present disclosure; and
[0024] Figure 10 shows a block diagram of an apparatus capable of implementing several embodiments of the present disclosure. Detailed Implementation
[0025] It is understood that before using the technical solutions disclosed in the various embodiments of this disclosure, users should be informed of the types, scope of use, and usage scenarios of the personal information involved in this disclosure in an appropriate manner in accordance with relevant laws and regulations, and user authorization should be obtained.
[0026] For example, upon receiving a user's active request, a prompt message is sent to the user to explicitly inform them that the requested operation will require the acquisition and use of the user's personal information. This allows the user to independently choose whether to provide personal information to the software or hardware, such as the electronic device, application, server, or storage medium performing the operations of this disclosed technical solution, based on the prompt message.
[0027] As an optional but non-limiting implementation, in response to a user's active request, sending a prompt message to the user can be done via a pop-up window, where the prompt message can be presented in text format. Furthermore, the pop-up window can also include a selection control allowing the user to choose "agree" or "disagree" to provide personal information to the electronic device.
[0028] It is understood that the above notification and user authorization process are merely illustrative and do not constitute a limitation on the implementation of this disclosure. Other methods that comply with relevant laws and regulations may also be applied to the implementation of this disclosure.
[0029] It is understood that the data involved in this technical solution (including but not limited to the data itself, the acquisition or use of the data) shall comply with the requirements of relevant laws, regulations and related provisions.
[0030] Embodiments of this disclosure will now be described in more detail with reference to the accompanying drawings. While some embodiments of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.
[0031] It should be noted that the headings of any section / subsection provided herein are not limiting. Various embodiments are described throughout this document, and embodiments of any type may be included under any section / subsection. Furthermore, embodiments described in any section / subsection may be combined in any way with any other embodiments described in the same section / subsection and / or different sections / subsections.
[0032] In this document, unless explicitly stated otherwise, performing a step in response to A does not mean that the step is performed immediately after A, but may include one or more intermediate steps.
[0033] In the description of embodiments of this disclosure, the term "comprising" and similar terms should be understood as open-ended inclusion, i.e., "including but not limited to". The term "based on" should be understood as "at least partially based on". The term "one embodiment" or "the embodiment" should be understood as "at least one embodiment". The term "some embodiments" should be understood as "at least some embodiments". Other explicit and implicit definitions may also be included below. The terms "first", "second", etc., may refer to different or the same objects. Other explicit and implicit definitions may also be included below.
[0034] As briefly mentioned earlier, with the rapid development of internet technology, live streaming has gradually become a new way of information exchange. Users can interact via electronic devices. For example, in a multi-person chorus, a "one-way" approach is typically used, meaning the audio stream between the lead singer and the backup singers is unidirectional, flowing only from the lead singer to the backup singers. This results in the lead singer not being able to hear the backup singers. Simultaneously, the backup singers cannot hear each other. Furthermore, in a "one-way" scenario, the backup singers' accompaniment originates from the lead singer; the backup singers do not play the accompaniment locally. This prevents them from adjusting the accompaniment volume and effects according to their own needs, thus reducing the interactive experience for users participating in the chorus during the live stream.
[0035] Embodiments of this disclosure propose an audio processing scheme. According to various embodiments of this disclosure, in response to a first participant joining a live interactive event associated with target music, a first data stream is pushed at a first terminal corresponding to the first participant. The first data stream includes the target music and first audio content collected by the first terminal. In response to the first participant being a first role in the live interactive event, a second data stream is pushed simultaneously with the first data stream. The second data stream includes the first audio content.
[0036] Therefore, the embodiments of this disclosure can, during live interactive sessions, create a first data stream including target music and first audio content and a second data stream including the first audio content, and send the first data stream and the second data stream to the second terminal respectively, thereby achieving the debinding of target music and first audio content, and thus improving the user experience.
[0037] The following description will focus on exemplary embodiments of the present disclosure with reference to the accompanying drawings.
[0038] Figure 1 illustrates a schematic diagram of an example environment 100 in which embodiments of the present disclosure can be implemented. As shown in Figure 1, environment 100 may include terminal devices 110-1, 110-2, ..., 110-N and server 120. For ease of discussion, terminal devices 110-1, 110-2, ..., 110-N may be collectively referred to as terminal device 110 or individually referred to as terminal device 110.
[0039] In environment 100, terminal device 110 may run an application that supports voice interaction. This application can be any suitable type of application for voice interaction, examples of which include, but are not limited to, video applications, social applications, live streaming applications, or other suitable applications. Users can perform activities related to live streaming based on this application. In some embodiments, activities related to live streaming can include initiating a live stream, participating in a live stream, and watching a live stream, etc. Users can interact with the application via terminal device 110 and / or its attached devices.
[0040] In environment 100, terminal device 110 can be any type of computing-capable device, including terminal devices or server devices. Terminal devices can be any type of mobile terminal, fixed terminal, or portable terminal, including mobile phones, desktop computers, laptop computers, notebook computers, netbook computers, tablet computers, media computers, multimedia tablets, personal communication system (PCS) devices, personal navigation devices, personal digital assistants (PDAs), audio / video players, digital cameras / camcorders, positioning devices, television receivers, radio receivers, e-book devices, gaming devices, or any combination of the foregoing, including accessories and peripherals of these devices or any combination thereof. Server devices may include, for example, computing systems / servers, such as mainframes, edge computing nodes, electronic devices in cloud environments, etc.
[0041] Server 120 can be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks, and big data and artificial intelligence platforms. Server 120 may include, for example, computing systems / servers such as mainframes, edge computing nodes, computing devices in a cloud environment, etc. Server 120 can provide backend services for applications supporting voice interaction in terminal device 110.
[0042] A communication connection can be established between the server 120 and the terminal device 110, and a communication connection can also be established between the terminal devices 110. The communication connection can be established via wired or wireless means. The communication connection can include, but is not limited to, Bluetooth connection, mobile network connection, Universal Serial Bus (USB) connection, Wireless Fidelity (WiFi) connection, etc., and the embodiments disclosed herein are not limited in this respect.
[0043] It should be understood that the structure and function of environment 100 are described for illustrative purposes only and do not imply any limitation on the scope of this disclosure.
[0044] The following will describe some exemplary embodiments of this disclosure with reference to a specific example.
[0045] Figure 2 illustrates a schematic diagram of the interaction process 200 involved in a live interactive event according to some embodiments of the present disclosure. The following description will use the execution of the interaction process 200 by a terminal device as an example. Next, the audio processing scheme provided by the embodiments of the present disclosure will be described in detail with reference to Figure 1.
[0046] Live interactive events can include appropriate types of interactive events that occur in the live room (also known as a live session). In some embodiments, live interactive events associated with target music can be, for example, multi-person chorus, multi-person poetry recitation, etc., and the associated target music can be the accompaniment of a song in a chorus, the background music for a poetry recitation, etc.
[0047] Taking a live interactive event as an example of a chorus event, as shown in Figure 2, the multiple participants in the chorus can include the first participant (lead singer A), the second participant (backup singers B and C), guest participants (guests on the microphone D (host) and E), and the live stream viewers (viewer 1, viewer 2, ..., viewer n). It should be understood that guest participants can also refer to user roles who did not participate in the chorus.
[0048] Each of the multiple participating parties may have a corresponding terminal device 110. For ease of discussion, the terminal device 110 corresponding to user A can be referred to as the "first terminal", the terminal devices 110 corresponding to user B and user C can be collectively referred to as the "second terminal" or individually, and the terminal devices 110 corresponding to user D and user E can be collectively referred to as the "third terminal" or individually.
[0049] In some embodiments, server 120 may have a first real-time communications (RTC) room and a second real-time communications (RTC) room. For ease of discussion, the first RTC room may be referred to as the main room 210, and the second RTC room may be referred to as the secondary room 220.
[0050] In some embodiments, the main room 210 can be represented by a room ID, and the secondary room 220 can be represented by the room ID of the main room 210 and a predetermined suffix. For example, the main room 210 can be represented as "roomid", and correspondingly, the secondary room 220 can be represented as "roomid_byte_aux". Here, "roomid" represents the ID of the main room 210, and "byte_aux" represents the predetermined suffix.
[0051] Referring to Figure 2, in a multi-person chorus scenario, the interaction between the first terminal, the second terminal, and the third terminal can be achieved through the main room 210 and the secondary room 220.
[0052] Specifically, the first terminal responds to the first participant joining the live interactive event associated with the target music, that is, the first terminal corresponding to user A joins the main room 210. After that, the first terminal can push the first data stream to the main room 210, so that the second terminal and the third terminal that have also joined the main room 210 can receive the first data stream.
[0053] In some embodiments, the first data stream can be pushed by the first terminal based on target music and first audio content. For example, the target music can be the background music (bgm) of a choral song played locally on the first terminal, and the first audio content can be the voice of user A captured by the microphone of the first terminal. For example, the first data stream can be represented as macA+bgm.
[0054] The first terminal responds to the first participant being the first role in the live interactive event, that is, the first terminal detects that user A initiates a duet request through the first terminal (at this time, user A can be called lead singer A), and the first terminal pushes a second data stream to the secondary room 220 based on the first audio content. The second terminal can receive the second data stream pushed by the first terminal by subscribing to the secondary room 220. For example, the second data stream can be represented as macA.
[0055] Therefore, it is possible to decouple vocals and background music at the first terminal.
[0056] In some embodiments, a first data stream corresponds to a first real-time communication channel associated with a chorus event, and a second data stream corresponds to a second real-time communication channel associated with a chorus event. The first and second real-time communication channels are associated with different roles in the chorus event.
[0057] In some embodiments, in response to macA being pushed to the second terminal, the first terminal can add timestamp information to macA+bgm. In some embodiments, the first terminal can add timestamp information to macA+bgm based on the Network Time Protocol (NTP) so that other devices can perform alignment processing based on the timestamp information. This part will be described below and will not be repeated here.
[0058] In some embodiments, in response to the target music being played on the first terminal, the first terminal may add collaboration information in macA, which indicates the playback progress and / or playback status of the target music on the first terminal.
[0059] In some embodiments, the collaboration information indicates the playback progress of the target music on the first terminal, which may be the decoding progress of the target music by the first terminal. In some embodiments, the collaboration information indicates the playback status of the target music on the first terminal, which may be a play, pause, or speed-up playback status.
[0060] The second terminal responds to the second participant joining the live interactive event associated with the target music, that is, user B and / or user C join the main room 210 through their respective second terminals. After that, the second terminal can push a third data stream to the main room 210, so that the first terminal and the third terminal that have also joined the main room 210 can receive the third data stream.
[0061] In some embodiments, the third data stream can be pushed by the second terminal based on the second audio content. For example, the second audio content can be the voice of user B or user C captured by the microphone of the second terminal. For example, the third data stream can be represented as macB and macC.
[0062] The second terminal responds to the second participant being the second role in the live interactive event, that is, the second terminal detects that user B and user C have initiated a request to join the chorus through their respective second terminals (at this time, user B can be called the backing vocalist B, and user C can be called the backing vocalist C). The second terminal subscribes to the data stream in the secondary room 220 to receive the macA created by the first terminal.
[0063] This allows the lead singer to hear the backing vocals, and vice versa, thus enabling real-time "dual-channel" chorus.
[0064] In some embodiments, the second terminal responds to the second participant joining a live interactive event associated with the target music, i.e., backing vocalist B and / or backing vocalist C join the main room 210 through their respective second terminals and receive macA+bgm created by the first terminal; the second terminal responds to the second participant being the second role in the live interactive event, i.e., the second terminal detects the second participant's request to join the chorus initiated by the second terminal, stops receiving macA+bgm, and only receives macA through the secondary room 220, thereby enabling the local playback of the target music (the accompaniment of the chorus song), and allowing backing vocalist B and backing vocalist C to adjust relevant parameters of the target music at the second terminal according to their own situation, such as loudness, pitch, reverb, etc.
[0065] In some embodiments, during the local playback of target music on the second terminal, the playback progress of the target music on the second terminal can be adjusted based on collaboration information.
[0066] For example, if the target music in the collaboration information is paused, the second terminal needs to pause the target music playing locally. Alternatively, the second terminal can adjust the playback progress of the target music based on the playback progress in the collaboration information, so that the target music is played synchronously on both the first and second terminals.
[0067] The third terminal responds to the live interactive event associated with the target music by the third participant joining the live stream. Specifically, the third terminal detects that the third terminals corresponding to users D and E have joined the main room 210. The third terminal subscribes to the macA+bgm pushed by the first terminal and the macB and macC pushed by the second terminal in the main room 210, as well as the data streams sent by other third terminals.
[0068] In some embodiments, the third terminal responds to the third participant being the third role in the live interactive event, that is, the third terminal detects that user D and user E have initiated a microphone request through their respective third terminals (at this time, user D can be called microphone guest D, and user E can be called microphone guest E), and pushes the fourth data stream to the main room 210.
[0069] In some embodiments, the fourth data stream can be pushed by a third terminal based on third audio content. For example, the third audio content can be the voice of user D or user E captured by the microphone of the third terminal. For example, the third data stream can be represented as macD and macE.
[0070] In some embodiments, if guest D is the host, the third terminal corresponding to the host can push the fourth data stream macD created by the third terminal to the main room 210, and the third terminal corresponding to the guest can push the fourth data stream macE created by the third terminal to the main room 210. In this case, the third terminal corresponding to the host can subscribe to macA+bgm, macB, macC, and macE in the main room 210; guest E can subscribe to macA+bgm, macB, macC, and macD in the main room 210; the first terminal can subscribe to macB, macC, and macD in the main room 210; and the second terminal can subscribe to macB, macC, and macD in the main room 210.
[0071] As can be seen from the foregoing, the first terminal added timestamp information to macA+bgm.
[0072] In some embodiments, the third terminal may perform time alignment processing on macA+bgm, macB, and macC based on timestamp information to ensure that the target music is played in alignment with the voices of lead vocalist A, backing vocalist B, and backing vocalist C.
[0073] In some embodiments, the third terminal can re-align macD and macE, as well as the aligned macA+bgm, macB and macC, and send the aligned macA+bgm, macB, macC, macD and macE to the Content Delivery Network (CDN), so that CDN230 can push the aligned macA+bgm, macB, macC, macD and macE to the terminal devices corresponding to each of the live stream viewers 1, 2, ..., n.
[0074] In some embodiments, the third terminal may also send macD and macE, as well as the aligned macA+bgm, macB and macC, to the server 120. The server 120 may align macA+bgm, macB, macC, macD and macE based on timestamp information, and send the aligned macA+bgm, macB, macC, macD and macE to the Content Delivery Network (CDN), so that the CDN 230 may send the aligned macA+bgm, macB, macC, macD and macE to the terminal devices corresponding to each of the live stream viewers 1, 2, ..., n.
[0075] In this way, the embodiments of this disclosure can push a first data stream including target music and first audio content and a second data stream including the first audio content to a second terminal during live interactive sessions, thereby debinding the target music and the first audio content and improving the user experience.
[0076] Figure 3 illustrates a schematic diagram of an example of an audio processing procedure 300 according to some embodiments of the present disclosure. Procedure 300 may be implemented at a first terminal.
[0077] In frame 310, the first terminal responds to the first participant joining the live interactive event associated with the target music by pushing the first data stream, which includes the target music and the first audio content collected by the first terminal.
[0078] In frame 320, the first terminal responds to the first participant being the first role in the live interactive event, and pushes a second data stream while pushing the first data stream. The second data stream includes the first audio content.
[0079] In some embodiments, process 300 further includes adding timestamp information to the first data stream.
[0080] In some embodiments, process 300 further includes: in response to the target music being played at the first terminal, the first terminal adds collaboration information to the second data stream, the collaboration information indicating the playback progress and / or playback status of the target music at the first terminal.
[0081] In some embodiments, process 300 further includes: a first terminal receiving a third data stream pushed by a second terminal, the third data stream including second audio content collected by the second terminal, the second terminal being a second participant in the live interactive time.
[0082] In some embodiments, the first data stream corresponds to a first real-time communication channel associated with a live interactive event, and the second data stream corresponds to a second real-time communication channel associated with a live interactive event. The first and second real-time communication channels are associated with different roles in the live interactive event.
[0083] Figure 4 illustrates a schematic diagram of another example of an audio processing procedure 400 according to some embodiments of the present disclosure. Procedure 400 can be implemented at a second terminal.
[0084] In frame 410, the second terminal responds to the second participant joining the live interactive event associated with the target music by pushing a third data stream, which includes the second audio content collected by the second terminal.
[0085] In frame 420, the second terminal responds to the fact that the second participant is the second role in the live interactive event, and receives the second data stream pushed by the first terminal. The second data stream includes the first audio content collected by the first terminal, and the first terminal corresponds to the first participant in the live interactive time.
[0086] In some embodiments, process 400 further includes: the second terminal receiving a first data stream created by the first terminal in response to the second participant joining a live interactive event associated with the target music, the first data stream including the target music and first audio content; and stopping receiving the first data stream in response to the second participant being a second role in the live interactive event.
[0087] In some embodiments, the second data stream further includes collaboration information indicating the playback progress and / or playback status of the target music at the first terminal. Process 400 further includes: the second terminal adjusting the playback progress of the target music based on the collaboration information.
[0088] Figure 5 illustrates a schematic diagram of yet another example of an audio processing procedure 500 according to some embodiments of the present disclosure. Procedure 500 can be implemented at a third terminal.
[0089] In box 510, the third terminal responds to the third participant joining the live interactive event associated with the target music by receiving the first data stream pushed by the first terminal and the third data stream pushed by the second terminal. The first data stream includes the target music and the first audio content collected by the first terminal, and the third data stream includes the second audio content collected by the second terminal. The first terminal corresponds to the first participant in the live interactive event, and the second terminal corresponds to the second participant in the live interactive event.
[0090] In some embodiments, process 500 further includes: the third terminal, in response to the third participant being a third role in the live interactive event, creating a fourth data stream, the fourth data stream including third audio content collected by the third terminal.
[0091] In some embodiments, the first data stream further includes timestamp information, and process 500 further includes: the third terminal performing time alignment processing on the first data stream and the third data stream based on the timestamp information.
[0092] In some embodiments, process 500 further includes: a third terminal sending a fourth data stream and an aligned first and second data stream to a content delivery network, so that the content delivery network sends the fourth data stream and the aligned first and second data streams to the fourth terminal corresponding to the live stream's viewer.
[0093] Figure 6 illustrates a schematic diagram of an example of an audio processing apparatus 600 according to some embodiments of the present disclosure. The apparatus 600 may be implemented as or included in a first terminal. The various modules / components in the apparatus 600 may be implemented by hardware, software, firmware, or any combination thereof.
[0094] As shown in Figure 6, the device 600 includes a first push module 610, configured to push a first data stream at a first terminal corresponding to the first participant in response to the first participant joining a live interactive event associated with the target music. The first data stream includes the target music content and first audio content collected by the first terminal. The device 600 also includes a second push module 620, configured to push a second data stream simultaneously with the first data stream in response to the first participant being a first role in the live interactive event. The second data stream includes the first audio content.
[0095] In some embodiments, the apparatus 600 further includes a first adding module configured to add timestamp information to the first data stream.
[0096] In some embodiments, the device 600 further includes a second adding module configured to add collaborative information to a second data stream in response to the target music being played at a first terminal. The collaborative information indicates the playback progress and / or playback status of the target music at the first terminal.
[0097] In some embodiments, the device 600 further includes a third receiving module configured to receive a third data stream created by the second terminal, the third data stream including second audio content acquired by the second terminal.
[0098] In some embodiments, the first data stream corresponds to a first real-time communication channel associated with a live interactive event, and the second data stream corresponds to a second real-time communication channel associated with a live interactive event. The first and second real-time communication channels are associated with different roles in the live interactive event.
[0099] Figure 7 illustrates a schematic diagram of another example of an audio processing apparatus 700 according to some embodiments of the present disclosure. The apparatus 700 may be implemented as or included in a second terminal. The various modules / components in the apparatus 700 may be implemented by hardware, software, firmware, or any combination thereof.
[0100] As shown in Figure 7, the device 700 includes a third push module 710, configured to push a third data stream at a second terminal corresponding to the second participant in response to a second participant joining a live interactive event associated with the target music. The third data stream includes second audio content collected by the second terminal. The device 700 also includes a first receiving module 720, configured to receive a second data stream pushed by a first terminal in response to the second participant being a second role in the live interactive event. The second data stream includes first audio content collected by the first terminal, and the first terminal corresponds to the first participant in the live interactive event.
[0101] In some embodiments, the device 700 further includes a fourth receiving module configured to receive a first data stream created by the first terminal, the first data stream including the target music and first audio content, in response to a second participant joining a live interactive event associated with the target music; and to stop receiving the first data stream in response to the second participant being a second role in the live interactive event.
[0102] In some embodiments, the second data stream further includes collaboration information indicating the playback progress and / or playback status of the target music at the first terminal, and the device 700 further includes an adjustment module configured to adjust the playback progress of the target music on the second terminal based on the collaboration information.
[0103] Figure 8 illustrates a schematic diagram of yet another example of an audio processing apparatus 800 according to some embodiments of the present disclosure. The apparatus 800 may be implemented as or included in a third terminal. The various modules / components in the apparatus 800 may be implemented by hardware, software, firmware, or any combination thereof.
[0104] As shown in Figure 8, the device 800 includes a second receiving module 810, configured to receive a first data stream pushed by a first terminal and a third data stream pushed by a second terminal at a third terminal corresponding to the third participant in response to a third participant joining a live interactive event associated with the target music. The first data stream includes the target music and the first audio content collected by the first terminal, and the third data stream includes the second audio content collected by the second terminal. The first terminal corresponds to the first participant in the live interactive event, and the second terminal corresponds to the second participant in the live interactive event.
[0105] In some embodiments, the device 800 further includes a fourth push module configured to push a fourth data stream in response to a third participant being a third role in a live interactive event. The fourth data stream includes third audio content collected by a third terminal.
[0106] In some embodiments, the first data stream further includes timestamp information, and the apparatus 800 further includes an alignment module configured to perform time alignment processing on the first data stream and the third data stream based on the timestamp information.
[0107] In some embodiments, the apparatus 800 further includes a transmission module configured to transmit a fourth data stream and an aligned first and second data stream to a content delivery network, such that the content delivery network transmits the fourth data stream and the aligned first and second data streams to a fourth terminal corresponding to a live stream viewer.
[0108] Figure 9 shows a schematic diagram of an example of an audio processing system 900 according to some embodiments of the present disclosure.
[0109] As shown in Figure 9, system 900 includes a first terminal, a second terminal, and a third terminal. The first terminal is configured to execute the above-described process 300, the second terminal is configured to execute the above-described process 400, and the third terminal is configured to execute the above-described process 500.
[0110] Figure 10 illustrates a block diagram of an electronic device 1000 in which one or more embodiments of the present disclosure may be implemented. It should be understood that the electronic device 1000 shown in Figure 10 is merely exemplary and should not be construed as limiting the functionality and scope of the embodiments described herein. The electronic device 1000 shown in Figure 10 can be used to implement the electronic device 110 of Figure 1.
[0111] As shown in Figure 10, the electronic device 1000 is in the form of a general-purpose electronic device. Components of the electronic device 1000 may include, but are not limited to, one or more processors or processing units 1010, memory 1020, storage device 1030, one or more communication units 1040, one or more input devices 1050, and one or more output devices 1060. The processor 1010 may be a physical or virtual processor and is capable of performing various processes according to programs stored in memory 1020. In a multiprocessor system, multiple processors execute computer-executable instructions in parallel to improve the parallel processing capability of the electronic device 1000.
[0112] Electronic device 1000 typically includes multiple computer storage media. Such media can be any accessible media that is accessible to electronic device 1000, including but not limited to volatile and non-volatile media, removable and non-removable media. Memory 1020 can be volatile memory (e.g., registers, cache, random access memory (RAM)), non-volatile memory (e.g., read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory), or some combination thereof. Storage device 1030 can be removable or non-removable media and may include machine-readable media, such as flash drives, disks, or any other media that can be used to store information and / or data and can be accessed within electronic device 1000.
[0113] Electronic device 1000 may further include additional removable / non-removable, volatile / non-volatile storage media. Although not shown in FIG. 10, disk drives for reading from or writing to removable, non-volatile disks (e.g., "floppy disks") and optical disk drives for reading from or writing to removable, non-volatile optical disks may be provided. In these cases, each drive may be connected to a bus (not shown) via one or more data media interfaces. Memory 1020 may include computer program product 1025 having one or more program modules configured to perform various methods or actions of various embodiments of this disclosure.
[0114] The communication unit 1040 enables communication with other electronic devices via a communication medium. Additionally, the functionality of the components of the electronic device 1000 can be implemented using a single computing cluster or multiple computing machines capable of communicating via communication connections. Therefore, the electronic device 1000 can operate in a networked environment using logical connections to one or more other servers, network personal computers (PCs), or another network node.
[0115] Input device 1050 can be one or more input devices, such as a mouse, keyboard, trackball, etc. Output device 1060 can be one or more output devices, such as a monitor, speaker, printer, etc. Electronic device 1000 can also communicate with one or more external devices (not shown) via communication unit 1040 as needed. These external devices include storage devices, display devices, etc., and can communicate with one or more devices that enable user interaction with electronic device 1000, or with any device that enables electronic device 1000 to communicate with one or more other electronic devices (e.g., network card, modem, etc.). Such communication can be performed via input / output (I / O) interface (not shown).
[0116] According to an exemplary implementation of this disclosure, a computer-readable storage medium is provided that stores computer-executable instructions thereon, wherein the computer-executable instructions are executed by a processor to implement the methods described above. According to an exemplary implementation of this disclosure, a computer program product is also provided, which is tangibly stored on a non-transitory computer-readable medium and includes computer-executable instructions, which are executed by a processor to implement the methods described above.
[0117] Various aspects of this disclosure are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatuses, devices, and computer program products implemented according to this disclosure. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.
[0118] These computer-readable program instructions can be provided to a processing unit of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine such that, when executed by the processing unit of the computer or other programmable data processing apparatus, they create means for implementing the functions / actions specified in one or more blocks of the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium that causes a computer, programmable data processing apparatus, and / or other device to operate in a particular manner. Thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing aspects of the functions / actions specified in one or more blocks of the flowchart and / or block diagram.
[0119] Computer-readable program instructions can be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions that execute on the computer, other programmable data processing apparatus, or other device to perform the functions / actions specified in one or more boxes of a flowchart and / or block diagram.
[0120] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of an instruction, which contains one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.
[0121] Various implementations of this disclosure have been described above. These descriptions are exemplary and not exhaustive, nor are they limited to the disclosed implementations. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described implementations. The terminology used herein is chosen to best explain the principles, practical applications, or improvements to technology in the market, or to enable others skilled in the art to understand the various implementations disclosed herein.
Claims
1. An audio processing method, comprising: In response to the first participant joining a live interactive event associated with the target music, a first data stream is pushed to the first terminal corresponding to the first participant. The first data stream includes the target music content and the first audio content collected by the first terminal. as well as In response to the first participant being the first role in the live interactive event, a second data stream is pushed simultaneously with the first data stream, the second data stream including the first audio content.
2. The method according to claim 1, further comprising: Add timestamp information to the first data stream.
3. The method according to claim 1 or 2, further comprising: In response to the target music being played at the first terminal, collaboration information is added to the second data stream, the collaboration information indicating the playback progress and / or playback status of the target music at the first terminal.
4. The method according to any one of claims 1 to 3, further comprising: The system receives a third data stream pushed by a second terminal, the third data stream including second audio content collected by the second terminal, the second terminal being the second participant in the live interactive event.
5. The method according to any one of claims 1 to 4, wherein the first data stream corresponds to a first real-time communication channel associated with the live interactive event, the second data stream corresponds to a second real-time communication channel associated with the live interactive event, and the first real-time communication channel and the second real-time communication channel are associated with different roles in the live interactive event.
6. An audio processing method, comprising: In response to a second participant joining a live interactive event associated with the target music, a third data stream is pushed to the second terminal corresponding to the second participant. The third data stream includes the second audio content collected by the second terminal. as well as In response to the fact that the second participant is the second role in the live interactive event, the first terminal receives the second data stream pushed by the first terminal. The second data stream includes the first audio content collected by the first terminal, and the first terminal corresponds to the first participant in the live interactive event.
7. The method according to claim 6, further comprising: In response to a second participant joining a live interactive event associated with the target music, the system receives a first data stream pushed by the first terminal, the first data stream including the target music and the first audio content; as well as In response to the fact that the second participant is the second role in the live interactive event, the reception of the first data stream is stopped.
8. The method according to claim 6 or 7, wherein the second data stream further includes collaboration information indicating the playback progress and / or playback status of the target music at the first terminal, and the method further includes: Based on the collaboration information, the playback progress of the target music on the second terminal is adjusted.
9. An audio processing method, comprising: In response to a third participant joining a live interactive event associated with the target music, a first data stream pushed by a first terminal and a third data stream pushed by a second terminal are received at a third terminal corresponding to the third participant. The first data stream includes the target music and first audio content collected by the first terminal, and the third data stream includes second audio content collected by the second terminal. The first terminal corresponds to the first participant in the live interactive event, and the second terminal corresponds to the second participant in the live interactive event.
10. The method of claim 9, further comprising: In response to the fact that the third participant is the third role in the live interactive event, a fourth data stream is pushed, the fourth data stream including the third audio content collected by the third terminal.
11. The method of claim 10, wherein the first data stream further includes timestamp information, and the method further includes: Based on the timestamp information, time alignment processing is performed on the first data stream and the third data stream.
12. The method of claim 11, further comprising: The fourth data stream, along with the aligned first and second data streams, is sent to the content delivery network so that the content delivery network sends the fourth data stream, along with the aligned first and second data streams, to the fourth terminal corresponding to the live stream's viewer.
13. An audio processing apparatus, comprising: The first push module is configured to respond to a first participant joining a live interactive event associated with the target music, and to push a first data stream at the first terminal corresponding to the first participant. The first data stream includes the target music content and the first audio content collected by the first terminal. as well as The second push module is configured to respond to the first participant being the first role in the live interactive event, and to push a second data stream, which includes the first audio content, while pushing the first data stream.
14. An audio processing apparatus, comprising: The third push module is configured to respond to a second participant joining a live interactive event associated with the target music and push a third data stream at the second terminal corresponding to the second participant. The third data stream includes the second audio content collected by the second terminal. as well as The first receiving module is configured to receive a second data stream pushed by the first terminal in response to the second participant being the second role in the live interactive event. The second data stream includes first audio content collected by the first terminal, and the first terminal corresponds to the first participant in the live interactive event.
15. An audio processing apparatus, comprising: The second receiving module is configured to, in response to a third participant joining a live interactive event associated with the target music, receive a first data stream created by a first terminal and a third data stream pushed by a second terminal at a third terminal corresponding to the third participant. The first data stream includes the target music and first audio content collected by the first terminal, and the third data stream includes second audio content collected by the second terminal. The first terminal corresponds to the first participant in the live interactive event, and the second terminal corresponds to the second participant in the live interactive event.
16. An electronic device comprising: At least one processor; as well as At least one memory coupled to the at least one processor and storing instructions for execution by the at least one processor, the instructions, when executed by the at least one processor, causing the terminal device to perform the method according to any one of claims 1 to 5, the method according to any one of claims 6 to 8, or the method according to any one of claims 9 to 12.
17. A computer-readable storage medium having stored thereon computer-executable instructions that can be executed by a processor to implement the method according to any one of claims 1 to 5, the method according to any one of claims 6 to 8, or the method according to any one of claims 9 to 12.
18. A computer program product tangibly stored in a computer storage medium and comprising computer-executable instructions, which, when executed by a device, perform the method according to any one of claims 1 to 5, any one of claims 6 to 8, or any one of claims 9 to 12.