Method and apparatus for managing audio data in livestreaming room, device, and medium
By acquiring and generating personalized audio data based on the type of activity users participate in during the live stream, the problem of participating guests being unable to obtain their own audio data is solved, thereby improving user immersion and focus.
Patent Information
- Application Number
- PCT/CN2025/103145
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-06-24
- Filing Date
- 2025-06-24
- Publication Date
- 2026-01-02
AI Technical Summary
In the live broadcast room, the guests participating in the event cannot access their own audio data, which reduces immersion and interferes with their concentration.
Based on whether users participate in the activities in the live broadcast room, the user type is determined, and the corresponding anchor audio data is obtained to generate and provide personalized audio data, including anchor voice, activity audio and other user audio, in order to improve the immersion and focus of participating users.
Personalized audio data management enhances user immersion and focus while reducing the disruptive impact of broadcaster audio on users.
Smart Images

Figure CN2025103145_02012026_PF_FP_ABST
Abstract
Description
Method, device, apparatus and medium for managing audio data in a live room
[0001] The present application claims priority to the Chinese patent application No. 202410823772.9, filed on June 24, 2024, entitled “Method, device, apparatus and medium for managing audio data in a live room”, the content of which is incorporated herein by reference in its entirety. TECHNICAL FIELD
[0002] Exemplary implementations of the present disclosure generally relate to the field of computers, and in particular, to a method, device, apparatus and computer readable storage medium for managing audio data in a live room. BACKGROUND
[0003] With the development of computer technology, more and more applications can provide live streaming functions. During the live streaming process, the host user can interact with the audience user, which can attract more audience users and provide more abundant information to the audience users. For example, some live streaming applications can create interactive activities in a live room, and users can participate in the interactive activities initiated by the host side. The host side can configure the interactive activities in the live room, and the host side, the users participating in the interactive activities, and other users watching the live room can interact in audio and / or video in the live room. SUMMARY
[0004] In a first aspect of the present disclosure, a method for managing audio data in a live room is provided. The method comprises: determining a type of a first user in the live room based on whether the first user participates in an activity in the live room; obtaining host audio data of a host user in the live room based on the type of the first user; and determining audio data for providing to the first user based on the host audio data.
[0005] In a second aspect of the present disclosure, an apparatus for managing audio data in a live room is provided. The apparatus comprises: a type determination module configured to determine a type of a first user in the live room based on whether the first user participates in an activity in the live room; a data obtaining module configured to obtain host audio data of a host user in the live room based on the type of the first user; and a data generation module configured to determine audio data for providing to the first user based on the host audio data.
[0006] In a third aspect of the present disclosure, an electronic device is provided. The electronic device comprises: at least one processor; and at least one memory coupled to the at least one processor and storing instructions for execution by the at least one processor, the instructions, when executed by the at least one processor, causing the electronic device to perform the method according to the first aspect of the present disclosure.
[0007] In a fourth aspect of the present disclosure, a computer-readable storage medium is provided, having stored thereon computer-executable instructions that, when executed by a processor, cause the processor to implement the method according to the first aspect of the present disclosure.
[0008] According to a fifth aspect of the present disclosure, a computer program product is provided, comprising computer-executable instructions, wherein the computer-executable instructions, when executed by a processor, implement the method according to the first aspect of the present disclosure.
[0009] It is to be understood that the details set forth herein are not intended to limit the key or critical features of the implementations of the present disclosure or to limit the scope of the present disclosure. Other features of the present disclosure will become apparent from the following description, which, taken in conjunction with the drawings, disclose various implementations. BRIEF DESCRIPTION OF DRAWINGS
[0010] The above and other features, aspects, and advantages of various implementations of the present disclosure will become more apparent from the following detailed description taken in conjunction with the accompanying drawings. In the drawings similar or common elements of the drawings are denoted by like reference numerals, in which:
[0011] FIG. 1 shows a schematic diagram of an example environment in which embodiments of the present disclosure can be implemented;
[0012] FIG. 2 shows a schematic diagram of an example of managing audio data in a live room in the prior art;
[0013] FIG. 3 shows a flowchart of a process for managing audio data in a live room according to some embodiments of the present disclosure;
[0014] FIG. 4 shows a schematic diagram of an example of managing audio data in a live room according to some embodiments of the present disclosure;
[0015] FIG. 5 shows a schematic diagram of a process for managing audio data in a live room according to some embodiments of the present disclosure;
[0016] FIG. 6 shows a schematic structural block diagram of an apparatus for managing audio data in a live room according to some embodiments of the present disclosure; and
[0017] FIG. 7 shows a block diagram of an electronic device in which one or more embodiments of the present disclosure can be implemented. DETAILED DESCRIPTION
[0018] Embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. While certain embodiments of the present disclosure are shown in the drawings, it is understood that the present disclosure can be embodied in various forms and should not be construed as being limited to the embodiments set forth herein; rather, these embodiments are provided so that the present disclosure can be more thoroughly and completely understood. It should be understood that the drawings and embodiments of the present disclosure are only for illustrative purposes and are not intended to limit the scope of protection of the present disclosure.
[0019] In the description of embodiments of the present disclosure, the term "comprising" and its conjugations should be understood to encompass the meanings of "consisting of" and "consisting essentially of", i.e., "including but not limited to". The term "based on" should be understood as "based at least in part on". The term "one embodiment" or "the embodiment" should be understood as "at least one embodiment". The term "some embodiments" should be understood as "at least some embodiments". The terms "first", "second" and the like can refer to different or identical objects. Other explicit and implicit definitions can also be included below.
[0020] In this document, unless explicitly stated otherwise, performing a step "in response to" an event does not mean that the step is performed immediately after the event, but can include one or more intermediate steps.
[0021] It can be understood that the data involved in the technical solutions of the present disclosure (including but not limited to the data itself, the obtaining or use of the data) should comply with the requirements of relevant laws and regulations and relevant provisions.
[0022] It can be understood that before using the technical solutions disclosed in the embodiments of the present disclosure, the type of personal information involved in the present disclosure, the scope of use, the scenario of use, etc. should be informed to the user and the authorization of the user should be obtained through appropriate means according to relevant laws and regulations.
[0023] For example, in response to receiving the active request of the user, a prompt information is sent to the user to explicitly prompt the user that the operation requested to be performed will require obtaining and using the personal information of the user, so that the user can voluntarily choose whether to provide the personal information to the software or hardware such as electronic device, application program, server or storage medium, etc. performing the operation of the technical solutions of the present disclosure according to the prompt information.
[0024] As an optional but non-limiting implementation manner, in response to receiving the active request of the user, the manner of sending the prompt information to the user may, for example, be the manner of pop-up window, and the prompt information may, for example, be presented in the form of text in the pop-up window. In addition, the pop-up window may also carry selection controls for the user to select "agree" or "disagree" to provide personal information to the electronic device.
[0025] It can be understood that the above notification and user authorization obtaining process is only illustrative, and does not limit the implementation of the present disclosure, and other ways that meet relevant laws and regulations can also be applied to the implementation of the present disclosure.
[0026] Example environment
[0027] FIG. 1 shows a schematic diagram of an example environment 100 in which embodiments of the present disclosure can be implemented. In the environment 100, a user 110 can establish a live room (for example, a live room in a live application can correspond to one live room) and provide live content and the like through an associated terminal device 120. In some scenarios, the user 110 is also referred to as a host, a live party or a management party of the live room, for example. The terminal device 120 can also be referred to as a host end of the live room.
[0028] One or more users 130-1, 130-2, …, 130-N can watch the live and participate in the interaction of the live room and the like through the respective associated terminal devices 140-1, 140-2, …, 140-N. For ease of discussion, the users 130-1, 130-2, …, 130-N can be collectively referred to or individually referred to as the user 130, and the terminal devices 140-1, 140-2, …, 140-N can be collectively referred to or individually referred to as the terminal device 140. In some scenarios, the user 130 can also be referred to as a viewer, a listener, a watching party or a participating party of the live room. The terminal device 140 can also be referred to as a viewer end of the live room. Alternatively and / or additionally, the user 130 can be invited to join the live, at which time the invited user can be converted into a guest user and talk on the mic to have a conversation with the host.
[0029] It should be understood that although only a single host user is shown in FIG. 1, in some embodiments, there can be multiple host users initiating a live in a certain live room.
[0030] In some embodiments, the terminal device 120 and the terminal device 140 can respectively install an application capable of providing a live service, or can access a website capable of providing a live service. The user 110 and the user 130 can operate the terminal device 120 and the terminal device 140 to access the corresponding application or website.
[0031] Correspondingly, the terminal device 120 and the terminal device 140 can present a corresponding live interface, which can provide live content of the live room, such as audio live content or video live content and the like, for example.
[0032] In some embodiments, the terminal device 120 and the terminal device 140 can also communicate with a server 150 through a network 152 to implement the provision of the live service. The server 150 can provide functions of management, configuration and maintenance and the like with respect to the application or the website.
[0033] The terminal devices 120 and 140 can be any type of mobile terminal, fixed terminal, or portable terminal including a mobile handset, a tablet computer, a laptop computer, a notebook computer, a subnotebook computer, an ultrabook computer, a netbook computer, a smartbook, a personal digital assistant (PDA), a media player, a media recording device, a television receiver, a radio broadcast receiver, an electronic book device, a game device, or any combination thereof, including the accessories and peripherals of these devices, or any combination thereof. In some embodiments, the terminal devices 120, 140 can also be able to support any type of interface to the user (such as "wearable" circuitry, etc.).
[0034] The server 150 can be various types of computing systems / servers capable of providing computing capabilities, including but not limited to mainframes, edge computing nodes, computing devices in a cloud environment, etc. The server 130 can provide background services for applications that provide live streaming services in the terminal devices 120 and 140, for example. In some embodiments, if a host user initiates an activity in a live streaming room, the server 120 can include a first server for providing live streaming services and a second server for providing activity services. Alternatively and / or additionally, the live streaming services and the activity services can be provided by a single server.
[0035] It should be understood that the structures and functions of the various elements in the environment 100 are described for illustrative purposes only, without implying any limitation on the scope of the present disclosure. Some example embodiments of the present disclosure will be described below with continued reference to the drawings.
[0036] In a live streaming scenario, a live streaming platform often uses both RTC (real-time communication) and CDN (content distribution network) services. When a host (e.g., the user 110) is in a live streaming session with guests (e.g., some of the users 130), the live streaming session with the guests will use the RTC service to ensure low-latency real-time interaction (the data stream corresponding to the RTC service can be referred to as an RTC stream), and the RTC service will then push the live streaming session with the guests to the CDN service, and the viewers who are not in the live streaming session with the guests (e.g., some of the users 130) will still obtain the live streaming content from the CDN service (the data stream corresponding to the CDN service can be referred to as a CDN stream).
[0037] The process of providing audio data to different users is described with reference to FIG. 2, which shows a schematic diagram of an example 200 of managing audio data in a live room in the prior art. A host can invite a part of guests to participate in an activity. As shown in FIG. 2, a spectator 201 can obtain a CDN stream 210 of the host (i.e., the user 110) based on a CDN service, which includes the host voice audio data of the user 110 in the live room, the activity audio data generated by the user 110 participating in the activity, and the guest audio data of each guest user in the live room. A guest 202 who does not participate in the activity and a guest 203 who participates in the activity can both obtain an RTC stream 220 of the user 110 based on an RTC service, which also includes the host voice audio data of the user 110 in the live room, the activity audio data generated by the user 110 participating in the activity, and the guest audio data of each guest user in the live room.
[0038] Conventionally, the content of the audio data heard by the guests participating in the activity, the guests not participating in the activity, and the spectators is the same. For the guests participating in the activity, they cannot obtain the activity audio data generated by themselves participating in the activity, and can only obtain the activity audio generated by the host participating in the activity, which affects the immersion of the guests participating in the activity and interferes with the focus of the guests participating in the activity.
[0039] Summary of audio data management
[0040] In view of this, embodiments of the present disclosure propose an improved scheme for managing audio data in a live room. According to the scheme, a type of a first user in the live room is determined based on whether the first user participates in an activity in the live room. Based on the type of the first user, host audio data of a host user in the live room is obtained. Based on the host audio data, audio data for providing to the first user is determined.
[0041] For ease of description, the process of audio data management is described below only as an example of a racing game as an activity. The host user can invite the guest user to participate in the racing game, at which time the host user can drive a car and the guest user can drive a motorcycle. For the guest user participating in the racing game, the user can be provided with the host voice data and the audio data of the motorcycle, thereby improving the immersion of such a guest user for the activity and exempting from the interference from the car audio of the host. For the guest user not participating in the racing game, the user can be provided with the host voice data and the audio data of the car, thereby participating in the activity from the perspective of the host.
[0042] In this way, in embodiments of the present disclosure, the anchor audio data of the corresponding anchor user can be acquired based on the type of the guest user in the live room, and then the audio data provided to the guest user is generated based on the acquired anchor audio data. This helps to reduce the influence of the anchor user on the guest user participating in the activity and / or other anchors participating in the activity, so that the guest user and / or other anchors can focus more on the activity content. Various example implementations of this scheme are described in detail below in combination with the drawings.
[0043] Detailed process of audio data management
[0044] FIG. 3 shows a flowchart of a process 300 for managing audio data in a live room according to some embodiments of the present disclosure. The process 300 can be implemented at the server 150, alternatively and / or additionally, the terminal devices 120 and 140 can invoke the functions of the server 150 to implement the process 300. The process 300 is described below with reference to FIG. 1. It can be understood that the process 300 can be performed only in the case where the guest user is determined to be invited to participate in the activity. In the case where the anchor user does not invite the guest user to participate in the activity, the anchor speech audio data of the anchor in the live room, the activity audio data generated by the anchor participating in the activity, and the guest audio data of each guest user in the live room in the live room can be directly provided to the guest user and the audience user.
[0045] In some example embodiments of the present disclosure, the process 300 of the present disclosure can be performed at any device with data processing capability, for example, can be performed at a terminal device, can be performed at any server device in a network. Specifically, the process 300 can be performed at a real-time communication server for managing audio data of the live room.
[0046] At block 310, the type of the first user is determined based on whether the first user participates in an activity in the live room. The activity can be an activity initiated by the anchor user. The activity may, for example, include a game. The first user may, for example, be a user in the live room who is in a mic-in session with the anchor user. The type of the first user may, for example, indicate whether the first user participates in the activity.
[0047] In embodiments of the present disclosure, the first user can include a guest user (also referred to as a first guest user) in the live room. Alternatively and / or additionally, in the case of multiple anchors, the first user can include another anchor user in the live room different from the anchor user. In this way, multiple types of users can be supported to participate in the activity, thereby improving the interactivity between users. For ease of description, the implementation process will be described below only with the guest user as an example of the first user, alternatively and / or additionally, in the case of multiple anchors, the first user can include other anchor users.
[0048] At block 320, based on the type of the first user, the anchor audio data of an anchor user in the live broadcast room is acquired.
[0049] In some embodiments, if it is determined that the type of the first user indicates that the first user is participating in the activity, the first anchor audio data of the anchor user can be received. The first anchor audio data may, for example, only include anchor speech audio data (e.g., human voice data) of the anchor user in the live broadcast room. As to the specific timing of receiving the first anchor audio data, in some embodiments, the first anchor audio data can be determined to be received in response to receiving a user operation (e.g., a trigger operation for a specific operation control) from the anchor user indicating that the first anchor audio data is to be received. Alternatively or additionally, in some embodiments, the first anchor audio data can also be received automatically in response to determining that there is a guest user participating in the activity in the live broadcast room.
[0050] In some embodiments, if it is determined that the type of the first user indicates that the first user is not participating in the activity, the terminal device 120 can receive second anchor audio data of the anchor user. This second anchor audio data may, for example, include anchor speech audio data of the anchor user in the live broadcast room, as well as anchor activity audio data generated by the anchor user participating in the activity. Taking the activity as a game as an example, the anchor activity audio data may, for example, include audio data of actions performed by the anchor user in the game (also referred to as corresponding audio data of the anchor user in the game) and / or background audio data in the game. In the above example of the racing game, the anchor activity audio data may, for example, include car engine audio generated by the anchor user driving the car, background music of the game, and the like. For another example, if the game is a chess game, the anchor activity audio data may, for example, include audio generated by the anchor user playing the game, dialogue of a virtual character corresponding to the anchor user in the game, background music of the game, and the like.
[0051] At block 330, based on the anchor audio data, audio data for providing to the first user is determined.
[0052] In some embodiments, the host audio data can be provided directly to the first user. Alternatively and / or additionally, audio data provided to the first user can be generated based on the host audio data and other audio data. For the first user participating in the activity, first activity audio data generated by the first user participating in the activity can also be obtained, and audio data can be generated based on the first activity audio data and the host voice audio data. Similar to the host activity audio data, the first activity audio data can include, for example, audio data of actions performed by the first user in the game (i.e., corresponding to the guest user in the game) and / or background audio data in the game. For example, in the above example of the racing game, the first activity audio data can include motorcycle engine audio generated by the first user driving the motorcycle, background music of the game, etc. For another example, if the game is a chess game, the first activity audio data can include audio generated by the first user playing the game, dialogues of the virtual character corresponding to the first user in the game, background music of the game, etc.
[0053] In some embodiments, for the first user not participating in the activity, audio data can be generated based on the received user audio data of the second user in the live room. Here, the user audio data of the second user in the live room can include voice audio data of the second user in the live room. Alternatively and / or additionally, the user audio data of the second user in the live room can include voice audio data and activity audio data of the second user in the live room.
[0054] In some embodiments, if there is a second user in the live room in addition to the first user (whether or not participating in the activity), user audio data of the second user in the live room (e.g., voice of the second user speaking in the live room) can be obtained, and audio data provided to the first user can be generated based on the user audio data of the second user in the live room. The second user here is a user who can speak in the live room, including, for example, at least one of the following: other guest users (whether or not participating in the activity) on the mic, and other hosts (whether or not participating in the activity) in a multi-host scenario.
[0055] In some examples, an audio receiving mode of the first user can be obtained. Here, the audio receiving mode can specify whether the first user allows to receive audio data from other users (i.e., second users) other than the host user. Specifically, the do-not-disturb mode can specify that the first user does not allow to receive audio data from other users, and the regular mode can specify that the first user allows to receive audio data from other users.
[0056] In determining the audio data provided to the first user, the audio data provided to the first user can be generated based on the audio receiving mode. Specifically, in the regular mode, in response to determining that the audio receiving mode indicates that audio data from users other than the anchor user is allowed to be received, user audio data of the second user in the live room can also be acquired. Further, the audio data provided to the first user can be generated based on the user audio data of the second user in the live room. At this time, the first user can hear the sound from the anchor and other users in the live room (including the guest and / or other anchors). Alternatively and / or additionally, in the do-not-disturb mode, the voice audio data from the second user can be excluded, at this time the first user only receives the audio data from the anchor, thereby the interference of the other users to the first user can be reduced.
[0057] Exemplarily, for the first user participating in the activity, the audio data provided to the first user participating in the activity can be generated based on the first anchor audio data of the anchor user, the first activity audio data generated by the first user participating in the activity, and the user audio data of the second user in the live room. That is, in the case of the racing game, the guest user participating in the activity can hear the anchor's voice, the audio of driving the motorcycle by himself, and the voice of other guests. For the first user not participating in the activity, the audio data provided to the first user not participating in the activity can be generated based on the second anchor audio data of the anchor user and the user audio data of the second user in the live room. That is, in the case of the racing game, the guest user not participating in the activity can hear the anchor's voice, the audio of driving the car by the anchor, and the voice of other guests.
[0058] In some embodiments, the guest audio data of each guest user in the live room can be acquired, and the guest audio data is provided to the users in the live room (including the anchor user, the guest user and / or the audience user in the live room). Exemplarily, for the audience user, the second anchor audio data of the anchor user and the guest audio data of each guest user in the live room (including the guest user participating in the activity and the guest user not participating in the activity) in the live room can be received, and the audio data for the audience user can be generated based on the second anchor audio data and the guest audio data. In some embodiments, to avoid the guest hearing the guest audio data from himself, for each guest (whether participating in the activity or not), the guest audio data of other guests other than the guest can be provided to the guest.
[0059] FIG. 4 shows a schematic diagram of an example 400 for managing audio data in a live room according to some embodiments of the present disclosure. As shown in FIG. 4, two RTC streams (i.e., RTC stream 410 and RTC stream 420) can be acquired for the host audio data from the host (i.e., user 110). The RTC stream 420 including the first host audio data (the first host audio data only including the host speech audio data of the user 110 in the live room) can be provided to the guest 203 participating in the activity. The RTC stream 410 including the second host audio data (the second host audio data including the host speech audio data of the user 110 in the live room, and the host activity audio data generated by the user 110 participating in the activity) can be provided to the guest 202 not participating in the activity. The RTC stream 410 and the RTC stream from the guest user (including the guest audio data of the guest user) can be CDN merged to obtain the CDN stream 210, which is provided to the spectator user 201 in the live room. Thus, the host activity audio data generated by the user 110 participating in the activity can not be heard by the guest participating in the activity, and the impact of the host activity audio data on the activity of the user 110 can be avoided.
[0060] FIG. 5 shows a schematic diagram of a process 500 for managing audio data in a live room according to some embodiments of the present disclosure. As shown in FIG. 5, the terminal device 120 corresponding to the host user, the terminal device 510 corresponding to the guest participating in the activity, and the terminal device 520 corresponding to the guest not participating in the live activity can be connected to the server 510, which can correspond to the server 150 in FIG. 1. The server 510 can include an activity server for providing background services for the activity. The terminal device 120 and the terminal device 510 can include cloud activity containers (i.e., cloud activity container 501 and cloud activity container 503), which can acquire an activity page from the activity server and present it at the corresponding terminal device. The server 510 can further include a live server for providing background services for the live mic-in 502-1.
[0061] When the terminal device 120, the terminal device 510, and the terminal device 520 are in the live streaming and live streaming microphone connection 502-1, 502-2, and 502-3 (collectively referred to as 502) state respectively, the terminal device 120 can obtain two RTC streams for the anchor user, one of which includes the anchor speech audio data of the anchor user in the live streaming room and the anchor activity audio data (i.e., second anchor audio data) generated by the anchor user participating in the activity, and the other of which includes the anchor speech audio data (i.e., first anchor audio data) of the anchor user in the live streaming room. The terminal device 510 can obtain an RTC stream of a guest participating in the activity, which includes the guest audio data of the guest participating in the activity in the live streaming room. The terminal device 520 can obtain an RTC stream of a guest not participating in the activity, which includes the guest audio data of the guest not participating in the activity in the live streaming room. It can be understood that there can be multiple guests participating in the activity and / or multiple guests not participating in the activity, and therefore there can be multiple terminal devices 510 and multiple terminal devices 520, which are not limited by the present disclosure.
[0062] The RTC system 530 can obtain the above-mentioned multiple RTC streams, and combine the RTC stream including the guest audio data of the guest participating in the activity in the live streaming room and the RTC stream including the guest audio data of the guest not participating in the activity in the live streaming room, and provide them to the anchor user (i.e., send them to the terminal device 120). The RTC system 530 can combine the RTC stream including the first anchor audio data and the RTC stream including the guest audio data of the other guests in the live streaming room except itself, and provide them to the guest participating in the activity (i.e., send them to the terminal device 510). The RTC system 530 can combine the RTC stream including the second anchor audio data and the RTC stream including the guest audio data of the other guests in the live streaming room except itself, and provide them to the guest participating in the activity (i.e., send them to the terminal device 520). The RTC system 530 can combine the RTC stream including the second anchor audio data, the RTC stream including the guest audio data of the guest participating in the activity in the live streaming room, and the RTC stream including the guest audio data of the guest not participating in the activity in the live streaming room, and provide them to the audience 201 in the form of the CDN stream 210.
[0063] In summary, according to the embodiments of the present disclosure, the anchor audio data of the corresponding anchor user can be obtained based on the type of the guest user in the live streaming room, and then the audio data provided to different types of guest users is generated based on the obtained anchor audio data. This helps to reduce the impact of the anchor activity audio data generated by the anchor user participating in the activity on the guest user participating in the activity, and thus the guest user can focus more on the activity content.
[0064] Example apparatus and devices
[0065] Embodiments of the present disclosure also provide a corresponding apparatus for implementing the above method or process. FIG. 6 shows a schematic structural block diagram of an apparatus 600 for managing audio data in a live room, according to certain embodiments of the present disclosure. Various modules / components in the apparatus 600 can be implemented by hardware, software, firmware, or any combination thereof.
[0066] As shown, the apparatus 600 includes a type determining module 610 configured to determine a type of a first user in a live room based on whether the first user participates in an activity in the live room. The apparatus 600 also includes a data obtaining module 620 configured to obtain anchor audio data of an anchor user in the live room based on the type of the first user. The apparatus 600 further includes a data generating module 630 configured to determine audio data for providing to the first user based on the anchor audio data.
[0067] In some embodiments, the data obtaining module 620 includes a first anchor audio data receiving module configured to, in response to determining that the type of the first user indicates that the first user participates in the activity, receive first anchor audio data of the anchor user, the first anchor audio data including anchor speech audio data of the anchor user in the live room.
[0068] In some embodiments, the data generating module 630 further includes a first activity audio data obtaining module configured to obtain first activity audio data generated by the first user participating in the activity, and a first data generating module configured to generate the audio data for providing to the first user based on the first anchor audio data and the first activity audio data.
[0069] In some embodiments, the data generating module 630 is further configured to, in response to determining that there is a second user in the live room, obtain user audio data of the second user in the live room, and generate the audio data for providing to the first user based on the user audio data of the second user in the live room.
[0070] In some embodiments, the apparatus 600 further includes a guest audio data obtaining module configured to obtain guest audio data of each guest user in the live room, and a guest audio data providing module configured to provide the guest audio data to a user in the live room, the user in the live room including at least any one of the following: the anchor user in the live room, the guest user, and a spectator user.
[0071] In some embodiments, the data obtaining module 620 includes a second anchor audio data receiving module configured to, in response to determining that the type of the first user indicates that the first user is not participating in the activity, receive second anchor audio data of the anchor user, the second anchor audio data including anchor speech audio data of the anchor user in the live broadcast room and anchor activity audio data generated by the anchor user participating in the activity.
[0072] In some embodiments, the activity is a game, and the anchor activity audio data includes at least one of the following: audio data corresponding to the anchor user in the game, and background audio data in the game.
[0073] In some embodiments, the apparatus 600 further includes a second anchor audio data providing module configured to provide the second anchor audio data to the spectator user in the live broadcast room.
[0074] In some embodiments, the apparatus 600 further includes an executing module configured to, in response to determining that the guest user is invited to participate in the activity, implement the apparatus at a real-time communication server for managing audio data of the live broadcast room.
[0075] In some embodiments, the apparatus 600 further includes an audio data providing module configured to, in response to determining that no guest user participates in the activity, provide the second anchor audio data to a user in the live broadcast room, the user in the live broadcast room including at least one of the following: the guest user in the live broadcast room, and the spectator user.
[0076] In some embodiments, the first user and the second user include at least one of the following: the guest user in the live broadcast room, and another anchor user in the live broadcast room different from the anchor user.
[0077] The units and / or modules included in the apparatus 600 can be implemented in various ways, including software, hardware, firmware, or any combination thereof. In some embodiments, one or more units and / or modules can be implemented using software and / or firmware, e.g., machine executable instructions stored on a storage medium. In addition to or alternatively, some or all of the units and / or modules in the apparatus 600 can be implemented at least partially by one or more hardware logic components. As an example and not by way of limitation, example types of hardware logic components that can be used include field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), system-on-a-chip (SOCs), complex programmable logic devices (CPLDs), etc.
[0078] It should be understood that one or more steps in the above methods can be performed by a suitable electronic device or combination of electronic devices. Such electronic device or combination of electronic devices may, for example, include the server 150, the terminal device 120 and / or the terminal device 140 in FIG. 1.
[0079] FIG. 7 illustrates a block diagram of an electronic device 700 in which one or more embodiments of the present disclosure can be implemented. It should be understood that the electronic device 700 illustrated in FIG. 7 is merely exemplary and should not be construed as limiting the functionality and scope of the embodiments described herein. The electronic device 700 illustrated in FIG. 7 can be used to implement the server 150, the terminal device 120 and / or the terminal device 140 of FIG. 1.
[0080] As illustrated in FIG. 7, the electronic device 700 is in the form of a general electronic device. The components of the electronic device 700 can include, but are not limited to, one or more processors or processing units 710, a memory 720, a storage device 730, one or more communication units 740, one or more input devices 750, and one or more output devices 760. The processing unit 710 can be a real or virtual processor and is capable of performing various processing according to programs stored in the memory 720. In a multi-processor system, multiple processing units perform computer-executable instructions in parallel to improve the parallel processing capability of the electronic device 700.
[0081] The electronic device 700 typically includes a number of computer storage media. Such media can be any available media that is accessible by the electronic device 700 and includes both volatile and non-volatile media, removable and non-removable media. The memory 720 can be a volatile memory (e.g., registers, cache, random access memory (RAM)), a non-volatile memory (e.g., read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory), or some combination thereof. The storage device 730 can be a removable or non-removable medium and can include machine-readable media such as a flash drive, a magnetic disk drive, or any other medium that can be used to store information and / or data and that can be accessed by the electronic device 700.
[0082] The electronic device 700 can further include additional detachable / non-detachable, volatile / non-volatile storage media. Although not shown in FIG. 7, a disk drive for reading from or writing to a detachable, non-volatile magnetic disk (e.g., a "floppy disk"), and an optical disk drive for reading from or writing to a detachable, non-volatile optical disk (e.g., a CD-ROM) can be provided. In these cases, each drive can be connected to the bus (not shown) by one or more data media interfaces. The memory 720 can include a computer program product 725 having one or more program modules configured to carry out the various methods or acts of the various embodiments of the present disclosure.
[0083] The communication unit 740 enables communication with other electronic devices through communication media. Additionally, the functionality of the components of the electronic device 700 can be implemented in a single computing cluster or a plurality of computer machines capable of communicating over a communication connection. As such, the electronic device 700 can operate in a networked environment using logical connections to one or more other servers, network personal computers (PCs), or another network nodes in the case of a distributed system environment.
[0084] The input device 750 can be one or more input devices, such as a mouse, a keyboard, a trackball, etc. The output device 760 can be one or more output devices, such as a display, a speaker, a printer, etc. The electronic device 700 can also communicate with one or more external devices (not shown) such as a storage device, a display device, etc., through the communication unit 740, as needed, with one or more devices that enable a user to interact with the electronic device 700, or with any devices (e.g., a network card, a modem, etc.) that enable the electronic device 700 to communicate with one or more other electronic devices. Such communication can be carried out via an input / output (I / O) interface (not shown).
[0085] According to an example implementation of the present disclosure, a computer readable storage medium having computer executable instructions stored thereon is provided, where the computer executable instructions are executed by a processor to implement the method described above. According to an example implementation of the present disclosure, a computer program product is also provided, which is tangibly stored on a non-transitory computer readable medium and includes computer executable instructions, where the computer executable instructions are executed by a processor to implement the method described above.
[0086] Various aspects of the disclosure are now described with reference to the drawings. In general, the drawings described herein relate to a method, apparatus, device, and computer program product implemented in accordance with the present disclosure. It should be understood that each block of the flowchart diagrams and / or block diagrams, and combinations of blocks in the flowchart diagrams and / or block diagrams, can be implemented by computer readable program instructions.
[0087] The computer readable program instructions can be provided to a processor of a general purpose computer, special purpose computer, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create means for implementing the functions / acts specified in the flowchart and / or block diagram block or blocks. These computer readable program instructions can also be stored in a computer readable storage medium that can include a non-transitory computer readable storage medium that can be a computer- readable storage medium having no data, programs, program modules, e.g., instructions for operation, or digital content stored thereon or therein for a short time or not at all. The computer readable storage medium can also have other meanings inhered thereby, which can include a computer- readable storage medium encoding computationally-precise instructions and / or having a fixed, distinct, non-heuristic structure.
[0088] The computer readable program instructions can also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable apparatus or other device to produce a computer implemented process such that the instructions which execute on the computer or other programmable apparatus provide processes for implementing the functions / acts specified in the flowchart and / or block diagram block or blocks.
[0089] The computer readable program instructions can also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable apparatus or other device to produce a computer implemented process such that the instructions which execute on the computer or other programmable apparatus provide processes for implementing the functions / acts specified in the flowchart and / or block diagram block or blocks.
[0090] The implementations of the disclosure have been described above with the intent to be illustrative rather than limiting. Although being shown and described in terms of certain implementations and overall functions, the implementations can be variously configured as forms having several advantages. Modifications and variations are capable of being made by persons of ordinary skill in the art without departing from the scope and spirit of the implementations. It is therefore intended that such modifications and variations be included within the scope of the implementations. The above descriptions are intended to be illustrative and not restrictive. Many modifications and variations of the implementations described herein will be apparent to those skilled in the art, in light of the before described teachings. It is, therefore, to be understood that changes can be made in the form, implementations, and details of the methods and apparatus described herein without departing from the scope and spirit of the implementations. It is therefore intended that the disclosed implementations be considered in all respects as illustrative and not restrictive, and that reference be made to the appended claims and their equivalents in determining the scope of the implementations.
Claims
1. A method for managing audio data in a live streaming room, comprising: The type of the first user is determined based on whether the first user in the live stream participates in the activities in the live stream. Based on the type of the first user, obtain the anchor audio data of the anchor user in the live broadcast room; as well as Based on the broadcaster's audio data, determine the audio data to be provided to the first user.
2. The method according to claim 1, wherein obtaining the broadcaster audio data based on the type of the first user includes: In response to determining the type of the first user indicating that the first user should participate in the activity, the first broadcaster audio data of the broadcaster user is received, the first broadcaster audio data including the broadcaster user's broadcaster voice audio data in the live broadcast room.
3. The method of claim 2, wherein determining the audio data to be provided to the first user based on the broadcaster's audio data further comprises: Obtain the first activity audio data generated by the first user participating in the activity; as well as Based on the first broadcaster's audio data and the first activity's audio data, the audio data to be provided to the first user is generated.
4. The method of claim 1 or 2, wherein determining the audio data to be provided to the first user further comprises: In response to determining that a second user exists in the live stream, the user audio data of the second user in the live stream is obtained; as well as Based on the user audio data of the second user in the live broadcast room, the audio data provided to the first user is generated.
5. The method according to claim 4, wherein generating the audio data provided to the first user based on the user audio data of the second user in the live broadcast room comprises: Obtain the audio receiving mode of the first user; as well as In response to determining that the audio reception mode indicates permission to receive audio data from users other than the broadcaster user, the audio data provided to the first user is generated based on the user audio data of the second user in the live broadcast room.
6. The method of claim 1, further comprising: Obtain the guest audio data of each guest user in the live broadcast room; as well as The guest audio data is provided to users in the live broadcast room, and the users in the live broadcast room include at least one of the following: the host user, the guest user, and the viewer user in the live broadcast room.
7. The method according to claim 1, wherein obtaining the broadcaster audio data based on the type of the first user comprises: In response to determining that the first user's type indicates that the first user did not participate in the activity, the system receives the second broadcaster audio data of the broadcaster user, the second broadcaster audio data including the broadcaster user's broadcaster voice audio data in the live broadcast room, and the broadcaster activity audio data generated by the broadcaster user participating in the activity.
8. The method according to claim 7, wherein the activity is a game, and the broadcaster activity audio data includes at least one of the following: audio data corresponding to the broadcaster user in the game, and background audio data in the game.
9. The method of claim 7, further comprising: The second anchor's audio data is provided to the viewers in the live broadcast room.
10. The method of claim 1, further comprising: In response to determining that the guest user has been invited to participate in the event, the method is executed on a real-time communication server used to manage the audio data of the live stream.
11. The method of claim 7, further comprising: In response to determining that no guest user is participating in the activity, the second anchor audio data is provided to the users in the live broadcast room, the users in the live broadcast room including at least one of the following: guest users in the live broadcast room and audience users.
12. The method according to claim 4, wherein the first user and the second user include at least one of the following: a guest user in the live broadcast room, and another host user in the live broadcast room who is different from the host user.
13. An apparatus for managing audio data in a live streaming room, comprising: The type determination module is configured to determine the type of the first user based on whether the first user in the live broadcast room participates in the activities of the live broadcast room; The data acquisition module is configured to acquire the anchor audio data of the anchor user in the live broadcast room based on the type of the first user; as well as The data generation module is configured to determine audio data to be provided to the first user based on the broadcaster's audio data.
14. An electronic device, comprising: At least one processor; as well as At least one memory coupled to the at least one processor and storing instructions for execution by the at least one processor, the instructions causing the electronic device to perform the method according to any one of claims 1 to 12 when executed by the at least one processor.
15. A computer-readable storage medium having stored thereon computer-executable instructions that can be executed by a processor to implement the method according to any one of claims 1 to 12.
16. A computer program product comprising computer-executable instructions, wherein the computer-executable instructions, when executed by a processor, implement the method according to any one of claims 1 to 12.
Citation Information
Patent Citations
Cloud game live broadcast method, cloud game server and computer readable storage medium
CN111818004A
Audio merging method, audio uploading method, equipment and program product
CN113542792A
Multi-person microphone connection method and related equipment
CN115065829A
Audio analysis and accessibility across applications and platforms
CN115733826A
Live broadcast stream control method and device of live broadcast room, equipment and medium
CN116828222A