Audio processing method and apparatus, device, and storage medium

By acquiring and processing audio stream evaluation information from live interactive events, and determining the processing mode, the problem of poor merging quality and high cost caused by uneven audio stream quality in live interactive events is solved, achieving higher quality audio stream merging and cost optimization.

WO2026067269A1PCT designated stage Publication Date: 2026-04-02BEIJING ZITIAO NETWORK TECH CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-09-19
Publication Date
2026-04-02

AI Technical Summary

Technical Problem

In live interactive events, the audio stream quality of multiple participants varies, resulting in poor audio stream quality after merging, and the transmission cost and performance overhead of merging and alignment operations are relatively large.

Method used

By acquiring audio streams from multiple terminal devices, determining their evaluation information, and determining the processing mode based on the evaluation information, the audio data of multiple audio streams are processed to improve the quality of the merged audio stream and reduce transmission costs and performance overhead of the merge alignment operation.

Benefits of technology

It improves the quality of the merged audio stream, reduces the transmission cost of the audio stream, and reduces the performance overhead of the merge alignment operation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2025122683_02042026_PF_FP_ABST
    Figure CN2025122683_02042026_PF_FP_ABST
Patent Text Reader

Abstract

The embodiments of the present disclosure provide an audio processing method and apparatus, a device, and a storage medium. The method comprises: acquiring multiple audio streams from multiple terminal devices, the multiple audio streams corresponding to multiple participants in a live interactive event; determining evaluation information associated with at least one audio stream among the multiple audio streams; at least on the basis of the evaluation information of the at least one audio stream, determining a processing mode for the multiple audio streams; and on the basis of the processing mode, processing audio data associated with the multiple audio streams. Thus, in one aspect, the quality of a merged audio stream can be improved, and in another aspect, the transmission cost of the audio stream and the performance overhead of a merging alignment operation can be reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Method, apparatus, device and storage medium for audio processing

[0001] The present application claims priority from the Chinese Patent Application No. 202411337363.4 filed on September 24, 2024 and entitled "Method, apparatus, device and storage medium for audio processing", the content of which is incorporated herein by reference in its entirety. TECHNICAL FIELD

[0002] Example embodiments of the present disclosure generally relate to the field of computers, and in particular, to a method, apparatus, device, computer readable storage medium and computer program product for audio processing. BACKGROUND

[0003] In a live interactive event, multiple audio streams corresponding to multiple participants need to be merged. However, the quality of the multiple audio streams corresponding to the multiple participants is different, which results in poor quality of the merged audio stream. SUMMARY

[0004] In a first aspect of the present disclosure, a method for audio processing is provided, comprising: obtaining multiple audio streams from multiple terminal devices, the multiple audio streams corresponding to multiple participants of a live interactive event; determining evaluation information associated with at least one audio stream of the multiple audio streams; determining a processing mode for the multiple audio streams based at least on the evaluation information of the at least one audio stream; and processing audio data associated with the multiple audio streams based on the processing mode.

[0005] In a second aspect of the present disclosure, an apparatus for audio processing is provided, comprising: an obtaining module configured to obtain multiple audio streams from multiple terminal devices, the multiple audio streams corresponding to multiple participants of a live interactive event; a first determining module configured to determine evaluation information associated with at least one audio stream of the multiple audio streams; a second determining module configured to determine a processing mode for the multiple audio streams based at least on the evaluation information of the at least one audio stream; and a processing module configured to process audio data associated with the multiple audio streams based on the processing mode.

[0006] In a third aspect of the present disclosure, an electronic device is provided. The device includes at least one processing unit; and at least one memory coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit. The instructions, when executed by the at least one processing unit, cause the device to perform the method of the first aspect.

[0007] In a fourth aspect of the present disclosure, a computer-readable storage medium is provided. The computer-readable storage medium has stored thereon a computer program, the computer program being executable by a processor to implement the method of the first aspect.

[0008] In a fifth aspect of the present disclosure, a computer program product is provided, comprising computer executable instructions, wherein the computer executable instructions, when executed by a processor, implement the method of the first aspect.

[0009] It should be understood that the content described in this part is not intended to limit the key features or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become apparent from the following description. BRIEF DESCRIPTION OF DRAWINGS

[0010] The above and other features, advantages, and aspects of embodiments of the present disclosure will become more apparent by describing in detail some embodiments thereof with reference to the attached drawings in which:

[0011] FIG. 1 shows a schematic diagram of an example environment in which embodiments of the present disclosure can be implemented;

[0012] FIG. 2 shows a flowchart of a process for audio processing according to some embodiments of the present disclosure;

[0013] FIG. 3 shows a schematic diagram of a human sound amount and score relationship curve according to some embodiments of the present disclosure;

[0014] FIG. 4 shows a block diagram of an apparatus for audio processing according to some embodiments of the present disclosure; and

[0015] FIG. 5 shows a block diagram of a device capable of implementing a plurality of embodiments of the present disclosure. DETAILED DESCRIPTION

[0016] It can be understood that, before using the technical solutions disclosed in the embodiments of the present disclosure, the type of personal information involved in the present disclosure, the scope of use, the scene of use, etc. should be informed to the user and the authorization of the user should be obtained through appropriate means according to relevant laws and regulations.

[0017] For example, in response to receiving the user's active request, a prompt message is sent to the user to explicitly prompt the user that the operation requested to be performed will require the acquisition and use of the user's personal information. Thus, the user can voluntarily choose whether to provide personal information to the electronic device, application program, server or storage medium, etc. software or hardware that performs the operation of the technical solutions of the present disclosure according to the prompt message.

[0018] As an optional but non-limiting implementation, in response to receiving the active request of the user, the manner of sending the prompt information to the user may be, for example, a pop-up window manner, in which the prompt information may be presented in the form of text. In addition, the pop-up window may also carry a selection control for the user to select "agree" or "disagree" to provide personal information to the electronic device.

[0019] It can be understood that the above notification and user authorization obtaining process is only illustrative, and does not limit the implementation of the present disclosure, and other manners meeting the relevant laws and regulations can also be applied to the implementation of the present disclosure.

[0020] It can be understood that the data involved in the technical solution (including but not limited to the data itself, the acquisition or use of the data) should comply with the requirements of the relevant laws and regulations and the relevant provisions.

[0021] Embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although certain embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms, and should not be interpreted as being limited to the embodiments set forth herein, rather, these embodiments are provided to make the present disclosure more thorough and complete. It should be understood that the drawings and embodiments of the present disclosure are only for exemplary purposes, and are not intended to limit the scope of protection of the present disclosure.

[0022] It should be noted that the titles of any section / subsection provided herein are not limiting. Various embodiments are described throughout this document, and any type of embodiment can be included under any section / subsection. Furthermore, embodiments described in any section / subsection can be combined with any other embodiment described in the same section / subsection and / or a different section / subsection in any manner.

[0023] In this document, unless explicitly stated otherwise, performing a step "in response to A" does not mean that the step is performed immediately after A, but can include one or more intermediate steps.

[0024] In the description of embodiments of the present disclosure, the term "comprising" and similar terms are to be interpreted as open-ended, i.e., "including but not limited to". The term "based on" is to be interpreted as "based, at least in part, on". The term "one embodiment" or "the embodiment" is to be interpreted as "at least one embodiment". The term "some embodiments" is to be interpreted as "at least some embodiments". Other explicit and implicit definitions can also be included below. The terms "first", "second", etc. can refer to different or the same objects. Other explicit and implicit definitions can also be included below.

[0025] As used herein, the term “model” can learn an association between a corresponding input and output from training data, so that after training is completed, a corresponding output can be generated for a given input. The generation of a model can be based on a machine learning technique. Deep learning is a machine learning algorithm that processes an input and provides a corresponding output by using multiple layers of processing units. In this document, a “model” can also be referred to as a “machine learning model”, a “machine learning network” or a “network”, which terms are used interchangeably herein. A model can in turn comprise different types of processing units or networks.

[0026] As used herein, a “unit”, “operational unit” or “sub-unit” can be composed of any suitable structure of machine learning model or network. As used herein, a group of elements or similar expression can comprise one or more such elements. For example, a “group of convolution units” can comprise one or more convolution units.

[0027] As briefly mentioned above, in a live interactive event, a plurality of audio streams corresponding to a plurality of participants need to be merged. However, due to the different qualities of the plurality of audio streams corresponding to the plurality of participants, the quality of the merged audio stream is poor.

[0028] For example, in a multi-person singing scenario, on the one hand, merging audio streams of different qualities will result in poor effects for the audience. On the other hand, for the merged audio stream, the audience can only distinguish a limited number (e.g. 2-3) of voices in the case of voice alignment. In the case of more than this number, from the perspective of the merging effect, no matter how many audio streams there are, the effect experienced by the audience is consistent or very small. However, the more audio streams there are, the higher the transmission cost and the greater the performance overhead of the merging and alignment operation.

[0029] Embodiments of the present disclosure provide a scheme for audio processing. According to various embodiments of the present disclosure, a plurality of audio streams are first obtained from a plurality of terminal devices, evaluation information associated with at least one of the plurality of audio streams is then determined, and a processing mode for the plurality of audio streams is determined based on at least the evaluation information of the at least one audio stream. Finally, audio data associated with the plurality of audio streams is processed based on the processing mode. In this way, embodiments of the present disclosure can improve the quality of the merged audio stream on the one hand, and reduce the transmission cost of the audio stream and the performance overhead of the merging and alignment operation on the other hand.

[0030] FIG. 1 illustrates a schematic diagram of an example environment 100 in which embodiments of the present disclosure can be implemented. As shown in FIG. 1, the environment 100 can include terminal devices 110-1, 110-2, …, 110-N and a server 120. For ease of discussion, the terminal devices 110-1, 110-2, …, 110-N can be collectively or individually referred to as terminal devices 110.

[0031] In the environment 100, the terminal devices 110 can run an application that supports voice interaction, which can be any suitable type of application for voice interaction, examples of which can include, but are not limited to, a video application, a social application, a live broadcast application, or other suitable application, based on which a user can perform a live broadcast related activity. In some embodiments, the live broadcast related activity can be an activity of initiating a live broadcast, an activity of participating in a live broadcast, an activity of watching a live broadcast, and the like. The user can interact with the application via the terminal device 110 and / or an attached device thereof.

[0032] In the environment 100, the terminal devices 110 can be any type of computing-enabled device, including a terminal device or a server device. The terminal device can be any type of mobile terminal, fixed terminal, or portable terminal including a mobile phone, a desktop computer, a laptop computer, a notebook computer, a netbook computer, a tablet computer, a media computer, a multimedia tablet, a personal communication system (PCS) device, a personal navigation device, a personal digital assistant (PDA), an audio / video player, a digital camera / camcorder, a positioning device, a television receiver, a radio broadcast receiver, an electronic book device, a game device, or any combination thereof, including accessories and peripherals of these devices or any combination thereof. The server device may, for example, include a computing system / server, such as a mainframe, an edge computing node, an electronic device in a cloud environment, and the like.

[0033] The server 120 can be a standalone physical server, a server cluster or a distributed system composed of multiple physical servers, or a cloud server providing cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content distribution networks, and basic cloud computing services such as big data and artificial intelligence platforms. The server 120 may, for example, include a computing system / server, such as a mainframe, an edge computing node, a computing device in a cloud environment, and the like. The server 120 can provide background services for the application in the terminal devices 110 that supports voice interaction.

[0034] A communication connection can be established between the server 120 and the terminal devices 110, and a communication connection can also be established between the terminal devices 110. The communication connection can be established in a wired manner or a wireless manner. The communication connection can include, but is not limited to, a Bluetooth connection, a mobile network connection, a Universal Serial Bus (USB) connection, a Wireless Fidelity (WiFi) connection, and the like, and embodiments of the present disclosure are not limited in this regard.

[0035] It should be understood that the structure and function of the environment 100 are described for illustrative purposes only, and do not imply any limitation on the scope of the present disclosure.

[0036] Some example embodiments of the present disclosure will be described below with reference to the accompanying drawings.

[0037] FIG. 2 shows a flowchart of a process 200 for audio processing according to some embodiments of the present disclosure. The following will be described taking an example that a server performs the interactive process 200. Next, the scheme provided by the present disclosure for audio processing will be described in detail in combination with FIG. 1.

[0038] At block 210, the server 120 acquires a plurality of audio streams from a plurality of terminal devices 110, the plurality of audio streams corresponding to a plurality of participants of a live interactive event.

[0039] The plurality of terminal devices 110 correspond to the plurality of participants of the live interactive event, that is, one terminal device 110 corresponds to one participant of the live interactive event. Each audio stream can be acquired by the corresponding terminal device 110 and transmitted to the server 120, or can be preprocessed by the corresponding terminal device 110 after being acquired and then transmitted to the server 120. For example, the preprocessing of the acquired audio stream by the terminal device 110 can be noise reduction, echo cancellation, and the like.

[0040] The live interactive event may, for example, be a live event that requires collaboration of a plurality of participants, such as a multi-person chorus, a multi-person poetry recitation, and the like. The server 120 can perform stream merging on the plurality of audio streams acquired from the plurality of terminal devices 110 to complete the live event that requires collaboration of a plurality of participants. For example, when the live interactive event is a multi-person chorus, the server 120 needs to perform stream merging on the plurality of audio streams acquired from the plurality of terminal devices 110 to present to the audience.

[0041] At block 220, the server 120 determines evaluation information associated with at least one of the plurality of audio streams.

[0042] The evaluation information can be used to evaluate the quality of the audio stream. For example, the higher the quality of the audio stream, the better the effect of the merged audio stream after the audio stream is merged; the lower the quality of the audio stream, the worse the effect of the merged audio stream after the audio stream is merged.

[0043] In some embodiments, the server 120 can obtain a set of evaluation parameters from at least one terminal device 110 corresponding to at least one audio stream, and determine the evaluation information associated with the at least one audio stream based on the set of evaluation parameters.

[0044] For example, each terminal device 110 of the at least one terminal device 110 can send the evaluation parameters to the server 120 at a predetermined frequency, and the server 120 can determine the evaluation information associated with the corresponding audio stream based on the evaluation parameters sent by each terminal device 110 each time. It should be understood that the process of determining the evaluation information associated with the corresponding audio stream by the server 120 each time should be completed before the next time the server 120 receives the evaluation parameters sent by the terminal device 110, so that the server 120 can evaluate the audio stream sent by the terminal device 110 in real time.

[0045] In some embodiments, the terminal device 110 can send the evaluation parameters to the server 120 in a long link manner.

[0046] The frequency at which the terminal device 110 sends the evaluation parameters to the server 120 can be set according to actual needs, and the embodiments of the present disclosure do not limit it. For example, in a multi-person singing scenario, the terminal device 110 can send the evaluation parameters to the server 120 once every n lyrics, or the terminal device 110 can send the evaluation parameters to the server 120 once every predetermined time period.

[0047] In some embodiments, the server 120 can also determine the evaluation parameters corresponding to the at least one audio stream respectively based on the at least one audio stream at a predetermined frequency, and determine the evaluation information associated with the at least one audio stream based on the evaluation parameters. It should be understood that the process of determining the evaluation information associated with the corresponding audio stream by the server 120 each time should also be completed before the next time the server 120 determines the evaluation parameters corresponding to the audio stream, so that the server 120 can evaluate the audio stream sent by the terminal device 110 in real time.

[0048] The set of evaluation parameters can include a plurality of evaluation parameters. In some embodiments, the server 120 can determine the evaluation information associated with the at least one audio stream based on the plurality of evaluation parameters and weight information, and the weight information can indicate the weights corresponding to the plurality of evaluation parameters.

[0049] For example, the server 120 can determine the evaluation information based on the following manner: evaluation information = first evaluation parameter x first weight + second evaluation parameter x second weight +... + mth evaluation parameter x mth weight. The weight values can be set according to the influence degree of different evaluation parameters on the audio stream.

[0050] It is worth noting that when performing the above calculation, the m evaluation parameters need to be pre-processed. The pre-processing of the m evaluation parameters may, for example, be normalization processing, so as to unify the dimensions of the m evaluation parameters, and avoid affecting the accuracy of the evaluation information due to the excessively large or small value corresponding to a certain evaluation parameter. For example, the m evaluation parameters can be normalized to convert the m evaluation parameters into values between 0 and 100.

[0051] In some embodiments, the set of evaluation parameters can include, but is not limited to, any one of the following: a first evaluation parameter indicating the network status of the terminal device 110 (for example, network transmission rate, network bandwidth, error code rate, delay, packet loss rate, etc.), a second evaluation parameter indicating the performance state of the human voice audio collected by the terminal device 110 (for example, the singing score of a user in a multi-person chorus scene), and a third evaluation parameter indicating the volume of the human voice audio collected by the terminal device 110.

[0052] For example, the normalization processing of the first evaluation parameter, the second evaluation parameter, and the third evaluation parameter can be performed according to a pre-constructed parameter-score relationship curve. Taking the normalization processing of the third evaluation parameter as an example, referring to FIG. 3, the horizontal coordinate represents the volume of the human voice audio collected by the terminal device 110, and the vertical coordinate represents the score of the volume of the human voice audio. As shown in FIG. 3, the excessively large (with noise) or excessively small (leading to unclear) volume of the human voice audio will result in a lower score.

[0053] In some embodiments, the at least one audio stream can include a first audio stream and a second audio stream. The first audio stream can be the audio stream corresponding to a participant of a live interactive event, and the second audio stream can be noise. The server 120 can determine the first evaluation information for the first audio stream based on a first set of evaluation parameters associated with the first audio stream and a second set of evaluation parameters associated with the second audio stream.

[0054] For example, the server 120 can determine the first evaluation information for the first audio stream based on the first set of evaluation parameters, the second set of evaluation parameters, the weight corresponding to the first set of evaluation parameters, and the weight corresponding to the second set of evaluation parameters. The determination of the second evaluation parameter of the second audio stream and the determination of the first evaluation information for the first audio stream can refer to the related content described above, which will not be described here.

[0055] In some embodiments, the terminal device 110 can send the weight information corresponding to the evaluation parameter to the server 120. In other embodiments, the server 120 can determine the weight information corresponding to the evaluation parameter. In other embodiments, the weight parameter can be sent to the server 120 by other devices (e.g., terminal devices not participating in the live interactive event), and the embodiments of the present disclosure do not limit.

[0056] Continuing to refer to FIG. 2, at block 230, the server 120 determines a processing mode for the plurality of audio streams based on at least the evaluation information of the at least one audio stream.

[0057] In some embodiments, the server 120 can sort the at least one audio stream based on the evaluation information, and determine the processing mode based on the sorted at least one audio stream. The processing mode indicates a set of audio streams to be pushed from the at least one audio stream, and the set of audio streams has a predetermined number. For example, a predetermined number of audio streams with the highest evaluation can be determined as the set of audio streams to be pushed, or the first predetermined number of audio streams can be determined as the set of audio streams to be pushed.

[0058] In some embodiments, since the evaluation information is determined by the server 120 at a predetermined frequency, the server 120 can determine the processing mode for the plurality of audio streams based on the obtained evaluation information each time the evaluation information is obtained. The server 120 can determine that an audio stream will not be pushed in a second time period in response to not receiving an evaluation parameter associated with the audio stream in a first time period, thereby improving the quality of the audio stream after stream merging.

[0059] In some embodiments, the first time period and the second time period can be determined by the predetermined frequency. In other embodiments, the first time period and the second time period can also be reasonably set as needed, for example, the first time period can include one or more time periods determined by the predetermined frequency, and the second time period can also include one or more time periods determined by the predetermined frequency, and the embodiments of the present disclosure do not limit.

[0060] In some embodiments, for the audio stream that is not pushed, the server 120 can continuously determine the evaluation information of the audio stream that is not pushed to form a global perception. For example, an audio stream is determined to be not pushed in a second time period, and the server 120 can also continue to determine the evaluation information of the audio stream in the second time period, to determine whether the audio stream is pushed in the next time period of the second time period.

[0061] At block 240, the server 120 processes the audio data associated with the plurality of audio streams based on the processing mode.

[0062] In some embodiments, the server 120 can send a control message to a target terminal device in the plurality of terminal devices 110 based on the processing mode, the control message indicating whether a target audio stream corresponding to the target terminal device is pushed or not.

[0063] In some embodiments, the server 120 can send a push result to the plurality of terminal devices 110, the push result indicating the pushed audio stream and the non-pushed audio stream. In some embodiments, the server 120 can send the push result to the plurality of terminal devices 110 through a long link.

[0064] According to embodiments of the present disclosure, a plurality of audio streams are first acquired from a plurality of terminal devices, evaluation information associated with at least one audio stream in the plurality of audio streams is then determined, a processing mode for the plurality of audio streams is then determined based on at least the evaluation information of the at least one audio stream, and finally, audio data associated with the plurality of audio streams is processed based on the processing mode.

[0065] In this way, embodiments of the present disclosure can select an audio stream to be pushed from the plurality of audio streams based on the evaluation information, which on one hand can improve the quality of the merged audio stream, and on the other hand can reduce the transmission cost of the audio stream and the performance overhead of the merging alignment operation.

[0066] FIG. 4 shows a schematic structural block diagram of an apparatus 400 for audio processing according to certain embodiments of the present disclosure. The apparatus 400 can be implemented as or included in the server 120. Various modules / components in the apparatus 400 can be implemented by hardware, software, firmware, or any combination thereof.

[0067] As shown in FIG. 4, the apparatus 400 includes an acquisition module 410 configured to acquire a plurality of audio streams from a plurality of terminal devices, the plurality of audio streams corresponding to a plurality of participants of a live interactive event. The apparatus 400 further includes a first determination module 420 configured to determine evaluation information associated with at least one audio stream in the plurality of audio streams. The apparatus 400 further includes a second determination module 430 configured to determine a processing mode for the plurality of audio streams based on at least the evaluation information of the at least one audio stream. The apparatus 400 further includes a processing module 440 configured to process audio data associated with the plurality of audio streams based on the processing mode.

[0068] In some embodiments, the first determination module 420 is further configured to acquire a set of evaluation parameters from at least one terminal device corresponding to the at least one audio stream; and determine the evaluation information associated with the at least one audio stream based on the set of evaluation parameters.

[0069] In some embodiments, the at least one audio stream includes a first audio stream and a second audio stream, and the first determining module 420 is further configured to determine first evaluation information for the first audio stream based on a first set of evaluation parameters associated with the first audio stream and a second set of evaluation parameters associated with the second audio stream.

[0070] In some embodiments, the apparatus 400 further includes a third determining module configured to determine, in response to not receiving evaluation parameters associated with a third audio stream within a first time period, that the third audio stream will not be pushed within a second time period.

[0071] In some embodiments, the set of evaluation parameters includes a plurality of evaluation parameters, and the first determining module 420 is further configured to determine the evaluation information associated with the at least one audio stream based on the plurality of evaluation parameters and weight information, the weight information indicating weights corresponding to the plurality of evaluation parameters.

[0072] In some embodiments, the set of evaluation parameters includes at least one of: a first evaluation parameter indicating a network status of the at least one terminal device; a second evaluation parameter indicating a performance status of human voice audio collected by the at least one terminal device; and a third evaluation parameter indicating a volume of human voice audio collected by the at least one terminal device.

[0073] In some embodiments, the second determining module 430 is further configured to sort the at least one audio stream based on the evaluation information, and determine the processing mode based on the sorted at least one audio stream, the processing mode indicating a set of audio streams to be pushed from the at least one audio stream, the set of audio streams having a predetermined number.

[0074] In some embodiments, the pushing module 440 is further configured to send, based on the processing mode, a control message to a target terminal device from the plurality of terminal devices, the control message indicating whether a target audio stream corresponding to the target terminal device is used for pushing.

[0075] FIG. 5 shows a block diagram of an electronic device 500 in which one or more embodiments of the disclosure can be implemented. It should be understood that the electronic device 500 shown in FIG. 5 is merely an example and should not be construed to limit the functionality and scope of the embodiments described herein. The electronic device 500 shown in FIG. 5 can be used to implement the server 120 of FIG. 1.

[0076] As shown in FIG. 5, electronic device 500 is in the form of a general-purpose electronic device. Components of electronic device 500 can include, but are not limited to, one or more processors or processing units 510, memory 520, storage 530, one or more communication units 540, one or more input devices 550, and one or more output devices 560. Processing unit(s) 510 can be actual or virtual processors and capable of executing various processing in accordance with programs stored in memory 520. In a multi-processing system, multiple processing units execute computer-executable instructions in parallel to improve the processing power of electronic device 500.

[0077] Electronic device 500 typically includes a plurality of computer storage media. Such media can be removable and / or non-removable, and can include volatile and / or nonvolatile media. Memory 520 can be volatile (such as, for example, registers, cache, random access memory (RAM)), non-volatile (such as, for example, read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory), or some combination thereof. Storage 530 can be removable or non-removable and can include machine-readable media, such as, for example, flash drives, disks, or any other media capable of storing information and / or data and accessible by electronic device 500.

[0078] Electronic device 500 can further include additional removable / non-removable, volatile / non-volatile storage media. Although not shown in FIG. 5, a disk drive or other computer-readable media drive can be provided for reading from or writing to a removable, non- volatile magnetic disk (e.g., a "hard drive"), and a disk drive or other computer-readable media drive can be provided for reading from or writing to a removable, non-volatile optical disk (such as a CD-ROM or other optical medium). In these instances, each drive can be connected to the bus (not shown) by one or more data media interfaces. Memory 520 can include a computer program product 525 having one or more program modules configured to carry out the various methods or actions of the various embodiments of the present disclosure.

[0079] Communication unit(s) 540 enable communication with other electronic devices via communication media. Additionally, functionality of components of electronic device 500 can be implemented in a single computing cluster or a plurality of computer machines capable of communication through a communication connection. Accordingly, electronic device 500 can operate in a networked environment using logical connections to one or more other servers, network personal computers (PCs), or another network nodes.

[0080] The input device 550 can be one or more input devices, such as a mouse, a keyboard, a trackball, etc. The output device 560 can be one or more output devices, such as a display, a speaker, a printer, etc. The electronic device 500 can also communicate with one or more external devices (not shown), such as storage devices, display devices, etc., one or more devices that enable a user to interact with the electronic device 500, or any devices (e.g., a network card, a modem, etc.) that enable the electronic device 500 to communicate with one or more other electronic devices, as desired via the communication unit 540. Such communication can be carried out via an input / output (I / O) interface (not shown).

[0081] According to an example implementation of the present disclosure, a computer readable storage medium is provided having computer executable instructions stored thereon, where the computer executable instructions are executed by a processor to implement the method described above. According to an example implementation of the present disclosure, a computer program product is also provided that is tangibly stored on a non-transitory computer readable medium and includes computer executable instructions, where the computer executable instructions are executed by a processor to implement the method described above.

[0082] The computer readable program instructions can also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable apparatus or other device to produce a computer implemented process such that the instructions which execute on the computer or other programmable apparatus provide processes for implementing the functions / acts specified in the flowchart and / or block diagram block or blocks.

[0083] The computer readable program instructions can also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable apparatus or other device to produce a computer implemented process such that the instructions which execute on the computer or other programmable apparatus provide processes for implementing the functions / acts specified in the flowchart and / or block diagram block or blocks.

[0084] The computer readable program instructions can also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable apparatus or other device to produce a computer implemented process such that the instructions which execute on the computer or other programmable apparatus provide processes for implementing the functions / acts specified in the flowchart and / or block diagram block or blocks.

[0085] The computer program product of the present disclosure can be a computer program product, which is a machine-readable medium (media) having instances of the software embodied thereon, such as computer software, firmware, wireless application protocol (WAP), middleware or microcode. For example, a computer program product can be a floppy disk, a CD-ROM, a DVD, a Blu-ray Disc™, a flash drive, a memory stick, a magnetic tape, or a hard disk drive. The machine-readable medium can be a single medium, or multiple media, of the same or different type. The computer program product can be one or more computer program components embodied in medium and / or transmission signals. The computer program product can have one or more computer program components embodied in medium and / or transmission signals.

[0086] The implementations of the disclosure have been described above with the intent to be illustrative rather than limiting. Although the implementations of the disclosure have been described with regard to one or more implementations, it will be recognized that a variety of modifications and changes can be made to these implementations without departing from the broader spirit and scope of the implementations as set forth in the preceding disclosure. For example, certain aspects of the implementations can be performed using hardware, software, and / or firmware, or any combination thereof. The above-described implementations should therefore be regarded as merely illustrative, and not as narrowing the scope of the disclosure, which is defined by the appended claims and their equivalents.

Claims

1. A method for audio processing, comprising: obtaining a plurality of audio streams from a plurality of terminal devices, the plurality of audio streams corresponding to a plurality of participants of a live interactive event; determining evaluation information associated with at least one audio stream of the plurality of audio streams; determining a processing mode for the plurality of audio streams based at least on the evaluation information of the at least one audio stream; and processing audio data associated with the plurality of audio streams based on the processing mode. 2.The method of claim 1, wherein determining evaluation information associated with at least one audio stream of the plurality of audio streams comprises: obtaining a set of evaluation parameters from at least one terminal device corresponding to the at least one audio stream; and determining the evaluation information associated with the at least one audio stream based on the set of evaluation parameters. 3.The method of claim 2, wherein the at least one audio stream comprises a first audio stream and a second audio stream, and determining the evaluation information associated with the at least one audio stream based on the set of evaluation parameters comprises: determining first evaluation information for the first audio stream based on a first set of evaluation parameters associated with the first audio stream and a second set of evaluation parameters associated with the second audio stream. 4.The method of claim 2, further comprising: determining that a third audio stream is not to be pushed in a second time period in response to not receiving evaluation parameters associated with the third audio stream in a first time period. 5.The method of claim 2, wherein the set of evaluation parameters comprises a plurality of evaluation parameters, and determining the evaluation information associated with the at least one audio stream based on the set of evaluation parameters comprises: determining the evaluation information associated with the at least one audio stream based on the plurality of evaluation parameters and weight information, the weight information indicating weights corresponding to the plurality of evaluation parameters. 6.The method of claim 2, wherein the set of evaluation parameters comprises at least one of: a first evaluation parameter indicating a network status of the at least one terminal device; a second evaluation parameter indicating a performance status of human voice audio captured by the at least one terminal device; a third evaluation parameter indicating a volume of human voice audio captured by the at least one terminal device. 7.The method of claim 1, wherein determining a processing mode for the plurality of audio streams based at least on the evaluation information of the at least one audio stream comprises: ranking the at least one audio stream based on the evaluation information; and determining the processing mode based on the ranked at least one audio stream, the processing mode indicating a set of audio streams to be pushed from the at least one audio stream, the set of audio streams having a predetermined number. 8.The method of claim 1, wherein processing audio data associated with the plurality of audio streams based on the processing mode comprises: sending a control message to a target terminal device of the plurality of terminal devices based on the processing mode, the control message indicating whether a target audio stream corresponding to the target terminal device is for pushing. ​ ​ ​ 9. An apparatus for audio processing, comprising: an obtaining module configured to obtain a plurality of audio streams from a plurality of terminal devices, the plurality of audio streams corresponding to a plurality of participants of a live interactive event; a first determining module configured to determine evaluation information associated with at least one audio stream of the plurality of audio streams; a second determining module configured to determine a processing mode for the plurality of audio streams based at least on the evaluation information of the at least one audio stream; and a processing module configured to process audio data associated with the plurality of audio streams based on the processing mode.

10. An electronic device, comprising: at least one processing unit; and at least one memory coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit, the instructions when executed by the at least one processing unit cause the terminal device to perform the method according to any one of claims 1-8.

11. A computer readable storage medium having stored thereon a computer program, the computer program being executable by a processor to implement the method according to any one of claims 1-8.

12. A computer program product comprising computer executable instructions, wherein the computer executable instructions implement the method according to any one of claims 1-8 when executed by a processor. ​ ​

Citation Information

Patent Citations

  • Methods and apparatuses for processing audio streams for use with multiple devices

    CN101553801A

  • Live broadcast audio processing method and device, computer equipment and storage medium

    CN112637613A

  • Audio stream processing method, computer device and computer program product

    CN116170613A

  • Audio processing method, computer equipment and computer storage medium

    CN117373480A

  • Systems and methods for facilitating side-channel communications during shared communication session

    US11153442B1