Audio processing method and apparatus, electronic device, user equipment, and storage medium
By using a music detection model to detect audio type and adjust virtual spatial position during interactive live streaming, the problem of reduced output quality caused by audio processing is solved, achieving more efficient spatial audio processing and improving audio output quality and immersive experience.
Patent Information
- Application Number
- PCT/CN2025/087915
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-04-10
- Filing Date
- 2025-04-09
- Publication Date
- 2025-10-16
AI Technical Summary
In interactive live streaming, spatial audio processing can reduce the audio output quality in certain scenarios, especially when live streaming users are using sound cards or playing music. Existing technology makes it difficult to accurately determine whether spatial audio processing should be performed to improve the immersive experience.
By acquiring audio streams from the first client during interactive live streaming, using a music detection model to detect the audio data type, determining whether to perform spatial audio processing based on the type, and adjusting the virtual spatial position as necessary to optimize the audio output effect.
It improves the audio output effect in interactive live streaming scenarios, reduces unnecessary resource consumption, improves the accuracy and efficiency of audio processing, and ensures the stereo and immersive experience of audio output.
Smart Images

Figure CN2025087915_16102025_PF_FP_ABST
Abstract
Description
Audio processing method and device, electronic device, user equipment and storage medium
[0001] Cross-reference to Related Applications
[0002] This application is based on and claims priority to the application with the Chinese application number 202410433363.8 and the filing date of April 10, 2024, the disclosure of which is incorporated herein in its entirety. TECHNICAL FIELD
[0003] The present disclosure relates to the technical field of computer, in particular to an audio processing method and device, electronic device, user equipment and storage medium. BACKGROUND
[0004] Interactive live streaming refers to multiple live streaming users joining a multi-person live streaming room for live streaming interaction, which can also be referred to as multi-user live streaming. Multiple client outputs of multiple live streaming users are combined into one live streaming for playing. In the case that the combined live streaming only includes an audio stream, it can present multiple live streaming user voice interaction chats, and in the case that the combined live streaming includes a video stream, it can simultaneously display multiple live streaming user live streaming pictures in the live streaming interface.
[0005] Processing the audio streams of multiple clients with spatial audio processing can make the sounds of different live streaming users sound obviously different in spatial orientation, and viewers can obtain a better immersive experience when watching interactive live streaming. SUMMARY
[0006] This summary is provided to introduce a selection of concepts, which are further described below in the detailed description. This summary is not intended to identify key or essential features of the claimed subject matter, nor is it intended to be used in limiting the scope of the claimed subject matter.
[0007] According to some embodiments of the present disclosure, an audio processing method is provided, comprising: collecting audio data from an audio stream obtained by a first client, wherein the first client is a client of a first user in multiple live streaming users participating in interactive live streaming; determining a type of the audio data according to the audio data; and determining whether to perform spatial audio processing on the audio stream obtained by the first client according to the type of the audio data.
[0008] According to other embodiments of the present disclosure, an audio processing device is provided, including: an acquisition module configured to acquire audio data from an audio stream obtained by a first client, wherein the first client is a client of a first user among multiple live broadcast users participating in an interactive live broadcast; a classification module configured to determine a type of audio data based on the audio data; and a spatial audio processing module configured to determine whether to perform spatial audio processing on the audio stream obtained by the first client based on the type of the audio data.
[0009] According to some further embodiments of the present disclosure, an electronic device is provided, including: a memory; and a processor coupled to the memory, wherein the processor is configured to execute the audio processing method of any embodiment of the present disclosure based on instructions stored in the memory.
[0010] According to some further embodiments of the present disclosure, a user device is provided, comprising: the audio processing apparatus of any of the aforementioned embodiments; and an audio input device configured to collect audio and input the audio into a first client.
[0011] According to some further embodiments of the present disclosure, a computer-readable storage medium is provided, on which a computer program is stored. When the program is executed by a processor, the audio processing method of any embodiment of the present disclosure is executed.
[0012] According to some further embodiments of the present disclosure, a computer program product is provided, comprising: instructions, wherein when the instructions are executed by a processor, the audio processing method of any embodiment of the present disclosure is implemented.
[0013] According to some further embodiments of the present disclosure, a computer program is provided, comprising: instructions, wherein when the instructions are executed by a processor, the audio processing method of any embodiment of the present disclosure is implemented.
[0014] Other features, aspects and advantages of the present disclosure will become apparent from the following detailed description of exemplary embodiments of the present disclosure with reference to the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] The preferred embodiments of the present disclosure are described below with reference to the accompanying drawings. The drawings described herein are used to provide a further understanding of the present disclosure. Each of the drawings, together with the following detailed description, is included in this specification and forms a part of the specification to explain the present disclosure. It should be understood that the drawings described below only relate to some embodiments of the present disclosure and do not constitute a limitation of the present disclosure. In the drawings:
[0016] FIG1 is a schematic flow chart showing an audio processing method according to some embodiments of the present disclosure;
[0017] FIG2 is a flowchart showing an audio processing method according to some other embodiments of the present disclosure;
[0018] FIG. 3 shows a flowchart of an audio processing method according to some embodiments of the present disclosure;
[0019] FIG. 4 shows a flowchart of an audio processing method according to some embodiments of the present disclosure;
[0020] FIG. 5 shows a structural diagram of an audio processing apparatus according to some embodiments of the present disclosure;
[0021] FIG. 6 shows a structural diagram of an electronic device according to some embodiments of the present disclosure;
[0022] FIG. 7 shows a structural diagram of an electronic device according to some embodiments of the present disclosure;
[0023] FIG. 8 shows a structural diagram of a user device according to some embodiments of the present disclosure.
[0024] It should be understood that the dimensions of the various portions shown in the attached drawings are shown schematically and are not necessarily true to scale. Identical or similar components are identified throughout the various figures with identical or similar reference numerals. As such, once a component is defined in one figure, further discussion of such component in subsequent figures can be omitted. DETAILED DESCRIPTION
[0025] The technical solutions in the embodiments of the present disclosure will be clearly and completely described below with reference to the drawings in the embodiments of the present disclosure. Obviously, the described embodiments are only some of the embodiments of the present disclosure, but not all the embodiments. The description of the embodiments below is actually only illustrative and should not be construed as any limitation on the present disclosure and its application or use. It should be understood that the present disclosure can be implemented in various forms and should not be construed as being limited to the embodiments described herein.
[0026] It should be understood that the various steps recited in the method embodiments of the present disclosure can be performed in different orders and / or in parallel. In addition, the method embodiments can include additional steps and / or omit the steps shown. The scope of the present disclosure is not limited in this respect. Unless otherwise specifically stated, the relative arrangement of components and steps, numerical expressions, and numerical values set forth in these embodiments should be interpreted as merely exemplary and not limiting to the scope of the present disclosure.
[0027] The term "comprise" and variations of the term, such as "comprising", "includes" and "including", used in the present disclosure means an open term that includes at least the elements listed thereafter, but does not exclude other elements. In other words, the term "comprising" and variations of the term, such as "comprising", "includes" and "including", means "including but not limited to". Further, the term "consisting of" and variations of the term, such as "consisting essentially of" and "consisting of", used in the present disclosure means an open term that includes at least the elements listed thereafter, but does not exclude other elements. In other words, the term "consisting of" and variations of the term, such as "consisting essentially of" and "consisting of", means "including but not limited to". Therefore, comprising and including are synonymous. The term "based on" means "based at least in part on".
[0028] Reference throughout this specification to "one embodiment", "some embodiments" or "an embodiment" means that a particular feature, structure or characteristic described in connection with the embodiment is included in at least one embodiment of the application. The appearances of the phrase "in one embodiment" in various places in the specification are not necessarily all referring to the same embodiment, although it can. As used herein, the term "or" as used herein, without further qualification, can be used to describe either a selective singular or plural scenario, such that, at least one of the group of items or at least one of the group of items can be present. For example, the phrase "A or B" can mean "A or B or both A and B".
[0029] It should be noted that the terms "first", "second", and the like, used in the description and in the claims of this disclosure, are used for distinguishing between similar elements and not necessarily for describing a sequential or chronological order. It is to be understood that the terms so used are interchangeable under appropriate circumstances such that the descriptive terms "first", "second", etc., can be interchanged with one another.
[0030] It should be noted that the terms "one", "another", "an", and the like, as used in this disclosure, are used in the sense that they mean "at least one", "any one", or "one or more", unless otherwise specified.
[0031] The names of the messages or information exchanged between the plurality of apparatuses in the embodiments of the present disclosure are used only for illustrative purposes, and are not intended to limit the scope of the messages or information.
[0032] The embodiments of the present disclosure will be described in detail below with reference to the drawings, but the present disclosure is not limited to these specific embodiments. The following specific embodiments can be combined with each other, and the same or similar concepts or processes can not be described in some embodiments. In addition, in one or more embodiments, specific features, structures, or characteristics can be combined by any suitable means apparent from the present disclosure to those of ordinary skill in the art.
[0033] The inventor finds that although processing the audio streams of multiple clients with spatial audio processing can provide a better immersive experience for the audience when watching the interactive live broadcast, spatial audio processing is not suitable for all scenarios. In some scenarios, spatial audio processing of audio can reduce the output effect of the audio. For example, when a live broadcast user is performing, the live broadcast user may play music using a sound card or other equipment. At this time, spatial audio processing of the audio input by the live broadcast user can cause problems such as heavy tone and monaural output, and reduce the sound effect of the performance of the live broadcast user. Therefore, how to more accurately and effectively perform spatial audio processing on the audio in the interactive live broadcast and improve the audio output effect in the scenario of the interactive live broadcast is a problem to be solved.
[0034] Therefore, the present disclosure provides an audio processing method, which is described below in conjunction with FIGS. 1-4.
[0035] FIG. 1 is a flowchart of some embodiments of the audio processing method of the present disclosure. As shown in FIG. 1, the method of this embodiment includes steps S102-S106. The method of this embodiment can be performed by a first client. The first client is the client of a first user in a plurality of live broadcast users participating in an interactive live broadcast. The client can be application software or hardware, for example, the client can be a live broadcast application, a hardware device in a user device, etc. The client corresponding to each live broadcast user participating in the interactive live broadcast can perform the method of this embodiment as the first client.
[0036] In step S102, audio data is collected from an audio stream obtained by the first client.
[0037] The audio data and other information related to the user collected in the present disclosure are authorized by the user. After collecting the audio data, the data is de-identified by technical means.
[0038] The plurality of live broadcast users participating in the interactive live broadcast can also be referred to as multi-user live broadcast. The plurality of live broadcast users can join a multi-person live broadcast room for live broadcast interaction. Each live broadcast user in the plurality of live broadcast users corresponds to a client, and the live broadcast streams of the plurality of clients can be combined into one live broadcast stream (co-stream) for playing. In the case of a video stream, the live broadcast pictures of the plurality of live broadcast users are displayed simultaneously in the live broadcast interface (user interface).
[0039] In some embodiments, the audio stream obtained by the first client includes at least one of audio input by the first client and audio sent by a second client received by the first client. The second client is one or more clients of one or more second users in the plurality of live broadcast users participating in the interactive live broadcast, except the first user. The audio input by the first client is audio collected by the user device to which the first client belongs through an audio input device.
[0040] For example, the audio data includes audio frames. In some embodiments, in response to detecting that audio is input by the first client, consecutive audio frames are collected from the audio stream obtained from the first client.
[0041] In the case that the user equipment to which the first client belongs collects audio on the first user side through an audio input device, the collection of audio data is performed. If there is no audio input on the first user side, for example, the first user does not speak, the audio input function is turned off, etc., the collection of audio data can not be performed, thereby saving resources and improving operation efficiency. Of course, the collection of audio data can also be performed in real time.
[0042] In some embodiments, in response to detecting that the music playing tool in the live broadcast application is not turned on, the audio data is collected from the audio stream obtained from the first client. If the first user directly plays music using the music playing tool in the live broadcast application to perform, the audio stream obtained from the first client can not be subjected to spatial audio processing.
[0043] In step S104, the type of the audio data is determined according to the audio data.
[0044] In some embodiments, the type of the audio data includes music and non-music. Of course, other types of audio data can also be distinguished from music and non-music, and whether to perform spatial audio processing can be selected, which can be determined according to actual needs. In subsequent embodiments, music and non-music are taken as examples for description.
[0045] In some embodiments, a music detection model is used to detect the audio data to determine the type of the audio data. The collected audio data can be detected in real time using the music detection model. The music detection model can extract features of the audio data, predict a probability that the audio data belongs to music according to the features of the audio data, and determine whether the audio data belongs to music according to a comparison between the probability that the audio data belongs to music and a probability threshold.
[0046] The music detection model can be a machine learning model, for example, a deep neural network model, such as an RNN-T (Recurrent Neural Network Transducer), an RNN-AED (Recurrent Neural Network-Attention-based Encoder-Decoder), a Transformer-AED (Transformer-Attention-based Encoder-Decoder), and the like, without being limited to the examples.
[0047] In step S106, it is determined whether to perform spatial audio processing on the audio stream obtained by the first client according to the type of the audio data.
[0048] In some embodiments, in response to determining that the type of the audio data is music, it is determined not to perform spatial audio processing on the audio stream obtained by the first client; and in response to determining that the type of the audio data is non-music, it is determined to perform spatial audio processing on the audio stream obtained by the first client.
[0049] If the type of the collected audio data is non-music, it indicates that the first user has not started performing, and can only be talking or the like, and in this case, spatial audio processing can be performed on the audio stream obtained by the first client, that is, the spatial audio processing function is turned on. If the type of the collected audio data is music, it indicates that the first user has started performing, and in order to avoid the problem of poor audio output effect caused by performing spatial audio processing on the audio stream obtained by the first client, it is selected not to perform spatial audio processing on the audio stream obtained by the first client, that is, the spatial audio processing function is turned off.
[0050] In some embodiments, performing spatial processing on the audio stream obtained by the first client includes: mapping the positions of one or more live users corresponding to the audio stream in a live picture (or user interface) to a three-dimensional spatial audio space to obtain virtual spatial positions of the one or more live users. The audio stream obtained by the first client is processed according to the virtual spatial positions of the one or more live users to obtain first audio with spatial audio effect, which can be played at the first client. The first audio can also be sent to the server, and the client of the audience plays the first audio received from the server, which can produce a sense of direction and stereoscopic effect in hearing. For example, the first user is a host. The second client can process the obtained audio stream by the same method, which will not be described here.
[0051] The live picture can be set with multiple positions according to preset rules, and the live pictures of live users of different clients are displayed, for example, in the case of live intercommunication of two live users, the live picture can be divided into two half screens on the left and right to display the live pictures of the two live users. Through spatial audio processing, the sound of the live user in the left picture can be heard on the left, and the sound of the live user in the right picture can be heard on the right.
[0052] The method of the above embodiment collects audio data from the audio stream obtained by the first client of the first user participating in the interactive live broadcast, determines the type of the audio data, and determines whether to perform spatial audio processing on the audio stream obtained by the first client. The method of the above embodiment can distinguish different types of audio data and select whether to perform spatial audio processing, more accurately and effectively perform spatial audio processing on the audio in the interactive live broadcast, and improve the audio output effect in the scene of the interactive live broadcast. In addition, spatial audio processing is not performed for all scenes, which reduces the overall overhead, saves resources, and improves the processing efficiency of the audio.
[0053] In some embodiments, collecting audio data from the audio stream obtained by the first client includes: in response to detecting that audio is input by the first client, collecting continuous audio frames from the audio stream obtained by the first client until the number of continuous audio frames reaches a threshold, and taking the continuous audio frames as the audio data.
[0054] Detecting whether the first client has audio input, that is, detecting whether the audio input device of the user equipment to which the first client belongs collects the audio input by the first user. The steps of detecting whether audio is input by the first client and collecting continuous audio frames from the audio stream obtained by the first client can be repeatedly performed until the number of continuous audio frames reaches a threshold.
[0055] In some embodiments, in response to the number of continuous audio frames not reaching the threshold, it is determined that spatial audio processing is performed on the audio stream obtained by the first client.
[0056] If the first user starts to perform and play background music, the first client can detect a continuous audio stream, that is, continuous audio frames, and if the number of continuous audio frames that reaches the threshold cannot be collected, it can be judged that no music is played. Therefore, it can be directly determined that spatial audio processing is performed on the audio stream obtained by the first client.
[0057] In some embodiments, the music detection model is used to detect the audio data and determine the type of the audio data.
[0058] For example, in a case where the number of collected continuous audio frames reaches a threshold value, the continuous audio frames are detected by using the music detection model to determine the type of the audio data, which can reduce the number of calls to the music detection model, reduce the complexity and overhead of the operation, and improve the detection efficiency. Of course, the music detection model can also be used to detect each audio frame.
[0059] In some embodiments, according to the characteristics of the audio data, it is determined whether to call the music detection model to detect the audio data and determine the type of the audio data, wherein the characteristics of the audio data include at least one of the input source of the audio frame in the audio data and the energy information of the audio frame. The input source of the audio frame in the audio data can only include the first client, only include the second client, or include the first client and the second client, and the second client can be one or more. The case where the input source of the audio frame in the audio data only includes the second client can be excluded by the aforementioned step of detecting whether the first client has audio input. The energy information of the audio frame can include at least one of the energy of the audio frame corresponding to the first client and the energy of the audio frame corresponding to the second client.
[0060] In some embodiments, in response to detecting that the input source of the audio frame includes the first client and the second client, the energy of the audio frame corresponding to the first client and the energy of the audio frame corresponding to the second client are determined; and in response to determining that the energy of the audio frame corresponding to the first client is greater than the energy of the audio frame corresponding to the second client, it is determined to call the music detection model to detect the audio data.
[0061] In a case where a live user performs, other live users usually do not speak or occasionally whisper. Therefore, the energy of the audio frame input by the performing live user is greater, and by comparing the energy of the audio frame corresponding to the first client and the second client, it can be preliminarily determined whether the first user of the first client is likely to play music and whether it is necessary to further accurately detect by the music detection model.
[0062] If the input source of the audio frame includes the first client and the second client, and the energy of the audio frame corresponding to the first client is the largest, the audio data can be marked, for example, a singer label is set, which is not limited to the examples. Only the audio data with the mark will be input into the music detection model for processing to determine whether it is music. The energy of each audio frame can be determined by the amplitude or peak value of each audio frame, and specific reference can be made to the existing calculation method, which will not be described here. For a plurality of audio frames, the energies of the plurality of audio frames can be summed.
[0063] In some embodiments, in response to detecting that the input source of the audio frame includes the first client and the second client, the energy of the audio frame corresponding to the first client and the energy of the audio frame corresponding to the second client are determined; in response to determining that the energy of the audio frame corresponding to the second client is greater than the energy of the audio frame corresponding to the first client, it is determined that the music detection model is not called to detect the audio data.
[0064] In some embodiments, in response to determining that the energy of the audio frame corresponding to the second client is greater than or equal to the energy of the audio frame corresponding to the first client, it is determined that the type of the audio data is non-music.
[0065] If the energy of the audio frame corresponding to the second client is greater than the energy of the audio frame corresponding to the first client, it can be determined that the first user of the first client does not play music, the type of the audio data can be directly determined as non-music, and the audio stream obtained by the first client can be directly determined to be subjected to spatial audio processing. If there are multiple second clients, as long as the energy of one second client is greater than the energy of the audio frame corresponding to the first client, it is determined that the music detection model is not called to detect the audio data, or the type of the audio data is determined as non-music.
[0066] In some embodiments, in response to detecting that the input source of the audio frame includes the first client and the second client, the energy of the audio frame corresponding to the first client and the energy of the audio frame corresponding to the second client are determined; in response to determining that the energy of the audio frame corresponding to the second client is greater than the energy of the audio frame corresponding to the first client, it is determined that the music detection model is not called to detect the audio data.
[0067] In the case of input audio of only the first client, it can be further determined whether the audio data is music.
[0068] The method of the above embodiments, in the case that the number of collected continuous audio frames reaches a threshold, further determines whether to call the music detection model to detect the audio data according to the characteristics of the audio data, so as to determine the type of the audio data, which can further reduce the number of calls to the music detection model, reduce the complexity and overhead of the operation, and improve the detection efficiency. Those skilled in the art can understand that the music detection model can also be directly called to detect the audio data without referring to the characteristics of the audio data.
[0069] In a case where the input source of the audio frame only includes the first client, a music detection model is called to further detect the type of the audio data, and in a case where the type of the audio data is non-music, it is determined to perform spatial audio processing on the audio stream obtained by the first client. However, in this case, it is likely that only the first user of the first client is continuously speaking, and other users rarely speak or do not speak. For this case, if the spatial audio processing is performed on the audio stream obtained by the first client, it is likely that the sound corresponding to the first client always seems to be on one side (left or right), and in this case, the spatial audio processing method is not accurately applied, and the output effect is reduced. To solve this problem, the disclosure further proposes a method for optimizing the virtual spatial position of the first client, which is described below.
[0070] In some embodiments, in response to determining that the type of the audio data is non-music, it is determined to perform spatial audio processing on the audio stream obtained by the first client; and in response to detecting that the input source of the audio frame only includes the first client, adjusting a current virtual spatial position corresponding to the first user according to an initial virtual spatial position corresponding to the first user, a number of times of calling the music detection model, and a preset step length, so that the current virtual spatial position corresponding to the first user gradually approaches a virtual center position until a difference between the current virtual spatial position corresponding to the first user and the virtual center position is not greater than a preset difference, wherein the initial virtual spatial position corresponding to the first user is a position of a position of the first user in a live picture mapped to a three-dimensional spatial audio space, and the virtual center position is a center position of the three-dimensional spatial audio space.
[0071] In a case where the input source of the audio frame only includes the first client, each time the music detection model is called to determine that the type of the audio data is non-music, the initial virtual spatial position corresponding to the first user can be adjusted according to the number of times of calling the music detection model and the preset step length. Through multiple adjustments, the virtual spatial position of the first client can gradually approach the virtual center position. In this way, the audio input from the first user side presents a stereo auditory effect of gradually moving from one side (or one corner) to the center when playing.
[0072] In some embodiments, the current virtual spatial position corresponding to the first user is adjusted according to a difference between coordinates of the initial virtual spatial position corresponding to the first user and coordinates of the virtual center position, the number of times of calling the music detection model, and the preset step length, wherein the difference between the coordinates of the current virtual spatial position corresponding to the first user and the coordinates of the virtual center position is inversely proportional to the number of times of calling the music detection model.
[0073] For example, the number of times of calling the music detection model is multiplied by the preset step length, and a difference between a preset value (for example, 1) and the product is taken as a current update parameter. A product of a difference between the coordinates of the initial virtual space position corresponding to the first user and the coordinates of the virtual center position and the current update parameter is taken as a current virtual space position corresponding to the first user.
[0074] For example, the current virtual space position corresponding to the first user can be represented as ((x-0)*(1-η*λ), (y-0)*(1-η*λ), (z-0)*(1-η*λ)), where (x, y, z) are the coordinates of the initial virtual space position corresponding to the first user, (0, 0, 0) are the coordinates of the virtual center position, η is the number of times of calling the music detection model, and λ is the preset step length. In some embodiments, λ can be a variable step length that increases as the number of times of calling the music detection model increases.
[0075] For example, if the coordinates of the initial virtual space position corresponding to the first user are (100, 100, 0), the current virtual space position after calling the music detection model once is (100*(1-λ), 100*(1-λ), 0).
[0076] The first client can perform spatial audio processing on the audio stream according to the current virtual space position to obtain audio with a spatial audio effect, which can be played at the first client. The audio with the spatial audio effect can also be sent to the server. In the case that the audience's client receives and plays the audio with the spatial audio effect, the sound of the first user can achieve a stereophonic effect of moving from one side (or one corner) to the center. The second client can determine the current virtual space position corresponding to the first user in the same way and perform the same processing and playing on the audio stream, which will not be described here.
[0077] Through the method of the above embodiments, the sound of the user can achieve a stereophonic effect of moving from one side (or one corner) to the center in the case that only one user continuously speaks, etc. The accuracy and effectiveness of spatial audio processing on the audio are improved, and the output effect of the audio is improved.
[0078] In some embodiments, the determination result of whether to perform spatial audio processing on the audio stream obtained by the first client is synchronized to the second client through signaling, so that the second client does not perform spatial audio processing on the received audio stream of the first client, where the second client is a client of a second user who is a second user in addition to the first user in the multiple live users.
[0079] In some embodiments, in response to switching from spatial audio processing of the audio stream acquired by the first client to no spatial audio processing of the audio stream acquired by the first client, a first signaling is sent to the second client to inform the second client to perform no spatial audio processing on the audio stream of the first client, and the second client can also perform no spatial audio processing on the audio stream acquired by the second client; in response to switching from no spatial audio processing of the audio stream acquired by the first client to spatial audio processing of the audio stream acquired by the first client, a second signaling is sent to the second client to inform the second client to perform spatial audio processing on the audio stream of the first client, and the second client can also perform spatial audio processing on the audio stream acquired by the second client.
[0080] In the case where the processing manner of the audio stream acquired by the first client changes, the second client can be informed by signaling to process the received audio stream of the first client in the corresponding processing manner. The case where the processing manner of the audio stream acquired by the first client is determined for the first time also belongs to the case where the processing manner changes. The signaling is, for example, an IMP (Internet Messaging Processor) signaling.
[0081] If the first client determines to perform spatial audio processing on the audio stream acquired by the first client, but receives the signaling sent by the second client to close the spatial audio processing mode, it is finally determined to perform no spatial audio processing on the audio stream acquired by the first client. The same applies to the second client. The decision priority of performing no spatial audio processing on the audio stream acquired by the first client is higher.
[0082] The method of the above embodiments can further improve the output effect of the audio by synchronizing the first client with the second client through signaling to instruct the second client whether to perform spatial audio processing on the audio stream.
[0083] Some other embodiments of the audio processing method of the present disclosure will be described below in combination with FIG. 2.
[0084] FIG. 2 is a flowchart of some other embodiments of the audio processing method of the present disclosure. As shown in FIG. 2, the method of the embodiments includes steps S202-S213. The method of the embodiments can be executed by the first client.
[0085] In step S202, it is detected whether audio is input by the first client, and if so, step S204 is executed, otherwise, step S202 is repeatedly executed.
[0086] In step S204, consecutive audio frames are collected from the audio stream acquired by the first client, and it is determined whether the number of the consecutive audio frames reaches a threshold, and if so, step S206 is executed, otherwise, step S205 is executed.
[0087] In step S205, it is determined to perform spatial audio processing on the audio stream obtained from the first client, and the step S202 is returned to start execution again.
[0088] In step S206, it is determined whether the input source of the continuous audio frame only includes the first client, if yes, step S207 is executed, otherwise, step S208 is executed.
[0089] In step S207, the music detection model is called to detect the continuous audio frame, and it is determined whether it is music, if yes, step S211 is executed, otherwise, step S213 is executed.
[0090] In step S208, it is determined whether the energy of the audio frame corresponding to the first client is the maximum, if yes, step S210 is executed, otherwise, step S209 is executed.
[0091] In step S209, it is determined to perform spatial audio processing on the audio stream obtained from the first client. Step S213 is executed after step S209.
[0092] In step S210, the music detection model is called to detect the continuous audio frame, and it is determined whether it is music, if yes, step S211 is executed, otherwise, step S212 is executed.
[0093] In step S211, it is determined not to perform spatial audio processing on the audio stream obtained from the first client.
[0094] In step S212, it is determined to perform spatial audio processing on the audio stream obtained from the first client.
[0095] In step S213, the current virtual space position corresponding to the first user is determined according to the initial virtual space position corresponding to the first user, the number of times of calling the music detection model, and the preset step length.
[0096] The method of the above embodiment, by collecting continuous audio frames from the audio stream obtained from the first client, determining whether to call the music detection model to detect the continuous audio frames according to the number of continuous audio frames, the input source of the audio frame, and the energy information of the audio frame, and determining the type of the continuous audio frame, so as to finally determine whether to perform spatial audio processing on the audio stream obtained from the first client. The audio in the interactive live broadcast can be more accurately and effectively processed in spatial audio, and the audio output effect in the scene of the interactive live broadcast is improved.
[0097] Some other embodiments of the audio processing method of the present disclosure will be described below in conjunction with FIG. 3.
[0098] FIG. 3 is a flowchart of some other embodiments of the audio processing method of the present disclosure. As shown in FIG. 3, the method of the embodiments includes steps S301-S312.
[0099] In step S301, in response to the first user starting to play music, the first client acquires a first audio stream.
[0100] The first client is a client of the first user in the plurality of live users participating in the interactive live broadcast. For example, the first user plays background music with the aid of an auxiliary device, and stabilizes the input audio frames.
[0101] In step S302, the first client collects continuous audio frames from the first audio stream, and determines whether a condition for calling a music detection model is met. If so, step S303 is performed.
[0102] For example, the condition for calling the music detection model includes that the continuous audio frames reach a threshold, and the energy of the audio frames corresponding to the first client is maximum.
[0103] In step S303, the first client inputs the continuous audio frames into the music detection model.
[0104] In step S304, the first client receives a detection result of the continuous audio frames returned by the music detection model. If the detection result is music, steps S305-S307 are performed.
[0105] If the detection result is non-music, the solutions of the foregoing embodiments can be referred to, which are not described herein again.
[0106] In step S305, the first client sends an instruction to close a spatial audio processing function to a first spatial audio processing module corresponding to the first client.
[0107] The first spatial audio processing module does not perform spatial audio processing on the audio stream acquired by the first client.
[0108] In step S306, the first client sends a first signaling to a second client, to notify the second client to close the spatial audio processing function.
[0109] The second client is a client of a user other than the first user in the plurality of live users participating in the interactive live broadcast. For example, the first client sends an IMP signaling to the second client, to notify to close the spatial audio processing function on the performer.
[0110] In step S307, the second client sends an instruction to close the spatial audio processing function on the audio stream of the first client to a second spatial audio processing module corresponding to the second client.
[0111] In step S308, in response to the end of the first user playing music, the first client acquires a second audio stream.
[0112] For example, the first user ends the playing of background music, and no longer provides stable input audio frames.
[0113] In step S309, the first client determines whether a performance end rule is met. If the performance end rule is met, steps S310-S312 are executed.
[0114] For example, the performance end rule is that the collected continuous audio frames cannot reach a threshold value or the energy of the audio frames corresponding to the second client is maximum.
[0115] In step S310, the first client sends an instruction to turn on a spatial audio processing function to the first spatial audio processing module corresponding thereto.
[0116] In step S311, the first client sends a second signaling to the second client to inform the second client to turn on the spatial audio processing function.
[0117] For example, the second client is informed of the end of the performance through the IMP signaling.
[0118] In step S312, the second client sends an instruction to turn on the spatial audio processing function of the audio stream of the first client to the second spatial audio processing module corresponding thereto.
[0119] Next, some other embodiments of the audio processing method of the present disclosure are described in combination with FIG. 4.
[0120] Fig. 4 is a flow chart of some other embodiments of the audio processing method of the present disclosure. As shown in Fig. 4, the flow chart is for the user device to process audio for real-time communication. In the uplink flow, the audio input device of the user device collects audio and performs ADC (Analog-to-Digital Converter), AEC (Acoustic Echo Cancellation), ANS (Automatic Noise Suppression), AGC (Automatic Gain Control) and other processing. Then the method of the foregoing embodiments can be used in the music detection process to determine whether the audio data is music, and the determination result is sent to the spatial audio processing strategy module. The audio detected by the music detection part includes locally collected audio and audio sent by the remote client. In the signal modification process, mixing and other operations can be performed. Then the audio is encoded and sent to the receiving end through the network.
[0121] In the downlink flow, encoded data sent by other clients can be received through the network, and locally collected audio can also be received. After receiving the encoded data, the jitter buffer can be used. Then all the obtained audio streams are processed by the PLC (Programmable Logic Controller), and then decoded. In the signal modification process, stretching, ducking, DRC (Dynamic Range Control), spatial audio processing, mixing, mic selection and other operations can be performed. Before the signal modification, the spatial audio processing strategy module can determine whether to perform spatial audio processing on the audio according to the music detection result. After the signal modification, DAC (Digital-to-Analog Converter) processing is performed before playing.
[0122] The client as the host can upload the video or audio stream after spatial audio processing to the CDN (Content Delivery Network) server, and the CDN server can send it to the client of each viewer for playing.
[0123] In addition to the case where the user plays music in live broadcast, there are some special cases in the actual live broadcast process, which cannot accurately apply the spatial audio processing method. For these special cases, the disclosure proposes a solution, which is described as follows.
[0124] In some embodiments, in response to receiving the audio streams of the plurality of clients and the length of the received audio streams reaching a preset length, the distribution of the basic virtual space positions corresponding to the plurality of live users of the plurality of clients is determined, wherein the basic virtual space positions corresponding to the plurality of live users are the positions of the plurality of live users in the live broadcast picture mapped to the positions in the three-dimensional spatial audio space; and whether to perform spatial audio processing on the audio streams of the plurality of clients is determined according to the distribution of the basic virtual space positions corresponding to the plurality of live users of the plurality of clients.
[0125] In some embodiments, in response to determining that the basic virtual space positions corresponding to the plurality of live users of the plurality of clients are distributed at different positions left and right of the center of the three-dimensional spatial audio space, it is determined to perform spatial audio processing on the audio streams of the plurality of clients; and in response to determining that the basic virtual space positions corresponding to the plurality of live users of the plurality of clients are not distributed at different positions left and right of the center of the three-dimensional spatial audio space, it is determined not to perform spatial audio processing on the audio streams of the plurality of clients.
[0126] In the layout of some live broadcast pictures, the left position or the right position is used to display fixed users. If the host and the fixed users are both located on the left or the right, when they have a conversation, if spatial audio processing is performed, it will cause the conversation to always sound on the left or the right, which reduces the playing effect. Therefore, in the case where the basic virtual space positions corresponding to the live users are distributed at different positions left and right of the center of the three-dimensional spatial audio space, the spatial audio effect is played for the audio streams reaching the preset length. In the process of spatial audio processing, the client will first obtain the basic virtual space positions corresponding to each live user, and then perform optimization processing such as normalization on the basic space positions to obtain the virtual space positions of each live user, and then obtain the audio with spatial audio effect. The specific process is not described again. The virtual space positions and the current virtual space positions in the foregoing embodiments can be the virtual space positions after optimization processing.
[0127] In some embodiments, in response to detecting that only the first client supports audio input, it is determined not to perform spatial audio processing on the audio stream obtained by the first client.
[0128] In some embodiments, the case where only the first client supports audio input includes at least one of the following: only the first client enables the audio input function, the first client enables the function of prohibiting the audio input of the second client, and only the first user participates in the interactive live broadcast.
[0129] For example, when a live user (e.g., a host) mutes the sound of a live room, only the live user in the live room can output audio. If spatial audio processing is performed on the audio of the live user, the sound will be located on one side (fixed direction), which affects the output effect of the sound. By detecting whether the function of prohibiting the audio input of the second client is enabled on the first client, it can be determined whether only the first client supports audio input, and thus whether to perform spatial audio processing on the audio stream obtained by the first client.
[0130] For example, in some live rooms, there will be a large amount of time when only one live user (e.g., a host) is live, or only one live user enables the audio input function, and other live users all disable the audio input function (mute). If spatial audio processing is performed on the audio of the live user, the sound will be located on one side (fixed direction), which affects the output effect of the sound. Therefore, by detecting whether only the first client enables the audio input function or only the first user participates in the interactive live, it is determined whether to perform spatial audio processing on the audio stream obtained by the first client.
[0131] The method of the above embodiment provides a solution to the problem that spatial audio processing may cause the output effect to be poor in some special cases, which can more accurately and effectively perform spatial audio processing on the audio, improve the output effect of the audio, and reduce the overall overhead, save resources, and improve the processing efficiency of the audio.
[0132] The present disclosure also provides an audio processing apparatus, which is described below in conjunction with FIG. 5.
[0133] FIG. 5 is a structural diagram of an audio processing apparatus according to some embodiments of the present disclosure. As shown in FIG. 5, the audio processing apparatus 50 according to the embodiment includes a collection module 510, a classification module 520, and a spatial audio processing module 530. The audio processing apparatus can be located in the first client.
[0134] The collection module 510 is configured to collect audio data from an audio stream obtained by the first client, where the first client is a client of a first user in a plurality of live users participating in an interactive live.
[0135] In some embodiments, the collection module 510 is configured to, in response to detecting that audio is input by the first client, collect continuous audio frames from the audio stream obtained by the first client until the number of continuous audio frames reaches a threshold, and take the continuous audio frames as the audio data.
[0136] The classification module 520 is configured to determine the type of the audio data according to the audio data.
[0137] In some embodiments, the type of the audio data includes music and non-music.
[0138] In some embodiments, the classification module 520 is configured to determine whether to invoke the music detection model to detect the audio data according to a feature of the audio data, determine the type of the audio data, wherein the feature of the audio data comprises at least one of an input source of an audio frame in the audio data, and energy information of the audio frame.
[0139] In some embodiments, the classification module 520 is configured to, in response to detecting that the input source of the audio frame comprises a first client and a second client, determine energy of the audio frame corresponding to the first client and energy of the audio frame corresponding to the second client, wherein the second client is a client of a second user in the plurality of live users other than the first user; and in response to determining that the energy of the audio frame corresponding to the first client is greater than or equal to the energy of the audio frame corresponding to the second client, determine to invoke the music detection model to detect the audio data.
[0140] In some embodiments, the classification module 520 is configured to, in response to the energy of the audio frame corresponding to the second client being greater than the energy of the audio frame corresponding to the first client, determine not to invoke the music detection model to detect the audio data; and in response to the energy of the audio frame corresponding to the second client being greater than the energy of the audio frame corresponding to the first client, determine that the type of the audio data is non-music.
[0141] In some embodiments, the classification module 520 is configured to, in response to detecting that the input source of the audio frame only comprises the first client, determine to invoke the music detection model to detect the audio data.
[0142] In some embodiments, the classification module 520 is configured to determine the type of the audio data by detecting the audio data using the music detection model.
[0143] The spatial audio processing module 530 is configured to determine whether to perform spatial audio processing on the audio stream obtained by the first client according to the type of the audio data.
[0144] In some embodiments, the spatial audio processing module 530 is configured to, in response to the number of continuous audio frames not reaching a threshold, determine to perform spatial audio processing on the audio stream obtained by the first client.
[0145] In some embodiments, in response to determining that the type of the audio data is non-music, the spatial audio processing module 530 is further configured to, in response to detecting that the input source of the audio frame only includes the first client, adjust the current virtual space position corresponding to the first user according to the initial virtual space position corresponding to the first user, the number of times of invoking the music detection model, and a preset step length, so that the current virtual space position corresponding to the first user gradually approaches the virtual center position until the difference between the current virtual space position corresponding to the first user and the virtual center position is not greater than a preset difference, wherein the initial virtual space position corresponding to the first user is a position in a three-dimensional spatial audio space to which a position of the first user in the live picture is mapped, and the virtual center position is a center position of the three-dimensional spatial audio space.
[0146] In some embodiments, the spatial audio processing module 530 is configured to adjust the current virtual space position corresponding to the first user according to the difference between the coordinates of the initial virtual space position corresponding to the first user and the coordinates of the virtual center position, the number of times of invoking the music detection model, and a preset step length, wherein the difference between the coordinates of the current virtual space position corresponding to the first user and the coordinates of the virtual center position is inversely proportional to the number of times of invoking the music detection model.
[0147] In some embodiments, the spatial audio processing module 530 is configured to, in response to determining that the type of the audio data is music, determine not to perform spatial audio processing on the audio stream obtained by the first client; and in response to determining that the type of the audio data is non-music, determine to perform spatial audio processing on the audio stream obtained by the first client.
[0148] In some embodiments, the audio processing apparatus further includes a sending module 540 configured to synchronize, through signaling, a determination result of whether to perform spatial audio processing on the audio stream obtained by the first client to a second client, so that the second client does not perform spatial audio processing on the received audio stream of the first client, wherein the second client is a client of a second user other than the first user among the multiple live users.
[0149] In some embodiments, the spatial audio processing module 530 is configured to, in response to receiving the audio streams of the multiple clients and the length of the received audio streams reaching a preset length, determine the distribution of the basic virtual space positions corresponding to the multiple live users of the multiple clients, wherein the basic virtual space positions corresponding to the multiple live users are positions in a three-dimensional spatial audio space to which positions of the multiple live users in a live picture are mapped; and determine whether to perform spatial audio processing on the audio streams of the multiple clients according to the distribution of the basic virtual space positions corresponding to the multiple live users of the multiple clients.
[0150] In some embodiments, the playing module 550 is configured to determine to perform spatial audio processing on the audio streams of the plurality of clients in response to determining that the plurality of live users of the plurality of clients correspond to the basic virtual space positions distributed at different positions left and right of the center of the three-dimensional spatial audio space; and determine not to perform spatial audio processing on the audio streams of the plurality of clients in response to determining that the plurality of live users of the plurality of clients correspond to the basic virtual space positions not distributed at different positions left and right of the center of the three-dimensional spatial audio space.
[0151] In some embodiments, the spatial audio processing module 530 is further configured to determine not to perform spatial audio processing on the audio stream acquired by the first client in response to detecting that only the first client supports audio input.
[0152] In some embodiments, the case that only the first client supports audio input includes at least one of only the first client enabling an audio input function, the first client enabling a function of prohibiting audio input of the second client, and only the first user participating in the interactive live broadcast.
[0153] It should be noted that each unit (module) described above is only a logical division according to the specific function implemented by it, and is not used to limit the specific implementation manner, for example, it can be implemented in software, hardware or a combination of software and hardware. In actual implementation, each unit described above can be implemented as an independent physical entity, or can also be implemented by a single entity (for example, a processor (CPU or DSP, etc.), an integrated circuit, etc.). In addition, each unit described above is indicated by a dashed line in the drawings, indicating that these units can not actually exist, and the operations / functions implemented by them can be implemented by the processing circuit itself.
[0154] In addition, although not shown, the device can also include a memory, which can store various information generated by the device, the units included in the device in operation, programs and data for operation, data to be transmitted by the communication unit, etc. The memory can be a volatile memory and / or a non-volatile memory. For example, the memory can include, but is not limited to, a random access memory (RAM), a dynamic random access memory (DRAM), a static random access memory (SRAM), a read-only memory (ROM), a flash memory. Of course, the memory can also be located outside the device. Alternatively, although not shown, the device can also include a communication unit, which can be used for communication with other devices. In one example, the communication unit can be implemented in a suitable manner known in the art, for example, including communication components such as an antenna array and / or a radio frequency link, various types of interfaces, communication units, etc. Here will not be described in detail. In addition, the device can also include other components not shown, such as radio frequency links, baseband processing units, network interfaces, processors, controllers, etc. Here will not be described in detail.
[0155] Some embodiments of the present disclosure also provide an electronic device (which can be an audio processing apparatus implementing the audio processing method of any of the preceding embodiments). FIG. 6 shows a block diagram of some embodiments of the electronic device of the present disclosure. For example, in some embodiments, the electronic device 6 can be various types of devices, for example, can include but is not limited to mobile terminals such as mobile phones, notebook computers, digital broadcast receivers, PDA (Personal Digital Assistants), PAD (Tablet PC), PMP (Portable Multimedia Player), car terminals (e.g., car navigation terminals), and the like, and fixed terminals such as digital TV, desktop computers, and the like. For example, the electronic device 6 can include a display panel for displaying data and / or execution results utilized in the scheme of the present disclosure. For example, the display panel can be various shapes, for example, a rectangular panel, an oval panel, or a polygonal panel, and the like. In addition, the display panel can not only be a flat panel, but also a curved panel, or even a spherical panel.
[0156] As shown in FIG. 6, the electronic device 6 of this embodiment includes a memory 61 and a processor 62 coupled to the memory 61. It should be noted that the components of the electronic device 60 shown in FIG. 6 are only exemplary and not limiting, and the electronic device 60 can also have other components according to actual application needs. The processor 62 can control other components in the electronic device 6 to perform the desired functions.
[0157] In some embodiments, the memory 61 is configured to store one or more computer readable instructions. When the processor 62 executes the computer readable instructions, the computer readable instructions are executed by the processor 62 to implement the audio processing method according to any of the preceding embodiments. For specific implementation of each step of the method and related explanations, please refer to the above embodiments, and repeated parts will not be described here.
[0158] For example, the processor 62 and the memory 61 can directly or indirectly communicate with each other. For example, the processor 62 and the memory 61 can communicate through a network. The network can include a wireless network, a wired network, and / or any combination of a wireless network and a wired network. The processor 62 and the memory 61 can also communicate with each other through a system bus, and the present disclosure does not limit this.
[0159] For example, the processor 62 can be embodied as various appropriate processors, processing devices, and the like, such as a central processing unit (CPU), a graphics processing unit (GPU), a network processing unit (NP), and the like; also can be a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic device, a discrete gate or transistor logic device, a discrete hardware component. The central processing unit (CPU) can be an X86 or ARM architecture, and the like. For example, the memory 61 can include any combination of various forms of computer readable storage media, such as a volatile memory and / or a non-volatile memory. The memory 61 may, for example, include a system memory, which stores, for example, an operating system, application programs, a boot loader, a database, and other programs, and the like. Various application programs and various data, and the like, can also be stored in the storage medium.
[0160] In addition, according to some embodiments of the present disclosure, various operations / processes according to the present disclosure, in the case of being implemented by software and / or firmware, programs constituting the software can be installed from a storage medium or a network to a computer system having a dedicated hardware structure, such as the computer system 700 shown in FIG. 7, which, when various programs are installed, is capable of performing various functions, including functions such as the foregoing, and the like. FIG. 7 is a block diagram showing an example structure of a computer system that can be employed according to embodiments of the present disclosure.
[0161] In FIG. 7, the central processing unit (CPU) 701 performs various processing according to programs stored in a read only memory (ROM) 702 or programs loaded from a storage section 708 to a random access memory (RAM) 703. In the RAM 703, data required when the CPU 701 performs various processing, and the like, is also stored as necessary. The central processing unit is merely exemplary, and can also be other types of processors, such as the various processors described above. The ROM 702, the RAM 703, and the storage section 708 can be various forms of computer readable storage media, as follows. Note that, although the ROM 702, the RAM 703, and the storage device 708 are shown separately in FIG. 7, one or more of them can be combined or located in the same or different memory or storage module.
[0162] The CPU 701, the ROM 702, and the RAM 703 are connected to each other via a bus 704. An input / output interface 705 is also connected to the bus 704.
[0163] The following components are connected to the input / output interface 705: an input portion 706, such as a touch panel, a touch pad, a keyboard, a mouse, an image sensor, a microphone, an accelerometer, a gyroscope, and the like; an output portion 707, including a display, such as a cathode ray tube (CRT), a liquid crystal display (LCD), a speaker, a vibrator, and the like; a storage portion 708, including a hard disk, a magnetic tape, and the like; and a communication portion 709, including a network interface card, such as a LAN card, a modem, and the like. The communication portion 709 allows communication processing to be performed via a network, such as the Internet. It is easily understood that, although the respective devices or modules in the electronic device 700 are shown in FIG. 7 as communicating through the bus 704, they can also communicate through a network or other means, where the network can include a wireless network, a wired network, and / or any combination of a wireless network and a wired network.
[0164] The drive 710 is also connected to the input / output interface 705 as necessary. A removable medium 711, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, and the like, is attached to the drive 710 as necessary, so that a computer program read therefrom is installed in the storage portion 708 as necessary.
[0165] In the case where the above series of processes are implemented by software, the program constituting the software can be installed from a network such as the Internet or a storage medium such as the removable medium 711.
[0166] According to an embodiment of the present disclosure, the processes described above with reference to the flowcharts can be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product including a computer program carried on a computer-readable medium, the computer program containing program codes for executing the methods shown in the flowcharts. In such an embodiment, the computer program can be downloaded and installed from a network by the communication portion 709, or installed from the storage portion 708, or installed from the ROM 702. When the computer program is executed by the CPU 701, the above-described functions defined in the methods of the embodiments of the present disclosure are executed.
[0167] Note that in the context of the present disclosure, a computer readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. The computer readable medium can be a computer readable signal medium or a computer readable storage medium or any combination thereof. The computer readable storage medium can be, for example, but not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of the computer readable storage medium can include, but are not limited to, an electrical connection having one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing. In the present disclosure, a computer readable storage medium can be any tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. In the present disclosure, a computer readable signal medium can include a propagated data signal with computer readable program code embodied therein, in baseband or as part of a carrier wave. Such a propagated signal can take any of a variety of forms, including, but not limited to, electro-magnetic, optical, or any suitable combination thereof. A computer readable signal medium can be any computer readable medium that is not a computer readable storage medium and that can communicate, propagate, or transport a program for use by or in connection with an instruction execution system, apparatus, or device. Program code embodied on a computer readable medium can be transmitted using any appropriate medium, including but not limited to wireless, wireline, optical fiber cable, RF, etc., or any suitable combination of the foregoing.
[0168] The computer readable medium described above can be included in the electronic device described above; or can exist separately from the electronic device and be not assembled into the electronic device.
[0169] In some embodiments, a computer program is also provided, comprising instructions which, when executed by a processor, cause the processor to perform the method of any one of the embodiments described above. For example, the instructions can be embodied as computer program code.
[0170] The present disclosure also provides a user device, which is described below in connection with Fig. 8.
[0171] Fig. 8 is a structural diagram of a user device according to some embodiments of the present disclosure. As shown in Fig. 8, the user device 8 of this embodiment comprises the audio processing apparatus 50 of any one of the embodiments described above, and
[0172] The audio input device 82 is configured to collect audio input to the first client. The audio input device can be a microphone or the like.
[0173] In some embodiments, the user device 8 can further comprise an audio output device 84 configured to play received audio. The audio output device 84 can be a speaker or the like.
[0174] In embodiments of the present disclosure, computer program code for carrying out operations of the present disclosure can be written in one or more programming languages, including object oriented programming languages, such as Java, Smalltalk, C++, and conventional procedural programming languages, such as the "C" programming language or similar programming languages. The program code can execute entirely on the user's computer, partly on the user's computer, as a stand-alone software package, partly on the user's computer and partly on a remote computer, or entirely on the remote computer or server. In the latter scenario, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection can be made to an external computer (for example, through the Internet using an Internet Service Provider).
[0175] The flow diagrams and the block diagrams in the drawings are illustrations of possible architectures, functions, and operations for systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flow diagrams or block diagrams can represent a module, a segment, or a portion of code, which comprises one or more executable instructions for implementing the specified logical function(s). It should also be noted that in some alternative implementations, the functions noted in the blocks can occur out of the order noted in the figures. For example, two blocks shown in succession may, in fact, be executed substantially concurrently or the blocks can sometimes be executed in the reverse order, depending upon the functionality involved. It will also be noted that each block of the block diagrams and / or flow diagrams, and combinations thereof, can be implemented by special purpose hardware-based systems that perform the specified functions or operations, or combinations of special purpose hardware and computer instructions.
[0176] The modules, components or units described in the embodiments of the present disclosure can be implemented by software or by hardware. In some cases, the name of the module, component or unit does not constitute a limitation on the module, component or unit itself.
[0177] The functionality described herein above can be performed, at least in part, by one or more hardware logic components. For example, and without limitation, example hardware logic components that can be used include field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), system-on-a-chip (SOCs), complex programmable logic devices (CPLDs), etc.
[0178] According to some embodiments of the present disclosure, an audio processing method is provided, comprising: collecting audio data from an audio stream obtained by a first client, wherein the first client is a client of a first user in a plurality of live users participating in an interactive live broadcast; determining a type of the audio data according to the audio data; and determining whether to perform spatial audio processing on the audio stream obtained by the first client according to the type of the audio data.
[0179] In some embodiments, the collecting of the audio data from the audio stream obtained by the first client comprises: in response to detecting that audio is input by the first client, collecting continuous audio frames from the audio stream obtained by the first client until a quantity of the continuous audio frames reaches a threshold, and taking the continuous audio frames as the audio data.
[0180] In some embodiments, the type of the audio data comprises music and non-music, and the determining of the type of the audio data according to the audio data comprises: determining whether to call a music detection model to detect the audio data according to a feature of the audio data, and determining the type of the audio data, wherein the feature of the audio data comprises at least one of an input source of an audio frame in the audio data and energy information of the audio frame.
[0181] In some embodiments, the determining of whether to call the music detection model to detect the audio data according to the feature of the audio data comprises: in response to detecting that the input source of the audio frame comprises a first client and a second client, determining energy of an audio frame corresponding to the first client and energy of an audio frame corresponding to the second client, wherein the second client is a client of a second user in the plurality of live users other than the first user; and in response to determining that the energy of the audio frame corresponding to the first client is greater than or equal to the energy of the audio frame corresponding to the second client, determining to call the music detection model to detect the audio data.
[0182] In some embodiments, the determining of whether to call the music detection model to detect the audio data according to the feature of the audio data comprises: in response to the energy of the audio frame corresponding to the second client being greater than the energy of the audio frame corresponding to the first client, determining not to call the music detection model to detect the audio data; and the determining of the type of the audio data according to the audio data further comprises: in response to the energy of the audio frame corresponding to the second client being greater than the energy of the audio frame corresponding to the first client, determining that the type of the audio data is non-music.
[0183] In some embodiments, the determining whether to invoke the music detection model to detect the audio data according to the characteristic of the audio data comprises: in response to detecting that the input source of the audio frame only includes the first client, determining to invoke the music detection model to detect the audio data.
[0184] In some embodiments, the audio processing method further comprises: in response to the number of continuous audio frames not reaching the threshold, determining to perform spatial audio processing on the audio stream obtained by the first client.
[0185] In some embodiments, the determining the type of the audio data according to the audio data comprises: detecting the audio data by using the music detection model to determine the type of the audio data.
[0186] In some embodiments, in response to determining that the type of the audio data is non-music, the determining to perform spatial audio processing on the audio stream obtained by the first client further comprises: in response to detecting that the input source of the audio frame only includes the first client, determining a current virtual space position corresponding to the first user according to an initial virtual space position corresponding to the first user, a number of times of invoking the music detection model, and a preset step length, so that the current virtual space position corresponding to the first user gradually approaches a virtual center position until a difference between the current virtual space position corresponding to the first user and the virtual center position is not greater than a preset difference, wherein the initial virtual space position corresponding to the first user is a position of a position of the first user in the live picture mapped to a three-dimensional spatial audio space, and the virtual center position is a center position of the three-dimensional spatial audio space.
[0187] In some embodiments, the adjusting the current virtual space position corresponding to the first user according to the initial virtual space position corresponding to the first user, the number of times of invoking the music detection model, and the preset step length comprises: adjusting the current virtual space position corresponding to the first user according to a difference between coordinates of the initial virtual space position corresponding to the first user and coordinates of the virtual center position, the number of times of invoking the music detection model, and the preset step length, wherein the difference between the coordinates of the current virtual space position corresponding to the first user and the coordinates of the virtual center position is inversely proportional to the number of times of invoking the music detection model.
[0188] In some embodiments, the determining whether to process the audio stream obtained by the first client by using the spatial audio processing mode according to the type of the audio data comprises: in response to determining that the type of the audio data is music, determining not to perform spatial audio processing on the audio stream obtained by the first client; and in response to determining that the type of the audio data is non-music, determining to perform spatial audio processing on the audio stream obtained by the first client.
[0189] In some embodiments, the audio processing method further includes: synchronizing, by signaling, a determination result of whether to perform spatial audio processing on the audio stream acquired by the first client to the second client, so that the second client does not perform spatial audio processing on the received audio stream of the first client, wherein the second client is a client of a second user in the plurality of live users other than the first user.
[0190] In some embodiments, the audio processing method further includes: in response to receiving the audio streams of the plurality of clients and the length of the received audio streams reaching a preset length, determining a distribution of the basis virtual spatial positions corresponding to the plurality of live users of the plurality of clients, wherein the basis virtual spatial positions corresponding to the plurality of live users are positions in a three-dimensional spatial audio space mapped from positions of the plurality of live users in a live picture; and determining whether to perform spatial audio processing on the audio streams of the plurality of clients according to the distribution of the basis virtual spatial positions corresponding to the plurality of live users of the plurality of clients.
[0191] In some embodiments, determining whether to perform spatial audio processing on the audio streams of the plurality of clients according to the distribution of the basis virtual spatial positions corresponding to the plurality of live users of the plurality of clients includes: in response to determining that the basis virtual spatial positions corresponding to the plurality of live users of the plurality of clients are distributed at different positions left and right of the center of the three-dimensional spatial audio space, determining to perform spatial audio processing on the audio streams of the plurality of clients; and in response to determining that the basis virtual spatial positions corresponding to the plurality of live users of the plurality of clients are not distributed at different positions left and right of the center of the three-dimensional spatial audio space, determining not to perform spatial audio processing on the audio streams of the plurality of clients.
[0192] In some embodiments, the audio processing method further includes: in response to detecting that only the first client supports audio input, determining not to perform spatial audio processing on the audio stream acquired by the first client.
[0193] In some embodiments, the case where only the first client supports audio input includes at least one of: only the first client has enabled an audio input function, the first client has enabled a function of prohibiting audio input of the second client, and only the first user participates in the interactive live broadcast.
[0194] According to some other embodiments of the present disclosure, an audio processing apparatus is provided, including: an acquisition module configured to acquire audio data from an audio stream acquired by a first client, wherein the first client is a client of a first user in a plurality of live users participating in an interactive live broadcast; a classification module configured to determine a type of the audio data according to the audio data; and a spatial audio processing module configured to determine whether to perform spatial audio processing on the audio stream acquired by the first client according to the type of the audio data.
[0195] According to still some embodiments of the present disclosure, an electronic device is provided, comprising: a memory; and a processor coupled to the memory, the processor configured to perform the audio processing method of any of the embodiments of the present disclosure based on instructions stored in the memory.
[0196] According to yet some embodiments of the present disclosure, a user device is provided, comprising: the audio processing apparatus of any of the preceding embodiments; and an audio input device configured to collect audio, input the first client.
[0197] In some embodiments, the user device further comprises an audio output device configured to play the received audio.
[0198] According to still some embodiments of the present disclosure, a computer readable storage medium is provided, having stored thereon a computer program, the program being executed by a processor to perform the audio processing method of any of the embodiments of the present disclosure.
[0199] According to yet some embodiments of the present disclosure, a computer program is provided, comprising: instructions that, when executed by a processor, cause the processor to perform the audio processing method of any of the embodiments of the present disclosure.
[0200] According to some embodiments of the present disclosure, a computer program product is provided, comprising instructions that, when executed by a processor, implement the audio processing method of any of the embodiments of the present disclosure.
[0201] The above description is merely some embodiments of the present disclosure and a description of the principles of the applied technology. Those skilled in the art should understand that the scope of the disclosure involved in the present disclosure is not limited to the technical solutions formed by the specific combinations of the above technical features, and should also cover other technical solutions formed by any combinations of the above technical features or their equivalent features without departing from the above disclosed concept. For example, the above features are replaced with each other to form technical solutions with similar functions disclosed in the present disclosure (but not limited to).
[0202] In the description provided herein, numerous specific details are set forth. However, it is understood that embodiments of the application can be practiced without these specific details. In other instances, well-known methods, structures and techniques have not been shown in detail in order not to obscure the understanding of this description.
[0203] Moreover, while operations are depicted in a particular order, this should not be understood as requiring such an order nor imposing a sequential order of executed operations. In certain circumstances, multitasking and parallel processing can be advantageous. Likewise, while several specific implementation details are contained in the above discussion, these should not be taken as limiting the scope of the present disclosure. Certain features described in the context of separate embodiments can also be implemented together in a single embodiment. Conversely, various features described in the context of a single embodiment can also be implemented in multiple embodiments separately or in any suitable sub-combination.
[0204] While certain aspects of the disclosure have been described with reference to particular embodiments, those skilled in the art will understand that the disclosure is illu- stative only and is not intended to be limiting. Many modifications and variations are possible in light of the above teachings. It is contemplated that, where any component or module is implemented in the embodiments described herein, there are a number of alternative ways to implement the same. The elements and acts of the various embodiments described above can be combined to provide further implementations. It is also contemplated that the components can be implemented using software, hardware, firmware, or combinations thereof that can be implemented in software and / or firmware. Additionally, one of ordinary skill will recognize that a plurality of hardware components can be implemented as software modules or components, or vice versa. Other structure and / or functionality can be combined, some can be split into separate components, and / or some can be omitted entirely. As such, non-limiting examples of the range of possible implementations have been set forth above. It is therefore contemplated to this end that the disclosure in its broader aspect is not limited to the specific details, representative apparatus, and illustrative examples shown and described above. Changes in form and detail can be made without departing from the spirit, and the general scope of equivalents of the present disclosure. Therefore, although the disclosure has been described in some detail with reference only to a few embodiments, those skilled in the art will understand that variations and modifications will occur to them to which they are intended to be included within the spirit and scope of the disclosure. Accordingly, all such modifications and variations are intended to be included within the scope of the appended claims. The disclosure has been provided with the understanding that embodiments of the disclosure are defined by the following claims, and equivalents thereto.
Claims
1. An audio processing method, comprising: Collecting audio data from an audio stream obtained by a first client, wherein the first client is a client of a first user among a plurality of live broadcast users participating in the interactive live broadcast; Determining the type of the audio data according to the audio data; Determine whether to perform spatial audio processing on the audio stream obtained by the first client according to the type of the audio data.
2. The audio processing method according to claim 1, wherein: The collecting audio data from the audio stream obtained from the first client includes: In response to detecting audio input from the first client, continuous audio frames are collected from the audio stream obtained from the first client until the number of the continuous audio frames reaches a threshold, and the continuous audio frames are used as the audio data.
3. The audio processing method according to claim 2, wherein: The type of the audio data includes music and non-music, and determining the type of the audio data according to the audio data includes: Based on the characteristics of the audio data, determine whether to call a music detection model to detect the audio data and determine the type of the audio data, wherein the characteristics of the audio data include at least one of the input source of the audio frame in the audio data and the energy information of the audio frame.
4. The audio processing method according to claim 3, wherein: The determining, based on the characteristics of the audio data, whether to call a music detection model to detect the audio data includes: In response to detecting that the input source of the audio frame includes the first client and the second client, determining the energy of the audio frame corresponding to the first client and the energy of the audio frame corresponding to the second client, wherein the second client is a client of a second user other than the first user among the multiple live broadcast users; In response to determining that the energy of the audio frame corresponding to the first client is greater than or equal to the energy of the audio frame corresponding to the second client, it is determined to call a music detection model to detect the audio data.
5. The audio processing method according to claim 4, wherein: The determining, based on the characteristics of the audio data, whether to call a music detection model to detect the audio data includes: In response to the energy of the audio frame corresponding to the second client being greater than the energy of the audio frame corresponding to the first client, determining not to call a music detection model to detect the audio data; The determining the type of the audio data according to the audio data further includes: In response to the energy of the audio frame corresponding to the second client being greater than the energy of the audio frame corresponding to the first client, the type of the audio data is determined to be non-music.
6. The audio processing method according to any one of claims 3 to 5, wherein: The determining, based on the characteristics of the audio data, whether to call a music detection model to detect the audio data includes: In response to detecting that the input source of the audio frame only includes the first client, it is determined to call a music detection model to detect the audio data.
7. The audio processing method according to any one of claims 2 to 6, further comprising: In response to the number of the continuous audio frames not reaching a threshold, determining to perform spatial audio processing on the audio stream obtained by the first client.
8. The audio processing method according to any one of claims 2 to 7, wherein: Determining the type of the audio data according to the audio data includes: The audio data is detected using a music detection model to determine the type of the audio data.
9. The audio processing method according to any one of claims 3 to 8, wherein: In response to determining that the type of the audio data is non-music, determining to perform spatial audio processing on the audio stream obtained by the first client, the audio processing method further includes: In response to detecting that the input source of the audio frame only includes the first client, the current virtual space position corresponding to the first user is adjusted according to the initial virtual space position corresponding to the first user, the number of times the music detection model is called and the preset step size, so that the current virtual space position corresponding to the first user gradually approaches the virtual center position until the difference between the current virtual space position corresponding to the first user and the virtual center position is no more than the preset difference, wherein the initial virtual space position corresponding to the first user is the position of the first user in the live broadcast picture mapped to the position in the three-dimensional space audio space, and the virtual center position is the center position of the three-dimensional space audio space.
10. The audio processing method according to claim 9, wherein: The adjusting the current virtual space position corresponding to the first user according to the initial virtual space position corresponding to the first user, the number of times the music detection model is called, and the preset step size includes: The current virtual space position corresponding to the first user is adjusted based on the difference between the coordinates of the initial virtual space position corresponding to the first user and the coordinates of the virtual center position, the number of times the music detection model is called and the preset step size, wherein the difference between the current virtual space position corresponding to the first user and the coordinates of the virtual center position is inversely proportional to the number of times the music detection model is called.
11. The audio processing method according to any one of claims 1 to 10, wherein: The determining, according to the type of the audio data, whether to use a spatial audio processing mode to process the audio stream obtained by the first client includes: In response to determining that the type of the audio data is music, determining not to perform spatial audio processing on the audio stream obtained by the first client; In response to determining that the type of the audio data is non-music, it is determined to perform spatial audio processing on the audio stream obtained by the first client.
12. The audio processing method according to any one of claims 1 to 11, further comprising: Through signaling, a determination result of whether to perform spatial audio processing on the audio stream obtained by the first client is synchronized to the second client, so that the second client does not perform spatial audio processing on the audio stream received from the first client, wherein the second client is a client of a second user other than the first user among multiple live broadcast users.
13. The audio processing method according to any one of claims 1 to 12, further comprising: In response to receiving audio streams from multiple clients and the length of the received audio streams reaching a preset length, determining a distribution of basic virtual space positions corresponding to multiple live broadcast users of the multiple clients, wherein the basic virtual space positions corresponding to the multiple live broadcast users are positions of the multiple live broadcast users in the live broadcast image mapped to positions in the three-dimensional audio space; Whether to perform spatial audio processing on the audio streams of the multiple clients is determined according to the distribution of basic virtual space positions corresponding to the multiple live broadcast users of the multiple clients. The audio processing method according to claim 13 , wherein: The determining whether to perform spatial audio processing on the audio streams of the multiple clients according to the distribution of basic virtual space positions corresponding to the multiple live broadcast users of the multiple clients includes: In response to determining that basic virtual space positions corresponding to multiple live broadcast users of the multiple clients are distributed in different directions to the left and right of the center of the three-dimensional spatial audio space, determining to perform spatial audio processing on the audio streams of the multiple clients; In response to determining that basic virtual space positions corresponding to multiple live broadcast users of the multiple clients are not distributed in different left and right directions of the center of the three-dimensional spatial audio space, it is determined not to perform spatial audio processing on the audio streams of the multiple clients.
15. The audio processing method according to any one of claims 1 to 14, further comprising: In response to detecting that only the first client supports audio input, it is determined not to perform spatial audio processing on the audio stream obtained by the first client. The audio processing method according to claim 15 , wherein: The situation where only the first client supports audio input includes: only the first client turns on the audio input function, the first client turns on the function of prohibiting the second client's audio input, and only the first user participates in the interactive live broadcast.
17. An audio processing device, comprising: a collection module configured to collect audio data from an audio stream obtained by a first client, wherein the first client is a client of a first user among a plurality of live broadcast users participating in the interactive live broadcast; a classification module, configured to determine a type of the audio data based on the audio data; The spatial audio processing module is configured to determine whether to perform spatial audio processing on the audio stream obtained by the first client according to the type of the audio data.
18. An electronic device comprising: processor; as well as A memory coupled to the processor, for storing instructions, wherein when the instructions are executed by the processor, the processor executes the audio processing method according to any one of claims 1 to 16.
19. A user equipment comprising: The audio processing device according to claim 17; as well as The audio input device is configured to collect audio and input the audio into the first client.
20. A computer-readable storage medium having a computer program stored thereon, wherein: When the program is executed by a processor, the steps of the audio processing method according to any one of claims 1 to 16 are implemented.
21. A computer program product comprising: Instruction, wherein when the instruction is executed by a processor, the audio processing method according to any one of claims 1 to 16 is implemented.
22. A computer program comprising: Instruction, wherein when the instruction is executed by a processor, the audio processing method according to any one of claims 1 to 16 is implemented.
Citation Information
Patent Citations
Audio stream processing method and mobile terminal
CN106126165A
Sound setting method and system
CN106658219A
Audio denoising method and device, electronic equipment and storage medium
CN112908352A
Audio processing method and device, electronic equipment and computer readable storage medium
CN114339297A
Audio playing method and device, storage medium, client and live broadcast system
CN115237250A