Audio signal processing method and conference equipment

By identifying audio signals from multiple channels, determining the target audio source category, and adopting corresponding processing strategies, the problem of uneven audio signal processing under the same environment is solved, achieving precise audio signal processing and effect enhancement.

CN121768418APending Publication Date: 2026-03-31ARASHI VISION INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411383708.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-09-30
Publication Date
2026-03-31

AI Technical Summary

Technical Problem

In existing technologies, the same processing strategy is often used for audio signals in the same environment, resulting in poor processing of audio signals from different sources, especially due to the uneven processing effect caused by the complexity and diversity of audio source identification.

Method used

By identifying audio signals from multiple channels, the target audio source category is determined, and corresponding audio processing strategies are adopted based on the audio source category, including adaptive echo cancellation, adaptive noise reduction, and adaptive gain control, to accurately process different audio source types.

Benefits of technology

It improves the accuracy of sound source category recognition, enhances the processing effect of different types of audio signals, such as sound enhancement or noise reduction, and improves the user interactivity and processing efficiency of conference equipment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121768418A_ABST
    Figure CN121768418A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses an audio signal processing method and conference equipment. The method comprises the following steps: acquiring a target audio signal; determining a corresponding target sound source category based on the target audio signal, wherein the target sound source category is obtained by identifying audio signals of a plurality of channels; and processing the target audio signal by adopting an audio processing strategy corresponding to the target sound source category. According to the audio signal processing method, the accuracy of sound source category recognition can be improved, further, the corresponding audio processing strategy is determined for the recognized target sound source category, and the processing effect of different categories of audio signals can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to, but is not limited to, the field of voice recognition technology, and particularly to an audio signal processing method and a conferencing device. Background Technology

[0002] In related technologies, the same processing strategy is often used for audio signals in the same environment. However, even in the same environment, there are often many different sound sources. If the same processing strategy is used for all of them, some audio signals will inevitably not be processed well.

[0003] Although existing technologies propose differentiating between different audio sources and then applying different noise reduction methods to each, audio source identification itself is complex. A single audio source often contains multiple audio signals, making accurate source identification a challenge. Current technologies require manual selection of the audio source type before processing, which is clearly not intelligent enough. Summary of the Invention

[0004] This application provides an audio signal processing method and a conference device.

[0005] According to a first aspect of the embodiments of this application, an audio signal processing method is provided, the method being applied in a conference device, the method comprising:

[0006] Acquire the target audio signal;

[0007] The target audio source category is determined based on the target audio signal, and the target audio source category is obtained by identifying audio signals from multiple channels;

[0008] The target audio signal is processed using an audio processing strategy corresponding to the target audio source category.

[0009] In some embodiments, the target sound source category is one of the following categories: no sound, pure vocals, light music, mixed vocals, song, electronic echo.

[0010] In some embodiments, the target audio signal includes audio signals from multiple channels; determining the corresponding target sound source category based on the target audio signal includes: identifying the sound source category corresponding to the audio signal of each channel to obtain an identification result for each channel; and determining the target sound source category based on the identification result for each channel.

[0011] In some embodiments, the method further includes: merging audio signals from at least two channels to obtain an audio signal from a merged channel; identifying the sound source category corresponding to the audio signal from the merged channel to obtain an identification result for the merged channel; and determining the target sound source category based on the identification result of each channel includes: determining the target sound source category based on the identification result of each channel and the identification result of the merged channel.

[0012] In some embodiments, determining the target audio source category based on the recognition results of each channel and the recognition results of the merged channels includes: determining the confidence level of the recognition results of each channel and the confidence level of the recognition results of the merged channels; determining the highest first confidence level among the confidence levels of the recognition results of the multiple channels and the confidence level of the recognition results of the merged channels; and determining the recognition result corresponding to the first confidence level as the target audio source category.

[0013] In some embodiments, processing the target audio signal using an audio processing strategy corresponding to the target audio source category includes: determining a target channel audio signal from the audio signals of the plurality of channels based on the target audio source category; and processing the target channel audio signal using the corresponding audio processing strategy.

[0014] In some embodiments, processing the target audio signal using an audio processing strategy corresponding to the target audio source category includes: when the target audio source category is pure human voice, determining the confidence level of the audio signals of the plurality of channels belonging to pure human voice; determining the highest confidence level of pure human voice among the confidence levels of the audio signals of the plurality of channels belonging to pure human voice; determining the audio signal of the channel corresponding to the highest confidence level of pure human voice as the target channel audio signal; and performing noise reduction processing on the target channel audio signal using a noise reduction strategy corresponding to the pure human voice category.

[0015] In some embodiments, the noise reduction processing of the target channel audio signal includes: adaptive echo cancellation, adaptive noise reduction, and adaptive gain control processing of the target channel audio signal.

[0016] In some embodiments, processing the target audio signal using an audio processing strategy corresponding to the target audio source category includes: when the target audio source category is light music, determining the audio signal of at least one channel among the plurality of channels as the target channel audio signal; and performing noise reduction processing on the target channel audio signal according to a preset music mode noise reduction scheme.

[0017] In some embodiments, processing the target audio signal using an audio processing strategy corresponding to the target audio source category includes: determining the target audio source category as electronic echo when the identification result of at least one of the plurality of channels is electronic echo; after determining the target audio source category, the method further includes: displaying a first prompt control, the first prompt control being used to provide information prompts.

[0018] In some embodiments, determining the target sound source category based on the recognition result of each channel includes: when the recognition results of the plurality of channels are all silent, determining the target sound source category as silent.

[0019] In some embodiments, the step of processing the target audio signal using an audio processing strategy corresponding to the target audio source category includes: when the target audio source category is mixed with human voices, calling a human voice and background sound separation algorithm to separate the target audio signal to obtain human voice signal and background sound signal, and performing sound enhancement processing on the human voice signal separately.

[0020] In some embodiments, the target audio signal includes audio signals from multiple channels; determining the corresponding target sound source category based on the target audio signal includes: merging the audio signals from the multiple channels to obtain a merged channel audio signal; identifying the sound source category corresponding to the merged channel audio signal to obtain the identification result of the merged channel; and determining the identification result of the merged channel as the target sound source category.

[0021] In some embodiments, acquiring the target audio signal includes: acquiring an initial audio signal acquired by an audio acquisition device; and determining at least one audio segment in the initial audio signal as the target audio signal.

[0022] In some embodiments, determining the corresponding target sound source category based on the target audio signal includes: extracting the spectral features of the target sound source signal; inputting the spectral features into a pre-trained sound source category recognition model; and using the sound source category recognition model to process the spectral features of at least one audio segment in the target audio signal to obtain the target sound source category.

[0023] According to a second aspect of the embodiments of this application, a conferencing device is provided, the conferencing device including a processor and a memory for storing a computer program capable of running on the processor; wherein the processor is used to run the computer program to perform any of the above-described audio signal processing methods.

[0024] As can be seen, the embodiments of this application can process the target audio signal according to the audio processing strategy corresponding to the target audio signal after determining the target audio source category. The target audio source category is identified and confirmed based on the audio signals of multiple channels. This method of determining the target audio source category can improve the accuracy of audio source category identification. Furthermore, determining the corresponding audio processing strategy for the identified target audio source category can improve the processing effect of different categories of audio signals. For example, it can improve the sound enhancement effect or noise reduction effect of different categories of audio signals.

[0025] It should be understood that the above general description and the following detailed description are merely exemplary and explanatory, and are not intended to limit the technical solutions of this disclosure. Attached Figure Description

[0026] To more clearly illustrate the technical solutions in the embodiments of the present invention, the drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0027] Figure 1 This is a flowchart of an audio signal processing method according to an embodiment of this application;

[0028] Figure 2 This is a schematic diagram illustrating the resampling of audio segments in an embodiment of this application;

[0029] Figure 3 A flowchart for determining the corresponding target audio source category based on the target audio signal, provided in an embodiment of this application;

[0030] Figure 4 Another flowchart for determining the corresponding target audio source category based on the target audio signal, provided in an embodiment of this application;

[0031] Figure 5 This is a schematic diagram of the conference equipment in an embodiment of this application. Detailed Implementation

[0032] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of this application.

[0033] In the following description, references are made to “some embodiments,” which describe a subset of all possible embodiments. However, it is understood that “some embodiments” may be the same subset or different subsets of all possible embodiments and may be combined with each other without conflict.

[0034] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application pertains. The terminology used herein is for descriptive purposes only and is not intended to limit the scope of this application.

[0035] This application provides an audio signal processing method applied in a conference device. Exemplarily, the conference device can be a mobile terminal such as a mobile phone or tablet, or a fixed terminal such as a personal computer. Audio acquisition devices such as microphones can be installed in the conference device.

[0036] Figure 1 This is a flowchart of an audio signal processing method according to an embodiment of this application, such as... Figure 1 As shown, the process includes:

[0037] Step 101: Acquire the target audio signal.

[0038] In this embodiment of the application, the target audio signal can be acquired by an audio acquisition device installed in the conference equipment.

[0039] In some embodiments, the process of acquiring a target audio signal includes: acquiring an initial audio signal acquired by an audio acquisition device; and determining at least one audio segment in the initial audio signal as the target audio signal.

[0040] For example, after acquiring the initial audio signal from the audio acquisition device, an audio segment of a specified length in the initial audio signal is sampled at set time intervals to obtain at least one audio segment in the initial audio signal.

[0041] Here, the time interval and length can be preset according to actual needs. For example, the time interval can be set between 2.2s and 2.8s, and the length can be set between 0.3s and 0.7s.

[0042] For example, after acquiring the initial audio signal from the audio acquisition device, an audio segment of a specified length in the initial audio signal can be sampled at a specified sampling frequency, such as 16kHz or other sampling frequencies.

[0043] Reference Figure 2 After acquiring the initial audio signal, a 2.4-second frequency segment of the initial audio signal is resampled every 0.5 seconds at a sampling frequency of 16 kHz. Figure 2 In the image, audio segment 1, audio segment 2, and audio segment 3 are three audio segments obtained through sampling.

[0044] As can be seen, the embodiments of this application can perform subsequent processing on at least one audio segment of the initial audio signal without performing overall processing on the initial audio signal, which can reduce the amount of data of the audio signal that needs to be processed.

[0045] For example, if the target audio signal is a multi-channel audio signal, the multi-channel audio signal can be split, and the resulting audio signal can be identified as the target audio signal. If the target audio signal includes multiple single-channel audio signals, the multiple single-channel audio signals can be merged to obtain a merged channel audio signal, which can then be identified as the target audio signal. For instance, if the target audio signal includes audio signals from channel A and channel B, the audio signals from channel A and channel B can be merged to obtain a merged channel audio signal. In this way, the audio signals from channel A, channel B, and the merged channel audio signal can all be identified as the target audio signal.

[0046] Step 102: Determine the corresponding target audio source category based on the target audio signal; the target audio source category is obtained by identifying audio signals from multiple channels.

[0047] In some embodiments, the target sound source category is one of the following: no sound, pure vocals, light music, mixed vocals, songs, and electronic echoes. Thus, the technical solution of this application embodiment is beneficial for achieving type identification of audio signals in indoor scenes.

[0048] In some embodiments, the process of determining the corresponding target sound source category based on the target audio signal includes: extracting the spectral features of the target sound source signal; inputting the spectral features into a pre-trained sound source category recognition model; and using the sound source category recognition model to process the spectral features of at least one audio segment in the target audio signal to obtain the target sound source category.

[0049] For example, a sound source category recognition model can be trained using training data in the following six categories: no sound, pure vocals, light music, mixed vocals, songs, and electronic echoes. For instance, training data for the no sound and pure vocals categories can be extracted from audio data of interviews, meetings, and daily conversations; training data for light music and mixed vocals categories can be extracted from audio and video data of restaurant dining, festival events, and reading aloud; training data for light music and song categories can be extracted from audio data of songs and music genres; and training data for electronic echoes can also be synthesized. Furthermore, audio data from publicly available databases can be used, thereby enriching the training data and the application scenarios of the sound source category recognition model.

[0050] After acquiring the training data, it can be labeled to identify the corresponding real sound source category. For example, if the training data is a multi-channel audio dataset, it can be split into multiple channels, and then each split audio data point can be labeled individually. If the training data includes multiple single-channel audio datasets, these single-channel datasets can be merged to obtain merged audio data, which can then be labeled.

[0051] For example, the framework of the sound source category recognition model can be a matchbox network or other network frameworks. After acquiring training data, gain interference and noise interference can be added to the training data to enhance the robustness of the sound source category recognition model to be trained. Then, the spectral features of the training data can be extracted and input into the sound source category recognition model. The sound source category recognition model processes the spectral features of the training data to obtain the prediction results corresponding to the training data. Based on the difference between the prediction results corresponding to the training data and the true values ​​corresponding to the training data, the network parameter values ​​of the sound source category recognition model are adjusted, thereby achieving the training of the sound source category recognition model.

[0052] For example, the spectral features of the training data can be features such as Mel spectrograms and Mel cepstral coefficients; the prediction results corresponding to the training data can include the confidence levels of the training data belonging to categories 1 to 6, where categories 1 to 6 are respectively no sound, pure vocals, light music, mixed vocals, songs, and electronic echoes.

[0053] As can be seen, the embodiments of this application can use the sound source category recognition model to process the spectral features of at least one audio segment to obtain the target sound source category; there is no need to use the sound source category recognition model to process the initial audio signal as a whole. Therefore, the data processing volume of the sound source category recognition model can be reduced, thereby improving the recognition efficiency of the sound source category recognition model.

[0054] Step 103: Process the target audio signal using an audio processing strategy corresponding to the target audio source category.

[0055] As can be seen, the embodiments of this application can process the target audio signal according to the audio processing strategy corresponding to the target audio signal after determining the target audio source category. The target audio source category is identified and confirmed based on the audio signals of multiple channels. This method of determining the target audio source category can improve the accuracy of audio source category identification. Furthermore, determining the corresponding audio processing strategy for the identified target audio source category can improve the processing effect of different categories of audio signals. For example, it can improve the sound enhancement effect or noise reduction effect of different categories of audio signals.

[0056] In some embodiments of this application, reference is made to Figure 3 When the target audio signal includes audio signals from multiple channels, the process of determining the corresponding target audio source category based on the target audio signal may include:

[0057] Step A1: Identify the audio source category corresponding to the audio signal of each channel to obtain the identification result for each channel.

[0058] Here, the spectral features of the audio signal of each channel can be extracted, and then the spectral features of the audio signal of each channel can be input into the sound source category recognition model. The sound source category recognition model outputs the recognition result of each channel. For example, the recognition result of each channel can include the confidence level of the audio signal of the corresponding channel belonging to the first to sixth categories respectively.

[0059] Step A2: Determine the target sound source category based on the recognition results of each channel.

[0060] As can be seen, the embodiments of this application can determine the target sound source category by comprehensively considering the recognition results of each channel in multiple channels. Compared with the scheme that only considers the recognition results of a single channel, this can improve the accuracy of the target sound source category.

[0061] In some embodiments of this application, after obtaining audio signals from multiple channels, the audio signals from at least two of the multiple channels can be merged to obtain an audio signal from a merged channel; the audio source category corresponding to the audio signal from the merged channel is identified to obtain the identification result of the merged channel.

[0062] Accordingly, the process of determining the target audio source category based on the recognition results of each channel includes: determining the target audio source category based on the recognition results of each channel and the recognition results of the merged channels.

[0063] As can be seen, the embodiments of this application can determine the target sound source category by comprehensively considering the recognition results of each channel in multiple channels as well as the recognition results of the merged channels. Compared with the scheme that only considers the recognition results of each channel in multiple channels, it can further improve the accuracy of the target sound source category.

[0064] In some embodiments of this application, the process of determining the target audio source category based on the recognition results of each channel and the recognition results of the merged channels may include: determining the confidence level of the recognition results of each channel and the confidence level of the recognition results of the merged channels; determining the highest first confidence level among the confidence levels of the recognition results of multiple channels and the confidence levels of the recognition results of the merged channels; and determining the recognition result corresponding to the first confidence level as the target audio source category.

[0065] Understandably, among the confidence levels of the recognition results from multiple channels and the confidence levels of the recognition results from merged channels, the highest first confidence level can represent the greatest probability of the target audio signal belonging to a category. Therefore, by determining the recognition result corresponding to the first confidence level as the target audio source category, the accuracy of the target audio source category can be improved.

[0066] In some embodiments of this application, the process of processing a target audio signal using an audio processing strategy corresponding to the target audio source category includes: determining a target channel audio signal from multiple channels of audio signals based on the target audio source category; and processing the target channel audio signal using the corresponding audio processing strategy.

[0067] Understandably, different audio source categories typically have different audio signal processing requirements. Therefore, the embodiments of this application can reasonably determine the target channel audio signal that meets the corresponding processing requirements based on the target audio source category, thereby facilitating the processing of the target channel audio signal according to the corresponding processing requirements.

[0068] In some embodiments of this application, the process of processing the target audio signal using an audio processing strategy corresponding to the target audio source category includes: when the target audio source category is pure human voice, determining the confidence level of multiple channels' audio signals belonging to pure human voice; among the confidence levels of multiple channels' audio signals belonging to pure human voice, determining the highest confidence level of pure human voice; determining the audio signal of the channel corresponding to the highest confidence level of pure human voice as the target channel audio signal; and performing noise reduction processing on the target channel audio signal using a noise reduction strategy corresponding to the pure human voice category.

[0069] For example, the target channel audio signal can be processed with adaptive echo cancellation (AEC), adaptive noise suppression (ANS), and adaptive gain control (AGC).

[0070] Understandably, among multiple channels of audio signals, the higher the confidence level, the purer the human voice signal and the less interference signal. In this case, the audio signal of the channel with the highest confidence level of pure human voice is determined as the target channel audio signal, and noise reduction processing is performed on the target channel audio signal to improve the noise reduction effect.

[0071] In some embodiments of this application, the process of processing the target audio signal using an audio processing strategy corresponding to the target audio source category includes: when the target audio source category is light music, determining the audio signal of at least one channel among multiple channels as the target channel audio signal; and performing noise reduction processing on the target channel audio signal according to a preset music mode noise reduction scheme.

[0072] In this embodiment, a music mode noise reduction scheme can usually be preset according to actual needs. Therefore, noise reduction processing of the target channel audio signal according to the preset music mode noise reduction scheme can meet the noise reduction requirements of the music mode and is conducive to enhancing the music signal of the target channel audio signal.

[0073] In summary, the embodiments of this application can distinguish between pure human voice and music by identifying the type of sound source, and thus provide different noise reduction processing schemes according to different sound source types. That is, the technical solutions of the embodiments of this application can be adapted to a wide range of noise reduction scenarios.

[0074] In some embodiments of this application, the process of processing the target audio signal using an audio processing strategy corresponding to the target audio source category includes: determining the target audio source category as electronic echo when the identification result of at least one channel among multiple channels is electronic echo; and displaying a first prompt control, the first prompt control being used to provide information prompts.

[0075] Here, the first prompt control is used to prompt the user to turn off the sound playback device. The first prompt control can be a text prompt such as a prompt box, and the sound playback device can be a speaker or other device in a conference system.

[0076] Understandably, when the target sound source is determined to be electronic echo, it is generally assumed that some sound playback devices are playing sound. In this case, by displaying the first prompt control, the user can be reminded to take appropriate actions to avoid electronic echo, making the conference equipment more user-friendly.

[0077] In some embodiments of this application, the process of determining the target sound source category based on the recognition result of each channel includes: when the recognition results of multiple channels are all silent, the target sound source category is determined to be silent.

[0078] Here, when the recognition results of multiple channels are all silent, the audio signals of multiple channels can be identified as silent audio signals.

[0079] It can be seen that when the recognition results of multiple channels are all silent, the audio signals of multiple channels are very likely to be silent audio signals. Therefore, the target sound source category can be reasonably determined as silent.

[0080] In some embodiments of this application, the process of processing the target audio signal using an audio processing strategy corresponding to the target audio source category includes: when the target audio source category is human voice mixed, calling a human voice and background sound separation algorithm to separate the target audio signal to obtain human voice signal and background sound signal, and performing sound enhancement processing on the human voice signal separately.

[0081] As can be seen, the technical solution of this application embodiment can distinguish the background music during normal speech. Through the separation algorithm of human voice and background sound, the background sound can be filtered out, and then the sound of human voice signal can be enhanced.

[0082] In some embodiments of this application, reference is made to Figure 4 The process of determining the corresponding target sound source category based on the target audio signal may include:

[0083] Step B1: Merge the audio signals from multiple channels to obtain the merged audio signal.

[0084] Step B2: Identify the audio source category corresponding to the audio signal of the merged channel to obtain the identification result of the merged channel.

[0085] Step B3: Determine the target audio source category based on the recognition results of the merged channels.

[0086] In scenarios involving online audio signal processing or other situations requiring rapid audio signal processing, it is possible to identify the audio source category only for the merged audio signals, thereby improving the timeliness of audio signal processing.

[0087] Those skilled in the art will understand that, in the above-described method of the specific implementation, the order in which each step is written does not imply a strict execution order and does not constitute any limitation on the implementation process. The specific execution order of each step should be determined by its function and possible internal logic.

[0088] It should be noted that, in the embodiments of this application, if the above methods are implemented as software functional modules and sold or used as independent products, they can also be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the embodiments of this application, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a terminal, server, etc.) to execute all or part of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), magnetic disks, or optical disks. Thus, the embodiments of this application are not limited to any specific hardware and software combination.

[0089] Correspondingly, this application embodiment further provides a computer program product, the computer program product including computer executable instructions, which are used to implement any of the audio signal processing methods provided in this application embodiment.

[0090] Accordingly, this application embodiment further provides a computer storage medium storing computer-executable instructions, which are used to implement any of the audio signal processing methods provided in the above embodiments.

[0091] This application also proposes a conference device. Figure 5 This is a schematic diagram of the composition structure of the conference equipment provided in the embodiments of this application, such as... Figure 5 As shown, the conference equipment 50 may include:

[0092] Memory 501 is configured to store executable instructions;

[0093] When the processor 502 is configured to execute executable instructions stored in the memory 501, it implements any of the above-described audio signal processing methods.

[0094] The processor 502 mentioned above can be at least one of ASIC, DSP, DSPD, PLD, FPGA, CPU, controller, microcontroller, and microprocessor.

[0095] The aforementioned computer-readable storage medium and memory 502 may be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), a magnetic random access memory (FRAM), a flash memory, a magnetic surface memory, an optical disc, or a compact disc read-only memory (CD-ROM), etc.; or it may be various terminals including one or any combination of the above-mentioned memories, such as mobile phones, computers, tablet devices, personal digital assistants, etc.

[0096] In some embodiments, the functions or modules of the apparatus provided in this application can be used to perform the methods described in the above method embodiments. The specific implementation can be referred to the description of the above method embodiments, and for the sake of brevity, it will not be repeated here.

[0097] The description of the various embodiments above tends to emphasize the differences between the various embodiments. The similarities or similarities between them can be referred to, and for the sake of brevity, they will not be repeated here.

[0098] The methods disclosed in the various method embodiments provided in this application can be arbitrarily combined to obtain new method embodiments without conflict.

[0099] The features disclosed in the various product embodiments provided in this application can be arbitrarily combined without conflict to obtain new product embodiments.

[0100] The features disclosed in the various method or device embodiments provided in this application can be arbitrarily combined without conflict to obtain new method or device embodiments.

[0101] The embodiments of this application have been described above with reference to the accompanying drawings. However, this application is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of this application without departing from the spirit and scope of the claims. All of these forms are within the protection scope of this application.

Claims

1. A method of processing an audio signal, characterized by, Applied to a conference device, the method comprises: obtaining a target audio signal; determining a corresponding target sound source category based on the target audio signal, the target sound source category being obtained based on identification of audio signals of multiple channels; processing the target audio signal using an audio processing strategy corresponding to the target sound source category.

2. The method of claim 1, wherein, The target sound source category is one of the following categories: no sound, pure human voice, light music, mixed human voice, song, electronic echo.

3. The method of claim 1, wherein, The target audio signal comprises audio signals of multiple channels; determining the corresponding target sound source category based on the target audio signal comprises: identifying the sound source category corresponding to the audio signal of each channel to obtain an identification result of each channel; determining the target sound source category based on the identification result of each channel.

4. The method of claim 3, wherein, The method further comprises: merging the audio signals of at least two channels to obtain a merged channel audio signal; identifying the sound source category corresponding to the merged channel audio signal to obtain a merged channel identification result; determining the target sound source category based on the identification result of each channel and the merged channel identification result. The method further comprises:

5. The method of claim 4, wherein, determining the confidence of the identification result of each channel and the confidence of the merged channel identification result; determining the highest first confidence from the confidence of the identification result of the multiple channels and the confidence of the merged channel identification result; determining the identification result corresponding to the first confidence as the target sound source category. The method further comprises:

6. The method of claim 1, wherein, determining a target channel audio signal from the audio signals of the multiple channels based on the target sound source category; processing the target channel audio signal using the corresponding audio processing strategy. The method further comprises:

7. The method of claim 1, wherein, when the target sound source category is pure human voice, determining the confidence that the audio signals of the multiple channels belong to pure human voice; determining the highest pure human voice confidence from the confidence that the audio signals of the multiple channels belong to pure human voice; determining the audio signal of the channel corresponding to the highest pure human voice confidence as the target channel audio signal; performing noise reduction processing on the target channel audio signal using a noise reduction strategy corresponding to the pure human voice category. The method further comprises:

8. The method of claim 7, wherein, performing adaptive echo cancellation, adaptive noise reduction, and adaptive gain control processing on the target channel audio signal. The method further comprises:

9. The method of claim 1, wherein, when the target sound source category is light music, determining the audio signal of at least one channel of the multiple channels as the target channel audio signal; performing noise reduction processing on the target channel audio signal according to a preset music mode noise reduction scheme. ​ 10. The method of claim 1, wherein, The processing of the target audio signal by using the audio processing strategy corresponding to the target audio source category comprises: When the identification result of at least one channel in the multiple channels is electronic echo, the target audio source category is determined as electronic echo; A first prompt control is displayed, and the first prompt control is used for information prompting.

11. The method of claim 1, wherein, The processing of the target audio signal by using the audio processing strategy corresponding to the target audio source category comprises: When the target audio source category is mixed voice, a voice and background sound separation algorithm is called to separate the target audio signal to obtain a voice signal and a background sound signal, and the voice signal is separately subjected to sound enhancement processing.

12. The method of claim 1, wherein, The target audio signal comprises multiple-channel audio signals. The target audio source category corresponding to the target audio signal is determined, which comprises: The multiple-channel audio signals are merged to obtain a merged-channel audio signal; An audio source category corresponding to the merged-channel audio signal is identified to obtain a merged-channel identification result; The merged-channel identification result is determined as the target audio source category.

13. The method according to any one of claims 1 to 12, characterized in that, The target audio signal is acquired, which comprises: An initial audio signal collected by an audio collection device is acquired; A signal of at least one audio segment in the initial audio signal is determined as the target audio signal.

14. The method of claim 1, wherein, The target audio source category corresponding to the target audio signal is determined, which comprises: Spectrum features of the target audio signal are extracted; The spectrum features are input into a pre-trained audio source category identification model, and the audio source category identification model is used to process the spectrum features of at least one audio segment in the target audio signal to obtain the target audio source category.

15. A conferencing device, characterized by The conference device comprises a processor and a memory for storing a computer program capable of running on the processor; wherein, The processor is used to run the computer program to execute the method in any one of claims 1 to 14.