Audio signal processing method, device, system and storage medium

The microphone array generates the spatial distribution information of the sound source and combines the historical transformation relationship to identify overlapping speech, which solves the problem of the reduction in the accuracy of overlapping speech detection in the microphone array scene, and achieves high-accurate overlapping speech recognition.

CN115019826BActive Publication Date: 2025-08-08ALIBABA (CHINA) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202110235834.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-03-03
Publication Date
2025-08-08
Estimated Expiration
2041-03-03

AI Technical Summary

Technical Problem

The existing overlapping voice detection technology has reduced accuracy in microphone array scenarios and cannot meet product-level detection needs.

Method used

The phase difference information of the audio signal is collected through the microphone array, the sound source spatial distribution information is generated, and the conversion relationship between a single voice and overlapping voice learned by the historical audio signal is identified whether the current audio signal is overlapping voice.

Benefits of technology

Improves the accuracy of identifying overlapping voices in microphone array scenarios to meet product-level detection needs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115019826B_ABST
    Figure CN115019826B_ABST
Patent Text Reader

Abstract

The embodiments of the present application provide an audio signal processing method, device, system, and storage medium. In the embodiments of the present application, a microphone array is used to collect audio signals, and the sound source spatial distribution information corresponding to the audio signal is generated based on the phase difference information of the audio signal collected by each microphone in the microphone array. Then, based on the sound source spatial distribution information and the conversion relationship between single speech and overlapping speech learned based on historical audio signals, it is identified whether the current audio signal is overlapping speech. Compared with single-channel audio, the audio signal collected by the microphone array contains the sound source spatial distribution information, so that it can accurately identify whether the current audio signal is overlapping speech, meeting product-level detection requirements.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of audio processing technology, and in particular to an audio signal processing method, device, system and storage medium. Background Art

[0002] A microphone array is a system composed of a specific number of microphones that samples and filters the spatial characteristics of a sound field. Microphone arrays have a strong ability to suppress far-field interference noise and are used in products with voice acquisition capabilities, such as microphones and voice recorders, to accurately capture voice signals in a variety of scenarios.

[0003] In some application scenarios, there may be both a single speaker and multiple speakers speaking simultaneously. The collected speech signals may include both single speech signals and overlapping speech signals from multiple speakers. To accurately identify the number of speakers in a meeting and the content of their respective speeches, it is necessary to identify overlapping speech signals and then perform speech recognition processing on them.

[0004] Existing technologies can detect overlapping speech signals based on a model trained with large amounts of audio data. However, existing overlapping speech detection is mostly based on single-channel audio. Directly applying existing overlapping speech detection technology to multi-channel audio scenarios using microphone arrays will reduce its accuracy and fail to meet product-level detection requirements. Summary of the Invention

[0005] Various aspects of the present application provide an audio signal processing method, device, system, and storage medium to improve the accuracy of identifying whether speech is overlapping speech, so as to meet product-level detection requirements.

[0006] An embodiment of the present application provides an audio signal processing method, comprising: obtaining a current audio signal collected by a microphone array, the microphone array including at least two microphones; generating current sound source spatial distribution information corresponding to the current audio signal based on phase difference information of the current audio signal collected by the at least two microphones; and identifying whether the current audio signal is overlapping speech based on the current sound source spatial distribution information and a conversion relationship between a single speech and overlapping speech learned based on historical audio signals.

[0007] An embodiment of the present application also provides an audio signal processing method, which is applicable to conference equipment, the conference equipment including a microphone array, the method including: obtaining a current conference signal collected by the microphone array in a conference scene, the microphone array including at least two microphones; generating current sound source spatial distribution information corresponding to the current conference signal based on phase difference information collected by at least two microphones; based on the current sound source spatial distribution information, combined with the conversion relationship between single speech and overlapping speech learned based on historical conference signals, identifying whether the current conference signal is overlapping speech.

[0008] An embodiment of the present application also provides an audio signal processing method, which is suitable for teaching equipment, the teaching equipment including a microphone array, the method including: obtaining a current classroom signal collected by the microphone array in a teaching environment, the microphone array including at least two microphones; generating current sound source spatial distribution information corresponding to the current classroom signal based on phase difference information of the current classroom signal collected by at least two microphones; based on the current sound source spatial distribution information, combined with the conversion relationship between single speech and overlapping speech learned based on historical classroom signals, identifying whether the current classroom signal is overlapping speech.

[0009] An embodiment of the present application also provides an audio signal processing method, which is applicable to an intelligent vehicle-mounted device, wherein the intelligent vehicle-mounted device includes a microphone acquisition array. The method includes: obtaining a current audio signal collected by the microphone array in a vehicle-mounted environment, wherein the microphone array includes at least two microphones; generating current sound source spatial distribution information corresponding to the current audio signal based on phase difference information of the current audio signal collected by at least two microphones; and identifying whether the current audio signal is an overlapping speech based on the current sound source spatial distribution information and the conversion relationship between a single speech and an overlapping speech learned based on historical audio signals.

[0010] The present application also provides a terminal device, comprising: a memory, a processor, and a microphone array; the memory is configured to store a computer program; the processor is coupled to the memory and configured to execute the computer program, and is configured to: obtain a current audio signal collected by the microphone array, the microphone array comprising at least two microphones; generate current sound source spatial distribution information corresponding to the current audio signal based on phase difference information of the current audio signal collected by the at least two microphones; and identify whether the current audio signal is overlapping speech based on the current sound source spatial distribution information and a conversion relationship between a single speech and overlapping speech learned based on historical audio signals. The present application also provides a conference device, comprising: a memory, a processor, and a microphone array; the memory is configured to store a computer program; the processor is coupled to the memory and configured to execute the computer program, and is configured to: obtain a current conference signal collected by the microphone array in a conference scene, the microphone array comprising at least two microphones; generate current sound source spatial distribution information corresponding to the current conference signal based on phase difference information of the current audio signal collected by the at least two microphones; and identify whether the current conference signal is overlapping speech based on the current sound source spatial distribution information and a conversion relationship between a single speech and overlapping speech learned based on historical conference signals.

[0011] An embodiment of the present application also provides a teaching device, including: a memory, a processor and a microphone array; the memory is used to store a computer program; the processor is coupled to the memory and is used to execute the computer program, so as to: obtain a current classroom signal collected by the microphone array in a teaching environment, the microphone array includes at least two microphones; based on the phase difference information of the current classroom signal collected by at least two microphones, generate current sound source spatial distribution information corresponding to the current classroom signal; based on the current sound source spatial distribution information, combined with the conversion relationship between single speech and overlapping speech learned based on historical classroom signals, identify whether the current classroom signal is overlapping speech.

[0012] An embodiment of the present application also provides an intelligent vehicle-mounted device, including: a memory, a processor and a microphone array; the memory is used to store a computer program; the processor is coupled to the memory and is used to execute the computer program, so as to: obtain a current audio signal collected by the microphone array in a vehicle-mounted environment, the microphone array including at least two microphones; based on the phase difference information of the current audio signal collected by at least two microphones, generate current sound source spatial distribution information corresponding to the current audio signal; based on the current sound source spatial distribution information, combined with the conversion relationship between a single voice and overlapping voice learned based on historical audio signals, identify whether the current audio signal is overlapping voice.

[0013] An embodiment of the present application also provides an audio signal processing system, including: a terminal device and a server device; the terminal device includes a microphone array, the microphone array includes at least two microphones, and is used to collect a current audio signal; the terminal device is used to upload the current audio signal collected by the at least two microphones to the server device; the server device is used to generate current sound source spatial distribution information corresponding to the current audio signal based on phase difference information of the current audio signal collected by the at least two microphones; based on the current sound source spatial distribution information, combined with the conversion relationship between a single voice and overlapping voice learned based on historical audio signals, identify whether the current audio signal is overlapping voice.

[0014] An embodiment of the present application also provides a server device, comprising: a memory and a processor; the memory is used to store a computer program; the processor is coupled to the memory and is used to execute the computer program, so as to: receive a current audio signal collected by at least two microphones in a microphone array uploaded by a terminal device; generate current sound source spatial distribution information corresponding to the current audio signal based on phase difference information of the current audio signal collected by at least two microphones; and identify whether the current audio signal is an overlapping speech based on the current sound source spatial distribution information and the conversion relationship between a single speech and an overlapping speech learned based on historical audio signals.

[0015] The embodiment of the present application further provides a computer-readable storage medium storing a computer program. When the computer program is executed by a processor, the processor implements the steps of the audio signal processing method provided in the embodiment of the present application.

[0016] An embodiment of the present application further provides a computer program product, including a computer program / instruction. When the computer program / instruction is executed by a processor, the processor is caused to implement the steps of the audio signal processing method provided in the embodiment of the present application.

[0017] In an embodiment of the present application, a microphone array is used to collect audio signals, and the sound source spatial distribution information corresponding to the audio signal is generated based on the phase difference information of the audio signal collected by each microphone in the microphone array. Then, based on the sound source spatial distribution information and combined with the conversion relationship between single speech and overlapping speech learned based on historical audio signals, it is identified whether the current audio signal is overlapping speech. Compared with single-channel audio, the audio signal collected by the microphone array contains the sound source spatial distribution information, so that it can accurately identify whether the current audio signal is overlapping speech, thereby meeting product-level detection requirements. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings:

[0019] Figure 1a A flowchart of an audio signal processing method provided by an exemplary embodiment of the present application;

[0020] Figure 1b A flowchart of another audio signal processing method provided by an exemplary embodiment of the present application;

[0021] Figure 2a A schematic diagram of the layout of microphones in a microphone array provided as an exemplary embodiment of the present application;

[0022] Figure 2b A schematic diagram of peak information of sound source spatial distribution information provided by an exemplary embodiment of the present application;

[0023] Figure 3a This is a schematic diagram of the usage status of conference equipment in a conference scenario;

[0024] Figure 3b This is a schematic diagram of the use of the sound pickup device in a business cooperation negotiation scenario;

[0025] Figure 3c A schematic diagram of the usage status of the teaching equipment in the teaching scenario;

[0026] Figure 3d This is a schematic diagram of the usage status of the intelligent vehicle-mounted device in the vehicle environment;

[0027] Figure 3e A flowchart of another audio signal processing method provided by an exemplary embodiment of the present application;

[0028] Figure 3f A flowchart of another audio signal processing method provided by an exemplary embodiment of the present application;

[0029] Figure 3g A flowchart of another audio signal processing method provided by an exemplary embodiment of the present application;

[0030] Figure 4 A schematic structural diagram of an audio signal processing system provided by an exemplary embodiment of the present application;

[0031] Figure 5 A schematic diagram of the structure of a terminal device provided by an exemplary embodiment of the present application;

[0032] Figure 6 A schematic diagram of the structure of a server device provided in an exemplary embodiment of the present application. DETAILED DESCRIPTION

[0033] To make the purpose, technical solutions, and advantages of this application more clear, the technical solutions of this application will be clearly and completely described below in conjunction with the specific embodiments of this application and the corresponding drawings. Obviously, the embodiments described are only part of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0034] The following describes in detail the technical solutions provided by various embodiments of the present application in conjunction with the accompanying drawings.

[0035] Figure 1a A flowchart of an audio signal processing method provided by an exemplary embodiment of the present application; Figure 1a As shown, the method includes:

[0036] 101a. Obtain a current audio signal collected by a microphone array, where the microphone array includes at least two microphones;

[0037] 102a. Generate current sound source spatial distribution information corresponding to the current audio signal based on phase difference information of the current audio signal collected by at least two microphones;

[0038] 103a. Identify whether the current audio signal is overlapping speech based on the current sound source spatial distribution information and the conversion relationship between single speech and overlapping speech learned based on historical audio signals.

[0039] In this embodiment, a sound source refers to an object that can produce sound through vibration. For example, the sound source can be a musical instrument, a vibrating tuning fork, a human vocal organ (such as vocal cords), or an animal vocal organ. The sound source can produce speech, which refers to sounds produced by human vocal organs and has certain social significance. The microphone array can collect audio signals emitted by the sound source. The audio signals may contain speech or other sounds besides speech, such as reverberation, echo, ambient noise, animal calls, or the sound of objects colliding.

[0040] In this embodiment, the microphone array includes at least two microphones, wherein the layout of the at least two microphones is not limited. Figure 2a As shown, this can be a linear array, a planar array, or a stereo array. Due to the specific layout of the microphones in a microphone array, the time it takes for the same audio signal to reach each microphone varies, resulting in a time delay. This time delay can be reflected as a phase difference when the same audio signal reaches each microphone, referred to as a phase difference.

[0041] In this embodiment, the microphone array may capture audio signals. Depending on the application scenario, "interruption" (i.e., one person interrupting another) may occur at any time. The current audio signal captured by the microphone array may be a single voice produced by a single speaker or an overlapping voice signal from multiple speakers. In this embodiment, the audio signal is assumed to exist in two states: a single voice signal and overlapping voice signals.

[0042] In this embodiment, the phase difference information present when each microphone in the microphone array collects the current audio signal is used to determine the state of the audio signal, namely, whether the audio signal is a single speech or overlapping speech, or whether the audio signal is overlapping speech. This phase difference information can, to a certain extent, reflect the spatial distribution of the sound source locations. Based on this spatial distribution of the sound source locations, the number and locations of valid sound sources can be identified. If the number of valid sound sources is identified, whether the audio signal is overlapping speech can be determined.

[0043] Specifically, the current audio signal collected by the microphone array can be obtained. The segmentation length of the audio signal is not limited and can be in units of signal frames. The current audio signal can be a signal frame, and each signal frame is usually in the millisecond level (for example, 20ms), which is usually less than the pronunciation duration of a single word or syllable during the speech process. Alternatively, several consecutive signal frames can be used as the current audio signal, and there is no limitation on this. Then, based on the phase difference information of the current audio signal collected by at least two microphones in the microphone array, the current sound source spatial distribution information corresponding to the current audio signal is generated. The current sound source spatial distribution information reflects the spatial distribution of the current sound source. According to the spatial distribution of the current sound source, the number of valid sound sources and the position of the valid sound source can be identified. When the number of valid sound sources is identified, it can be determined whether the audio signal is overlapping speech.

[0044] In practical applications, given the continuity of audio signals, there are certain rules for transitioning from one state to another. For example, the state of the current audio signal may be related to the state corresponding to the previous audio signal, or it may be related to the states corresponding to the previous two or previous N (N>2) audio signals. Based on this, under the initialization probability of single speech and overlapping speech, the transition relationship between single speech and overlapping speech is continuously learned based on the state of historical audio signals. This transition relationship refers to the transition probability between the states corresponding to the audio signals, which may include: the transition probability between single speech and single speech, the transition rate between single speech and overlapping speech, the transition probability between overlapping speech and single speech, and the transition probability between overlapping speech and overlapping speech. Based on the above, when determining whether the current audio signal is an overlapping signal, the learned transition relationship between single speech and overlapping speech can be relied upon, and combined with the current sound source spatial distribution information, to identify whether the current audio signal is an overlapping signal. Compared with single-channel audio, the audio signal collected by the microphone array contains the sound source spatial distribution information, so that it can accurately identify whether the audio signal at any moment is an overlapping speech, meeting product-level detection requirements.

[0045] In this embodiment, the phase difference information can reflect the spatial distribution of the sound source position to a certain extent. In order to better reflect the spatial distribution of the sound source position, in some optional embodiments of the present application, the phase difference information of the current audio signal collected by at least two microphones can be used to calculate the wave spectrum corresponding to the current audio signal. The wave spectrum can reflect the spatial distribution of the current sound source.

[0046] Optionally, for any position in the position space, the phase difference information of the current audio signal collected by any two microphones is accumulated to obtain the probability of each position being the current sound source position. Based on the probability of each position in the position space being the current sound source position, a wave arrival spectrogram corresponding to the current audio signal is generated. Specifically, the probability of each position being the current sound source position can be obtained using a sound source localization algorithm based on a phase-weighted controllable response power (SRP-PHAT). The basic principle of the SRP-PHAT algorithm is: assuming that any position in the position space is the position of the sound source, the microphone array collects the audio signal emitted by the sound source at that position, and uses the Generalized Cross Correlation-PHAseTransformation (GCC-PHAT) algorithm to calculate the cross-correlation function between the audio signals collected by any two microphones, and weight the cross-power spectral density function of the cross-correlation function. Then, the calculated GCC-PHAT values between all any two microphones are accumulated to obtain the SRP-PHAT value corresponding to any position. Furthermore, based on the SRP-PHAT value corresponding to any position, the probability of each position being the current sound source position can be obtained. Based on the probability of each position in the position space being the current sound source position, the wave arrival spectrogram corresponding to the current audio signal is generated. For example, the SRP-PHAT value corresponding to each orientation can be directly used as the probability that each orientation is the current sound source location. The arrival spectrogram can record each orientation and its corresponding SRP-PHAT value. A larger SRP-PHAT value indicates a greater probability that the orientation corresponding to the SRP-PHAT value is the sound source location. For another example, the ratio of the SRP-PHAT value at each orientation to the sum of the SRP-PHAT values for all orientations can be used as the probability that each orientation is the current sound source location. The arrival spectrogram can directly reflect the probability of each orientation being the current sound source location.

[0047] In an optional embodiment, a Hidden Markov Model (HMM) can be used to identify whether the current audio signal is overlapping speech. Specifically, the state of the audio signal, i.e., single speech and overlapping speech, can be used as two hidden states of the HMM, and the peak information of the current sound source spatial distribution information of the audio signal can be calculated as the observed state of the HMM. Optionally, the peak information of the current sound source spatial distribution information can be calculated using a Kurtosis algorithm or an Excessive Mass algorithm. The peak information can be the number of peaks, such as Figure 2bAs shown in FIG, there are three forms of peak information, that is, three forms of observation states, namely unimodal, bimodal, and multimodal.

[0048] In this embodiment, after calculating the observed state for the current audio signal, the current observed state can be input into an HMM. Combined with the transition relationship between two hidden states learned by the HMM, the probability of the current observed state corresponding to the hidden state is calculated, conditioned on the historical observed state. Specifically, initialization probabilities for the hidden states can be set, for example, 0.6 for single speech and 0.4 for overlapping speech. With these initialization probabilities set, the transition relationship between hidden states and the transmission relationship from hidden states to observed states are continuously learned based on the historical audio signal state to obtain an HMM model. After the observed state is input into the HMM model, the HMM model outputs the probability that the current observed state is the hidden state, conditioned on the historical observed state. For example, if the historical observed state contains five consecutive single peaks, the HMM model identifies the hidden state corresponding to the five consecutive single peaks as a single speech. If the current observed state is a bimodal observation state, assuming the presence of five consecutive single peaks, the HMM model outputs the probabilities of the current observed state being overlapping speech and single speech, respectively, with the probability of the current observed state being overlapping speech being greater than the probability of the current observed state being a single speech.

[0049] In this embodiment, after the HMM model outputs the probability of the current observed state corresponding to the hidden state, whether the current audio signal is overlapping speech can be identified based on the probability of the current observed state corresponding to the hidden state. If the probability of the current observed state corresponding to overlapping speech is greater than the probability of the current observed state corresponding to a single speech, the current audio signal is considered to be overlapping speech; if the probability of the current observed state corresponding to overlapping speech is less than or equal to the probability of the current observed state corresponding to a single speech, the current audio signal is considered to be a single speech.

[0050] In an optional embodiment, if the current audio signal is identified as overlapping speech, at least two valid sound source locations are determined based on the current sound source spatial distribution information. For example, if the current sound source spatial distribution information includes the probabilities of each location being the current sound source location, the two locations with the highest probabilities of being the current sound source location are selected as valid sound source locations. For another example, if the current sound source spatial distribution information is represented by an arrival spectrogram, and the arrival spectrogram includes SRP-PHAT values for each location, the two locations with the highest SRP-PHAT values from the arrival spectrogram can be selected as valid sound source locations. Then, the audio signals in at least two effective sound source directions can be speech enhanced. Specifically, the beamforming (BF) technology can be used to form a beam on the audio signal in the effective sound source direction. The beam can be used to effectively perform speech enhancement on the audio signal, while suppressing the audio signals in other directions outside the effective sound source direction, thereby achieving the effect of speech separation. On this basis, speech recognition is performed separately on the enhanced audio signals in at least two effective sound source directions, which can improve the accuracy of speech recognition and enhance the user experience.

[0051] In another optional embodiment, if the current audio signal is identified as a single speech, the direction with the highest probability of being the current sound source location is determined as the effective sound source direction; speech enhancement is performed on the audio signal at the effective sound source direction, and speech recognition is performed on the enhanced audio signal at the effective sound source direction. The implementation of speech enhancement for a single speech is the same or similar to the implementation of speech enhancement for overlapping speech in the above embodiment and is not further described here.

[0052] In some application scenarios of the present application, such as conference scenarios, teacher teaching scenarios, or business cooperation negotiation scenarios, it is often necessary to recognize voice signals, and non-voice signals, such as environmental noise, animal calls, or object collision sounds, are not of much concern. Based on this, before identifying whether the current audio signal is overlapping voice, it is also possible to determine whether the current audio signal is a voice signal. If the current audio signal is not a voice signal, the current audio signal is not recognized to improve the efficiency of audio processing. If the current audio signal is a voice signal, it is identified whether the current audio signal is overlapping voice.

[0053] Based on the above, the embodiment of the present application further provides an audio signal processing method, such as Figure 1b As shown, the method includes:

[0054] 101b. Obtain a current audio signal collected by a microphone array, where the microphone array includes at least two microphones;

[0055] 102b. Generate current sound source spatial distribution information corresponding to the current audio signal based on phase difference information of the current audio signal collected by at least two microphones;

[0056] 103b. Calculate the direction of arrival of the current audio signal based on the current sound source spatial distribution information;

[0057] 104b. Selecting a microphone from at least two microphones as a target microphone based on the direction of arrival;

[0058] 105b. Perform voice endpoint detection (VAD) on the current audio signal collected by the target microphone to determine whether the current audio signal is a voice signal.

[0059] 106b. If the current audio signal is a speech signal, execute step 107b; otherwise, terminate the processing of the current audio signal.

[0060] 107b. Identify whether the current audio signal is overlapping speech based on the current sound source spatial distribution information and the conversion relationship between single speech and overlapping speech learned based on historical audio signals.

[0061] In this embodiment, for the contents of steps 101b, 102b and 107b, reference may be made to the detailed contents of steps 101a, 102a and 103a in the aforementioned embodiment, which will not be repeated here.

[0062] In this embodiment, the direction of arrival (DOA) of the current audio signal is calculated based on the current sound source spatial distribution information. The direction of arrival refers to the directional angle at which the current audio signal reaches the microphone array. The direction of arrival may be the same as or different from the directional angle at which each microphone in the microphone array receives the audio signal, depending on the specific layout of the microphones. Where the sound source spatial distribution information includes the probabilities of each orientation as the current sound source location, the orientation with the highest probability of being the current sound source location may be used directly as the direction of arrival, or the orientation at a set angle to the orientation with the highest probability of being the current sound source location may be used as the direction of arrival. There is no limitation on this.

[0063] After calculating the direction of arrival, a microphone can be selected from at least two microphones as the target microphone based on the direction of arrival. For example, the direction angle at which each microphone receives the current audio signal can be calculated, and the direction angle consistent with the direction of arrival can be selected from multiple direction angles, and the microphone corresponding to the direction angle can be used as the target microphone. After determining the target microphone, voice activity detection (VAD) can be performed on the current audio signal collected by the target microphone to determine whether the current audio signal is a voice signal. The basic principle of VAD is to accurately locate the start and end endpoints of the voice signal from the noisy audio signal, thereby determining whether the current audio signal is a voice signal. That is, if the start and end endpoints of the voice signal can be detected from the audio signal, the audio signal is considered to be a voice signal. If the start and end endpoints of the voice signal cannot be detected from the audio signal, the audio signal is considered not to be a voice signal.

[0064] It should be noted that, in this embodiment, the implementation method of VAD for the current audio signal collected by the target microphone is not limited. In an optional embodiment, the software VAD function can be used to perform VAD on the current audio signal collected by the target microphone. The software VAD function refers to the implementation of the VAD function by software, and the software that implements the VAD function is not limited. For example, it can be a neural network VAD (Nature Network-VAD, NN-VAD) model trained by a human voice model. In another optional embodiment, the hardware VAD function can be used to perform VAD on the current audio signal collected by the target microphone. The hardware VAD function refers to the implementation of the VAD function by a built-in VAD module in a voice chip or device. The VAD module can be solidified on the voice chip, and the VAD function can be modified by configuring parameters.

[0065] The audio signal processing method provided by this embodiment can be applied to various multi-person speaking scenarios, such as multi-person conference scenarios, court trial scenarios, or teaching scenarios. In these application scenarios, the terminal devices of this embodiment will be deployed in these scenarios to collect audio signals in the application scenarios and implement the other functions described in the above-mentioned method embodiments of this application and the following system embodiments. Among them, the terminal device can be implemented as a sound pickup device such as a voice recorder, a recording stick, a recorder, or a microphone, or it can be implemented as a terminal device with a recording function, such as conference equipment, teaching equipment, a robot, a smart set-top box, a smart TV, a smart speaker, and a smart car-mounted device. In order to have a better acquisition effect and facilitate the identification of whether the audio signal is overlapping speech, further, according to whether the audio signal is overlapping speech, the audio signal is subjected to speech enhancement and speech recognition, and the placement of the terminal device can be reasonably determined according to the specific deployment of the multi-person speaking scenario. Figure 3a As shown, in a multi-person conference scenario, the terminal device is a conference device as an example. The conference device includes a microphone array with a sound pickup function. Considering that multiple speakers are distributed in different directions of the conference device, it is preferred to deploy the conference device in the center of the conference table; Figure 3b As shown, in a business cooperation negotiation scenario, the terminal device is a sound pickup device as an example. The first business party and the second business party are seated opposite each other. The conference organizer is located between the first business party and the second business party and is responsible for organizing the negotiation between the two parties. The sound pickup device is deployed in a central position between the conference organizer, the first business party, and the second business party. The sound pickup devices of the first business party, the second business party, and the conference organizer are located in different directions to facilitate the sound pickup device to pick up the sound. Figure 3c As shown in the figure, in the teaching scenario, the terminal device is a teaching device as an example. The teaching device is deployed on the lecture table, and the teacher and the student are located at different positions of the teaching device, which is convenient for picking up the voice of the teacher and the student at the same time. Figure 3d As shown, in the in-vehicle scenario, the terminal device is implemented as an intelligent in-vehicle device on the vehicle equipment. The intelligent in-vehicle device is located in the center of the car, and the passengers in seats A, B, C, and D are located in different directions of the intelligent in-vehicle device, which is convenient for picking up the voices of different passengers.

[0066] The following describes in detail the audio and video signal processing process in different application scenarios.

[0067] against Figure 3a In the conference scenario shown in FIG. 1 , the present application embodiment provides an audio signal processing method suitable for conference equipment, such as Figure 3e As shown, the method includes:

[0068] 301e. Obtain a current conference signal collected by a microphone array in a conference scene, where the microphone array includes at least two microphones;

[0069] 302e. Generate current sound source spatial distribution information corresponding to the current conference signal based on phase difference information of the current conference signal collected by at least two microphones;

[0070] 303e. Based on the current sound source spatial distribution information and the conversion relationship between single speech and overlapping speech learned based on historical conference signals, identify whether the current conference signal is overlapping speech.

[0071] For details about steps 301e to 303e, please refer to the previous Figure 1a and Figure 1b The embodiments shown are not described in detail here.

[0072] against Figure 3cIn the teaching scenario shown in FIG, the embodiment of the present application provides an audio signal processing method suitable for teaching equipment, such as Figure 3f As shown, the method includes:

[0073] 301f. Obtain a current classroom signal collected by a microphone array in a teaching environment, where the microphone array includes at least two microphones;

[0074] 302f. Generate current sound source spatial distribution information corresponding to the current classroom signal based on phase difference information of the current classroom signal collected by at least two microphones;

[0075] 303f. According to the current sound source spatial distribution information and the conversion relationship between single speech and overlapping speech learned based on historical classroom signals, identify whether the current classroom signal is overlapping speech.

[0076] For details about steps 301f to 303f, please refer to the previous Figure 1a and Figure 1b The embodiments shown are not described in detail here.

[0077] against Figure 3d In the vehicle-mounted scenario shown in FIG, the embodiment of the present application provides an audio signal processing method suitable for intelligent vehicle-mounted devices, such as Figure 3g As shown, the method includes:

[0078] 301g. Obtain a current audio signal collected by a microphone array in a vehicle environment, where the microphone array includes at least two microphones;

[0079] 302g. Generate current sound source spatial distribution information corresponding to the current audio signal based on phase difference information of the current audio signal collected by at least two microphones;

[0080] 303g. Identify whether the current audio signal is overlapping speech based on the current sound source spatial distribution information and the conversion relationship between single speech and overlapping speech learned based on historical audio signals.

[0081] For details about steps 301g to 303g, please refer to the previous Figure 1a and Figure 1b The embodiments shown are not described in detail here.

[0082] It should be noted that the method provided in the embodiment of the present application can be completed entirely by the terminal device, or a portion of the functions can be implemented on the server device, and there is no limitation on this. Based on this, the present embodiment provides an audio signal processing system, which describes the process of implementing the audio signal processing method based on the terminal device and the server device. Figure 4As shown, the audio signal processing system 400 includes: a terminal device 401 and a server device 402. The audio signal processing system 400 can be applied to a multi-person speaking scenario, for example Figure 3a The multi-person conference scenario shown in the figure is Figure 3b The business cooperation meeting scene shown, Figure 3c The teaching scenario shown and Figure 3d In these scenarios, the terminal device 401 can cooperate with the server device 402 to implement the above-mentioned embodiments of the method of the present application. Figures 3a to 3d The server device 402 is not shown in the multi-person speaking scenario.

[0083] The terminal device 401 of this embodiment has functional modules such as a power button, an adjustment button, a microphone array, and a speaker, wherein the microphone array includes at least two microphones and, optionally, a display screen. The terminal device 401 can realize functions such as automatic recording, MP3 playback, FM radio, digital camera function, phone recording, timed recording, external transcription, repeater or editing. Figure 4 As shown, the terminal device 401 can use at least two microphones in the microphone array to collect the current audio signal, and upload the current audio signal collected by the at least two microphones to the server device 402; the server device 402 receives the current audio signal collected by the at least two microphones, and generates the current sound source spatial distribution information corresponding to the current audio signal based on the phase difference information of the current audio signal collected by the at least two microphones; based on the current sound source spatial distribution information, combined with the conversion relationship between single speech and overlapping speech learned based on historical audio signals, it is identified whether the current audio signal is overlapping speech.

[0084] In some optional embodiments of the present application, if the server device 402 identifies the current audio signal as overlapping speech, it determines at least two valid sound source locations based on the current sound source spatial distribution information; performs speech enhancement on the audio signals at the at least two valid sound source locations; and performs speech recognition on each of the enhanced audio signals at the at least two valid sound source locations. Furthermore, optionally, if the current sound source spatial distribution information includes the probabilities of each location being the current sound source location, the two locations with the highest probabilities of being the current sound source location are selected as the valid sound source locations.

[0085] In some optional embodiments of the present application, if the server device 402 recognizes that the current audio signal is a single voice, it will use the direction with the highest probability of being the current sound source position as the effective sound source direction; perform voice enhancement on the audio signal at the effective sound source direction, and perform voice recognition on the enhanced audio signal at the effective sound source direction.

[0086] It should be noted that the implementation form of the terminal device varies depending on the application of the audio signal processing system in different scenarios. For example, in a conference setting, the terminal device is implemented as a conference device; in a business cooperation meeting, the terminal device is implemented as a sound pickup device; in a teaching setting, the terminal device is implemented as a teaching device; and in an in-vehicle environment, the terminal device is implemented as an intelligent in-vehicle device.

[0087] It should be noted that the execution entity of each step of the method provided in the above embodiment can be the same device, or the method can be executed by different devices. For example, the execution entity of steps 101a to 103a can be device A; for another example, the execution entity of steps 101a and 102a can be device A, and the execution entity of step 103a can be device B; and so on.

[0088] In addition, some of the processes described in the above embodiments and accompanying drawings include multiple operations that appear in a specific order. However, it should be understood that these operations may not be executed in the order in which they appear herein or may be executed in parallel. The sequence numbers of the operations, such as 101a, 102a, etc., are merely used to distinguish between different operations and do not represent any execution order. In addition, these processes may include more or fewer operations, and these operations may be executed sequentially or in parallel.

[0089] Figure 5 This is a schematic diagram of the structure of a terminal device provided by an exemplary embodiment of the present application. Figure 5 As shown, the terminal device includes: a microphone array 53, a memory 54 and a processor 55.

[0090] The memory 54 is used to store computer programs and can be configured to store various other data to support operations on the terminal device. Examples of such data include instructions for any application program or method operating on the terminal device.

[0091] The memory 54 may be implemented by any type of volatile or non-volatile memory device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk, or optical disk.

[0092] The processor 55 is coupled to the memory 54 and is used to execute the computer program in the memory 54 to: obtain the current audio signal collected by the microphone array 53, the microphone array 53 includes at least two microphones; generate the current sound source spatial distribution information corresponding to the current audio signal based on the phase difference information of the current audio signal collected by the at least two microphones; and identify whether the current audio signal is an overlapping speech based on the current sound source spatial distribution information and the conversion relationship between a single speech and an overlapping speech learned based on historical audio signals.

[0093] In an optional embodiment, when the processor 55 generates the current sound source spatial distribution information corresponding to the current audio signal based on the phase difference information of the current audio signal collected by at least two microphones, it is specifically used to: calculate the wave spectrum corresponding to the current audio signal based on the phase difference information of the current audio signal collected by at least two microphones, and the wave spectrum reflects the spatial distribution of the current sound source.

[0094] In an optional embodiment, when the processor 55 calculates the arrival spectrogram corresponding to the current audio signal based on the phase difference information of the current audio signal collected by at least two microphones, it is specifically used to: accumulate the phase difference information of the current audio signal collected by any two microphones for any direction in the position space to obtain the probability of the direction being the current sound source position; and generate the arrival spectrogram corresponding to the current audio signal based on the probability of each direction in the position space being the current sound source position.

[0095] In an optional embodiment, when the processor 55 identifies whether the current audio signal is overlapping speech based on the current sound source spatial distribution information and the conversion relationship between single speech and overlapping speech learned based on historical audio signals, it is specifically used to: calculate the peak information of the current sound source spatial distribution information as the current observation state of the hidden Markov model HMM, and use the single speech and overlapping speech as the two hidden states of the HMM; input the current observation state into the HMM, and calculate the probability that the current observation state corresponds to the hidden state based on the jump relationship between the two hidden states learned by the HMM and the historical observation state as a prerequisite; and identify whether the current audio signal is overlapping speech based on the probability that the current observation state corresponds to the hidden state.

[0096] In an optional embodiment, the processor 55 is further used to: if the current audio signal is identified as overlapping speech, determine at least two valid sound source directions based on the current sound source spatial distribution information; perform speech enhancement on the audio signals at the at least two valid sound source directions, and perform speech recognition on the enhanced audio signals at the at least two valid sound source directions respectively.

[0097] In an optional embodiment, when the processor 55 determines at least two valid sound source directions based on the current sound source spatial distribution information, it is specifically used to: when the current sound source spatial distribution information includes the probability of each direction being the current sound source position, use the two directions with the highest probability of being the current sound source position as the valid sound source directions.

[0098] In an optional embodiment, the processor 55 is further used to: if the current audio signal is recognized as a single voice, then the direction with the highest probability of being the current sound source position is used as the effective sound source direction; perform speech enhancement on the audio signal at the effective sound source direction, and perform speech recognition on the enhanced audio signal at the effective sound source direction.

[0099] In an optional embodiment, before identifying whether the current audio signal is overlapping speech, the processor 55 is further used to: calculate the direction of arrival of the current audio signal based on the current sound source spatial distribution information; select one microphone from at least two microphones as the target microphone based on the direction of arrival; and perform voice endpoint detection (VAD) on the current audio signal collected by the target microphone to determine whether the current audio signal is a speech signal.

[0100] In an optional embodiment, the terminal device is a conference device, a sound pickup device, a robot, a smart set-top box, a smart TV, a smart speaker, and a smart car device.

[0101] The terminal device provided in the embodiment of the present application can use a microphone array to collect audio signals, and generate sound source spatial distribution information corresponding to the audio signal based on the phase difference information of the audio signal collected by each microphone in the microphone array. Then, based on the sound source spatial distribution information and combined with the conversion relationship between single speech and overlapping speech learned based on historical audio signals, it can identify whether the current audio signal is overlapping speech. Compared with single-channel audio, the audio signal collected by the microphone array contains sound source spatial distribution information, so that it can accurately identify whether the current audio signal is overlapping speech, thereby meeting product-level detection requirements.

[0102] Further, if Figure 5 As shown, the terminal device also includes: a communication component 56, a display 57, a power component 58, a speaker 59 and other components. Figure 5 Only some components are shown schematically, which does not mean that the terminal equipment only includes Figure 5 The components shown. It should be noted that Figure 5 The components in the dotted box are optional components, not mandatory components, and the specific components depend on the product form of the terminal device.

[0103] In an optional embodiment, the above-mentioned terminal device can be applied to different application scenarios, and when applied to different application scenarios, it is specifically implemented as different device forms.

[0104] For example, the terminal device can be implemented as a conference device, and the implementation structure of the conference device is similar to Figure 5 The implementation structure of the terminal device shown is the same or similar, please refer to Figure 5 The terminal device is implemented as shown in the figure. Figure 5 The differences between the terminal devices in the illustrated embodiments primarily lie in the different functions implemented by the processor executing the computer program stored in the memory. For the conferencing device, its processor executing the computer program stored in the memory can be used to: obtain the current conference signal collected by a microphone array in the conference scene, where the microphone array includes at least two microphones; generate current sound source spatial distribution information corresponding to the current conference signal based on the phase difference information collected by the at least two microphones; and identify whether the current conference signal is overlapping speech based on the current sound source spatial distribution information and the conversion relationship between single speech and overlapping speech learned based on historical conference signals.

[0105] For another example, the terminal device can be implemented as a teaching device, and the implementation structure of the teaching device is similar to Figure 5 The implementation structure of the terminal device shown is the same or similar, please refer to Figure 5 The structure of the terminal device shown is realized. Figure 5 The differences between the terminal devices in the illustrated embodiments primarily lie in the different functions implemented by the processor executing the computer program stored in the memory. For teaching devices, the processor executing the computer program stored in the memory can be used to: obtain the current classroom signal collected by a microphone array in the teaching environment, where the microphone array includes at least two microphones; generate current sound source spatial distribution information corresponding to the current classroom signal based on the phase difference information collected by the at least two microphones; and identify whether the current classroom signal is overlapping speech based on the current sound source spatial distribution information and the conversion relationship between single speech and overlapping speech learned based on historical classroom signals.

[0106] For another example, the terminal device can be implemented as an intelligent vehicle-mounted device, and the implementation structure of the intelligent vehicle-mounted device is similar to Figure 5 The implementation structure of the terminal device shown is the same or similar, please refer to Figure 5 The structure of the terminal device shown is realized. Figure 5The differences between the terminal devices in the illustrated embodiments primarily lie in the different functions implemented by the processor executing the computer program stored in the memory. For the intelligent in-vehicle device, its processor executing the computer program stored in the memory can be used to: obtain a current audio signal collected by a microphone array in the in-vehicle environment, where the microphone array includes at least two microphones; generate current sound source spatial distribution information corresponding to the current audio signal based on phase difference information collected by the at least two microphones; and identify whether the current audio signal is overlapping speech based on the current sound source spatial distribution information and the conversion relationship between single speech and overlapping speech learned based on historical audio signals.

[0107] Accordingly, an embodiment of the present application further provides a computer-readable storage medium storing a computer program. When the computer program is executed by a processor, the processor is enabled to implement each step in each method embodiment provided in the embodiment of the present application.

[0108] Accordingly, an embodiment of the present application also provides a computer program product, including a computer program / instruction. When the computer program / instruction is executed by a processor, the processor is caused to implement the steps in the various methods provided in the embodiments of the present application.

[0109] Figure 6 This is a schematic diagram of the structure of a server device provided by an exemplary embodiment of the present application. Figure 6 As shown, the server device includes: a memory 64 and a processor 65.

[0110] The memory 64 is used to store computer programs and can be configured to store various other data to support operations on the server device. Examples of such data include instructions for any application or method operating on the server device.

[0111] The memory 64 may be implemented by any type of volatile or non-volatile memory device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk, or optical disk.

[0112] The processor 65 is coupled to the memory 64 and is used to execute the computer program in the memory 64 to: receive a current audio signal collected by at least two microphones in a microphone array uploaded by a terminal device; generate current sound source spatial distribution information corresponding to the current audio signal based on phase difference information of the current audio signal collected by the at least two microphones; and identify whether the current audio signal is overlapping speech based on the current sound source spatial distribution information and the conversion relationship between single speech and overlapping speech learned based on historical audio signals.

[0113] In an optional embodiment, when the processor 65 generates the current sound source spatial distribution information corresponding to the current audio signal based on the phase difference information of the current audio signal collected by at least two microphones, it is specifically used to: calculate the wave spectrum corresponding to the current audio signal based on the phase difference information of the current audio signal collected by at least two microphones, and the wave spectrum reflects the spatial distribution of the current sound source.

[0114] In an optional embodiment, when the processor 65 calculates the arrival spectrogram corresponding to the current audio signal based on the phase difference information of the current audio signal collected by at least two microphones, the processor is specifically used to: accumulate the phase difference information of the current audio signal collected by any two microphones for any direction in the position space to obtain the probability of the direction being the current sound source position; and generate the arrival spectrogram corresponding to the current audio signal based on the probability of each direction in the position space being the current sound source position.

[0115] In an optional embodiment, when the processor 65 identifies whether the current audio signal is overlapping speech based on the current sound source spatial distribution information and the conversion relationship between single speech and overlapping speech learned based on historical audio signals, it is specifically used to: calculate the peak information of the current sound source spatial distribution information as the current observation state of the hidden Markov model HMM, and use the single speech and overlapping speech as the two hidden states of the HMM; input the current observation state into the HMM, combine the jump relationship between the two hidden states learned by the HMM, and calculate the probability that the current observation state corresponds to the hidden state with the historical observation state as a prerequisite; and identify whether the current audio signal is overlapping speech based on the probability that the current observation state corresponds to the hidden state.

[0116] In an optional embodiment, the processor 65 is further used to: if the current audio signal is identified as overlapping speech, determine at least two valid sound source directions based on the current sound source spatial distribution information; perform speech enhancement on the audio signals at the at least two valid sound source directions, and perform speech recognition on the enhanced audio signals at the at least two valid sound source directions respectively.

[0117] In an optional embodiment, when the processor 65 determines at least two valid sound source directions based on the current sound source spatial distribution information, it is specifically used to: when the current sound source spatial distribution information includes the probability of each direction being the current sound source position, use the two directions with the highest probability of being the current sound source position as the valid sound source directions.

[0118] In an optional embodiment, the processor 65 is also used to: if the current audio signal is recognized as a single voice, then the direction with the highest probability of being the current sound source position is used as the effective sound source direction; perform speech enhancement on the audio signal at the effective sound source direction, and perform speech recognition on the enhanced audio signal at the effective sound source direction.

[0119] In an optional embodiment, before identifying whether the current audio signal is overlapping speech, the processor 65 is further used to: calculate the direction of arrival of the current audio signal based on the current sound source spatial distribution information; select one microphone from at least two microphones as the target microphone based on the direction of arrival; and perform voice endpoint detection (VAD) on the current audio signal collected by the target microphone to determine whether the current audio signal is a speech signal.

[0120] The server device provided in the embodiment of the present application can use a microphone array to collect audio signals, and generate sound source spatial distribution information corresponding to the audio signal based on the phase difference information of the audio signal collected by each microphone in the microphone array. Then, based on the sound source spatial distribution information and combined with the conversion relationship between single speech and overlapping speech learned based on historical audio signals, it can identify whether the current audio signal is overlapping speech. Compared with single-channel audio, the audio signal collected by the microphone array contains sound source spatial distribution information, so that it can accurately identify whether the current audio signal is overlapping speech, meeting product-level detection requirements.

[0121] Further, if Figure 6 As shown, the server device also includes: a communication component 66, a power supply component 68 and other components. Figure 6 Only some components are shown schematically, which does not mean that the server device only includes Figure 6 The components shown. It should be noted that Figure 6 The components in the dotted box are optional components, not mandatory components, and the specific components depend on the product form of the server device.

[0122] Accordingly, an embodiment of the present application further provides a computer-readable storage medium storing a computer program. When the computer program is executed by a processor, the processor is enabled to implement each step in each method embodiment provided in the embodiment of the present application.

[0123] Accordingly, an embodiment of the present application also provides a computer program product, including a computer program / instruction. When the computer program / instruction is executed by a processor, the processor is caused to implement the steps in the various methods provided in the embodiments of the present application.

[0124] above Figure 5 and Figure 6The communication component is configured to facilitate wired or wireless communication between the device where the communication component is located and other devices. The device where the communication component is located can access a wireless network based on a communication standard, such as WiFi, 2G, 3G, 4G / LTE, 5G and other mobile communication networks, or a combination thereof. In an exemplary embodiment, the communication component receives a broadcast signal or broadcast-related information from an external broadcast management system via a broadcast channel. In an exemplary embodiment, the communication component also includes a near field communication (NFC) module to facilitate short-range communication. For example, the NFC module can be implemented based on radio frequency identification (RFID) technology, infrared data association (IrDA) technology, ultra-wideband (UWB) technology, Bluetooth (BT) technology and other technologies.

[0125] above Figure 5 The display in the embodiment includes a screen, which may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen may be implemented as a touch screen to receive input signals from a user. The touch panel includes one or more touch sensors to sense touches, slides, and gestures on the touch panel. The touch sensor may not only sense the boundaries of a touch or slide action, but also detect the duration and pressure associated with the touch or slide operation.

[0126] above Figure 5 and Figure 6 The power supply component in a device provides power to various components of the device in which the power supply component is located. The power supply component may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power to the device in which the power supply component is located.

[0127] Those skilled in the art will appreciate that the embodiments of the present application can be provided as methods, systems, or computer program products. Therefore, the present application can adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment in combination with software and hardware. Moreover, the present application can adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.

[0128] The present application is described with reference to the flow chart and / or block diagram of the method, device (system), and computer program product according to the embodiment of the present application. It should be understood that each flow process and / or box in the flow chart and / or block diagram and the combination of the flow process and / or box in the flow chart and / or block diagram can be realized by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processing machine or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device produce a device for realizing the function specified in one flow chart flow or multiple flows and / or one box or multiple boxes of the block diagram.

[0129] These computer program instructions may also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce a product including an instruction device that implements the functions specified in one or more processes in the flowchart and / or one or more boxes in the block diagram.

[0130] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, so that the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one or more processes in the flowchart and / or one or more boxes in the block diagram.

[0131] In a typical configuration, a computing device includes one or more processors (CPUs), input / output interfaces, network interfaces, and memory.

[0132] Memory may include non-permanent storage in a computer-readable medium, random access memory (RAM) and / or non-volatile memory in the form of read-only memory (ROM) or flash RAM. Memory is an example of a computer-readable medium.

[0133] Computer-readable media includes permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. The information can be computer-readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer-readable media does not include transitory computer-readable media (transitory media), such as modulated data signals and carrier waves.

[0134] It should also be noted that the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, commodity, or apparatus that includes a series of elements includes not only those elements but also other elements not explicitly listed, or includes elements inherent to such process, method, commodity, or apparatus. In the absence of further limitations, an element defined by the phrase "comprises a ..." does not exclude the presence of other identical elements in the process, method, commodity, or apparatus that includes the element.

[0135] The foregoing is merely an embodiment of the present application and is not intended to limit the present application. For those skilled in the art, the present application may have various changes and variations. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present application should all be included within the scope of the claims of the present application.

Claims

1. A method for processing an audio signal, characterized in that: include: Acquire a current audio signal collected by a microphone array, where the microphone array includes at least two microphones; generating current sound source spatial distribution information corresponding to the current audio signal based on phase difference information of the current audio signal collected by the at least two microphones; Identifying whether the current audio signal is overlapping speech based on the current sound source spatial distribution information and the conversion relationship between single speech and overlapping speech learned based on historical audio signals includes: Calculating peak information of the current sound source spatial distribution information as the current observation state and inputting it into a hidden Markov model HMM, and using the HMM to identify whether the current audio signal is overlapping speech; The HMM is obtained by continuously learning the transition relationship between hidden states and the emission relationship from hidden states to observed states based on the state of historical audio signals.

2. The method according to claim 1, characterized in that Generating current sound source spatial distribution information corresponding to the current audio signal according to phase difference information of the current audio signal collected by the at least two microphones, including: According to the phase difference information of the current audio signal collected by the at least two microphones, a wave arrival spectrogram corresponding to the current audio signal is calculated, and the wave arrival spectrogram reflects the spatial distribution of the current sound source.

3. The method according to claim 2, characterized in that Calculating a wave arrival spectrogram corresponding to the current audio signal based on phase difference information of the current audio signal collected by the at least two microphones, including: For any position in the position space, the phase difference information of the current audio signal collected by any two microphones is accumulated to obtain the probability that the position is the current sound source position; According to the probability of each direction in the position space being the current sound source position, a wave arrival spectrogram corresponding to the current audio signal is generated.

4. The method according to any one of claims 1 to 3, characterized in that Using the HMM to identify whether the current audio signal is overlapping speech includes: Use single speech and overlapping speech as two hidden states of HMM; Input the current observation state into HMM, combine the jump relationship between the two hidden states learned by HMM, and calculate the probability that the current observation state corresponds to the hidden state based on the historical observation state; According to the probability that the current observed state corresponds to the hidden state, it is identified whether the current audio signal is overlapping speech.

5. The method according to any one of claims 1 to 3, characterized in that Also includes: If the current audio signal is identified as overlapping speech, determining at least two valid sound source directions based on the current sound source spatial distribution information; Speech enhancement is performed on the audio signals in the at least two effective sound source directions, and speech recognition is performed on the enhanced audio signals in the at least two effective sound source directions respectively.

6. The method according to claim 5, characterized in that Determining at least two valid sound source positions according to the current sound source spatial distribution information includes: In a case where the current sound source spatial distribution information includes the probabilities of each direction being the current sound source position, the two directions with the highest probabilities of being the current sound source position are taken as valid sound source directions.

7. The method according to claim 6, characterized in that Also includes: If the current audio signal is recognized as a single voice, the direction with the highest probability of being the current sound source position is taken as the effective sound source direction; Speech enhancement is performed on the audio signal in the effective sound source direction, and speech recognition is performed on the enhanced audio signal in the effective sound source direction.

8. The method according to any one of claims 1 to 3, characterized in that Before identifying whether the current audio signal is overlapping speech, the method further includes: Calculating the direction of arrival of the current audio signal based on the current sound source spatial distribution information; selecting a microphone from the at least two microphones as a target microphone according to the direction of arrival; Perform voice endpoint detection (VAD) on the current audio signal collected by the target microphone to determine whether the current audio signal is a voice signal.

9. An audio signal processing method, characterized in that: Applicable to a conference device, the conference device including a microphone array, the method comprising: Acquire a current conference signal collected by the microphone array in a conference scene, where the microphone array includes at least two microphones; Generating current sound source spatial distribution information corresponding to the current conference signal based on phase difference information of the current conference signal collected by the at least two microphones; Identifying whether the current conference signal is overlapping speech based on the current sound source spatial distribution information and the conversion relationship between single speech and overlapping speech learned based on historical conference signals includes: Calculating the peak value information of the current sound source spatial distribution information as the current observation state and inputting it into the Hidden Markov Model HMM, and using the HMM to identify whether the current conference signal is overlapping speech; The HMM is obtained by continuously learning the conversion relationship between hidden states and the emission relationship from hidden states to observation states based on the state of historical conference signals.

10. An audio signal processing method, characterized in that: Applicable to a teaching device, the teaching device including a microphone array, the method comprising: Acquiring a current classroom signal collected by the microphone array in a teaching environment, wherein the microphone array includes at least two microphones; Generating current sound source spatial distribution information corresponding to the current classroom signal based on phase difference information of the current classroom signal collected by the at least two microphones; According to the current sound source spatial distribution information, combined with the conversion relationship between single speech and overlapping speech learned based on historical classroom signals, identifying whether the current classroom signal is overlapping speech includes: Calculate the peak value information of the current sound source spatial distribution information as the current observation state and input it into the Hidden Markov Model HMM, and use the HMM to identify whether the current classroom signal is overlapping speech; The HMM is obtained by continuously learning the conversion relationship between hidden states and the emission relationship from hidden states to observation states based on the state of historical classroom signals.

11. An audio signal processing method, characterized in that: Applicable to an intelligent vehicle-mounted device, the intelligent vehicle-mounted device including a microphone array, the method comprising: Acquire a current audio signal collected by the microphone array in a vehicle-mounted environment, wherein the microphone array includes at least two microphones; generating current sound source spatial distribution information corresponding to the current audio signal based on phase difference information of the current audio signal collected by the at least two microphones; Identifying whether the current audio signal is overlapping speech based on the current sound source spatial distribution information and the conversion relationship between single speech and overlapping speech learned based on historical audio signals includes: Calculating peak information of the current sound source spatial distribution information as the current observation state and inputting it into a hidden Markov model HMM, and using the HMM to identify whether the current audio signal is overlapping speech; The HMM is obtained by continuously learning the transition relationship between hidden states and the emission relationship from hidden states to observed states based on the state of historical audio signals.

12. A terminal device, characterized in that: include: Memory, processor, and microphone array; The memory is used to store computer programs; The processor is coupled to the memory and is configured to execute the computer program to: obtain a current audio signal collected by a microphone array, the microphone array including at least two microphones; generate current sound source spatial distribution information corresponding to the current audio signal based on phase difference information of the current audio signal collected by the at least two microphones; and identify whether the current audio signal is overlapping speech based on the current sound source spatial distribution information and in combination with a conversion relationship between a single speech and overlapping speech learned based on historical audio signals, specifically for: Calculating peak information of the current sound source spatial distribution information as the current observation state and inputting it into a hidden Markov model HMM, and using the HMM to identify whether the current audio signal is overlapping speech; The HMM is obtained by continuously learning the transition relationship between hidden states and the emission relationship from hidden states to observed states based on the state of historical audio signals.

13. The terminal device according to claim 12, characterized in that The terminal devices include conference equipment, sound pickup equipment, robots, smart set-top boxes, smart TVs, smart speakers and smart car-mounted equipment.

14. A conference device, characterized in that: include: Memory, processor, and microphone array; The memory is used to store computer programs; The processor is coupled to the memory and is configured to execute the computer program to: obtain a current conference signal collected by the microphone array in a conference scene, the microphone array comprising at least two microphones; generate current sound source spatial distribution information corresponding to the current conference signal based on phase difference information of the current conference signal collected by the at least two microphones; identify whether the current conference signal is overlapping speech based on the current sound source spatial distribution information and in combination with a conversion relationship between a single speech and overlapping speech learned based on historical conference signals, specifically for: calculating peak information of the current sound source spatial distribution information as a current observation state input into a hidden Markov model HMM, and using the HMM to identify whether the current conference signal is overlapping speech; The HMM is obtained by continuously learning the conversion relationship between hidden states and the emission relationship from hidden states to observation states based on the state of historical conference signals.

15. A teaching device, characterized in that: include: Memory, processor, and microphone array; The memory is used to store computer programs; The processor is coupled to the memory and is used to execute the computer program to: obtain a current classroom signal collected by the microphone array in a teaching environment, the microphone array including at least two microphones; generate current sound source spatial distribution information corresponding to the current classroom signal based on phase difference information of the current classroom signal collected by the at least two microphones; identify whether the current classroom signal is overlapping speech based on the current sound source spatial distribution information and the conversion relationship between single speech and overlapping speech learned based on historical classroom signals, specifically for: calculating peak information of the current sound source spatial distribution information as the current observation state input into the hidden Markov model HMM, and using the HMM to identify whether the current classroom signal is overlapping speech; The HMM is obtained by continuously learning the conversion relationship between hidden states and the emission relationship from hidden states to observation states based on the state of historical classroom signals.

16. An intelligent vehicle-mounted device, characterized in that: include: Memory, processor, and microphone array; The memory is used to store computer programs; The processor is coupled to the memory and is configured to execute the computer program to: obtain a current audio signal collected by the microphone array in a vehicle-mounted environment, the microphone array comprising at least two microphones; generate current sound source spatial distribution information corresponding to the current audio signal based on phase difference information of the current audio signal collected by the at least two microphones; identify whether the current audio signal is overlapping speech based on the current sound source spatial distribution information and a conversion relationship between a single speech and overlapping speech learned based on historical audio signals, specifically to: calculate peak information of the current sound source spatial distribution information as a current observation state input into a hidden Markov model HMM, and use the HMM to identify whether the current audio signal is overlapping speech; The HMM is obtained by continuously learning the transition relationship between hidden states and the emission relationship from hidden states to observed states based on the state of historical audio signals.

17. An audio signal processing system, characterized in that: include: Terminal equipment and server equipment; The terminal device includes a microphone array, wherein the microphone array includes at least two microphones for collecting a current audio signal; The terminal device is used to upload the current audio signal collected by the at least two microphones to the server device; The server device is configured to generate current sound source spatial distribution information corresponding to the current audio signal based on phase difference information of the current audio signal collected by the at least two microphones; identify whether the current audio signal is overlapping speech based on the current sound source spatial distribution information and the conversion relationship between single speech and overlapping speech learned based on historical audio signals, specifically for: calculating peak information of the current sound source spatial distribution information as the current observation state input into a hidden Markov model HMM, and using the HMM to identify whether the current audio signal is overlapping speech; The HMM is obtained by continuously learning the transition relationship between hidden states and the emission relationship from hidden states to observed states based on the state of historical audio signals.

18. A server device, characterized in that: include: memory and processor; The memory is used to store computer programs; The processor, coupled to the memory, is configured to execute the computer program to: The method includes receiving a current audio signal collected by at least two microphones in a microphone array uploaded by a terminal device; generating current sound source spatial distribution information corresponding to the current audio signal based on phase difference information of the current audio signal collected by the at least two microphones; identifying whether the current audio signal is overlapping speech based on the current sound source spatial distribution information and a conversion relationship between single speech and overlapping speech learned based on historical audio signals, specifically for: calculating peak information of the current sound source spatial distribution information as the current observation state input into a hidden Markov model HMM, and using the HMM to identify whether the current audio signal is overlapping speech; The HMM is obtained by continuously learning the transition relationship between hidden states and the emission relationship from hidden states to observed states based on the state of historical audio signals.

19. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the processor is caused to implement the steps of the method according to any one of claims 1 to 11.

20. A computer program product comprising a computer program / instructions, characterized in that When the computer program / instructions are executed by a processor, the processor is caused to implement the steps in the method according to any one of claims 1 to 11.

Citation Information

Patent Citations

  • Superimposed sound detection method, device and equipment

    CN111640456A

  • Speaker recognition / location using neural network

    CN112088403A