Sound source positioning method and device, equipment and storage medium
By utilizing beam parameters and speech signal features in audio frame signals to filter and cluster directional pickup signals, the problems of noise and multi-source interference are solved, and more accurate sound source localization is achieved.
Patent Information
- Application Number
- CN202511456349.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-13
- Publication Date
- 2025-11-07
- Estimated Expiration
- 2045-10-13
AI Technical Summary
In the prior art, noise or multi-source interference in the audio frame signal leads to poor accuracy in sound source localization.
By acquiring preset audio frame signals, directional sound pickup is performed using the first beam parameters. Combining the probability, energy information, and similarity of the speech signals, a second directional sound pickup signal is determined from multiple directional sound pickup signals, reducing noise and multi-source interference and improving positioning accuracy.
It effectively reduces noise and multi-source interference, improving the accuracy and convenience of sound source localization for audio frame signals.
Smart Images

Figure CN120908752A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of sound source positioning, and particularly relates to a sound source positioning method and device, equipment and a storage medium. BACKGROUND
[0002] In related technologies, an electronic device can perform sound source positioning on a corresponding audio frame signal to determine positioning information of a sound source. However, in the case that noise exists in the audio frame signal or multiple sound sources exist in the audio frame signal, the noise or multiple sound sources existing in the audio frame signal will interfere with the sound source positioning process of the audio frame signal, and thus lead to poor sound source positioning accuracy of the audio frame signal. Therefore, it is urgent to improve the sound source positioning accuracy of the audio frame signal. SUMMARY
[0003] The main purpose of the present application is to provide a sound source positioning method, device, equipment and storage medium, aiming at solving the technical problem that the sound source positioning process of the audio frame signal is easily interfered by noise or multiple sound sources, and thus leading to poor sound source positioning accuracy of the audio frame signal.
[0004] In a first aspect, the present application provides a sound source positioning method, comprising: obtaining a preset audio frame signal, wherein a speech signal exists in the preset audio frame signal; directively picking up the preset audio frame signal according to a first beam parameter to obtain a plurality of first directively picked-up signals corresponding to the preset audio frame signal; determining at least one second directively picked-up signal from the plurality of first directively picked-up signals according to at least two of a first probability that a speech signal exists in the plurality of first directively picked-up signals, first energy information of the plurality of first directively picked-up signals, and a first similarity between the plurality of first directively picked-up signals; determining first positioning information of a sound source corresponding to the speech signal in the preset audio frame signal according to the at least one second directively picked-up signal.
[0005] In a second aspect, the present application provides a sound source positioning device, comprising: an audio acquisition module configured to obtain a preset audio frame signal, wherein a speech signal exists in the preset audio frame signal; a first signal determination module configured to directively pick up the preset audio frame signal according to a first beam parameter to obtain a plurality of first directively picked-up signals corresponding to the preset audio frame signal; determine at least one second directional sound pickup signal from the plurality of first directional sound pickup signals according to at least two of a first probability that a voice signal exists in the plurality of first directional sound pickup signals, first energy information of the plurality of first directional sound pickup signals, and a first similarity between the plurality of first directional sound pickup signals; determine first positioning information of a sound source corresponding to the voice signal in the preset audio frame signal according to the at least one second directional sound pickup signal.
[0006] In a third aspect, the present application provides an electronic device, which comprises a memory and a processor; The memory is configured to store a computer program. The processor is configured to execute the computer program and implement the steps of the sound source positioning method when the computer program is executed.
[0007] In a fourth aspect, the present application provides a computer readable storage medium, which stores a computer program. When the computer program is executed by a processor, the steps of the sound source positioning method are implemented.
[0008] The present application provides a sound source positioning method, device, equipment and storage medium. The sound source positioning method comprises: obtaining a preset audio frame signal, and a voice signal exists in the preset audio frame signal; performing directional sound pickup on the preset audio frame signal according to first beam parameters to obtain a plurality of first directional sound pickup signals corresponding to the preset audio frame signal; determining at least one second directional sound pickup signal from the plurality of first directional sound pickup signals according to at least two of a first probability that a voice signal exists in the plurality of first directional sound pickup signals, first energy information of the plurality of first directional sound pickup signals, and a first similarity between the plurality of first directional sound pickup signals; and determining first positioning information of a sound source corresponding to the voice signal in the preset audio frame signal according to the at least one second directional sound pickup signal.
[0009] In the case of determining the second directional sound pickup signal by comprehensively considering at least two of the first probability, the first energy information and the first similarity, the second directional sound pickup signal is equivalent to being determined by comprehensively considering at least two of voice signal detection, energy information judgment and similarity judgment on the preset audio frame signal. The second directional sound pickup signal can be used to distinguish the voice signal from noise signal in the preset audio frame signal, and can be used to determine whether the sound sources corresponding to the voice signals are the same sound source, thereby reducing noise interference or multiple sound source interference in the process of positioning the sound source of the preset audio frame signal. Based on the reduction of noise interference or multiple sound source interference in the process of positioning the sound source of the preset audio frame signal, the sound source positioning accuracy of the preset audio frame signal is improved. BRIEF DESCRIPTION OF DRAWINGS
[0010] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings needed to be used in the embodiments description will be briefly introduced. Obviously, the drawings in the following description are some embodiments of the present application, and other drawings can be obtained by those skilled in the art without any creative effort on the basis of these drawings.
[0011] Figure 1 is a flow diagram of a sound source positioning method provided by an embodiment of the present application; Figure 2 is a schematic block diagram of a sound source positioning device provided by an embodiment of the present application; Figure 3 is a schematic block diagram of an electronic device provided by an embodiment of the present application. DETAILED DESCRIPTION
[0012] The technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are some embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all the other embodiments obtained by those skilled in the art without any creative effort belong to the protection scope of the present application.
[0013] The flow diagram shown in the drawings is only an example description, and it is not necessary to include all the contents and operations / steps, and it is not necessary to execute in the described order. For example, some operations / steps can be decomposed, combined or partially merged, and thus the actual execution order can be changed according to the actual situation.
[0014] Embodiments of the present application provide a sound source positioning method, device, equipment and storage medium. The sound source positioning method can be applied to an electronic device. The electronic device can include a near-eye display device, a wearable device, a terminal device, etc., without limitation. The near-eye display device can include an augmented reality (AR) glasses, a virtual reality (VR) glasses, a mixed reality (MR) glasses, an AR helmet, a VR helmet, a MR helmet, etc., without limitation. The wearable device includes a smart watch, a smart ring, a smart bracelet, etc., without limitation. The terminal device includes a mobile phone, a television, a computer, etc., without limitation. The sound source positioning method can also be applied to a server, which can be a separate server, or a cloud server providing cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery network (CDN), and basic cloud computing services such as big data and artificial intelligence platforms.
[0015] Some embodiments of the present application will be described in detail below with reference to the accompanying drawings. The following examples and features in the examples can be combined with each other without conflict.
[0016] Please refer to Figure 1 , Figure 1 is a flowchart of a sound source positioning method provided by an embodiment of the present application. It should be noted that the sound source positioning method provided by the embodiments of the present application can be used in an electronic device, or can be used in a server, without limitation.
[0017] As Figure 1 shown, the sound source positioning method includes steps S101 to S104.
[0018] S101, obtaining a preset audio frame signal, the preset audio frame signal containing a speech signal.
[0019] Exemplarily, the preset audio frame signal is used to indicate an audio frame signal capable of sound source positioning processing. The speech signal is used to indicate a signal with human voice characteristics. For example, the speech signal includes at least one of biological human voice and non-biological human voice. The biological human voice can be understood as the sound emitted by the human himself. The non-biological human voice can be understood as the human voice played by the electronic device. For example, the non-biological voice can include human voice speech played by a sound box, human music, etc., without limitation.
[0020] In some embodiments, the electronic device can obtain an initial audio frame signal. Accordingly, the electronic device can detect whether a speech signal exists in the initial audio frame signal. If a speech signal exists in the initial audio frame signal, the electronic device can determine that the initial audio frame signal can be used for subsequent sound source positioning, and can then determine the initial audio frame signal as a preset audio frame signal. Accordingly, if a speech signal does not exist in the initial audio frame signal, the electronic device can determine that the initial audio frame signal does not need to be positioned, and then does not need to be determined as a preset audio frame signal.
[0021] In the case where the preset audio frame signal is obtained and a speech signal exists in the preset audio frame signal, the preset audio frame signal is used to determine the target positioning information of the sound source corresponding to the speech signal.
[0022] S102, according to the first beam parameter, the preset audio frame signal is directionally picked up to obtain a plurality of first directional pickup signals corresponding to the preset audio frame signal.
[0023] In the case where the preset audio frame signal is obtained, the preset audio frame signal can also include noise signals other than speech signals or a plurality of sound sources each corresponding to a speech signal. In order to reduce the noise signal interference in the preset audio frame signal and the adverse effects of mutual interference of a plurality of sound sources on the positioning of the preset audio frame signal, the electronic device can directionally pick up the preset audio frame signal to achieve spatial filtering of the preset audio frame signal, thereby selectively enhancing the sound source to be positioned and suppressing noise signals.
[0024] For example, the electronic device can directionally pick up the preset audio frame signal based on a fixed beam algorithm. In some embodiments, when the preset audio frame signal is directionally picked up based on the fixed beam algorithm, the preset audio frame signal can be directionally picked up according to the first beam parameter to obtain a plurality of first directional pickup signals corresponding to the preset audio frame signal.
[0025] The first beam parameter may comprise, for example, a first angle change step. The first angle change step may be used to determine a plurality of first scanning angles. For example, a sound source corresponding to a speech signal in a preset audio frame signal may exist at any extremely small angle in a space in which the electronic device is located. Under the limitation of system resources of the electronic device, the electronic device cannot perform directional sound pickup of the preset audio frame signal at each infinitesimal angle. Based on this, the electronic device may discretize the space in which the electronic device is located according to the first angle change step to obtain a plurality of first scanning angles. The plurality of first scanning angles may systematically cover the space in which the electronic device is located, and thus may be used by the electronic device to perform sound source positioning of the preset audio frame signal. The first angle change step may be used to indicate an angle difference between adjacent first scanning angles.
[0026] The electronic device may perform sound source positioning of the preset audio frame signal according to different first scanning angles. For example, the electronic device may perform directional sound pickup of the preset audio frame signal at a first scanning angle to obtain a first directional sound pickup signal of the preset audio frame signal corresponding to the first scanning angle. The first directional sound pickup signal corresponding to the first scanning angle is obtained by processing the preset audio frame signal with the first scanning angle as a signal enhancement direction and a direction other than the first scanning angle as a signal suppression direction. Similarly, for a plurality of first scanning angles, a first directional sound pickup signal of the preset audio frame signal corresponding to each of the plurality of first scanning angles may be determined. The first scanning angle and the first directional sound pickup signal correspond to each other.
[0027] Correspondingly, since each first directional sound pickup signal is obtained by performing directional sound pickup of the preset audio frame signal according to a different first scanning angle, the electronic device may determine first positioning information of a sound source corresponding to a speech signal in the preset audio frame signal according to a first scanning angle corresponding to each first directional sound pickup signal. For example, the electronic device may determine a first sound source angle of the sound source corresponding to the speech signal in the preset audio frame signal according to the first scanning angle. The first sound source angle may be used to determine the first positioning information of the sound source.
[0028] In a case where the preset audio frame signal is directional sound picked up according to the first beam parameter to obtain a plurality of first directional sound pickup signals corresponding to the preset audio frame signal, the electronic device may directly perform directional sound pickup of the preset audio frame signal by using the first beam parameter, which is beneficial to improving the convenience of directional sound pickup of the preset audio frame signal. The plurality of first directional sound pickup signals may be used to subsequently evaluate the first positioning information of the sound source corresponding to the speech signal in the preset audio frame signal, which is beneficial to improving the convenience of determining the first positioning information of the sound source corresponding to the speech signal in the preset audio frame signal.
[0029] S103, determining at least one second directional sound signal from the plurality of first directional sound signals according to at least two of the first probability of the presence of the speech signal in the plurality of first directional sound signals, the first energy information of the plurality of first directional sound signals, and the first similarity between the plurality of first directional sound signals.
[0030] In a case where the plurality of first directional sound signals corresponding to the preset audio frame signal are acquired, the electronic device can screen and cluster the plurality of first directional sound signals, to determine the first positioning information of the sound source corresponding to the speech signal in the preset audio frame signal subsequently.
[0031] For example, the electronic device can perform voice activity detection on each first directional sound signal to determine the first probability of the presence of the speech signal in each first directional sound signal, and then determine whether the speech signal is present in each first directional sound signal. For example, based on a preset voice activity detection (VAD) model, voice activity detection is performed on the first directional sound signal to obtain the first probability of the presence of the speech signal in the first directional sound signal; when the first probability is greater than or equal to a first probability threshold, it is determined that the speech signal is present in the first directional sound signal; when the first probability is less than the first probability threshold, it is determined that the speech signal is not present in the first directional sound signal. The first probability threshold can be preset or set by the user, which is not limited herein. In a case where the first directional sound signal is input into the VAD model, the VAD model can perform voice activity detection on the first directional sound signal to determine whether the speech signal is present in the first directional sound signal. For example, the VAD model can perform feature extraction on the first directional sound signal to obtain the feature value of the speech signal corresponding to the first directional sound signal. In a case where the feature value of the speech signal in the first directional sound signal is greater than or equal to a first feature value threshold, the VAD model can determine that the first probability of the presence of the speech signal in the first directional sound signal is greater than or equal to the first probability threshold, and then determine that the speech signal is present in the first directional sound signal. Correspondingly, in a case where the feature value of the speech signal in the first directional sound signal is less than the first feature value threshold, the VAD model can determine that the first probability of the presence of the speech signal in the first directional sound signal is less than the first probability threshold, and then determine that the speech signal is not present in the first directional sound signal.
[0032] The electronic device can determine the first energy information of each first directional sound pickup signal to perform energy judgment on each first directional sound pickup signal, measure the energy of each first directional sound pickup signal, and further assist in determining whether the sound source is concentrated in the first scanning angle corresponding to each first directional sound pickup signal. For example, the electronic device can perform time-domain energy calculation on the first directional sound pickup signal to obtain the first energy information of the first directional sound pickup signal, or the electronic device can perform frequency-domain energy calculation on the first directional sound pickup signal to obtain the first energy information of the first directional sound pickup signal. In the process of performing time-domain energy calculation on the first directional sound pickup signal, the amplitude of the first directional sound pickup signal can be directly operated, such as determining the sum of squares of all sample point amplitudes of the first directional sound pickup signal to obtain the first energy information of the first directional sound pickup signal. In the process of performing frequency-domain energy calculation on the first directional sound pickup signal, the first directional sound pickup signal can be subjected to fast Fourier transform to convert the first directional sound pickup signal from the time domain to the frequency domain, and then the power spectrum of the first directional sound pickup signal after fast Fourier transform is calculated, and then the frequency band energy of the first directional sound pickup signal after fast Fourier transform is calculated according to the power spectrum, so as to determine the first energy information of the first directional sound pickup signal. For a plurality of first directional sound pickup signals, if the first energy information of the first directional sound pickup signal is larger, the possibility that the sound source is concentrated in the first scanning angle corresponding to the first directional sound pickup signal is greater; if the first energy information of the first directional sound pickup signal is smaller, the possibility that the sound source is concentrated in the first scanning angle corresponding to the first directional sound pickup signal is smaller. For example, when the first energy information of the first directional sound pickup signal is greater than or equal to the first energy threshold, it is determined that the sound source is concentrated in the first scanning angle corresponding to the first directional sound pickup signal; when the first energy information of the first directional sound pickup signal is less than the first energy threshold, it is determined that the sound source is not concentrated in the first scanning angle corresponding to the first directional sound pickup signal. The first energy threshold can be pre-set or set by the user, which is not limited herein.
[0033] Correspondingly, the electronic device can perform similarity determination on the multiple first directional sound pickup signals to determine the first similarity between the multiple first directional sound pickup signals, and further determine whether the sound source is concentrated in the first scanning angle corresponding to the multiple first directional sound pickup signals. For example, the voice signals existing in the preset audio frame signal can correspond to different sound sources, i.e., the preset audio frame signal includes voice signals corresponding to multiple sound sources respectively. In order to reduce the mutual interference of the multiple sound sources in the process of sound source positioning of the preset audio frame signal, and further improve the sound source positioning accuracy of the preset audio frame signal, the electronic device can determine whether the voice signals existing in the multiple first directional sound pickup signals correspond to the same sound source by determining the first similarity between the multiple first directional sound pickup signals. Taking multiple first directional sound pickup signals including signal A, signal B, and signal C as an example. The electronic device can determine the first similarity between signal A and signal B and signal C respectively. The greater the first similarity between signal A and signal B, the greater the possibility that the voice signals existing in signal A and signal B correspond to the same sound source. Correspondingly, the smaller the first similarity between signal A and signal C, the smaller the possibility that the voice signals existing in signal A and signal C correspond to the same sound source. For example, the multiple first directional sound pickup signals further include signal D. In the case that the greater the first similarity between signal D and signal A, and the greater the first similarity between signal D and signal B, the electronic device can infer that the greater the possibility that the voice signals existing in signal A, signal B, and signal D correspond to the same sound source. For example, when the first similarity between at least two first directional sound pickup signals is greater than or equal to a first similarity threshold, it is determined that the at least two first directional sound pickup signals correspond to the same sound source; when the first similarity between at least two first directional sound pickup signals is less than the first similarity threshold, it is determined that the at least two first directional sound pickup signals do not correspond to the same sound source. The first similarity threshold can be pre-set or set by the user, which is not limited herein.
[0034] The electronic device can screen and cluster the multiple first directional sound pickup signals according to at least two of the first probability of the multiple first directional sound pickup signals, the first energy information of the multiple first directional sound pickup signals, and the first similarity between the multiple first directional sound pickup signals, to determine the first directional sound pickup signals satisfying the first screening condition as the second directional sound pickup signals. The first screening condition includes at least two of the first probability being greater than or equal to a first probability threshold, the first energy information being greater than or equal to a first energy threshold, and the first similarity being greater than or equal to a first similarity threshold.
[0035] In some example embodiments, the electronic device can perform voice signal detection on the first directional sound pickup signal first, and then perform at least one of energy information judgment and similarity judgment on the first directional sound pickup signal; the electronic device can also perform similarity judgment on the first directional sound pickup signal first, and then perform at least one of energy information judgment and voice signal detection on the first directional sound pickup signal; the electronic device can also perform energy information judgment on the first directional sound pickup signal first, and then perform at least one of voice signal detection and similarity judgment on the first directional sound pickup signal, without limitation.
[0036] Based on this, in the case of determining at least one second directional sound pickup signal from the plurality of first directional sound pickup signals, it can be determined that there is a voice signal in the second directional sound pickup signal, or the possibility of a voice signal in the second directional sound pickup signal is relatively large, and further the second directional sound pickup signal can be used to determine the target positioning information of the sound source corresponding to the voice signal in the preset audio frame signal.
[0037] Since the second directional sound pickup signal is determined by comprehensively determining at least two of the voice signal detection, the energy information judgment, and the similarity judgment on the first directional sound pickup signal, the second directional sound pickup signal can be used to distinguish the voice signal from the noise signal present in the preset audio frame signal, and can be used to determine whether the sound sources corresponding to the voice signals are the same sound source, and further to reduce noise interference or multiple sound source interference in the process of sound source positioning of the preset audio frame signal, so as to reduce the possibility of excluding the voice signal due to noise signal misjudgment or incorrectly identifying the noise signal as the voice signal. Therefore, it can be known that the determination of the second directional sound pickup signal is beneficial to subsequent improvement of the convenience and accuracy of the sound source positioning of the preset audio frame signal.
[0038] S104, determining first positioning information of a sound source corresponding to a voice signal in the preset audio frame signal according to the at least one second directional sound pickup signal.
[0039] For example, the first positioning information includes a first sound source angle of the sound source corresponding to the voice signal in the preset audio frame signal.
[0040] In some embodiments, the first scanning angle corresponding to the second directional sound pickup signal is determined as the first sound source angle of the sound source.
[0041] Since the first directional sound pickup signal corresponding to the second directional sound pickup signal meets at least two of the first probability being greater than or equal to the first probability threshold, the first energy information being greater than or equal to the first energy threshold, and the first similarity being greater than or equal to the first similarity threshold, the electronic device can infer that the first scanning angle corresponding to the first directional sound pickup signal can cover the sound source corresponding to the voice signal in the preset audio frame signal, and further the first scanning angle corresponding to the second directional sound pickup signal can be determined as the first sound source angle of the sound source.
[0042] Correspondingly, in a case where the plurality of second directional sound signals are determined, if the first similarity between the first directional sound signals corresponding to the plurality of second directional sound signals respectively is greater than or equal to the first similarity threshold, it can be determined that the plurality of second directional sound signals correspond to the same sound source, and then the first scanning angle corresponding to the plurality of second directional sound signals respectively is determined as the corresponding first scanning angle set, and the first scanning angle set is determined as the first sound source angle of the same sound source. If the first similarity between the first directional sound signals corresponding to the plurality of second directional sound signals respectively is less than the first similarity threshold, it can be determined that the plurality of second directional sound signals correspond to different sound sources, and then the first sound source angle of the different sound sources is determined according to the first scanning angle corresponding to the plurality of second directional sound signals respectively.
[0043] In a case where the first positioning information of the sound source corresponding to the speech signal in the preset audio frame signal is determined according to the at least one second directional sound signal, since the second directional sound signal can be used to distinguish the speech signal from the noise signal existing in the preset audio frame signal, and can be used to determine whether the sound source corresponding to the speech signal is the same sound source, the first positioning information of the sound source determined according to the at least one second directional sound signal can be used to reduce noise interference or multiple sound source interference in the process of positioning the sound source in the preset audio frame signal, thereby improving the sound source positioning accuracy of the preset audio frame signal.
[0044] In some embodiments, after the at least one second directional sound signal is determined from the plurality of first directional sound signals according to at least two of the first probability of the speech signal existing in the plurality of first directional sound signals, the first energy information of the plurality of first directional sound signals, and the first similarity between the plurality of first directional sound signals, the method further includes: performing directional sound picking on the second directional sound signal according to a second beam parameter to obtain a plurality of third directional sound signals corresponding to the second directional sound signal; the second beam parameter is different from the first beam parameter; determining a fourth directional sound signal from the at least one second directional sound signal according to at least two of a second probability of the speech signal existing in the plurality of third directional sound signals corresponding to the at least one second directional sound signal respectively, second energy information of the plurality of third directional sound signals, and second similarity between the plurality of third directional sound signals.
[0045] Determining the first positioning information of the sound source corresponding to the speech signal in the preset audio frame signal according to the at least one second directional sound signal includes: determining the second positioning information of the sound source corresponding to the speech signal in the preset audio frame signal according to the fourth directional sound signal.
[0046] In a case where the second directional pickup signal is acquired, the electronic device can perform directional pickup on the second directional pickup signal based on a fixed beam algorithm to achieve multi-level directional pickup on the preset audio frame signal in combination with directional pickup performed by the electronic device on the preset audio frame signal. In some embodiments, in a case where directional pickup is performed on the second directional pickup signal based on the fixed beam algorithm, directional pickup can be performed on the second directional pickup signal according to the second beam parameter to obtain a plurality of third directional pickup signals corresponding to the second directional pickup signal. The second beam parameter is different from the first beam parameter.
[0047] For example, the second beam parameter includes a second angle change step. The second angle change step can be used to determine a plurality of second scanning angles. The second angle change step is smaller than the first angle change step included in the first beam parameter. The second angle change step can be used to indicate the angle difference between adjacent second scanning angles. The description of the plurality of second scanning angles determined according to the second angle change step can refer to the description of the plurality of first scanning angles determined according to the first angle change step, which will not be repeated here. For example, in a case where the second angle change step is smaller than the first angle change step, the degree of overlap between every two adjacent second scanning angles is greater than the degree of overlap between every two adjacent first scanning angles, and the scanning density between every two adjacent second scanning angles is greater than the scanning density between every two adjacent first scanning angles, then compared with directional pickup on the preset audio frame signal according to the first beam parameter, the second beam parameter can be used to perform more refined directional pickup on the second directional pickup signal to refine and adjust the positioning information of the potential sound source, such as the first scanning angle corresponding to the second directional pickup signal, thereby improving the sound source positioning accuracy in the sound source positioning process of the preset audio frame signal.
[0048] In a case where the plurality of third directional pickup signals corresponding to the at least one second directional pickup signal are acquired, the electronic device can again evaluate the at least one second directional pickup signal to determine a fourth directional pickup signal from the at least one second directional pickup signal. The fourth directional pickup signal can be used to determine the second positioning information of the sound source corresponding to the voice signal in the preset audio frame signal.
[0049] For example, the electronic device can determine a second probability that a voice signal exists in each third directional pickup signal corresponding to each second directional pickup signal, to determine whether a voice signal exists in each third directional pickup signal corresponding to each second directional pickup signal. For example, based on a preset VAD model, voice activity detection is performed on the third directional pickup signal to obtain a second probability that a voice signal exists in the third directional pickup signal; when the second probability is greater than or equal to a second probability threshold, it is determined that a voice signal exists in the third directional pickup signal; when the second probability is less than the second probability threshold, it is determined that a voice signal does not exist in the third directional pickup signal. The second probability threshold can be preset or set by the user, which is not limited herein. In an exemplary embodiment, the second probability threshold is greater than the first probability threshold, so that the electronic device can select a fourth directional pickup signal with a higher possibility of existing voice information from at least one second directional pickup signal, thereby improving the sound source positioning accuracy of the preset audio frame signal.
[0050] The electronic device can determine second energy information of the plurality of third directional pickup signals corresponding to each second directional pickup signal, to perform energy judgment on the plurality of third directional pickup signals corresponding to each second directional pickup signal, measure the energy of each third directional pickup signal, and further assist in determining whether the sound source is concentrated in the second scanning angle corresponding to each third directional pickup signal. The relevant description of determining the second energy information of the plurality of third directional pickup signals corresponding to the second directional pickup signal can refer to the relevant description of determining the first energy information of the first directional pickup signal, which will not be repeated here. For the plurality of third directional pickup signals corresponding to each second directional pickup signal, the greater the second energy information of the third directional pickup signal, the greater the possibility that the sound source is concentrated in the second scanning angle corresponding to the third directional pickup signal; the smaller the second energy information of the third directional pickup signal, the smaller the possibility that the sound source is concentrated in the second scanning angle corresponding to the third directional pickup signal. For example, when the second energy information corresponding to the third directional pickup signal is greater than or equal to a second energy threshold, it is determined that the sound source is concentrated in the second scanning angle corresponding to the third directional pickup signal; when the second energy information corresponding to the third directional pickup signal is less than the second energy threshold, it is determined that the sound source is not concentrated in the second scanning angle corresponding to the third directional pickup signal. The second energy threshold can be preset or set by the user, which is not limited herein. In an exemplary embodiment, the second energy threshold is greater than the first energy threshold, so that the electronic device can select a fourth directional pickup signal with a higher possibility of being concentrated in the corresponding second scanning angle from at least one second directional pickup signal, thereby improving the sound source positioning accuracy of the preset audio frame signal.
[0051] Correspondingly, the electronic device can determine the second similarity between the plurality of third directional pickup signals corresponding to each second directional pickup signal, to determine whether the sound source is concentrated in the second scanning angle corresponding to the plurality of third directional pickup signals. For example, because the speech signal existing in the preset audio frame signal can correspond to different sound sources, that is, the preset audio frame signal includes speech signals corresponding to multiple sound sources respectively, in order to reduce the mutual interference of multiple sound sources in the process of sound source positioning of the preset audio frame signal, and further improve the sound source positioning accuracy of the preset audio frame signal, the electronic device can further determine the second similarity between the plurality of third directional pickup signals corresponding to the second directional pickup signal on the basis of the at least one second directional pickup signal, and further evaluate whether the speech signals existing in the plurality of third directional pickup signals correspond to the same sound source. Taking an example that the at least one second directional pickup signal includes signal A, and the plurality of third directional pickup signals corresponding to signal A include signal A1, signal A2 and signal A3. The electronic device can determine the second similarity between signal A1 and signal A2, signal A3 respectively. The greater the second similarity between signal A1 and signal A2, the greater the possibility that the speech signals existing in signal A1 and signal A2 correspond to the same sound source. Correspondingly, the smaller the second similarity between signal A1 and signal A3, the smaller the possibility that the speech signals existing in signal A1 and signal A3 correspond to the same sound source. For example, the plurality of third directional pickup signals further include signal A4, and the greater the second similarity between signal A4 and signal A1, and the greater the second similarity between signal A4 and signal A2, the greater the possibility that the speech signals existing in signal A1, signal A2 and signal A4 correspond to the same sound source. For each plurality of third directional pickup signals corresponding to each second directional pickup signal, the greater the second similarity between different third directional pickup signals, the greater the possibility that the speech signals existing in the different third directional pickup signals correspond to the same sound source; the smaller the second similarity between different third directional pickup signals, the smaller the possibility that the speech signals existing in the different third directional pickup signals correspond to the same sound source. For example, when the second similarity between at least two third directional pickup signals corresponding to the second directional pickup signal is greater than or equal to a second similarity threshold, it is determined that the at least two third directional pickup signals correspond to the same sound source; when the second similarity between at least two third directional pickup signals corresponding to the second directional pickup signal is less than the second similarity threshold, it is determined that the at least two third directional pickup signals do not correspond to the same sound source. The second similarity threshold can be pre-set or set by the user, which is not limited herein.In an example implementation, the second similarity threshold is greater than the first similarity threshold, so that the electronic device filters the second directional pickup signals to obtain the fourth directional pickup signal with a higher possibility of corresponding to the same sound source from the at least one second directional pickup signal, thereby improving the sound source positioning accuracy of the preset audio frame signal.
[0052] The electronic device can determine the second directional pickup signal satisfying the second filtering condition as the fourth directional pickup signal from the at least one second directional pickup signal according to at least two of the second probabilities of the plurality of third directional pickup signals corresponding to each of the second directional pickup signal, the second energy information of the plurality of third directional pickup signals, and the second similarity between the plurality of third directional pickup signals. The second filtering condition includes at least two of the second probability of each of the plurality of third directional pickup signals being greater than or equal to a second probability threshold, the second energy information of each of the plurality of third directional pickup signals being greater than or equal to a second energy threshold, and the second similarity between the plurality of third directional pickup signals being greater than or equal to a second similarity threshold.
[0053] For example, in a case where the second probability of each of the plurality of third directional pickup signals corresponding to the same second directional pickup signal is greater than or equal to the second probability threshold, the electronic device can determine that the voice signal exists in each of the plurality of third directional pickup signals corresponding to the second directional pickup signal. Since the second scanning angle corresponding to each of the plurality of third directional pickup signals is determined according to the first scanning angle corresponding to the second directional pickup signal, the electronic device can infer that the voice signal existing in the plurality of third directional pickup signals has a higher possibility of corresponding to the same sound source.
[0054] In a case where the second energy information of each of the plurality of third directional pickup signals corresponding to the same second directional pickup signal is greater than or equal to the second energy threshold, the electronic device can determine that the sound source corresponding to each of the plurality of third directional pickup signals concentrates on the second scanning angle corresponding to the corresponding third directional pickup signal. Since the second scanning angle corresponding to each of the plurality of third directional pickup signals is determined according to the first scanning angle corresponding to the second directional pickup signal, the electronic device can infer that the voice signal existing in the plurality of third directional pickup signals has a higher possibility of corresponding to the same sound source.
[0055] In a case where the second similarity between the plurality of third directional pickup signals corresponding to the same second directional pickup signal is greater than or equal to the second similarity threshold, the electronic device can determine that the voice signal existing in the plurality of third directional pickup signals corresponding to the second directional pickup signal has a higher possibility of corresponding to the same sound source.
[0056] Based on this, the electronic device can determine that the second directional pickup signal is the fourth directional pickup signal when the voice signal existing in the multiple directional pickup signals corresponding to the same second directional pickup signal is more likely to correspond to the same sound source, such as when at least two of the following conditions are met: the second probability of the second directional pickup signal corresponding to the multiple third directional pickup signals is greater than or equal to the second probability threshold, the second energy information of the second directional pickup signal corresponding to the multiple third directional pickup signals is greater than or equal to the second energy threshold, and the second similarity between the second directional pickup signals corresponding to the multiple third directional pickup signals is greater than or equal to the second similarity threshold.
[0057] In a case where the fourth directional pickup signal is determined, the electronic device can determine, according to the fourth directional pickup signal, second positioning information of the sound source corresponding to the voice signal in the preset audio frame signal.
[0058] For example, the second positioning information includes a second sound source angle of the sound source corresponding to the voice signal in the preset audio frame signal.
[0059] In some embodiments, the second scanning angle corresponding to the fourth directional pickup signal is determined as the second sound source angle of the sound source.
[0060] Since the multiple third directional pickup signals corresponding to the second directional pickup signal corresponding to the fourth directional pickup signal all meet at least two of the following conditions: the second probability is greater than or equal to the second probability threshold, the second energy information is greater than or equal to the second energy threshold, and the second similarity is greater than or equal to the second similarity threshold, the electronic device can infer that the second scanning angle corresponding to the fourth directional pickup signal can cover the sound source corresponding to the voice signal in the preset audio frame signal, and thus the second scanning angle corresponding to the fourth directional pickup signal can be determined as the second sound source angle of the sound source.
[0061] Correspondingly, in a case where multiple fourth directional pickup signals are determined, if the second similarity between the second directional pickup signals corresponding to the multiple fourth directional pickup signals is greater than or equal to the second similarity threshold, it can be determined that the multiple fourth directional pickup signals correspond to the same sound source, and thus the second scanning angle corresponding to the multiple fourth directional pickup signals is determined as a second scanning angle set, and the second scanning angle set is determined as the second sound source angle of the same sound source. If the second similarity between the second directional pickup signals corresponding to the multiple fourth directional pickup signals is less than the second similarity threshold, it can be determined that the multiple fourth directional pickup signals correspond to different sound sources, and thus the second sound source angle of the different sound sources is determined according to the second scanning angle corresponding to the multiple fourth directional pickup signals.
[0062] Since the second scanning angle corresponding to the fourth directional sound pickup signal is determined according to the second beam parameter, and the second beam parameter is different from the first beam parameter, the second scanning angle corresponding to the fourth directional sound pickup signal is actually obtained by angle fine-tuning the first scanning angle corresponding to the second directional sound pickup signal. The higher the degree of angle fine-tuning of the first scanning angle, the higher the accuracy of determining the second sound source angle of the sound source according to the second scanning angle, which is conducive to improving the sound source positioning accuracy of the preset audio frame signal.
[0063] Correspondingly, since the fourth directional sound pickup signal is determined by comprehensively determining at least two of the voice signal detection, the energy information judgment, and the similarity judgment of the second directional sound pickup signal, the fourth directional sound pickup signal can be used to distinguish the voice signal and the noise signal existing in the preset audio frame signal. Moreover, the fourth directional sound pickup signal is determined by comprehensively determining whether the sound sources corresponding to the voice signals existing in the plurality of third directional sound pickup signals corresponding to each of the at least one second directional sound pickup signal are the same sound source, which is conducive to reducing noise interference or multiple sound source interference in the process of sound source positioning of the preset audio frame signal, thereby reducing the possibility that the voice signal is excluded due to noise signal misjudgment or the noise signal is incorrectly identified as the voice signal, and further improving the sound source positioning accuracy of the preset audio frame signal.
[0064] In some embodiments, an initial audio frame signal is obtained; voice activity detection is performed on the initial audio frame signal based on a preset voice activity detection model to obtain a third probability that a voice signal exists in the initial audio frame signal; and when the third probability is greater than or equal to a third probability threshold, the initial audio frame signal is determined as the preset audio frame signal.
[0065] The initial audio frame signal is taken as a signal , and a voice activity detection (VAD) model is represented as For example. The electronic device can input the signal to the VAD model, so that the VAD model performs voice activity detection on the signal to obtain a third probability that a voice signal exists in the signal .
[0066] The process that the VAD model performs voice activity detection on the signal to obtain the third probability that a voice signal exists in the signal may be represented as:
[0067] , wherein is used to indicate the third probability that a voice signal exists in the signal .
[0068] In the process of determining the third probability In the case that the third probability is compared with a third probability threshold value to determine whether the voice signal exists in the signal.
[0069] When the third probability satisfies:
[0070] It can be determined that the voice signal exists in the signal , and further the signal can be determined as the preset audio frame signal for subsequent sound source positioning of the preset audio frame signal to obtain the first positioning information of the sound source corresponding to the voice signal in the preset audio frame signal, or to obtain the second positioning information of the sound source corresponding to the voice signal in the preset audio frame signal. The third probability threshold value can be pre-set or set by the user, which is not limited herein. In an exemplary embodiment, the third probability threshold value is less than or equal to the first probability threshold value, so that the electronic device can identify as many preset audio frame signals as possible from the initial audio frame signal, and further can be used for subsequent sound source positioning of the preset audio frame signal by the electronic device.
[0071] Exemplarily, the first beam parameter includes a first angle change step. The first angle change step can be used to preliminarily determine the approximate direction of the sound source corresponding to the voice signal in the preset audio frame signal. Moreover, the first angle change step affects the number of identifiable sound sources and the angle resolution. Of course, the first beam parameter is not limited to this, and the first beam parameter can also include the half-power beamwidth (HPBW) of the beam. The HPBW of the beam can be used to determine the width of a single beam, and the first angle change step can be used to determine the degree of overlap between adjacent beams and the scanning density.
[0072] In some embodiments, according to the preset angle corresponding to the preset audio frame signal and the first angle change step, a plurality of first scanning angles corresponding to the preset angle are determined; and according to the first scanning angle, directional sound pickup is performed on the preset audio frame signal to obtain a first directional sound pickup signal corresponding to the preset audio frame signal at the first scanning angle.
[0073] Taking the preset audio frame signal as the signal , the preset angle as the angle , and the first angle change step as the step for example.
[0074] According to the angle and the step , the angle corresponding to the first scan angle. The plurality of first scan angles can be uniformly represented as an angle , and the plurality of angles may constitute a first scan angle set. The first scan angle set can be represented as:
[0075] wherein i is used to indicate the ordinal number of the first scan angle, i.e., the i-th first scan angle.
[0076] The signal corresponding to the angle is the starting scan angle when the signal is directionally picked up according to the first beam parameter. The smaller the step value, the higher the angle resolution when the signal is directionally picked up, and correspondingly, the greater the calculation amount.
[0077] For example, different angles in the first scan angle set can correspond to different beam directions, and then the preset audio frame signal can be directionally picked up according to the first scan angle to obtain a first directionally picked up signal corresponding to the first scan angle. Correspondingly, for the plurality of first scan angles, the first directionally picked up signals corresponding to different first scan angles of the preset audio frame signal can be determined, i.e., the plurality of first directionally picked up signals corresponding to the preset audio frame signal can be determined.
[0078] In an exemplary embodiment, the angle may be set to -90° or 0°, but is not limited thereto.
[0079] In an exemplary embodiment, the step may be set to greater than 0° and less than or equal to 180°, or greater than 0° and less than or equal to 360°, but is not limited thereto.
[0080] In the case of determining the plurality of first directionally picked up signals corresponding to the preset audio frame signal, it is equivalent to realizing the first level directional picking of the preset audio frame signal, and the plurality of first directionally picked up signals can be used for subsequent determination of the first positioning information of the sound source corresponding to the voice signal in the preset audio frame signal, or determination of the second positioning information of the sound source corresponding to the voice signal in the preset audio frame signal.
[0081] In some embodiments, when there are at least two of the first directionally picked up signals that meet the first probability greater than or equal to the first probability threshold, the first energy greater than or equal to the first energy threshold, and the first similarity greater than or equal to the first similarity threshold, the first directionally picked up signal is determined as the second directionally picked up signal.
[0082] Based on the consideration of reducing noise interference or multiple sound source interference in the process of sound source positioning on the preset audio frame signal, the electronic device can screen and preliminarily cluster the plurality of first directional sound pickup signals, such as comparing the first directional sound pickup signals corresponding to the plurality of first scanning angles with each other, to screen out the first directional sound pickup signals with higher signal-to-noise ratio and stronger human sound source characteristics as the second directional sound pickup signals.
[0083] The plurality of first directional sound pickup signals constitute a set The first directional sound pickup signals in the set include signal and signal , and For example. Signal and signal may be used to indicate the first directional sound pickup signals under different first scanning angles.
[0084] The electronic device can calculate the first similarity between signal and signal , to make a similarity judgment on signal and signal , and calculate the first energy information of signal and signal respectively, to make an energy judgment on signal and signal . Accordingly, by comprehensively judging the similarity and energy of signal and signal , the electronic device can judge whether the sound source is concentrated in the first scanning angle corresponding to signal and signal respectively, and further evaluate the angle concentration of the sound source.
[0085] For example, the first similarity between signal and signal may include the Pearson correlation coefficient between signal and signal , without limitation. The calculation process of the first similarity between signal and signal may be represented as:
[0086] Wherein, is used to indicate the first similarity between signal and signal ; is used to indicate signal and signal a corresponding expected value, a mean value of the indication signal , a mean value of the indication signal ; a standard deviation of the indication signal , a standard deviation of the indication signal .
[0087] The first similarity between the signals and the signals is not limited to the Pearson correlation coefficient between the signals and the signals , and is not limited herein.
[0088] The first similarity between the signals and the signals can be used to measure the similarity between the signals and the signals to exclude irrelevant noise signals.
[0089] Exemplarily, the process of the electronic device calculating the first energy information of the first directional pickup signal can be represented as:
[0090] wherein, the first energy information of the first directional pickup signal is used to indicate. The first directional pickup signal may include the signals and the signals , and the first energy information of the signals may be represented as , and so on. The first energy information of the signals may be represented as .
[0091] For example, the first similarity is compared with the first similarity threshold , and the first energy information , the first energy information are compared with the first energy threshold , respectively, in the case of , and , it can be inferred that the sound sources corresponding to the speech signals in the preset audio frame signal are concentrated in the respective first scanning angles of the signals and the signals . The first energy information can be determined according to an average value or a variance of the first energy information of the first directional sound pickup signals corresponding to the plurality of first scanning angles, and is not limited to this. The first energy information can also be preset or set by a user, and is not limited in this regard.
[0092] Of course, the electronic device can also jointly determine the similarity and the energy of the plurality of first directional sound pickup signals to determine that the sound source corresponding to the voice signal in the preset audio frame signal is concentrated in the corresponding first scanning angle, and then determine the corresponding first scanning angle as a potential sound source candidate angle of the sound source, and determine the first directional sound pickup signal corresponding to the corresponding first scanning angle as a high-confidence directional sound pickup signal. The first scanning angle as the potential sound source candidate angle of the sound source can constitute an angle set Each first scanning angle in the angle set can be uniformly represented as an angle . Correspondingly, the first directional sound pickup signal corresponding to each first scanning angle in the angle set can constitute a signal set Each first directional sound pickup signal in the signal set can be uniformly represented as a signal , and the signal is the high-confidence directional sound pickup signal.
[0093] The signal obtained through screening and preliminary clustering can be further subjected to voice activity detection to determine whether a voice signal exists in the signal , and then ensure that the directional sound pickup signal involved in subsequent sound source positioning of the preset audio frame signal includes human voice characteristics.
[0094] For example, the signal can be input into a VAD model to obtain a first probability that a voice signal exists in the signal . Correspondingly, the first probability can be compared with a first probability threshold to determine whether a voice signal exists in the signal .
[0095] When the first probability satisfies:
[0096] It can be determined that a voice signal exists in the signal .
[0097] Based on this, it can be determined that the signal meets the condition that the first probability is greater than or equal to the first probability threshold ,Signal The first energy information is greater than or equal to the first energy threshold. ,Signal The first similarity is greater than or equal to the first similarity threshold. At least two of them, thus enabling the signal to Determine the second directional pickup signal.
[0098] The second directional pickup signal can be used to subsequently determine the first localization information of the sound source corresponding to the speech signal in the preset audio frame signal.
[0099] For example, the second beam parameters include at least a second angle variation step size and a first angle difference threshold; the second angle variation step size is smaller than the first angle variation step size included in the first beam parameters. The description related to the second angle variation step size can be referred to the foregoing description of the first angle variation step size, and will not be repeated here. Of course, the second beam parameters are not limited to this; the second beam parameters may also include the HPBW of the beam, which is not restricted here.
[0100] In some implementations, multiple second scanning angles corresponding to the first scanning angle are determined based on the first scanning angle corresponding to the second directional pickup signal and the second angle change step size; the absolute value of the angle difference between each second scanning angle and the first scanning angle is less than or equal to the first angle difference threshold; the second directional pickup signal is directionally picked up based on the second scanning angle to obtain a third directional pickup signal corresponding to the second directional pickup signal at the second scanning angle.
[0101] Using the second directional pickup signal as the signal ,Signal The corresponding first scanning angle is angle. The second angle change step size is the step size. The first angle difference threshold is The first angle change step size is the step size. For example.
[0102] In step length Smaller than step size In this case, electronic devices can target signals corresponding angle At the angle Set a smaller fine-tuning step size nearby, such as step size. And a fine-tuning range, such as the first angle difference threshold. This allows for the search of more precise second location information of the sound source corresponding to the speech signal in the preset audio frame signal.
[0103] According to the angle and step size The angle can be determined. corresponding to the plurality of second scanning angles. The plurality of second scanning angles can be uniformly represented as an angle , and the plurality of second scanning angles can constitute a second scanning angle set. The second scanning angle set can be represented as:
[0104] wherein n and m are integers, is used to indicate an angle difference between the angle and the angle , that is, an angle difference between the first scanning angle and the second scanning angle, and an absolute value of the angle difference is less than or equal to a first angle difference threshold .
[0105] For example, different second scanning angles in the second scanning angle set may correspond to different beam directions, and then the second directional sound pickup signal can be directionally sound picked according to the second scanning angle to obtain a third directional sound pickup signal corresponding to the second scanning angle. Accordingly, for the plurality of second scanning angles, the third directional sound pickup signal corresponding to each second scanning angle of the second directional sound pickup signal can be determined, that is, the plurality of third directional sound pickup signals corresponding to the second directional sound pickup signal can be determined.
[0106] Accordingly, in the case where there is at least one second directional sound pickup signal, the plurality of third directional sound pickup signals corresponding to each second directional sound pickup signal can be determined.
[0107] In the case where the plurality of third directional sound pickup signals corresponding to each second directional sound pickup signal is determined, it is equivalent to realizing the secondary directional sound pickup of the preset audio frame signal. Based on the determination of the first directional sound pickup signal and the determination of the third directional sound pickup signal, the multi-level directional sound pickup of the preset audio frame signal can be realized. The plurality of third directional sound pickup signals corresponding to each second directional sound pickup signal can be used for subsequent determination of the second positioning information of the sound source corresponding to the voice signal in the preset audio frame signal, so as to further refine the first positioning information of the sound source, and thus improve the sound source positioning accuracy of the preset audio frame signal.
[0108] In some embodiments, when at least two of the following conditions are met, the second directional sound pickup signal is determined to be the fourth directional sound pickup signal: the second probability of the second directional sound pickup signal corresponding to each of the plurality of third directional sound pickup signals is greater than or equal to a second probability threshold, the second energy information of the second directional sound pickup signal corresponding to each of the plurality of third directional sound pickup signals is greater than or equal to a second energy threshold, and the second similarity between the plurality of third directional sound pickup signals is greater than or equal to a second similarity threshold.
[0109] Based on the consideration of reducing noise interference or multiple sound source interference in the process of sound source positioning on the preset audio frame signal, the electronic device can perform at least two of voice activity detection, similarity judgment and energy judgment on each second directional sound pickup signal respectively corresponding to a plurality of third directional sound pickup signals, to further determine a fourth directional sound pickup signal. The fourth directional sound pickup signal is equivalent to the screening of at least one second directional sound pickup signal, which can improve the accuracy of sound source positioning on the preset audio frame signal.
[0110] The plurality of third directional sound pickup signals corresponding to the second directional sound pickup signal form a set , the third directional sound pickup signals in the set can be uniformly represented as For example.
[0111] For each third directional sound pickup signal in the set , a similarity judgment can be performed thereon to obtain a second similarity between each two third directional sound pickup signals in the set ; an energy judgment can also be performed thereon to obtain second energy information of each third directional sound pickup signal ; and a voice activity detection can also be performed thereon to obtain a second probability of each third directional sound pickup signal .
[0112] In the case where each third directional sound pickup signal in the set meets at least two of the following conditions: the second probability of each third directional sound pickup signal is greater than or equal to a second probability threshold , the second energy information of each third directional sound pickup signal is greater than or equal to a second energy threshold , and the second similarity between each two third directional sound pickup signals is greater than or equal to a second similarity threshold , the electronic device can determine that the plurality of third directional sound pickup signals in the set sustain high similarity, high energy and high VAD probability, and further determine that the second directional sound pickup signal corresponding to the set is the fourth directional sound pickup signal.
[0113] Of course, it is not limited to this. When the second probability that the second directional pickup signal meets at least one third directional pickup signal is less than the second probability threshold, the second energy information of the at least one third directional pickup signal is less than the second energy threshold, and the second similarity between the at least two third directional pickup signals is less than the second similarity threshold, the electronic device cannot determine that the second directional pickup signal is the fourth directional pickup signal. Based on this, the electronic device can continue to determine, from the plurality of third directional pickup signals corresponding to the second directional pickup signal, at least two third directional pickup signals that meet at least two of the following conditions: the second probability is greater than or equal to the second probability threshold, the second energy information is greater than or equal to the second energy threshold, and the second similarity is greater than or equal to the second similarity threshold, as fifth directional pickup signals. Accordingly, the electronic device can continue to perform directional pickup on the fifth directional pickup signal according to the third beam parameter to obtain a plurality of sixth directional pickup signals corresponding to the fifth directional pickup signal. The third beam parameter is different from the second beam parameter. For example, the third beam parameter includes a third angle change step and a second angle difference threshold. The third angle change step is less than or equal to the second scanning angle change step, and the second angle difference threshold is less than or equal to the first angle difference threshold. Of course, it is not limited to this, and is not limited herein. The electronic device can determine the fifth directional pickup signal as a seventh directional pickup signal when at least two of the following conditions exist: the fifth directional pickup signal meets the respective third probability that the corresponding plurality of sixth directional pickup signals each exist a voice signal is greater than or equal to a third probability threshold, the respective third energy information of the corresponding plurality of sixth directional pickup signals is greater than or equal to a third energy threshold, and the third similarity between the corresponding plurality of fifth directional pickup signals is greater than or equal to a third similarity threshold. The seventh directional pickup signal can be used by the electronic device to determine the third positioning information of the sound source corresponding to the voice signal in the preset audio frame signal. By analogy, the directional pickup signals corresponding to the voice signal in the preset audio frame signal, such as the second directional pickup signal, the fourth directional pickup signal, the seventh directional pickup signal, and the like, are iteratively optimized and verified to determine the final directional pickup signal, so as to determine the positioning information of the sound source corresponding to the voice signal in the preset audio frame signal according to the final directional pickup signal, such as one of the first positioning information, the second positioning information, the third positioning information, and the like of the sound source corresponding to the voice signal in the preset audio frame signal.
[0114] The fourth directional pickup signal can be used to determine the second positioning information of the sound source corresponding to the voice signal in the preset audio frame signal.
[0115] For example, the second positioning information includes a second sound source angle of the sound source corresponding to the voice signal in the preset audio frame signal. The second scanning angle corresponding to the fourth directional pickup signal is determined as the second sound source angle of the sound source, and the second positioning information of the sound source can be determined.
[0116] Since the second scanning angle corresponding to the fourth directional pickup signal is determined according to the second beam parameter, and the second beam parameter is different from the first beam parameter, the second scanning angle corresponding to the fourth directional pickup signal is equivalent to the first scanning angle corresponding to the second directional pickup signal corresponding to the fourth directional pickup signal being angle-tuned, thereby facilitating improvement of the sound source positioning accuracy of the preset audio frame signal.
[0117] The sound source positioning method provided by the above embodiments comprises the following steps: obtaining a preset audio frame signal, wherein the preset audio frame signal contains a speech signal; performing directional pickup on the preset audio frame signal according to a first beam parameter to obtain a plurality of first directional pickup signals corresponding to the preset audio frame signal; determining at least one second directional pickup signal from the plurality of first directional pickup signals according to at least two of the following: a first probability of the speech signal existing in the plurality of first directional pickup signals, first energy information of the plurality of first directional pickup signals, and a first similarity between the plurality of first directional pickup signals; and determining first positioning information of a sound source corresponding to the speech signal in the preset audio frame signal according to the at least one second directional pickup signal.
[0118] In the case of determining the second directional pickup signal by comprehensively considering at least two of the first probability, the first energy information, and the first similarity, the second directional pickup signal is equivalent to being determined by comprehensively considering at least two of the following: speech signal detection, energy information judgment, and similarity judgment of the preset audio frame signal. The second directional pickup signal can be used to distinguish the speech signal from the noise signal in the preset audio frame signal, and can be used to determine whether the sound sources corresponding to the speech signals are the same sound source, thereby reducing noise interference or multiple sound source interference in the process of positioning the sound source of the preset audio frame signal. Based on the reduction of noise interference or multiple sound source interference in the process of positioning the sound source of the preset audio frame signal, the sound source positioning accuracy of the preset audio frame signal is improved.
[0119] Please refer to Figure 2 , Figure 2is a schematic block diagram of a sound source positioning device provided by an embodiment of the present application. The sound source positioning device can be configured in an electronic device or a server, and is used to execute the sound source positioning method described above. The electronic device can include a near-eye display device, a wearable device, a terminal device, etc., without limitation. The near-eye display device can include AR glasses, VR glasses, MR glasses, an AR helmet, a VR helmet, an MR helmet, etc., without limitation. The wearable device includes a smart watch, a smart ring, a smart bracelet, etc., without limitation. The terminal device includes a mobile phone, a television, a computer, etc., without limitation. The server can be a separate server, or a cloud server providing cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content distribution networks, and basic cloud computing services such as big data and artificial intelligence platforms.
[0120] As shown in Figure 2 , the sound source positioning device includes an audio acquisition module 110, a first signal determination module 120, a second signal determination module 130, and a sound source positioning module 140.
[0121] The audio acquisition module 110 is configured to acquire a preset audio frame signal, wherein the preset audio frame signal contains a speech signal. The first signal determination module 120 is configured to perform directional sound pickup on the preset audio frame signal according to a first beam parameter to obtain a plurality of first directional sound pickup signals corresponding to the preset audio frame signal. The second signal determination module 130 is configured to determine at least one second directional sound pickup signal from the plurality of first directional sound pickup signals according to at least two of a first probability that the speech signal exists in the plurality of first directional sound pickup signals, first energy information of the plurality of first directional sound pickup signals, and a first similarity between the plurality of first directional sound pickup signals. The sound source positioning module 140 is configured to determine first positioning information of a sound source corresponding to the speech signal in the preset audio frame signal according to the at least one second directional sound pickup signal.
[0122] For example, the first beam parameter includes a first angle change step; and the first signal determination module 120 includes a first angle determination submodule and a first directional sound pickup submodule.
[0123] The first angle determination submodule is configured to determine a plurality of first scanning angles corresponding to a preset angle of the preset audio frame signal according to the preset angle and the first angle change step. The first directional sound pickup sub-module is configured to perform directional sound pickup on the preset audio frame signal according to the first scanning angle to obtain a first directional sound pickup signal corresponding to the first scanning angle of the preset audio frame signal.
[0124] In an example, the second signal determination module 130 includes a first clustering sub-module.
[0125] The first clustering sub-module is configured to determine the first directional sound pickup signal as a second directional sound pickup signal when at least two of the following conditions are met: the first directional sound pickup signal meets the first probability threshold, the first energy information is greater than or equal to the first energy threshold, and the first similarity is greater than or equal to the first similarity threshold.
[0126] In an example, the sound source positioning device further includes a third signal determination sub-module and a fourth signal determination sub-module.
[0127] The third signal determination sub-module is configured to perform directional sound pickup on the second directional sound pickup signal according to a second beam parameter to obtain a plurality of third directional sound pickup signals corresponding to the second directional sound pickup signal; the second beam parameter is different from the first beam parameter. The fourth signal determination sub-module is configured to determine a fourth directional sound pickup signal from at least one of the second directional sound pickup signals according to at least two of the following: a second probability of a speech signal existing in the plurality of third directional sound pickup signals corresponding to each of the second directional sound pickup signals, second energy information of the plurality of third directional sound pickup signals, and a second similarity between the plurality of third directional sound pickup signals.
[0128] The sound source positioning module 140 includes a sound source positioning sub-module.
[0129] The sound source positioning sub-module is configured to determine second positioning information of a sound source corresponding to a speech signal in the preset audio frame signal according to the fourth directional sound pickup signal.
[0130] In an example, the fourth signal determination sub-module includes a second clustering sub-module.
[0131] The second clustering sub-module is configured to determine the second directional sound pickup signal as a fourth directional sound pickup signal when at least two of the following conditions are met: the second probability corresponding to each of the plurality of third directional sound pickup signals is greater than or equal to a second probability threshold, the second energy information corresponding to each of the plurality of third directional sound pickup signals is greater than or equal to a second energy threshold, and the second similarity between the plurality of third directional sound pickup signals is greater than or equal to a second similarity threshold.
[0132] Exemplarily, the second beam parameter comprises at least a second angle change step and a first angle difference threshold; the second angle change step is smaller than a first angle change step comprised in the first beam parameter; the third signal determination submodule comprises a second angle determination submodule and a second directional sound pickup submodule.
[0133] The second angle determination submodule is configured to determine a plurality of second scanning angles corresponding to the first scanning angle according to the first scanning angle corresponding to the second directional sound pickup signal and the second angle change step; an absolute value of an angle difference between each of the second scanning angles and the first scanning angle is smaller than or equal to the first angle difference threshold. The second directional sound pickup submodule is configured to perform directional sound pickup on the second directional sound pickup signal according to the second scanning angle, to obtain a third directional sound pickup signal corresponding to the second scanning angle.
[0134] Exemplarily, the audio acquisition module 110 comprises an initial signal acquisition submodule, a voice activity detection submodule and a preset signal determination submodule.
[0135] The initial signal acquisition submodule is configured to acquire an initial audio frame signal. The voice activity detection submodule is configured to perform voice activity detection on the initial audio frame signal based on a preset voice activity detection model, to obtain a third probability that a voice signal exists in the initial audio frame signal. The preset signal determination submodule is configured to determine the initial audio frame signal as a preset audio frame signal when the third probability is greater than or equal to a third probability threshold.
[0136] It should be noted that, for the convenience and brevity of description, the specific working processes of the above-described apparatuses and modules and units can refer to the corresponding processes in the foregoing method embodiments, which will not be described herein again.
[0137] The methodologies of the present application can be employed in a variety of computer system contexts. For example, the methodologies can be employed in a personal computer, a server computer, a handheld device or portable device, a tablet device, a multiprocessor system, a microprocessor-based system, a set top box, programmable consumer electronics, network PC, minicomputer, mainframe computer, distributed computing environments that include any of the above systems or devices, or the like. The present application can be described in the general context of computer-executable instructions, such as program modules, being executed by a computer. Generally, program modules include routines, programs, objects, components, data structures, and the like, that perform particular tasks or implement particular abstract data types. The present application can also be practiced in distributed computing environments where tasks are performed by remote processing devices that are linked through a communications network. In a distributed computing environment, program modules can be located in local and remote computer storage media including memory storage devices.
[0138] Exemplarily, the above method and device can be implemented in the form of a computer program, which can be run on an electronic device or a server to control the electronic device and thus perform sound source positioning. Exemplarily, the electronic device can include a near-eye display device, a wearable device, a terminal device, and the like, which are not limited herein. The near-eye display device can include AR glasses, VR glasses, MR glasses, an AR helmet, a VR helmet, an MR helmet, and the like, which are not limited herein. The wearable device includes a smart watch, a smart ring, a smart bracelet, and the like, which are not limited herein. The terminal device includes a mobile phone, a television, a computer, and the like, which are not limited herein. The server can be a separate server, or a cloud server providing cloud services, a cloud database, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content distribution networks, and basic cloud computing services such as big data and artificial intelligence platforms.
[0139] Please refer to Figure 3 , Figure 3 is a structural schematic block diagram of an electronic device provided by an embodiment of the present application.
[0140] As Figure 3 indicated, the electronic device includes a memory and a processor. The memory and the processor can be connected through a system bus. The memory can include a storage medium and an internal memory.
[0141] The storage medium can store an operating system and a computer program. The computer program, when executed, can enable the processor to perform any sound source positioning method.
[0142] The processor is configured to provide computing and control capabilities to support the operation of the entire electronic device.
[0143] The internal memory provides an environment for the running of a computer program in a storage medium, and the computer program, when executed by the processor, can enable the processor to perform any sound source positioning method.
[0144] Those skilled in the art can understand that, Figure 3 The structure shown in the figure is only a block diagram of part of the structure related to the scheme of the present application, and does not constitute a limitation on the near-eye display device to which the scheme of the present application is applied. A specific near-eye display device can include more or fewer components than those shown in the figure, or combine certain components, or have a different component arrangement.
[0145] It should be understood that the processor can be a central processing unit (CPU), and the processor can also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor.
[0146] In one embodiment, the processor is configured to execute the computer program and implement the following steps when executing the computer program: obtain a preset audio frame signal, the preset audio frame signal containing a speech signal; directively pick up the preset audio frame signal according to the first beam parameter to obtain a plurality of first directivity picked signals corresponding to the preset audio frame signal; determine at least one second directivity picked signal from the plurality of first directivity picked signals according to at least two of a first probability of the speech signal existing in the plurality of first directivity picked signals, first energy information of the plurality of first directivity picked signals, and a first similarity between the plurality of first directivity picked signals; determine first positioning information of a sound source corresponding to the speech signal in the preset audio frame signal according to the at least one second directivity picked signal.
[0147] It should be noted that those skilled in the art can clearly understand that, for the convenience and brevity of description, the above description of the specific working process of sound source positioning can refer to the corresponding process in the foregoing sound source positioning method embodiments, which will not be described herein.
[0148] The embodiment of the present application further provides a computer readable storage medium, and the computer readable storage medium stores a computer program.
[0149] The computer readable storage medium can be an internal storage unit of the electronic device, for example, a hard disk or a memory of the electronic device. The computer readable storage medium can also be an external storage device of the electronic device, for example, a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc.
[0150] It should be understood that the terms used herein in the specification and the appended claims are merely used for the purpose of describing particular embodiments and are not intended to limit the present application. As used in this specification and the appended claims, the singular forms "a," "an," and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise.
[0151] It should also be understood that the term "and / or" as used herein refers to any or all possible combinations of one or more of the associated listed items, and all possible combinations thereof. It should be noted that the terms "comprising," "including," or any other variant thereof are intended to cover non-exclusive inclusions, such that processes, methods, articles, or systems that include a series of elements are not limited to those elements, but can also include other elements not expressly listed, or other elements inherent to such processes, methods, articles, or systems. Without more limitations, the element defined by the phrase "comprising a" does not exclude the presence of additional identical elements in the process, method, article, or system that includes the element.
[0152] The above-mentioned serial numbers of the embodiments of the present application are only for description, and do not represent the advantages or disadvantages of the embodiments. The above description is only a specific implementation of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art can easily think of various equivalent modifications or replacements within the technical scope disclosed by the present application, and these modifications or replacements should be covered within the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.
Claims
1. A method of acoustic source localization, the method comprising: The method comprises: acquiring a preset audio frame signal, wherein the preset audio frame signal contains a speech signal; directed sound pickup is performed on the preset audio frame signal according to a first beam parameter, to obtain a plurality of first directed sound pickup signals corresponding to the preset audio frame signal; at least one second directed sound pickup signal is determined from the plurality of first directed sound pickup signals according to at least two of a first probability that the speech signal exists in the plurality of first directed sound pickup signals, first energy information of the plurality of first directed sound pickup signals, and a first similarity between the plurality of first directed sound pickup signals; first positioning information of a sound source corresponding to the speech signal in the preset audio frame signal is determined according to the at least one second directed sound pickup signal.
2. The acoustic source positioning method of claim 1, wherein, The first beam parameter comprises a first angle change step; The directed sound pickup performed on the preset audio frame signal according to the first beam parameter to obtain the plurality of first directed sound pickup signals corresponding to the preset audio frame signal comprises: a plurality of first scanning angles corresponding to a preset angle of the preset audio frame signal are determined according to the preset angle and the first angle change step; directed sound pickup is performed on the preset audio frame signal according to the first scanning angle, to obtain a first directed sound pickup signal corresponding to the preset audio frame signal at the first scanning angle.
3. The acoustic source localization method of claim 2, wherein, The determination of the at least one second directed sound pickup signal from the plurality of first directed sound pickup signals according to at least two of the first probability that the speech signal exists in the plurality of first directed sound pickup signals, the first energy information of the plurality of first directed sound pickup signals, and the first similarity between the plurality of first directed sound pickup signals comprises: when there is at least one first directed sound pickup signal that meets at least two of the first probability being greater than or equal to a first probability threshold value, the first energy information being greater than or equal to a first energy threshold value, and the first similarity being greater than or equal to a first similarity threshold value, the first directed sound pickup signal is determined as a second directed sound pickup signal.
4. The acoustic source positioning method according to any one of claims 1 to 3, characterized in that, After the determination of the at least one second directed sound pickup signal from the plurality of first directed sound pickup signals according to at least two of the first probability that the speech signal exists in the plurality of first directed sound pickup signals, the first energy information of the plurality of first directed sound pickup signals, and the first similarity between the plurality of first directed sound pickup signals, the method further comprises: directed sound pickup is performed on the second directed sound pickup signal according to a second beam parameter, to obtain a plurality of third directed sound pickup signals corresponding to the second directed sound pickup signal; the second beam parameter is different from the first beam parameter; a fourth directed sound pickup signal is determined from the at least one second directed sound pickup signal according to at least two of a second probability that the speech signal exists in a plurality of third directed sound pickup signals corresponding to each of the at least one second directed sound pickup signal, second energy information of the plurality of third directed sound pickup signals, and a second similarity between the plurality of third directed sound pickup signals; The determination of the first positioning information of the sound source corresponding to the speech signal in the preset audio frame signal according to the at least one second directed sound pickup signal comprises: According to the fourth directional sound pickup signal, second positioning information of a sound source corresponding to a speech signal in the preset audio frame signal is determined.
5. The acoustic source localization method of claim 4, wherein, The fourth directional sound pickup signal is determined from the at least one second directional sound pickup signal according to at least two of a second probability that a speech signal exists in each of a plurality of third directional sound pickup signals corresponding to the second directional sound pickup signal, second energy information of the plurality of third directional sound pickup signals, and a second similarity between the plurality of third directional sound pickup signals, and the method comprises: When at least two of the second probability that a speech signal exists in each of the plurality of third directional sound pickup signals corresponding to the second directional sound pickup signal is greater than or equal to a second probability threshold value, the second energy information of the plurality of third directional sound pickup signals corresponding to the second directional sound pickup signal is greater than or equal to a second energy threshold value, and the second similarity between the plurality of third directional sound pickup signals corresponding to the second directional sound pickup signal is greater than or equal to a second similarity threshold value, the second directional sound pickup signal is determined to be the fourth directional sound pickup signal.
6. The acoustic source localization method of claim 4, wherein, The second beam parameter comprises at least a second angle change step and a first angle difference threshold value; the second angle change step is smaller than a first angle change step comprised in the first beam parameter; The second directional sound pickup signal is directionally sound picked according to the second beam parameter, to obtain a plurality of third directional sound pickup signals corresponding to the second directional sound pickup signal, and the method comprises: According to the first scanning angle corresponding to the second directional sound pickup signal and the second angle change step, a plurality of second scanning angles corresponding to the first scanning angle are determined; an absolute value of an angle difference between each second scanning angle and the first scanning angle is less than or equal to the first angle difference threshold value; The second directional sound pickup signal is directionally sound picked according to the second scanning angle, to obtain a third directional sound pickup signal corresponding to the second scanning angle of the second directional sound pickup signal.
7. The acoustic source positioning method according to any one of claims 1 to 3, characterized in that, The preset audio frame signal is obtained, and the method comprises: An initial audio frame signal is obtained; A speech activity detection is performed on the initial audio frame signal based on a preset speech activity detection model, to obtain a third probability that a speech signal exists in the initial audio frame signal; When the third probability is greater than or equal to a third probability threshold value, the initial audio frame signal is determined to be the preset audio frame signal.
8. A sound source positioning apparatus characterized by comprising: The sound source positioning device comprises: An audio acquisition module is configured to obtain a preset audio frame signal in which a speech signal exists; A first signal determination module is configured to directionally sound pick the preset audio frame signal according to a first beam parameter, to obtain a plurality of first directional sound pickup signals corresponding to the preset audio frame signal; A second signal determination module is configured to determine at least one second directional sound pickup signal from the plurality of first directional sound pickup signals according to at least two of a first probability that a speech signal exists in each of the plurality of first directional sound pickup signals, first energy information of the plurality of first directional sound pickup signals, and a first similarity between the plurality of first directional sound pickup signals; A sound source positioning module is configured to determine first positioning information of a sound source corresponding to a speech signal in the preset audio frame signal according to at least one second directional sound pickup signal.
9. An electronic device, comprising: The electronic device comprises a memory and a processor; The memory is configured to store a computer program; The processor is configured to execute the computer program and implement the steps of the sound source positioning method according to any one of claims 1 to 7 when executing the computer program.
10. A computer-readable storage medium having stored thereon a computer program, characterized in that, The computer program, when executed by the processor, implements the steps of the sound source positioning method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Audio processing method, audio processing device, system and medium
CN111833901A
Reception process recording method, related device, server, system and storage medium
CN116320886A
Sound source positioning method and device, medium and equipment
CN117409813A
Multi-microphone array beamforming signal enhancement method and device
CN119811408A
Sound source positioning method and apparatus, and electronic device
WO2022135131A1