Sound source positioning method, device, equipment and storage medium

By performing directional sound pickup and signal filtering in audio frame signals, and combining the probability, energy, and similarity of speech signals, the problems of noise and multi-source interference are solved, improving the accuracy and convenience of sound source localization.

CN120908752BActive Publication Date: 2026-01-02ZHUHAI MOJIE TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511456349.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-10-13
Publication Date
2026-01-02
Estimated Expiration
2045-10-13

AI Technical Summary

Technical Problem

In the prior art, noise or interference from multiple sound sources in the audio frame signal leads to poor accuracy in sound source localization.

Method used

By acquiring a preset audio frame signal, directional sound pickup is performed using the first beam parameters. Combining the probability, energy information, and similarity of the speech signal, a second directional sound pickup signal is determined from multiple directional sound pickup signals, thereby determining the location information of the sound source.

Benefits of technology

It reduces noise and interference from multiple sound sources, improving the accuracy and convenience of sound source localization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120908752B_ABST
    Figure CN120908752B_ABST
Patent Text Reader

Abstract

The application relates to the field of sound source positioning, and provides a sound source positioning method, device and equipment and a storage medium. The sound source positioning method comprises the following steps: acquiring a preset audio frame signal, wherein the preset audio frame signal contains a speech signal; performing directional sound pickup on the preset audio frame signal according to first beam parameters to obtain a plurality of first directional sound pickup signals corresponding to the preset audio frame signal; determining at least one second directional sound pickup signal from the plurality of first directional sound pickup signals according to at least two of the following: a first probability that the speech signal exists in the plurality of first directional sound pickup signals, first energy information of the plurality of first directional sound pickup signals and a first similarity between the plurality of first directional sound pickup signals; and determining target positioning information of a sound source corresponding to the speech signal in the preset audio frame signal according to the at least one second directional sound pickup signal, so as to improve the sound source positioning accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of sound source positioning, and particularly relates to a sound source positioning method and device, equipment and a storage medium. BACKGROUND

[0002] In the related art, an electronic device can perform sound source positioning on a corresponding audio frame signal to determine positioning information of a sound source. However, in the case that noise exists in the audio frame signal or multiple sound sources exist in the audio frame signal, the noise or multiple sound sources existing in the audio frame signal will interfere with the sound source positioning process of the audio frame signal, and thus lead to poor sound source positioning accuracy of the audio frame signal. Therefore, there is an urgent need to improve the sound source positioning accuracy of the audio frame signal. SUMMARY

[0003] The main purpose of the present application is to provide a sound source positioning method, device, equipment and storage medium, which aims to solve the technical problem of poor sound source positioning accuracy of an audio frame signal due to the sound source positioning process of the audio frame signal being easily interfered by noise or multiple sound sources.

[0004] In a first aspect, the present application provides a sound source positioning method, comprising:

[0005] obtaining a preset audio frame signal, wherein a speech signal exists in the preset audio frame signal;

[0006] directly picking up the preset audio frame signal according to a first beam parameter to obtain a plurality of first direct pickup signals corresponding to the preset audio frame signal;

[0007] determining at least one second direct pickup signal from the plurality of first direct pickup signals according to at least two of a first probability that a speech signal exists in the plurality of first direct pickup signals, first energy information of the plurality of first direct pickup signals, and a first similarity between the plurality of first direct pickup signals;

[0008] determining first positioning information of a sound source corresponding to the speech signal in the preset audio frame signal according to the at least one second direct pickup signal.

[0009] In a second aspect, the present application provides a sound source positioning device, comprising:

[0010] an audio acquisition module configured to obtain a preset audio frame signal, wherein a speech signal exists in the preset audio frame signal;

[0011] a first signal determination module configured to directly pick up the preset audio frame signal according to a first beam parameter to obtain a plurality of first direct pickup signals corresponding to the preset audio frame signal;

[0012] The second signal determination module is configured to determine at least one second directional sound pickup signal from the plurality of first directional sound pickup signals according to at least two of a first probability of the presence of a speech signal in the plurality of first directional sound pickup signals, first energy information of the plurality of first directional sound pickup signals, and a first similarity between the plurality of first directional sound pickup signals.

[0013] The sound source positioning module is configured to determine first positioning information of a sound source corresponding to the speech signal in the preset audio frame signal according to the at least one second directional sound pickup signal.

[0014] In a third aspect, the present application provides an electronic device, which comprises a memory and a processor.

[0015] The memory is configured to store a computer program.

[0016] The processor is configured to execute the computer program and implement the steps of the sound source positioning method as described above when executing the computer program.

[0017] In a fourth aspect, the present application provides a computer readable storage medium, which stores a computer program. The computer program is executed by a processor to implement the steps of the sound source positioning method as described above.

[0018] The present application provides a sound source positioning method, device, equipment and storage medium. The sound source positioning method comprises: obtaining a preset audio frame signal, and the preset audio frame signal contains a speech signal; performing directional sound pickup on the preset audio frame signal according to first beam parameters to obtain a plurality of first directional sound pickup signals corresponding to the preset audio frame signal; determining at least one second directional sound pickup signal from the plurality of first directional sound pickup signals according to at least two of a first probability of the presence of a speech signal in the plurality of first directional sound pickup signals, first energy information of the plurality of first directional sound pickup signals, and a first similarity between the plurality of first directional sound pickup signals; and determining first positioning information of a sound source corresponding to the speech signal in the preset audio frame signal according to the at least one second directional sound pickup signal.

[0019] In the case that the second directional sound pickup signal is determined by integrating at least two of the first probability, the first energy information, and the first similarity, the second directional sound pickup signal is equivalent to being determined by integrating at least two of the speech signal detection, the energy information judgment, and the similarity judgment on the preset audio frame signal. The second directional sound pickup signal can be used to distinguish the speech signal and the noise signal existing in the preset audio frame signal, and can be used to determine whether the sound sources corresponding to the speech signals are the same sound source, thereby reducing noise interference or multiple sound source interference in the process of sound source positioning on the preset audio frame signal. Based on the reduction of noise interference or multiple sound source interference in the process of sound source positioning on the preset audio frame signal, the sound source positioning accuracy of the preset audio frame signal is improved. BRIEF DESCRIPTION OF DRAWINGS

[0020] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings needed in the embodiment description will be briefly introduced. Obviously, the drawings in the following description are some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.

[0021] Figure 1 is a flowchart of a sound source positioning method provided by an embodiment of the present application;

[0022] Figure 2 is a schematic block diagram of a sound source positioning device provided by an embodiment of the present application;

[0023] Figure 3 is a schematic block diagram of an electronic device provided by an embodiment of the present application. DETAILED DESCRIPTION

[0024] The technical solutions in the embodiments of the present application will be described clearly and completely in combination with the drawings in the embodiments of the present application. Obviously, the described embodiments are some embodiments of the present application, not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor are within the scope of protection of the present application.

[0025] The flowchart shown in the drawings is only an example and does not necessarily include all contents and operations / steps, nor does it necessarily be executed in the described order. For example, some operations / steps can be decomposed, combined or partially combined, so the actual execution order may be changed according to the actual situation.

[0026] Embodiments of the present application provide a sound source positioning method, device, equipment and storage medium. The sound source positioning method can be applied to an electronic device. The electronic device can include a near-eye display device, a wearable device, a terminal device, etc., without limitation. The near-eye display device can include an augmented reality (AR) glasses, a virtual reality (VR) glasses, a mixed reality (MR) glasses, an AR helmet, a VR helmet, a MR helmet, etc., without limitation. The wearable device includes a smart watch, a smart ring, a smart bracelet, etc., without limitation. The terminal device includes a mobile phone, a television, a computer, etc., without limitation. The sound source positioning method can also be applied to a server, which can be a separate server, or a cloud server providing cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery network (CDN), and basic cloud computing services such as big data and artificial intelligence platforms.

[0027] Some embodiments of the present application will be described in detail below with reference to the accompanying drawings. The following examples and features in the examples can be combined with each other without conflict.

[0028] Please refer to Figure 1 , Figure 1 is a flowchart of a sound source positioning method provided by an embodiment of the present application. It should be noted that the sound source positioning method provided by the embodiments of the present application can be used in an electronic device, or can be used in a server, without limitation.

[0029] As Figure 1 indicated, the sound source positioning method includes steps S101 to S104.

[0030] S101, obtaining a preset audio frame signal, the preset audio frame signal including a speech signal.

[0031] Exemplarily, the preset audio frame signal is used to indicate an audio frame signal capable of sound source positioning processing. The speech signal is used to indicate a signal with human voice characteristics. For example, the speech signal includes at least one of biological human voice and non-biological human voice. The biological human voice can be understood as a sound emitted by a human himself. The non-biological human voice can be understood as a human voice played by an electronic device. For example, the non-biological voice can include a human voice played by a sound box, human voice music, etc., without limitation.

[0032] In some embodiments, the electronic device can obtain an initial audio frame signal. Accordingly, the electronic device can detect whether a speech signal exists in the initial audio frame signal. If a speech signal exists in the initial audio frame signal, the electronic device can determine that the initial audio frame signal can be used for subsequent sound source positioning, and can then determine the initial audio frame signal as a preset audio frame signal. Accordingly, if a speech signal does not exist in the initial audio frame signal, the electronic device can determine that the initial audio frame signal does not need to be positioned, and then does not need to be determined as a preset audio frame signal.

[0033] In the case where the preset audio frame signal is obtained and a speech signal exists in the preset audio frame signal, the preset audio frame signal is used to determine the target positioning information of the sound source corresponding to the speech signal.

[0034] S102, according to the first beam parameter, the preset audio frame signal is directionally picked up to obtain a plurality of first directional pickup signals corresponding to the preset audio frame signal.

[0035] In the case where the preset audio frame signal is obtained, the preset audio frame signal can also include noise signals other than speech signals or a plurality of sound sources each corresponding to a speech signal. In order to reduce the noise signal interference in the preset audio frame signal and the adverse effects of mutual interference of a plurality of sound sources on the positioning of the preset audio frame signal, the electronic device can directionally pick up the preset audio frame signal to achieve spatial filtering of the preset audio frame signal, thereby selectively enhancing the sound source to be positioned and suppressing noise signals.

[0036] For example, the electronic device can directionally pick up the preset audio frame signal based on a fixed beam algorithm. In some embodiments, when the preset audio frame signal is directionally picked up based on the fixed beam algorithm, the preset audio frame signal can be directionally picked up according to the first beam parameter to obtain a plurality of first directional pickup signals corresponding to the preset audio frame signal.

[0037] The first beam parameter may comprise, for example, a first angle change step. The first angle change step may be used to determine a plurality of first scanning angles. For example, a sound source corresponding to a speech signal in a preset audio frame signal may exist at any extremely small angle in a space in which the electronic device is located. Under the limitation of system resources of the electronic device, the electronic device cannot perform directional sound pickup of the preset audio frame signal at each infinitesimal angle. Based on this, the electronic device may discretize the space in which the electronic device is located according to the first angle change step to obtain a plurality of first scanning angles. The plurality of first scanning angles may systematically cover the space in which the electronic device is located, and thus may be used by the electronic device to perform sound source positioning of the preset audio frame signal. The first angle change step may be used to indicate an angle difference between adjacent first scanning angles.

[0038] The electronic device may perform sound source positioning of the preset audio frame signal according to different first scanning angles. For example, the electronic device may perform directional sound pickup of the preset audio frame signal at a first scanning angle to obtain a first directional sound pickup signal of the preset audio frame signal corresponding to the first scanning angle. The first directional sound pickup signal corresponding to the first scanning angle is obtained by processing the preset audio frame signal with the first scanning angle as a signal enhancement direction and a direction other than the first scanning angle as a signal suppression direction. Similarly, for a plurality of first scanning angles, a first directional sound pickup signal of the preset audio frame signal corresponding to each of the plurality of first scanning angles may be determined. The first scanning angle and the first directional sound pickup signal correspond to each other.

[0039] Correspondingly, since each first directional sound pickup signal is obtained by performing directional sound pickup of the preset audio frame signal according to a different first scanning angle, the electronic device may determine first positioning information of a sound source corresponding to a speech signal in the preset audio frame signal according to a first scanning angle corresponding to each first directional sound pickup signal. For example, the electronic device may determine a first sound source angle of the sound source corresponding to the speech signal in the preset audio frame signal according to the first scanning angle. The first sound source angle may be used to determine the first positioning information of the sound source.

[0040] In a case where the preset audio frame signal is directional sound picked up according to the first beam parameter to obtain a plurality of first directional sound pickup signals corresponding to the preset audio frame signal, the electronic device may directly perform directional sound pickup of the preset audio frame signal using the first beam parameter, which is beneficial to improving the convenience of directional sound pickup of the preset audio frame signal. The plurality of first directional sound pickup signals may be used to subsequently evaluate the first positioning information of the sound source corresponding to the speech signal in the preset audio frame signal, which is beneficial to improving the convenience of determining the first positioning information of the sound source corresponding to the speech signal in the preset audio frame signal.

[0041] S103, determining at least one second directional sound signal from the plurality of first directional sound signals according to at least two of the first probability of the presence of the speech signal in the plurality of first directional sound signals, the first energy information of the plurality of first directional sound signals, and the first similarity between the plurality of first directional sound signals.

[0042] In a case where the plurality of first directional sound signals corresponding to the preset audio frame signal are acquired, the electronic device can screen and cluster the plurality of first directional sound signals, to determine the first positioning information of the sound source corresponding to the speech signal in the preset audio frame signal subsequently.

[0043] For example, the electronic device can perform voice activity detection on each first directional sound signal to determine the first probability of the presence of the speech signal in each first directional sound signal, and then determine whether the speech signal is present in each first directional sound signal. For example, based on a preset voice activity detection (VAD) model, voice activity detection is performed on the first directional sound signal to obtain the first probability of the presence of the speech signal in the first directional sound signal; when the first probability is greater than or equal to a first probability threshold, it is determined that the speech signal is present in the first directional sound signal; when the first probability is less than the first probability threshold, it is determined that the speech signal is not present in the first directional sound signal. The first probability threshold can be preset or set by the user, which is not limited herein. In a case where the first directional sound signal is input into the VAD model, the VAD model can perform voice activity detection on the first directional sound signal to determine whether the speech signal is present in the first directional sound signal. For example, the VAD model can perform feature extraction on the first directional sound signal to obtain the feature value of the speech signal corresponding to the first directional sound signal. In a case where the feature value of the speech signal in the first directional sound signal is greater than or equal to a first feature value threshold, the VAD model can determine that the first probability of the presence of the speech signal in the first directional sound signal is greater than or equal to the first probability threshold, and then determine that the speech signal is present in the first directional sound signal. Correspondingly, in a case where the feature value of the speech signal in the first directional sound signal is less than the first feature value threshold, the VAD model can determine that the first probability of the presence of the speech signal in the first directional sound signal is less than the first probability threshold, and then determine that the speech signal is not present in the first directional sound signal.

[0044] The electronic device can determine the first energy information of each first directional sound pickup signal to perform energy judgment on each first directional sound pickup signal, measure the energy of each first directional sound pickup signal, and further assist in determining whether the sound source is concentrated in the first scanning angle corresponding to each first directional sound pickup signal. For example, the electronic device can perform time-domain energy calculation on the first directional sound pickup signal to obtain the first energy information of the first directional sound pickup signal, or the electronic device can perform frequency-domain energy calculation on the first directional sound pickup signal to obtain the first energy information of the first directional sound pickup signal. In the process of performing time-domain energy calculation on the first directional sound pickup signal, the amplitude of the first directional sound pickup signal can be directly operated, such as determining the sum of squares of all sample point amplitudes of the first directional sound pickup signal to obtain the first energy information of the first directional sound pickup signal. In the process of performing frequency-domain energy calculation on the first directional sound pickup signal, the first directional sound pickup signal can be subjected to fast Fourier transform to convert the first directional sound pickup signal from the time domain to the frequency domain, and then the power spectrum of the first directional sound pickup signal after fast Fourier transform is calculated, and then the frequency band energy of the first directional sound pickup signal after fast Fourier transform is calculated according to the power spectrum, so as to determine the first energy information of the first directional sound pickup signal. For a plurality of first directional sound pickup signals, if the first energy information of the first directional sound pickup signal is larger, the possibility that the sound source is concentrated in the first scanning angle corresponding to the first directional sound pickup signal is greater; if the first energy information of the first directional sound pickup signal is smaller, the possibility that the sound source is concentrated in the first scanning angle corresponding to the first directional sound pickup signal is smaller. For example, when the first energy information of the first directional sound pickup signal is greater than or equal to the first energy threshold, it is determined that the sound source is concentrated in the first scanning angle corresponding to the first directional sound pickup signal; when the first energy information of the first directional sound pickup signal is less than the first energy threshold, it is determined that the sound source is not concentrated in the first scanning angle corresponding to the first directional sound pickup signal. The first energy threshold can be pre-set or set by the user, which is not limited herein.

[0045] Correspondingly, the electronic device can perform similarity determination on the multiple first directional sound pickup signals to determine the first similarity between the multiple first directional sound pickup signals, and further determine whether the sound source is concentrated in the first scanning angle corresponding to the multiple first directional sound pickup signals. For example, the voice signals existing in the preset audio frame signal can correspond to different sound sources, i.e., the preset audio frame signal includes voice signals corresponding to multiple sound sources respectively. In order to reduce the mutual interference of the multiple sound sources in the process of sound source positioning of the preset audio frame signal, and further improve the sound source positioning accuracy of the preset audio frame signal, the electronic device can determine whether the voice signals existing in the multiple first directional sound pickup signals correspond to the same sound source by determining the first similarity between the multiple first directional sound pickup signals. Taking multiple first directional sound pickup signals including signal A, signal B, and signal C as an example. The electronic device can determine the first similarity between signal A and signal B and signal C respectively. The greater the first similarity between signal A and signal B, the greater the possibility that the voice signals existing in signal A and signal B correspond to the same sound source. Correspondingly, the smaller the first similarity between signal A and signal C, the smaller the possibility that the voice signals existing in signal A and signal C correspond to the same sound source. For example, the multiple first directional sound pickup signals further include signal D. In the case that the greater the first similarity between signal D and signal A, and the greater the first similarity between signal D and signal B, the electronic device can infer that the greater the possibility that the voice signals existing in signal A, signal B, and signal D correspond to the same sound source. For example, when the first similarity between at least two first directional sound pickup signals is greater than or equal to a first similarity threshold, it is determined that the at least two first directional sound pickup signals correspond to the same sound source; when the first similarity between at least two first directional sound pickup signals is less than the first similarity threshold, it is determined that the at least two first directional sound pickup signals do not correspond to the same sound source. The first similarity threshold can be pre-set or set by the user, which is not limited herein.

[0046] The electronic device can screen and cluster the multiple first directional sound pickup signals according to at least two of the first probability of the multiple first directional sound pickup signals, the first energy information of the multiple first directional sound pickup signals, and the first similarity between the multiple first directional sound pickup signals, to determine the first directional sound pickup signals satisfying the first screening condition as the second directional sound pickup signals. The first screening condition includes at least two of the first probability being greater than or equal to a first probability threshold, the first energy information being greater than or equal to a first energy threshold, and the first similarity being greater than or equal to a first similarity threshold.

[0047] In some example embodiments, the electronic device can perform voice signal detection on the first directional sound pickup signal first, and then perform at least one of energy information judgment and similarity judgment on the first directional sound pickup signal; the electronic device can also perform similarity judgment on the first directional sound pickup signal first, and then perform at least one of energy information judgment and voice signal detection on the first directional sound pickup signal; the electronic device can also perform energy information judgment on the first directional sound pickup signal first, and then perform at least one of voice signal detection and similarity judgment on the first directional sound pickup signal, without limitation.

[0048] Based on this, in the case of determining at least one second directional sound pickup signal from the plurality of first directional sound pickup signals, it can be determined that there is a voice signal in the second directional sound pickup signal, or the possibility of a voice signal in the second directional sound pickup signal is relatively large, and further the second directional sound pickup signal can be used to determine the target positioning information of the sound source corresponding to the voice signal in the preset audio frame signal.

[0049] Since the second directional sound pickup signal is determined by comprehensively determining at least two of the voice signal detection, the energy information judgment, and the similarity judgment on the first directional sound pickup signal, the second directional sound pickup signal can be used to distinguish the voice signal from the noise signal present in the preset audio frame signal, and can be used to determine whether the sound sources corresponding to the voice signals are the same sound source, and further to reduce noise interference or multiple sound source interference in the process of sound source positioning of the preset audio frame signal, so as to reduce the possibility of excluding the voice signal due to noise signal misjudgment or incorrectly identifying the noise signal as the voice signal. Therefore, it can be known that the determination of the second directional sound pickup signal is beneficial to subsequent improvement of the convenience and accuracy of the sound source positioning of the preset audio frame signal.

[0050] S104, determining first positioning information of a sound source corresponding to a voice signal in the preset audio frame signal according to the at least one second directional sound pickup signal.

[0051] For example, the first positioning information includes a first sound source angle of the sound source corresponding to the voice signal in the preset audio frame signal.

[0052] In some embodiments, the first scanning angle corresponding to the second directional sound pickup signal is determined as the first sound source angle of the sound source.

[0053] Since the first directional sound pickup signal corresponding to the second directional sound pickup signal meets at least two of the first probability being greater than or equal to the first probability threshold, the first energy information being greater than or equal to the first energy threshold, and the first similarity being greater than or equal to the first similarity threshold, the electronic device can infer that the first scanning angle corresponding to the first directional sound pickup signal can cover the sound source corresponding to the voice signal in the preset audio frame signal, and further the first scanning angle corresponding to the second directional sound pickup signal can be determined as the first sound source angle of the sound source.

[0054] Correspondingly, in a case where the plurality of second directional sound signals are determined, if the first similarity between the first directional sound signals corresponding to the plurality of second directional sound signals respectively is greater than or equal to the first similarity threshold, it can be determined that the plurality of second directional sound signals correspond to the same sound source, and then the first scanning angle corresponding to the plurality of second directional sound signals respectively is determined as the corresponding first scanning angle set, and the first scanning angle set is determined as the first sound source angle of the same sound source. If the first similarity between the first directional sound signals corresponding to the plurality of second directional sound signals respectively is less than the first similarity threshold, it can be determined that the plurality of second directional sound signals correspond to different sound sources, and then the first sound source angle of the different sound sources is determined according to the first scanning angle corresponding to the plurality of second directional sound signals respectively.

[0055] In a case where the first positioning information of the sound source corresponding to the speech signal in the preset audio frame signal is determined according to the at least one second directional sound signal, since the second directional sound signal can be used to distinguish the speech signal from the noise signal existing in the preset audio frame signal, and can be used to determine whether the sound sources corresponding to the speech signals are the same sound source, the first positioning information of the sound source determined according to the at least one second directional sound signal can be used to reduce noise interference or multiple sound source interference in the process of positioning the sound source in the preset audio frame signal, thereby improving the sound source positioning accuracy of the preset audio frame signal.

[0056] In some embodiments, after the at least one second directional sound signal is determined from the plurality of first directional sound signals according to at least two of the first probability of the speech signal existing in the plurality of first directional sound signals, the first energy information of the plurality of first directional sound signals, and the first similarity between the plurality of first directional sound signals, the method further includes: performing directional sound picking on the second directional sound signal according to a second beam parameter to obtain a plurality of third directional sound signals corresponding to the second directional sound signal; the second beam parameter is different from the first beam parameter; determining a fourth directional sound signal from the at least one second directional sound signal according to at least two of a second probability of the speech signal existing in the plurality of third directional sound signals corresponding to the at least one second directional sound signal respectively, second energy information of the plurality of third directional sound signals, and second similarity between the plurality of third directional sound signals.

[0057] Determining the first positioning information of the sound source corresponding to the speech signal in the preset audio frame signal according to the at least one second directional sound signal includes: determining the second positioning information of the sound source corresponding to the speech signal in the preset audio frame signal according to the fourth directional sound signal.

[0058] In a case where the second directional pickup signal is acquired, the electronic device can perform directional pickup on the second directional pickup signal based on a fixed beam algorithm to achieve multi-level directional pickup on the preset audio frame signal in combination with directional pickup performed by the electronic device on the preset audio frame signal. In some embodiments, in a case where directional pickup is performed on the second directional pickup signal based on the fixed beam algorithm, directional pickup can be performed on the second directional pickup signal according to the second beam parameter to obtain a plurality of third directional pickup signals corresponding to the second directional pickup signal. The second beam parameter is different from the first beam parameter.

[0059] For example, the second beam parameter includes a second angle change step. The second angle change step can be used to determine a plurality of second scanning angles. The second angle change step is smaller than the first angle change step included in the first beam parameter. The second angle change step can be used to indicate the angle interval between adjacent second scanning angles. The description of the plurality of second scanning angles determined according to the second angle change step can refer to the description of the plurality of first scanning angles determined according to the first angle change step, which will not be repeated here. For example, in a case where the second angle change step is smaller than the first angle change step, the degree of overlap between every two adjacent second scanning angles is greater than the degree of overlap between every two adjacent first scanning angles, and the scanning density between every two adjacent second scanning angles is greater than the scanning density between every two adjacent first scanning angles, then compared with directional pickup on the preset audio frame signal according to the first beam parameter, the second beam parameter can be used to perform more refined directional pickup on the second directional pickup signal to refine and adjust the positioning information of the potential sound source, such as the first scanning angle corresponding to the second directional pickup signal, thereby improving the sound source positioning accuracy in the sound source positioning process of the preset audio frame signal.

[0060] In a case where the plurality of third directional pickup signals corresponding to the at least one second directional pickup signal are acquired, the electronic device can again evaluate the at least one second directional pickup signal to determine a fourth directional pickup signal from the at least one second directional pickup signal. The fourth directional pickup signal can be used to determine the second positioning information of the sound source corresponding to the speech signal in the preset audio frame signal.

[0061] For example, the electronic device can determine a second probability that a voice signal exists in each third directional pickup signal corresponding to each second directional pickup signal, to determine whether a voice signal exists in each third directional pickup signal corresponding to each second directional pickup signal. For example, based on a preset VAD model, voice activity detection is performed on the third directional pickup signal to obtain a second probability that a voice signal exists in the third directional pickup signal; when the second probability is greater than or equal to a second probability threshold, it is determined that a voice signal exists in the third directional pickup signal; when the second probability is less than the second probability threshold, it is determined that a voice signal does not exist in the third directional pickup signal. The second probability threshold can be preset or set by the user, which is not limited herein. In an exemplary embodiment, the second probability threshold is greater than the first probability threshold, so that the electronic device can select a fourth directional pickup signal with a higher possibility of existing voice information from at least one second directional pickup signal, thereby improving the sound source positioning accuracy of the preset audio frame signal.

[0062] The electronic device can determine second energy information of the plurality of third directional pickup signals corresponding to each second directional pickup signal, to perform energy judgment on the plurality of third directional pickup signals corresponding to each second directional pickup signal, measure the energy of each third directional pickup signal, and further assist in determining whether the sound source is concentrated in the second scanning angle corresponding to each third directional pickup signal. The relevant description of determining the second energy information of the plurality of third directional pickup signals corresponding to the second directional pickup signal can refer to the relevant description of determining the first energy information of the first directional pickup signal, which will not be repeated here. For the plurality of third directional pickup signals corresponding to each second directional pickup signal, the greater the second energy information of the third directional pickup signal, the greater the possibility that the sound source is concentrated in the second scanning angle corresponding to the third directional pickup signal; the smaller the second energy information of the third directional pickup signal, the smaller the possibility that the sound source is concentrated in the second scanning angle corresponding to the third directional pickup signal. For example, when the second energy information corresponding to the third directional pickup signal is greater than or equal to a second energy threshold, it is determined that the sound source is concentrated in the second scanning angle corresponding to the third directional pickup signal; when the second energy information corresponding to the third directional pickup signal is less than the second energy threshold, it is determined that the sound source is not concentrated in the second scanning angle corresponding to the third directional pickup signal. The second energy threshold can be preset or set by the user, which is not limited herein. In an exemplary embodiment, the second energy threshold is greater than the first energy threshold, so that the electronic device can select a fourth directional pickup signal with a higher possibility of being concentrated in the corresponding second scanning angle from at least one second directional pickup signal, thereby improving the sound source positioning accuracy of the preset audio frame signal.

[0063] Correspondingly, the electronic device can determine the second similarity between the plurality of third directional pickup signals corresponding to each second directional pickup signal, to determine whether the sound source is concentrated in the second scanning angle corresponding to the plurality of third directional pickup signals. For example, because the speech signal existing in the preset audio frame signal can correspond to different sound sources, that is, the preset audio frame signal includes speech signals corresponding to a plurality of sound sources respectively, in order to reduce the mutual interference of the plurality of sound sources in the process of sound source positioning of the preset audio frame signal, and further improve the sound source positioning accuracy of the preset audio frame signal, the electronic device can further determine the second similarity between the plurality of third directional pickup signals corresponding to the second directional pickup signal on the basis of the at least one second directional pickup signal, and further evaluate whether the speech signal existing in the plurality of third directional pickup signals corresponds to the same sound source. Taking an example that the at least one second directional pickup signal includes signal A, and the plurality of third directional pickup signals corresponding to signal A include signal A1, signal A2 and signal A3. The electronic device can determine the second similarity between signal A1 and signal A2, signal A3 respectively. The greater the second similarity between signal A1 and signal A2, the greater the possibility that the speech signals existing in signal A1 and signal A2 correspond to the same sound source. Correspondingly, the smaller the second similarity between signal A1 and signal A3, the smaller the possibility that the speech signals existing in signal A1 and signal A3 correspond to the same sound source. For example, the plurality of third directional pickup signals further include signal A4, and the greater the second similarity between signal A4 and signal A1, and the greater the second similarity between signal A4 and signal A2, the greater the possibility that the speech signals existing in signal A1, signal A2 and signal A4 correspond to the same sound source. For each second directional pickup signal, the greater the second similarity between different third directional pickup signals, the greater the possibility that the speech signals existing in the different third directional pickup signals correspond to the same sound source. For example, when the second similarity between the at least two third directional pickup signals corresponding to the second directional pickup signal is greater than or equal to the second similarity threshold, it is determined that the at least two third directional pickup signals correspond to the same sound source; when the second similarity between the at least two third directional pickup signals corresponding to the second directional pickup signal is less than the second similarity threshold, it is determined that the at least two third directional pickup signals do not correspond to the same sound source. The second similarity threshold can be pre-set or set by the user, which is not limited herein.In an example implementation, the second similarity threshold is greater than the first similarity threshold, so that the electronic device filters the second directional pickup signals to obtain the fourth directional pickup signal with a higher possibility of corresponding to the same sound source from the at least one second directional pickup signal, thereby improving the sound source positioning accuracy of the preset audio frame signal.

[0064] The electronic device can determine the second directional pickup signal satisfying the second filtering condition as the fourth directional pickup signal from the at least one second directional pickup signal according to at least two of the second probabilities of the plurality of third directional pickup signals corresponding to each of the second directional pickup signal, the second energy information of the plurality of third directional pickup signals, and the second similarity between the plurality of third directional pickup signals. The second filtering condition includes at least two of the second probability of each of the plurality of third directional pickup signals being greater than or equal to a second probability threshold, the second energy information of each of the plurality of third directional pickup signals being greater than or equal to a second energy threshold, and the second similarity between the plurality of third directional pickup signals being greater than or equal to a second similarity threshold.

[0065] For example, in a case where the second probability of each of the plurality of third directional pickup signals corresponding to the same second directional pickup signal is greater than or equal to the second probability threshold, the electronic device can determine that the voice signal exists in the plurality of third directional pickup signals corresponding to the second directional pickup signal. Since the second scanning angle corresponding to each of the plurality of third directional pickup signals is determined according to the first scanning angle corresponding to the second directional pickup signal, the electronic device can infer that the voice signal existing in the plurality of third directional pickup signals has a higher possibility of corresponding to the same sound source.

[0066] In a case where the second energy information of each of the plurality of third directional pickup signals corresponding to the same second directional pickup signal is greater than or equal to the second energy threshold, the electronic device can determine that the sound source corresponding to each of the plurality of third directional pickup signals corresponding to the second directional pickup signal is concentrated in the second scanning angle corresponding to the corresponding third directional pickup signal. Since the second scanning angle corresponding to each of the plurality of third directional pickup signals is determined according to the first scanning angle corresponding to the second directional pickup signal, the electronic device can infer that the voice signal existing in the plurality of third directional pickup signals has a higher possibility of corresponding to the same sound source.

[0067] In a case where the second similarity between the plurality of third directional pickup signals corresponding to the same second directional pickup signal is greater than or equal to the second similarity threshold, the electronic device can determine that the voice signal existing in the plurality of third directional pickup signals corresponding to the second directional pickup signal has a higher possibility of corresponding to the same sound source.

[0068] Based on this, the electronic device can determine that the second directional pickup signal is the fourth directional pickup signal when the voice signal existing in the multiple directional pickup signals corresponding to the same second directional pickup signal is more likely to correspond to the same sound source, such as when at least two of the following conditions are met: the second probability of the second directional pickup signal corresponding to the multiple third directional pickup signals respectively is greater than or equal to the second probability threshold, the second energy information of the second directional pickup signal corresponding to the multiple third directional pickup signals respectively is greater than or equal to the second energy threshold, and the second similarity between the second directional pickup signals corresponding to the multiple third directional pickup signals is greater than or equal to the second similarity threshold.

[0069] In a case where the fourth directional pickup signal is determined, the electronic device can determine, according to the fourth directional pickup signal, second positioning information of the sound source corresponding to the voice signal in the preset audio frame signal.

[0070] For example, the second positioning information includes a second sound source angle of the sound source corresponding to the voice signal in the preset audio frame signal.

[0071] In some embodiments, the second scanning angle corresponding to the fourth directional pickup signal is determined as the second sound source angle of the sound source.

[0072] Since the multiple third directional pickup signals corresponding to the second directional pickup signal corresponding to the fourth directional pickup signal all meet at least two of the following conditions: the second probability is greater than or equal to the second probability threshold, the second energy information is greater than or equal to the second energy threshold, and the second similarity is greater than or equal to the second similarity threshold, the electronic device can infer that the second scanning angle corresponding to the fourth directional pickup signal can cover the sound source corresponding to the voice signal in the preset audio frame signal, and thus the second scanning angle corresponding to the fourth directional pickup signal can be determined as the second sound source angle of the sound source.

[0073] Correspondingly, in a case where multiple fourth directional pickup signals are determined, if the second similarity between the second directional pickup signals corresponding to the multiple fourth directional pickup signals respectively is greater than or equal to the second similarity threshold, it can be determined that the multiple fourth directional pickup signals correspond to the same sound source, and thus the second scanning angle corresponding to the multiple fourth directional pickup signals respectively is determined as a second scanning angle set, and the second scanning angle set is determined as the second sound source angle of the same sound source. If the second similarity between the second directional pickup signals corresponding to the multiple fourth directional pickup signals respectively is less than the second similarity threshold, it can be determined that the multiple fourth directional pickup signals correspond to different sound sources, and thus the second sound source angle of the different sound sources is determined according to the second scanning angle corresponding to the multiple fourth directional pickup signals respectively.

[0074] Since the second scanning angle corresponding to the fourth directional sound pickup signal is determined according to the second beam parameter, and the second beam parameter is different from the first beam parameter, the second scanning angle corresponding to the fourth directional sound pickup signal is actually obtained by performing angle fine adjustment on the first scanning angle corresponding to the second directional sound pickup signal. The higher the degree of angle fine adjustment on the first scanning angle, the higher the accuracy of determining the second sound source angle of the sound source according to the second scanning angle, which is conducive to improving the sound source positioning accuracy of the preset audio frame signal.

[0075] Correspondingly, since the fourth directional sound pickup signal is determined by comprehensively determining at least two of the voice signal detection, the energy information judgment, and the similarity judgment of the second directional sound pickup signal, the fourth directional sound pickup signal can be used to distinguish the voice signal and the noise signal existing in the preset audio frame signal. Moreover, the fourth directional sound pickup signal is determined by comprehensively determining whether the sound sources corresponding to the voice signals existing in the plurality of third directional sound pickup signals corresponding to each of the at least one second directional sound pickup signal are the same sound source, which is conducive to reducing noise interference or multiple sound source interference in the process of positioning the sound source of the preset audio frame signal, thereby reducing the possibility that the voice signal is excluded due to noise signal misjudgment or the noise signal is incorrectly identified as the voice signal, and further improving the sound source positioning accuracy of the preset audio frame signal.

[0076] In some embodiments, an initial audio frame signal is obtained; voice activity detection is performed on the initial audio frame signal based on a preset voice activity detection model to obtain a third probability that a voice signal exists in the initial audio frame signal; and when the third probability is greater than or equal to a third probability threshold, the initial audio frame signal is determined as the preset audio frame signal.

[0077] The initial audio frame signal is taken as a signal , and a voice activity detection (VAD) model is represented as For example. The electronic device can input the signal to the VAD model, so that the VAD model performs voice activity detection on the signal to obtain a third probability that a voice signal exists in the signal .

[0078] The process that the VAD model performs voice activity detection on the signal to obtain the third probability that a voice signal exists in the signal may be represented as:

[0079]

[0080] wherein, is used to indicate the third probability that a voice signal exists in the signal .

[0081] In a case where the third probability is determined , the third probability may be compared with a third probability threshold value to determine whether the voice signal exists in the signal .

[0082] When the third probability satisfies:

[0083]

[0084] It can be determined that the voice signal exists in the signal , and the signal may be determined as the preset audio frame signal for subsequent sound source positioning of the preset audio frame signal, to obtain the first positioning information of the sound source corresponding to the voice signal in the preset audio frame signal, or to obtain the second positioning information of the sound source corresponding to the voice signal in the preset audio frame signal. The third probability threshold value can be pre-set or set by the user, which is not limited herein. In an exemplary embodiment, the third probability threshold value is less than or equal to the first probability threshold value, so that the electronic device can identify as many preset audio frame signals as possible from the initial audio frame signal, and the electronic device can subsequently perform sound source positioning on the preset audio frame signal.

[0085] For example, the first beam parameter includes a first angle change step. The first angle change step can be used to preliminarily determine the approximate direction of the sound source corresponding to the voice signal in the preset audio frame signal. Moreover, the first angle change step affects the number of identifiable sound sources and the angle resolution. Of course, the first beam parameter is not limited to this, and the first beam parameter can also include the half-power beamwidth (HPBW) of the beam. The HPBW of the beam can be used to determine the width of a single beam, and the first angle change step can be used to determine the degree of overlap between adjacent beams and the scanning density.

[0086] In some embodiments, a plurality of first scanning angles corresponding to the preset angle are determined according to the preset angle corresponding to the preset audio frame signal and the first angle change step; and the preset audio frame signal is directionally picked up according to the first scanning angle, to obtain a first directional pickup signal of the preset audio frame signal corresponding to the first scanning angle.

[0087] For example, the preset audio frame signal is taken as the signal , the preset angle is angle , and the first angle change step is step .

[0088] According to the angle and the step , a plurality of first scanning angles are determined.The angle can be determined The plurality of first scanning angles can be uniformly represented as an angle , and the plurality of angles can constitute a first scanning angle set. The first scanning angle set can be represented as:

[0089]

[0090] Wherein, i is used to indicate the ordinal number of the first scanning angle, i.e. the i-th first scanning angle.

[0091] The signal corresponds to the angle , which is the starting scanning angle when the signal is directionally picked up according to the first beam parameter. The smaller the value of the step , the higher the angle resolution when the signal is directionally picked up, and correspondingly, the larger the amount of calculation.

[0092] For example, different angles in the first scanning angle set can correspond to different beam directions, and then the preset audio frame signal can be directionally picked up according to the first scanning angle to obtain a first directional pickup signal corresponding to the first scanning angle. Correspondingly, for the plurality of first scanning angles, the first directional pickup signal corresponding to the different first scanning angles of the preset audio frame signal can be determined, i.e. the plurality of first directional pickup signals corresponding to the preset audio frame signal can be determined.

[0093] In an exemplary embodiment, the angle can be set to -90° or 0°, of course, not limited thereto, which is not limited herein.

[0094] In an exemplary embodiment, the step can be set to greater than 0° and less than or equal to 180°, or set to greater than 0° and less than or equal to 360°, of course, not limited thereto, which is not limited herein.

[0095] In the case of determining the plurality of first directional pickup signals corresponding to the preset audio frame signal, it is equivalent to realizing the first directional pickup of the preset audio frame signal, and then the plurality of first directional pickup signals can be used to determine the first positioning information of the sound source corresponding to the voice signal in the preset audio frame signal, or determine the second positioning information of the sound source corresponding to the voice signal in the preset audio frame signal.

[0096] In some implementations, when a first directional pickup signal meets at least two of the following conditions: a first probability is greater than or equal to a first probability threshold, a first energy is greater than or equal to a first energy threshold, and a first similarity is greater than or equal to a first similarity threshold, the first directional pickup signal is determined to be a second directional pickup signal.

[0097] Based on the consideration of reducing noise interference or multiple sound source interference during the sound source localization process of preset audio frame signals, electronic devices can screen and preliminarily cluster multiple first directional sound pickup signals, such as comparing the first directional sound pickup signals corresponding to multiple first scanning angles in pairs, so as to screen out the first directional sound pickup signal with higher signal-to-noise ratio and stronger human sound source characteristics as the second directional sound pickup signal.

[0098] A set consisting of multiple first directional pickup signals ,gather The first directional pickup signal in the signal includes the signal and signals ,and For example, the signal. and signals It can be used to indicate the first directional pickup signal at different first scanning angles.

[0099] Electronic devices can calculate signals and signals The first similarity between them, to the signals and signals Perform similarity judgment and calculate signals separately. and signals Each's first energy information, in order to respond to the signal and signals Energy determination is performed. Correspondingly, the signal is comprehensively analyzed. and signals Similarity and energy assessments can be used by electronic devices to determine whether a sound source is concentrated in a signal. and signals Each has its corresponding first scanning angle, which is then used to assess the angular concentration of the sound source.

[0100] For example, signals and signals The first similarity between them can include signals and signals The Pearson correlation coefficient between them is not restricted here. and signals The calculation process for the first similarity between them can be expressed as:

[0101]

[0102] in, Used to indicate signals and signals The first similarity between them; Used to indicate signals and signals The common expected value, Used to indicate signals The mean, Used to indicate signals The mean; Used to indicate signals standard deviation Used to indicate signals The standard deviation.

[0103] Of course, the signal and signals The first similarity between them is not limited to signals and signals The Pearson correlation coefficient between them is not restricted here.

[0104] Signal and signals First similarity between Measurable signals and signals The degree of similarity is used to exclude irrelevant noise signals.

[0105] For example, the process by which an electronic device calculates the first energy information of the first directional pickup signal can be represented as:

[0106]

[0107] in, Used to indicate the first directional pickup signal The first energy information. The first directional pickup signal. Can include signals and signals Then the signal The first energy information can be represented as And so on, signal The first energy information can be represented as .

[0108] For example, the first similarity Similarity threshold Comparison, and the first energy information First energy information Respectively compared with the first energy threshold In , and the case, it can be inferred that the sound source corresponding to the speech signal in the preset audio frame signal is concentrated in the signal and the first scanning angle corresponding to the signal . may be determined according to the average or variance of the first energy information of each of the plurality of first directional pickup signals corresponding to the plurality of first scanning angles, and of course is not limited thereto, but can also be pre-set or set by the user, which is not limited herein.

[0109] Of course, it is not limited thereto, and for the plurality of first directional pickup signals, the electronic device can jointly perform similarity judgment and energy judgment on the plurality of first directional pickup signals to determine that the sound source corresponding to the speech signal in the preset audio frame signal is concentrated in the corresponding first scanning angle, and then take the corresponding first scanning angle as the potential sound source candidate angle of the sound source, and take the first directional pickup signal corresponding to the corresponding first scanning angle as the high-confidence directional pickup signal. The first scanning angle as the potential sound source candidate angle of the sound source can constitute an angle set , and each first scanning angle in the angle set may be uniformly represented as an angle . Correspondingly, each first scanning angle in the angle set corresponds to a first directional pickup signal, which can constitute a signal set , and each first directional pickup signal in the signal set may be uniformly represented as a signal , and the signal is the high-confidence directional pickup signal.

[0110] For the signal obtained through screening and preliminary clustering, speech activity detection can be further performed to determine whether there is a speech signal in the signal , and then to ensure that the directional pickup signal involved in the subsequent sound source positioning of the preset audio frame signal includes human voice features.

[0111] For example, the signal may be input into a VAD model to obtain a first probability that there is a speech signal in the signal . Correspondingly, the first probability may be compared with a first probability threshold to determine whether there is a speech signal in the signal .

[0112] When the first probability satisfies:

[0113]

[0114] determining the signal contains a speech signal.

[0115] Based on this, the signal meets at least two of the following conditions: the first probability is greater than or equal to a first probability threshold value , the first energy information of the signal is greater than or equal to a first energy threshold value , and the first similarity of the signal is greater than or equal to a first similarity threshold value , and the signal is determined as a second directional sound pickup signal.

[0116] The second directional sound pickup signal can be used to subsequently determine first positioning information of a sound source corresponding to a speech signal in a preset audio frame signal.

[0117] For example, the second beam parameter at least includes a second angle change step and a first angle difference threshold value; the second angle change step is smaller than the first angle change step included in the first beam parameter. Descriptions related to the second angle change step can refer to the foregoing descriptions of the first angle change step, which will not be repeated here. Of course, the second beam parameter is not limited to this, and the second beam parameter can also include the HPBW of the beam, which is not limited here.

[0118] In some embodiments, according to the first scanning angle corresponding to the second directional sound pickup signal and the second angle change step, a plurality of second scanning angles corresponding to the first scanning angle are determined; the absolute value of the angle difference between each second scanning angle and the first scanning angle is less than or equal to the first angle difference threshold value; and according to the second scanning angle, the second directional sound pickup signal is directionally picked up to obtain a third directional sound pickup signal corresponding to the second scanning angle.

[0119] Taking the second directional sound pickup signal as the signal , the first scanning angle corresponding to the signal is the angle , the second angle change step is the step , the first angle difference threshold value is , and the first angle change step is the step , for example.

[0120] In the case where the step is less than the step , the electronic device can perform directional sound pickup on the signal corresponding to the angle at the angle a smaller fine tuning step size, such as a step size and a fine tuning range, such as a first angle difference threshold to search for a second positioning information of the sound source corresponding to the speech signal in the preset audio frame signal.

[0121] According to the angle and the step size , a plurality of second scanning angles corresponding to the angle can be determined. The plurality of second scanning angles can be uniformly expressed as an angle , and the plurality of second scanning angles can constitute a second scanning angle set. The second scanning angle set can be expressed as:

[0122]

[0123] wherein n and m are integers, to indicate the angle difference between the angle and the angle , that is, the angle difference between the first scanning angle and the second scanning angle, and the absolute value of is less than or equal to the first angle difference threshold .

[0124] For example, different second scanning angles in the second scanning angle set can correspond to different beam directions, and then directional sound pickup can be performed on the second directional sound pickup signal according to the second scanning angle to obtain a third directional sound pickup signal corresponding to the second scanning angle. Accordingly, for a plurality of second scanning angles, a third directional sound pickup signal corresponding to each second scanning angle of the second directional sound pickup signal can be determined, that is, a plurality of third directional sound pickup signals corresponding to the second directional sound pickup signal can be determined.

[0125] Accordingly, in the case where there is at least one second directional sound pickup signal, a plurality of third directional sound pickup signals corresponding to each second directional sound pickup signal can be determined.

[0126] In the case where a plurality of third directional sound pickup signals corresponding to each second directional sound pickup signal are determined, a two-level directional sound pickup of the preset audio frame signal is realized. Based on the determination of the first directional sound pickup signal and the determination of the third directional sound pickup signal, a multi-level directional sound pickup of the preset audio frame signal can be realized. The plurality of third directional sound pickup signals corresponding to each second directional sound pickup signal can be used for subsequent determination of the second positioning information of the sound source corresponding to the speech signal in the preset audio frame signal, so as to further refine the first positioning information of the sound source, and thus improve the sound source positioning accuracy of the preset audio frame signal.

[0127] In some embodiments, the second directional pickup signal is determined as the fourth directional pickup signal when at least two of the following conditions are met: the second probability of the second directional pickup signal is greater than or equal to a second probability threshold, the second energy information of the second directional pickup signal is greater than or equal to a second energy threshold, and the second similarity between the second directional pickup signal and the corresponding plurality of third directional pickup signals is greater than or equal to a second similarity threshold.

[0128] Based on the consideration of reducing noise interference or multiple sound source interference in the process of sound source positioning of the preset audio frame signal, the electronic device can perform at least two of the voice activity detection, the similarity judgment, and the energy judgment on the plurality of third directional pickup signals corresponding to each second directional pickup signal, to further determine the fourth directional pickup signal. The fourth directional pickup signal is obtained by screening at least one second directional pickup signal, and thus the sound source positioning accuracy of the preset audio frame signal can be improved.

[0129] The plurality of third directional pickup signals corresponding to the second directional pickup signal form a set , and the third directional pickup signals in the set may be uniformly represented as For example.

[0130] For each third directional pickup signal in the set , a similarity judgment can be performed thereon to obtain a second similarity between each two third directional pickup signals in the set , an energy judgment can be performed thereon to obtain second energy information of each third directional pickup signal , and a voice activity detection can be performed thereon to obtain a second probability of each third directional pickup signal .

[0131] When each third directional pickup signal in the set meets at least two of the following conditions: the second probability of each third directional pickup signal is greater than or equal to a second probability threshold , the second energy information of each third directional pickup signal is greater than or equal to a second energy threshold , and the second similarity between each two third directional pickup signals is greater than or equal to a second similarity threshold , the electronic device can determine that the plurality of third directional pickup signals in the set .They consistently exhibit high similarity, high energy, and high VAD probability, thus allowing us to determine the set. The corresponding second directional pickup signal is the fourth directional pickup signal.

[0132] Of course, this is not the only possibility. When the second directional pickup signal meets the following conditions: the second probability of at least one third directional pickup signal is less than the second probability threshold, the second energy information of at least one third directional pickup signal is less than the second energy threshold, and the second similarity between at least two third directional pickup signals is less than the second similarity threshold, the electronic device cannot determine that the second directional pickup signal is the fourth directional pickup signal. Based on this, the electronic device can continue to select at least two third directional pickup signals from the multiple third directional pickup signals corresponding to the second directional pickup signal that meet at least two of the following conditions: the second probability is greater than or equal to the second probability threshold, the second energy information is greater than or equal to the second energy threshold, and the second similarity is greater than or equal to the second similarity threshold, as the fifth directional pickup signal. Accordingly, the electronic device can continue to perform directional pickup on the fifth directional pickup signal according to the third beam parameters to obtain multiple sixth directional pickup signals corresponding to the fifth directional pickup signal. The third beam parameters are different from the second beam parameters. Taking the third beam parameters including the third angle change step size and the second angle difference threshold as an example: the third angle change step size is less than or equal to the second scanning angle change step size, and the second angle difference threshold is less than or equal to the first angle difference threshold. Of course, this is not limited to these limitations. An electronic device can determine a fifth directional pickup signal as a seventh directional pickup signal if at least two of the following conditions are met: the third probability of a speech signal in each of the corresponding plurality of sixth directional pickup signals is greater than or equal to a third probability threshold; the third energy information of each of the corresponding plurality of sixth directional pickup signals is greater than or equal to a third energy threshold; and the third similarity among the corresponding plurality of fifth directional pickup signals is greater than or equal to a third similarity threshold. The seventh directional pickup signal can be used by the electronic device to determine the third localization information of the sound source corresponding to the speech signal in a preset audio frame signal. Similarly, directional pickup signals corresponding to the speech signal in the preset audio frame signal, such as the second directional pickup signal, the fourth directional pickup signal, the seventh directional pickup signal, etc., are iteratively optimized and verified to determine the final directional pickup signal. Based on the final directional pickup signal, the localization information of the sound source corresponding to the speech signal in the preset audio frame signal is determined, such as determining one of the first localization information, second localization information, third localization information, etc., of the sound source corresponding to the speech signal in the preset audio frame signal.

[0133] The fourth directional pickup signal can be used to determine the second localization information of the sound source corresponding to the speech signal in the preset audio frame signal.

[0134] For example, the second positioning information includes a second sound source angle of a sound source corresponding to the speech signal in the preset audio frame signal. The second scan angle corresponding to the fourth directional sound pickup signal is determined as the second sound source angle of the sound source, so that the second positioning information of the sound source can be determined.

[0135] Since the second scan angle corresponding to the fourth directional sound pickup signal is determined according to the second beam parameter, and the second beam parameter is different from the first beam parameter, the second scan angle corresponding to the fourth directional sound pickup signal is equivalent to the first scan angle corresponding to the second directional sound pickup signal corresponding to the fourth directional sound pickup signal, which is angle fine-tuned. This is beneficial to improve the sound source positioning accuracy of the preset audio frame signal.

[0136] The sound source positioning method provided in the above embodiments comprises the following steps: obtaining a preset audio frame signal, wherein the preset audio frame signal contains a speech signal; performing directional sound pickup on the preset audio frame signal according to a first beam parameter to obtain a plurality of first directional sound pickup signals corresponding to the preset audio frame signal; determining at least one second directional sound pickup signal from the plurality of first directional sound pickup signals according to at least two of a first probability of the speech signal existing in the plurality of first directional sound pickup signals, first energy information of the plurality of first directional sound pickup signals, and a first similarity between the plurality of first directional sound pickup signals; and determining first positioning information of a sound source corresponding to the speech signal in the preset audio frame signal according to the at least one second directional sound pickup signal.

[0137] In the case of determining the second directional sound pickup signal by comprehensively considering at least two of the first probability, the first energy information, and the first similarity, the second directional sound pickup signal is equivalent to being determined by comprehensively considering at least two of speech signal detection, energy information judgment, and similarity judgment on the preset audio frame signal. The second directional sound pickup signal can be used to distinguish the speech signal from the noise signal in the preset audio frame signal, and can be used to determine whether the sound sources corresponding to the speech signals are the same sound source, thereby reducing noise interference or multiple sound source interference in the process of positioning the sound source of the preset audio frame signal. Based on the reduction of noise interference or multiple sound source interference in the process of positioning the sound source of the preset audio frame signal, it is beneficial to improve the sound source positioning accuracy of the preset audio frame signal.

[0138] Please refer to Figure 2 , Figure 2is a schematic block diagram of a sound source positioning device provided by an embodiment of the present application. The sound source positioning device can be configured in an electronic device or a server, and is used to execute the sound source positioning method described above. The electronic device can include a near-eye display device, a wearable device, a terminal device, etc., without limitation. The near-eye display device can include AR glasses, VR glasses, MR glasses, an AR helmet, a VR helmet, an MR helmet, etc., without limitation. The wearable device includes a smart watch, a smart ring, a smart bracelet, etc., without limitation. The terminal device includes a mobile phone, a television, a computer, etc., without limitation. The server can be a separate server, or a cloud server providing cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content distribution networks, and basic cloud computing services such as big data and artificial intelligence platforms.

[0139] As shown in Figure 2 , the sound source positioning device includes an audio acquisition module 110, a first signal determination module 120, a second signal determination module 130, and a sound source positioning module 140.

[0140] The audio acquisition module 110 is configured to acquire a preset audio frame signal, wherein the preset audio frame signal contains a speech signal.

[0141] The first signal determination module 120 is configured to perform directional sound pickup on the preset audio frame signal according to a first beam parameter, to obtain a plurality of first directional sound pickup signals corresponding to the preset audio frame signal.

[0142] The second signal determination module 130 is configured to determine at least one second directional sound pickup signal from the plurality of first directional sound pickup signals according to at least two of a first probability that the speech signal exists in the plurality of first directional sound pickup signals, first energy information of the plurality of first directional sound pickup signals, and a first similarity between the plurality of first directional sound pickup signals.

[0143] The sound source positioning module 140 is configured to determine first positioning information of a sound source corresponding to the speech signal in the preset audio frame signal according to the at least one second directional sound pickup signal.

[0144] For example, the first beam parameter includes a first angle change step; and the first signal determination module 120 includes a first angle determination submodule and a first directional sound pickup submodule.

[0145] The first angle determination submodule is configured to determine a plurality of first scanning angles corresponding to a preset angle of the preset audio frame signal according to the preset angle and the first angle change step.

[0146] The first directional sound pickup sub-module is configured to perform directional sound pickup on the preset audio frame signal according to the first scanning angle to obtain a first directional sound pickup signal corresponding to the first scanning angle.

[0147] The second signal determination module 130 includes a first clustering sub-module.

[0148] The first clustering sub-module is configured to determine the first directional sound pickup signal as a second directional sound pickup signal when at least two of the following conditions are met: the first directional sound pickup signal meets the first probability threshold, the first energy information is greater than or equal to the first energy threshold, and the first similarity is greater than or equal to the first similarity threshold.

[0149] The sound source positioning device further includes a third signal determination sub-module and a fourth signal determination sub-module.

[0150] The third signal determination sub-module is configured to perform directional sound pickup on the second directional sound pickup signal according to a second beam parameter to obtain a plurality of third directional sound pickup signals corresponding to the second directional sound pickup signal; the second beam parameter is different from the first beam parameter.

[0151] The fourth signal determination sub-module is configured to determine a fourth directional sound pickup signal from at least one of the second directional sound pickup signals according to at least two of the following: a second probability of a speech signal existing in the plurality of third directional sound pickup signals corresponding to each of the second directional sound pickup signals, second energy information of the plurality of third directional sound pickup signals, and a second similarity between the plurality of third directional sound pickup signals.

[0152] The sound source positioning module 140 includes a sound source positioning sub-module.

[0153] The sound source positioning sub-module is configured to determine second positioning information of a sound source corresponding to a speech signal in the preset audio frame signal according to the fourth directional sound pickup signal.

[0154] The fourth signal determination sub-module includes a second clustering sub-module.

[0155] The second clustering sub-module is configured to determine the second directional sound pickup signal as a fourth directional sound pickup signal when at least two of the following conditions are met: the second probability of each of the plurality of third directional sound pickup signals corresponding to the second directional sound pickup signal is greater than or equal to a second probability threshold, the second energy information of each of the plurality of third directional sound pickup signals corresponding to the second directional sound pickup signal is greater than or equal to a second energy threshold, and the second similarity between the plurality of third directional sound pickup signals corresponding to the second directional sound pickup signal is greater than or equal to a second similarity threshold.

[0156] Exemplarily, the second beam parameter comprises at least a second angle change step and a first angle difference threshold; the second angle change step is smaller than a first angle change step comprised in the first beam parameter; the third signal determination submodule comprises a second angle determination submodule and a second directional sound pickup submodule.

[0157] The second angle determination submodule is configured to determine a plurality of second scanning angles corresponding to the first scanning angle according to the first scanning angle corresponding to the second directional sound pickup signal and the second angle change step; an absolute value of an angle difference between each of the second scanning angles and the first scanning angle is smaller than or equal to the first angle difference threshold.

[0158] The second directional sound pickup submodule is configured to perform directional sound pickup on the second directional sound pickup signal according to the second scanning angle, to obtain a third directional sound pickup signal of the second directional sound pickup signal corresponding to the second scanning angle.

[0159] Exemplarily, the audio acquisition module 110 comprises an initial signal acquisition submodule, a voice activity detection submodule and a preset signal determination submodule.

[0160] The initial signal acquisition submodule is configured to acquire an initial audio frame signal.

[0161] The voice activity detection submodule is configured to perform voice activity detection on the initial audio frame signal based on a preset voice activity detection model, to obtain a third probability that a voice signal exists in the initial audio frame signal.

[0162] The preset signal determination submodule is configured to determine the initial audio frame signal as a preset audio frame signal when the third probability is greater than or equal to a third probability threshold.

[0163] It should be noted that, for the convenience and brevity of description, the specific working processes of the above-described apparatuses and modules and units can refer to the corresponding processes in the foregoing method embodiments, which will not be described herein.

[0164] The methodologies of the present application can be employed in a variety of computer system contexts. For example, the methodologies can be employed in a personal computer, a server computer, a handheld device or portable device, a tablet device, a multiprocessor system, a microprocessor-based system, a set top box, programmable consumer electronics, network PC, minicomputer, mainframe computer, distributed computing environments that include any of the above systems or devices, or the like. The present application can be described in the general context of computer-executable instructions, such as program modules, being executed by a computer. Generally, program modules include routines, programs, objects, components, data structures, and the like, that perform particular tasks or implement particular abstract data types. The present application can also be practiced in distributed computing environments where tasks are performed by remote processing devices that are linked through a communications network. In a distributed computing environment, program modules can be located in local and remote computer storage media including memory storage devices.

[0165] Exemplarily, the above method and device can be implemented in the form of a computer program, which can be run on an electronic device or a server to control the electronic device and thus perform sound source positioning. Exemplarily, the electronic device can include a near-eye display device, a wearable device, a terminal device, and the like, which are not limited herein. The near-eye display device can include AR glasses, VR glasses, MR glasses, an AR helmet, a VR helmet, an MR helmet, and the like, which are not limited herein. The wearable device includes a smart watch, a smart ring, a smart bracelet, and the like, which are not limited herein. The terminal device includes a mobile phone, a television, a computer, and the like, which are not limited herein. The server can be a separate server, or a cloud server providing cloud services, a cloud database, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content distribution networks, and basic cloud computing services such as big data and artificial intelligence platforms.

[0166] Please refer to Figure 3 , Figure 3 is a structural schematic block diagram of an electronic device provided by an embodiment of the present application.

[0167] As Figure 3 shown, the electronic device includes a memory and a processor. The memory and the processor can be connected through a system bus. The memory can include a storage medium and an internal memory.

[0168] The storage medium can store an operating system and a computer program. The computer program, when executed, can enable the processor to perform any sound source positioning method.

[0169] The processor is configured to provide computing and control capabilities to support the operation of the entire electronic device.

[0170] The internal memory provides an environment for the running of a computer program in a storage medium, and the computer program, when executed by the processor, can enable the processor to perform any sound source positioning method.

[0171] Those skilled in the art can understand that, Figure 3 The structure shown in the figure is only a block diagram of part of the structure related to the scheme of the present application, and does not constitute a limitation on the near-eye display device to which the scheme of the present application is applied. A specific near-eye display device can include more or fewer components than those shown in the figure, or combine certain components, or have a different arrangement of components.

[0172] It should be understood that the processor can be a central processing unit (CPU), and the processor can also be other general-purpose processors, digital signal processors (DSP), application specific integrated circuits (ASIC), field programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor.

[0173] In one embodiment, the processor is configured to execute the computer program and implement the following steps when executing the computer program:

[0174] obtaining a preset audio frame signal, wherein the preset audio frame signal contains a speech signal;

[0175] directing sound pickup on the preset audio frame signal according to the first beam parameter to obtain a plurality of first directional sound pickup signals corresponding to the preset audio frame signal;

[0176] determining at least one second directional sound pickup signal from the plurality of first directional sound pickup signals according to at least two of a first probability of the presence of a speech signal in the plurality of first directional sound pickup signals, first energy information of the plurality of first directional sound pickup signals, and a first similarity between the plurality of first directional sound pickup signals;

[0177] determining first positioning information of a sound source corresponding to the speech signal in the preset audio frame signal according to the at least one second directional sound pickup signal.

[0178] It should be noted that the skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working process of the above-mentioned sound source positioning can refer to the corresponding process in the foregoing sound source positioning method embodiments, and will not be described here.

[0179] The embodiments of the present application further provide a computer readable storage medium, and the computer readable storage medium stores a computer program. The method realized by the computer program executed by a processor can refer to each embodiment of the sound source positioning method of the present application.

[0180] The computer readable storage medium can be an internal storage unit of the electronic device, such as a hard disk or a memory of the electronic device. The computer readable storage medium can also be an external storage device of the electronic device, such as a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc.

[0181] It should be understood that the terms used in the specification of the present application are only for the purpose of describing specific embodiments and are not intended to limit the present application. As used in the specification and the appended claims of the present application, unless otherwise clearly indicated by the context, the singular forms "a", "an" and "the" are intended to include the plural forms.

[0182] It should also be understood that the term "and / or" used in the specification and the appended claims of the present application means any combination of one or more of the associated listed items and all possible combinations, and includes these combinations. It should be noted that in this document, the terms "include", "contain" or any other variants thereof are intended to cover non-exclusive inclusion, so that the process, method, article or system including a series of elements not only includes those elements, but also includes other elements not explicitly listed or inherent to such process, method, article or system. Without more limitations, the element defined by the statement "including a" does not exclude the presence of other identical elements in the process, method, article or system including the element.

[0183] The above-mentioned serial numbers of the embodiments of the present application are only for description and do not represent the advantages or disadvantages of the embodiments. The above description is only a specific implementation of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art can easily think of various equivalent modifications or replacements within the technical scope disclosed by the present application, and these modifications or replacements should be covered within the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

Claims

1. A method for locating a sound source, characterized in that, include: Acquire a preset audio frame signal, wherein the preset audio frame signal contains a speech signal; Based on the first beam parameters, directional sound pickup is performed on the preset audio frame signal to obtain multiple first directional sound pickup signals corresponding to the preset audio frame signal; At least one second directional pickup signal is determined from the plurality of first directional pickup signals based on at least two of the following: a first probability of the presence of a speech signal among the plurality of first directional pickup signals, first energy information of the plurality of first directional pickup signals, and a first similarity between the plurality of first directional pickup signals. Based on at least one of the second directional pickup signals, the first localization information of the sound source corresponding to the speech signal in the preset audio frame signal is determined; The first beam parameters include a first angle variation step size; The step of performing directional sound pickup on the preset audio frame signal according to the first beam parameters to obtain multiple first directional sound pickup signals corresponding to the preset audio frame signal includes: Based on the preset angle corresponding to the preset audio frame signal and the first angle change step size, determine a plurality of first scanning angles corresponding to the preset angle; Based on the first scanning angle, directional sound pickup is performed on the preset audio frame signal to obtain the first directional sound pickup signal corresponding to the preset audio frame signal at the first scanning angle.

2. The sound source localization method according to claim 1, characterized in that, The step of determining at least one second directional pickup signal from the plurality of first directional pickup signals based on at least two of the following: a first probability of the presence of a speech signal among the plurality of first directional pickup signals, first energy information of the plurality of first directional pickup signals, and a first similarity among the plurality of first directional pickup signals, includes: When a first directional sound pickup signal meets at least two of the following criteria: a first probability greater than or equal to a first probability threshold, a first energy information greater than or equal to a first energy threshold, and a first similarity greater than or equal to a first similarity threshold, the first directional sound pickup signal is determined to be a second directional sound pickup signal.

3. The sound source localization method according to any one of claims 1 to 2, characterized in that, After determining at least one second directional pickup signal from the plurality of first directional pickup signals based on at least two of the following: a first probability of the presence of a speech signal among the plurality of first directional pickup signals, first energy information of the plurality of first directional pickup signals, and a first similarity between the plurality of first directional pickup signals, the method further includes: Based on the second beam parameters, the second directional pickup signal is used for directional pickup to obtain multiple third directional pickup signals corresponding to the second directional pickup signal; the second beam parameters are different from the first beam parameters; A fourth directional pickup signal is determined from at least two of the following: a second probability of the presence of a speech signal among a plurality of third directional pickup signals corresponding to at least one second directional pickup signal; second energy information of the plurality of third directional pickup signals; and a second similarity between the plurality of third directional pickup signals. The step of determining the first localization information of the sound source corresponding to the speech signal in the preset audio frame signal based on at least one of the second directional pickup signals includes: Based on the fourth directional pickup signal, the second location information of the sound source corresponding to the speech signal in the preset audio frame signal is determined.

4. The sound source localization method according to claim 3, characterized in that, The step of determining a fourth directional pickup signal from at least one second directional pickup signal based on at least two of the following: a second probability of the presence of a speech signal among a plurality of third directional pickup signals corresponding to at least one second directional pickup signal; second energy information of the plurality of third directional pickup signals; and a second similarity among the plurality of third directional pickup signals: When there exists a second directional pickup signal that meets at least two of the following conditions: the second probability of each of the corresponding plurality of third directional pickup signals is greater than or equal to a second probability threshold, the second energy information of each of the corresponding plurality of third directional pickup signals is greater than or equal to a second energy threshold, and the second similarity between the corresponding plurality of third directional pickup signals is greater than or equal to a second similarity threshold, the second directional pickup signal is determined to be a fourth directional pickup signal.

5. The sound source localization method according to claim 3, characterized in that, The second beam parameters include at least a second angle change step size and a first angle difference threshold; the second angle change step size is smaller than the first angle change step size included in the first beam parameters; The step involves performing directional sound pickup on the second directional pickup signal according to the second beam parameters to obtain multiple third directional pickup signals corresponding to the second directional pickup signal, including: Based on the first scanning angle corresponding to the second directional pickup signal and the second angle change step size, a plurality of second scanning angles corresponding to the first scanning angle are determined; the absolute value of the angle difference between each second scanning angle and the first scanning angle is less than or equal to the first angle difference threshold. Based on the second scanning angle, the second directional pickup signal is used for directional pickup to obtain a third directional pickup signal corresponding to the second directional pickup signal at the second scanning angle.

6. The sound source localization method according to any one of claims 1 to 2, characterized in that, The acquisition of the preset audio frame signal includes: Obtain the initial audio frame signal; Based on a preset speech activity detection model, speech activity detection is performed on the initial audio frame signal to obtain a third probability that a speech signal exists in the initial audio frame signal. When the third probability is greater than or equal to the third probability threshold, the initial audio frame signal is determined to be a preset audio frame signal.

7. A sound source localization device, characterized in that, The sound source localization device includes: An audio acquisition module is used to acquire a preset audio frame signal, wherein the preset audio frame signal contains a speech signal; The first signal determination module is used to perform directional sound pickup on the preset audio frame signal according to the first beam parameters, and obtain multiple first directional sound pickup signals corresponding to the preset audio frame signal; The second signal determination module is used to determine at least one second directional pickup signal from the plurality of first directional pickup signals based on at least two of the following: a first probability of the presence of a speech signal in the plurality of first directional pickup signals, first energy information of the plurality of first directional pickup signals, and a first similarity between the plurality of first directional pickup signals. The sound source localization module is used to determine the first localization information of the sound source corresponding to the speech signal in the preset audio frame signal based on at least one of the second directional pickup signals; The first beam parameters include a first angle variation step size; The step of performing directional sound pickup on the preset audio frame signal according to the first beam parameters to obtain multiple first directional sound pickup signals corresponding to the preset audio frame signal includes: Based on the preset angle corresponding to the preset audio frame signal and the first angle change step size, determine a plurality of first scanning angles corresponding to the preset angle; Based on the first scanning angle, directional sound pickup is performed on the preset audio frame signal to obtain the first directional sound pickup signal corresponding to the preset audio frame signal at the first scanning angle.

8. An electronic device, characterized in that, The electronic device includes a memory and a processor; The memory is used to store computer programs; The processor is configured to execute the computer program and, in executing the computer program, implement the steps of the sound source localization method as described in any one of claims 1 to 6.

9. A computer-readable storage medium storing a computer program thereon, characterized in that, When the computer program is executed by the processor, it implements the steps of the sound source localization method as described in any one of claims 1 to 6.

Citation Information

Patent Citations

  • Audio processing method, audio processing device, system and medium

    CN111833901A

  • Reception process recording method, related device, server, system and storage medium

    CN116320886A