Sound source orientation method and apparatus thereof, sound source separation and tracking method and chip

CN115825853BActive Publication Date: 2026-09-11SHENZHEN SYNSENSE TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202310109065.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-02-14
Publication Date
2026-09-11
Estimated Expiration
2043-02-14

AI Technical Summary

Technical Problem

[0005]此外,现有声源定位/定向方法大多需要借助奇异值分解(SVD)、子空间(subspace)、波束赋形(beamforming)、广义互相关相位变换等算法来提高声源方法的准确度,这将增加数据处理量,对设备的计算性能要求较高,复杂的计算不仅消耗大量的存储资源和功耗,更难以在低功耗的硬件中实现

Benefits of technology

1)本发明的声源方向估计方案不需要波束成形、子空间等复杂的算法,使用基于事件驱动(event-based)的脉冲神经网络实现声源估计,方法简单、定向容易、功耗低且硬件实现容易。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115825853B_ABST
    Figure CN115825853B_ABST
Patent Text Reader

Abstract

The application discloses a sound source orientation method and device, a sound source separation and tracking method and a chip, and aims at solving the technical problems of complex calculation, poor anti-interference and difficulty in hardware implementation of the existing sound source orientation method. The application obtains a pulse signal of the sound source data to be processed by performing zero-crossing pulse coding on the sound source data to be processed. Direction estimation is performed on the pulse signal obtained through the zero-crossing pulse coding based on a pulse neural network, so that the target sound source direction of the sound source data to be processed is obtained. The method is simple, has good real-time performance and low cost, can be easily implemented in a low-power hardware, and the chip test result is almost consistent with the computer simulation result, so that the method has commercial application value. The application is suitable for the field of brain-like computing.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to a sound source localization method and apparatus, a sound source separation and tracking method and chip, and specifically to a sound source localization method and apparatus, a sound source separation and tracking method and chip based on low power consumption and low cost of spiking neural networks (SNN). Background Technology

[0002] Sound localization is an instinct evolved from biological evolution, enabling the rapid and effective identification of sound sources in noisy or complex environments. With the development of artificial intelligence technology, biomimetic machine vision and machine hearing are finding wide application in cutting-edge fields such as video conferencing, intelligent robots, smart homes, quality video surveillance systems, and the Internet of Things.

[0003] Some existing methods are based on deep learning artificial neural networks (ANN or RNN, etc.) for sound source localization (SSL). However, these techniques lack the internal dynamics of the neural network, are not biomimetic / intelligent enough, and their real-time performance needs improvement. Furthermore, they have high energy consumption and storage requirements, and are mainly used for high-computing-power terminals connected to the network, making them unsuitable for edge computing and IoT scenarios.

[0004] Because the distance and direction between the sound source and each microphone in the microphone array are different, but each microphone in the array may receive the speech signal from that sound source, and because the sound source moves, room reverberation, interference from other sound sources, and noise (including but not limited to environmental noise and / or internal noise of electronic devices) inevitably degrade the quality of the speech signal, speech intelligibility, and the accuracy of sound source localization. Current sound source localization technology is not biomimetic and lacks high sensitivity and robustness. These factors increase the difficulty of sound source localization and reduce its real-time performance, affecting audiovisual effects and degrading the performance of electronic devices that use voice as an interaction method. Therefore, after determining the location of the sound source, it is usually necessary to perform speech signal noise reduction and sound source separation processing.

[0005] In addition, most existing sound source localization / direction methods require algorithms such as singular value decomposition (SVD), subspace, beamforming, and generalized cross-correlation phase transformation to improve the accuracy of sound source methods. This will increase the amount of data processing and place high demands on the computing performance of the equipment. The complex calculations not only consume a lot of storage resources and power consumption, but are also difficult to implement in low-power hardware.

[0006] If a sound source localization scheme can be developed in a biological or biomimetic manner, and is sensitive to the relative delay of incoming signals from different microphones, and can detect sound sources in real time and quickly and effectively identify the location or direction of the sound source, with low consumption of computing or storage resources, and is low-power, low-cost and easy to implement, it will be a major advancement in the commercial application of machine hearing in the field of edge intelligent computing. Summary of the Invention

[0007] To solve or alleviate some or all of the above-mentioned technical problems, the present invention is achieved through the following technical solution: A first type of sound source localization method, the method comprising: performing zero-crossing pulse coding on the sound source data to be processed to obtain a pulse signal of the sound source data to be processed; Based on a pulse neural network, the direction of the pulse signal obtained by zero-crossing pulse coding is estimated to obtain the target sound source direction of the sound source data to be processed.

[0008] In one embodiment, based on the sound source data to be processed received by the microphone, the sound source data received by each microphone is preprocessed and then subjected to zero-crossing pulse coding; wherein, the zero-crossing pulse coding includes: Zero-crossing point detection is performed on the pre-processed sound source data of each microphone to obtain the zero-crossing points in the sound source data to be processed received by each microphone and the time information corresponding to each zero-crossing point. Based on the zero-crossing points in the sound source data to be processed received by each microphone and the time information corresponding to each zero-crossing point, pulse coding is performed to obtain the pulse signal of the sound source data to be processed received by each microphone.

[0009] In one embodiment, zero-crossing point detection is performed on the preprocessed sound source data of each microphone to obtain the zero-crossing points in the sound source data to be processed received by each microphone and the time information corresponding to each zero-crossing point, including: For the pre-processed sound source data of each microphone, a set of multiple target signal points of the pre-processed sound source data of the microphone is determined based on the signal value of each signal point in the pre-processed sound source data of the microphone. Based on the signal values ​​of each signal point in each set of target signal points, the sum of the corresponding signal values ​​of each signal point in each set of target signal points is obtained; The sum of the signal values ​​corresponding to each signal point in each set of target signal points is compared to determine the target signal points with local maxima in each set of target signal points and the time information corresponding to the target signal points. Based on the target signal points with local maxima in each set of target signal points and the time information corresponding to each target signal point, the zero-crossing points in the sound source data to be processed received by the microphone and the time information corresponding to each zero-crossing point are determined.

[0010] In one embodiment, determining a set of multiple target signal points in the preprocessed sound source data of the microphone based on the signal values ​​of each signal point in the preprocessed sound source data includes: By comparing the signal values ​​of each signal point in the preprocessed sound source data of the microphone, the signal points in the preprocessed sound source data of the microphone whose signal values ​​continuously decrease are determined. Based on the time information corresponding to the signal points with continuously decreasing signal values ​​in the preprocessed sound source data of the microphone, the signal points with continuously decreasing signal values ​​in the preprocessed sound source data of the microphone are grouped to obtain multiple target signal point sets.

[0011] In one embodiment, comparing the sum of the signal values ​​corresponding to each signal point in each set of target signal points to determine the target signal point with a local maximum in each set of target signal points includes: For each set of target signal points, the sum of the corresponding signal values ​​of each signal point in the set of target signal points is compared to determine the candidate target signal points in the set of target signal points that have the sum of the initial maximum signal values; Based on the time information corresponding to the candidate target signal points, determine the candidate time period with local maxima in the set of target signal points; The signal point with the maximum sum of signal values ​​is determined by comparing the sum of signal values ​​of each time information within the candidate time period. The signal point with the maximum sum of signal values ​​is then identified as the target signal point with a local maximum in the target signal point set.

[0012] In one embodiment, preprocessing includes channel decomposition of the sound source data to be processed received by each microphone, resulting in the sound source data to be processed received by each microphone being decomposed into multiple frequency channels.

[0013] In one embodiment, the preprocessing further includes activity detection based on the channel components of the sound source data to be processed received by each microphone after channel decomposition, in order to obtain a target frequency; the target frequency is one or more frequencies. The sound source component in the target frequency channel of the sound source data to be processed received by each microphone is determined as the preprocessed sound source data of each microphone.

[0014] In one embodiment, the channel decomposition of the sound source data to be processed received by each microphone includes: filtering the sound source data to be processed received by each microphone using a bandpass filter bank, and dividing the sound source data to be processed received by the microphone into multiple frequency channels.

[0015] In one embodiment, the energy or energy and / or average energy of the channel components after channel decomposition of the sound source data received by each microphone at different frequencies are calculated within the same time window to obtain the target frequency that meets the preset conditions.

[0016] In one type of embodiment, the preset condition is that the energy or energy and / or average energy is greater than or equal to a first threshold; or, the energy or energy and / or average energy is greater than or equal to the first threshold and less than or equal to a second threshold.

[0017] In one embodiment, based on a pulse neural network, the direction of the pulse signal obtained by zero-crossing pulse coding is estimated to obtain the target sound source direction of the sound source data to be processed, including: The pulse signal of the sound source data to be processed is input to the feature extraction module, which extracts features from the pulse signal of the sound source data to be processed to obtain a pulse feature sequence; the feature extraction module is constructed based on a long short-term memory network. The pulse feature sequence is input into the pulse neural network for direction estimation to obtain the target direction of the sound source data to be processed.

[0018] In one embodiment, the sound source data may be replaced with electromagnetic waves and / or seismic waves and / or radar and / or physiological signals, and correspondingly, the microphone may be replaced with a sensor corresponding to electromagnetic waves or seismic waves or radar or physiological signals.

[0019] The first type of sound source directional device includes: an encoding module, used to perform zero-crossing encoding on the sound source data to be processed to obtain a pulse signal of the sound source data to be processed; The estimation module, based on a pulse neural network, performs direction estimation on the pulse signal obtained by zero-crossing pulse coding to obtain the target sound source direction of the sound source data to be processed.

[0020] In one embodiment, the sound source directional device further includes: a preprocessing module, used to preprocess the sound source data received by the microphone to obtain preprocessed sound source data for each microphone; The encoding module performs zero-crossing encoding on the pre-processed sound source data from each microphone to obtain the pulse signal of the sound source data to be processed received by each microphone.

[0021] In one embodiment, the preprocessing module includes: a channel decomposition module, which performs channel decomposition on the sound source data to be processed received by each microphone; An activity detection module, coupled to a channel decomposition module, performs activity detection based on the channel components of the sound source data to be processed received by each microphone after channel decomposition, in order to obtain a target frequency; the target frequency is one or more frequencies. The sound source component in the target frequency channel of the sound source data to be processed received by each microphone is determined as the preprocessed sound source data of each microphone.

[0022] In one embodiment, the energy or energy and / or average energy of the channel components after channel decomposition of the sound source data received by each microphone at different frequencies are calculated within the same time window to obtain the target frequency that meets the preset conditions.

[0023] In one type of embodiment, the preset condition is that the energy or energy and / or average energy is greater than or equal to a first threshold; or, the energy or energy and / or average energy is greater than or equal to the first threshold and less than or equal to a second threshold.

[0024] In one embodiment, the encoding module is used to detect zero-crossing points in the sound source data to be processed received by each microphone and the time information corresponding to each zero-crossing point; Pulse coding is performed based on the zero-crossing points in the sound source data to be processed received by each microphone and the time information corresponding to each zero-crossing point. Based on the signal values ​​of each signal point in the preprocessed sound source data of each microphone, determine the set of multiple target signal points in the preprocessed sound source data of the microphone; Based on the signal values ​​of each signal point in each set of target signal points, the sum of the corresponding signal values ​​of each signal point in each set of target signal points is obtained; The sum of the signal values ​​corresponding to each signal point in each set of target signal points is compared to determine the target signal points with local maxima in each set of target signal points and the time information corresponding to the target signal points. Based on the target signal points with local maxima in each set of target signal points and the time information corresponding to each target signal point, the zero-crossing points in the sound source data to be processed received by the microphone and the time information corresponding to each zero-crossing point are determined.

[0025] In one embodiment, based on the signal values ​​of each signal point in the preprocessed sound source data of the microphone, a set of multiple target signal points of the preprocessed sound source data of the microphone is determined, including: By comparing the signal values ​​of each signal point in the preprocessed sound source data of the microphone, the signal points in the preprocessed sound source data of the microphone whose signal values ​​continuously decrease are determined. Based on the time information corresponding to the signal points with continuously decreasing signal values ​​in the preprocessed sound source data of the microphone, the signal points with continuously decreasing signal values ​​in the preprocessed sound source data of the microphone are grouped to obtain multiple target signal point sets.

[0026] In one embodiment, comparing the sum of the signal values ​​corresponding to each signal point in each set of target signal points to determine the target signal point with a local maximum in each set of target signal points includes: For each set of target signal points, the sum of the corresponding signal values ​​of each signal point in the set of target signal points is compared to determine the candidate target signal points in the set of target signal points that have the sum of the initial maximum signal values; Based on the time information corresponding to the candidate target signal points, determine the candidate time period with local maxima in the set of target signal points; The signal point with the maximum sum of signal values ​​is determined by comparing the sum of signal values ​​of each time information within the candidate time period. The signal point with the maximum sum of signal values ​​is then identified as the target signal point with a local maximum in the target signal point set.

[0027] In one embodiment, the sound source directional device further includes: a feature extraction module, coupled between the encoding module and the pulse neural network, for extracting features from the pulse signals of the sound source data to be processed received by each microphone generated by the encoding module, to obtain a pulse feature sequence; The spiking neural network estimates the direction based on the pulse feature sequence to obtain the target direction of the sound source data to be processed.

[0028] A sound source separation method includes: estimating the sound source direction of the sound source data to be separated using the first type of sound source orientation method as described above, and determining the candidate sound sources corresponding to the sound source data to be separated and the target sound source direction of each candidate sound source; Based on the location of each sound channel in the collected sound source data to be separated and the target sound source direction of each candidate sound source, sound source separation is performed to obtain the sound signal of each candidate sound source. The target sound source is determined from the plurality of candidate sound sources based on the sound signals of each candidate sound source.

[0029] A sound source tracking method includes: determining the target sound source direction of sound source data using the first type of sound source orientation method as described above; and performing sound source tracking based on the target sound source direction of the sound source data.

[0030] The first type of chip includes the first type of sound source directional device as described above.

[0031] The first type of electronic device includes either the first type of sound source directional device as described above, or the first type of chip as described above.

[0032] A preprocessing apparatus is provided for preprocessing sound source data received by a microphone to obtain preprocessed sound source data for each microphone. The preprocessing apparatus includes: a channel decomposition module for performing channel decomposition on the sound source data received by each microphone; and an activity detection module coupled to the channel decomposition module for performing activity detection based on the channel components of the channel-decomposed sound source data received by each microphone to obtain a target frequency; the target frequency is one or more frequencies; and the sound source components of the sound source data received by each microphone in the target frequency channel are determined as the preprocessed sound source data for each microphone.

[0033] In one type of embodiment, the activity detection module performs independent activity detection, joint activity detection, or partial joint activity detection; The independent activity detection is as follows: statistically analyze the energy or average energy of the sound source data to be processed received by each microphone at different frequency channels after channel decomposition, and obtain the initial activity frequency that meets the first preset condition; based on the initial activity frequency corresponding to each microphone, obtain the target frequency; The joint activity detection is as follows: sum the energy of all microphones in different frequency channels, or sum the average signal energy of all microphones in different frequency channels, so as to obtain the target frequency that meets the second preset condition. The local joint activity detection is as follows: grouping the sound source data to be processed received by all microphones to obtain at least two sound source data combinations; each sound source data combination includes at least one sound source data to be processed; calculating the sum of the energy of all sound source data in each sound source data combination at different frequency channels, or calculating the average signal energy of all sound source data in each sound source data combination at different frequency channels, to obtain the activity frequency of each sound source data combination; and determining the target frequency based on the activity frequency of each sound source data combination.

[0034] In one type of embodiment, the first preset condition for the independent activity detection is that the energy or the average energy value is the maximum, or the energy or the average energy value is greater than or equal to a second threshold. The second preset condition for the joint activity detection is that the energy and / or the average energy value is the maximum, or the energy and / or the average energy value is greater than or equal to the second threshold.

[0035] In one type of embodiment, obtaining the target frequency based on the initial activity frequency corresponding to each microphone includes one of the following methods: Based on the frequency of occurrence of each initial activity frequency, at least one initial activity frequency is selected as the target frequency. Based on the magnitude of the frequency values ​​of each initial activity frequency, at least one initial activity frequency is selected as the target frequency; Cluster the initial activity frequencies corresponding to each microphone and select at least one initial activity frequency as the target frequency.

[0036] In one type of embodiment, determining the target frequency based on the active frequency of each combination of sound source data includes one of the following methods: Based on the frequency of occurrence of each activity, at least one activity frequency is selected as the target frequency. Based on the frequency values ​​of each activity frequency, at least one activity frequency is selected as the target frequency; Cluster the active frequencies corresponding to each of the sound source data combinations, and select at least one active frequency as the target frequency.

[0037] In one embodiment, the channel decomposition module includes two or more filter groups for preprocessing the sound source data to be processed received by each microphone in a microphone array composed of two or more microphones; wherein the number of filter groups is equal to or less than the number of microphones in the microphone array, and each filter group is coupled to one microphone in the microphone array. The filter bank performs filtering processing, dividing the sound source data to be processed received by the corresponding microphone into multiple frequency channels.

[0038] In some embodiments, the microphone array is linear, circular, spherical, cross-shaped, or spiral.

[0039] In one embodiment, the microphone array is a circular array containing eight microphones.

[0040] In one embodiment, the sound source data may be replaced with electromagnetic waves and / or seismic waves and / or radar and / or physiological signals, and correspondingly, the microphone may be replaced with a sensor corresponding to electromagnetic waves or seismic waves or radar or physiological signals.

[0041] A preprocessing method is used to preprocess sound source data received by a microphone to obtain sound source data after preprocessing steps for each microphone; the preprocessing method includes: The sound source data to be processed received by each microphone is decomposed into multiple frequency channels. Activity detection is performed on the channel components after channel decomposition of the sound source data received by each microphone to obtain the target frequency; the target frequency is one or more frequencies. The sound source component in the target frequency channel of the sound source data to be processed received by each microphone is determined as the preprocessed sound source data of each microphone.

[0042] In one embodiment, the channel decomposition of the sound source data to be processed received by each microphone includes: filtering the sound source data to be processed received by each microphone using a bandpass filter bank, and dividing the sound source data to be processed received by the microphone into multiple frequency channels.

[0043] In one embodiment, the step of obtaining the target frequency by activity detection based on the channel components after channel decomposition of the sound source data to be processed received by each microphone includes: performing independent activity detection on the channel components after channel decomposition of the sound source data to be processed received by each microphone to obtain the initial activity frequency corresponding to each microphone; wherein, independent activity detection is to separately count the energy or average energy of each microphone in different frequency channels to obtain the initial activity frequency that satisfies the first preset condition. The target frequency is obtained based on the initial activity frequency corresponding to each microphone.

[0044] In one type of embodiment, the first preset condition is that the energy or the average energy value is the maximum, or the energy or the average energy value is greater than or equal to a second threshold.

[0045] In one type of embodiment, obtaining the target frequency based on the initial activity frequency corresponding to each microphone includes one of the following methods: Based on the frequency of occurrence of each initial activity frequency, at least one initial activity frequency is selected as the target frequency. Based on the magnitude of the frequency values ​​of each initial activity frequency, at least one initial activity frequency is selected as the target frequency; Cluster the initial activity frequencies corresponding to each microphone and select at least one initial activity frequency as the target frequency.

[0046] In one embodiment, the step of obtaining the target frequency by activity detection based on the channel components after channel decomposition of the sound source data to be processed received by each microphone includes: performing joint activity detection on the channel components after channel decomposition of the sound source data to be processed received by all microphones. The joint activity detection involves summing the energy of all microphones at different frequency channels, or averaging the signal energy of all microphones at different frequency channels, to obtain the target frequency that meets the second preset condition.

[0047] In one type of embodiment, the second preset condition is that the energy and / or the average energy value is the maximum, or the energy and / or the average energy value is greater than or equal to a second threshold.

[0048] In one embodiment, obtaining the target frequency by activity detection based on the channel components after channel decomposition of the sound source data received by each microphone includes: performing local joint activity detection on the channel components after channel decomposition of the sound source data received by all microphones, wherein the local joint activity detection is: The sound source data received by all microphones is grouped to obtain at least two sound source data combinations; each sound source data combination includes at least one sound source data to be processed. The activity frequency of each sound source data combination is obtained by summing the energy of all sound source data in different frequency channels in each sound source data combination, or by summing the average signal energy of all sound source data in different frequency channels in each sound source data combination. Based on the activity frequency of each of the sound source data combinations, a target frequency is determined; based on the target frequency, the target frequency channel of the sound source data to be processed received by each microphone is determined.

[0049] In one embodiment, the sound source data may be replaced with electromagnetic waves and / or seismic waves and / or radar and / or physiological signals, and correspondingly, the microphone may be replaced with a sensor corresponding to electromagnetic waves or seismic waves or radar or physiological signals.

[0050] The second type of sound source directional device includes the preprocessing device as described above; and... The encoding module, coupled to the preprocessing device, is used to perform pulse encoding on the preprocessed sound source data of each microphone to obtain a pulse signal corresponding to the sound source data to be processed received by each microphone. The estimation module, coupled to the encoding module, estimates the direction of the pulse signal obtained by pulse encoding based on a pulse neural network to obtain the target sound source direction of the sound source data to be processed.

[0051] In one embodiment, the encoding module performs zero-crossing pulse coding.

[0052] In one embodiment, the zero-crossing point of the pre-processed sound source data of each microphone on the target frequency channel is detected, and a pulse is generated at the zero-crossing point; wherein the zero-crossing point is an upward zero-crossing point and / or a downward zero-crossing point; The upward zero-crossing point is the signal point where the signal amplitude changes from negative to positive, and the downward zero-crossing point is the signal point where the signal amplitude changes from positive to negative.

[0053] In one type of embodiment, the encoding module performs the following steps: Zero-crossing point detection is performed on the pre-processed sound source data of each microphone to obtain the zero-crossing points in the sound source data to be processed received by each microphone and the time information corresponding to each zero-crossing point. Based on the zero-crossing points in the sound source data to be processed received by each microphone and the time information corresponding to each zero-crossing point, pulse coding is performed to obtain the pulse signal of the sound source data to be processed received by each microphone.

[0054] The second type of sound source localization method includes the preprocessing method described above, which obtains preprocessed sound source data for each microphone. The pre-processed sound source data from each microphone is pulse-coded to obtain a pulse signal corresponding to the sound source data to be processed received by each microphone. Based on a pulse neural network, the direction of the pulse signal obtained by pulse coding is estimated to obtain the target sound source direction of the sound source data to be processed.

[0055] In one embodiment, the pulse coding is zero-crossing pulse coding.

[0056] The second type of chip includes the preprocessing device as described above, or the second type of sound source directional device as described above.

[0057] The second type of electronic device includes the second type of sound source directional device as described above, or includes the second type of chip as described above.

[0058] Some or all of the embodiments of the present invention have the following beneficial technical effects: 1) The sound source direction estimation scheme of the present invention does not require complex algorithms such as beamforming and subspace. It uses an event-based spiking neural network to achieve sound source estimation. The method is simple, easy to orient, has low power consumption and is easy to implement in hardware.

[0059] 2) The zero-crossing pulse coding method of this invention can effectively capture the phase information required for sound source direction estimation. It performs sound source direction estimation based on relative time delay information, improving the real-time performance, anti-interference capabilities, and accuracy of sound source estimation. Furthermore, robust zero-crossing pulse coding is employed to enhance robustness.

[0060] 3) This invention applies adaptive broadband orientation technology for sound source orientation, that is, it identifies the active frequency components in the signal in each time interval and uses these components for localization. While enhancing real-time performance, it can effectively overcome the problem of unstable speech signals, has strong environmental adaptability, and can be applied to a variety of complex environments.

[0061] 4) The activity detection of the present invention has multiple implementation methods, including independent activity detection, joint activity detection, or partial joint activity detection, which is highly flexible.

[0062] 5) This invention can effectively estimate the DOA of narrowband and broadband audio signals. In addition, the sound source direction estimation scheme of this invention responds very quickly to sudden changes in DOA (e.g., changes in the speaker in a conference room), can quickly output the DOA angle after the switch, and can track the sound source quickly, effectively and accurately when it moves (e.g., the speaker moves).

[0063] 6) The sound source localization technology of this invention achieves good sound source localization results in the chip, and the test results of sound source mutation and tracking using this chip are very similar to the computer simulation results, which can be ignored. The sound source localization technology of this invention can be effectively implemented in hardware and has commercial application value.

[0064] Further beneficial effects will be described in the preferred embodiments.

[0065] The technical solutions / features disclosed above are intended to summarize the technical solutions and features described in the Detailed Embodiments section, and therefore the scope of the description may not be entirely the same. However, these new technical solutions disclosed in this section are also part of the numerous technical solutions disclosed in this invention document. The technical features disclosed in this section, together with the technical features disclosed in the subsequent Detailed Embodiments section and some contents in the drawings not explicitly described in the specification, disclose more technical solutions in a reasonable combination.

[0066] The technical solution formed by combining all the technical features disclosed at any position in this invention is used to support the summary of the technical solution, the modification of the patent document, and the disclosure of the technical solution. Attached Figure Description

[0067] Figure 1 This is a schematic diagram of the direction of arrival in a circular array; Figure 2 This is a schematic flowchart of a sound source localization method provided in a certain embodiment of the present invention; Figure 3 This is a schematic diagram of the pulse sequence of the low-frequency channel after zero-crossing pulse coding provided in an embodiment of the present invention; Figure 4 This is a schematic diagram of the pulse sequence of the high-frequency channel after zero-crossing pulse encoding provided in an embodiment of the present invention; Figure 5 This is a schematic diagram of a spiking neural network; Figure 6 This is a schematic diagram of the array resolution of the linear microphone array provided in an embodiment of the present invention; Figure 7 This is a time-frequency diagram of a speech signal; Figure 8 This is a schematic diagram illustrating the principle of sound source localization in a certain embodiment of the present invention; Figure 9 This is a schematic diagram of sound source localization after the signal received by the microphone in a certain embodiment of the present invention has been preprocessed.

[0068] Figure 10 This is a schematic diagram of independent activity detection provided in an embodiment of the present invention; Figure 11 This is a schematic diagram of the joint activity detection provided in an embodiment of the present invention; Figure 12 This is a schematic diagram of the sound source localization model based on long short-term memory network and spiking neural network provided in an embodiment of the present invention; Figure 13 These are the results of sound source direction simulation tests in the low-frequency channel of this invention; Figure 14 This is a test comparison between the low-power hardware implementation of the neuromorphic chip and the simulation model of the present invention, using the low-frequency channel for sound source localization. Figure 15 These are the results of the sound source direction simulation test under the high-frequency channel of this invention; Figure 16 This is a test comparison between the low-power hardware implementation of the neuromorphic chip and the simulation model of the present invention, using high-frequency channels for sound source localization. Figure 17 This is a schematic flowchart of the sound source signal separation method provided in an embodiment of the present invention; Figure 18 This is a flowchart illustrating the sound source tracking method provided in an embodiment of the present invention; Figure 19 The results are from a test of sound source tracking using a neuromorphic chip implemented in low-power hardware according to this invention. Detailed Implementation

[0069] Since it is impossible to exhaustively describe all alternative solutions, the key points of the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Other technical solutions and details not disclosed in detail below generally belong to technical objectives or features that can be achieved by conventional means in the art, and due to space limitations, they will not be described in detail here.

[0070] Unless it refers to division, the " / " in any position in this invention represents logical "OR". The serial numbers "first", "second", etc., in any position in this invention are merely descriptive distinguishing marks and do not imply an absolute temporal or spatial order, nor do they imply that terms prefixed with such serial numbers necessarily refer to different things than the same terms prefixed with other modifiers.

[0071] This invention describes various key points used to combine into various specific embodiments, which will be incorporated into various methods and products. In this invention, even if a key point is described only when introducing a method / product solution, it means that the corresponding product / method solution also explicitly includes that technical feature.

[0072] The description of the existence or inclusion of a step, module, or feature at any location in this invention does not imply that such existence is exclusive or unique. Those skilled in the art can obtain other embodiments by supplementing the technical solutions disclosed in this invention with other technical means. The embodiments disclosed in this invention are generally for the purpose of disclosing preferred embodiments, but this does not imply that opposite embodiments of the preferred embodiments are excluded by this invention. As long as such opposite embodiments at least solve some technical problem of this invention, they are intended to be covered by this invention. Based on the key points described in the specific embodiments of this invention, those skilled in the art can apply substitution, deletion, addition, combination, or reordering of certain technical features to obtain a technical solution that still follows the concept of this invention. These solutions that do not depart from the technical concept of this invention are also within the protection scope of this invention. Explanation of some important terms and symbols: Neuromorphic (mimicry) chips: These chips are event-driven, performing computations or processing only when an event occurs, achieving ultra-high real-time performance and ultra-low power consumption in their hardware circuitry. Based on type, neuromorphic chips are categorized into those based on analog, digital, or mixed-signal circuits.

[0073] Spiking neural networks (SNNs) are a type of event-driven neuromorphic chip and represent the third generation of artificial neural networks. They possess rich spatiotemporal dynamics, diverse encoding mechanisms, and event-driven characteristics, resulting in low computational cost and low power consumption. Compared to artificial neural networks (ANNs), SNNs are more biomimetic and advanced. Brain-inspired computing or neuromorphic computing based on SNNs outperforms traditional AI chips in terms of performance and computational overhead. It should be noted that this invention does not specifically limit the type of spiking neural network. Any neural network driven by pulse signals or events can be applied to the sound source localization method provided in this invention. Spiking neural networks can be built according to actual application scenarios, such as spiking convolutional neural networks (SCNNs), spiking recurrent neural networks (SRNNs), and long short-term neural networks (LSTMs).

[0074] Direction of Arrival (DoA): The directional angle at which the audio signal from the sound source arrives at the microphone array. For sound sources with different DoAs, the delay of the audio signal arriving at the microphone array varies. Microphone arrays can have different spatial shapes, such as circular, linear, spherical, cross-shaped, and spiral. For example, a circular microphone array will be used as an example. Figure 1 As shown, Figure 1 As shown, understandably Figure 1 The microphone array shown is merely an example, and the embodiments of the present invention do not specifically limit the microphone array. Figure 1 This is a schematic diagram of the direction of arrival (DoA) in a circular array, where the array elements are projected along the DoA as a measure of the relative time at which the signal is received at the array element. Here, the array element refers to a microphone in the microphone array.

[0075] It should be noted that the DoA estimation method of the present invention is not only applicable to audio waves, but also to electromagnetic waves, seismic waves, radar waves and similar waves or one-dimensional waves, in order to find the direction or location of the target.

[0076] Narrowband: The bandwidth of a signal is much lower than its center frequency. For example, the bandwidth of a narrowband signal is 10-100 MHz.

[0077] Narrowband positioning: Narrowband signals have a relatively simple spectrum and can be considered as single-frequency signals. In the narrowband case, located at... Different microphone arrays produce microphones that generate sound from the microphones of different microphone arrays. Given the phase shift of the incident harmonic signal, the effect of the microphone array arrangement on the signal can be determined by... The defined M-dim array response vector a(n) is encoded, where n is a unit norm vector representing the DoA direction of the incident wave, within the unit circle. The DoA vector on the array, where λ is the wavelength and M is the number of microphones in the array. It is the k-th microphone in the array arrangement. Understandably, the array response as a function of the DoA vector is indeed a spatial harmonic signal whose frequency depends on the geometry of the microphone array arrangement. .

[0078] Array resolution: Characterizes the distance at which two targets can be distinguished by a DoA (Domain of Arrival) in the presence of noise. This resolution depends on the array geometry, and more importantly, on the spatial size of the array. For example, the angular resolution of a linear array of size L is... Generally speaking, the larger the array, the better its angular resolution.

[0079] Grating Lobes: While larger arrays can produce better resolution, they can lead to grating lobe problems when the number of array elements and the overall spatial span of the array are limited. When this problem occurs, aliasing occurs in the array response vectors; that is, two different DoA vectors n1 and n2 may have the same array response vector, i.e., a(n1) = a(n2). This makes it impossible to determine the angle of the sound source and thus impossible to distinguish and find the correct DoA. Therefore, for microphone arrays with a limited number of microphones, there is a trade-off between increasing array resolution and avoiding grating lobes. For example, one could determine the appropriate DoA based on... The array resolution of the microphone array is obtained by determining the distance between the microphones in the microphone array, where λ is the wavelength of the audio signal.

[0080] Wideband positioning: Wideband signals have a relatively complex spectrum, containing rich frequency components. Wideband positioning can be seen as a generalization of the narrowband case, meaning it can perform positioning by processing multiple received frequency signals. A fixed array with a given array element configuration can only process signals within a limited frequency range: when the frequency exceeds... When the signal wavelength is very small, especially smaller than the spacing between array elements, the resulting grating lobe effect will limit the positioning effect; when the frequency is lower than... When the signal wavelength is very large, especially when the wavelength is larger than the span of the entire array, the angular resolution of the array will be limited, which will make it impossible to locate a single target with sufficient accuracy in the presence of measurement noise.

[0081] As described in the background section, existing sound source localization methods mainly rely on beamforming technology, where beamforming technology defines... This represents the signal received at the microphone, where M is the number of microphones in the array. Indicates microphone The relative delay time at a given point is a function of the DoA (DoA) of the audio signal n. In beamforming, different DoAs (DoAs) can be obtained by weighting and delaying the accumulated received signals at different microphones (called spatial matched filtering). The DoA with the highest power is then identified, which represents the direction of the target sound source. This can be achieved using algorithms such as Delay and Sum, Minimum Variance Distortionless Response (MVDR), and SRP-PHAT (Supported Response Power Phase Transform). These methods are primarily applied to narrowband signals. In narrowband cases, the input audio signal is concentrated around the carrier frequency. By calculating the energy of the received audio signals at different microphones in the microphone array (e.g., calculating the power after Fourier transform of the received signals from each microphone), the DoA with the highest power is found based on the calculated power, thus determining the direction of the target sound source.

[0082] It is evident that existing sound source localization methods are not only computationally complex and demanding on the computing performance of the equipment, but also consume a large amount of storage resources and power, affecting real-time performance and making them difficult to implement in low-power hardware.

[0083] Therefore, to provide a low-power, low-cost, real-time, and easily hardware-implemented sound source localization scheme, this invention provides a sound source localization method, device, chip, and electronic device. It captures relative delay information from the sound source data to be processed using pulse coding, and estimates the sound source direction based on this relative delay information, thereby improving the accuracy of sound source estimation. Specifically, this invention uses a zero-crossing pulse coding method to capture the delay information required for sound source direction estimation, thus ensuring the accuracy of the sound source direction estimation. Furthermore, it uses a spiking neural network for direction estimation to obtain the target direction of the sound source data to be processed. While ensuring the accuracy of sound source direction estimation, it reduces power consumption and has better robustness and faster processing speed.

[0084] To facilitate understanding of the technical solution of the present invention, the sound source direction finding method, device, chip and electronic device provided by the present invention will be introduced below in conjunction with actual application scenarios.

[0085] To improve the real-time performance of sound source localization, reduce its power consumption and complexity, ensure that the sound source localization method can be easily applied to low-power hardware, and further enhance its localization performance when the sound source switches or / and changes or / and moves, this invention uses a spiking neural network (SNN) for sound source direction estimation, transforming DoA estimation into a classification task for the spiking neural network, which includes at least... Figure 2 The steps shown, wherein Figure 2This is a schematic flowchart of a sound source localization method provided in an embodiment of the present invention.

[0086] In step S100, the sound source data to be processed is subjected to zero-crossing pulse coding to obtain the pulse signal of the sound source data to be processed. Considering the pulse communication mechanism of the spiking neural network, the sound source data to be processed needs to be converted into a set of pulse features in advance.

[0087] In one type of embodiment, the sound source data to be processed can be a speech signal in the time domain or an audio signal in the frequency domain.

[0088] In one type of embodiment, the sound source data to be processed may be a voice signal collected in the current environment, which includes the voice signal of the sound source and / or noise present in the environment.

[0089] Preferably, the sound source data to be processed is a speech signal in the current environment acquired in real time; alternatively, the sound source data to be processed can also be a speech signal in the current environment acquired over a past period of time. The past period of time can be the past 1 second, the past 1 minute, etc., and this embodiment of the invention does not specifically limit this.

[0090] Optionally, speech signals from the surrounding environment can be collected using a microphone array. The microphone array can be, for example,... Figure 1 The circular microphone array shown can also be a linear microphone array, a distributed microphone array, or a cross-shaped microphone array. Each microphone array includes at least one microphone, which can be a noise-canceling microphone; this invention is not limited thereto. Furthermore, the sound source data to be processed can be obtained by filtering, denoising, and performing time-frequency analysis on the acquired speech signal from the current environment; this invention is not limited in this regard.

[0091] Considering that the sound source data acquired in practical applications is a broadband signal, directly using the sound source data for zero-crossing coding may introduce interference, thereby reducing the accuracy of the sound source direction estimation results. Therefore, to improve the accuracy of sound source direction estimation, in one embodiment, the sound source data to be processed undergoes zero-crossing pulse coding after preprocessing. The preprocessing includes channel decomposition, where the broadband signal is decomposed into multiple frequency channels, each containing a narrowband signal with a different frequency range. Specifically, channel decomposition is performed on the sound source data received by each microphone in the microphone array, decomposing the broadband sound source data into multiple narrowband signals, and zero-crossing coding is performed on the narrowband signals of each frequency channel.

[0092] Optionally, the sound source data received by each microphone can be decomposed into channels using a filter bank, which includes, but is not limited to, bandpass filter banks and narrowband filter banks. Furthermore, channel decomposition may also include other modules, such as low-noise amplifiers (LNAs) coupled to the filter bank; this invention is not limiting.

[0093] Considering that zero-crossing coding would be applied to each of the multiple narrowband signals obtained after channel decomposition, resulting in a large data volume, this would increase the processing time for sound source localization and thus reduce its real-time performance. Furthermore, each of the multiple narrowband signals obtained after channel decomposition has a different frequency range, while the sound signal corresponding to the sound source direction has high energy and a specific frequency. That is, the energy of the sound source data received by each microphone in the microphone array is mainly concentrated in one or more frequency ranges. Therefore, by using activity detection, the target frequency channel with the highest energy among the multiple narrowband signals obtained after channel decomposition can be determined, or a target frequency channel with processing energy greater than or equal to a preset energy threshold can be selected from the multiple narrowband signals obtained after channel decomposition, and zero-crossing coding is applied to the narrowband signal of the target frequency channel. Based on this, in order to achieve rapid sound source localization, in one embodiment, the preprocessing includes channel decomposition and activity detection. Specifically, channel decomposition is performed on the sound source data to be processed received by each microphone in the microphone array, decomposing the broadband sound source data to be processed into multiple frequency channels. Activity detection is performed based on the energy of the narrowband signals of the multiple frequency channels obtained after channel decomposition to determine the target frequency channel, and zero-crossing coding is performed on the narrowband signal of the target frequency channel.

[0094] Activity detection can be achieved by selecting the target frequency channel with the highest energy from multiple narrowband signals obtained after channel decomposition, or by selecting the top few (two or more) energy channels from the multiple narrowband signals obtained after channel decomposition as the target frequency channels, or by selecting the target frequency channels whose processing energy is greater than or equal to a preset energy threshold from the multiple narrowband signals obtained after channel decomposition. The preset energy threshold can be any one of the average, median, mode, second maximum, and third maximum energy values ​​of the narrowband signals of the multiple frequency channels; this embodiment of the invention does not specifically limit its application.

[0095] Pulse coding includes frequency coding, temporal coding, burst coding, and population coding. Pulse coding can be viewed as a feature extractor, and the features it generates are processed by a Sound Neural Network (SNN). However, some features extracted by pulse coding from the input sound source signal are incoherent (or correlated) features, such as the short time-frequency transform (STFT) intensity of the input signal. Because incoherent features cannot capture phase information, are not sensitive enough to the relative time of signals received under different microphones, and in practical applications such as conference rooms with reflection propagation (or reverberation propagation), they cause significant frequency domain distortion to the input sound source signal, thus disrupting the extracted features. Therefore, incoherent features of the sound source signal cannot be used for sound source direction estimation.

[0096] Due to these issues, this invention employs a different type of pulse coding—zero-crossing pulse coding. Based on zero-crossing pulse coding, the sound source data to be processed is encoded to obtain a pulse signal of the sound source data. Zero-crossing pulse coding is sensitive to the relative delay of signals transmitted from different microphones and can quickly capture this information. It does not cause significant disturbances in reflection propagation environments, thereby improving the real-time performance and accuracy of sound source localization.

[0097] Specifically, the pulse coding method based on zero-crossing coding includes at least steps SA111~SA113: Step SA111: Zero-crossing point detection is performed on the sound source data to be processed to obtain the zero-crossing points in each of the sound source data to be processed and the time information corresponding to each zero-crossing point.

[0098] In addition, the sound source data to be processed received by each microphone can be preprocessed according to step S100. After preprocessing, the sound source data to be processed received by each microphone is obtained as preprocessed sound source data of each microphone.

[0099] Step SA112 involves performing zero-crossing detection on the sound source data to be processed or the pre-processed sound source data received by each microphone to obtain the zero-crossing points in each sound source data to be processed and the time information corresponding to each zero-crossing point.

[0100] Among them, a zero-crossing point can be a signal point with a signal value of 0, or a signal point where the signal value changes abruptly, such as a signal point where the signal value changes from positive to negative or / and from negative to positive. A zero-crossing point can also be a signal point where the product of the signal values ​​of adjacent signal points is less than 0.

[0101] In one embodiment of the present invention, a zero-crossing point can be determined by multiplying the sum of the signal values ​​at each time point by the signal value at the next adjacent time point.

[0102] Optionally, the sound source data received by each microphone is preprocessed to obtain preprocessed sound source data. Starting from the first time point of the preprocessed sound source data of each microphone, the product of the value of the current time point and the value of the next time point adjacent to the current time point is calculated sequentially. If the product of the value of the current time point and the value of the next time point adjacent to the current time point is less than 0, it is determined that there is at least one zero-crossing point between the current time point and the next time point adjacent to the current time point; if the product of the value of the current time point and the value of the next time point adjacent to the current time point is greater than 0, it is determined that there is no zero-crossing point between the current time point and the next time point adjacent to the current time point; if the product of the value of the current time point and the value of the next time point adjacent to the current time point is equal to 0, and the value of the current time point is not 0, it is determined that the next time point adjacent to the current time point is the zero-crossing point of that microphone.

[0103] Optionally, if the product of the value at the current time point and the value at the next time point adjacent to the current time point is less than 0, then a zero-crossing point with a value of 0 is determined between the current time point and the next time point adjacent to the current time point using a bisection method, and the time corresponding to the zero-crossing point is determined as the time information corresponding to the zero-crossing point.

[0104] In one type of implementation, to achieve sparsity processing and improve real-time performance, only downward zero-crossing points are detected, i.e., signal points where the signal value changes from positive to negative. Alternatively, only upward zero-crossing points are detected, i.e., signal points where the signal value changes from negative to positive. Here, a downward zero-crossing point can be a zero-crossing point within a time period of decreasing signal value, or a signal point where the signal value changes from positive to negative; an upward zero-crossing point can be a zero-crossing point within a time period of increasing signal value, or a signal point where the signal value changes from negative to positive.

[0105] To improve the accuracy of zero-crossing points in the sound source data to be processed, and thus improve the accuracy of the pulse signal, in a certain embodiment of the present invention, upward and / or downward zero-crossing points are detected to determine the zero-crossing points. For example, taking the detection of downward zero-crossing points as an example, candidate sound source data with a rate of change of signal value less than or equal to 0 are selected from the narrowband signal of the target frequency channel. Based on the signal value corresponding to each time point in the candidate sound source data, the cumulative sum of values ​​corresponding to each time point in the candidate sound source data is obtained. Based on the cumulative sum of values ​​corresponding to each time point in the candidate sound source data, a target time point with the maximum cumulative sum is selected from the cumulative sum of values ​​corresponding to each time point in the candidate sound source data. Based on the target time point with the maximum cumulative sum, the zero-crossing point and the corresponding time information are determined. Here, the cumulative sum of values ​​refers to the cumulative sum of values ​​at each time point in the candidate sound source data and all time points before that time point. It can be understood that for the first time point in the candidate sound source data, the cumulative sum of values ​​at the first time point is the value at the first time point.

[0106] Optionally, there may be at least one candidate sound source data, and each candidate sound source data includes a set of monotonically decreasing time points and the signal value corresponding to each time point.

[0107] Optionally, the signal point with the maximum cumulative sum in the candidate sound source data can be determined as the zero-crossing point in the candidate sound source data, and the target time point with the maximum cumulative sum in the candidate sound source data is the time information corresponding to the zero-crossing point.

[0108] Optionally, zero-crossing points and their corresponding time information can be determined from the time period consisting of the target time point with the maximum cumulative sum, the time point before the target time point, and the time point after the target time point in the candidate sound source data.

[0109] Specifically, for each candidate sound source data, the candidate time period containing the zero-crossing point in the candidate sound source data can be determined based on the target time point with the maximum cumulative sum, the time point preceding the target time point, and the time point following the target time point. This candidate time period is then divided into multiple candidate times using a preset time window. Based on the value corresponding to each candidate time, the zero-crossing point within the candidate time period containing the zero-crossing point in the candidate sound source data, as well as the corresponding time information, are determined. Understandably, the length of the preset time window is shorter than the time length between the target time point and the time point preceding the target time point.

[0110] For example, the cumulative sum of values ​​corresponding to each candidate time can be obtained based on the value corresponding to each candidate time in the candidate time period where the zero crossover point is located in the candidate sound source data. The signal point corresponding to the candidate time corresponding to the maximum cumulative sum is determined as the zero crossover point in the candidate sound source data. The candidate time corresponding to the maximum cumulative sum is the time information corresponding to the zero crossover point.

[0111] Alternatively, when using the zero-crossing detection method to detect upward zero-crossing points, multiple monotonically increasing time point sets are selected from the preprocessed sound source data. For each monotonically increasing time point set, the cumulative sum of the values ​​of each time point in the monotonically increasing time point set is obtained based on the value of each time point in the monotonically increasing time point set. Based on the cumulative sum of the values ​​of each time point in the monotonically increasing time point set, a target time point with the minimum cumulative sum is selected from the cumulative sum of the values ​​of each time point in the monotonically increasing time point set. Based on the target time point with the minimum cumulative sum in the monotonically increasing time point set, the zero-crossing point and the corresponding time information are determined.

[0112] Step SA113: Based on the zero-crossing points in each sound source data to be processed and the time information corresponding to each zero-crossing point, pulse coding is performed to obtain the pulse signal of the sound source data to be processed from each microphone.

[0113] In one embodiment, after determining the zero-crossing point of the sound source data to be processed by each microphone in the microphone sequence and the time information corresponding to each zero-crossing point, a pulse can be generated according to the time information corresponding to each zero-crossing point. For example, a pulse can be generated at the zero-crossing point to obtain the pulse signal of the sound source data to be processed by the microphone.

[0114] Optionally, after determining the zero-crossing points of the preprocessed sound source data for each microphone and the corresponding time information, the pulse mode of the zero-crossing points in the preprocessed sound source data is set to 1, and the pulse mode of the non-zero-crossing points in the preprocessed sound source data is set to 0 to generate pulses. The time information corresponding to each zero-crossing point is used as the pulse time. Based on the time information of the zero-crossing points in the preprocessed sound source data, the zero-crossing points in the preprocessed sound source data are arranged on the same time axis to obtain the pulse signal of the preprocessed sound source data for each microphone.

[0115] Optionally, after determining the zero-crossing points in the preprocessed sound source data of each microphone and the corresponding time information of each zero-crossing point, the pulse triggering time point is determined according to the time information corresponding to each zero-crossing point, and a pulse is generated according to the pulse triggering time point to obtain the pulse signal of the preprocessed sound source data of each microphone.

[0116] Optionally, after determining the zero-crossing points in the preprocessed sound source data of each microphone and the corresponding time information, the frequency value corresponding to the zero-crossing point in the preprocessed sound source data is compared with a preset threshold. If the frequency value corresponding to the zero-crossing point in the preprocessed sound source data is greater than the preset threshold, the pulse mode of the zero-crossing point is set to 1; if the frequency value corresponding to the zero-crossing point in the preprocessed sound source data is less than or equal to the preset threshold, the pulse mode of the zero-crossing point is set to 0. In this way, the zero-crossing points in the preprocessed sound source data are arranged on the same time axis according to the chronological order to obtain the pulse signal of the preprocessed sound source data of each microphone.

[0117] In a multi-microphone array, the zero-crossing points of the preprocessed sound source data from each microphone are detected, and zero-crossing pulse coding is performed to obtain the pulse sequence corresponding to that microphone. Since the sound signals corresponding to the sound source arrive at different microphones in the array at different times, the relative delay times between different microphones in the array are different. This difference in relative delay times leads to different time information of the zero-crossing points of the preprocessed sound source data from different microphones, resulting in different delay times for the pulse signals from different microphones. Therefore, zero-crossing pulse coding can capture the delay times between different microphones.

[0118] For example, a circular array containing eight microphones will be used as an example for illustration. Figure 3 and Figure 4 As shown, Figure 3 This is a schematic diagram of the pulse sequence of the low-frequency channel after zero-crossing pulse coding provided in an embodiment of the present invention. Figure 4 This is a schematic diagram of the pulse sequence of the high-frequency channel after zero-crossing pulse encoding, provided in an embodiment of the present invention. Through... Figure 3 and Figure 4 It can be seen that the zero-crossing pulse from Mic 1 is generated at time 12, while the zero-crossing pulse from Mic 2 is generated at time 20. Compared with Mic 1, the zero-crossing pulse from Mic 2 is delayed by 8 units of time.

[0119] Step S200: Based on the pulse neural network, the direction of the pulse signal obtained by zero-crossing pulse coding is estimated to obtain the target sound source direction of the sound source data to be processed.

[0120] Figure 5 This is a schematic diagram of a spiking neural network, which includes an input layer, an intermediate layer, and an output layer. Neurons are present in all three layers, and each intermediate layer includes at least one hidden layer, with multiple neurons in each hidden layer.

[0121] The input layer is used to excite pulses based on the number of pulses in the input pulse signal within a preset clock cycle to obtain a pulse sequence, and then transmits the pulse sequence to the intermediate layer. The intermediate layer is used to excite pulses based on the number of pulses in the received pulse sequence within a preset time period to obtain a new pulse sequence, and then transmits the new pulse sequence to the output layer. The output layer is used to generate a target pulse signal based on the pulse sequence transmitted by the intermediate layer, and to make a direction decision based on the target pulse signal to determine the direction of the sound source corresponding to the target pulse signal.

[0122] Spiking neural networks on chips or hardware cannot function directly (perform accurate inference based on input environmental signals). The neurons and synaptic modules / units are merely hardware implementations; further steps are needed to group and define connections between them, as well as define the weights and corresponding time constants stored in the synaptic circuits. Therefore, pre-training is required to obtain the necessary parameters, such as supervised or unsupervised training, and on-chip or off-chip training. This invention uses supervised training as an example, but is not limited to it. The network configuration parameters obtained through training are mapped to hardware, such as a chip. After receiving signals from the environment, the chip runs its internal spiking neural network to automatically complete the inference process based on the received signals.

[0123] In a preferred embodiment of the present invention, a spiking neural network is trained based on sample sound source data. For example, a pulse coding sequence of a specific frequency is used as the input to the spiking neural network, and the DoA direction is used as the target to train the SNN, thereby obtaining a sound source localization model. The specific frequency is at least one frequency value, such as... f 1 or / and f 2 or / and f 3rd grade f 1 and f 2 and f 3. Not equal. The sample sound source data includes audio signals from each microphone in the microphone array at different DoA directions.

[0124] For each sample sound source data, the audio signals at different microphones in the sample sound source data are decomposed into channels. The audio components (sound source components) on each frequency channel or target frequency channel (e.g., obtained by activity detection as described above) are encoded with zero-crossing pulses to obtain sample pulse signals or pulse sequences. The sample pulse signals can then be input into the input layer of the spiking neural network model. After processing by the spiking neural network, the output layer of the spiking neural network model estimates the sound source direction and outputs the prediction sector corresponding to the sample sound source data.

[0125] Then, based on the sector labels and predicted sectors corresponding to the sample sound source data, the localization loss of the spiking neural network is determined. Based on the localization loss, the configuration parameters of the spiking neural network are adjusted. This process is iterated until the spiking neural network meets the preset convergence condition, at which point the sound source localization model is obtained.

[0126] The configuration parameters of the spiking neural network include one or more of the following parameters: synaptic weights, firing times, thresholds, and decay time constants of the neurons corresponding to the input, intermediate, and output layers of the spiking neural network, etc., but this invention is not limited thereto. The preset convergence condition can be that the localization loss is less than or equal to a preset localization loss threshold, or that the number of iterations is less than or equal to a preset number of iterations threshold.

[0127] In a preferred embodiment, taking a circular microphone array as an example, the microphone array includes M microphones. The entire range of the direction of arrival (DoA) is quantized into η, and these quantized DoA values ​​are used as candidate classification classes. Here, η is the angular oversampling parameter, η = 1, 2, 3, ... The ηM labels correspond to the ηM angular spaces quantized by the DoA. Therefore, the DoA resolution is approximately... .

[0128] In a preferred embodiment, taking a circular array as an example, a spatial coordinate system can be established with the geometric center of the microphone array as the origin. Using this origin as the center and a preset distance as the radius, the array rotates clockwise or counterclockwise. Within this region, a position is selected for each rotation angle resolution, resulting in multiple selected positions. The azimuth angle of each selected position is set as the direction of the sample sound source. The angular resolution can be determined based on the type of microphone array and the number of microphones in the array. For example, for a circular microphone array with 8 microphones, η = 1, the DoA can be quantized into 8 sectors, each with an angular resolution of 45 degrees. Furthermore, the microphone array can also have other shapes, which are not limited by this invention.

[0129] Optionally, to improve the accuracy of the target sound source direction and reduce the influence of grating lobes, the angular resolution can be determined based on the type of microphone array, the number of microphones in the array, and preset oversampling parameters. Preferably, for a circular microphone array with 8 microphones, when the oversampling parameter η = 2, the DoA can be quantized into 2 x 8 = 16 sectors, each with an angular resolution of 22.5 degrees.

[0130] Based on a sound source localization model trained using a spiking neural network, or spiking neural network hardware configured with the same or similar parameters as the trained network, direction estimation is performed on the sound source data to be processed, yielding the target sound source direction. The trained network configuration parameters can be directly deployed to the spiking neural network hardware, or deployed after processing, such as quantization. For each microphone array...

[0131] For example, taking a circular microphone array as an example, the direction of arrival (DoAs) can be divided into sectors according to a preset angular resolution and a preset oversampling parameter. The divided sectors are used as labels, and sample sound source data for each sector is collected. Thus, each sample sound source data corresponds to a sector label. For example, for a circular microphone array with 8 microphones, when the oversampling parameter is 2, DoAs can be quantized into 2 x 8 = 16 sectors, each with an angular resolution of 22.5 degrees.

[0132] In an optional implementation, the localization loss of the spiking neural network is determined based on the sector labels corresponding to the sample sound source data and the predicted sectors. This loss can employ either the mean squared error function (MSE) or the MSE surround loss. With the MSE function, adjacent sectors can be distinguished; that is, the distance between class-0 and class-1 is the same as the distance between class-0 and class-2. With the MSE surround loss, geometrically adjacent sectors can be distinguished. For example, in a circular microphone array, class-0 and class-1 have the same distance as class-0 and class-15 (considering the circular shape of the array), while the distance between class-0 and class-2 is greater.

[0133] In one embodiment of the present invention, considering that it is difficult to train the SNN directly, an ANN can be built based on the parameters of the spiking neural network. By training the ANN, the target parameters of the spiking neural network are obtained. Based on the target parameters of the spiking neural network, the parameters of the SNN in the electronic device are adjusted to obtain the sound source localization model.

[0134] Optionally, during ANN training, to ensure the consistency of parameters between the SNN and ANN in the electronic device, the parameters of the ANN can be synchronized to the SNN in the electronic device after each iteration of training. The sample pulse signal is then input into the spiking neural network of the electronic device. After processing by the spiking neural network, the predicted sector corresponding to the sample sound source data is output. Based on the predicted sector and the sector label corresponding to the sample sound source data, the localization loss of the spiking neural network is determined. The parameters of the ANN are adjusted based on the localization loss, and the adjusted parameters of the ANN are synchronized to the SNN in the electronic device.

[0135] Optionally, the training method described in the applicant's earlier application (Chinese patent application with publication number CN114861892A) can be used directly for training. The contents of that earlier application are incorporated herein by reference in their entirety.

[0136] The sound source localization method provided in this embodiment of the invention estimates the direction of the pulse signal of the sound source data to be processed based on a spiking neural network, thereby obtaining the target direction of the sound source data to be processed. It can reduce power consumption while ensuring the accuracy of the sound source direction estimation, and has better robustness, faster processing speed, and is easy to implement in low-power hardware.

[0137] Considering that the geometry of a microphone array affects its array resolution, and array resolution is related to the accuracy of direction-of-arrival (DOA) estimation, while a linear microphone array with all microphones aligned in a line can achieve good array resolution, the alignment of all microphones results in asymmetrical resolution. For example, ... Figure 6 As shown, Figure 6 This is a schematic diagram illustrating the array resolution of a linear microphone array provided in an embodiment of the present invention. When the sound source is in front of the linear microphone array, the linear microphone array can achieve a high array resolution; however, when the sound source is to the side of the linear microphone array, the array resolution is lower. In other words, the positional relationship between the sound source and the linear microphone array affects the array resolution of the linear microphone array. Similarly, distributed microphone arrays and cross-shaped microphone arrays can achieve good array resolution at certain angles, but they all have angles with low array resolution, which makes the applicability of linear microphones, distributed microphone arrays, and cross-shaped microphone arrays poor. As mentioned above, for a circular microphone array, its array resolution is related to the number of microphones in the circular microphone array and the oversampling parameters, and is not affected by the positional relationship between the sound source and the circular microphone array. Therefore, to ensure the accuracy of sound source directionality, in some embodiments, by means of... Figure 1 The circular microphone array shown collects speech signals from the surrounding environment.

[0138] Furthermore, in practical applications such as voice conferencing in enclosed rooms, microphone arrays cannot use pre-designed waveforms for positioning; they need to use the speaker's voice signal to locate the source or determine the direction of arrival (DOA). However, voice signals are highly unstable, exhibiting various abrupt jumps in the time-frequency domain. This means the speaker's voice signal is not a fixed-frequency harmonic signal; rather, it is caused by the change in the dominant frequency of the sound source over time. Figure 7 The diagram shows the time-frequency representation of the speech signal.

[0139] In a preferred embodiment of the present invention, an adaptive wide-band localization / direction technique is applied to locate the sound source, that is, the active frequency components in the signal are identified in each time interval, and these components are used for localization.

[0140] Specifically, such as Figure 8 As shown, Figure 8 This is a schematic diagram illustrating the principle of sound source localization in a certain embodiment of the present invention. It includes a preprocessor module, a pulse coding module, and a spiking neural network processor, which are sequentially coupled. The preprocessor module preprocesses signals received from multiple microphones, such as through time-domain and / or frequency-domain analysis. Furthermore, the preprocessor module can identify active frequency components, also known as target frequency components (or target frequency channels in other parts of this document). The pulse coding module, coupled to the preprocessor module, performs zero-crossing pulse coding to obtain pulse signals. The spiking neural network processor deploys a spiking neural network, which performs sound source localization based on multiple pulse signals.

[0141] In one embodiment, the preprocessing module includes a channel decomposition module, which includes multiple (two or more) filter groups, each of which is coupled to a different microphone for time-frequency analysis of the audio signals received by each microphone.

[0142] Optionally, the number of filter banks is less than or equal to the number of microphones. Optionally, the number of pulse signals is less than or equal to the number of filter banks. In the embodiments described herein, the number of filter banks is equal to the number of microphones as an example, but the present invention is not limited thereto.

[0143] Figure 9 This is a schematic diagram illustrating the sound source localization after preprocessing the signal received by the microphone in a certain embodiment of the present invention. For example... Figure 9 As shown, the preprocessing module includes multiple (two or more) filter banks and an activity detection module. The activity detection module is coupled to each filter bank and performs activity detection across all microphones, detecting the activity of the signals received by all microphones at different frequencies to identify the target frequency component (also known as the target frequency, activity frequency, or activity frequency component), that is, identifying one or more frequency channels with stronger activity in the signal.

[0144] In this embodiment, each filter bank corresponds to a microphone. The filter bank performs channel decomposition on the signal received by the corresponding microphone, decomposing the sound source data to be processed into multiple frequency channels. The activity detection module performs activity detection based on the energy of the narrowband signal after channel decomposition to determine the target frequency channel.

[0145] The pulse coding module is coupled to the activity detection module. Based on the target frequency, it performs zero-crossing pulse coding on the filtered signals of each filter bank to obtain a pulse signal or pulse sequence corresponding to the signal received by the microphone of each filter bank. In other words, the pulse coding module performs zero-crossing pulse coding on the audio components of the signals received by each microphone at the target frequency. Specifically, the pulse coding module includes at least one zero-crossing pulse coding unit. The input of each zero-crossing pulse coding unit is coupled to the output of the corresponding filter bank via the corresponding activity detection unit. Each zero-crossing pulse coding unit performs zero-crossing pulse coding on the components of each filter bank output at the target frequency channel to generate the corresponding pulse signal or pulse sequence.

[0146] In one embodiment of the invention, other circuitry, such as a low-noise amplifier (LNA), can be coupled between the filter bank and the microphone. The LNA amplifies the input audio with low noise. Furthermore, each parallel channel may include a rectifier coupled after the filter of that channel to rectify the output of the channel filter.

[0147] Furthermore, the filter can be a bandpass filter, a bandstop filter, a narrowband filter, etc. Optionally, channel decomposition of the sound source data to be processed can also be performed by dividing the sound source data to be processed into different frequency ranges using a bandpass filter bank, with each frequency range corresponding to a frequency channel; alternatively, channel decomposition of the sound source data to be processed can also be performed by dividing the sound source data to be processed into multiple narrowbands according to the bandwidth of the sound source data to be processed using a narrowband filter bank, with each narrowband corresponding to a frequency channel.

[0148] For example, taking the number of filter banks as the same as the number of microphones in the microphone array as an illustration, when the filter bank includes 16 parallel bandpass filters (BPFs), where each BPF corresponds to one channel, each filter bank performs channel decomposition on the audio signal of the sound source data to be processed received by the microphone corresponding to the filter bank, obtaining multiple frequency channels. Multiple parallel channels are filtered by frequency band and the signal activity changing over time in different frequency bands is detected. Each channel's BPF retains only the signal matching the center frequency of that channel's BPF. After the filter bank performs time-frequency analysis on the audio signal received by the microphone corresponding to the filter bank, the audio components on the 16 channels are obtained. For example, the audio signal received by microphone Mic 1 is processed by BPF0 (frequency... BPF1 (frequency) ...and thus obtain N1_0, N1_1...

[0149] Optionally, based on the signals in the frequency channels of different frequency ranges output by each filter bank, the activity detection module selects the target frequency channel that meets the preset energy threshold from the signals in the frequency components of different frequency ranges output by each filter bank. Specifically, the filter bank in the preprocessing module performs channel decomposition on the sound source data to be processed received by the corresponding microphone to obtain multiple frequency channels of each sound source data to be processed. Based on the energy of the narrowband signals of the multiple frequency channels obtained after channel decomposition, activity detection is performed to determine the target frequency channel, and zero-crossing encoding is performed on the narrowband signal of the target frequency channel. The audio components in the target frequency channels of each sound source data to be processed are determined as the preprocessed sound source data of each microphone.

[0150] In one embodiment, the activity detection module includes multiple (or more) activity detection units, with one activity detection unit corresponding to each filter bank, and each activity detection unit coupled to its corresponding filter bank. Each activity detection unit independently performs activity detection on the signals decomposed from each filter bank channel to determine the target frequency channel of the sound source data to be processed received by each microphone.

[0151] Figure 10 This is a schematic diagram of independent activity detection provided in an embodiment of the present invention. The activity detection unit performs activity detection on the audio components of the sound source data to be processed by each microphone across multiple frequency channels, and selects one or more (two or more) channels with the highest / relatively high energy in the sound source data received by each microphone as the initial target frequency channel for that microphone.

[0152] In one embodiment, after obtaining the initial target frequency channel of each microphone, the audio components of the sound source data received by each microphone after channel decomposition are pulse-coded on the target frequency channel to improve data sparsity and real-time performance without reducing the accuracy of sound source directionality.

[0153] In one embodiment, the initial target frequency channel of each microphone can be obtained based on the activity detection unit described above. The frequency range of each microphone's initial target frequency channel is then filtered to determine the activity frequency (i.e., the target frequency). Based on the activity frequency, the target frequency channel of each microphone is determined, thereby ensuring that each microphone's target frequency channel has the same center frequency. The activity frequency can be the center frequency of the signal in the frequency channel.

[0154] Optionally, based on the center frequency of the initial target frequency channel of each microphone, the frequency of each center frequency can be counted, and the one or several center frequencies with the highest / relatively high frequency can be determined as the active frequencies. Alternatively, the one or several center frequencies with the largest frequency values ​​can be determined as the active frequencies.

[0155] Optionally, based on the center frequency of the initial target frequency channel of each microphone, one or more center frequencies with the highest / highest frequency values ​​are selected as the active frequency.

[0156] Optionally, the center frequencies of the initial target frequency channels of each microphone are clustered, the initial target frequency channels are divided into at least one cluster, the target cluster with the largest number of elements in the cluster is selected, and the center frequency corresponding to the target cluster is determined as the active frequency.

[0157] In one embodiment, after determining the activity frequency, the center frequency of the initial target frequency channel of each microphone can be compared with the activity frequency. If the center frequency of the initial target frequency channel of the microphone matches the activity frequency, then the initial target frequency channel of the microphone is determined as the target frequency channel of the microphone. If the center frequency of the initial target frequency channel of the microphone does not match the activity frequency, then a target frequency channel whose center frequency matches the activity frequency is selected from the multiple frequency channels of the microphone. Matching can mean that the center frequency of the initial target frequency channel of the microphone is the same as the activity frequency, or that the difference between the center frequency of the initial target frequency channel of the microphone and the activity frequency is less than or equal to a preset frequency difference.

[0158] Considering that individual activity detection is performed on multiple frequency channels of each microphone, it is difficult to ensure that the initial target frequency channels of each microphone have the same center frequency, and subsequent processing steps are required to ensure that the initial target frequency channels of each microphone have the same center frequency, which will increase the preprocessing steps. Based on this, in order to ensure the accuracy of pulse coding and further simplify the operation, in a certain preferred embodiment, the activity detection module performs joint activity detection after multiple filter banks.

[0159] Figure 11 This is a schematic diagram of joint activity detection provided in an embodiment of the present invention. The filter bank performs channel decomposition on the sound source data to be processed received by the corresponding microphone. The activity detection module performs joint activity detection on multiple frequency channels after channel decomposition of all microphone signals to determine the target frequency. That is, the activity detection module performs joint activity detection among all microphones corresponding to the filter bank, identifying the target frequency based on the activity of the audio signals received by all microphones corresponding to the filter bank at different frequencies. Further, the target frequency is one or more frequencies obtained from the joint activity detection that satisfy preset conditions. Based on the activity frequency, the target frequency channel of each filter bank or the microphone corresponding to the filter bank is determined.

[0160] Optionally, joint activity detection can be achieved by summing the energy of all microphones across different frequency channels; in other words, by calculating the sum of the audio signals received by all microphones at different frequencies. The preset condition can be that the cumulative energy sum is greater than or equal to a preset cumulative energy sum threshold, or it can be the maximum cumulative energy sum. The preset threshold can be any one of the average, median, mode, second-highest, third-highest, or fourth-highest values ​​of the energy sums of all microphones across different frequency channels.

[0161] Optionally, joint activity detection can be to statistically analyze the average signal energy of all microphones at different frequency channels; the preset condition can be that the average signal energy is greater than or equal to a preset average threshold, or the preset condition can be the maximum average signal energy.

[0162] Optionally, taking joint activity detection as an example of statistically summing the energy of all microphones at different frequency channels, the activity detection module accumulates the energy of all microphones at different frequencies according to the frequency corresponding to each frequency channel in the multiple frequency channels of each of the sound source data to be processed, and obtains the energy sum corresponding to each frequency. Based on the energy sum corresponding to each frequency, the activity frequency that satisfies the preset condition is determined. Based on the activity frequency, the target frequency channel of each of the sound source data to be processed is determined.

[0163] Specifically, the frequency corresponding to each frequency channel in the multiple frequency channels of the sound source data to be processed received by each microphone can be compared with the active frequency, and the frequency channel in the sound source data to be processed received by each microphone that matches the active frequency can be determined as the target frequency channel of the sound source data to be processed.

[0164] For example, taking a circular microphone array with eight microphones as an example, the system performs channel decomposition on the sound source data received by each microphone based on the filter bank corresponding to that microphone, obtaining the audio components of the sound source data to be processed in multiple frequency channels. The activity detection module performs joint activity detection, selecting the activity frequency that meets the preset conditions based on the sum of the energy of the audio components of the signals received by all microphones in each frequency channel, and obtaining the target frequency channel for each microphone based on the activity frequency.

[0165] For example, for each frequency channel f, via Detect the cumulative energy of all microphones within a window of size w at the given frequency channel f, where M is the number of microphones in the microphone array and t is the discrete time. Let f represent the signal component with center frequency f received by the m-th microphone at time t. After calculating the cumulative energy of all microphones at different frequency channels, the signal is then processed through... , where k is a positive integer greater than or equal to 1 and less than or equal to M.

[0166] Optionally, when k=1, the active frequency with the largest cumulative energy sum is selected from the cumulative energy sums of all microphones in each frequency channel. The frequency channel whose center frequency matches this active frequency among the multiple frequency channels of each microphone is determined as the target frequency channel for each microphone. Zero-crossing pulse coding is performed on the signal in the target frequency channel of each microphone to obtain a pulse sequence. The number of pulse signals in the pulse sequence is the same as the number of microphones in the microphone array. For example, for a circular array containing 8 microphones, if k=1, 8 pulse signals with different delays are obtained.

[0167] Optionally, when k ≥ 2, at least two active frequencies with larger cumulative energy sums are selected from the cumulative energy sums of all microphones in each frequency channel. The frequency channel whose center frequency matches the active frequency among the multiple frequency channels of each microphone is determined as the target frequency channel for each microphone. Zero-crossing pulse coding is performed on the signal in the target frequency channel of each microphone to obtain a pulse sequence. The number of pulse signals in the pulse sequence is equal to the number of microphones in the microphone array * k. For example, when k = 2, 8 × 2 = 16 pulse signals are obtained, of which 8 are pulse signals with different delays related to the first active frequency component, and the other 8 are pulse signals with different delays related to the second active frequency component.

[0168] Considering that when there are many filter banks, or a large number of filters in a filter bank, joint activity detection of multiple frequency channels of all microphones will increase the data processing load of the activity detection module, increase the preprocessing time, and thus affect the real-time performance of the sound source localization method.

[0169] To ensure the real-time performance of the sound source localization method, in one embodiment, the activity detection module includes at least two activity detection units, each activity detection unit being coupled to at least one filter bank. Each activity detection unit performs local joint activity detection on multiple frequency channels output by the at least one filter bank corresponding to the activity detection unit to obtain the activity frequencies in the multiple frequency channels output by the at least one filter bank corresponding to the activity detection unit. Based on the activity frequencies in the multiple frequency channels output by each filter bank, the target frequency is determined, and the target frequency channel of each microphone is obtained based on the target frequency.

[0170] Optionally, the microphones can be grouped according to the number of activity detection units in the activity detection module to obtain at least two microphone combinations. Each microphone combination corresponds to one activity detection unit. For each activity detection unit, the target frequency of each microphone in the corresponding microphone combination is determined according to the center frequencies of multiple frequency channels obtained by channel decomposition of each microphone in the microphone combination, following the joint activity detection method described above. Each microphone combination includes at least one microphone, and the number of microphones in each microphone combination can be the same or different.

[0171] Optionally, after obtaining the target frequency of each microphone, clustering can be performed based on the target frequency of each microphone, dividing the target frequencies of multiple microphones into a cluster, selecting the target cluster with the largest number of elements in the cluster, and determining the target frequency corresponding to the target cluster as the activity frequency.

[0172] Optionally, based on the target frequency of each microphone, the frequency of each target frequency can be counted, and the target frequency with the highest frequency can be determined as the activity frequency.

[0173] Optionally, the maximum target frequency can be determined as the activity frequency based on the target frequency of each microphone.

[0174] In one embodiment of the present invention, a filter bank is used to decompose the sound source data to be processed received by each microphone into multiple frequency channels, and the active frequency component is identified by joint activity detection. By performing zero-crossing pulse coding on the audio components of each microphone in the active frequency component channel, the speed and accuracy of sound source localization are improved, and power consumption and cost are reduced.

[0175] For example, taking one microphone in a microphone array as an example, the process includes at least the following steps: Channel decomposition of the sound source data received by the microphone to obtain audio components (or sound source components) on multiple channels. Based on the target frequency identified by activity detection (a frequency that meets preset conditions), a channel corresponding to the activity frequency is selected from the multiple channels; this channel can be called the target frequency channel. Zero-crossing pulse coding is performed on the audio components on the target frequency channel corresponding to each microphone to obtain a pulse signal or pulse sequence corresponding to the signal received by that microphone. Optionally, the activity frequency is identified based on independent activity detection, joint activity detection, or local activity detection.

[0176] For example, if the activity detection selects an activity frequency, at the beginning t1, The frequency was identified as the active frequency, therefore, across all microphones, based on the BPF4 (center frequency) of each microphone. The audio components on the channel generate pulse signals. Later at t2, The activity frequency was identified based on the BPF5 (center frequency) of each microphone. The audio components on the channel generate pulse signals. These pulses connect together over time to form a pulse sequence.

[0177] Optionally, the preset condition can be one or more frequency channels with the highest sum of energy, or one or more channels with a sum of energy greater than a preset energy threshold. Understandably, the frequency channel that meets the preset condition is the target frequency channel. As described above, whether the preset condition is met can be determined based on a first threshold, a second threshold, and a third threshold. Based on the target frequency channel, the audio component of the signal received by each microphone on that target frequency channel is the effective frequency component.

[0178] When estimating the direction of a sound source based on a spiking neural network, the sound source data to be processed needs to be converted into a set of pulse features, that is, the sound source data to be processed needs to be converted into pulse signals. Then, the direction of the pulse signals is estimated by using a spiking neural network to obtain the target sound source direction of the sound source data to be processed. To improve the accuracy of the sound source direction, in some embodiments, a local maximum method can be used to perform zero-crossing coding on the preprocessed sound source data. Specifically, the zero-crossing coding method based on local maxima includes at least steps SB111~SB115: Step SB111: Based on the signal values ​​of each signal point in the preprocessed sound source data of each microphone, determine multiple target signal point sets of the preprocessed sound source data of the microphone. Optionally, the target signal point set includes a combination of signal points whose signal values ​​continuously decrease within a preset time period.

[0179] In a certain embodiment of the present invention, in order to reduce the amount of data processing, downward zero-crossing points can be extracted from the preprocessed sound source data. For example, signal points with continuously decreasing signal values ​​can be selected from the preprocessed sound source data to obtain a target signal point set. Specifically, the method for determining the target signal point set based on downward zero-crossing points includes: By comparing the signal values ​​of each signal point in the preprocessed sound source data of the microphone, the signal points in the preprocessed sound source data of the microphone whose signal values ​​continuously decrease are determined.

[0180] Based on the time information corresponding to the signal points with continuously decreasing signal values ​​in the preprocessed sound source data of the microphone, the signal points with continuously decreasing signal values ​​in the preprocessed sound source data of the microphone are grouped to obtain multiple target signal point sets.

[0181] In one embodiment, the sound source data to be processed received by each microphone can be decomposed into channels and activity detected according to the above preprocessing method. For the signal values ​​of each signal point in the audio component of the target frequency channel of each microphone, a set of multiple target signal points of the preprocessed sound source data is determined.

[0182] Optionally, the target signal point set includes a combination of signal points whose signal values ​​continuously increase over a preset time period.

[0183] In a certain embodiment of the present invention, in order to reduce the amount of data processing, upward zero-crossing points can be extracted from the preprocessed sound source data. For example, signal points with continuously increasing signal values ​​can be selected from the preprocessed sound source data to obtain a target signal point set. Specifically, the method for determining the target signal point set based on upward zero-crossing points includes: By comparing the signal values ​​of each signal point in the preprocessed sound source data of the microphone, the signal points in the preprocessed sound source data of the microphone whose signal values ​​continuously increase are determined.

[0184] Based on the time information corresponding to the signal points with continuously increasing signal values ​​in the preprocessed sound source data of the microphone, the signal points with continuously increasing signal values ​​in the preprocessed sound source data of the microphone are grouped to obtain multiple target signal point sets.

[0185] Step SB112: Based on the signal values ​​(not absolute values, i.e., they may be positive, negative, or zero) of each signal point in each set of target signal points, obtain the sum of the corresponding signal values ​​of each signal point in each set of target signal points.

[0186] Step SB113: Compare the sum of the signal values ​​corresponding to each signal point in each target signal point set to determine the target signal point with a local maximum value in each target signal point set and the time information corresponding to the target signal point.

[0187] In one embodiment, the absolute values ​​of the sum of the signal values ​​corresponding to each signal point in each set of target signal points can be compared. The absolute value of the sum of the sums of the signal values ​​corresponding to each signal point is determined as the target signal point with a local maximum. The time information corresponding to the target signal point is obtained based on the time point corresponding to the target signal point.

[0188] In one embodiment, the sum of the signal values ​​corresponding to each signal point in each set of target signal points can be compared, and the signal point with the largest sum of signal values ​​can be determined as the target signal point with a local maximum. Based on the time point corresponding to the target signal point, the time information corresponding to the target signal point can be obtained.

[0189] In one embodiment, the sum of the signal values ​​corresponding to each signal point in the target signal point set can be compared, and the signal point with the smallest sum of signal values ​​can be determined as the target signal point with a local maximum. Based on the time point corresponding to the target signal point, the time information corresponding to the target signal point can be obtained.

[0190] For example, taking the determination of the signal point with the maximum sum of signal values ​​as the target signal point with a local maximum as an example, step SB113 includes: For each set of target signal points, the sum of the corresponding signal values ​​of each signal point in the set of target signal points is compared to determine the candidate target signal point in the set of target signal points that has the sum of the initial maximum signal values.

[0191] Based on the time information corresponding to the candidate target signal points, candidate time periods with local maxima are determined in the set of target signal points.

[0192] The signal point with the maximum sum of signal values ​​is determined by comparing the sum of signal values ​​of each time information within the candidate time period. The signal point with the maximum sum of signal values ​​is then identified as the target signal point with a local maximum in the target signal point set.

[0193] Step SB114: Compare the sum of the signal values ​​corresponding to each signal point in each target signal point set to determine the target signal point with a local maximum value in each target signal point set and the time information corresponding to the target signal point.

[0194] Step SB115: Based on the target signal points with local maxima in each set of target signal points and the time information corresponding to each target signal point, determine the zero-crossing points in the sound source data to be processed received by the microphone and the time information corresponding to each zero-crossing point.

[0195] In a certain embodiment of the present invention, pulse coding can be performed based on the zero-crossing points in the preprocessed sound source data and the time information corresponding to each zero-crossing point, and the pulse signal of the sound source data to be processed can be obtained according to the above step SA113.

[0196] The embodiments of the present invention perform zero-crossing pulse coding based on local maxima, which can eliminate zero-crossing electron avalanches caused by noise and improve the quality of pulse signals, thereby ensuring the accuracy of the target sound source method.

[0197] To improve the performance of the spiking neural network and ensure the accuracy of sound source direction estimation, in some embodiments, a long short-term memory network can be used to correct the pulse signal after zero-crossing pulse coding, and the corrected pulse sequence can be input into the spiking neural network for direction estimation to obtain the target sound source direction of the sound source data to be processed, thereby ensuring the accuracy of sound source direction estimation.

[0198] Figure 12 This is a schematic diagram of a structure for sound source localization based on a long short-term memory network and a spiking neural network, provided by an embodiment of the present invention. The step of obtaining the target sound source direction from the sound source data to be processed based on this embodiment may further include: The pulse signal of the sound source data to be processed is input into a preset feature extraction module. The preset feature extraction module extracts features from the pulse signal of the sound source data to be processed, and obtains a pulse feature sequence.

[0199] The pulse feature sequence is input into a pulse neural network for direction estimation to obtain the target direction of the sound source data to be processed.

[0200] The preset feature extraction module is constructed based on a long short-term memory network. When the pulse signal of the sound source data to be processed is input into the feature extraction module, the hidden state is extracted from the pulse signal of the sound source data according to the input gate, output gate, and forget gate in the feature extraction module. Based on the hidden state, the pulse signal of the sound source data to be processed is corrected to generate a pulse feature sequence. In a certain embodiment of the present invention, the feature extraction module can be trained through a supervised implementation method.

[0201] After obtaining the pulse feature sequence, the pulse feature sequence can be input into the spiking neural network. The input layer, intermediate layer and output layer of the spiking neural network are used to estimate the direction of the pulse feature sequence to obtain the target direction of the sound source data to be processed.

[0202] To better illustrate the sound source localization technology provided in the embodiments of the present invention, a preferred embodiment is applied to an application scenario of audio conferencing in a closed room to obtain the target direction of the sound source data to be processed in this application scenario, such as... Figures 13 to 16 As shown.

[0203] in, Figure 13 The results are simulation test results of sound source orientation in the low-frequency channel, showing the effect of the sound source orientation method provided in this embodiment on the low-frequency channel when the DoA of the sound source signal suddenly changes from 90 degrees to -90 degrees. The results are expressed in radians, where 1 radian equals 60°. Figure 13From bottom to top, the diagrams show the signal diagrams of the corresponding sound source data received by each microphone, the direction of the target sound source detected by the algorithm, and the pulse signals after zero-crossing pulse coding corresponding to each microphone.

[0204] Figure 14 This is a test comparison between a low-power hardware-implemented neuromorphic chip and a simulation model using low-frequency channels for sound source localization. Figure 14 From bottom to top, the results show test results of sound source localization using a low-frequency channel in a computer device's sound source localization model and test results of sound source localization using a low-power hardware-implemented chip. The results are expressed in radians, where 1 radian equals 60°. Figure 14 As can be seen, in a real-world environment, the sound source localization technology provided by this invention achieves good sound source localization results in the chip, and the difference compared with the computer simulation results is very small and negligible.

[0205] Figure 15 The results are simulation test results of sound source orientation in the high-frequency channel, showing the effect of the sound source orientation method provided by the embodiment of the present invention in the high-frequency channel when the DoA of the sound source signal suddenly changes from 90 degrees to -90 degrees. The results are expressed in radians, where 1 radian equals 60°. Figure 15 From bottom to top, the diagrams show the signal diagrams of the corresponding sound source data received by each microphone, the direction of the target sound source detected by the algorithm, and the pulse signals after zero-crossing pulse coding corresponding to each microphone.

[0206] Figure 16 This is a comparison of test results between a low-power hardware-implemented neuromorphic chip and a simulation model using high-frequency channels for sound source localization. Figure 16 From bottom to top, the results show test results for sound source localization using a high-frequency channel in a computer device's sound source localization model and test results for sound source localization using a low-power hardware-implemented chip. All results are expressed in radians, with 1 radian equal to 60°. Figure 16 As can be seen, in a real-world environment, the sound source localization technology provided by this invention achieves good sound source localization results in the chip, and the difference compared with the computer simulation results is very small and negligible.

[0207] from Figures 13 to 16 It can be seen that, in both high-frequency and low-frequency channels, the sound source localization technology implemented by the low-power hardware of this invention responds quite quickly to changes in DoA.

[0208] The sound source localization method provided in this invention estimates the direction of the pulse signal of the sound source data to be processed based on a pulse neural network, thereby obtaining the target direction of the sound source data. This method can reduce power consumption while ensuring the accuracy of sound source direction estimation, and has better robustness and faster processing speed. Furthermore, it captures the relative time delay information in the sound source data to be processed through pulse coding, and performs sound source direction estimation based on the relative time delay information, thereby improving the accuracy of sound source estimation. The pulse coding method based on zero crossover points can effectively capture the phase information required for sound source direction estimation, thus ensuring the accuracy of sound source direction estimation.

[0209] To better implement the sound source localization method provided in the embodiments of the present invention, a sound source localization device is provided based on the sound source localization method. The sound source localization device includes: The encoding module is used to perform zero-crossing encoding on the sound source data to be processed to obtain the pulse signal of the sound source data to be processed; The estimation module is used to estimate the direction of the pulse signal obtained by zero-crossing pulse coding based on the pulse neural network, so as to obtain the target sound source direction of the sound source data to be processed.

[0210] In a certain embodiment of the present invention, the encoding module includes: The zero-crossing detection unit is used to perform zero-crossing detection on the pre-processed sound source data of each microphone, and obtain the zero-crossing points in the sound source data to be processed received by each microphone and the time information corresponding to each zero-crossing point. The encoding unit is used to perform pulse encoding based on the zero-crossing points in the sound source data to be processed received by each of the microphones and the time information corresponding to each zero-crossing point, to obtain the pulse signal of the sound source data to be processed received by each of the microphones.

[0211] In one embodiment of the present invention, a zero-crossing point detection unit is configured to, for the preprocessed sound source data of each microphone, determine multiple sets of target signal points in the preprocessed sound source data of the microphone based on the signal values ​​of each signal point in the preprocessed sound source data of the microphone; obtain the sum of the signal values ​​corresponding to each signal point in each set of target signal points based on the signal values ​​of each signal point in each set of target signal points; compare the sum of the signal values ​​corresponding to each signal point in each set of target signal points to determine the target signal points with local maxima in each set of target signal points and the time information corresponding to the target signal points; and determine the zero-crossing points in the sound source data to be processed received by the microphone and the time information corresponding to each zero-crossing point based on the target signal points with local maxima in each set of target signal points and the time information corresponding to each target signal point.

[0212] In one embodiment of the present invention, the zero-crossing detection unit is used to compare the signal values ​​of each signal point in the preprocessed sound source data of the microphone, determine the signal points in the preprocessed sound source data of the microphone whose signal values ​​decrease continuously, and group the signal points in the preprocessed sound source data of the microphone according to the time information corresponding to the signal points in the preprocessed sound source data of the microphone, to obtain multiple target signal point sets.

[0213] In one embodiment of the present invention, the zero-crossing point detection unit is configured to, for each set of target signal points, compare the sum of the signal values ​​corresponding to each signal point in the set of target signal points to determine candidate target signal points in the set of target signal points that have an initial maximum sum of signal values; determine candidate time periods in the set of target signal points that have local maxima based on the time information corresponding to the candidate target signal points; compare the sum of the signal values ​​corresponding to each signal point in the candidate time periods to determine the signal point with the maximum sum of signal values; and determine the signal point with the maximum sum of signal values ​​as the target signal point with a local maxima in the set of target signal points.

[0214] In one embodiment of the present invention, the sound source directional device further includes: The preprocessing module is used to decompose the sound source data to be processed received by each microphone into multiple frequency channels.

[0215] In one embodiment of the present invention, a preprocessing module is used to perform activity detection based on the channel components of the sound source data to be processed received by each microphone after channel decomposition, so as to obtain a target frequency; the target frequency is one or more frequencies; and the sound source components of the sound source data to be processed received by each microphone in the target frequency channel are determined as the preprocessed sound source data of each microphone.

[0216] In one embodiment of the present invention, an estimation module is used to input the pulse signal of the sound source data to be processed to a feature extraction module, and to extract features from the pulse signal of the sound source data to be processed by the feature extraction module to obtain a pulse feature sequence; the feature extraction module is constructed based on a long short-term memory network; the pulse feature sequence is input to the sound source localization model for direction estimation to obtain the target direction of the sound source data to be processed.

[0217] The sound source orientation device provided in this embodiment of the invention estimates the direction of the pulse signal of the sound source data to be processed based on a pulse neural network, thereby obtaining the target direction of the sound source data. This device can reduce power consumption while ensuring the accuracy of the sound source direction estimation, and has better robustness and faster processing speed. Furthermore, it captures the relative time delay information in the sound source data to be processed through pulse coding, and performs sound source direction estimation based on the relative time delay information, thereby improving the accuracy of the sound source estimation. The pulse coding method based on zero crossover points can effectively capture the phase information required for sound source direction estimation, thus ensuring the accuracy of the sound source direction estimation.

[0218] To better implement the sound source localization method provided in the embodiments of the present invention, a sound source signal separation method is provided based on the sound source localization method, specifically, as follows: Figure 17 As shown, Figure 17 This is a schematic flowchart of a sound source signal separation method provided in an embodiment of the present invention. The sound source signal separation method shown includes at least steps S300 to S500: Step S300: Perform sound source direction estimation on the sound source data to be separated, and determine the candidate sound sources corresponding to the sound source data to be separated and the target sound source direction of each candidate sound source.

[0219] In a certain embodiment of the present invention, the sound source direction of the sound source data to be separated can be estimated according to the sound source orientation method in any of the above embodiments, and the candidate sound sources corresponding to the sound source data to be separated and the target sound source direction of each candidate sound source can be determined.

[0220] Step S400: Based on the position of each sound channel of the collected sound source data to be separated and the target sound source direction of each candidate sound source, the sound source is separated to obtain the sound signal of each candidate sound source.

[0221] Candidate sound sources refer to sound sources that may exist in the current environment, estimated from the sound source data to be separated, including target sound sources.

[0222] In certain embodiments of the present invention, there are multiple ways to perform sound source separation processing on the sound source data to be separated, including, for example: (1) The sound source can be separated by the location of each sound channel and the target sound source direction of each candidate sound source by the separation method based on independent subspace analysis, so as to obtain the sound signal of each candidate sound source.

[0223] (2) The sound source can be separated by the position of each sound channel and the target sound source direction of each candidate sound source by the separation method based on non-negative matrix decomposition, so as to obtain the sound signal of each candidate sound source.

[0224] (3) The sound source separation can be performed on the location of each sound channel and the target sound source direction of each candidate sound source by the separation method based on independent vector analysis with auxiliary function optimization, so as to obtain the sound signal of each candidate sound source.

[0225] It should be noted that the above sound source separation processing method is only an illustrative example and does not constitute a limitation on the sound signal processing method provided in the embodiments of the present invention. For example, a separation method based on overdetermined independent vector analysis with auxiliary function optimization can also be used to perform sound source separation processing according to the position of each sound channel of the collected sound source data to be separated and the target sound source direction of each candidate sound source to obtain the sound signal of each candidate sound source.

[0226] Step S500: Based on the sound signals of each candidate sound source, determine the target sound source from multiple candidate sound sources.

[0227] In one embodiment of the present invention, a target sound source can be determined from a plurality of candidate sound sources by evaluating the sound quality scores of the sound signals of each candidate sound source. For example, the target sound source with the highest sound quality score can be selected from a plurality of candidate sound sources based on the sound quality scores.

[0228] Optionally, there are multiple ways to evaluate the quality of the sound signal from each candidate sound source, including, for example: (1) The sound quality score of each candidate sound source can be determined by calculating the signal interference ratio of the sound signal of each candidate sound source.

[0229] (2) The sound quality score of each candidate sound source can be determined by calculating the signal distortion ratio of the sound signal of each candidate sound source.

[0230] (3) The sound quality score of each candidate sound source can be determined by calculating the maximum likelihood ratio of the sound signal of each candidate sound source.

[0231] (4) The sound quality score of each candidate sound source can be determined by calculating the cepstral clustering of the sound signal of each candidate sound source.

[0232] (5) The sound quality score of each candidate sound source can be determined by calculating the frequency-weighted segmented signal-to-noise ratio of the sound signal of each candidate sound source.

[0233] (6) The sound quality score of each candidate sound source can be determined by calculating the speech quality perception evaluation score of the sound signal of each candidate sound source.

[0234] (7) The sound quality score of each candidate sound source can be determined by calculating the kurtosis value of the sound signal of each candidate sound source.

[0235] (8) The sound quality score of each candidate sound source can be determined by calculating the probability score corresponding to the speech feature vector of each candidate sound source's sound signal. The probability score is used to characterize the probability that the sound signal of each candidate sound source is the speech signal of the target sound source.

[0236] It should be noted that the method for evaluating the sound signal quality of each candidate sound source described above is merely illustrative and does not constitute a limitation on the sound signal processing method provided in the embodiments of the present invention. In practical applications, the corresponding sound quality score determination method can be selected based on the computational power of the electronic device in the actual application scenario.

[0237] The sound source signal separation method provided in this embodiment of the invention improves the accuracy of the direction of candidate sound sources by estimating the direction of the sound source data to be separated, and determines the final target sound source by evaluating the values ​​of each candidate sound source, thereby further improving the accuracy of the separated sound source signal and improving the problem of low stability of signal separation.

[0238] To better implement the sound source localization method provided in the embodiments of the present invention, based on the sound source localization method and the application scenario of audio conferencing in a closed room, a sound source tracking method is provided, specifically, as follows: Figure 18 As shown, Figure 18 This is a schematic flowchart of a sound source tracking method provided in an embodiment of the present invention. The sound source tracking method shown includes at least steps S600 to S800: In step S600, sound source data from the conference scene is received via a microphone array.

[0239] Step S700: Estimate the direction of the sound source data to determine the direction of the target sound source corresponding to the sound source data.

[0240] In a certain embodiment of the present invention, the direction of the sound source data can be estimated by the sound source orientation method in any of the above embodiments to determine the direction of the target sound source corresponding to the sound source data.

[0241] Step S800: Adjust the direction parameters of the sound source tracking device according to the direction of the target sound source.

[0242] Among them, the sound source tracking device is used to track target sound sources in an audio conference in a room, including but not limited to cameras, microphones, etc.

[0243] Understandably, the sound source tracking device and the microphone array receiving the sound source are on the same plane.

[0244] In one embodiment of the present invention, the current direction parameters of the sound source tracking device can be obtained, the target direction parameters of the sound source tracking device can be obtained according to the target sound source direction corresponding to the sound source data, and the direction parameters of the sound source tracking device can be adjusted according to the target direction parameters and the current direction parameters. The direction parameters include, but are not limited to, azimuth and pitch angles. For example,... Figure 19 As shown, Figure 19 This is the test result of sound source tracking performed by the neuromorphic chip implemented in low-power hardware according to the present invention. The results are expressed in radians, where 1 radian equals 60°, and 1 radian is 90°. Figure 19 The first image, from top to bottom, represents the target sound source direction of a real sound source at different times. The second and third images are categorized as the sound source tracking effect before and after smoothing. Figure 19 As can be seen, the sound source tracking method provided in this embodiment of the invention can respond quickly according to the direction of the sound source, thereby achieving rapid sound source tracking.

[0245] The sound source tracking method provided in this embodiment of the invention improves the accuracy of the sound source direction by estimating the sound source direction from the sound source data, and quickly adjusts the direction parameters of the sound source tracking device based on the target sound source direction to achieve rapid sound source tracking.

[0246] This invention also provides a chip that utilizes any of the sound source orientation methods described above, or uses a sound source signal separation method.

[0247] The chip is a neuromorphic chip or a neuromorphic chip, meaning it can be developed by simulating the working mechanism of biological neurons. It is typically event-triggered and features low power consumption, low latency response, and no privacy breaches. Existing neuromorphic chips include Intel's Loihi, IBM's TrueNorth, and Synsense's Dynap-CNN, but this invention is not limited to these.

[0248] This invention also provides an electronic device that stores the sound source direction finding method, the sound source signal separation method, or the chip described above.

[0249] Although the invention has been described with reference to specific features and embodiments, various modifications, combinations, and substitutions can be made therein without departing from the invention. The scope of protection of this invention is not limited to the specific embodiments of processes, machines, manufactures, material compositions, apparatuses, methods, and steps described in the specification, and these methods and modules may also be implemented in one or more related, interdependent, cooperative, or upstream / downstream products or methods.

[0250] Therefore, the specification and drawings should be simply regarded as a description of some embodiments of the technical solutions defined by the appended claims, and thus the appended claims should be interpreted in accordance with the principle of the greatest reasonable interpretation, and are intended to cover as much as possible all modifications, variations, combinations or equivalents within the scope of the invention, while avoiding unreasonable interpretations.

[0251] To achieve better technical effects or for the needs of certain applications, those skilled in the art may make further improvements to the technical solution based on this invention. However, even if such improvements / designs are inventive and / or progressive, as long as they rely on the technical concept of this invention and cover the technical features defined in the claims, the technical solution should also fall within the protection scope of this invention.

[0252] The technical features mentioned in the appended claims may have alternative technical features, or the order of certain technical processes or material organization may be rearranged. Those skilled in the art, upon learning of this invention, will readily conceive of these alternative means, or alter the order of the technical processes or material organization, and then employ substantially the same means to solve substantially the same technical problems and achieve substantially the same technical effects. Therefore, even if the claims explicitly define the aforementioned means and / or order, these modifications, alterations, and substitutions should all fall within the scope of protection of the claims based on the principle of equivalents.

[0253] The method steps or modules described in the embodiments disclosed herein can be implemented in hardware, software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the steps and components of each embodiment have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application or design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementations should not be considered outside the scope of protection claimed by this invention.

Claims

1. A method of sound source localization, the method comprising: The method includes: Based on the sound source data to be processed received by the microphone, zero-crossing pulse coding is performed to obtain the pulse signal of the sound source data to be processed; A spiking neural network is used to estimate the direction of the target sound source in the sound source data to be processed by using a pulse signal obtained by zero-crossing pulse coding. The zero-crossing pulse encoding includes: Based on the sound source data to be processed received by each microphone, zero-crossing point detection is performed to obtain the zero-crossing points in the sound source data to be processed received by each microphone and the time information corresponding to each zero-crossing point. Based on the zero-crossing points in the sound source data to be processed received by each microphone and the time information corresponding to each zero-crossing point, pulse coding is performed to obtain the pulse signal of the sound source data to be processed received by each microphone. The zero-crossing point detection includes: For the pre-processed sound source data of each microphone, a set of multiple target signal points of the pre-processed sound source data of the microphone is determined based on the signal value of each signal point in the pre-processed sound source data of the microphone. Based on the signal values ​​of each signal point in each set of target signal points, the sum of the corresponding signal values ​​of each signal point in each set of target signal points is obtained; The sum of the signal values ​​corresponding to each signal point in each set of target signal points is compared to determine the target signal points with local maxima in each set of target signal points and the time information corresponding to the target signal points. Based on the target signal points with local maxima in each set of target signal points and the time information corresponding to each target signal point, the zero-crossing points in the sound source data to be processed received by the microphone and the time information corresponding to each zero-crossing point are determined. The sound source data received by each microphone is preprocessed and then subjected to zero-crossing pulse coding; The preprocessing includes channel decomposition of the sound source data to be processed received by each microphone, decomposing the sound source data to be processed received by each microphone into multiple frequency channels.

2. The sound source localization method according to claim 1, characterized in that, The step of determining a set of multiple target signal points in the preprocessed sound source data of the microphone based on the signal values ​​of each signal point in the preprocessed sound source data of the microphone includes: By comparing the signal values ​​of each signal point in the preprocessed sound source data of the microphone, the signal points in the preprocessed sound source data of the microphone whose signal values ​​continuously decrease are determined. Based on the time information corresponding to the signal points with continuously decreasing signal values ​​in the preprocessed sound source data of the microphone, the signal points with continuously decreasing signal values ​​in the preprocessed sound source data of the microphone are grouped to obtain multiple target signal point sets.

3. The sound source localization method according to claim 1, characterized in that, The step of comparing the sum of the signal values ​​corresponding to each signal point in each set of target signal points to determine the target signal points with local maxima in each set of target signal points includes: For each set of target signal points, the sum of the signal values ​​corresponding to each signal point in the set of target signal points is compared to determine the candidate target signal point in the set of target signal points that has the sum of the initial maximum signal values; Based on the time information corresponding to the candidate target signal points, determine the candidate time period with local maxima in the set of target signal points; The signal point with the maximum sum of signal values ​​is determined by comparing the sum of signal values ​​of each time information within the candidate time period. The signal point with the maximum sum of signal values ​​is then identified as the target signal point with a local maximum in the target signal point set.

4. The sound source localization method according to any one of claims 1-3, characterized in that: in, The zero-crossing point is the upward zero-crossing point and / or the downward zero-crossing point; The upward zero-crossing point is the signal point where the signal amplitude changes from negative to positive, and the downward zero-crossing point is the signal point where the signal amplitude changes from positive to negative.

5. The sound source localization method according to claim 4, characterized in that: Preprocessing also includes activity detection of the channel components of the sound source data to be processed received by each microphone after channel decomposition, in order to obtain the target frequency; The target frequency is one or more frequencies; The sound source component in the target frequency channel of the sound source data to be processed received by each microphone is determined as the preprocessed sound source data of each microphone.

6. The sound source localization method according to claim 5, characterized in that, The channel decomposition of the sound source data received by each microphone includes: For the sound source data to be processed received by each microphone, a bandpass filter bank is used to filter the sound source data received by the microphone and divide the sound source data received by the microphone into multiple frequency channels.

7. The sound source localization method according to claim 5, characterized in that... : Within the same time window, the energy or energy and / or average energy of the channel components after channel decomposition of the sound source data received by each microphone at different frequencies are calculated to obtain the target frequency that meets the preset conditions.

8. The sound source localization method according to claim 7, characterized in that... : The preset condition is that the energy or energy and / or average energy is greater than or equal to a first threshold. Alternatively, the energy or energy and / or average energy is greater than or equal to the first threshold and less than or equal to the second threshold.

9. The sound source localization method according to claim 5, characterized in that: The zero-crossing point of the pre-processed sound source data of each microphone on the target frequency channel is detected, and a pulse is generated at the zero-crossing point.

10. The sound source localization method according to claim 8, characterized in that: The pulse signal of the sound source data to be processed is input to the feature extraction module, and the feature extraction module performs feature extraction on the pulse signal of the sound source data to be processed to obtain a pulse feature sequence. The feature extraction module is constructed based on a long short-term memory network. The pulse feature sequence is input into the pulse neural network for direction estimation to obtain the target direction of the sound source data to be processed.

11. The sound source localization method according to claim 10, characterized in that: The sound source data can be replaced with electromagnetic waves and / or seismic waves and / or radar and / or physiological signals, and correspondingly, the microphone is replaced with a sensor corresponding to electromagnetic waves or seismic waves or radar or physiological signals.

12. A sound source directional device, characterized in that, The sound source directional device includes: The encoding module performs zero-crossing pulse coding on the sound source data to be processed received by the microphone to obtain the pulse signal of the sound source data to be processed; A spiking neural network is used to estimate the direction of the target sound source in the sound source data to be processed by using a pulse signal obtained by zero-crossing pulse coding. The zero-crossing pulse coding performs zero-crossing point detection and pulse coding. Based on the sound source data to be processed received by each microphone, zero-crossing point detection is performed to obtain the zero-crossing points in the sound source data to be processed received by each microphone and the time information corresponding to each zero-crossing point. Based on the zero-crossing points in the sound source data to be processed received by each microphone and the time information corresponding to each zero-crossing point, pulse coding is performed to obtain the pulse signal of the sound source data to be processed received by each microphone. The encoding module is used to detect zero-crossing points in the sound source data to be processed received by each microphone and the time information corresponding to each zero-crossing point; Pulse coding is performed based on the zero-crossing points in the sound source data to be processed received by each microphone and the time information corresponding to each zero-crossing point. Based on the signal values ​​of each signal point in the preprocessed sound source data of each microphone, determine the set of multiple target signal points in the preprocessed sound source data of the microphone; Based on the signal values ​​of each signal point in each set of target signal points, the sum of the corresponding signal values ​​of each signal point in each set of target signal points is obtained; The sum of the signal values ​​corresponding to each signal point in each set of target signal points is compared to determine the target signal points with local maxima in each set of target signal points and the time information corresponding to the target signal points. Based on the target signal points with local maxima in each set of target signal points and the time information corresponding to each target signal point, the zero-crossing points in the sound source data to be processed received by the microphone and the time information corresponding to each zero-crossing point are determined. The sound source directional device further includes: a preprocessing module, used to preprocess the sound source data received by the microphone to obtain preprocessed sound source data for each microphone; The encoding module performs zero-crossing pulse encoding on the pre-processed sound source data from each microphone to obtain the pulse signal of the sound source data to be processed received by each microphone.

13. The sound source directional device according to claim 12, characterized in that: The preprocessing module includes: The channel decomposition module performs channel decomposition on the sound source data to be processed received by each microphone; An activity detection module, coupled to a channel decomposition module, performs activity detection based on the channel components of the sound source data to be processed received by each microphone after channel decomposition, in order to obtain a target frequency; the target frequency is one or more frequencies. The sound source component in the target frequency channel of the sound source data to be processed received by each microphone is determined as the preprocessed sound source data of each microphone.

14. The sound source directional device according to claim 13, characterized in that: Detect the zero-crossing point of the pre-processed sound source data of each microphone on the target frequency channel, and generate a pulse at the zero-crossing point; Wherein, the zero crossing point is the upward zero crossing point and / or the downward zero crossing point; The upward zero-crossing point is the signal point where the signal amplitude changes from negative to positive, and the downward zero-crossing point is the signal point where the signal amplitude changes from positive to negative.

15. The sound source directional device according to claim 13, characterized in that: Within the same time window, the energy or energy and / or average energy of the channel components after channel decomposition of the sound source data received by each microphone at different frequencies are calculated to obtain the target frequency that meets the preset conditions.

16. The sound source directional device according to claim 15, characterized in that... : The preset condition is that the energy or energy and / or average energy is greater than or equal to a first threshold.

17. The sound source directional device according to claim 15, characterized in that: The preset condition is that the energy or energy and / or average energy is greater than or equal to a first threshold and less than or equal to a second threshold.

18. The sound source directional device according to claim 12, characterized in that, Based on the signal values ​​of each signal point in the preprocessed sound source data of the microphone, a set of multiple target signal points in the preprocessed sound source data of the microphone is determined, including: By comparing the signal values ​​of each signal point in the preprocessed sound source data of the microphone, the signal points in the preprocessed sound source data of the microphone whose signal values ​​continuously decrease are determined. Based on the time information corresponding to the signal points with continuously decreasing signal values ​​in the preprocessed sound source data of the microphone, the signal points with continuously decreasing signal values ​​in the preprocessed sound source data of the microphone are grouped to obtain multiple target signal point sets.

19. The sound source directional device according to claim 12, characterized in that, The step of comparing the sum of the signal values ​​corresponding to each signal point in each set of target signal points to determine the target signal points with local maxima in each set of target signal points includes: For each set of target signal points, the sum of the signal values ​​corresponding to each signal point in the set of target signal points is compared to determine the candidate target signal point in the set of target signal points that has the sum of the initial maximum signal values; Based on the time information corresponding to the candidate target signal points, determine the candidate time period with local maxima in the set of target signal points; The signal point with the maximum sum of signal values ​​is determined by comparing the sum of signal values ​​of each time information within the candidate time period. The signal point with the maximum sum of signal values ​​is then identified as the target signal point with a local maximum in the target signal point set.

20. The sound source directional device according to any one of claims 12 to 16, characterized in that, The sound source directional device further includes: The feature extraction module, coupled between the encoding module and the pulse neural network, extracts features from the pulse signals of the sound source data to be processed received by each microphone generated by the encoding module, and obtains a pulse feature sequence. The spiking neural network estimates the direction based on the pulse feature sequence to obtain the target direction of the sound source data to be processed.

21. A method for separating sound sources, characterized in that, The sound source separation method includes: The sound source orientation method according to any one of claims 1 to 11 is used to estimate the sound source direction of the sound source data to be separated, and to determine the candidate sound sources corresponding to the sound source data to be separated and the target sound source direction of each candidate sound source. Based on the location of each sound channel in the collected sound source data to be separated and the target sound source direction of each candidate sound source, sound source separation is performed to obtain the sound signal of each candidate sound source. The target sound source is determined from the plurality of candidate sound sources based on the sound signals of each candidate sound source.

22. A sound source tracking method, characterized in that, The sound source tracking method includes: The target sound source direction of the sound source data is determined by the sound source orientation method according to any one of claims 1 to 11. Sound source tracking is performed based on the target sound source direction of the sound source data.

23. A chip, characterized in that, The chip includes the sound source directional device as described in any one of claims 12 to 20.

24. An electronic device, characterized in that, The electronic device includes a sound source directional device as described in any one of claims 12 to 20, or includes a chip as described in claim 23.

Citation Information

Patent Citations

  • Chip-in-loop proxy training method and device, chip and electronic device

    CN114861892A

  • Sound source detection method, sound source separation method, and apparatus for executing them

    JP2004325127A

  • Sound source localizing / identifying apparatus

    JP2008085472A