Sound source localization method and related device, equipment and storage medium

By extracting and processing the phase characteristics of the audio to be tested in the target device, and combining the microphone array attribute information for feature fusion, the problem of device differences limiting the accuracy of sound source positioning is solved, and the universality and accuracy of sound source positioning between different devices is improved.

CN119199741BActive Publication Date: 2025-05-13IFLYTEK CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411740129.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-29
Publication Date
2025-05-13
Estimated Expiration
2044-11-29

AI Technical Summary

Technical Problem

In the prior art, the optimization of the sound source positioning algorithm is limited by device differences, resulting in a decrease in the accuracy of sound source positioning, especially the lack of universality of sound source positioning among different devices.

Method used

By extracting the phase characteristics of the audio to be tested in the target device, sampling it to a unified target dimension, and feature extraction and fusion are performed in combination with the attribute information of the target microphone array to obtain the target feature, thereby realizing sound source positioning.

Benefits of technology

This method can realize the universality of sound source positioning among different devices, while improving the accuracy of sound source positioning and reducing feature loss due to equipment differences.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119199741B_ABST
    Figure CN119199741B_ABST
Patent Text Reader

Abstract

The present application discloses a sound source localization method and related devices, equipment and storage media, wherein the sound source localization method includes: extracting phase features based on the audio to be tested collected by the target microphone array in the target device; performing feature sampling to the target dimension based on the phase features to obtain the first feature, and performing feature extraction based on the attribute information of the target microphone array in the target device to obtain the second feature of the target dimension; wherein the attribute information at least includes the arrangement mode and the number of array elements of the target microphone array, and the target dimension is a unified feature dimension when performing sound source localization for different devices; based on the first feature and the second feature, the target feature is obtained by fusion; based on the target feature, the sound source localization result of the audio to be tested is obtained. The above scheme can improve the accuracy of sound source localization while achieving the universality of sound source localization for different devices.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of audio processing technology, and in particular to a sound source localization method and related devices, equipment and storage media. Background Art

[0002] Sound source localization is the problem of estimating the position of one or more sound sources relative to a microphone array based on a recorded multi-channel audio signal. It is widely used in ship and vehicle detection, localization of major noise sources in machines, etc.

[0003] In the prior art, sound source localization is usually achieved by using a sound source detection algorithm based on a neural network. However, in actual application scenarios, due to the inevitable differences in the devices used for sound source localization, especially the sound pickup components in the devices, the optimization of the sound source detection algorithm is limited, which reduces the accuracy of sound source localization. In view of this, how to improve the accuracy of sound source localization while achieving universality of sound source localization on different devices has become an urgent problem to be solved. Summary of the invention

[0004] The main technical problem solved by the present application is to provide a sound source localization method and related devices, equipment and storage media, which can improve the accuracy of sound source localization while achieving the universality of sound source localization of different devices.

[0005] In order to solve the above technical problems, the first aspect of the present application provides a sound source localization method, including: extracting phase features based on the audio to be tested collected by a target microphone array in a target device; performing feature sampling to a target dimension based on the phase features to obtain a first feature, and performing feature extraction based on the attribute information of the target microphone array in the target device to obtain a second feature of the target dimension; wherein the attribute information includes at least the arrangement mode and the number of array elements of the target microphone array, and the target dimension is a unified feature dimension when performing sound source localization for different devices; based on the first feature and the second feature, fusing to obtain a target feature; based on the target feature, obtaining a sound source localization result of the audio to be tested.

[0006] In order to solve the above technical problems, the second aspect of the present application provides a sound source localization device, including: an extraction module, a sampling module, a fusion module and a positioning module, the extraction module is used to extract phase features based on the audio to be tested collected by the target microphone array in the target device; the sampling module is used to perform feature sampling to a target dimension based on the phase feature to obtain a first feature, and to perform feature extraction based on the attribute information of the target microphone array in the target device to obtain a second feature of the target dimension; wherein the attribute information includes at least the arrangement mode and the number of array elements of the target microphone array, and the target dimension is a unified feature dimension when performing sound source localization for different devices; the fusion module is used to fuse the first feature and the second feature to obtain the target feature; the positioning module is used to obtain the sound source localization result of the audio to be tested based on the target feature.

[0007] In order to solve the above technical problem, the third aspect of the present application provides an electronic device, including a memory and a processor coupled to each other, the memory stores program instructions, and the processor is used to execute the program instructions to implement the sound source localization method in the above first aspect.

[0008] In order to solve the above technical problem, the fourth aspect of the present application provides a computer-readable storage medium, which stores program instructions that can be executed by a processor, and the program instructions are used to implement the sound source localization method described in the first aspect above.

[0009] The above scheme extracts the phase feature of the audio to be tested based on the audio to be tested collected by the target microphone array in the target device, and phase samples the phase feature to the target dimension, which is a unified feature dimension when locating the sound source of different devices, and extracts features based on the attribute information of the target microphone array in the target device to obtain a second feature in the target dimension, the attribute information of the target microphone array at least includes the arrangement mode and the number of array elements of the target microphone array, and fuses the first feature and the second feature to obtain the target feature, and obtains the sound source localization result of the audio to be tested based on the target feature. Therefore, even if the target microphone array in the target device used for sound source localization has different numbers and arrangements of array elements, resulting in the extracted phase feature of the audio to be tested being located in different feature dimensions, the first feature obtained based on the phase feature is sampled to a unified feature dimension, which can achieve the universality of sound source localization of different devices, and the second feature obtained based on the attribute information of the target microphone array can improve the specificity of the fused target feature, and the phase feature of the audio to be tested lost in the target feature is reduced as much as possible based on the attribute information of the target microphone array, so that the accuracy of the sound source localization result obtained based on the target feature can be improved. Therefore, the accuracy of sound source positioning can be improved while achieving the universality of sound source positioning of different devices. BRIEF DESCRIPTION OF THE DRAWINGS

[0010] Figure 1 It is a flow chart of an embodiment of a sound source localization method of the present application;

[0011] Figure 2 It is a schematic diagram of the framework of an embodiment of the model training process of the sound source localization method of the present application;

[0012] Figure 3 It is a schematic diagram of the framework of an embodiment of a sound source localization device of the present application;

[0013] Figure 4 It is a schematic diagram of the framework of an embodiment of the electronic device of the present application;

[0014] Figure 5 It is a schematic diagram of a framework of an embodiment of a computer-readable storage medium of the present application. DETAILED DESCRIPTION

[0015] The scheme of the embodiment of the present application is described in detail below in conjunction with the drawings of the specification.

[0016] In the following description, for the purpose of explanation rather than limitation, specific details such as specific system structures, interfaces, and technologies are provided to facilitate a thorough understanding of the present application.

[0017] The terms "system" and "network" are often used interchangeably in this article. The term "and / or" in this article is only a description of the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can mean: A exists alone, A and B exist at the same time, and B exists alone. In addition, the fragment " / " in this article generally indicates that the associated objects before and after are in an "or" relationship. In addition, "many" in this article means two or more than two.

[0018] See also Figure 1 , Figure 1 This is a flow chart of an embodiment of the sound source localization method of the present application. Specifically, it may include the following steps:

[0019] Step S10: extracting phase features based on the audio to be tested collected by the target microphone array in the target device.

[0020] In the disclosed embodiments, the phase feature is used to characterize the time difference and phase difference between the audio to be measured and the target microphone array of the target device. The difference can provide information about the direction of the audio source to be measured. Specifically, the phase feature can be extracted based on relative delay estimation, beamforming, signal subspace and other methods. For example, based on the beamforming method, all angle compensation phases are used for each element of the target microphone array to achieve scanning of the target area, and then the weighted summation is performed on each signal, and feature extraction is performed based on the beam output power to obtain the phase feature of the audio source information to be measured.

[0021] In one implementation scenario, the characteristic dimension of the phase feature is determined by the number of array elements of the target microphone array in the target device and the phase dimension of each array element. For example, the number of array elements of the target microphone array is N, and the phase dimension of each array element is D. Therefore, the product of the number of array elements and the phase dimension is taken as the characteristic dimension of the phase feature, that is, N×D. It can be understood that the phase feature is N×D matrix data.

[0022] In one implementation scenario, in order to achieve sound source localization of the audio to be measured, the target microphone array of the target device includes at least three array elements. It should be noted that the specific number and arrangement of the array elements are not limited in this application.

[0023] Step S20: performing feature sampling to a target dimension based on the phase feature to obtain a first feature, and performing feature extraction based on attribute information of a target microphone array in a target device to obtain a second feature of the target dimension.

[0024] In one implementation scenario, feature sampling is performed based on the phase feature to a target dimension to obtain a first feature, and the target dimension is a unified feature dimension when performing sound source localization on different devices. For example, still taking the above embodiment as an example, the feature dimension of the phase feature is , the target dimension is 1, that is Therefore, the feature is downsampled to the target dimension based on the phase feature. In the above scheme, even if the target microphone array in the target device used for sound source localization has different numbers of array elements and arrangements, resulting in the extracted phase features of the audio to be tested being located in different feature dimensions, the first features obtained based on the phase features are sampled to a unified feature dimension, which can achieve the universality of sound source localization of different devices.

[0025] In a specific implementation scenario, when the characteristic dimension of the phase feature is determined by the number of array elements of the target microphone array in the target device and the phase dimension of each array element, the phase dimension of each array element is sampled to a preset phase dimension, and the product of the value of the preset phase dimension and the number of array elements of the target microphone array is a unified preset value when locating the sound source for different devices. The preset value can represent the value of the target dimension in terms of data. Based on the preset phase dimension, the phase feature is feature sampled to the target dimension to obtain the first feature. Still taking the aforementioned embodiment as an example, regarding the characteristic dimension of the phase feature of the audio to be measured collected by the target device A , the target dimension is 1, adjust for ,and 1. For another example, the characteristic dimension of the phase characteristic of the audio to be measured collected by the target device B is , the target dimension is still 1, adjust for ,and 1. The above scheme can achieve the universality of sound source localization of different devices.

[0026] In a specific implementation scenario, the feature sampling of the phase feature to the target dimension is not limited in this application. It can be understood that in different implementation scenarios, the phase feature can be used for feature upsampling, feature downsampling, etc. Specific implementation methods include phase inverse tangent method, phase reference spectrum method, etc. Please refer to the detailed description of the phase feature sampling technology. For the sake of brevity, it will not be repeated here.

[0027] In one implementation scenario, feature extraction is performed based on the attribute information of the target microphone array in the target device to obtain a second feature of the target dimension, and the attribute information at least includes the arrangement mode and the number of array elements of the target microphone array. In the above scheme, the second feature obtained based on the attribute information of the target microphone array can improve the specificity of the target feature obtained by fusion, and the phase feature of the audio to be measured lost in the target feature can be reduced as much as possible based on the attribute information of the target microphone array, so the accuracy of the sound source localization result obtained based on the target feature can be improved.

[0028] In a specific implementation scenario, a feature set is set that is mapped one-to-one with the device ID of the target device. The features in the feature set corresponding to the device ID of the target device can be used to characterize the attribute information of the target microphone array in the target device. Based on the device ID of the target device, a quick selection is made to obtain the features of the attribute information of the target microphone array in the target device. When each feature in the feature set is in the target dimension, the features of the attribute information of the target microphone array in the target device in the feature set are directly used as the second feature. When each feature in the feature set is not in the target dimension, the features of the attribute information of the target microphone array in the feature set are sampled to the target dimension to obtain the second feature. The above scheme can improve the convenience of obtaining the second feature.

[0029] In a specific implementation scenario, as a possible implementation method, a second feature extraction model can be pre-trained. The second feature extraction model can include but is not limited to a network model of an Encoder-Decoder architecture, etc. After obtaining the attribute information of the target microphone array in the target device, the attribute information is input into the second feature extraction model, and the output result of the second feature extraction model is used as the second feature. In order to ensure the recognition accuracy of the second feature extraction model as much as possible, the attribute information of the sample microphone array can be collected, and the real features can be annotated on the sample microphone array. On this basis, the attribute information of the sample microphone array can be processed based on the second feature extraction model to obtain the predicted features of the sample microphone array, so that the network parameters of the second feature extraction model can be adjusted based on the difference between the real features and the predicted features until the second feature extraction model training converges, and the attribute information of the target microphone array can be processed based on the second feature extraction model with converged training to obtain the second feature. It should be noted that the specific processing process of the second feature extraction model can refer to the technical details such as the network model of the Encoder-Decoder architecture, which will not be repeated here.

[0030] Step S30: Based on the first feature and the second feature, a target feature is obtained by fusing.

[0031] In one implementation scenario, the target feature is obtained by fusing the first feature and the second feature, which can improve the specificity of the fused target feature, and the phase feature of the audio to be measured lost in the target feature is reduced as much as possible based on the attribute information of the target microphone array, so the accuracy of the sound source localization result obtained based on the target feature can be improved. Therefore, the accuracy of sound source localization can be improved while achieving the universality of sound source localization of different devices.

[0032] In a specific implementation scenario, since the first feature and the second feature are both sampled to the target dimension, that is, have the same feature size, the first feature and the second feature are feature spliced ​​to obtain the spliced ​​feature as the fusion feature.

[0033] Step S40: obtaining a sound source localization result of the audio to be tested based on the target features.

[0034] In one implementation scenario, the target feature has both a phase feature related to the audio to be measured and a second feature related to the property information of the target microphone array of the target device, thereby improving the accuracy of sound source positioning.

[0035] In a specific implementation scenario, the sound source localization result is predicted by a sound source localization model. It should be noted that the network structure of the sound source localization model is not limited in this application.

[0036] In a specific implementation scenario, the sound source localization model includes at least an encoder integrated with a self-attention mechanism. The self-attention mechanism is a mechanism used in deep learning models, which is particularly effective when processing sequence data. It allows each element of the input sequence to be compared with other elements in the sequence to calculate the representation of the sequence. The self-attention mechanism captures the dependencies between any positions by calculating the weights of the query, key, and value.

[0037] In a specific implementation scenario, the sound source localization model is trained based on several sample audio sets, and any sample audio set includes: each sample audio obtained by collecting the sound source at the same sample position after different array elements in the candidate microphone array in the same candidate device are turned off. For example, the candidate microphone array in candidate device A has n array elements, namely array element 1, array element 2... array element n. The sample audio set of candidate device A includes sample audio collected by turning off array element 1, sample audio collected by turning off array element 2... sample audio collected by turning off array element n, and the sample audio is located at the same sample position. It can be understood that the sample audio collected in the same sample audio set can be the same sample audio or different sample audio, which is not limited in this application. The above scheme trains the sound source localization model based on the sample audio set including each sample audio collected by the sound source at the same sample position after different array elements in the candidate microphone array in the same candidate device are turned off, which can improve the stability of the sound source localization model after training and improve the accuracy of the sound source localization result.

[0038] In a specific implementation scenario, the sound source localization model also includes a testing process. During the training process, different array elements in the candidate microphone array in the same candidate device are randomly turned off to obtain a sample audio set for training the sound source localization model. During the testing process, different array elements in the candidate microphone array in the same candidate device are sequentially turned off to obtain a sample audio set for testing the sound source localization model, thereby improving the positioning quality of the sound localization model training and accelerating the convergence speed of the sound source localization model.

[0039] In a specific implementation scenario, the sound source localization result includes an azimuth angle of the audio to be measured, and the position information of the audio to be measured can be obtained based on the azimuth angle.

[0040] In a specific implementation scenario, the specific training process of the sound source localization model includes extracting features based on each sample audio in the same sample audio set, obtaining sample target features of each sample audio, making predictions based on the sample target features, obtaining predicted azimuths of the sound source, and adjusting the network parameters of the sound source localization model based on the predicted azimuths corresponding to each sample audio in the same sample audio set and the sample azimuths marked by the corresponding sample audio set. The above scheme obtains the sound source localization result of the audio to be tested based on the sound source localization model that has converged in training, which can improve the accuracy of sound source localization and improve the versatility of sound source localization.

[0041] See also Figure 2 , Figure 2 Schematic diagram of the framework of an embodiment of the sound source localization method model training process of the present application. Figure 2 As shown, in a specific implementation scenario, each candidate device that collects a number of sample audio sets includes at least a target device. For example, a sound source localization model is trained based on the sample audio sets corresponding to candidate device 1, candidate device 2...candidate device m. Specifically, candidate device 1 corresponds to sample audio set 1, candidate device 2 corresponds to sample audio set 2, and candidate device m corresponds to sample audio set m. Then, based on the sound source localization model that has converged in the training, the sound source localization result of the audio to be tested is obtained. The target device for collecting the audio to be tested can be any candidate device among device 1, device 2...device m. It can be understood that the same type of devices have the same microphone array, that is, the number of array elements in the microphone array is the same as the arrangement of the array elements.

[0042] Please continue to refer to Figure 2 In a specific implementation scenario, the sample phase features of each sample audio in the same sample audio set are extracted, and feature sampling is performed to the target dimension based on the sample phase features to obtain the first sample feature, and feature extraction is performed based on the attribute information of the candidate microphone array in the candidate device when any sample audio is collected to obtain the second sample feature of the target dimension. The first sample feature and the second sample feature of the same sample audio are fused to obtain the sample target feature of the corresponding sample audio. In the above scheme, the sample target features have a consistent target dimension, so the sound source localization model improves the accuracy of sound source localization while achieving the universality of sound source localization of different devices.

[0043] In a specific implementation scenario, based on the predicted azimuths corresponding to each sample audio in the same sample audio set, the final azimuth of the corresponding sample audio set is determined, and based on the difference between the final azimuth and the sample azimuth of the same sample audio set, the network parameters of the sound source localization model are adjusted.

[0044] In a specific implementation scenario, the distribution of the predicted azimuths corresponding to each sample audio in the same sample audio set is counted, and the predicted azimuth with the largest distribution in the distribution is used as the final azimuth of the corresponding sample audio set. For example, the final azimuth is determined based on a voting mechanism. For details, please refer to the technical details of the voting mechanism technology. For the sake of brevity, it will not be repeated here.

[0045] In a specific implementation scenario, the network parameters of the sound source localization model are adjusted based on the difference between the final azimuth and the sample azimuth of the same sample audio set and a preset loss function. The specific type of the loss function is not limited in this application.

[0046] The above scheme extracts the phase feature of the audio to be tested based on the audio to be tested collected by the target microphone array in the target device, and phase samples the phase feature to the target dimension, which is a unified feature dimension when locating the sound source of different devices, and extracts features based on the attribute information of the target microphone array in the target device to obtain a second feature in the target dimension, the attribute information of the target microphone array at least includes the arrangement mode and the number of array elements of the target microphone array, and fuses the first feature and the second feature to obtain the target feature, and obtains the sound source localization result of the audio to be tested based on the target feature. Therefore, even if the target microphone array in the target device used for sound source localization has different numbers and arrangements of array elements, resulting in the extracted phase feature of the audio to be tested being located in different feature dimensions, the first feature obtained based on the phase feature is sampled to a unified feature dimension, which can achieve the universality of sound source localization of different devices, and the second feature obtained based on the attribute information of the target microphone array can improve the specificity of the fused target feature, and the phase feature of the audio to be tested lost in the target feature is reduced as much as possible based on the attribute information of the target microphone array, so that the accuracy of the sound source localization result obtained based on the target feature can be improved. Therefore, the accuracy of sound source positioning can be improved while achieving the universality of sound source positioning of different devices.

[0047] See also Figure 3 , Figure 3 Schematic diagram of the framework of an embodiment of the sound source localization device 30 of the present application. Figure 3As shown, the sound source localization device 30 includes: an extraction module 31, a sampling module 32, a fusion module 33 and a positioning module 34, the extraction module 31 is used to extract the phase feature based on the audio to be tested collected by the target microphone array in the target device; the sampling module 32 is used to perform feature sampling to the target dimension based on the phase feature to obtain the first feature, and to perform feature extraction based on the attribute information of the target microphone array in the target device to obtain the second feature of the target dimension; wherein the attribute information at least includes the arrangement mode and the number of array elements of the target microphone array, and the target dimension is a unified feature dimension when performing sound source localization for different devices; the fusion module 33 is used to obtain the target feature by fusing the first feature and the second feature; the positioning module 34 is used to obtain the sound source localization result of the audio to be tested based on the target feature.

[0048] In the above scheme, the sound source localization device 30 extracts the phase feature of the audio to be measured based on the audio to be measured collected by the target microphone array in the target device, phase samples the phase feature to the target dimension, the target dimension is a unified feature dimension when performing sound source localization on different devices, and extracts features based on the attribute information of the target microphone array in the target device to obtain a second feature in the target dimension, the attribute information of the target microphone array at least includes the arrangement mode and the number of array elements of the target microphone array, fuses the first feature and the second feature to obtain the target feature, and obtains the sound source localization result of the audio to be measured based on the target feature. Therefore, even if the target microphone array in the target device used for sound source localization has different numbers and arrangements of array elements, resulting in the phase feature of the audio to be measured being extracted and located in different feature dimensions, the first feature obtained based on the phase feature is sampled to a unified feature dimension, which can achieve the universality of sound source localization of different devices, and the second feature obtained based on the attribute information of the target microphone array can improve the specificity of the fused target feature, and the phase feature of the audio to be measured lost in the target feature is reduced as much as possible based on the attribute information of the target microphone array, so that the accuracy of the sound source localization result obtained based on the target feature can be improved. Therefore, the accuracy of sound source positioning can be improved while achieving the universality of sound source positioning of different devices.

[0049] In some disclosed embodiments, the sound source localization device 30 also includes a localization model module (not shown), and the sound source localization result is predicted by the sound source localization model, which is trained based on several sample audio sets. Any sample audio set includes: after different array elements in the candidate microphone array in the same candidate device are turned off, each sample audio is collected from the sound source at the same sample position.

[0050] In some disclosed embodiments, the positioning model module (not shown) also includes a sample target feature extraction module (not shown), which is used to perform feature extraction based on each sample audio in the same sample audio set to obtain sample target features of each sample audio; the positioning model module (not shown) also includes a sample prediction module (not shown), which is used to perform prediction based on the sample target features to obtain a predicted azimuth of the sound source; the positioning model module (not shown) also includes a parameter adjustment module (not shown), which is used to adjust the network parameters of the sound source localization model based on the predicted azimuths corresponding to each sample audio in the same sample audio set and the sample azimuths marked by the corresponding sample audio set.

[0051] In some disclosed embodiments, the sample target feature extraction module (not shown) also includes a first feature extraction submodule (not shown) for extracting sample phase features of each sample audio in the same sample audio set; the sample target feature extraction module (not shown) also includes a second feature extraction submodule (not shown) for performing feature sampling to a target dimension based on the sample phase features to obtain a first sample feature, and performing feature extraction based on attribute information of a candidate microphone array in a candidate device when any sample audio is collected to obtain a second sample feature of the target dimension; the sample target feature extraction module (not shown) also includes a sample feature fusion module (not shown) for fusing the first sample features and the second sample features of the same sample audio to obtain a sample target feature of the corresponding sample audio.

[0052] In some disclosed embodiments, the parameter adjustment module (not shown) also includes an azimuth determination module (not shown) for determining the final azimuth of the corresponding sample audio set based on the predicted azimuths corresponding to each sample audio in the same sample audio set; the parameter adjustment module (not shown) also includes a difference adjustment module (not shown) for adjusting the network parameters of the sound source localization model based on the difference between the final azimuth and the sample azimuth of the same sample audio set.

[0053] In some disclosed embodiments, the azimuth angle determination module (not shown) also includes a statistical module (not shown) for counting the distribution of predicted azimuth angles corresponding to each sample audio in the same sample audio set; the azimuth angle determination module (not shown) also includes a distribution determination module (not shown) for taking the most distributed predicted azimuth angle in the distribution as the final azimuth angle of the corresponding sample audio set.

[0054] In some disclosed embodiments, the positioning model module (not shown) further includes a configuration module (not shown), the sound source localization model includes at least an encoder integrated with a self-attention mechanism; and / or, each candidate device from which a plurality of sample audio sets are collected includes at least a target device; and / or, the candidate microphone array in the candidate device includes at least three array elements.

[0055] In some disclosed embodiments, the characteristic dimension of the phase feature is determined by the number of array elements of the target microphone array in the target device and the phase dimension of each array element, and the sampling module 32 also includes a dimension preset module (not shown) for sampling the phase dimension of each array element to a preset phase dimension; wherein the product of the value of the preset phase dimension and the number of array elements of the target microphone array is a unified preset value when locating the sound source for different devices; the sampling module 32 also includes a dimension sampling module (not shown) for feature sampling the phase feature to a target dimension based on the preset phase dimension to obtain a first feature.

[0056] See also Figure 4 , Figure 4 4 is a schematic diagram of a framework of an embodiment of an electronic device 40 of the present application. The electronic device 40 includes a memory 41 and a processor 42. The memory 41 stores program instructions, and the processor 42 is used to execute the program instructions to implement the steps in any of the above-mentioned sound source positioning method embodiments. For details, please refer to the aforementioned disclosed embodiments, which will not be repeated here. The electronic device 40 may specifically include but is not limited to: a server, a smart phone, a laptop computer, a tablet computer, a self-service machine, etc., which are not limited here.

[0057] Specifically, the processor 42 is used to control itself and the memory 41 to implement the steps in any of the above-mentioned sound source localization method embodiments. The processor 42 can also be called a CPU (Central Processing Unit). The processor 42 may be an integrated circuit chip with signal processing capabilities. The processor 42 can also be a general-purpose processor, a digital signal processor (Digital Signal Processor, DSP), an application-specific integrated circuit (Application Specific Integrated Circuit, ASIC), a field-programmable gate array (Field-Programmable Gate Array, FPGA) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc. In addition, the processor 42 can be implemented by an integrated circuit chip.

[0058] In the above scheme, the electronic device 40 extracts the phase feature of the audio to be measured based on the audio to be measured collected by the target microphone array in the target device, phase samples the phase feature to the target dimension, the target dimension is a unified feature dimension when locating the sound source of different devices, and extracts features based on the attribute information of the target microphone array in the target device to obtain a second feature in the target dimension, the attribute information of the target microphone array at least includes the arrangement mode and the number of array elements of the target microphone array, fuses the first feature and the second feature to obtain the target feature, and obtains the sound source localization result of the audio to be measured based on the target feature. Therefore, even if the target microphone array in the target device used for sound source localization has different numbers and arrangements of array elements, resulting in the extracted phase feature of the audio to be measured being located in different feature dimensions, the first feature obtained based on the phase feature is sampled to a unified feature dimension, which can achieve the universality of sound source localization of different devices, and the second feature obtained based on the attribute information of the target microphone array can improve the specificity of the fused target feature, and the phase feature of the audio to be measured lost in the target feature is reduced as much as possible based on the attribute information of the target microphone array, so that the accuracy of the sound source localization result obtained based on the target feature can be improved. Therefore, the accuracy of sound source positioning can be improved while achieving the universality of sound source positioning of different devices.

[0059] See also Figure 5 , Figure 5 1 is a schematic diagram of a framework of an embodiment of a computer-readable storage medium 50 of the present application. The computer-readable storage medium 50 stores program instructions 51 that can be executed by a processor, and the program instructions 51 are used to implement the steps in any of the above-mentioned sound source localization method embodiments.

[0060] In the above scheme, the computer-readable storage medium 50 extracts the phase feature of the audio to be measured based on the audio to be measured collected by the target microphone array in the target device, phase samples the phase feature to the target dimension, the target dimension is a unified feature dimension when locating the sound source of different devices, and extracts features based on the attribute information of the target microphone array in the target device to obtain a second feature in the target dimension, the attribute information of the target microphone array at least includes the arrangement mode and the number of array elements of the target microphone array, fuses the first feature and the second feature to obtain the target feature, and obtains the sound source localization result of the audio to be measured based on the target feature. Therefore, even if the target microphone array in the target device used for sound source localization has different numbers and arrangements of array elements, resulting in the extracted phase feature of the audio to be measured being located in different feature dimensions, the first feature obtained based on the phase feature is sampled to a unified feature dimension, which can achieve the universality of sound source localization of different devices, and the second feature obtained based on the attribute information of the target microphone array can improve the specificity of the fused target feature, and the phase feature of the audio to be measured lost in the target feature is reduced as much as possible based on the attribute information of the target microphone array, so that the accuracy of the sound source localization result obtained based on the target feature can be improved. Therefore, the accuracy of sound source positioning can be improved while achieving the universality of sound source positioning of different devices.

[0061] In some embodiments, the functions or modules included in the device provided by the embodiments of the present disclosure can be used to execute the method described in the above method embodiments. The specific implementation can refer to the description of the above method embodiments, and for the sake of brevity, it will not be repeated here.

[0062] The above description of various embodiments tends to emphasize the differences between the various embodiments. The same or similar aspects can be referenced to each other, and for the sake of brevity, they will not be repeated herein.

[0063] In the several embodiments provided in the present application, it should be understood that the disclosed methods and devices can be implemented in other ways. For example, the device implementation described above is only schematic. For example, the division of modules or units is only a logical function division. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, and the indirect coupling or communication connection of devices or units can be electrical, mechanical or other forms.

[0064] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the present embodiment.

[0065] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of software functional units.

[0066] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product. The computer software product is stored in a storage medium, including several instructions to enable a computer device (which can be a personal computer, server, or network device, etc.) or a processor to execute all or part of the steps of each implementation method of the present application. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk and other media that can store program code.

[0067] If the technical solution of this application involves personal information, the product using the technical solution of this application has clearly informed the personal information processing rules and obtained the individual's voluntary consent before processing the personal information. If the technical solution of this application involves sensitive personal information, the product using the technical solution of this application has obtained the individual's separate consent before processing the sensitive personal information, and at the same time meets the "explicit consent" requirement. For example, on personal information collection devices such as cameras, clear and prominent signs are set to inform that the personal information collection scope has been entered and personal information will be collected. If the individual voluntarily enters the collection scope, it is deemed that he or she agrees to the collection of his or her personal information; or on the device that processes personal information, the personal information processing rules are notified by obvious signs / information, and the individual's authorization is obtained through pop-up information or by asking the individual to upload his or her personal information; among them, the personal information processing rules may include information such as the personal information processor, the purpose of personal information processing, the processing method, and the type of personal information processed.

Claims

1. A sound source localization method, characterized in that: include: Extracting a phase feature based on the audio to be tested collected by the target microphone array in the target device; wherein the feature dimension of the phase feature is determined by the number of array elements of the target microphone array in the target device and the phase dimension of each array element; Performing feature sampling to a target dimension based on the phase feature to obtain a first feature, and performing feature extraction based on attribute information of the target microphone array in the target device to obtain a second feature of the target dimension; wherein the attribute information at least includes an arrangement mode and a number of array elements of the target microphone array, and the target dimension is a unified feature dimension when performing sound source localization on different devices; Based on the first feature and the second feature, a target feature is obtained by fusing; Based on the target feature, a sound source localization result of the audio to be measured is obtained.

2. The method according to claim 1, characterized in that The sound source localization result is predicted by a sound source localization model, and the sound source localization model is trained based on a plurality of sample audio sets, wherein any of the sample audio sets includes: each sample audio collected from a sound source at the same sample position after different array elements in a candidate microphone array in the same candidate device are turned off.

3. The method according to claim 2, characterized in that The training steps of the sound source localization model include: Extracting features of the sample audios in the same sample audio set to obtain sample target features of the sample audios; Predicting based on the sample target features to obtain a predicted azimuth of the sound source; Based on the predicted azimuths respectively corresponding to the sample audios in the same sample audio set and the sample azimuths marked corresponding to the sample audio set, the network parameters of the sound source localization model are adjusted.

4. The method according to claim 3, characterized in that The extracting features of each sample audio in the same sample audio set to obtain sample target features of each sample audio includes: Extracting sample phase features of each of the sample audios in the same sample audio set; Performing feature sampling to the target dimension based on the sample phase feature to obtain a first sample feature, and performing feature extraction based on the attribute information of the candidate microphone array in the candidate device when any of the sample audios is collected to obtain a second sample feature of the target dimension; Based on the first sample feature and the second sample feature of the same sample audio, a sample target feature corresponding to the sample audio is obtained by fusion.

5. The method according to claim 3, characterized in that: The adjusting the network parameters of the sound source localization model based on the predicted azimuths respectively corresponding to the sample audios in the same sample audio set and the sample azimuths marked corresponding to the sample audio set includes: Determine a final azimuth corresponding to the sample audio set based on the predicted azimuths corresponding to the respective sample audios in the same sample audio set; Based on the difference between the final azimuth and the sample azimuth of the same sample audio set, the network parameters of the sound source localization model are adjusted.

6. The method according to claim 5, characterized in that The determining, based on the predicted azimuths respectively corresponding to the respective sample audios in the same sample audio set, a final azimuth corresponding to the sample audio set comprises: Counting the distribution of the predicted azimuth angles corresponding to the respective sample audios in the same sample audio set; The predicted azimuth angle with the largest distribution in the distribution situations is used as the final azimuth angle corresponding to the sample audio set.

7. The method according to claim 2, characterized in that The sound source localization model at least includes an encoder integrated with a self-attention mechanism; And / or, each of the candidate devices from which the plurality of sample audio sets are collected includes at least the target device; And / or, the candidate microphone array in the candidate device includes at least three array elements.

8. The method according to claim 1, characterized in that The performing feature sampling to a target dimension based on the phase feature to obtain a first feature includes: The phase dimension of each array element is sampled to a preset phase dimension; wherein the product of the value of the preset phase dimension and the number of array elements of the target microphone array is a unified preset value when performing sound source localization on different devices; Based on the preset phase dimension, the phase feature is feature sampled to the target dimension to obtain a first feature.

9. A sound source localization device, characterized in that: include: An extraction module, configured to extract a phase feature based on the audio to be tested collected by a target microphone array in a target device; wherein the feature dimension of the phase feature is determined by the number of array elements of the target microphone array in the target device and the phase dimension of each array element; A sampling module, configured to perform feature sampling to a target dimension based on the phase feature to obtain a first feature, and perform feature extraction based on the attribute information of the target microphone array in the target device to obtain a second feature of the target dimension; wherein the attribute information at least includes an arrangement mode and a number of array elements of the target microphone array, and the target dimension is a unified feature dimension when performing sound source localization on different devices; A fusion module, used for fusing the first feature and the second feature to obtain a target feature; A positioning module is used to obtain a sound source positioning result of the audio to be measured based on the target feature.

10. An electronic device, characterized in that: The invention comprises a memory and a processor coupled to each other, wherein the memory stores program instructions, and the processor is used to execute the program instructions to implement the sound source localization method according to any one of claims 1 to 8.

11. A computer-readable storage medium having program instructions stored thereon, characterized in that: When the program instructions are executed by a processor, the sound source localization method according to any one of claims 1 to 8 is implemented.

Citation Information

Patent Citations

  • Sound source positioning method and device

    CN108231085A

  • Sound source positioning method and device, readable storage medium and electronic equipment

    CN111161757A