Sound source localization method and apparatus

By dividing the multi-microphone array into a more optimal array and setting weight information, the problems of hardware irregularities and performance differences are solved, and the accuracy of sound source localization is improved.

WO2026156510A1PCT designated stage Publication Date: 2026-07-30BEIJING XIAOMI MOBILE SOFTWARE CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
BEIJING XIAOMI MOBILE SOFTWARE CO LTD
Filing Date
2025-01-21
Publication Date
2026-07-30

AI Technical Summary

Technical Problem

In existing technologies, multi-microphone array terminals suffer from hardware irregularities, performance differences, and inconsistent hardware configurations when locating sound sources. This causes the algorithm performance to fluctuate in different angles and directions, increasing the demand for computing resources and errors.

Method used

By dividing the microphone array into better segments and setting corresponding weight information, the sound source to be located is located by using the audio signals and weight information of multiple microphone arrays.

Benefits of technology

It effectively reduces positioning calculation errors and improves the accuracy of sound source localization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2025073791_30072026_PF_FP_ABST
    Figure CN2025073791_30072026_PF_FP_ABST
Patent Text Reader

Abstract

A sound source localization method and apparatus. The method comprises: acquiring audio signals of a sound source to be localized collected by microphones (mic1, mic2, mic3, and mic4) in a plurality of microphone arrays (S201); and localizing said sound source on the basis of the audio signals and weight information corresponding to the plurality of microphone arrays (S202). Therefore, said sound source is localized by partitioning microphones into more optimal microphone arrays and providing corresponding weight information, which can effectively reduce localization calculation errors and improve the localization accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

Sound source localization method and device Technical Field

[0001] This disclosure relates to the field of communication technology, and in particular to a method and apparatus for sound source localization. Background Technology

[0002] With the rapid advancement and development of science and technology, stereo playback technology has become widely used in the field of multimedia equipment technology. Almost all modern mobile phones, televisions, headphones, tablets and other devices support stereo playback, and the market demand for immersive sound formats has shown a continuous growth trend. Summary of the Invention

[0003] This disclosure provides a sound source localization method and apparatus, which can locate the sound source to be located by dividing a microphone array into better parts and setting corresponding weight information, thereby effectively reducing localization calculation errors and improving localization accuracy.

[0004] This disclosure presents a method and apparatus for sound source localization.

[0005] According to a first aspect of the present disclosure, a sound source localization method is proposed, comprising: acquiring audio signals of a sound source to be located collected by microphones in a plurality of microphone arrays; and locating the sound source to be located based on the audio signals corresponding to the plurality of microphone arrays and weight information.

[0006] In the above embodiments, by dividing the microphone array into better sections and setting corresponding weight information, the sound source to be located can be located, which can effectively reduce the positioning calculation error and improve the positioning accuracy.

[0007] According to a second aspect of the present disclosure, a sound source localization device is provided, comprising: a signal acquisition module for acquiring audio signals of a sound source to be located collected by microphones in a plurality of microphone arrays; and a localization processing module for locating the sound source to be located based on the audio signals and weight information corresponding to the plurality of microphone arrays.

[0008] According to a third aspect of the present disclosure, a communication device is provided, comprising: one or more processors, wherein the communication device is configured to perform the method described in the first aspect.

[0009] According to a fourth aspect of the present disclosure, a communication device is provided, comprising: one or more processors; and a memory coupled to the processors, the memory storing instructions which, when executed by the processors, cause the communication device to perform the method described in the first aspect.

[0010] According to a fifth aspect of the present disclosure, a computer storage medium is provided, wherein the computer storage medium stores computer-executable instructions; the computer-executable instructions, when executed by a processor, can implement the method described in the first aspect.

[0011] According to a sixth aspect of the present disclosure, a computer program product is provided, wherein the computer program product stores a computer program; after being executed by a processor, the computer program is able to implement the method described in the first aspect.

[0012] Additional aspects and advantages of this disclosure will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of this disclosure. Attached Figure Description

[0013] To more clearly illustrate the technical solutions in the embodiments of this disclosure, the accompanying drawings required for the description of the embodiments are introduced below. The following drawings are only some embodiments of this disclosure and do not impose specific limitations on the protection scope of this disclosure.

[0014] Figure 1 is an architecture diagram of a communication system provided in an embodiment of this disclosure;

[0015] Figure 2 is a flowchart of a sound source localization method provided in an embodiment of this disclosure;

[0016] Figure 3 is a schematic diagram of the microphone layout in a sound source localization device provided in an embodiment of this disclosure;

[0017] Figure 4 is a schematic diagram of a specified angle region provided in an embodiment of this disclosure;

[0018] Figure 5 is a structural diagram of a sound source localization device provided in an embodiment of this disclosure;

[0019] Figure 6A is a structural diagram of a communication device provided in an embodiment of this disclosure;

[0020] Figure 6B is a structural diagram of a chip provided in an embodiment of this disclosure. Detailed Implementation

[0021] This disclosure presents a method and apparatus for sound source localization.

[0022] In a first aspect, embodiments of this disclosure propose a sound source localization method, comprising: acquiring audio signals of a sound source to be located collected by microphones in a plurality of microphone arrays; and locating the sound source to be located based on the audio signals and weight information corresponding to the plurality of microphone arrays.

[0023] In the above embodiments, by dividing the microphone array into better sections and setting corresponding weight information, the sound source to be located can be located, which can effectively reduce the positioning calculation error and improve the positioning accuracy.

[0024] In conjunction with some embodiments of the first aspect, in some embodiments, the weight information corresponding to the microphone array is the target weight corresponding to the target angle region where the microphone array can perform sound source localization.

[0025] In the above embodiments, a better microphone array can be divided, and a target weight corresponding to the target angle region can be set when performing sound source localization. By combining the target weight corresponding to the target angle region, the sound source can be located, which can effectively reduce calculation errors and improve localization accuracy.

[0026] In conjunction with some embodiments of the first aspect, in some embodiments, the performance error of locating the sample sound source located at the sample angle is determined for each of the M candidate microphone arrays.

[0027] Based on the performance error of each candidate microphone array in locating the sample sound source located at the sample angle, determine the weights of each angle in the specified angle region for sound source localization of each candidate microphone array.

[0028] Based on the weights of each angle in the specified angular region for sound source localization using M candidate microphone arrays, determine multiple microphone arrays and the target weights corresponding to the target angular region for sound source localization of each microphone array.

[0029] The M candidate microphone arrays are obtained by grouping the N candidate microphones used to locate the sound source, where N is an integer greater than 2.

[0030] In conjunction with some embodiments of the first aspect, in some embodiments, the sum of the angle regions formed by the target angle regions for sound source localization by the multiple microphone arrays is equal to the specified angle region, and the sum of the target weights corresponding to each angle in the specified angle region for sound source localization by the multiple microphone arrays is 1.

[0031] In the above embodiments, for N candidate microphones with irregular distribution, different performance, and different hardware configurations, M candidate microphone arrays can be determined by grouping them. Based on the performance error of each candidate microphone array in locating the sample sound source at the sample angle, the weights corresponding to each angle in the specified angle region for sound source localization of each candidate microphone array are determined. Then, multiple microphone arrays that can locate sound sources at any angle position in the specified angle region are determined, as well as the target weights corresponding to the target angle regions for sound source localization of each microphone array. This can effectively reduce calculation errors and improve localization accuracy.

[0032] In conjunction with some embodiments of the first aspect, in some embodiments, the performance error of localizing a sample sound source located at a sample angle for each of the M candidate microphone arrays includes:

[0033] Acquire the sample audio signal from the sample sound source located at the sample angle, collected by each candidate microphone array;

[0034] Based on the sample audio signal, determine the sample positioning angle of the located sample sound source;

[0035] Based on the sample angle and sample positioning angle corresponding to each candidate microphone array, determine the performance error of each candidate microphone array in locating the sample sound source located at the sample angle.

[0036] In conjunction with some embodiments of the first aspect, in some embodiments, the number of sample angles is multiple, and the sample angles include various angles of a specified angle region; or the sample angles include a portion of the various angles of the specified angle region.

[0037] In conjunction with some embodiments of the first aspect, in some embodiments, based on the performance error of each candidate microphone array in locating sample sound sources located at sample angles, the weights corresponding to each angle in the specified angular region for sound source localization by each candidate microphone array are determined, including:

[0038] Based on the performance error of each candidate microphone array in locating the sample sound source at the sample angle, determine the weight corresponding to the sample angle for sound source localization of each candidate microphone array.

[0039] Based on the weights corresponding to the sample angles used for sound source localization by each candidate microphone array, determine the weights corresponding to each angle in the specified angular region for sound source localization by each candidate microphone array.

[0040] In conjunction with some embodiments of the first aspect, in some embodiments, the weight corresponding to the sample angle for sound source localization by each candidate microphone array is determined based on the performance error of each candidate microphone array in locating the sample sound source located at the sample angle, including:

[0041] If the performance error of the candidate microphone array in locating the sample sound source at the sample angle is less than the first threshold, the weight corresponding to the sample angle for sound source localization by the candidate microphone array is determined as the first value.

[0042] If the performance error of the candidate microphone array in locating the sample sound source at the sample angle is greater than or equal to the first threshold and less than the second threshold, the weight corresponding to the sample angle for sound source localization by the candidate microphone array is determined as the second value.

[0043] If the performance error of the candidate microphone array in locating the sample sound source at the sample angle is greater than or equal to the second threshold, the weight corresponding to the sample angle for sound source localization by the candidate microphone array is determined as the third value.

[0044] In conjunction with some embodiments of the first aspect, in some embodiments, the weight corresponding to the sample angle for sound source localization by each candidate microphone array is determined based on the performance error of each candidate microphone array in locating the sample sound source located at the sample angle, including:

[0045] Based on the performance error of each candidate microphone array in locating the sample sound source at the sample angle, and the mapping relationship between the performance error and the weight, the weight corresponding to the sample angle for sound source localization of each candidate microphone array is determined.

[0046] In conjunction with some embodiments of the first aspect, in some embodiments, based on the weights corresponding to the sample angles used for sound source localization by each candidate microphone array, the weights corresponding to each angle in a specified angular region for sound source localization by each candidate microphone array are determined, including:

[0047] Based on the weights corresponding to the angles of two adjacent samples in the sound source localization of each candidate microphone array, determine the weights corresponding to each angle between two adjacent sample angles in the specified angle region for sound source localization of each candidate microphone array.

[0048] In the above embodiments, the performance error of each candidate microphone array in locating the sample sound source located at the sample angle can be determined, and then the weights corresponding to each angle in the specified angle region for sound source localization by each candidate microphone array can be determined according to the performance error, so as to reduce calculation errors and improve localization accuracy.

[0049] Secondly, embodiments of this disclosure propose a sound source localization device, comprising: a signal acquisition module for acquiring audio signals of a sound source to be located collected by microphones in a plurality of microphone arrays; and a localization processing module for locating the sound source to be located based on the audio signals and weight information corresponding to the plurality of microphone arrays.

[0050] Thirdly, a communication device is proposed, comprising: one or more processors, wherein the communication device is used to execute the method described in the first aspect.

[0051] Fourthly, embodiments of this disclosure provide a communication device, which includes: one or more processors; and a memory coupled to the processors, the memory storing instructions that, when executed by the processors, cause the communication device to perform the method described in the first aspect.

[0052] Fifthly, embodiments of this disclosure provide a storage medium storing instructions that, when executed on a communication device, cause the communication device to perform the method described in the first aspect.

[0053] In a sixth aspect, embodiments of this disclosure provide a program product that, when executed by a communication device, causes the communication device to perform the method described in the first aspect.

[0054] In a seventh aspect, embodiments of this disclosure provide a computer program that, when run on a computer, causes the computer to perform the method described in the first aspect.

[0055] Eighthly, embodiments of this disclosure provide a chip or chip system. The chip or chip system includes processing circuitry configured to perform the method described in the first aspect.

[0056] It is understood that the aforementioned communication equipment, communication system, storage medium, program product, etc., are all used to execute the methods proposed in the embodiments of this disclosure. Therefore, the beneficial effects they can achieve can be referred to the beneficial effects in the corresponding methods, and will not be repeated here.

[0057] This disclosure provides a sound source localization method and apparatus. In some embodiments, the terms "sound source localization method" and "information processing method," "localization method," etc., can be used interchangeably.

[0058] This disclosure is not exhaustive, but merely illustrative of some embodiments, and is not intended to limit the scope of protection of this disclosure. Unless otherwise specified, each step in a particular embodiment can be implemented as an independent embodiment, and the steps can be arbitrarily combined. For example, a solution after removing some steps in a particular embodiment can also be implemented as an independent embodiment, and the order of the steps in a particular embodiment can be arbitrarily interchanged. Furthermore, the optional implementation methods in a particular embodiment can be arbitrarily combined; moreover, the embodiments can be arbitrarily combined, for example, some or all steps of different embodiments can be arbitrarily combined, and a particular embodiment can be arbitrarily combined with the optional implementation methods of other embodiments. In all embodiments of this disclosure, unless otherwise specified or logically conflicting, the terminology and / or descriptions between the embodiments are consistent and can be mutually referenced. Technical features in different embodiments can be combined to form new embodiments based on their inherent logical relationships.

[0059] The terminology used in the embodiments of this disclosure is for the purpose of describing particular embodiments only and is not intended to limit the scope of this disclosure.

[0060] In this embodiment of the disclosure, unless otherwise stated, elements expressed in the singular form, such as "a," "an," "the," "the," "the," "the," "the," "the," "this," etc., can mean "one and only one," or "one or more," "at least one," etc. For example, when using articles such as "a," "an," "the," etc. in translation, the noun following the article can be understood as either a singular expression or a plural expression.

[0061] In the embodiments disclosed herein, "multiple" refers to two or more.

[0062] In some embodiments, the terms “at least one of A or B, at least one of A and B”, “one or more”, “a plurality of”, “multiple”, etc., may be used interchangeably.

[0063] In some embodiments, the notation "at least one of A and B", "A and / or B", "A in one case, B in another", "in response to one case A, in response to another case B", etc., may include the following technical solutions depending on the situation: in some embodiments, A (execute A regardless of whether there is a branch B); in some embodiments, B (execute B regardless of whether there is a branch A); in some embodiments, execution is selected from A and B (A and B are selectively executed); in some embodiments, both A and B are executed. The same applies when there are more branches such as A, B, C, etc.

[0064] In some embodiments, the notation "A or B" may include the following technical solutions, depending on the situation: in some embodiments, A (execute A regardless of whether a branch B exists); in some embodiments, B (execute B regardless of whether a branch A exists); in some embodiments, execution is selected from A and B (A and B are selectively executed). The same applies when there are more branches such as A, B, and C.

[0065] The prefixes "first," "second," etc., used in the embodiments of this disclosure are merely for distinguishing different descriptive objects and do not impose restrictions on the position, order, priority, quantity, or content of the descriptive objects. The description of the descriptive objects is found in the claims or the context of the embodiments, and the use of prefixes should not constitute unnecessary restrictions. For example, if the descriptive object is a "field," the ordinal numbers preceding "field" in "first field" and "second field" do not restrict the position or order of the "fields." "First" and "second" do not restrict whether the "fields" they modify are in the same message, nor do they restrict the order of "first field" and "second field." Similarly, if the descriptive object is a "level," the ordinal numbers preceding "level" in "first level" and "second level" do not restrict the priority between "levels." Furthermore, the number of descriptive objects is not limited by ordinal numbers and can be one or more. For example, in "first device," the number of "devices" can be one or more. Furthermore, the objects modified by different prefixes can be the same or different. For example, if the object being described is "device", then "first device" and "second device" can be the same device or different devices, and their types can be the same or different. Similarly, if the object being described is "information", then "first information" and "second information" can be the same information or different information, and their content can be the same or different.

[0066] In some embodiments, “including A,” “containing A,” “for indicating A,” and “carrying A” can be interpreted as directly carrying A or indirectly indicating A.

[0067] In some embodiments, terms such as "time / frequency" and "time-frequency domain" refer to the time domain and / or frequency domain.

[0068] In some embodiments, terms such as “in response to…”, “in response to determining…”, “in the case of…”, “when…”, “when…”, “if…”, etc. can be used interchangeably. These descriptions all refer to the device making a corresponding action under certain objective circumstances. They do not necessarily limit the time, nor do they require the device to make a judgment action when implementing it, nor do they mean that there must be other limitations.

[0069] In some embodiments, the terms “greater than,” “greater than or equal to,” “not less than,” “more than,” “more than or equal to,” “not less than,” “higher than,” “higher than or equal to,” “not lower than,” and “above” can be used interchangeably, as can the terms “less than,” “less than or equal to,” “not greater than,” “less than,” “less than or equal to,” “not more than,” “lower than,” “lower than or equal to,” “not higher than,” and “below”.

[0070] In some embodiments, devices, etc., may be interpreted as physical or virtual, and their names are not limited to those described in the embodiments. Terms such as “device,” “equipment,” “circuit,” “network element,” “network function,” “network device,” “function,” “node,” “unit,” “section,” “system,” “network,” “chip,” “chip system,” “entity,” and “subject” are interchangeable.

[0071] In some embodiments, "network" can be interpreted as devices included in a network (e.g., access network devices, core network devices, etc.).

[0072] In some embodiments, the terms "access network device (AN device)," "radio access network device (RAN device)," "base station (BS)," "radio base station," "fixed station," "node," "access point," "transmission point (TP)," "reception point (RP)," "transmission / reception point (TRP)," "panel," "antenna panel," "antenna array," "cell," "macro cell," "small cell," "femto cell," "pico cell," "sector," "cell group," "serving cell," "carrier," "component carrier," and "bandwidth part (BWP)" can be used interchangeably.

[0073] In some embodiments, the terms "terminal", "terminal device", "user equipment (UE)", "user terminal", "mobile station (MS)", "mobile terminal (MT)", "subscriber station", "mobile unit", "subscriber unit", "wireless unit", "remote unit", "mobile device", "wireless device", "wireless communication device", "remote device", "mobile subscriber station", "access terminal", "mobile terminal", "wireless terminal", "remote terminal", "handset", "user agent", "mobile client", and "client" can be used interchangeably.

[0074] In some embodiments, access network devices, core network devices, or network devices can be replaced with terminals. For example, embodiments of this disclosure can also be applied to structures where communication between access network devices, core network devices, or network devices and terminals is replaced with communication between multiple terminals (e.g., device-to-device (D2D), vehicle-to-everything (V2X), etc.). In this case, the structure can also be configured such that the terminal has all or part of the functions of the access network device. Furthermore, terms such as "uplink" and "downlink" can be replaced with terms corresponding to communication between terminals (e.g., "sidelink"). For example, uplink channel, downlink channel, etc., can be replaced with sidelink channel, uplink link, downlink link, etc., can be replaced with sidelink link.

[0075] In some embodiments, the terminal may be replaced by an access network device, a core network device, or a network device. In this case, the access network device, core network device, or network device may also be configured to have all or some of the functions of the terminal.

[0076] In some embodiments, the acquisition of data, information, etc., may comply with the laws and regulations of the country where the location is situated.

[0077] In some embodiments, data, information, etc., may be obtained with the user's consent.

[0078] Furthermore, each element, each row, or each column in the table of this disclosure can be implemented as an independent embodiment, and any combination of any element, any row, or any column can also be implemented as an independent embodiment.

[0079] Figure 1 is a schematic diagram of the architecture of a communication system according to an embodiment of the present disclosure.

[0080] Figure 1 is an architecture diagram of a communication system provided in an embodiment of this disclosure.

[0081] As shown in Figure 1, the communication system 100 includes a terminal 101 and a network device 102.

[0082] In some embodiments, terminal 101 includes, but is not limited to, at least one of the following: mobile phone, wearable device, Internet of Things device, car with communication function, smart car, tablet computer, computer with wireless transceiver function, virtual reality (VR) terminal, augmented reality (AR) terminal, wireless terminal in industrial control, wireless terminal in self-driving, wireless terminal in remote medical surgery, wireless terminal in smart grid, wireless terminal in transportation safety, wireless terminal in smart city, and wireless terminal in smart home.

[0083] In some embodiments, network device 102 may include at least one of access network device and core network device.

[0084] In some embodiments, the access network device is, for example, a node or device that connects a terminal to a wireless network. The access network device may include, but is not limited to, at least one of the following in a 5G communication system: evolved Node B (eNB), next-generation eNB (ng-eNB), next-generation Node B (gNB), node B (NB), home node B (HNB), home evolved node B (HeNB), radio backhaul device, radio network controller (RNC), base station controller (BSC), base transceiver station (BTS), base band unit (BBU), mobile switching center, base station in a 6G communication system, open RAN, cloud RAN, base station in other communication systems, and access node in a Wi-Fi system.

[0085] In some embodiments, the access network device may be a satellite.

[0086] In some embodiments, the core network equipment may be a single device, multiple devices, or a group of devices, including all or part of a first network element, a second network element, a third network element, a fourth network element, etc. Network elements may be virtual or physical. The core network may include, for example, at least one of an Evolved Packet Core (EPC), a 5G Core Network (5GCN), and a Next Generation Core (NGC).

[0087] In some embodiments, the first network element is, for example, an access and mobility management function (AMF) network element.

[0088] In some embodiments, the second network element is, for example, a Location Management Function (LMF) network element.

[0089] In some embodiments, the third network element is, for example, a sensing function (SF) network element.

[0090] In some embodiments, the fourth network element is, for example, a network repository function (NRF) network element.

[0091] In some embodiments, the first network element is used to implement terminal access management and mobility management. It is responsible for terminal state maintenance, terminal reachability management, mobility management (MM), forwarding of non-access stratum (NAS) messages, and forwarding of session management (SM) N2 messages.

[0092] In some embodiments, the second network element is used to coordinate and schedule the resources required for the location of the terminal.

[0093] In some embodiments, the third network element is used to perform wireless sensing using access network equipment or terminals to realize sensing services.

[0094] In some embodiments, the fourth network element is used for dynamic registration of network function service capabilities and network function discovery.

[0095] In some embodiments, at least one of the first network element, the second network element, and the third network element can be independent of the core network equipment.

[0096] In some embodiments, at least one of the first network element, the second network element, and the third network element may be part of the core network equipment.

[0097] It is understood that the communication system described in this disclosure is for the purpose of more clearly illustrating the technical solutions of this disclosure, and does not constitute a limitation on the technical solutions proposed in this disclosure. As those skilled in the art will know, with the evolution of system architecture and the emergence of new business scenarios, the technical solutions proposed in this disclosure are also applicable to similar technical problems.

[0098] The following embodiments of this disclosure can be applied to the communication system 100 shown in FIG1, or to some of the main bodies, but are not limited thereto. The main bodies shown in FIG1 are illustrative. The communication system may include all or some of the main bodies in FIG1, or may include other main bodies outside of FIG1. ​​The number and form of each main body are arbitrary. Each main body may be physical or virtual. The connection relationship between the main bodies is illustrative. The main bodies may not be connected or may be connected. The connection can be in any way, it can be a direct connection or an indirect connection, it can be a wired connection or a wireless connection.

[0099] The embodiments disclosed herein can be applied to Long Term Evolution (LTE), LTE-Advanced (LTE-A), LTE-Beyond (LTE-B), Super 3G, IMT-Advanced, 4th Generation Mobile Communication System (4G), 5th Generation Mobile Communication System (5G), 5G New Radio (NR), Future Radio Access (FRA), New-Radio Access Technology (RAT), New Radio (NR), New Radio Access (NX), Future Generation Radio Access (FX), Global System for Mobile Communications (GSM), CDMA2000, Ultra Mobile Broadband (UMB), IEEE 802.11 (Wi-Fi), IEEE 802.16 (WiMAX), IEEE 802.20, and Ultra-Wideband. The technologies used include UWB (Ultra-Wideband), Bluetooth (a registered trademark), public land mobile network (PLMN) networks, device-to-device (D2D) systems, machine-to-machine (M2M) systems, Internet of Things (IoT) systems, vehicle-to-everything (V2X) systems, systems utilizing other communication methods, and next-generation systems built upon them. Furthermore, multiple systems can be combined (e.g., a combination of LTE or LTE-A with 5G).

[0100] To meet the ever-increasing demand for audio, devices, especially smartphones, are commonly equipped with multiple microphones, such as three to four. This increase in the number of microphones creates larger microphone arrays, providing richer data sources for audio algorithms. Theoretically, this should help improve algorithm performance, but it also increases algorithm complexity and computational resource requirements. Therefore, for devices equipped with multiple microphones, there is an urgent need for an efficient data integration and processing solution to balance algorithm performance and resource consumption.

[0101] In immersive audio processing algorithms, sound source localization technology has attracted much attention due to its crucial role in audio acquisition. For example, the direction of arrival (DOA) parameter, as an important indicator in monophonic speech recognition and multi-channel audio processing, is not only defined as standard metadata for Immersive Voice and Audio Services (IVAS) codecs in the 3rd Generation Partnership Project (3GPP) standard, but is also widely used in various algorithms to accurately describe the location of sound sources.

[0102] In conclusion, with the widespread adoption of stereo playback technology in multimedia devices, the importance of sound source localization technology is becoming increasingly prominent. Facing the challenges brought by the increasing number of microphones, developing efficient data integration and processing solutions is crucial for improving algorithm performance and reducing resource consumption.

[0103] In some embodiments, a sound source localization algorithm is an algorithm that determines the location of a sound source by analyzing sound signals. This algorithm has applications in many fields, including robot navigation, environmental monitoring, and security surveillance. The basic principle of a sound source localization algorithm is to utilize the characteristics of sound propagation in space, and to deduce the location of the sound source by measuring the differences (relationships) in reception time or intensity of sound at different locations (mics at different locations).

[0104] In some embodiments, the sound source localization algorithm processes the following steps: One is a cross-correlation analysis method: The sound source localization algorithm first performs cross-correlation processing on each pair of microphones in the microphone array. By analyzing the correlation between the received signals, the relative position of the sound source is determined. This step utilizes the basic principle of signal processing, namely, estimating the angle between the sound source and the receiver by comparing the phase difference of the signals. Another method is multi-angle scanning: To comprehensively assess the location of the sound source, the sound source localization algorithm performs a detailed scan of all possible angles. By evaluating the signal response at each angle, the sound source localization algorithm can identify the peaks of the spatial spectrum. The location and intensity of these peaks provide valuable information about the direction of signal arrival.

[0105] In some embodiments, the sound source localization algorithm divides the microphone into small arrays of two microphones each, and performs localization processing at different angles.

[0106] In related technologies, due to limitations in cost and hardware space of terminals such as mobile phones, there are generally irregular microphone arrays, poor microphone hardware performance, and inconsistencies in performance between different microphones and even the same microphone in different directions caused by diffraction and refraction effects due to terminal structure obstruction. This leads to:

[0107] 1. Performance differences between different microphone pairs and inconsistencies in hardware errors between different microphone pairs lead to inconsistent algorithm performance between different microphone pairs.

[0108] 2. Irregular microphone arrays, coupled with the performance variations of microphones in different directions, result in inconsistent array errors across different orientations. Conversely, breaking down a single microphone array (more than two) into multiple smaller arrays will lead to variations in error and performance across different orientations. Consequently, algorithm performance will be inconsistent across different angles.

[0109] Therefore, how to solve the problem of algorithm performance fluctuations in different angles and directions when using multiple microphones with irregular distribution, different performance and different hardware configurations for sound source localization, as well as the problem of increased computing power requirements after forming multiple microphone arrays, are urgent issues that need to be addressed.

[0110] Based on this, this disclosure provides a sound source localization method and apparatus. The method includes: acquiring audio signals of a sound source to be located collected by microphones in a plurality of microphone arrays; and locating the sound source to be located based on the audio signals corresponding to the plurality of microphone arrays and weight information. Thus, by dividing the microphone arrays into optimal arrays and setting corresponding weight information, the sound source to be located can be effectively located, thereby reducing localization calculation errors and improving localization accuracy.

[0111] Figure 2 is a schematic diagram illustrating a sound source localization method according to an embodiment of the present disclosure. As shown in Figure 2, the present disclosure relates to a sound source localization method, which includes:

[0112] S201, acquire the audio signal of the sound source to be located collected by the microphones in the multiple microphone array.

[0113] In this embodiment of the disclosure, the sound source localization method is executed by a sound source localization device, wherein the sound source localization device can be a communication device, and the communication device can be a terminal, such as a mobile phone, computer, tablet, vehicle, etc.

[0114] In this embodiment of the disclosure, the communication device can execute the sound source localization method by installing a specific application. For example, if the communication device is a terminal, a specific application can be installed on the terminal to execute the sound source localization method.

[0115] In some embodiments, the sound source localization device has multiple microphones, which can be grouped into multiple candidate microphone arrays, wherein each candidate microphone array includes multiple microphones.

[0116] For example, as shown in Figure 3, taking the sound source localization device as an example, the terminal has 4 microphones (mic), and the layout of the four microphones is shown in Figure 3. The 4 microphones include mic1, mic2, mic3 and mic4.

[0117] In some embodiments, the four microphones can be grouped into candidate microphone arrays, wherein mic1 and mic2 are candidate microphone array 1, mic1 and mic3 are candidate microphone array 2, mic1 and mic4 are candidate microphone array 3, mic2 and mic3 are candidate microphone array 4, mic2 and mic4 are candidate microphone array 5, mic3 and mic4 are candidate microphone array 6, mic1, mic2, and mic3 are candidate microphone array 7, mic1, mic2, and mic4 are candidate microphone array 8, mic1, mic3, and mic4 are candidate microphone array 9, mic2, mic3, and mic4 are candidate microphone array 10, and mic1, mic2, mic3, and mic4 are candidate microphone array 11.

[0118] It should be noted that in the above embodiment, the 4 microphones are grouped to obtain 11 candidate microphone arrays. Of course, they can also be grouped into fewer than 11 candidate microphone arrays, such as grouping into 5 candidate microphone arrays, etc. This disclosure does not impose specific limitations on this.

[0119] In this embodiment of the present disclosure, for the multiple candidate microphone arrays grouped into multiple microphone groups in the sound source localization device, all of them can be determined as microphone arrays for locating the sound source to be located, or some of the candidate microphone arrays among the multiple candidate microphone arrays can be determined as microphone arrays for locating the sound source to be located.

[0120] In some embodiments, a subset of the candidate microphone arrays is selected as the microphone array based on the performance error of the sound source localization using multiple candidate microphone arrays. Here, the localization performance error refers to the localization accuracy.

[0121] For example, a sound source at a known location is located using multiple candidate microphone arrays, and the location results are obtained respectively. The accuracy of the location is determined by comparing the location results obtained from the multiple candidate microphone arrays with the known location of the sound source.

[0122] It is understandable that multiple candidate microphone arrays participate in the localization of the sound source to be located, and each candidate microphone array obtains a localization result. Further, based on the localization results obtained from the multiple candidate microphone arrays, the final localization result of the sound source to be located is obtained.

[0123] In some embodiments, the localization results obtained by locating the sound source to be located from multiple candidate microphone arrays are averaged, and the average value is determined as the final localization result of the sound source to be located.

[0124] In some embodiments, the maximum value among the localization results obtained by locating the sound source to be located from multiple candidate microphone arrays is determined as the final localization result of the sound source to be located.

[0125] In some embodiments, the minimum value among the localization results obtained by multiple candidate microphone arrays in locating the sound source to be located is determined as the final localization result of the sound source to be located.

[0126] In some embodiments, the median value of the localization results obtained by locating the sound source to be located using multiple candidate microphone arrays is determined as the final localization result of the sound source to be located.

[0127] In some embodiments, the positioning results obtained by locating the sound source to be located from multiple candidate microphone arrays are weighted and summed to obtain the final positioning result of the sound source to be located.

[0128] In some embodiments, the weighting information for the weighted summation of the positioning results obtained by locating the sound source to be located from multiple candidate microphone arrays can be that the weights corresponding to the multiple candidate microphone arrays are all equal, or different weights can be set according to the positioning performance of different candidate microphone arrays, or different weights can be set according to the different positioning performance errors of different candidate microphone arrays in different angular regions. Here, the positioning performance error can be the positioning accuracy.

[0129] In this embodiment of the disclosure, after determining the multiple microphone arrays for locating the sound source to be located, and the weight information corresponding to each microphone array, the microphones in the determined multiple microphone arrays can be used to collect the audio signal of the sound source to be located.

[0130] S202, based on the audio signals and weight information corresponding to multiple microphone arrays, locates the sound source to be located.

[0131] In some embodiments, multiple microphone arrays for locating the sound source to be located, and weight information corresponding to each microphone array are determined.

[0132] Each microphone array includes multiple microphones.

[0133] In some embodiments, the weight information corresponding to the microphone array is the target weight corresponding to the target angle region where the microphone array can perform sound source localization.

[0134] For example, taking the sound source localization device shown in Figure 3 as an example, based on the localization performance error of the 11 microphone arrays, microphone array 1 composed of mic1 and mic2, microphone array 3 composed of mic1 and mic4, and microphone array 4 composed of mic2 and mic3 are determined to be microphone arrays. Among them, the sound source localization device can locate the sound source located in a specified angle area, which is a horizontal angle [-180°, 180°] and a pitch angle [-180°, 180°], or a specified angle area is a horizontal angle [-90°, 90°] and a pitch angle [-90°, 90°].

[0135] As shown in Figure 4, the Z-axis is determined by the plane containing the two longest sides of the sound source locator, i.e., the ZY plane. The plane in which the sound source locator rotates along the perpendicular line of the longest side is the central axis, i.e., the line where the XZ plane intersects with the ZY plane. The plane in which the sound source locator rotates along the Z-axis is the XY plane. The line where the XY plane intersects with the XZ plane is the X-axis. The line where the XY plane intersects with the ZY plane is the Y-axis.

[0136] It should be noted that the sound source localization device can locate a sound source located in a specified angular region. The specified angular region can also be represented by an angular region in a coordinate system of other directions, such as a solid angle. This embodiment of the present disclosure does not impose any specific limitations on this.

[0137] In some embodiments, the sum of the angle regions formed by the target angle regions for sound source localization by the multiple microphone arrays is equal to the specified angle region, and the sum of the target weights corresponding to each angle in the specified angle region for sound source localization by the multiple microphone arrays is 1.

[0138] In this embodiment of the disclosure, the sum of the angle regions formed by the target angle regions for sound source localization by multiple microphone arrays is equal to the specified angle region.

[0139] In some embodiments, the specified angular region is the maximum angular range that can be measured to locate the sound source.

[0140] In some embodiments, a sound source can only be located if it is located within a specified angular region.

[0141] In some embodiments, the specified angle region is the entire angle region.

[0142] In some embodiments, the specified angle region is the maximum angle region that the sound source localization device can locate when positioning the sound source.

[0143] In some embodiments, the specified angular region is the entire angular range surrounding the sound source localization device.

[0144] In this embodiment of the disclosure, in the designated angular region where multiple microphone arrays can perform sound source localization, the sum of the target weights corresponding to each angle is 1.

[0145] For example, the specified angle range is horizontal angle [-90°, 90°] and pitch angle [-90°, 90°], and the microphone array 1 composed of mic1 and mic2, the microphone array 3 composed of mic1 and mic4, and the microphone array 4 composed of mic2 and mic3 are determined to be microphone arrays.

[0146] Among them, the microphone array 1 composed of mic1 and mic2 can perform sound source localization at a target angle of horizontal angle [-90°, 90°], with a target weight of 1.

[0147] The microphone array 3, composed of mic1 and mic4, can perform sound source localization at target angles of pitch [-90°, 10°], where the target weight corresponding to the pitch angle [-90°, -10°] is 1, and the target weight corresponding to the pitch angle [-10°, 10°] is 0.5.

[0148] The microphone array 4, composed of mic2 and mic3, can perform sound source localization at target angles of pitch [-10°, 90°], where the target weight corresponding to the pitch angle [-10°, 10°] is 0.5, and the target weight corresponding to the pitch angle [10°, 90°] is 1.

[0149] In this case, at a pitch angle of 10°, the target weight corresponding to microphone array 3 composed of mic1 and mic4 is 0.5, and the target weight corresponding to microphone array 4 composed of mic2 and mic3 is 0.5. That is, at a pitch angle of 10°, the sum of the target weights corresponding to multiple microphone arrays is 1.

[0150] For example, the specified angle range is horizontal angle [-180°, 180°] and pitch angle [-180°, 180°], and the microphone array 1 composed of mic1 and mic2, the microphone array 3 composed of mic1 and mic4, and the microphone array 4 composed of mic2 and mic3 are determined to be microphone arrays.

[0151] Among them, the target angle for sound source localization of the microphone array 1 composed of mic1 and mic2 is the horizontal angle [-180°, 90°], the target weight corresponding to the horizontal angle [-180°, 0°] is 1, and the target weight corresponding to the horizontal angle [0°, 90°] is 0.5.

[0152] The microphone array 3, composed of mic1 and mic4, can perform sound source localization at target angles of [0°, 180°] and [-180°, 10°]. The target weight for the horizontal angle [0°, 90°] is 0.5, the target weight for the horizontal angle [90°, 180°] is 1, the target weight for the pitch angle [-180°, -10°] is 1, and the target weight for the pitch angle [-10°, 10°] is 0.5.

[0153] The microphone array 4, composed of mic2 and mic3, can perform sound source localization at target angles of [-10°, 180°], where the target weight corresponding to [-10°, 10°] is 0.5 and the target weight corresponding to [10°, 180°] is 1.

[0154] Specifically, at a horizontal angle of 90°, the target weight corresponding to microphone array 1 (composed of mic1 and mic2) is 0.5, and the target weight corresponding to microphone array 3 (composed of mic1 and mic4) is also 0.5. In other words, at a horizontal angle of 90°, the sum of the target weights corresponding to multiple microphone arrays is 1. Similarly, at a pitch angle of 10°, the target weight corresponding to microphone array 3 (composed of mic1 and mic4) is 0.5, and the target weight corresponding to microphone array 4 (composed of mic2 and mic3) is also 0.5. Again, at a pitch angle of 10°, the sum of the target weights corresponding to multiple microphone arrays is 1.

[0155] It should be noted that the above examples are for illustrative purposes only and are not intended to limit the specific embodiments of this disclosure. The specified angle region can also be represented by a solid angle, the determined microphone array can also be 5, 10, etc., and the target weight can also be 0.1, 0.4, 0.6, 0.9, etc.

[0156] In some embodiments, determining multiple microphone arrays for locating the sound source to be located, and weight information corresponding to each microphone array, includes:

[0157] The N candidate microphones that can be used to locate the sound source are grouped to determine M candidate microphone arrays, where each candidate microphone array includes multiple candidate microphones, and N is an integer greater than 2.

[0158] Determine the performance error of each candidate microphone array in locating sample sound sources located at the sample angle;

[0159] Based on the performance error of each candidate microphone array in locating the sample sound source located at the sample angle, determine the weights of each angle in the specified angle region for sound source localization of each candidate microphone array.

[0160] Based on the weights corresponding to each angle in the specified angular region for sound source localization using M candidate microphone arrays, multiple microphone arrays are determined, along with the target weights corresponding to the target angular regions for sound source localization of each microphone array.

[0161] In this embodiment of the disclosure, the sound source localization device can use N candidate microphones to locate the sound source to be located, where N is an integer greater than 2. The N candidate microphones are grouped to determine an M candidate microphone array.

[0162] in, It is expressed as the number of combinations of choosing N items from N items. It is expressed as the number of combinations of choosing N-1 items from N items. It is represented as the number of combinations of choosing 2 from N.

[0163] For example, the sound source localization device can use 5 candidate microphones to locate the sound source, and group the 5 candidate microphones to determine 2 candidate microphone groups, or determine no more than 10 candidate microphone groups. One candidate microphone group.

[0164] After determining M candidate microphone groups, the performance error of each candidate microphone array in locating the sample sound source at the sample angle can be determined.

[0165] In this embodiment of the disclosure, the positioning angle result is obtained by locating the sample sound source located at a specified angle by each candidate microphone array. By comparing the positioning angle result with the specified angle, the performance error of the candidate microphone array in locating the sample sound source located at the sample angle can be determined.

[0166] In some embodiments, determining the performance error of each candidate microphone array in locating a sample sound source located at a sample angle includes:

[0167] Acquire the sample audio signal from the sample sound source located at the sample angle, collected by each candidate microphone array;

[0168] Based on the sample audio signal of the sample sound source collected by each candidate microphone array, determine the sample positioning angle of the located sample sound source.

[0169] Based on the sample angle and sample positioning angle corresponding to each candidate microphone array, determine the performance error of each candidate microphone array in locating the sample sound source located at the sample angle.

[0170] In this embodiment of the disclosure, sample audio signals of sample sound sources located at sample angles are acquired by each candidate microphone array. Based on the sample audio signals of sample sound sources located at sample angles acquired by each candidate microphone array, the sample positioning angle of the located sample sound source is determined. Based on the sample angle and sample positioning angle corresponding to each candidate microphone array, the performance error of each candidate microphone array in locating the sample sound source located at the sample angle is determined.

[0171] In some embodiments, the number of sample angles is multiple, and the sample angles include all angles of a specified angle region; or the sample angles include a portion of the angles of a specified angle region.

[0172] In this embodiment of the disclosure, sample audio signals of sample sound sources located at sample angles are acquired by each candidate microphone array. Based on the sample audio signals of sample sound sources located at sample angles acquired by each candidate microphone array, the sample positioning angle of the located sample sound source is determined. Based on the sample angle and sample positioning angle corresponding to each candidate microphone array, the performance error of each candidate microphone array in locating the sample sound source located at the sample angle is determined.

[0173] In cases where the sample angles include various angles within a specified angle region, the performance error of each candidate microphone array in locating sample sound sources at various angles within the specified angle region can be determined.

[0174] In cases where the sample angles include a portion of the angles within a specified angular region, the performance error of each candidate microphone array in locating sample sound sources located at a portion of the angles within the specified angular region can be determined.

[0175] In some embodiments, based on the performance error of each candidate microphone array in locating sample sound sources located at sample angles, the weights corresponding to each angle in the specified angular region for sound source localization by each candidate microphone array are determined, including:

[0176] Based on the performance error of each candidate microphone array in locating the sample sound source at the sample angle, determine the weight corresponding to the sample angle for sound source localization of each candidate microphone array.

[0177] Based on the weights corresponding to the sample angles used for sound source localization by each candidate microphone array, determine the weights corresponding to each angle in the specified angular region for sound source localization by each candidate microphone array.

[0178] In this embodiment of the disclosure, when the performance error of each candidate microphone array in locating the sample sound source located at the sample angle is determined, the weights corresponding to each angle in the specified angular region for sound source localization by each candidate microphone array can be determined based on the performance error of each candidate microphone array in locating the sample sound source located at the sample angle. Specifically, the weights corresponding to the sample angles for sound source localization by each candidate microphone array are determined based on the performance error of each candidate microphone array in locating the sample sound source located at the sample angle, and then the weights corresponding to each angle in the specified angular region for sound source localization by each candidate microphone array are determined based on the weights corresponding to the sample angles for sound source localization by each candidate microphone array.

[0179] In some embodiments, the performance error is the absolute value of the difference between the sample angle and the sample positioning angle. It can be understood that a smaller performance error indicates higher positioning accuracy, thus allowing for a higher weight to be assigned. Conversely, a larger performance error indicates lower positioning accuracy, allowing for a lower weight to be assigned.

[0180] In some embodiments, the performance error is the ratio of the sample positioning angle to the sample angle. It can be understood that the closer the performance error is to 1, the higher the positioning accuracy, and thus a higher weight can be assigned. Conversely, the farther the performance error is from 1, the lower the positioning accuracy, and a lower weight can be assigned.

[0181] In some embodiments, a performance error and weight mapping table is set up, where different performance errors correspond to different weights, or performance errors within different ranges correspond to different weights, and performance errors within the same range correspond to the same weight.

[0182] In some embodiments, determining the weight corresponding to the sample angle for sound source localization by each candidate microphone array based on the performance error of each candidate microphone array in localizing the sample sound source at the sample angle includes: determining the weight corresponding to the sample angle for sound source localization by each candidate microphone array based on the performance error of each candidate microphone array in localizing the sample sound source at the sample angle, and the mapping relationship between performance error and weight.

[0183] In this embodiment of the disclosure, if the mapping relationship between performance error and weight is predetermined, and the performance error of the candidate microphone array in locating the sample sound source at the sample angle is determined, the weight corresponding to the sample angle for sound source localization by the candidate microphone array can be determined according to the mapping relationship between performance error and weight.

[0184] In some embodiments, the weight corresponding to the sample angle for sound source localization by each candidate microphone array is determined based on the performance error of each candidate microphone array in locating the sample sound source at the sample angle, including:

[0185] If the performance error of the candidate microphone array in locating the sample sound source at the sample angle is less than the first threshold, the weight corresponding to the sample angle for sound source localization by the candidate microphone array is determined as the first value.

[0186] If the performance error of the candidate microphone array in locating the sample sound source at the sample angle is greater than or equal to the first threshold and less than the second threshold, the weight corresponding to the sample angle for sound source localization by the candidate microphone array is determined as the second value.

[0187] If the performance error of the candidate microphone array in locating the sample sound source at the sample angle is greater than or equal to the second threshold, the weight corresponding to the sample angle for sound source localization by the candidate microphone array is determined as the third value.

[0188] In some embodiments, the first value is 1, the second value is 0.5, and the third value is 0.

[0189] In some embodiments, the first value is 1, the second value is 0.8, and the third value is 0.5.

[0190] In some embodiments, the first value is 1, the second value is 0.7, and the third value is 0.2.

[0191] In some embodiments, the first threshold is 10° and the second threshold is 50°, or the first threshold is 5° and the second threshold is 30°.

[0192] In some embodiments, the values ​​of the first threshold and the second threshold are related to the M performance errors of the M candidate microphone arrays in locating the sample sound source located at the sample angle.

[0193] In some embodiments, the first threshold is the median of the M performance errors of the M candidate microphone arrays in locating the sample sound source at the sample angle, sorted from best to worst, and the second threshold is the worst value among the M performance errors of the M candidate microphone arrays in locating the sample sound source at the sample angle.

[0194] In some embodiments, the first threshold is the value of the (M-1)th position of the M performance errors of the M candidate microphone arrays in locating the sample sound source at the sample angle, sorted from best to worst, and the second threshold is the value of the Mth position of the M performance errors of the M candidate microphone arrays in locating the sample sound source at the sample angle, sorted from best to worst.

[0195] It should be noted that the values ​​of the first value, the second value, the third value, the first threshold, and the second threshold in the above embodiments can also be other values ​​besides those in the above embodiments, and this disclosure does not impose specific restrictions on them.

[0196] In some embodiments, when the sample angles include a portion of the angles in the specified angle region, the performance error of each candidate microphone array in locating sample sound sources located at a portion of the angles in the specified angle region can be determined. Then, based on the performance error of each candidate microphone array in locating sample sound sources located at a portion of the angles in the specified angle region, the weights corresponding to each angle in the specified angle region for sound source localization by each candidate microphone array can be determined.

[0197] In some embodiments, determining the weights of each angle in a specified angular region for sound source localization of each candidate microphone array based on the weights of the sample angles for sound source localization of each candidate microphone array includes: determining the weights of each angle between two adjacent sample angles in a specified angular region for sound source localization of each candidate microphone array based on the weights of the adjacent two sample angles for sound source localization of each candidate microphone array.

[0198] In this embodiment of the present disclosure, the weights corresponding to each angle between two adjacent sample angles in the specified angular region for sound source localization of each candidate microphone array are determined based on the weights corresponding to the adjacent sample angles for sound source localization of each candidate microphone array.

[0199] In some embodiments, there are also K angles between adjacent sample angles, where K is an integer greater than 0.

[0200] For example, the weights corresponding to two adjacent sample angles are 0.8 and 0.4. There are also three angles between these two adjacent sample angles: angle 1, angle 2, and angle 3. It can be determined that the weight corresponding to angle 1 is 0.7, the weight corresponding to angle 2 is 0.6, and the weight corresponding to angle 3 is 0.5.

[0201] For example, the weights corresponding to two adjacent sample angles are 0.8 and 0.8. There are also three angles between these two adjacent sample angles: angle 1, angle 2, and angle 3. It can be determined that the weight corresponding to angle 1 is 0.8, the weight corresponding to angle 2 is 0.8, and the weight corresponding to angle 3 is 0.8.

[0202] It should be noted that the above examples are for illustrative purposes only and are not intended to limit the specific embodiments of this disclosure. The weight values ​​corresponding to two adjacent sample angles may also be other cases, and there may be more than three angles between two adjacent sample angles, such as four angles.

[0203] In this embodiment of the disclosure, after the audio signal of the sound source to be located is collected by the microphones in a determined array of multiple microphones, the sound source to be located can be located based on the audio signal and weight information corresponding to the multiple microphone arrays.

[0204] For example, when the specified angle range is horizontal angle [-180°, 180°] and pitch angle [-180°, 180°], multiple microphone arrays are determined to include microphone array 1, microphone array 2 and microphone array 3.

[0205] Specifically, the weight information for microphone array 1 is as follows: horizontal angle [-180°, 180°] = 1, pitch angle [-180°, 180°] = 0; the weight information for microphone array 2 is as follows: horizontal angle [-180°, 180°] = 0, pitch angle [-180°, 0°] = 1, pitch angle [0°, 90°] = 0.5; the weight information for microphone array 3 is as follows: horizontal angle [-180°, 180°] = 0, pitch angle [0°, 90°] = 0.5, pitch angle [90°, 180°] = 1.

[0206] In this process, after the audio signals of the sound source to be located are collected using the microphones in microphone arrays 1, 2, and 3, microphone array 1 processes the collected audio signals of the sound source to be located using a sound source localization algorithm to determine the horizontal angle as 85° and the pitch angle as 0°; microphone array 2 processes the collected audio signals of the sound source to be located using a sound source localization algorithm to determine the horizontal angle as 0° and the pitch angle as 60°; and microphone array 3 processes the collected audio signals of the sound source to be located using a sound source localization algorithm to determine the horizontal angle as 0° and the pitch angle as 80°.

[0207] Based on the audio signals and weight information corresponding to microphone arrays 1, 2, and 3, the location result of the sound source to be located is determined as follows: horizontal angle is 1*85°+0°*0+0°*0=85°, and pitch angle is 0°*0+60°*0.5+80°*0.5=70°. Therefore, the location result of the sound source to be located is determined to be 85° horizontal angle and 70° pitch angle.

[0208] It should be noted that the above examples are for illustrative purposes only and are not intended to limit the specific embodiments of this disclosure. The weight information can also be represented by the weight corresponding to the solid angle.

[0209] By implementing the embodiments of this disclosure, and by dividing the microphone array into a better configuration and setting corresponding weight information, the sound source to be located can be located, which can effectively reduce the positioning calculation error and improve the positioning accuracy.

[0210] It is understandable that multi-microphone devices face challenges such as performance fluctuations in algorithms at different angles and directions, as well as increased computational demands when multiple microphones are arranged in an array. In practical mobile phone array applications, diffraction and refraction effects caused by device obstruction, as well as differences in microphone aperture direction and microphone performance, severely impact the accuracy of positioning algorithms. To overcome this challenge, the embodiments disclosed in this paper introduce a weighting mechanism, allowing the contribution of each microphone to vary according to the characteristics of its region. By carefully selecting a microphone combination with better regional characteristics, computational errors can be effectively reduced, and positioning accuracy can be improved.

[0211] Furthermore, embodiments of this disclosure employ optimization algorithms to refine the partition boundaries. By analyzing the estimation and measurement errors of sound sources at different spatial locations, the configuration of the microphone array is adjusted to achieve the best partitioning effect. This strategy can significantly improve the robustness of the positioning algorithm, providing users with more accurate and reliable positioning services.

[0212] Meanwhile, the partitioning scheme has less dependence on the array shape. Arrays of different shapes can be partitioned into multiple smaller arrays with similar shapes, which reduces the need for modifications when porting the algorithm to different arrays.

[0213] For example, taking Figure 3 as an example, when running a sound source localization algorithm on a 4-mic terminal device (mobile phone), it is necessary to obtain the pitch angle [-90, 90] and horizontal angle [-90, 90] of the sound source in front, so as to obtain the angle information of the sound source relative to the device.

[0214] By running a sound source localization algorithm on an array of two microphones, the angle between the sound source and the line connecting the microphones can be obtained.

[0215] Assume the microphone layout of a 4-mic terminal device is as shown in Figure 3:

[0216] The device's microphones are divided into multiple sub-arrays in pairs. The performance and error of each array in different angular regions are analyzed and calculated. For example, mic1 & mic2 perform better in region a, mic2 & mic3 perform better in region b, and mic1 & mic4 perform better in region c.

[0217] Region a: Pitch angle [-90, 90].

[0218] Region b: Horizontal angle [-90, 10].

[0219] Region c: Horizontal angle [-10, 90].

[0220] Therefore, the mic is divided into 3 pairs of mics: mic1 & mic2 calculation area a, mic2 & mic3 calculation area b, and mic1 & mic4 calculation area c. By combining the results, all the required angles can be located.

[0221] At the same time, the results of the transitional regions are combined, and the contribution of different mics to the results is adjusted to obtain more accurate results.

[0222] This disclosure provides a method for processing audio signals:

[0223] 1. Zoning of sound sources at different azimuth angles.

[0224] 2. Divide the large array of microphones into multiple smaller arrays of microphones for computation.

[0225] 3. Based on the actual algorithm performance, divide the mic array into combinations and assign weights to different arrays.

[0226] This disclosure also proposes an apparatus (also referred to as a communication device, etc.) for implementing any of the above methods. For example, an apparatus is proposed, which includes units or modules for implementing the steps performed by the terminal in any of the above methods.

[0227] It should be understood that the division of units or modules in the above device is only a logical functional division. In actual implementation, they can be fully or partially integrated into a single physical entity, or they can be physically separated. Furthermore, the units or modules in the device can be implemented by a processor calling software: for example, the device includes a processor connected to a memory containing instructions. The processor calls the instructions stored in the memory to implement any of the above methods or to implement the functions of the units or modules in the above device. The processor can be, for example, a general-purpose processor, such as a Central Processing Unit (CPU) or a microprocessor, and the memory can be internal or external to the device. Alternatively, the units or modules in the device can be implemented in the form of hardware circuits. The functionality of some or all of the units or modules can be achieved through the design of these hardware circuits, which can be understood as one or more processors. For example, in one implementation, the hardware circuit is an application-specific integrated circuit (ASIC). The functionality of some or all of the units or modules is achieved through the design of the logical relationships between the components within the circuit. In another implementation, the hardware circuit can be implemented using a programmable logic device (PLD). Taking a field-programmable gate array (FPGA) as an example, it can include a large number of logic gates. The connection relationships between the logic gates are configured through a configuration file, thereby achieving the functionality of some or all of the units or modules. All units or modules of the above device can be implemented entirely through processor-called software, entirely through hardware circuits, or partially through processor-called software with the remaining parts implemented through hardware circuits.

[0228] In this embodiment, the processor is a circuit with signal processing capabilities. In one implementation, the processor can be a circuit with instruction read and execute capabilities, such as a Central Processing Unit (CPU), a microprocessor, a graphics processing unit (GPU) (which can be understood as a microprocessor), or a digital signal processor (DSP). In another implementation, the processor can implement certain functions through the logical relationships of hardware circuits. The logical relationships of the aforementioned hardware circuits are fixed or reconfigurable. For example, the processor is a hardware circuit implemented using an application-specific integrated circuit (ASIC) or a programmable logic device (PLD), such as an FPGA. In a reconfigurable hardware circuit, the process of the processor loading a configuration document and configuring the hardware circuit can be understood as the process of the processor loading instructions to implement the functions of some or all of the above units or modules. Furthermore, it can also be a hardware circuit designed for artificial intelligence, which can be understood as an ASIC, such as a Neural Network Processing Unit (NPU), a Tensor Processing Unit (TPU), or a Deep Learning Processing Unit (DPU).

[0229] Figure 5 is a schematic diagram of the structure of the sound source localization device proposed in an embodiment of this disclosure. As shown in Figure 5, the sound source localization device 10 may include at least one of a signal acquisition module 11 and a localization processing module 12.

[0230] In some embodiments, the signal acquisition module 11 is used to acquire the audio signal of the sound source to be located collected by the microphones in a plurality of microphone arrays;

[0231] The positioning processing module 12 is used to locate the sound source to be located based on the audio signals and weight information corresponding to multiple microphone arrays.

[0232] In some embodiments, the communication device is the sound source localization device described in the above embodiments.

[0233] In some embodiments, the weight information corresponding to the microphone array is the target weight corresponding to the target angle region where the microphone array can perform sound source localization.

[0234] In some embodiments, the positioning processing module 12 is specifically used for:

[0235] Determine the performance error of locating the sample sound source located at the sample angle for each of the M candidate microphone arrays;

[0236] Based on the performance error of each candidate microphone array in locating the sample sound source located at the sample angle, determine the weights of each angle in the specified angle region for sound source localization of each candidate microphone array.

[0237] Based on the weights of each angle in the specified angular region for sound source localization using M candidate microphone arrays, determine multiple microphone arrays and the target weights corresponding to the target angular region for sound source localization of each microphone array.

[0238] The M candidate microphone arrays are obtained by grouping the N candidate microphones used to locate the sound source, where N is an integer greater than 2.

[0239] In some embodiments, the sum of the angle regions formed by the target angle regions for sound source localization by the multiple microphone arrays is equal to the specified angle region, and the sum of the target weights corresponding to each angle in the specified angle region for sound source localization by the multiple microphone arrays is 1.

[0240] In some embodiments, the positioning processing module 12 is specifically used to: acquire sample audio signals of sample sound sources located at sample angles collected by each candidate microphone array; determine the sample positioning angle of the located sample sound source based on the sample audio signals; and determine the performance error of each candidate microphone array in locating the sample sound source located at the sample angle based on the sample angle and sample positioning angle corresponding to each candidate microphone array.

[0241] In some embodiments, the number of sample angles is multiple, and the sample angles include all angles of a specified angle region; or the sample angles include a portion of the angles of a specified angle region.

[0242] In some embodiments, the localization processing module 12 is specifically used to: determine the weight corresponding to the sample angle for sound source localization of each candidate microphone array based on the performance error of each candidate microphone array in locating the sample sound source located at the sample angle.

[0243] Based on the weights corresponding to the sample angles used for sound source localization by each candidate microphone array, determine the weights corresponding to each angle in the specified angular region for sound source localization by each candidate microphone array.

[0244] In some embodiments, the positioning processing module 12 is specifically used for:

[0245] If the performance error of the candidate microphone array in locating the sample sound source at the sample angle is less than the first threshold, the weight corresponding to the sample angle for sound source localization by the candidate microphone array is determined as the first value.

[0246] If the performance error of the candidate microphone array in locating the sample sound source at the sample angle is greater than or equal to the first threshold and less than the second threshold, the weight corresponding to the sample angle for sound source localization by the candidate microphone array is determined as the second value.

[0247] If the performance error of the candidate microphone array in locating the sample sound source at the sample angle is greater than or equal to the second threshold, the weight corresponding to the sample angle for sound source localization by the candidate microphone array is determined as the third value.

[0248] In some embodiments, the positioning processing module 12 is specifically used to: determine the weight corresponding to the sample angle for sound source positioning of each candidate microphone array based on the performance error of each candidate microphone array in locating the sample sound source at the sample angle, and the mapping relationship between the performance error and the weight.

[0249] In some embodiments, the positioning processing module 12 is specifically used to: determine the weights of each angle between two adjacent sample angles in a specified angular region for sound source localization of each candidate microphone array, based on the weights corresponding to the adjacent sample angles for sound source localization of each candidate microphone array.

[0250] Figure 6A is a schematic diagram of the structure of the communication device 5100 proposed in an embodiment of this disclosure. The communication device 5100 can be a sound source localization device (e.g., a terminal), or a chip, chip system, or processor that supports the sound source localization device in implementing any of the above methods. The communication device 5100 can be used to implement the methods described in the above method embodiments, and for details, please refer to the description in the above method embodiments.

[0251] As shown in Figure 6A, the communication device 5100 is used to execute any of the above methods. In some embodiments, the communication device 5100 includes one or more processors 5101. The processor 5101 may be a general-purpose processor or a special-purpose processor, such as a baseband processor or a central processing unit. The baseband processor may be used to process communication protocols and communication data, and the central processing unit may be used to control communication devices (e.g., base stations, baseband chips, terminals, terminal chips, DUs or CUs, etc.), execute programs, and process program data. Optionally, the communication device 5100 is used to execute any of the above methods. Optionally, one or more processors 5101 are used to invoke instructions to cause the communication device 5100 to execute any of the above methods.

[0252] In some embodiments, the communication device 5100 further includes one or more transceivers 5102. When the communication device 5100 includes one or more transceivers 5102, the transceiver 5102 performs at least one of the communication steps such as sending and / or receiving in the above-described method (e.g., the sending and / or receiving steps in S201-S203, but not limited thereto), and the processor 5101 performs at least one of other steps (e.g., steps other than sending and / or receiving in S201-S203, but not limited thereto). In optional embodiments, the transceiver may include a receiver and / or a transmitter, which may be separate or integrated. Optionally, the terms transceiver, transceiver unit, transceiver, transceiver circuit, interface circuit, interface, etc., can be used interchangeably; the terms transmitter, sending unit, transmitter, sending circuit, etc., can be used interchangeably; the terms receiver, receiving unit, receiver, receiving circuit, etc., can be used interchangeably.

[0253] In some embodiments, the communication device 5100 further includes one or more memories 5103 for storing data and / or instructions. Optionally, one or more processors 5101 are used to invoke instructions stored in the memory 5103 to cause the communication device 5100 to perform any of the above methods. Optionally, all or part of the memory 5103 may also be located outside the communication device 5100. In an optional embodiment, the communication device 5100 may include one or more interface circuits 5104. Optionally, the interface circuit 5104 is connected to the memory 5103 and can be used to receive data and / or instructions from the memory 5103 or other devices, and can be used to send data and / or instructions to the memory 5103 or other devices. For example, the interface circuit 5104 can read data and / or instructions stored in the memory 5103 and send the data and / or instructions to the processor 5101.

[0254] The communication device 5100 described in the above embodiments may be a network device or a terminal, but the scope of the communication device 5100 described in this disclosure is not limited thereto, and the structure of the communication device 5100 may not be limited by FIG. 6A. The communication device may be a standalone device or a part of a larger device. For example, the communication device may be: (1) a standalone integrated circuit IC, or chip, or chip system or subsystem; (2) a collection of one or more ICs, optionally, the IC collection may also include storage components for storing data, programs and / or instructions; (3) an ASIC, such as a modem; (4) a module that can be embedded in other devices; (5) a receiver, terminal, smart terminal, cellular phone, wireless device, handheld device, mobile unit, vehicle device, network device, cloud device, artificial intelligence device, etc.; (6) others, etc.

[0255] Figure 6B is a schematic diagram of the structure of the chip 5200 proposed in an embodiment of this disclosure. For cases where the communication device 5100 can be a chip or a chip system, the schematic diagram of the chip 5200 shown in Figure 6B can be referenced, but is not limited thereto.

[0256] Chip 5200 includes one or more processors 5201. Chip 5200 is used to perform any of the methods described above.

[0257] In some embodiments, chip 5200 further includes one or more interface circuits 5202. Optionally, terms such as interface circuit, interface, and transceiver pin can be used interchangeably. In some embodiments, chip 5200 further includes one or more memories 5203 for storing data and / or instructions. Optionally, all or part of the memories 5203 may be located outside of chip 5200. Optionally, the interface circuit 5202 is connected to the memories 5203, and the interface circuit 5202 can be used to receive data and / or instructions from the memories 5203 or other devices, and the interface circuit 5202 can be used to send data and / or instructions to the memories 5203 or other devices. For example, the interface circuit 5202 can read data and / or instructions stored in the memories 5203 and send the data and / or instructions to the processor 5201.

[0258] In some embodiments, the interface circuit 5202 performs at least one of the communication steps such as sending and / or receiving in the above-described method (e.g., the sending and / or receiving steps in S201 to S203, but not limited thereto). The interface circuit 5202 performing the communication steps such as sending and / or receiving in the above-described method refers, for example, to the interface circuit 5202 performing data and / or instruction interaction between the processor 5201, the chip 5200, the memory 5203, or the transceiver device. In some embodiments, the processor 5201 performs at least one of other steps (e.g., steps other than sending and / or receiving in S201 to S203, but not limited thereto).

[0259] The modules and / or devices described in the various embodiments, such as virtual devices, physical devices, and chips, can be combined or separated arbitrarily as needed. Optionally, some or all steps can also be performed collaboratively by multiple modules and / or devices, which is not limited here.

[0260] This disclosure also proposes a storage medium storing instructions that, when executed on a communication device, cause the communication device to perform any of the above methods. Optionally, the storage medium is an electronic storage medium. Optionally, the storage medium is a computer-readable storage medium, but not limited thereto; it may also be a storage medium readable by other devices. Optionally, the storage medium may be a non-transitory storage medium, but not limited thereto; it may also be a temporary storage medium.

[0261] This disclosure also proposes a program product, including a program and / or instructions, which, when executed by a communication device, cause the communication device to perform any of the above methods. Optionally, the program product is a computer program product. Optionally, the program product is stored on the storage medium.

[0262] This disclosure also proposes a computer program that, when run on a computer, causes the computer to perform any of the above methods.

[0263] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this disclosure.

[0264] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.

[0265] The above description is merely a specific embodiment of this disclosure, but the scope of protection of this disclosure is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this disclosure should be included within the scope of protection of this disclosure. Therefore, the scope of protection of this disclosure should be determined by the scope of the claims.

Claims

1. A method for locating a sound source, characterized in that, include: Acquire audio signals from the microphones in a multi-microphone array that are collected from the sound source to be located; The sound source to be located is located based on the audio signals and weight information corresponding to the multiple microphone arrays.

2. The method as described in claim 1, characterized in that, The weight information corresponding to the microphone array is the target weight corresponding to the target angle region where the microphone array can locate the sound source.

3. The method as described in claim 2, characterized in that, The method further includes: Determine the performance error of locating the sample sound source located at the sample angle for each of the M candidate microphone arrays; Based on the performance error of each candidate microphone array in locating the sample sound source located at the sample angle, the weights corresponding to each angle in the specified angle region for sound source localization by each candidate microphone array are determined. Based on the weights corresponding to each angle in the specified angle region for sound source localization using the M candidate microphone arrays, the plurality of microphone arrays are determined, as well as the target weights corresponding to the target angle regions for sound source localization of each microphone array. The M candidate microphone arrays are obtained by grouping the N candidate microphones used to locate the sound source into groups, where N is an integer greater than 2.

4. The method as described in claim 3, characterized in that, The sum of the angle regions formed by the target angle regions for sound source localization by the multiple microphone arrays is equal to the specified angle region, and the sum of the target weights corresponding to each angle in the specified angle region for sound source localization by the multiple microphone arrays is 1.

5. The method as described in claim 3 or 4, characterized in that, The performance error of determining each of the M candidate microphone arrays for locating the sample sound source at the sample angle includes: Acquire the sample audio signal of the sample sound source located at the sample angle collected by each of the candidate microphone arrays; Based on the sample audio signal, determine the sample positioning angle of the located sample sound source; Based on the sample angle and the sample positioning angle corresponding to each candidate microphone array, the performance error of each candidate microphone array in locating the sample sound source located at the sample angle is determined.

6. The method according to any one of claims 3 to 5, characterized in that, The number of sample angles is multiple, and the sample angles include all angles of the specified angle region; or the sample angles include a portion of the angles of the specified angle region.

7. The method according to any one of claims 3 to 6, characterized in that, The step of determining the weights corresponding to each angle in the specified angular region for sound source localization by each candidate microphone array based on the performance error of localizing the sample sound source at the sample angle for each candidate microphone array includes: Based on the performance error of each candidate microphone array in locating the sample sound source at the sample angle, the weight corresponding to the sample angle for sound source localization of each candidate microphone array is determined. Based on the weights corresponding to the sample angles used for sound source localization by each candidate microphone array, the weights corresponding to each angle in the specified angle region for sound source localization by each candidate microphone array are determined.

8. The method as described in claim 7, characterized in that, The step of determining the weight corresponding to the sample angle for sound source localization by each candidate microphone array based on the performance error of localizing the sample sound source at the sample angle of each candidate microphone array includes: If the performance error of the candidate microphone array in locating the sample sound source at the sample angle is less than a first threshold, the weight corresponding to the sample angle for sound source localization by the candidate microphone array is determined to be a first value. If the performance error of the candidate microphone array in locating the sample sound source at the sample angle is greater than or equal to the first threshold and less than the second threshold, the weight corresponding to the sample angle for sound source localization by the candidate microphone array is determined to be the second value. If the performance error of the candidate microphone array in locating the sample sound source at the sample angle is greater than or equal to the second threshold, the weight corresponding to the sample angle for sound source localization by the candidate microphone array is determined as the third value.

9. The method as described in claim 7, characterized in that, The step of determining the weight corresponding to the sample angle for sound source localization by each candidate microphone array based on the performance error of localizing the sample sound source at the sample angle of each candidate microphone array includes: Based on the performance error of each candidate microphone array in locating a sample sound source at a sample angle, and the mapping relationship between the performance error and the weight, the weight corresponding to the sample angle for sound source localization by each candidate microphone array is determined.

10. The method according to any one of claims 7 to 9, characterized in that, The step of determining the weights corresponding to each angle in a specified angular region for sound source localization of each candidate microphone array based on the weights corresponding to the sample angles used for sound source localization of each candidate microphone array includes: Based on the weights corresponding to the two adjacent sample angles for sound source localization of each candidate microphone array, determine the weights corresponding to each angle between two adjacent sample angles in the specified angle region for sound source localization of each candidate microphone array.

11. A sound source localization device, characterized in that, The device includes: The signal acquisition module is used to acquire the audio signal of the sound source to be located collected by the microphones in the multiple microphone arrays; The positioning processing module is used to locate the sound source to be located based on the audio signals and weight information corresponding to the multiple microphone arrays.

12. A storage medium storing instructions, characterized in that, When the instructions are executed on a communication device, the communication device performs the method as described in any one of claims 1 to 10.

13. A program product comprising at least one of a program and instructions, characterized in that, When at least one of the programs or instructions is executed by a communication device, it implements the method of any one of claims 1 to 10.