Sound source positioning method, electronic equipment and storage medium
By setting up microphone arrays in the head-mounted device and necklace, and combining relative pose and sound source localization auxiliary data, an enhanced stereo microphone array is constructed, which solves the problem of inaccurate vertical and rear positioning of the head-mounted device and achieves high-precision sound source localization in the entire space range.
Patent Information
- Application Number
- CN202511599068.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-04
- Publication Date
- 2025-12-05
- Estimated Expiration
- 2045-11-04
AI Technical Summary
Head-mounted devices such as AR/VR devices have a limited number of microphones and narrow spacing, making it difficult to form an effective aperture in the vertical direction. This makes it impossible to accurately distinguish the height information of the sound source, and the user's head occlusion effect affects the accuracy of the sound source positioning behind when wearing the device.
By setting up microphone arrays in the head-mounted device and the necklace, and utilizing the relative pose and sound source localization auxiliary data provided by the necklace, a unified spatial reference system is established, forming an enhanced stereo microphone array to compensate for the head-mounted device's insufficient positioning in the vertical direction and rear.
It achieves high-precision sound source localization across the entire space range, solving the problems of fuzzy vertical positioning and inaccurate rear positioning of head-mounted devices, and significantly improving the accuracy of sound source localization.
Smart Images

Figure CN121069312A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of acoustic signal processing technology, and in particular to a sound source localization method, electronic device and storage medium. Background Technology
[0002] Currently, head-mounted devices such as AR (Augmented Reality) / VR (Virtual Reality) devices and smart glasses have limited capacity due to size and design constraints, resulting in a limited number of microphones and narrow spacing between them. This leads to a microphone array geometry that is typically close to a planar ring array, making it difficult to create an effective aperture in the vertical direction. Consequently, it becomes impossible to accurately distinguish the height information of a sound source, resulting in vertical positioning ambiguity. Secondly, when worn, head-mounted devices are close to the user's head, which significantly obstructs sound waves. This obstruction particularly affects sound signals from behind the wearer, causing distortion or attenuation of the rear sound source signals received by the microphone array, thus significantly reducing the accuracy of rear sound source positioning.
[0003] The above content is only used to help understand the technical solution of this application and does not represent an admission that the above content is prior art. Summary of the Invention
[0004] The main objective of this application is to provide a sound source localization method, electronic device, and storage medium, aiming to solve the technical problem of low sound source localization accuracy in current head-mounted devices.
[0005] To achieve the above objectives, this application proposes a sound source localization method, which is applied to a head-mounted device. The head-mounted device includes a first microphone array, and a necklace connected to the head-mounted device has a second microphone array. The sound source localization method includes: Acquire the first microphone array signal collected by the first microphone array; The relative pose between the necklace and the head-mounted device is obtained, as well as the sound source localization assistance data sent by the necklace is obtained, wherein the sound source localization assistance data is obtained based on the second microphone array signal collected by the second microphone array; The sound source is located based on the first microphone array signal, the sound source localization auxiliary data, and the relative pose, and the sound source localization result is obtained.
[0006] Optionally, the sound source localization auxiliary data includes a candidate direction set obtained by localizing the sound source based on the signal from the second microphone array; The step of performing sound source localization based on the first microphone array signal, the sound source localization auxiliary data, and the relative pose to obtain the sound source localization result includes: Based on the signal from the first microphone array, the sound source is located to obtain a preliminary positioning direction; After transforming the preliminary positioning direction and each candidate direction in the candidate direction set to the same coordinate system based on the relative pose, the preliminary positioning direction and each candidate direction are weighted and summed to obtain the sound source localization result.
[0007] Optionally, the sound source localization auxiliary data further includes a second TDOA matrix calculated based on the signal of the second microphone array, wherein the second TDOA matrix includes the measured time difference of arrival for each microphone pair in the second microphone array; The step of obtaining the sound source localization result by weighted summation of the preliminary localization direction and each of the candidate directions includes: For each direction pair formed by combining the initial positioning direction with each candidate direction, a consistency score for the direction pair is calculated based on the second TDOA matrix and the first TDOA matrix calculated based on the signal of the first microphone array. The first TDOA matrix includes the measured time difference of arrival corresponding to each microphone pair in the first microphone array. The consistency score of the direction pair is used as the weight of the candidate direction in the direction pair, and the candidate directions are weighted and summed to obtain the first fusion direction; The sound source localization result is obtained by weighted summation of the first fusion direction and the preliminary positioning direction.
[0008] Optionally, the step of weighted summing of the first fusion direction and the preliminary positioning direction to obtain the sound source localization result includes: If both the first fusion direction and the preliminary positioning direction fall within the target direction range, then the first weight is used as the weight corresponding to the first fusion direction, and the first fusion direction and the preliminary positioning direction are weighted and summed to obtain the sound source localization result. The target direction range is a cone-shaped or fan-shaped region around the target direction in the coordinate system of the head-mounted device, and is a subset of the full-directional range of the sound source localization of the head-mounted device. The target direction faces the back of the wearer's head when the head-mounted device is worn. If at least one of the first fusion direction and the preliminary positioning direction does not fall within the target direction range, then the second weight is used as the weight corresponding to the first fusion direction, and the first fusion direction and the preliminary positioning direction are weighted and summed to obtain the sound source localization result, wherein the second weight is less than the first weight.
[0009] Optionally, the step of obtaining the sound source localization result by weighted summation of the preliminary positioning direction and each of the candidate directions includes: The preliminary fusion direction is obtained by weighted summation of the preliminary positioning direction and each of the candidate directions; The third TDOA matrix is calculated based on the initial fusion direction, wherein the third TDOA matrix includes the theoretical arrival time difference of each target microphone pair under the assumption that the sound source comes from the initial fusion direction, and each target microphone pair is a microphone pair composed of two microphones in the first microphone array and two microphones in the second microphone array; The spectral characteristics, spatial filter response map, and ambient noise spectrum are calculated based on the signals from the first microphone array. The third TDOA matrix, the spectral features, the spatial filter response map, and the environmental noise spectrum are input into a preset deep learning model for processing to obtain the sound source localization result.
[0010] Optionally, the sound source localization method further includes: If the preliminary positioning direction is obtained, but the sound source positioning auxiliary data sent by the necklace has not yet been received, the preliminary positioning direction is output as the sound source positioning result. Upon receiving the sound source localization assistance data sent by the necklace, the sound source localization result is output by weighted summation of the preliminary localization direction and each of the candidate directions.
[0011] Optionally, the sound source localization method further includes: Obtain the clock offset from the necklace; The step of performing sound source localization based on the first microphone array signal, the sound source localization auxiliary data, and the relative pose to obtain the sound source localization result includes: Data time alignment is performed based on the acquisition time of the first microphone array signal, the timestamp carried in the sound source localization auxiliary data, and the clock offset. Sound source localization is then performed based on the aligned data and the relative pose to obtain the sound source localization result.
[0012] Furthermore, to achieve the above objectives, this application also proposes a sound source localization method, which is applied to a necklace, wherein the necklace is provided with a second microphone array, and a head-mounted device that establishes a communication connection with the necklace is provided with a first microphone array. The sound source localization method includes: Acquire the second microphone array signal collected by the second microphone array; Sound source localization auxiliary data is obtained based on the signals from the second microphone array; The sound source localization auxiliary data is sent to the head-mounted device so that the head-mounted device can perform sound source localization based on the first microphone array signal, the sound source localization auxiliary data, and the relative pose between the head-mounted device and the necklace, and obtain the sound source localization result, wherein the first microphone array signal is obtained based on the first microphone array.
[0013] In addition, to achieve the above objectives, this application also proposes an electronic device, which further includes: a memory, a processor, and a computer program stored in the memory and executable on the processor, the computer program being configured to implement the steps of the sound source localization method as described above.
[0014] In addition, to achieve the above objectives, this application also proposes a storage medium, which is a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it implements the steps of the sound source localization method described above.
[0015] In addition, to achieve the above objectives, this application also provides a computer program product, which includes a computer program that, when executed by a processor, implements the steps of the sound source localization method described above.
[0016] One or more technical solutions proposed in this application have at least the following technical effects: By acquiring the signal from the first microphone array of the head-mounted device, basic data for sound source localization is established. By acquiring the relative pose of the necklace and the head-mounted device and auxiliary data for sound source localization, a unified spatial reference system is established and auxiliary spatial information of the necklace is introduced. By performing sound source localization based on the first microphone array signal, auxiliary data for sound source localization, and relative pose, the microphone arrays of the head-mounted device and the necklace are virtually integrated into an enhanced array with a better spatial layout. Furthermore, the natural height difference of the necklace in the vertical direction can be used to solve the problem that the head-mounted device cannot distinguish between upper and lower sounds, and the microphone unit of the necklace located at the back of the neck can be used to avoid head obstruction, significantly improving the localization accuracy of rear sound sources, and ultimately achieving more accurate sound source localization in the entire space. Attached Figure Description
[0017] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.
[0018] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0019] Figure 1 This is a schematic diagram showing the layout of a microphone array in a head-mounted device related to the sound source localization method of this application; Figure 2 This is a schematic diagram showing the placement of a microphone array in a necklace, which is involved in the sound source localization method of this application. Figure 3 This is a flowchart illustrating the first embodiment of the sound source localization method of this application; Figure 4 This is a flowchart illustrating the fourth embodiment of the sound source localization method of this application; Figure 5 This is a schematic diagram of the device structure of the hardware operating environment involved in the sound source localization method in the embodiments of this application.
[0020] The purpose, features, and advantages of this application will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation
[0021] It should be understood that the specific embodiments described herein are merely illustrative of the technical solutions of this application and are not intended to limit this application.
[0022] To better understand the technical solution of this application, a detailed description will be provided below in conjunction with the accompanying drawings and specific implementation methods.
[0023] Microphone arrays typically have a geometry close to a planar ring array, making it difficult to create an effective aperture in the vertical direction. This makes it impossible to accurately distinguish the height information of a sound source, resulting in vertical positioning ambiguity. Secondly, when worn, head-mounted devices are close to the user's head, which significantly obstructs sound waves. This obstruction particularly affects sound signals from behind the wearer, causing distortion or attenuation of the rear sound source signal received by the microphone array, thus significantly reducing the accuracy of rear sound source localization.
[0024] This application provides a solution that establishes basic data for sound source localization by acquiring the first microphone array signal of a head-mounted device; establishes a unified spatial reference system and introduces auxiliary spatial information of the necklace by acquiring the relative pose of the necklace and the head-mounted device and the auxiliary spatial information of the necklace; and virtually integrates the microphone arrays of the head-mounted device and the necklace into an enhanced array with a better spatial layout by performing sound source localization based on the first microphone array signal, the auxiliary sound source localization data, and the relative pose. Furthermore, the natural height difference of the necklace in the vertical direction can be used to solve the problem that the head-mounted device cannot distinguish between upper and lower sounds, and the microphone unit of the necklace located at the back of the neck can be used to avoid head obstruction, significantly improving the localization accuracy of rear sound sources, and ultimately achieving more accurate sound source localization in the entire space.
[0025] In other words, the embodiments of this application combine a head-mounted device with a smart necklace to form a large-scale stereo microphone array system, which can overcome the above-mentioned problems caused by using head-mounted devices alone, achieve high-precision sound source localization in all directions, and solve the problems of inability to distinguish between up and down and insufficient accuracy in backward positioning.
[0026] The following presents a first embodiment of the sound source localization method of this application. In this embodiment, the sound source localization method is executed by a head-mounted device. It is understood that there are many types of head-mounted devices, such as smart glasses, VR (Virtual Reality) / AR (Augmented Reality) devices, etc. Each type of head-mounted device has many different hardware architectures and software system implementations. The embodiments of this application do not limit the type of head-mounted device, hardware architecture, or software system implementation of the sound source localization method. In this embodiment, a microphone array (hereinafter referred to as the first microphone array for distinction) is provided in the head-mounted device, and a second microphone array (hereinafter referred to as the second microphone array for distinction) is provided in the necklace. The head-mounted device can establish a communication connection with the necklace and use the second microphone array in the necklace to assist in sound source localization, thereby improving the accuracy of sound source localization. A microphone array is a collection of at least two microphones arranged in a certain manner. In this embodiment, the number of microphones contained in the first microphone array and the second microphone array is not limited, nor is the arrangement of the first microphone array in the head-mounted device or the arrangement of the second microphone array in the necklace limited. In one feasible implementation, the first microphone array on the head-mounted device can adopt a circular layout, for example, arranging 3 to 5 microphones on the frame of smart glasses. These microphones are arranged in a circular pattern on a horizontal plane, without needing to deliberately create a height difference in the vertical direction. This layout requires as few as 3 microphones to achieve accurate sound source localization, thereby reducing constraints on the appearance design of the glasses and allowing for more flexible and aesthetically pleasing designs, meeting the aesthetic requirements of glasses as decorative items. For example, Figure 1 An example of distributing four microphones on the frame of smart glasses is given. Figure 1 A microphone is placed at positions A, B, C, and D, forming a first microphone array. In one feasible embodiment, the second microphone array on the necklace can be designed to contain 8 to 12 microphones, distributed around the user's neck when worn. Because the necklace wraps around the neck, some microphones are located at the back of the neck, thus enabling the capture of sound from behind, compensating for the head-mounted device's inaccurate localization of rear sound sources. Simultaneously, the necklace naturally creates a vertical height difference when worn, which helps distinguish the vertical direction of sound, solving the problem of ambiguous vertical localization in head-mounted devices. For example, Figure 2An example of distributing eight microphones on a necklace is given.
[0027] This embodiment does not limit the material or appearance design of the necklace; it can be made of either flexible or rigid materials.
[0028] In one feasible implementation, the head-mounted device, such as smart glasses, can be equipped with a high-performance AI chip that supports real-time neural network inference, multi-microphone signal fusion processing, and user interaction, and features a 3-5 mic ring-shaped planar array. The necklace can utilize an integrated low-power MCU and a small DSP, featuring an 8-12 mic ring-shaped planar array, responsible for raw audio acquisition and front-end preprocessing.
[0029] Reference Figure 3 , Figure 3 This is a flowchart illustrating the first embodiment of the sound source localization method of this application. In this embodiment, the sound source localization method includes steps S10 to S30: Step S10: Obtain the first microphone array signal collected by the first microphone array.
[0030] The microphone array signal acquired by the first microphone array is referred to as the first microphone array signal for distinction. The first microphone array signal may include the sound signals acquired by each microphone. These signals are raw audio data, represented as a time series, and are used for subsequent sound source localization processing, such as estimating the direction of the sound source by analyzing the time difference or phase difference between the microphones.
[0031] Step S20: Obtain the relative pose between the necklace and the head-mounted device, and obtain the sound source localization auxiliary data sent by the necklace, wherein the sound source localization auxiliary data is obtained based on the second microphone array signal collected by the second microphone array.
[0032] Relative pose refers to the relative position and orientation between the head-mounted device and the necklace. Specifically, it can be represented by translation vectors and rotation matrices. Translation vectors describe the spatial displacement between the two, while rotation matrices describe the difference in their orientation.
[0033] There are several ways to obtain relative pose: for example, a default value can be preset, assuming that the relative pose of the head-mounted device and the necklace remains constant when the user wears them; or, joint calibration can be performed using the inertial measurement unit (IMU) and ultra-wideband (UWB) module in both the head-mounted device and the necklace. The joint calibration process may include: using UWB to measure the three-dimensional spatial distance and orientation between the head-mounted device and the necklace, obtaining the position coordinates of the necklace in the head-mounted device coordinate system, i.e., the translation vector, based on the three-dimensional spatial distance and orientation; using the orientation assistance information provided by UWB and the poses of the head-mounted device and the necklace themselves measured by the IMU, calculating the relative rotation matrix required to transform the orientation from the necklace coordinate system to the head-mounted device coordinate system using a sensor fusion algorithm.
[0034] In one feasible implementation, with joint calibration using an IMU and UWB, the IMU in the head-mounted device can monitor the intensity of the user's movement, for example, by determining the intensity of movement based on the acceleration amplitude. When the movement is intense, the relative pose may change rapidly, thus increasing the update frequency; conversely, the update frequency is decreased to conserve resources. Simultaneously, the head-mounted device's battery level can also affect the update frequency: when the battery is low, the update frequency is reduced to extend battery life. The accuracy of relative pose is crucial for sound source localization; therefore, periodic calibration or adjustment based on environmental changes is necessary.
[0035] The microphone array signal acquired by the second microphone array is referred to as the second microphone array signal for distinction. The second microphone array signal can include the sound signals acquired by each microphone.
[0036] In one feasible implementation, the sound source localization auxiliary data may be the signal of the second microphone array itself, i.e., the raw audio data collected by the second microphone array in the necklace; or, the sound source localization auxiliary data may be a sound source direction estimate calculated based on the signal of the second microphone array, such as a set of candidate directions.
[0037] In one feasible implementation, Bluetooth 5.3 can be used to achieve high-speed, low-latency data transmission between the head-mounted device and the necklace. The data transmission from the necklace to the head-mounted device uses LE Audio LC3 encoding to transmit sound source localization assistance data (such as candidate direction sets, TDOA matrices, timestamps, etc.), enabling the transmission of low-bitrate, high-fidelity compressed audio files. The data channel can use GATT+L2CAP custom services to transmit structured data. LE Audio refers to Low Energy Audio, a next-generation Bluetooth audio technology standard developed by the Bluetooth Special Interest Group (SIG). LC3 encoding is a new type of audio encoder introduced in the LE Audio standard. GATT (General Attribute Profile) defines a structured data service model for creating custom services (such as "sound source localization assistance data service") and feature values to encapsulate and identify the structured data to be transmitted (such as candidate direction sets, TDOA matrices, etc.). L2CAP (Logical Link Control and Adaptation Protocol) manages the data channel between Bluetooth devices, providing reliable or unreliable data packet transmission. By combining the GATT-defined data structure and L2CAP to establish an efficient data channel, reliable and orderly transmission of sound source localization assistance data can be achieved.
[0038] In one feasible implementation, the necklace can run a Voice Activity Detection (VAD) algorithm, which analyzes audio signals collected by a second microphone array in real time to determine whether human voice is present in the current environment. The necklace only initiates subsequent sound source localization auxiliary data calculation processes (such as generating candidate direction sets or TDOA matrices) and sends this data to the head-mounted device when the VAD algorithm detects voice activity. When silent or in the absence of ambient noise, it remains in sleep or low-power listening mode. Because there is no need for continuous, computationally demanding sound source localization data generation and wireless transmission during periods without voice, the necklace's processor and communication module can remain in a low-power state most of the time, effectively extending the necklace's battery life as a portable device. Secondly, it also reduces unnecessary wireless data transmission between the head-mounted device and the necklace.
[0039] The necklace only performs simple preprocessing, reducing power consumption; while the main data processing tasks are handled by the head-mounted device, which ensures both efficient system operation and extended battery life.
[0040] Step S30: Perform sound source localization based on the first microphone array signal, the sound source localization auxiliary data, and the relative pose to obtain the sound source localization result.
[0041] The sound source localization result can be the direction of the sound source obtained from the final localization, for example, the direction in three-dimensional space represented by a unit vector, or represented by spherical coordinates such as azimuth and elevation.
[0042] There are many specific implementation methods for sound source localization based on the first microphone array signal, sound source localization auxiliary data, and relative pose, and this embodiment is not limited to any particular method. For example, in one feasible embodiment, the sound source localization auxiliary data can be the second microphone array signal itself, and the relative pose can be a translation vector and a rotation matrix: using the relative pose, especially the rotation matrix, the spatial coordinates of all microphones in the second microphone array are transformed and unified into the coordinate system of the head-mounted device. In this way, all the microphones in the first and second microphone arrays seem to constitute a larger "virtual microphone array" in the coordinate system of the head-mounted device; the first microphone array signal and the second microphone array signal are considered as a whole signal, and a sound source localization algorithm (such as beamforming based on the entire virtual array or the TDOA method) is used to calculate the sound source direction, thereby obtaining the sound source localization result.
[0043] Using sound source localization auxiliary data can yield more accurate sound source localization results because the second microphone array provides additional spatial information, especially by capturing sound from different locations. This reduces blind spots or errors that occur when the head-mounted device is used for localization alone. For example, the microphone array in a necklace covers the rear and vertical directions, compensating for the limitations of the head-mounted device. Fusion of multi-microphone array data can improve robustness, especially in noisy environments.
[0044] In one feasible implementation, the sound source localization auxiliary data includes a candidate direction set obtained by sound source localization based on the second microphone array signal. Each candidate direction in the candidate direction set can be obtained by applying a sound source localization algorithm such as the generalized cross-correlation function (GCC-PHAT) to the second microphone array signal. The algorithm calculates a set of possible sound source directions based on the time difference or phase difference between the microphones, and each direction corresponds to an estimated value. Step S30 includes S301~S302: Step S301: Based on the signal from the first microphone array, perform sound source localization to obtain a preliminary localization direction.
[0045] The initial localization direction is calculated by applying a sound source localization algorithm such as SRP-PHAT (Steered Response Power-Phase Transform) to the signal from the first microphone array. The first microphone array may generate multiple candidate directions, and the initial localization direction can be the one with the highest confidence, for example, by selecting based on signal strength or consistency score, to ensure that usable localization results are provided even when using the head-mounted device alone.
[0046] Step S302: After transforming the preliminary positioning direction and each candidate direction in the candidate direction set to the same coordinate system according to the relative pose, the preliminary positioning direction and each candidate direction are weighted and summed to obtain the sound source localization result.
[0047] In one feasible implementation, the relative pose can be represented using translation vectors and rotation matrices. When transforming the candidate direction set to the head-mounted device coordinate system, a rotation matrix is applied to rotate each candidate direction so that all directions are represented in the head-mounted device coordinate system. Then, the initial positioning direction and each candidate direction are weighted and summed, for example, according to a pre-set weight: Let the initial positioning direction be P, the candidate direction set be D={d1,d2,...,dn}, the weight allocation be w_p for P, and the weight of each candidate direction be w_di, and w_p+Σw_di=1, then the sound source localization result R=w_p*P+Σ(w_di*di); the weight w_p can be set to be larger than Σw_di, and the weight w_di of each candidate direction can be determined according to the confidence of each candidate direction, with a larger weight corresponding to a candidate direction with higher confidence. Setting the weight w_p of the initial localization direction P to be larger than the sum of the weights Σw_di of all candidate directions means that the system is more inclined to trust the initial result calculated by the head-mounted device itself. This helps maintain low processing latency because the head-mounted device's data acquisition and processing are usually more direct and faster, avoiding communication overhead or delays that may be introduced by waiting for auxiliary data. This ensures real-time response of sound source localization, which is especially important in application scenarios requiring rapid interaction (such as AR rendering). High-confidence candidate directions often originate from more reliable signal processing results (such as direction estimates calculated by the TDOA algorithm), which can compensate for the head-mounted device's insufficient localization in specific directions (such as the rear or vertical direction), thereby calibrating and refining the initial localization direction and improving the overall localization accuracy.
[0048] In other feasible implementations, the weights can be dynamically adjusted based on the consistency between the candidate direction and the initial positioning direction to improve accuracy.
[0049] By combining the candidate direction set provided by the necklace's second microphone array, the head-mounted device can be calibrated based on the initial direction of localization. Because the necklace's microphone array has better spatial distribution, such as rear and vertical coverage, the candidate direction set can compensate for the shortcomings of the head-mounted device's individual localization, thereby improving the overall accuracy of sound source localization, especially in the rear and vertical directions.
[0050] In one feasible embodiment, the sound source localization method further includes steps S40-S50: Step S40: If the preliminary positioning direction has been obtained, but the sound source positioning auxiliary data sent by the necklace has not yet been received, the preliminary positioning direction is output as the sound source positioning result.
[0051] In practical applications, there may be situations where the head-mounted device has obtained a preliminary positioning direction based on the first microphone array, but the necklace has not sent sound source localization auxiliary data, or the transmission of sound source localization auxiliary data is delayed or fails, for example, due to unstable communication connection, insufficient necklace battery, or network latency. Outputting the preliminary positioning direction as the sound source localization result is to provide it to subsequent processing modules, such as AR rendering algorithms, as an input parameter for the sound source direction, to ensure that basic functions are available and maintain real-time performance.
[0052] Step S50: Upon receiving the sound source localization auxiliary data sent by the necklace, the sound source localization result is output by weighted summation of the preliminary localization direction and each of the candidate directions.
[0053] Outputting the sound source localization result provides it to subsequent processing modules, serving as a sound direction parameter for later processing, such as AR rendering, but with greater accuracy. This case-by-case output mechanism balances the real-time nature and accuracy of sound source localization. When data is unavailable, a preliminary result is output promptly to ensure real-time performance; when data is available, a fused result is output to improve accuracy, thereby optimizing the user experience and avoiding delays caused by waiting for data. Furthermore, it provides a gradual clarification process from initial rapid localization to final precise calibration, which not only improves the accuracy of sound source localization but also enhances user immersion and experience. Users can perceive the sound source gradually becoming clearer from vague, and this progressive precision localization enhances the realism and naturalness of the interaction.
[0054] In one feasible embodiment, the sound source localization method further includes step S60: Step S60: Obtain the clock offset between the necklace and the clock.
[0055] Clock offset is obtained through a two-way timestamp synchronization mechanism: the head-mounted device sends a synchronization request (including its local timestamp) at time t1, the necklace receives and records the timestamp at time t2, and then sends back a response at time t3 (including t2 and t3). The head-mounted device receives the response at time t4. The round-trip delay is calculated as: delay = (t4 - t1) - (t3 - t2), and the clock offset is: offset = [(t2 - delay / 2) - t1]. The system can periodically perform timestamp alignment, for example, every 5 seconds, to compensate for clock errors and ensure data time synchronization.
[0056] Step S30 includes: Step S303: Perform data time alignment based on the acquisition time of the first microphone array signal, the timestamp carried in the sound source localization auxiliary data, and the clock offset; perform sound source localization based on the aligned data and the relative pose to obtain the sound source localization result.
[0057] The data time alignment process may include: adjusting the time axis of the sound source localization auxiliary data according to the acquisition time T1 of the first microphone array signal, the timestamp T2 carried in the sound source localization auxiliary data, and the clock offset, for example, correcting T2 to T2' = T2 - offset to align with T1. Then, sound source localization is performed using the aligned data and relative pose, for example, by calculating the sound source direction using TDOA or beamforming algorithms. For example, in one feasible embodiment, when the sound source localization auxiliary data is the second microphone array signal itself, the aligned first and second microphone array signals can be combined, and the sound source direction can be calculated after transforming the relative pose to the same coordinate system. For example, in one feasible implementation, when the sound source localization auxiliary data includes a candidate direction set and a second TDOA matrix calculated based on the second microphone array signal, the sound source localization auxiliary data may further include timestamps of the candidate direction set and the second TDOA matrix, which are determined according to the acquisition time of the second microphone array signal. The head-mounted device may also determine the timestamps of the preliminary localization direction and the first TDOA matrix calculated based on the acquisition time of the first microphone array signal, and align the candidate direction set and the second TDOA matrix with the preliminary localization direction and the first TDOA matrix based on the timestamps of the candidate direction set and the second TDOA matrix, and perform subsequent sound source localization based on the aligned candidate direction set, the second TDOA matrix, the preliminary localization direction, and the first TDOA matrix.
[0058] In this embodiment, a distributed computing architecture is employed, with the head-mounted device and the necklace each processing their own collected data, and then transmitting the data wirelessly for comprehensive analysis. This approach ensures both real-time response speed and the ability to complete complex sound source localization calculations with limited hardware resources. Due to size limitations, the head-mounted device has a limited number of microphones arranged in a near-circular pattern, resulting in fuzzy vertical positioning and inaccurate rearward positioning. The necklace's second microphone array compensates for these deficiencies by providing rear microphones and vertical height differences. By combining relative pose and sound source localization auxiliary data, the head-mounted device can obtain more comprehensive sound source information, significantly improving positioning accuracy, especially in the rearward and vertical directions.
[0059] Based on the first embodiment described above, a second embodiment of the sound source localization method of this application is proposed. In this embodiment, content that is the same as or similar to that in the first embodiment can be referred to the above description and will not be repeated hereafter. In this embodiment, the sound source localization auxiliary data further includes a second TDOA matrix calculated based on the signals of the second microphone array. The second TDOA matrix includes the measured time difference of arrival (TDOA) corresponding to each microphone pair in the second microphone array. The second TDOA matrix is a matrix composed of the measured TDOA between each microphone pair in the second microphone array. It is calculated by performing cross-correlation analysis on the signals of the second microphone array to estimate the time difference of the sound signals received by each pair of microphones, for example, by using a generalized cross-correlation function to calculate the time difference and organizing it into a matrix form.
[0060] The step S302, which involves weighted summation of the preliminary positioning direction and each of the candidate directions to obtain the sound source localization result, includes steps S3021-S3023: Step S3021: For each direction pair formed by the preliminary positioning direction and each candidate direction, the consistency score of the direction pair is calculated according to the second TDOA matrix and the first TDOA matrix calculated based on the signal of the first microphone array. The first TDOA matrix includes the measured time difference of arrival corresponding to each microphone pair in the first microphone array.
[0061] The first TDOA matrix is the measured time difference of arrival matrix for each microphone pair in the first microphone array. Its calculation method is the same as the second TDOA matrix, obtained through cross-correlation analysis. The consistency score is calculated by comparing the theoretical time difference between the first and second TDOA matrices for a given direction pair. For example, using a multi-view TDOA fusion algorithm: let the direction pair be (P, di), where P is the initial positioning direction and di is the candidate direction. The theoretical time difference is calculated based on the sound wave propagation model. The consistency score s_i = exp(-Σ|Δt_actual - Δt_theoretical|^2), where Δt_actual is the measured time difference and Δt_theoretical is the theoretical time difference. A higher score indicates better consistency between the direction pair in the two arrays, meaning the candidate direction is more reliable.
[0062] Step S3022: The consistency score of the direction pair is used as the weight of the candidate direction in the direction pair, and the candidate directions are weighted and summed to obtain the first fusion direction.
[0063] Let the candidate direction set be D={d1,d2,...,dn}, and the consistency score of each candidate direction di be s_i. Then the first fusion direction F=Σ(s_i*d_i) / Σs_i, which is the weighted average direction.
[0064] Step S3023: The first fusion direction and the preliminary positioning direction are weighted and summed to obtain the sound source localization result.
[0065] Let the initial localization direction be P, the first fusion direction be F, and the weight be w. Then the sound source localization result R = w * F + (1 - w) * P.
[0066] By using a second TDOA matrix and a candidate direction set, a consistency score is calculated for weighted fusion, making the sound source localization results more reliable. This calibrates the initial localization direction, especially in backward localization, leveraging the advantages of necklace data to improve accuracy, as the necklace's microphone array provides better backward coverage.
[0067] In one feasible embodiment, step S3023 includes S30231~S30232: Step S30231: If both the first fusion direction and the preliminary positioning direction fall within the target direction range, then the first weight is used as the weight corresponding to the first fusion direction, and the first fusion direction and the preliminary positioning direction are weighted and summed to obtain the sound source localization result. The target direction range is a cone-shaped or fan-shaped region around the target direction in the coordinate system of the head-mounted device, and is a subset of the full-directional range of the sound source localization of the head-mounted device. The target direction faces the back of the wearer's head when the head-mounted device is worn.
[0068] The target direction is a specific direction in the head-mounted device's coordinate system, facing the back of the wearer's head when the device is worn. The target direction range is a cone-shaped area centered on this direction, with an angle range of ±30 degrees. When both the first fusion direction and the initial positioning direction fall within the target direction range, the weight of the initial positioning direction in the weighted summation is 1-w1, where w1 is the first weight. For example, w1=0.7 indicates greater confidence in the accuracy of the necklace data in the rear direction.
[0069] Step S30232: If at least one of the first fusion direction and the preliminary positioning direction does not fall within the target direction range, then the second weight is used as the weight corresponding to the first fusion direction, and the first fusion direction and the preliminary positioning direction are weighted and summed to obtain the sound source localization result, wherein the second weight is less than the first weight.
[0070] When at least one of the first fusion direction and the preliminary positioning direction does not fall within the target direction range, during weighted summation, the weight of the preliminary positioning direction is 1 - w2, where w2 is the second weight and w2 < w1. For example, w2 = 0.3. The second weight is set smaller than the first weight because the second microphone array on the necklace has a microphone behind the neck. So when both the preliminary positioning direction and the first fusion direction for positioning are from behind the wearer's head, the first fusion direction is a direction determined with the assistance of the necklace data. Increasing its weight can more accurately compensate for the problem of inaccurate rearward positioning of the head-mounted device; while when the direction does not fall within the rearward range, the data of the head-mounted device may be more reliable. Therefore, the weight of the necklace data is reduced to prioritize the data of the head-mounted device.
[0071] Based on the above first and / or second embodiments, a third embodiment of the sound source localization method of the present application is proposed. In this embodiment, for the content that is the same as or similar to the above first and second embodiments, reference can be made to the above introduction and will not be repeated hereinafter. In this embodiment, the step of obtaining the sound source localization result by performing weighted summation on the preliminary positioning direction and each of the candidate directions in step S302 includes S3024 to S3027: Step S3024, perform weighted summation on the preliminary positioning direction and each of the candidate directions to obtain a preliminary fusion direction.
[0072] Compared with the method in the foregoing embodiments of taking the direction obtained by performing weighted summation on the preliminary positioning direction and each candidate direction as the final sound source localization result, in this embodiment, this direction is taken as the preliminary fusion direction, and on the basis of the preliminary fusion direction, further correction is performed through a deep learning model.
[0073] Step S3025, calculate a third TDOA matrix according to the preliminary fusion direction, where the third TDOA matrix includes the theoretical arrival time differences of each target microphone pair assuming that the sound source is from the preliminary fusion direction, and each of the target microphone pairs is a microphone pair formed by pairwise combination of the microphones in the first microphone array and the microphones in the second microphone array.
[0074] The calculation of the third TDOA matrix is based on the initial fusion direction as the assumed sound source direction, utilizing the sound wave propagation model and the geometric relationship between the microphone pairs. Specifically, let the initial fusion direction be a unit vector d, and the relative pose between the head-mounted device and the necklace be known (e.g., represented by translation vectors and rotation matrices). The positions of all microphones in the first and second microphone arrays are transformed to the same coordinate system (e.g., the head-mounted device coordinate system). For each target microphone pair (i.e., each pair of microphones in the first and second microphone arrays), the path difference of the sound wave propagation from the assumed sound source direction to the two microphones is calculated: path difference Δl = d·(p_i - p_j), where p_i and p_j are the position vectors of the two microphones, and · represents the dot product. Then, the theoretical time difference of arrival Δt_ij = Δl / c, where c is the speed of sound (e.g., 340 m / s). The third TDOA matrix is composed of all these Δt_ij, forming a matrix or vector representing the expected time difference (theoretical time difference of arrival) between the microphone pairs under the assumed sound source direction.
[0075] Step S3026: Calculate the spectral characteristics, spatial filter response map, and ambient noise spectrum based on the signal from the first microphone array.
[0076] Spectral characteristics are obtained through frequency domain analysis of the first microphone array signal, such as using Short-Time Fourier Transform (STFT) to calculate the power spectrum or Mel-frequency cepstral coefficients (MFCC) to capture the spectral characteristics of the sound source, such as frequency distribution and timbre information. The spatial filtering response map is generated using beamforming algorithms, such as a delayed summation beamformer, to calculate the spatial response based on the geometry of the first microphone array and the initial fusion direction, forming a map representing the signal intensity in different directions, highlighting the energy distribution in the direction of the sound source. The ambient noise spectrum is obtained by statistically analyzing the spectrum of the first microphone array signal during periods of no sound source activity or silence, such as calculating the average power spectrum of the background noise, to characterize the ambient noise characteristics for noise suppression in subsequent processing.
[0077] Step S3027: Input the third TDOA matrix, the spectral features, the spatial filter response map, and the environmental noise spectrum into a preset deep learning model for processing to obtain the sound source localization result.
[0078] Deep learning models include Conv-TasNet and Spatial Transformer Network. Conv-TasNet is an end-to-end speech separation network, but here it is used for feature extraction: the signal from the first microphone array is input into Conv-TasNet, its original output layer is removed, retaining the encoder and separation module, and outputting a high-dimensional signal feature vector that captures the time-frequency and spatial information of the signal. Then, this signal feature is concatenated with the third TDOA matrix (flattened into a vector), spectral features (such as MFCC vectors), spatial filter response map (flattened into a vector), and ambient noise spectrum (flattened into a vector) to form a comprehensive feature vector. The concatenated feature vector is input into the Spatial Transformer Network, which learns spatial transformation parameters through convolutional and fully connected layers, and finally outputs the sound source direction (such as a unit vector or spherical coordinates). The Spatial Transformer Network is improved here to be a regression network, outputting continuous direction values instead of the traditional classification output, to adapt to the continuous direction estimation of sound source localization. Specifically, improvements to Conv-TasNet include removing the output layer and adjusting the input dimension to accept multimodal features; improvements to Spatial Transformer Network include adding a regression head (such as a fully connected layer) to output a direction vector.
[0079] Deep learning models can be pre-trained and deployed in head-mounted devices. These models can be trained using supervised learning. Training data can be constructed as follows: multiple sound source directions are defined, specific audio signals are played from these directions, and microphone array signals under different sound source directions are collected using a first and second microphone array as samples. For each sample, a corresponding third TDOA matrix, spectral features, spatial filter response map, and environmental noise spectrum are generated as input features of the model, following the same method as in the previous embodiments; the sound source direction corresponding to the sample (i.e., the true sound source direction) is used as the training label. The training objective is to optimize loss functions such as mean squared error to make the sound source direction predicted by the model as close as possible to the true label, thereby learning the complex mapping relationship from multimodal input features to continuous sound source directions.
[0080] In this embodiment, by combining a deep learning model, the third TDOA matrix, spectral features, spatial filter response map, and ambient noise spectrum can further calibrate the initial fusion direction, thereby obtaining more accurate sound source localization results. This is because the third TDOA matrix provides theoretical constraints based on geometric relationships, the spectral features and spatial filter response map capture the frequency and spatial characteristics of the signal, the ambient noise spectrum helps with noise reduction, and the deep learning model, through end-to-end learning and fusion of these multimodal features, can model nonlinear relationships in complex acoustic environments (such as multipath effects and noise interference), compensate for possible errors in the initial fusion direction, and thus improve the robustness and accuracy of localization, especially when the sound source direction changes or the environment is noisy.
[0081] Based on the first, second, and / or third embodiments described above, a fourth embodiment of the sound source localization method of this application is proposed. In this embodiment, content that is the same as or similar to the first, second, and third embodiments described above can be referred to the above description and will not be repeated hereafter. In this embodiment, the executing entity of the sound source localization method is a necklace. A second microphone array is provided in the necklace, and a first microphone array is provided in the head-mounted device that establishes a communication connection with the necklace. (Refer to...) Figure 4 , Figure 4 This is a flowchart illustrating the first embodiment of the sound source localization method of this application, which includes steps A10 to A30: Step A10: Acquire the second microphone array signal collected by the second microphone array.
[0082] Step A20: Obtain sound source localization auxiliary data based on the signal from the second microphone array.
[0083] Step A30: The sound source localization auxiliary data is sent to the head-mounted device so that the head-mounted device can perform sound source localization based on the first microphone array signal, the sound source localization auxiliary data, and the relative pose between the head-mounted device and the necklace, and obtain the sound source localization result, wherein the first microphone array signal is obtained based on the first microphone array.
[0084] The explanation and specific implementation of steps A10 to A30 in this embodiment can be found in the explanation and specific implementation of steps S10 to S30 above, and will not be repeated here.
[0085] This application provides an electronic device. The electronic device may be a head-mounted device or a necklace. The electronic device includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform the sound source localization method in the above embodiments.
[0086] The following is for reference. Figure 5 It shows a schematic diagram of the structure of an electronic device suitable for implementing the embodiments of this application. Figure 5 The electronic device shown is merely an example and should not impose any limitation on the functionality and scope of use of the embodiments of this application.
[0087] like Figure 5 As shown, the electronic device may include a processing unit 1001 (e.g., a central processing unit, a graphics processing unit, etc.), which can perform various appropriate actions and processes according to a program stored in a read-only memory 1002 or a program loaded from a storage device 1003 into a random access memory 1004. The random access memory 1004 also stores various programs and data required for the operation of the electronic device. The processing unit 1001, the read-only memory 1002, and the random access memory 1004 are interconnected via a bus 1005. An input / output interface 1006 is also connected to the bus. Typically, the following systems can be connected to the input / output interface 1006: input devices 1007 including, for example, touchscreens, touchpads, keyboards, mice, image sensors, microphones, accelerometers, gyroscopes, etc.; output devices 1008 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 1003 including, for example, magnetic tapes, hard disks, etc.; and communication devices 1009. The communication device 1009 allows the electronic device to communicate wirelessly or wiredly with other devices to exchange data. Although the diagrams show electronic devices with various systems, it should be understood that it is not required to implement or have all of the systems shown. More or fewer systems may be implemented alternatively.
[0088] Specifically, according to the embodiments disclosed in this application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments disclosed in this application include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device, or installed from storage device 1003, or installed from read-only memory 1002. When the computer program is executed by processing device 1001, it performs the functions defined in the methods of the embodiments disclosed in this application.
[0089] Compared with the prior art, the beneficial effects of the electronic device provided in this application embodiment are the same as those of the sound source localization method provided in the above embodiment, and other technical features of the electronic device are the same as those disclosed in the method of the previous embodiment, which will not be repeated here.
[0090] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
[0091] This application provides a computer-readable storage medium having computer-readable program instructions (i.e., a computer program) stored thereon, which are used to execute the sound source localization method in the above embodiments.
[0092] The computer-readable storage medium provided in this application embodiment may be, for example, a USB flash drive, but is not limited to, electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems or devices, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to: electrical connections having one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof. In this embodiment, the computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system or device. The program code contained on the computer-readable storage medium may be transmitted using any suitable medium, including but not limited to: wires, optical cables, RF (Radio Frequency), etc., or any suitable combination thereof.
[0093] The aforementioned computer-readable storage medium may be included in an electronic device or may exist independently without being assembled into an electronic device.
[0094] The aforementioned computer-readable storage medium carries one or more programs, which, when executed by an electronic device, cause the electronic device to perform the functions defined in the methods of the embodiments disclosed in this application.
[0095] Computer program code for performing the operations of this application can be written in one or more programming languages or a combination thereof, including object-oriented programming languages such as Java, Smalltalk, and C++, and conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a Local Area Network (LAN) or a Wide Area Network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0096] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0097] The modules described in the embodiments of this application can be implemented in software or hardware. The names of the modules do not necessarily limit the functionality of the unit itself.
[0098] The readable storage medium provided in this application embodiment is a computer-readable storage medium, which stores computer-readable program instructions (i.e., computer programs) for executing the above-described sound source localization method. Compared with the prior art, the beneficial effects of the computer-readable storage medium provided in this application embodiment are the same as the beneficial effects of the sound source localization method provided in the above-described embodiments, and will not be repeated here.
[0099] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the steps of the sound source localization method described above.
[0100] Compared with the prior art, the beneficial effects of the computer program product provided in this application embodiment are the same as the beneficial effects of the sound source localization method provided in the above embodiments, and will not be repeated here.
[0101] The above description is only a part of the embodiments of this application and does not limit the patent scope of this application. All equivalent structural transformations made under the technical concept of this application and using the contents of the specification and drawings of this application, or direct / indirect applications in other related technical fields, are included in the patent protection scope of this application.
Claims
1. A method for locating a sound source, characterized in that, The sound source localization method is applied to a head-mounted device, which is equipped with a first microphone array, and a necklace that establishes a communication connection with the head-mounted device is equipped with a second microphone array. The sound source localization method includes: Acquire the first microphone array signal collected by the first microphone array; The relative pose between the necklace and the head-mounted device is obtained, as well as the sound source localization assistance data sent by the necklace is obtained, wherein the sound source localization assistance data is obtained based on the second microphone array signal collected by the second microphone array; The sound source is located based on the first microphone array signal, the sound source localization auxiliary data, and the relative pose, and the sound source localization result is obtained.
2. The sound source localization method as described in claim 1, characterized in that, The sound source localization auxiliary data includes a candidate direction set obtained by sound source localization based on the second microphone array signal; The step of performing sound source localization based on the first microphone array signal, the sound source localization auxiliary data, and the relative pose to obtain the sound source localization result includes: Based on the signal from the first microphone array, the sound source is located to obtain a preliminary positioning direction; After transforming the preliminary positioning direction and each candidate direction in the candidate direction set to the same coordinate system based on the relative pose, the preliminary positioning direction and each candidate direction are weighted and summed to obtain the sound source localization result.
3. The sound source localization method as described in claim 2, characterized in that, The sound source localization auxiliary data also includes a second TDOA matrix calculated based on the signal of the second microphone array, the second TDOA matrix including the measured time difference of arrival for each microphone pair in the second microphone array; The step of obtaining the sound source localization result by weighted summation of the preliminary localization direction and each of the candidate directions includes: For each direction pair formed by combining the initial positioning direction with each candidate direction, a consistency score for the direction pair is calculated based on the second TDOA matrix and the first TDOA matrix calculated based on the signal of the first microphone array. The first TDOA matrix includes the measured time difference of arrival corresponding to each microphone pair in the first microphone array. The consistency score of the direction pair is used as the weight of the candidate direction in the direction pair, and the candidate directions are weighted and summed to obtain the first fusion direction; The sound source localization result is obtained by weighted summation of the first fusion direction and the preliminary positioning direction.
4. The sound source localization method as described in claim 3, characterized in that, The step of weighted summing of the first fusion direction and the preliminary positioning direction to obtain the sound source localization result includes: If both the first fusion direction and the preliminary positioning direction fall within the target direction range, then the first weight is used as the weight corresponding to the first fusion direction, and the first fusion direction and the preliminary positioning direction are weighted and summed to obtain the sound source localization result. The target direction range is a cone-shaped or fan-shaped region around the target direction in the coordinate system of the head-mounted device, and is a subset of the full-directional range of the sound source localization of the head-mounted device. The target direction faces the back of the wearer's head when the head-mounted device is worn. If at least one of the first fusion direction and the preliminary positioning direction does not fall within the target direction range, then the second weight is used as the weight corresponding to the first fusion direction, and the first fusion direction and the preliminary positioning direction are weighted and summed to obtain the sound source localization result, wherein the second weight is less than the first weight.
5. The sound source localization method as described in claim 2, characterized in that, The step of obtaining the sound source localization result by weighted summation of the preliminary localization direction and each of the candidate directions includes: The preliminary fusion direction is obtained by weighted summation of the preliminary positioning direction and each of the candidate directions; The third TDOA matrix is calculated based on the initial fusion direction, wherein the third TDOA matrix includes the theoretical arrival time difference of each target microphone pair under the assumption that the sound source comes from the initial fusion direction, and each target microphone pair is a microphone pair composed of two microphones in the first microphone array and two microphones in the second microphone array; The spectral characteristics, spatial filter response map, and ambient noise spectrum are calculated based on the signals from the first microphone array. The third TDOA matrix, the spectral features, the spatial filter response map, and the environmental noise spectrum are input into a preset deep learning model for processing to obtain the sound source localization result.
6. The sound source localization method as described in claim 2, characterized in that, The sound source localization method further includes: If the preliminary positioning direction is obtained, but the sound source positioning auxiliary data sent by the necklace has not yet been received, the preliminary positioning direction is output as the sound source positioning result. Upon receiving the sound source localization assistance data sent by the necklace, the sound source localization result is output by weighted summation of the preliminary localization direction and each of the candidate directions.
7. The sound source localization method according to any one of claims 1 to 6, characterized in that, The sound source localization method further includes: Obtain the clock offset from the necklace; The step of performing sound source localization based on the first microphone array signal, the sound source localization auxiliary data, and the relative pose to obtain the sound source localization result includes: Data time alignment is performed based on the acquisition time of the first microphone array signal, the timestamp carried in the sound source localization auxiliary data, and the clock offset. Sound source localization is then performed based on the aligned data and the relative pose to obtain the sound source localization result.
8. A method for locating a sound source, characterized in that, The sound source localization method is applied to a necklace, which is equipped with a second microphone array, and a head-mounted device that establishes a communication connection with the necklace is equipped with a first microphone array. The sound source localization method includes: Acquire the second microphone array signal collected by the second microphone array; Sound source localization auxiliary data is obtained based on the signals from the second microphone array; The sound source localization auxiliary data is sent to the head-mounted device so that the head-mounted device can perform sound source localization based on the first microphone array signal, the sound source localization auxiliary data, and the relative pose between the head-mounted device and the necklace, and obtain the sound source localization result, wherein the first microphone array signal is obtained based on the first microphone array.
9. An electronic device, characterized in that, The electronic device further includes: a memory, a processor, and a computer program stored in the memory and executable on the processor, the computer program being configured to implement the steps of the sound source localization method as described in any one of claims 1 to 8.
10. A storage medium, characterized in that, The storage medium is a computer-readable storage medium, and a computer program is stored on the storage medium. When the computer program is executed by a processor, it implements the steps of the sound source localization method as described in any one of claims 1 to 8.
Citation Information
Patent Citations
Live broadcast method and live broadcast device based on multiple microphones
CN108307268A
Multi-mode man-machine interaction control system and method of humanoid robot
CN116931728A
Arrangement of illumination sources inside and outside finger-obscured area of top cover of handheld controller to assist artificial reality system in tracking position of controller, and systems and methods of use thereof
CN118871175A
Neck-worn device and wearable device
CN120584498A
Optimization of microphone array geometry for direction of arrival estimation
US10638222B1