Sound source positioning method and system, electronic equipment and storage medium

By establishing the mapping relationship between the user's gaze direction and the microphone space vector, the problem that traditional technology cannot accurately locate the sound source in complex environments is solved, and higher sound source positioning accuracy and user experience are improved.

CN120143052APending Publication Date: 2025-06-13BEIJING SUPERHEXA CENTURY TECH CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510393579.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-31
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

Traditional sound source positioning technology cannot accurately focus on the sound sources that users pay attention to in complex environments, resulting in messy audio information and poor user experience.

Method used

By determining the gaze space vector based on the eye movement of the target user, determining the microphone space vector based on the microphone array position, and establishing a mapping relationship between the two, accurately positioning the sound source in the direction of the user's gaze.

Benefits of technology

In complex acoustic environments, attention to different microphone sound signals can be dynamically adjusted, and the sound sources that are of interest to users can be accurately focused, improving the accuracy and user experience of sound source positioning.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120143052A_ABST
    Figure CN120143052A_ABST
Patent Text Reader

Abstract

The invention provides a sound source localization method and system, an electronic device and a storage medium, and belongs to the technical field of sound source localization, the method is applied to a wearable device, and the method comprises the following steps: determining a gaze space vector corresponding to a gaze direction of a target user based on an eye action of the target user, the target user being a user wearing the wearable device; based on the position coordinates of a microphone array, determining a microphone space vector corresponding to each microphone, the microphone array being arranged in the wearable device; establishing a mapping relation based on the fixation space vector and the microphone space vector; and positioning the sound source in the gazing direction of the target user based on the mapping relation. According to the sound source positioning method and system, the electronic equipment and the storage medium provided by the invention, the user experience can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure belongs to the technical field of sound source localization, and more specifically, relates to a sound source localization method and system, an electronic device, and a storage medium. Background Art

[0002] Today, with the increasing popularity of wearable devices, the diversity of their functions and the optimization of user experience have become important development directions. In the field of audio processing, traditional sound source localization technologies often fail to meet the user's demand for focusing on specific sound sources in complex environments.

[0003] Most of the existing sound source localization methods are based on fixed algorithms and models, without considering the user's subjective attention direction. In actual scenarios, such as in noisy shopping malls, multi-person meetings and other environments, users may only be interested in sound sources in a specific direction, while traditional technologies are difficult to accurately focus on these target sound sources, resulting in messy audio information and poor user experience. Summary of the Invention

[0004] The purpose of the present disclosure is to provide a sound source localization method and system, an electronic device, and a storage medium to improve the user experience.

[0005] In the first aspect of the embodiments of the present disclosure, a sound source localization method is provided, which is applied to a wearable device and includes: Determining a gaze space vector corresponding to the gaze direction of a target user based on the eye movements of the target user, where the target user is the user wearing the wearable device; Determining a microphone space vector corresponding to each microphone based on the position coordinates of a microphone array, where the microphone array is disposed on the wearable device; Establishing a mapping relationship based on the gaze space vector and the microphone space vector; Locating the sound source in the direction gazed by the target user based on the mapping relationship.

[0006] In the second aspect of the embodiments of the present disclosure, a sound source localization system is provided, which is applied to a wearable device and includes: A first vector determination module for determining a gaze space vector corresponding to the gaze direction of a target user based on the eye movements of the target user, where the target user is the user wearing the wearable device; A second vector determination module for determining a microphone space vector corresponding to each microphone based on the position coordinates of a microphone array, where the microphone array is disposed on the wearable device; A mapping module for establishing a mapping relationship based on the gaze space vector and the microphone space vector; A sound source localization module for locating the sound source in the direction gazed by the target user based on the mapping relationship.

[0007] In a third aspect of the embodiments of the present disclosure, an electronic device is provided, including a memory, a processor, and a computer program stored in the memory and running on the processor. When the processor executes the computer program, the steps of the above-mentioned sound source localization method are implemented.

[0008] In a fourth aspect of the embodiments of the present disclosure, a computer-readable storage medium is provided. The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the above-mentioned sound source localization method are implemented.

[0009] The beneficial effects of the sound source localization method, system, electronic device, and storage medium provided by the embodiments of the present disclosure are as follows: By determining the gaze space vector based on the eye movements of the target user, combining the positions of the microphone arrays to determine the microphone space vector, and establishing a mapping relationship between the two, the embodiments of the present disclosure achieve accurate sound source localization focusing on the direction of the user's attention. In a complex acoustic environment, the wearable device can dynamically adjust the attention to the sound signals of different microphones according to the user's gaze direction, and concentrate the attention on the sound source that the user is interested in. This not only improves the accuracy of sound source localization but also enhances the user experience. BRIEF DESCRIPTION OF THE DRAWINGS

[0010] In order to more clearly illustrate the technical solutions in the embodiments of the present disclosure, the following will briefly introduce the drawings required for use in the embodiments or the description of the prior art. Obviously, the following drawings are only some embodiments of the present disclosure. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0011] Figure 1 It is a schematic flowchart of the sound source localization method provided by an embodiment of the present disclosure; Figure 2 It is a structural block diagram of the sound source localization system provided by an embodiment of the present disclosure; Figure 3 It is a schematic block diagram of the electronic device provided by an embodiment of the present disclosure. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0012] In the following description, specific details such as specific system structures and technologies are presented for the purpose of illustration rather than limitation, so as to thoroughly understand the embodiments of the present disclosure. However, those skilled in the art should clearly understand that the present disclosure can also be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, systems, circuits, and methods are omitted to avoid unnecessary details from interfering with the description of the present disclosure.

[0013] To make the objectives, technical solutions, and advantages of the present disclosure clearer, the following will be described through specific embodiments with reference to the drawings.

[0014] Please refer to Figure 1 , Figure 1 which is a schematic flowchart of a sound source localization method provided by an embodiment of the present disclosure. This method is applied to a wearable device and includes: S101: Based on the eye movements of the target user, determine the gaze space vector corresponding to the gaze direction of the target user. The target user is the user wearing the wearable device.

[0015] In this embodiment, the wearable device can be a portable electronic device directly worn on the body. The wearable device can be an Augmented Reality (AR) device, such as a head-mounted glasses.

[0016] In this embodiment, the wearable device can capture the eye movement data of the target user by means of a built-in eye tracking device. The eye tracking device can include an infrared sensor, a camera, etc. The infrared sensor can emit infrared rays and detect the situation of its reflection from the eyes, so as to obtain the position and movement information of the eyes; the camera can directly capture images of the eyes.

[0017] Analyze and process the obtained eye movement data. Convert the eye movement data into a gaze direction in three-dimensional space to obtain the corresponding gaze space vector. The gaze space vector represents the actual gaze direction of the target user in space.

[0018] For example, according to the rotation angle of the eyeball and the position change information of the pupil, it can be mapped into a three-dimensional space coordinate system to obtain a vector with direction and magnitude, so as to calculate the actual gaze direction of the target user in space and represent it in the form of a vector.

[0019] S102: Based on the position coordinates of the microphone array, determine the microphone space vector corresponding to each microphone. The microphone array is arranged on the wearable device.

[0020] In this embodiment, the position coordinates of the microphone array can be determined during the design process of the wearable device to clarify the position coordinates of each microphone in the device coordinate system.

[0021] Based on the position coordinates of each microphone, taking the origin of the wearable device coordinate system as the starting point and pointing to the position of each microphone, the corresponding microphone space vector can be generated.

[0022] S103: Establish a mapping relationship based on the gaze space vector and the microphone space vector.

[0023] In this embodiment, the cosine value of the included angle between the gaze space vector and each microphone space vector can be calculated. The smaller the included angle, the closer the microphone is to the user's gaze direction.

[0024] Assign weights to each microphone based on the size of the included angle. The smaller the included angle, the greater the weight, thereby establishing a mapping relationship between the gaze direction and each microphone. The mapping relationship reflects the importance of each microphone for sound source localization under the user's gaze direction.

[0025] S104: Locate the sound source in the direction where the target user is gazing based on the mapping relationship.

[0026] In this embodiment, since the sound emitted by the sound source will reach each microphone at different times and intensities, there will be differences in the sound signals received by each microphone.

[0027] The microphone array collects the sound signals in the surrounding environment, and based on the established mapping relationship, performs weighted processing on the sound signals received by each microphone to obtain weighted values. The microphones closer to the user's gaze direction have larger corresponding weights; the microphones farther away from the gaze direction have smaller weights, and the influence of their signals is relatively small.

[0028] Based on the weighted values, judge the time difference of the sound signals reaching different microphones, and determine the position of the sound source in the direction where the target user is gazing.

[0029] Exemplarily, assume that the wearable device is an AR glasses. When user A wears this AR glasses and enters a concert, the built-in eye tracking device in the glasses is quickly activated. For example, by precisely detecting the infrared situation reflected from user A's eyes, obtain the position and movement information of user A's eyes; based on the rotation angle of user A's eyeball and the change in pupil position, map this information accurately into a three-dimensional space coordinate system, thereby quickly determining the gaze space vector corresponding to the direction where user A is gazing. Assume that the performance of the band on the stage attracts user A's attention, and this gaze space vector clearly points to the stage direction.

[0030] In the design stage of the AR glasses, the position coordinates of each microphone in the microphone array have been determined. Calculate the cosine value of the included angle between the gaze space vector and each microphone space vector. The microphones facing the stage direction have a smaller included angle with the gaze space vector, while the microphones facing away from the stage have a larger included angle. Based on this difference in the size of the included angle, assign corresponding weights to each microphone. The smaller the included angle, the higher the weight. In this way, a mapping relationship is established between the gaze direction of user A and each microphone. This mapping relationship clearly and accurately reflects the importance of each microphone for sound source localization under the current gaze direction of user A.

[0031] The microphone array collects the sound signals of the concert. According to the previously established mapping relationship, the sound signals received by each microphone are weighted. In this process, the sound signals collected by the microphones with larger weights in the direction of the stage dominate in the processing. The AR glasses can effectively highlight the sound source information in the direction of the stage and filter out other irrelevant or interfering sounds. Finally, user A can more clearly and attentively feel the wonderful performance of the band, greatly enhancing the immersive experience of user A at the party scene.

[0032] As can be seen from the above, in this embodiment, by determining the gaze space vector based on the eye movements of the target user, combining the microphone array position to determine the microphone space vector, and establishing the mapping relationship between the two, accurate sound source localization in the direction of the user's attention is achieved. In a complex acoustic environment, the wearable device can dynamically adjust the attention to the sound signals of different microphones according to the user's gaze direction, focusing the attention on the sound sources that the user is interested in. This not only improves the accuracy of sound source localization but also enhances the user's experience.

[0033] In an embodiment of the present disclosure, based on the eye movements of the target user, determining the gaze space vector corresponding to the gaze direction of the target user includes: Obtaining the eye movement visual features and electrooculogram signal features of the target user; Performing feature fusion on the eye movement visual features and electrooculogram signal features to obtain a comprehensive eye movement feature vector; Mapping the comprehensive eye movement feature vector to a gaze space vector.

[0034] In practical applications, a single eye movement visual feature or electrooculogram signal feature may have limitations and is easily affected by external interference or its own physiological factors.

[0035] In this embodiment, a camera and an electromyography sensor can be integrated in the wearable device.

[0036] The camera can capture images of the target user's eyes. Based on computer vision technology, eye movement visual features such as the rotation angle of the eyeball, the size and position changes of the pupil are extracted from these images. These features can intuitively reflect the movement state and direction of the eyes in space.

[0037] The electromyography sensor can detect the electrical signals generated by the muscles around the eyes. The activities of the eye muscles are closely related to the eye movements. When the eyeball rotates, the corresponding eye muscles can generate electrical signals with specific patterns. By performing preprocessing such as filtering and amplifying on the electrical signals, features that can reflect the intention of eye movement are extracted, such as the intensity and frequency changes of the electromyography signals.

[0038] By integrating eye movement visual features and electrooculogram signal features, a more comprehensive and accurate representation of eye movement information can be obtained.

[0039] Based on the weighted summation or principal component analysis method, the eye movement visual features and electrooculogram signal features can be combined into a comprehensive eye movement feature vector. The comprehensive eye movement feature vector contains complementary information of the two features and can more accurately describe the eye movement state of the target user.

[0040] In this embodiment, a mapping relationship between the comprehensive eye movement feature vector and the gaze space vector can be established through a machine learning algorithm. Training can be carried out through models such as deep neural networks to learn the complex non-linear relationship between the two. The obtained comprehensive eye movement feature vector is input into the trained mapping model, and the model outputs the corresponding gaze space vector according to the learned relationship. This vector represents the actual gaze direction of the target user in three-dimensional space and can provide key information for subsequent applications such as sound source localization based on the user's gaze direction.

[0041] It can be concluded from the above that in this embodiment, by simultaneously acquiring eye movement visual features and electrooculogram signal features and utilizing their complementarity, the eye movement information of the target user can be captured more comprehensively and accurately. The comprehensive eye movement feature vector obtained by feature fusion effectively improves the accuracy of the description of the eye movement state. Mapping it to the gaze space vector can greatly improve the accuracy of determining the user's gaze direction.

[0042] In an embodiment of the present disclosure, establishing a mapping relationship based on the gaze space vector and the microphone space vector includes: Performing feature extraction on the gaze space vector to obtain gaze space vector features; Performing feature extraction on the microphone space vector to obtain microphone space vector features; Fusing the gaze space vector features and the microphone space vector features to obtain high-dimensional fusion features; Encoding the high-dimensional fusion features to obtain the mapping relationship between the gaze space vector and the microphone space vector.

[0043] In this embodiment, features are extracted from the gaze space vector and the microphone space vector respectively to obtain information that can represent their essential characteristics.

[0044] For the gaze space vector, the direction cosine angle features of the gaze space vector are extracted, and these features can accurately describe the gaze direction of the target user. For the microphone space vector, geometric features such as the three-dimensional space coordinate position of the microphone space vector and the relative distance and angle from other microphones are extracted, and these geometric features can reflect the layout and mutual relationship of the microphones in space. Through feature extraction, the original vectors are transformed into more representative and discriminative feature representations.

[0045] In this embodiment, the extracted gaze space vector features and microphone space vector features are fused to form high-dimensional fusion features.

[0046] Since the gaze space vector features and microphone space vector features come from different data sources respectively and each contains important information related to sound source localization, fusing them can comprehensively utilize this information. The high-dimensional fusion features can be encoded by a deep autoencoder and mapped to a low-dimensional latent feature space. In the low-dimensional space, the features of the two vectors can be better correlated and interacted with each other, removing redundant information while retaining key features, making the potential relationship between the two vectors clearer.

[0047] After the encoding process, the mapping relationship between the gaze space vector and the microphone space vector can be obtained. The mapping relationship reflects the degree of association between the user's gaze direction and each microphone in space.

[0048] Exemplarily, the eye movement tracking module in the wearable device continuously collects eye images and infrared reflection data. After preprocessing (such as denoising, image enhancement), the rotation angle and direction of the eyeball are calculated through computer vision algorithms.

[0049] Calculate the direction cosine values of the gaze space vector according to the rotation angle, and use them as the features of the gaze space vector. Assume the gaze space vector is , then its direction cosine is , where:

[0050]

[0051]

[0052] Read the three-dimensional space coordinates of each microphone (i = 1, 2,..., n, n is the number of microphones).

[0053] Calculate the relative distance between any two microphones : , ( ) Calculate the relative angle between any two microphones , the relative angle can be calculated through the vector dot product and cross product formulas. These coordinate, distance, and angle information constitute the geometric features of the microphone space vector.

[0054] Concatenate the extracted gaze space vector features and microphone space vector features to form a high-dimensional fusion feature vector. For example, arrange the three direction cosine values of the gaze space vector and the coordinates, relative distances, and relative angle values of all microphones in sequence to form a long vector.

[0055] Construct a deep autoencoder model, including an encoder and a decoder. The encoder consists of multiple fully connected layers and is used to map the high-dimensional fusion feature vector to a low-dimensional latent feature space; the decoder also consists of multiple fully connected layers and is used to reconstruct the low-dimensional feature vector back to the high-dimensional space. Use a large amount of sample data to train the deep autoencoder to minimize the reconstruction error. During the training process, the encoder learns an effective representation of the high-dimensional fusion feature vector, that is, the low-dimensional latent feature. After training, input the real-time obtained high-dimensional fusion feature vector into the trained encoder to obtain the corresponding low-dimensional latent feature vector. The low-dimensional vector represents the mapping relationship between the gaze space vector and the microphone space vector.

[0056] It can be concluded from the above that in this embodiment, features are extracted from the gaze and microphone space vectors respectively, and their essential information can be accurately obtained, and the original vectors are transformed into more representative and discriminative representations. The two features are fused to obtain a high-dimensional fusion feature, which can comprehensively utilize important information related to different data sources and sound source localization. Encoding the high-dimensional features into a low-dimensional latent space removes redundancy and retains key features, making the latent relationship between vectors clearer.

[0057] In an embodiment of the present disclosure, localize the sound source in the direction where the target user is looking based on the mapping relationship, including: Calculate the dynamic weight of each microphone in the microphone array based on the angle between the gaze space vector and each microphone space vector; Determine the position of the sound source in the direction where the target user is looking based on the dynamic weight and the time difference of the audio signals received by different microphones.

[0058] In this embodiment, the dynamic weight of each microphone is determined according to the angle between the gaze space vector and each microphone space vector. The gaze space vector represents the gaze direction of the target user, and the microphone space vector reflects the position of each microphone in space. The size of the angle reflects the degree of deviation of each microphone relative to the user's gaze direction. When the angle is small, it means that the microphone is closer to the user's gaze direction, and its importance in sound source localization is relatively higher, so a larger dynamic weight can be assigned; when the angle is large, the microphone is far from the gaze direction, and its dynamic weight is smaller. By calculating the dynamic weight, an importance measure related to the user's gaze direction is given to each microphone, realizing differential treatment of different microphones.

[0059] In this embodiment, after obtaining the dynamic weights of each microphone, the time difference of the audio signals received by different microphones is combined to determine the sound source position. Sound takes time to travel in the air. Since the distances between the sound source and each microphone are different, there will be differences in the times when the audio signals reach different microphones. When calculating the sound source position, the dynamic weights can affect the contribution degree of the time difference information provided by each microphone to the final positioning result. For a microphone with a larger weight, the time difference information of the audio signal it receives will play a more important role in the calculation, while the influence of a microphone with a smaller weight is relatively weak. This can make the positioning process more focused on the user's gaze direction, improving the accuracy and reliability of locating the sound source in the direction where the target user is looking.

[0060] It can be concluded from the above that this embodiment calculates the dynamic weights of the microphones based on the included angle, making the positioning more in line with the user's gaze direction and highlighting the key points of concern. Combining the dynamic weights and the time difference of the audio signals to locate the sound source can effectively improve the positioning accuracy, avoid interference from irrelevant sound sources, enable the wearable device to better meet the user's need to focus on a specific sound source, and enhance the usage experience.

[0061] In an embodiment of the present disclosure, calculating the dynamic weight of each microphone in the microphone array based on the included angle between the gaze space vector and the space vectors of each microphone includes: Calculating the dynamic weight of each microphone in the microphone array based on the first formula; The first formula is:

[0062] Wherein, represents the dynamic weight of the i-th microphone, represents the adjustment parameter, represents the included angle between the gaze space vector and the space vector of the i-th microphone, represents the sound signal intensity received by the i-th microphone, represents the average value of the sound signal intensities received by all microphones.

[0063] In this embodiment, Based on the Gaussian function form, the influence of the spatial position relationship between the microphone and the user's gaze direction on the dynamic weight can be reflected.

[0064] When the included angle is small, it indicates that the i-th microphone is close to the user's gaze direction, the value of approaches 0, approaches 1; when the included angle increases, the value of becomes more negative, The value decreases rapidly. As a result, the microphone closer to the user's line of sight obtains a higher weight, while the microphone farther from the line of sight has a lower weight. Thus, during the sound source localization process, more attention is paid to the sound information received by the microphones in the direction that the user may be interested in.

[0065] γ represents an adjustment parameter used to control the rate at which the weight changes with the included angle. The larger the value of γ, the faster the weight decreases as the included angle increases, that is, the stronger the inhibitory effect on the weight of the microphone deviating from the line of sight; the smaller the value of γ, the smoother the change of the weight with the included angle.

[0066] The influence of the sound signal intensity on the dynamic weight is considered. When is greater than , it indicates that the sound signal received by the i-th microphone is relatively strong, and is greater than 2; when is less than , is between 1 and 2.

[0067] It can be concluded from the above that in this embodiment, the dynamic weight is calculated by the first formula, comprehensively considering the included angle between the microphone and the line of sight and the sound signal intensity. It not only allows the microphone closer to the line of sight to obtain a higher weight and focus on the direction of attention, but also enhances the role of the microphone receiving strong signals, making the localization more scientific. It can effectively improve the accuracy and pertinence of sound source localization and enhance the usage experience of wearable devices.

[0068] In an embodiment of the present disclosure, it further includes: Calculating the delay time of the audio signal reaching each microphone relative to the reference microphone; Performing phase adjustment on the audio signal received by each microphone based on the delay time, and performing amplitude adjustment on the audio signal received by each microphone based on the environmental parameters to obtain an independent audio signal corresponding to each microphone; Superposing the independent audio signals of each microphone to obtain a target audio signal.

[0069] In this embodiment, the distances between the sound source and the microphones in the microphone array are different, and there is a time difference for the audio signal to reach different microphones. Taking the reference microphone as a reference, the wave path difference is calculated through geometric relationships (such as in a linear array in a two-dimensional plane ), and then based on the sound propagation speed c in the air, the delay time of the audio signal reaching each microphone relative to the reference microphone is calculated.

[0070] For a signal with a frequency of f, the phase adjustment amount is calculated according to the formula In actual processing, for the audio signal received by the m-th microphone Multiply by a complex exponential factor To achieve phase adjustment, that is . This can make the phase of the target direction signal consistent at each microphone, create conditions for in-phase superposition, and enhance the target direction signal.

[0071] In this embodiment, the amplitude is adjusted based on environmental parameters. Individual differences of microphones and environmental factors will cause the amplitudes of the signals received by each microphone to be different. The gain of each microphone can be measured through experiments as the amplitude adjustment coefficient to calibrate the amplitudes of the signals of each microphone, making the signals received by each microphone comparable in amplitude and improving the subsequent signal superposition effect.

[0072] By superimposing the independent audio signals of each microphone to obtain the target audio signal, after phase and amplitude adjustment, the target direction signal meets the conditions for in-phase superposition, while the signals in other directions are cancelled out or greatly attenuated due to inconsistent phase and amplitude during superposition. Through the formula ( is the weighting coefficient of the m-th microphone) to superimpose the signals of each microphone, the target direction signal can be highlighted, the interference signals in other directions can be suppressed, and the enhanced target audio signal can be obtained.

[0073] It can be concluded from the above that in this embodiment, by calculating the delay time and adjusting the phase and amplitude, the in-phase superposition of the target direction audio signal can be enhanced and the interference in other directions can be suppressed. Considering the environmental parameters to calibrate the amplitude can improve the signal comparability. Finally, by superimposing to obtain the target audio signal, the quality of the audio signal can be effectively improved, the clarity and accuracy of the target audio can be enhanced, and the audio processing effect can be optimized.

[0074] In an embodiment of the present disclosure, phase adjustment is performed on the audio signal received by each microphone based on the delay time, including: Determining the phase adjustment amount of the audio signal based on the delay time; Performing phase adjustment on the audio signal received by each microphone based on the phase adjustment amount.

[0075] In this embodiment, traditionally, for a signal with a frequency of f, the phase adjustment amount can be calculated according to the formula , where is the delay time of the m-th microphone relative to the reference microphone.

[0076] In this embodiment, the influence of environmental factors on the phase can be considered. When sound propagates in different environments, its propagation characteristics will be affected by factors such as temperature and humidity, which will in turn affect the phase of the signal. Assume that the influence of environmental factors on the phase can be represented by a comprehensive environmental factor , and the environmental factor It is related to environmental parameters such as temperature and humidity and can be determined through experiments or empirical formulas.

[0077] The creative formula for determining the phase adjustment amount of the audio signal based on the delay time can be expressed as:

[0078] Wherein, represents the phase adjustment amount of the audio signal received by the m-th microphone, is the frequency of the audio signal, is the delay time of the m-th microphone relative to the reference microphone, is the comprehensive environmental factor. This calculation formula takes into account the influence of environmental factors on the signal phase. When the environmental conditions change, the value of will change accordingly, so that the phase adjustment amount can more accurately reflect the actual situation and improve the effect of audio signal processing.

[0079] It can be concluded from the above that in this embodiment, the phase adjustment amount is determined based on the delay time and the audio signal is phase-adjusted, which can effectively make the phases of the target-direction audio signals received by each microphone consistent. When the signals are superimposed, in-phase enhancement of the target-direction signals can be achieved, while interfering signals in other directions are suppressed, improving the accuracy and quality of audio processing and optimizing the user's audio experience.

[0080] Corresponding to the sound source localization method in the above embodiment, Figure 2 is the structural block diagram of the sound source localization system provided by an embodiment of the present disclosure. For the sake of convenience of description, only the parts related to the embodiments of the present disclosure are shown. Refer to Figure 2 The sound source localization system 20 includes: a first vector determination module 21, a second vector determination module 22, a mapping module 23, and a sound source localization module 24.

[0081] Among them, the first vector determination module 21 is used to determine the gaze space vector corresponding to the gaze direction of the target user based on the eye movements of the target user, where the target user is a user wearing a wearable device; The second vector determination module 22 is used to determine the microphone space vector corresponding to each microphone based on the position coordinates of the microphone array, and the microphone array is arranged on the wearable device; The mapping module 23 is used to establish a mapping relationship based on the gaze space vector and the microphone space vector; The sound source localization module 24 is used to localize the sound source in the direction gazed by the target user based on the mapping relationship.

[0082] In an embodiment of the present disclosure, the first vector determination module 21 is specifically used for: Obtain the eye movement visual features and electrooculogram signal features of the target user; Feature fusion is performed on the eye movement visual features and the electrooculogram signal features to obtain a comprehensive eye movement feature vector; The comprehensive eye movement feature vector is mapped to a fixation space vector.

[0083] In an embodiment of the present disclosure, the mapping module 23 is specifically configured to: Feature extraction is performed on the fixation space vector to obtain fixation space vector features; Feature extraction is performed on the microphone space vector to obtain microphone space vector features; Feature fusion is performed on the fixation space vector features and the microphone space vector features to obtain high-dimensional fusion features; Encoding is performed on the high-dimensional fusion features to obtain the mapping relationship between the fixation space vector and the microphone space vector.

[0084] In an embodiment of the present disclosure, the sound source localization module 24 is specifically configured to: Calculate the dynamic weight of each microphone in the microphone array based on the included angle between the fixation space vector and each microphone space vector; Determine the sound source position in the direction where the target user is looking based on the dynamic weight and the time difference of the audio signals received by different microphones.

[0085] In an embodiment of the present disclosure, the sound source localization module 24 is further specifically configured to: Calculate the dynamic weight of each microphone in the microphone array based on the first formula; The first formula is:

[0086] Wherein, represents the dynamic weight of the i-th microphone, represents the adjustment parameter, represents the included angle between the fixation space vector and the i-th microphone space vector, represents the sound signal intensity received by the i-th microphone, represents the average value of the sound signal intensities received by all microphones.

[0087] In an embodiment of the present disclosure, the sound source localization system 20 further includes: an audio adjustment module, which is specifically configured to: Calculate the delay time of the audio signal reaching each microphone relative to the reference microphone; Perform phase adjustment on the audio signal received by each microphone based on the delay time, and perform amplitude adjustment on the audio signal received by each microphone based on the environmental parameters to obtain an independent audio signal corresponding to each microphone; Superimpose the independent audio signals of each microphone to obtain a target audio signal.

[0088] In one embodiment of the present disclosure, the audio adjustment module is further specifically configured to: Determine the phase adjustment amount of the audio signal based on the delay time; Perform phase adjustment on the audio signals received by each microphone based on the phase adjustment amount.

[0089] Refer to Figure 3 , Figure 3 which is a schematic block diagram of an electronic device provided in an embodiment of the present disclosure. As Figure 3 shown, the electronic device 300 in this embodiment may include: one or more processors 301, one or more input devices 302, one or more output devices 303, and one or more memories 304. The above-mentioned processors 301, input devices 302, output devices 303, and memories 304 communicate with each other through a communication bus 305. The memory 304 is used to store computer programs, and the computer programs include program instructions. The processor 301 is used to execute the program instructions stored in the memory 304. Among them, the processor 301 is configured to call the program instructions to execute the functions of each module in the above-mentioned system embodiments, such as Figure 2 the functions of the modules 21 to 24 shown.

[0090] It should be understood that in the embodiments of the present disclosure, the so-called processor 301 may be a central processing unit (CPU), and this processor may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor, or this processor may also be any conventional processor, etc.

[0091] The input device 302 may include a touchpad, a fingerprint acquisition sensor (for acquiring the fingerprint information and the direction information of the fingerprint of the user), a microphone, etc., and the output device 303 may include a display (such as an LCD), a speaker, etc.

[0092] The memory 304 may include a read-only memory and a random access memory, and provide instructions and data to the processor 301. A part of the memory 304 may also include a non-volatile random access memory. For example, the memory 304 may also store information about the device type.

[0093] In a specific implementation, the processor 301, input device 302, and output device 303 described in the embodiments of the present disclosure may implement the implementation manners described in the first and second embodiments of the sound source localization method provided by the embodiments of the present disclosure, and may also implement the implementation manner of the electronic device described in the embodiments of the present disclosure, which will not be elaborated herein.

[0094] In another embodiment of the present disclosure, a computer-readable storage medium is provided. The computer-readable storage medium stores a computer program, and the computer program includes program instructions. When the program instructions are executed by a processor, all or part of the processes in the methods of the above embodiments are implemented. It can also be completed by instructing relevant hardware through the computer program. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by the processor, the steps of the above method embodiments can be implemented. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file or some intermediate form, etc. The computer-readable medium may include: any entity or system capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), electrical carrier signal, telecommunication signal, and software distribution medium, etc.

[0095] The computer-readable storage medium may be an internal storage unit of the electronic device in any of the foregoing embodiments, such as the hard disk or memory of the electronic device. The computer-readable storage medium may also be an external storage device of the electronic device, such as a plug-in hard disk, smart media card (SMC), secure digital (SD) card, flash card, etc. equipped on the electronic device. Further, the computer-readable storage medium may also include both the internal storage unit and the external storage device of the electronic device. The computer-readable storage medium is used to store the computer program and other programs and data required by the electronic device. The computer-readable storage medium may also be used to temporarily store the data that has been output or will be output.

[0096] Those of ordinary skill in the art can realize that the units and algorithm steps of the examples described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the composition and steps of the examples have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this disclosure.

[0097] Those skilled in the art can clearly understand that for the convenience and conciseness of description, the specific working processes of the above-described electronic devices and units can refer to the corresponding processes in the foregoing method embodiments, and will not be elaborated herein.

[0098] In several embodiments provided in this application, it should be understood that the disclosed electronic devices and methods can be implemented in other ways. For example, the system embodiments described above are merely illustrative. For example, the division of units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the displayed or discussed mutual coupling, direct coupling, or communication connection may be an indirect coupling or communication connection through some interfaces or units, or may be an electrical, mechanical, or other form of connection.

[0099] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of the embodiments of this disclosure.

[0100] In addition, the functional units in the various embodiments of this disclosure can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional units.

[0101] The above is only the specific implementation manner of this disclosure, but the protection scope of this disclosure is not limited thereto. Any person skilled in the art can easily think of various equivalent modifications or substitutions within the technical scope disclosed by this disclosure, and these modifications or substitutions should be covered by the protection scope of this disclosure. Therefore, the protection scope of this disclosure should be subject to the protection scope of the claims.

Claims

1. A sound source localization method, applied to a wearable device, characterized in that: include: Determining a gaze space vector corresponding to a gaze direction of the target user based on an eye movement of the target user, wherein the target user is a user wearing the wearable device; Determining a microphone space vector corresponding to each microphone based on the position coordinates of a microphone array, wherein the microphone array is disposed on the wearable device; Establishing a mapping relationship based on the gaze space vector and the microphone space vector; The sound source in the direction the target user is looking at is located based on the mapping relationship.

2. The sound source localization method according to claim 1, characterized in that: The determining, based on the eye movement of the target user, a gaze space vector corresponding to the gaze direction of the target user includes: Obtain the target user's eye movement visual characteristics and eye electromyography signal characteristics; Performing feature fusion on the eye movement visual features and the eye electromyography signal features to obtain a comprehensive eye movement feature vector; The comprehensive eye movement feature vector is mapped into the gaze space vector.

3. The sound source localization method according to claim 1, characterized in that: The establishing a mapping relationship based on the gaze space vector and the microphone space vector comprises: Extracting features from the gaze space vector to obtain gaze space vector features; Performing feature extraction on the microphone space vector to obtain microphone space vector features; fusing the gaze space vector feature and the microphone space vector feature to obtain a high-dimensional fusion feature; The high-dimensional fusion feature is encoded to obtain a mapping relationship between the gaze space vector and the microphone space vector.

4. The sound source localization method according to claim 1, characterized in that: The sound source in the direction where the target user is looking is located based on the mapping relationship, including: Calculating a dynamic weight of each microphone in the microphone array based on an angle between a gaze space vector and each microphone space vector; The position of the sound source in the direction the target user is looking at is determined based on the dynamic weight and the time difference between the audio signals received by different microphones.

5. The sound source localization method according to claim 4, characterized in that: The calculating the dynamic weight of each microphone in the microphone array based on the angle between the gaze space vector and the space vectors of each microphone includes: Calculating a dynamic weight of each microphone in the microphone array based on a first formula; The first formula is: in, represents the dynamic weight of the i-th microphone, represents the adjustment parameter, represents the angle between the gaze space vector and the i-th microphone space vector, represents the sound signal strength received by the i-th microphone, Indicates the average strength of the sound signals received by all microphones.

6. The sound source localization method according to claim 1, characterized in that: Also includes: Calculate the delay time of the audio signal arriving at each microphone relative to the reference microphone; The phase of the audio signal received by each microphone is adjusted based on the delay time, and the amplitude of the audio signal received by each microphone is adjusted based on the environmental parameters to obtain an independent audio signal corresponding to each microphone; The independent audio signals of each microphone are superimposed to obtain the target audio signal.

7. The sound source localization method according to claim 6, characterized in that: The phase adjustment of the audio signal received by each microphone based on the delay time includes: determining an audio signal phase adjustment amount based on the delay time; The phase of the audio signal received by each microphone is adjusted based on the phase adjustment amount.

8. A sound source localization system, applied to a wearable device, characterized in that: include: a first vector determination module, configured to determine a gaze space vector corresponding to a gaze direction of a target user based on an eye movement of the target user, wherein the target user is a user wearing the wearable device; A second vector determination module, configured to determine a microphone space vector corresponding to each microphone based on a position coordinate of a microphone array, wherein the microphone array is disposed on the wearable device; A mapping module, configured to establish a mapping relationship based on the gaze space vector and the microphone space vector; The sound source localization module is used to locate the sound source in the direction where the target user is looking based on the mapping relationship.

9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 7 are implemented.

10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.

Citation Information

Cited By

  • Voice recording method, device, equipment and medium

    CN120909437A