Sound source positioning and recognition method and device based on multi-modal fusion and storage medium

By combining millimeter-wave radar and microphone arrays to obtain spatial and directional information of the sound source, and combining visual features to confirm identity, the problem of accuracy and reliability of sound source localization and recognition in complex scenarios is solved, and stable speaker localization and identity recognition are achieved.

CN121410646BActive Publication Date: 2026-05-15VALUEHD CORP
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
VALUEHD CORP
Filing Date
2025-12-24
Publication Date
2026-05-15

AI Technical Summary

Technical Problem

Existing sound source localization technologies struggle to accurately and reliably identify speakers in complex scenarios, especially in noisy environments where it is difficult to distinguish between multiple targets moving in the same direction. Furthermore, visual aids are limited by lighting conditions and target occlusion.

Method used

By combining millimeter-wave radar to detect lip movements and vocal cord vibrations, the spatial coordinates of biological targets are obtained. The direction angle of the sound source is obtained using a microphone array. Based on spatial relationships, valid sound-emitting targets are identified. Visual features are obtained through a close-up pan-tilt unit and a camera for identity verification.

Benefits of technology

It can stably and reliably complete speaker localization and identification in complex environments, effectively eliminate interference from non-speaking personnel, and improve the accuracy and reliability of localization and identification.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121410646B_ABST
    Figure CN121410646B_ABST
Patent Text Reader

Abstract

The application discloses a sound source positioning and identification method and device based on multi-modal fusion and a storage medium, relates to the technical field of sound source positioning, and discloses a sound source positioning and identification method based on multi-modal fusion, which comprises the following steps: detecting a biological target with lip movement characteristics and / or vocal cord vibration characteristics in a scene through a millimeter wave radar, and acquiring the spatial coordinates of the biological target; acquiring the sound source direction angle of a sound source through a microphone array, and generating a corresponding sound source direction vector based on the sound source direction angle; judging whether the biological target is an effective sound production target based on the spatial relationship between the spatial coordinates and the sound source direction vector; and if the biological target is the effective sound production target, determining the identity information of the effective sound production target. The application effectively fuses multi-modal data such as radars and audio, significantly improves the accuracy of sound source positioning and identification, and realizes the technical effects of stable and reliable speaker positioning and identity recognition in a complex environment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of sound source localization technology, and in particular to a sound source localization and identification method, device and storage medium based on multimodal fusion. Background Technology

[0002] Direction of Arrival (DOA) estimation technology aims to determine the spatial location of a sound source by processing audio signals received by a microphone array. Currently, DOA estimation mainly employs single-audio modality-based methods or dual-modality methods combining audio and vision. Single-audio modality methods are susceptible to interference in complex noisy environments and struggle to distinguish between multiple targets facing the same direction. While dual-modality methods combining audio and vision can assist in localization through visual information, their performance is limited by ambient lighting conditions and target occlusion, making it difficult to guarantee the stability and reliability of localization accuracy in complex scenes.

[0003] The above content is only used to help understand the technical solution of this application and does not represent an admission that the above content is prior art. Summary of the Invention

[0004] The main purpose of this application is to provide a sound source localization and recognition method, device and storage medium based on multimodal fusion, which aims to solve the technical problem that it is difficult to accurately associate sound sources with visual targets in complex scenes, resulting in the inability to accurately and reliably identify the speaker's identity.

[0005] To achieve the above objectives, embodiments of this application provide a sound source localization and identification method based on multimodal fusion, the sound source localization and identification method based on multimodal fusion comprising:

[0006] The system detects biological targets with lip movement and / or vocal cord vibration characteristics in a scene using millimeter-wave radar and obtains the spatial coordinates of the biological targets.

[0007] The sound source direction angle is obtained by using a microphone array, and a corresponding sound source direction vector is generated based on the sound source direction angle.

[0008] Based on the spatial relationship between the spatial coordinates and the sound source direction vector, it is determined whether the biological target is a valid sound-emitting target;

[0009] If the biological target is the valid vocal target, then the identity information of the valid vocal target is determined.

[0010] In one embodiment, the step of determining whether the biological target is a valid sound-emitting target based on the spatial relationship between the spatial coordinates and the sound source direction vector includes:

[0011] Obtain the target distance of the biological target relative to the millimeter-wave radar;

[0012] Based on the sound source direction vector and the target distance, determine the estimated location of the sound source corresponding to the sound source;

[0013] Calculate the spatial deviation between the spatial coordinates and the estimated location of the sound source;

[0014] If the spatial deviation is less than or equal to a preset threshold, the biological target is determined to be the effective vocal target.

[0015] In one embodiment, the step of determining the identity information of the valid vocal target if the biological target is the valid vocal target includes:

[0016] Obtain the horizontal and vertical angles of the effective sound-emitting target relative to the radar coordinate system;

[0017] Based on the horizontal and vertical angles, the close-up gimbal is driven to adjust the shooting direction so that the close-up lens mounted on the close-up gimbal is aimed at the effective sound target.

[0018] The visual image corresponding to the effective sound-emitting target is obtained through the close-up shot.

[0019] Based on the visual image, extract the target facial features and target human shape features of the effective vocal target;

[0020] The target facial features and target humanoid features are matched with detected targets or pre-registered identity information in the global image to determine the target identity information of the valid vocal target.

[0021] In one embodiment, the step of acquiring the visual image corresponding to the effective sound-emitting target through the close-up shot includes:

[0022] Determine the real-time imaging status of the effective sound-emitting target in the current imaging frame of the close-up shot;

[0023] Based on the real-time imaging status, the zoom level of the close-up shot is adjusted so that the effective sound-emitting target is in the optimal imaging area;

[0024] When the effective sound-emitting target is in the optimal imaging area, the close-up shot is controlled to zoom in and capture the visual image corresponding to the effective sound-emitting target.

[0025] In one embodiment, before the step of matching the target facial features and the target humanoid features with detected targets or pre-registered identity information in the global image to determine the target identity information of the valid sound-emitting target, the sound source localization and recognition method based on multimodal fusion further includes:

[0026] A global image of the target scene is obtained through the panoramic lens in a binocular camera;

[0027] Target detection is performed on the global image to identify multiple detected targets;

[0028] Visual features corresponding to each of the detection targets are extracted, and an association mapping between each detection target and the corresponding visual features is established, wherein the visual features include facial features and human-shaped features.

[0029] In one embodiment, prior to the step of driving the close-up gimbal to adjust the shooting direction based on the horizontal and vertical angles so that the close-up lens mounted on the close-up gimbal is aligned with the effective sound target, the sound source localization and identification method based on multimodal fusion further includes:

[0030] Obtain the reference direction angle corresponding to the current pointing of the close-up gimbal;

[0031] Calculate the angular deviation between the sound source direction angle of the effective sound-emitting target and the reference direction angle;

[0032] If the angle deviation is greater than a preset angle threshold, then the step of adjusting the shooting direction of the close-up gimbal is executed; otherwise, the current direction of the close-up gimbal remains unchanged.

[0033] In one embodiment, the step of detecting biological targets with lip movement characteristics and / or vocal cord vibration characteristics in a scene using millimeter-wave radar and obtaining the spatial coordinates of the biological targets includes:

[0034] The millimeter-wave radar is controlled to transmit FMCW signals and receive reflected echoes from each target.

[0035] Based on the reflected echo, the target distance, horizontal angle and vertical angle of each target relative to the millimeter-wave radar are extracted, and the micro-motion characteristics of each target are identified;

[0036] Based on the micro-motion characteristics, biological targets possessing the lip movement characteristics and / or the vocal cord vibration characteristics are selected from among the targets;

[0037] The spatial coordinates of the biological target are determined based on the target distance, the horizontal angle, and the vertical angle.

[0038] In one embodiment, the step of acquiring the sound source direction angle of the sound source through a microphone array and generating a corresponding sound source direction vector based on the sound source direction angle includes:

[0039] The sound source signal is acquired by the microphone array, and the arrival time difference of the sound source signal received by each microphone is obtained.

[0040] Based on the time difference of arrival, determine the sound source direction angle relative to the microphone array;

[0041] The sound source direction vector is determined based on the sound source direction angle.

[0042] This application embodiment also provides a sound source localization and recognition device based on multimodal fusion. The sound source localization and recognition device based on multimodal fusion includes: a memory, a processor, and a computer program stored in the memory and executable on the processor. The computer program is configured to implement the steps of the sound source localization and recognition method based on multimodal fusion as described above.

[0043] This application embodiment also provides a storage medium, which is a computer-readable storage medium, and stores a computer program on the storage medium. When the computer program is executed by a processor, it implements the steps of the sound source localization and recognition method based on multimodal fusion as described above.

[0044] One or more technical solutions proposed in this application have at least the following technical effects:

[0045] This application uses millimeter-wave radar to detect biological targets with lip movement and / or vocal cord vibration characteristics in a scene and obtains their corresponding spatial coordinates. Leveraging the high sensitivity of millimeter-wave radar to micro-motion physiological signals, it filters out candidate targets that are actually emitting sound, effectively eliminating interference from silent individuals or non-speech sources. Next, it obtains the sound source direction angle through a microphone array and generates a corresponding sound source direction vector based on this angle, thereby obtaining the spatial pointing information of the sound source. Furthermore, based on the spatial relationship between the spatial coordinates and the sound source direction vector, it determines whether the biological target is a valid emitting target, effectively overcoming the problem of traditional single-audio methods struggling to distinguish between multiple speakers in the same direction, while also avoiding the dependence of audio-visual fusion schemes on lighting conditions and target visibility. In addition, after confirming a valid emitting target, its identity information is specifically determined, avoiding invalid identification of non-vocal individuals and significantly improving the accuracy of sound source localization and identification. This application, by fusing radar and audio multimodal data, achieves the technical effect of stable and reliable speaker localization and identification in complex environments. Attached Figure Description

[0046] Figure 1 This is a flowchart illustrating an embodiment of the sound source localization and identification method based on multimodal fusion provided in this application.

[0047] Figure 2 This is a flowchart illustrating Embodiment 2 of the sound source localization and identification method based on multimodal fusion in this application.

[0048] Figure 3This is a simplified flowchart illustrating Embodiment 2 of the sound source localization and identification method based on multimodal fusion in this application.

[0049] Figure 4 This is a schematic diagram of the structure of the sound source localization and recognition device based on multimodal fusion, which is the hardware operating environment involved in the sound source localization and recognition method based on multimodal fusion in the embodiments of this application.

[0050] The purpose, features, and advantages of this application will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation

[0051] It should be understood that the specific embodiments described herein are merely illustrative of this application and are not intended to limit this application.

[0052] To better understand the technical solution of this application, a detailed description will be provided below in conjunction with the accompanying drawings and specific implementation methods.

[0053] Direction of Arrival (DOA) estimation technology aims to determine the spatial location of a sound source by processing audio signals received by a microphone array. Currently, DOA estimation mainly employs single-audio modality-based methods or dual-modality methods combining audio and vision. Single-audio modality methods are susceptible to interference in complex noisy environments and struggle to distinguish between multiple targets facing the same direction. While dual-modality methods combining audio and vision can assist in localization through visual information, their performance is limited by ambient lighting conditions and target occlusion, making it difficult to guarantee the stability and reliability of localization accuracy in complex scenes.

[0054] In view of the above problems, this application proposes a sound source localization and identification method based on multimodal fusion. This method detects biological targets with lip movement characteristics and / or vocal cord vibration characteristics in a scene using millimeter-wave radar and obtains the spatial coordinates of the biological targets. It then obtains the sound source direction angle using a microphone array and generates a corresponding sound source direction vector based on the sound source direction angle. Based on the spatial relationship between the spatial coordinates and the sound source direction vector, it determines whether the biological target is a valid sound-emitting target. If the biological target is a valid sound-emitting target, its identity information is determined.

[0055] It should be noted that the sound source localization and recognition method based on multimodal fusion proposed in this application can be applied to multimodal sound source localization and recognition devices. These devices include a millimeter-wave radar module, a microphone array, and a binocular camera. The three components are integrated into the same hardware device and work collaboratively under a unified device clock and device coordinate system. They can simultaneously collect millimeter-wave radar point cloud data, DOA acoustic data, and visual data for the same scene.

[0056] Specifically, the millimeter-wave radar module is used to transmit FMCW (Frequency Modulated Continuous Wave) signals and receive reflected echoes from various targets in the scene, thereby obtaining spatial information of biological targets relative to multimodal sound source localization and identification devices (or millimeter-wave radar), including target distance (R), horizontal angle (θ), and vertical angle (φ); and further extracting the micro-motion features of biological targets from them to identify physiological activities related to vocalization, such as lip movements or vocal cord vibrations.

[0057] A microphone array consists of multiple microphones arranged in a preset geometric structure, such as a 6-unit microphone array, used to collect multi-channel audio signals and perform voice activity detection (VAD) and direction of arrival estimation based on the audio signals. It outputs the sound source direction angle (α) and the corresponding confidence level (C, 0≤C≤1). The higher the confidence level, the less noise interference the sound source is affected by.

[0058] The binocular camera includes a panoramic lens and a zoomable close-up lens. The panoramic lens captures a global image of the scene; when the millimeter-wave radar determines that a biological target detected by the radar matches the sound source estimated by the DOA in space (i.e., they belong to the same valid sound source), the close-up lens is triggered. Then, based on the spatial coordinates of the valid sound source, the close-up lens's pan-tilt unit rotates and its focal length is adjusted so that the close-up lens is aligned with the valid sound source, thereby acquiring the corresponding close-up image for subsequent identity recognition.

[0059] This application provides a solution that uses millimeter-wave radar to detect biological targets with lip movement and / or vocal cord vibration characteristics in a scene and obtains their corresponding spatial coordinates. Leveraging the high sensitivity of millimeter-wave radar to micro-motion physiological signals, candidate targets that are actually emitting sound are screened, effectively eliminating interference from silent individuals or non-speech sources. Next, the sound source direction angle is obtained through a microphone array, and a corresponding sound source direction vector is generated based on this angle, thereby obtaining the spatial pointing information of the sound source. Furthermore, based on the spatial relationship between the spatial coordinates and the sound source direction vector, it is determined whether the biological target is a valid emitting target, effectively overcoming the problem of traditional single-audio methods struggling to distinguish between multiple speakers in the same direction, while also avoiding the dependence of audio-visual fusion schemes on lighting conditions and target visibility. In addition, after confirming a valid emitting target, its identity information is specifically determined, avoiding invalid identification of non-vocal individuals and significantly improving the accuracy of sound source localization and identification. This application, by fusing radar and audio multimodal data, achieves the technical effect of stable and reliable speaker localization and identification in complex environments.

[0060] It should be noted that the executing entity in this embodiment can be a computing service device with data processing, network communication, and program execution functions, such as an integrated microphone array, millimeter-wave radar, a close-up gimbal for binocular cameras, a smart conference terminal, etc., or an electronic device capable of realizing the above functions, a multimodal sound source localization and recognition device, etc. The following description uses a multimodal sound source localization and recognition device as an example to illustrate this embodiment and the subsequent embodiments.

[0061] The sound source localization and identification method based on multimodal fusion proposed in the first embodiment of this application is described in reference [link to relevant documentation]. Figure 1 The method includes steps S10 to S40:

[0062] Step S10: Detect biological targets with lip movement characteristics and / or vocal cord vibration characteristics in the scene using millimeter-wave radar, and obtain the spatial coordinates of the biological targets.

[0063] It should be noted that millimeter-wave radar is a type of sensing radar operating in the millimeter-wave band (typically 30 GHz to 300 GHz). Its emitted FMCW (Frequency Modulated Continuous Wave) signals possess high resolution, strong anti-jamming capabilities, and penetration, enabling the detection of target range, angle, lip movement, and vocal cord vibration characteristics. Lip movement and vocal cord vibration are both types of micro-motion characteristics used to distinguish whether a target is speaking. Lip movement refers to the minute movements produced by the opening and closing of the lips during vocalization, while vocal cord vibration refers to the minute vibrations caused by the opening and closing of the vocal cords in the larynx. Both of these micro-motion characteristics produce unique phase modulations in the radar echo, which can be identified through time-frequency analysis.

[0064] As one possible implementation, step S10 includes steps S110 to S140:

[0065] Step S110: Control the millimeter-wave radar to transmit FMCW signals and receive reflected echoes from each target.

[0066] Step S120: Based on the reflected echo, extract the target distance, horizontal angle and vertical angle of each target relative to the millimeter-wave radar, and identify the micro-motion characteristics of each target.

[0067] Step S130: Based on the micro-motion characteristics, select the biological targets that have the lip movement characteristics and / or the vocal cord vibration characteristics from among the targets.

[0068] Step S140: Determine the spatial coordinates of the biological target based on the target distance, the horizontal angle, and the vertical angle.

[0069] It should be noted that the target distance represents the straight-line spatial distance from the biological target to the millimeter-wave radar, the horizontal angle represents the azimuth angle of the biological target in the horizontal plane relative to the front of the multimodal sound source localization and identification device, and the vertical angle represents the elevation angle of the biological target in the vertical plane relative to the horizontal plane; these three parameters together constitute the position description of the biological target in the polar coordinate system with the millimeter-wave radar as the origin.

[0070] It is understood that, since the millimeter-wave radar module, microphone array, and binocular camera are integrated into the same hardware device, the spatial parameters of the target relative to the millimeter-wave radar and microphone array, such as distance, azimuth, and elevation angle, mentioned in this embodiment and subsequent embodiments can be equivalent to the corresponding spatial parameters of the target relative to the multimodal sound source localization and identification device.

[0071] In this embodiment, the millimeter-wave radar is first controlled to periodically transmit FMCW signals while simultaneously receiving reflected echoes from each target. By performing time-frequency analysis on the reflected echoes, such as Fast Fourier Transform (FFT), the target distance, horizontal angle, and vertical angle of each target relative to the millimeter-wave radar are extracted, and the micro-motion characteristics corresponding to each target, such as lip movement characteristics and vocal cord vibration characteristics, are identified. Targets with lip movement characteristics and / or vocal cord vibration characteristics are then identified as biological targets.

[0072] Based on this, the spatial coordinates of the biological target are determined according to the target distance (R), horizontal angle (θ), and vertical angle (φ) extracted by the millimeter-wave radar. Specifically, a device coordinate system is established with the installation location of the multimodal sound source localization and identification device as the origin. The X-axis points horizontally to the right, the Y-axis points directly in front of the device, and the Z-axis points vertically upwards. Next, based on the polar coordinate representation (R, θ, φ) of the biological target in the millimeter-wave radar, it is transformed into Cartesian coordinates in the device coordinate system through coordinate transformation to obtain the spatial coordinates (X1, Y1, Z1) of the biological target. The calculation formula is: X1 = R Sinθ, Y1=R cosθ, Z1=R tanφ.

[0073] Traditional sound source localization methods rely solely on audio signals and are susceptible to interference from background noise, reverberation, or multiple speakers. In contrast, this implementation method uses millimeter-wave radar to pre-screen biological targets with genuine vocal physiological characteristics, which can avoid misidentifying non-vocal personnel (such as those who are present but do not speak) or environmental noise sources as sound sources, thereby improving the accuracy of subsequent sound source localization and identification.

[0074] Step S20: Obtain the sound source direction angle of the sound source through the microphone array, and generate the corresponding sound source direction vector based on the sound source direction angle.

[0075] It should be noted that a microphone array is an audio acquisition device consisting of multiple microphones arranged in a specific geometric structure (such as a linear, circular, or uniform circular array) to capture the time difference of sound waves arriving at each microphone in space.

[0076] The sound source azimuth angle refers to the incident azimuth angle of the sound source relative to the microphone array, usually with the horizontal plane as the reference.

[0077] The sound source direction vector is a unit vector that maps the sound source direction angle to the device coordinate system. In this embodiment, the sound source direction vector (X2, Y2) can be represented as (cosα, sinα), where α is the sound source direction angle; the vertical component Z2 is temporarily set to 0 and will be supplemented by millimeter-wave radar data later.

[0078] In this embodiment, multi-channel sound source signals can be acquired through a 6-unit uniform circular array microphone array. Then, non-speech segments, such as fan noise and page turning sounds, are first filtered out using VAD (Voice Activity Detection). Then, based on the improved MUSIC algorithm (which adds a noise suppression module to the traditional MUSIC spectrum estimation), the sound source direction angle (α) is calculated, and the corresponding confidence level (C, 0≤C≤1) is output simultaneously to characterize the reliability of the direction estimation in the current noise environment. The higher the confidence level, the less noise interference the sound source is subjected to. Subsequently, the sound source direction angle is converted into a two-dimensional direction vector as the sound source direction vector.

[0079] As one possible implementation, step S20 includes steps S210 to S40:

[0080] Step S210: Acquire sound source signals through the microphone array and obtain the arrival time difference of the sound source signals received by each microphone.

[0081] Step S220: Determine the sound source direction angle relative to the microphone array based on the arrival time difference.

[0082] Step S230: Determine the sound source direction vector based on the sound source direction angle.

[0083] It should be noted that the sound source signal refers to the sound wave signal generated by the target sound source and propagated through the air medium, and finally received by each unit in the microphone array; the time difference of arrival refers to the time difference that the same sound source signal wavefront takes to reach different microphones in the microphone array, which is essentially the time delay caused by the difference in path length when the sound wave propagates at a finite speed.

[0084] In this embodiment, the GCC-PHAT (Generalized Cross-Correlation-Phase Transform) algorithm can be used to calculate the time difference of arrival. Specifically, firstly, a Fast Fourier Transform is performed on the sound source signals collected by all microphones to convert the time-domain signals to the frequency domain; then, for the frequency domain signals corresponding to any two microphones, their cross-power spectra are calculated and phase normalization is performed; next, an Inverse Fourier Transform is performed on the normalized frequency domain results to obtain the time-domain cross-correlation function characterizing the similarity between the two signals; by locating the peak position of this cross-correlation function, the corresponding time delay is determined, thereby determining the time difference of arrival of the sound source signals received by each microphone.

[0085] Optionally, a time delay estimation method based on phase difference can also be used to determine the time difference of arrival. Specifically, firstly, a Fourier transform is performed on the sound source signal collected by the microphone array to obtain the complex spectrum of each microphone in the frequency domain; then, within the effective frequency band with a high signal-to-noise ratio (e.g., 1kHz-4kHz), the phase difference between different microphones at the same frequency component is extracted, and the phase difference is de-coiled to eliminate phase jumps caused by periodicity; since the phase difference is linearly related to the time difference of sound wave arrival within this effective frequency band, that is, the higher the frequency, the greater the phase shift caused by the same time difference, the corresponding time delay can be deduced from this linear relationship, thereby obtaining the time difference of arrival of the sound source signal received by each microphone.

[0086] It is important to note that, in order to avoid phase ambiguity caused by spatial aliasing, the spacing between each microphone should be less than half the wavelength of the sound wave corresponding to the highest frequency in the target frequency band, so as to ensure that the phase difference uniquely corresponds to a physical time difference.

[0087] Step S30: Based on the spatial relationship between the spatial coordinates and the sound source direction vector, determine whether the biological target is a valid sound-emitting target.

[0088] It should be noted that an effective vocal target refers to a biological target that simultaneously meets the physiological characteristics of vocalization and whose spatial location is consistent with the direction of the sound source.

[0089] In this embodiment, the sound source direction vector is a unit vector that only contains direction information and lacks a distance dimension; therefore, the two cannot be directly compared. Thus, when determining the spatial relationship between spatial coordinates and the sound source direction vector, the target distance provided by millimeter-wave radar is used to "project" the direction vector to the corresponding distance, forming an estimated sound source location that can be compared with the spatial coordinates.

[0090] As one possible implementation, step S30 includes steps S310 to S340:

[0091] Step S310: Obtain the target distance of the biological target relative to the millimeter-wave radar.

[0092] Step S320: Determine the estimated location of the sound source corresponding to the sound source based on the sound source direction vector and the target distance.

[0093] Step S330: Calculate the spatial deviation between the spatial coordinates and the estimated location of the sound source.

[0094] Step S340: If the spatial deviation is less than or equal to a preset threshold, then the biological target is determined to be the effective sound-emitting target.

[0095] In this embodiment, the target distance R of the biological target relative to the millimeter-wave radar is obtained directly from the data measured by the millimeter-wave radar. Since it is assumed that the biological target detected by the millimeter-wave radar is the current speaker, and that its sound source location largely coincides with its body location in space, the target distance R measured by the millimeter-wave radar is used as a reasonable approximation of the sound source distance, denoted as K=R. Based on this, the estimated sound source location (K0) in the device coordinate system is obtained. X2, K Y2).

[0096] Based on this, the spatial coordinates (X1, Y1, Z1) provided by the millimeter-wave radar and the estimated location of the sound source (K) are calculated. X2, K The Euclidean distance between Y2 and Y3, as the spatial deviation (diff), is expressed as follows:

[0097]

[0098] Where K is the estimated distance to the sound source, which is approximated by the target distance R measured by millimeter-wave radar, i.e., K=R.

[0099] If the spatial deviation is less than or equal to a preset threshold (e.g., 0.5m), the biological target is determined to be from the same target as the current sound source, and it is confirmed as a valid sound source. Conversely, if the spatial deviation exceeds the preset threshold, the DOA sound source localization result is considered to be spatially aligned with the biological target detected by the millimeter-wave radar, and may be environmental noise, echo, or other interfering sound source. Therefore, the sound source information is ignored to avoid misidentification.

[0100] In multi-person scenarios, the above matching can be performed sequentially on all biological targets detected by millimeter-wave radar, calculating the spatial deviation between the spatial coordinates of each biological target and the sound source direction vector. By comparing the deviation values ​​corresponding to each biological target, the biological target with the smallest spatial deviation and meeting the preset threshold condition is selected as the final effective sound-emitting target.

[0101] Specifically, in multi-person scenarios, such as conference rooms or smart classrooms, multiple biological targets exist simultaneously within the detection range of millimeter-wave radar, and there may be situations where multiple people speak simultaneously or take turns speaking. In this case, the millimeter-wave radar detects multiple biological targets with lip movement characteristics and / or vocal cord vibration characteristics and obtains their respective spatial coordinates. Simultaneously, the microphone array calculates the sound source direction angle based on the DOA estimation algorithm and generates the corresponding sound source direction vector. Since the sound source direction vector only contains direction information and not distance information, it cannot be directly matched with spatial coordinates. Therefore, using the distance information of each biological target provided by the millimeter-wave radar, the sound source direction vector is extended along that direction to the corresponding distance, forming the estimated sound source location point. Subsequently, the spatial deviation between the spatial coordinates of each biological target and the estimated sound source location is calculated, for example, using Euclidean distance as the deviation index. When the deviation between the spatial coordinates of a biological target and the estimated sound source location is less than a preset threshold (e.g., 0.5 meters), and it possesses obvious physiological characteristics of vocalization (e.g., continuous vocal cord vibration or periodic lip movement), then the biological target is determined to be a valid vocal target. If multiple biological targets meet the above conditions, the one with the smallest deviation is selected as the current speaker. Finally, after confirming the current speaker, the identity of the current speaker is matched and verified by combining visual features.

[0102] This implementation method deeply integrates millimeter-wave radar, DOA direction, and visual data, enabling the device to accurately identify the real speaker and associate their identity information in complex multi-target environments. This effectively solves the problem of misjudgment caused by sound reflection, visual obstruction, or interference from non-speaking personnel in traditional methods, and achieves end-to-end accurate association from sound source localization to identity recognition.

[0103] In this embodiment, while DOA estimation can provide azimuth information of the sound source, it cannot distinguish the specific speaker from multiple targets located in similar directions. Although millimeter-wave radar can accurately measure the spatial position of biological targets, it is difficult to directly confirm whether the target is emitting sound. Therefore, this embodiment spatially registers the spatial coordinates of the biological target measured by millimeter-wave radar with the DOA direction vector, calculates the spatial deviation between the position and the sound source direction for each biological target, and determines whether it is a valid speaker based on whether the spatial deviation meets a preset threshold. In multi-person application scenarios, the target with the smallest spatial deviation and meeting the threshold condition is preferentially selected as the final speaker, thereby achieving reliable identification of specific speaker targets in complex environments and significantly improving the accuracy of sound source localization and identification.

[0104] Step S40: If the biological target is the valid vocal target, then determine the identity information of the valid vocal target.

[0105] In this embodiment, the gimbal can be controlled to point according to the horizontal angle θ and vertical angle φ measured by millimeter-wave radar, and a high-definition close-up image of the effective sound-emitting target is acquired through a zoom lens. Subsequently, facial features and human figure features are extracted from the close-up image, and the facial features and human figure features are matched and associated with the detected target in the global image captured by the panoramic lens, or compared with pre-recorded identity information to determine the specific identity information of the effective sound-emitting target.

[0106] It should be noted that, in order to protect user privacy, all biometric data (including facial features and human features) are collected only after obtaining the user's explicit authorization, and the acquisition methods strictly comply with relevant laws and regulations and privacy protection standards.

[0107] Optionally, voiceprint recognition technology can be combined to assist in identity verification by analyzing the voice features of a valid voice target.

[0108] This implementation establishes a unique mapping from close-up targets to specific identities by accurately matching high-resolution individual features captured in close-up shots with multiple detection targets or pre-stored identity records in the global image.

[0109] This application effectively overcomes the limitation of single audio modality in distinguishing the real speaker in noisy environments and multi-target scenarios with DOA (Directional Occlusion) information by spatially registering precise spatial coordinates provided by millimeter-wave radar with the DOA information. It also avoids the problem of inaccurate positioning caused by environmental factors such as lighting conditions and target occlusion in visual-assisted DOA positioning schemes. After spatially locating the effective speaker target, this application further drives a close-up pan-tilt unit to align with the effective speaker target, extracting visual features such as facial features and human figures from a high-resolution close-up image. Based on this, combined with the detected target from a panoramic view or pre-recorded identity information, the identity of the effective speaker target is identified through feature comparison and matching mechanisms. By fusing multimodal data from millimeter-wave radar, acoustic DOA, and vision, this application constructs a cross-modal collaborative speaker verification closed loop, significantly improving the accuracy of sound source localization and identity recognition in complex environments such as multi-person interaction, providing reliable technical support for applications such as intelligent conferencing systems and remote collaboration platforms.

[0110] Based on the above embodiments of this application, in the second embodiment of this application, the content that is the same as or similar to that in the first embodiment can be referred to the above description, and will not be repeated hereafter. Based on this, please refer to... Figure 2 The sound source localization and identification method based on multimodal fusion includes steps S410 to S450 in step S40:

[0111] Step S410: Obtain the horizontal and vertical angles of the effective sound-emitting target relative to the radar coordinate system.

[0112] Step S420: Based on the horizontal angle and the vertical angle, drive the close-up gimbal to adjust the shooting direction so that the close-up lens mounted on the close-up gimbal is aimed at the effective sound target.

[0113] Step S430: Obtain the visual image corresponding to the effective sound-emitting target through the close-up shot.

[0114] Step S440: Based on the visual image, extract the target facial features and target human shape features of the effective vocal target.

[0115] Step S450: Match the target facial features and the target humanoid features with the detected targets or pre-registered identity information in the global image to determine the target identity information of the valid vocal target.

[0116] It should be noted that a close-up gimbal is an electric platform with horizontal rotation and vertical pitch freedom, which integrates a binocular camera. The close-up lens is used for close-range visual capture of specific biological targets. Adjusting the shooting direction refers to controlling the gimbal motor to rotate to a specified azimuth and pitch angle so that the binocular camera is precisely aimed at the target. The pre-registered identity information refers to a set of pre-stored characteristic templates and identity tags of known personnel.

[0117] In this embodiment, the multimodal sound source localization and identification device is integrated into the close-up pan-tilt unit. The millimeter-wave radar and microphone array are fixedly installed to provide stable global sound source direction and target spatial coordinates, which do not move with the pan-tilt unit. Meanwhile, the close-up lens in the binocular camera rotates with the close-up pan-tilt unit to perform close-range visual capture of the confirmed valid sound-emitting target.

[0118] In this embodiment, the horizontal angle θ and vertical angle φ measured by the millimeter-wave radar are converted into target control commands for the close-up gimbal to drive the close-up gimbal to adjust the shooting direction so that the close-up lens mounted on the close-up gimbal is aimed at the effective sound target.

[0119] Optionally, to avoid frequent movements of the gimbal due to minor fluctuations in the DOA azimuth angle or radar angle, a minimum turning threshold (e.g., angle change ≥ 10°) can be set. When new sound source information is available, turning is only triggered if the direction of the corresponding effective sound source deviates from the minimum turning threshold.

[0120] After the gimbal completes its pointing adjustment, a trigger signal controls the close-up camera to acquire images. Specifically, based on the target distance detected by the millimeter-wave radar, the focal length of the close-up camera is automatically adjusted to ensure that the effective sound-emitting target maintains a suitable imaging ratio in the close-up image.

[0121] As one possible implementation, step S430 includes steps S4310 to S4330:

[0122] Step S4310: Determine the real-time imaging status of the effective sound-emitting target in the current imaging frame of the close-up shot.

[0123] Step S4320: Based on the real-time imaging state, adjust the zoom level of the close-up shot to place the effective sound-emitting target in the optimal imaging area.

[0124] Step S4330: When the effective sound-emitting target is in the optimal imaging area, control the close-up shot to zoom in and capture the visual image corresponding to the effective sound-emitting target.

[0125] It should be noted that real-time imaging status refers to the target's visual attributes obtained through real-time processing of video frames in the current close-up shot. These attributes include the center offset of the effective sound-emitting target in the image, the proportion of the face area to the image height, and image sharpness. The optimal imaging area refers to a pre-defined area of ​​the image that is beneficial for subsequent face recognition and feature extraction.

[0126] In this embodiment, after driving the close-up gimbal to initially align the close-up lens with the effective sound-emitting target, the face region of the effective sound-emitting target is first located using a face detection algorithm, and its real-time imaging state in the current frame is evaluated. Based on the real-time imaging state, the pixel offset between the center of the face region and the center of the frame is determined. When the pixel offset between the center of the face region and the center of the frame is detected to exceed a preset offset threshold, the optical zoom module is controlled to adjust the zoom magnification of the close-up lens so that the effective sound-emitting target is in the optimal imaging region. When the face region is detected to be in the optimal imaging region of the frame, the close-up lens is triggered to acquire a visual image.

[0127] Optionally, when the pixel offset between the center of the face region and the center of the image exceeds a preset offset threshold, the angle compensation value of the close-up gimbal is calculated based on the pixel offset; then, the pointing angle of the close-up gimbal is finely adjusted based on the angle compensation value so that the effective sound target is in the optimal imaging area, and the close-up shot is triggered to acquire a visual image.

[0128] Optionally, after driving the close-up gimbal to initially align the close-up with the effective sound target, the face region of the effective sound target is first located using a face detection algorithm, and its real-time imaging status in the current frame is evaluated. Based on the real-time imaging status, the proportion of the face region height to the total frame height is determined. When the proportion of the face region height to the total frame height is detected to be lower than a preset proportion threshold, the optical zoom module is controlled to adjust the zoom magnification of the close-up until the face region enters the preset optimal imaging area. When it is confirmed that the effective sound target is in the optimal imaging area, the close-up captures the corresponding visual image.

[0129] Optionally, after driving the close-up gimbal to initially align the close-up lens with the effective sound target, the face region of the effective sound target is first located using a face detection algorithm, and its real-time imaging status in the current frame is evaluated. Based on the real-time imaging status, the image sharpness corresponding to the face region is determined. When the image sharpness is detected to be lower than a preset sharpness threshold, the autofocus function is activated, and the lens focus position or exposure parameters are dynamically adjusted to improve image sharpness. When the image sharpness is detected to reach or exceed the preset sharpness threshold, the close-up lens is triggered to acquire the visual image.

[0130] As one possible implementation, steps S417 to S419 are included before step S420:

[0131] Step S417: Obtain the reference direction angle corresponding to the current pointing of the close-up gimbal.

[0132] Step S418: Calculate the angular deviation between the sound source direction angle of the effective sound-emitting target and the reference direction angle.

[0133] Step S419: If the angle deviation is greater than the preset angle threshold, then execute the step of adjusting the shooting direction of the close-up gimbal; otherwise, keep the current direction of the close-up gimbal unchanged.

[0134] It should be noted that the reference orientation angle refers to the angular coordinates of the close-up gimbal's current shooting direction in the horizontal direction, which can be obtained in real time through the gimbal's built-in orientation encoder. The preset angle threshold is used to prevent the close-up gimbal from moving frequently due to small angle changes.

[0135] In this embodiment, after confirming a valid sound target, the current reference angle of the close-up gimbal is first obtained, and then the reference angle is compared with the sound source direction angle of the valid sound target. If the deviation between the two exceeds a preset angle threshold, the close-up gimbal is driven to adjust the shooting direction to align with the new valid sound target, avoiding repeated responses to small angle fluctuations of the same speaker. If the angle deviation does not exceed the preset angle threshold, it is determined that the close-up gimbal is now aligned with the valid sound target, and there is no need to repeatedly adjust the shooting direction, thereby reducing mechanical wear, saving power consumption and improving equipment stability.

[0136] As one possible implementation, steps S447 to S449 are included before step S450:

[0137] Step S447: Obtain a global image of the target scene using the panoramic lens in the binocular camera.

[0138] Step S448: Perform target detection on the global image to identify multiple detected targets.

[0139] Step S449: Extract the visual features corresponding to each of the detection targets and establish an association mapping between each of the detection targets and the corresponding visual features, wherein the visual features include face features and human shape features.

[0140] It should be noted that in this embodiment, when acquiring a global image and identifying the detected target through the panoramic lens, an identity association process is initiated simultaneously. By matching the visual features of each detected target with a pre-stored identity feature database in real time, the corresponding identity information is automatically associated with each identifiable detected target. Therefore, in subsequent steps, once the visual features of a valid speaking target are successfully matched with a detected target in the global image, the identity information already bound to that detected target can be directly obtained, thereby completing the rapid determination of the speaker's identity.

[0141] This embodiment constructs an association mapping containing global detection targets and their visual features before close-up identity matching. This allows the target face features and target human shape features corresponding to the effective voice targets extracted from close-up shots to be compared with the detection targets in the global image, thereby completing the accurate identification of the voicer's identity.

[0142] For example, to help understand the implementation process of the sound source localization and recognition method based on multimodal fusion obtained by combining the above embodiments, please refer to... Figure 3 , Figure 3 A simplified flowchart of a sound source localization and identification method based on multimodal fusion is provided, specifically:

[0143] After the device is started, it simultaneously performs acoustic DOA detection and millimeter-wave radar detection. First, it collects sound source signals through a microphone array and uses a sound source localization algorithm to determine the direction of sound arrival, thus identifying the location of the sound source. At the same time, the millimeter-wave radar scans the target scene to filter out biological targets that are actually making sounds, effectively eliminating interference from silent people or non-speech sources. Subsequently, the acoustic DOA detection determines whether the collected sound source signal is human voice; if it is determined to be non-human voice, the sound source signal is ignored and monitoring continues; if it is determined to be human voice, the sound source direction vector of the sound source signal is spatially matched with the spatial coordinates of the biological target detected by the millimeter-wave radar, and the spatial deviation between the two is further determined to meet a preset threshold.

[0144] If the spatial deviation is greater than a preset threshold, it indicates that the biological target and the current sound source are from different targets and are therefore invalid sound sources. The acoustic DOA detection loop is then restarted. If the spatial deviation is less than or equal to the preset threshold, the biological target and the current sound source are determined to be the same target and are considered valid sound-emitting targets.

[0145] Next, based on the horizontal and vertical angles of the biological target detected by the millimeter-wave radar relative to the radar coordinate system, the gimbal is driven to turn to the spatial coordinate position corresponding to the biological target, so that the close-up lens is aimed at the effective sound-emitting target. Once the close-up lens is in position, a visual image corresponding to the effective sound-emitting target is acquired, and the target's facial features and humanoid features are extracted from the visual image.

[0146] Prior to this, the panoramic lens in the binocular camera had acquired a global image of the target scene. By performing target detection on the global image, multiple targets were identified, and the visual features (including facial features and human-shaped features) corresponding to each target were extracted. A correlation mapping between the target and its visual features was established, and each target was automatically associated with its corresponding identity information based on its visual features.

[0147] Based on this, the target facial features and target human shape features extracted from the visual image are compared with the visual features of each detected target in the global image. Once a match is successful, the identity information corresponding to the detected target can be directly obtained, thereby completing the accurate identification and location of the valid speaker. Alternatively, these features can be matched with pre-registered identity information, which can also confirm the specific identity information of the valid speaker, thus accurately identifying and locating the valid speaker.

[0148] This embodiment effectively overcomes the problem of traditional single-audio methods' inability to distinguish between multiple speakers speaking in the same direction by deeply fusing millimeter-wave radar, acoustic DOA, and visual modalities. It also avoids the dependence of audio-visual fusion schemes on lighting conditions and target visibility. By determining whether the biological target with lip-movement vocal cord vibration characteristics detected by millimeter-wave radar and the current sound source detected by acoustic DOA are the same target, this application effectively eliminates interference from false sound sources, avoids invalid identification of non-speaking personnel, and significantly improves the accuracy and reliability of sound source localization and identification. This provides reliable technical support for applications such as intelligent conferencing systems and remote collaboration platforms.

[0149] This application provides a sound source localization and recognition device based on multimodal fusion. The sound source localization and recognition device based on multimodal fusion includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the sound source localization and recognition method based on multimodal fusion in the above embodiment 1.

[0150] The following is for reference. Figure 4 The diagram illustrates a structural schematic of a sound source localization and identification device based on multimodal fusion suitable for implementing embodiments of this application. The sound source localization and identification device based on multimodal fusion in the embodiments of this application may include various hardware and software components for implementing the sound source localization and identification method based on multimodal fusion. Figure 4 The sound source localization and identification device based on multimodal fusion shown is merely an example and should not impose any limitations on the functionality and scope of use of the embodiments of this application.

[0151] like Figure 4As shown, the sound source localization and recognition device based on multimodal fusion may include a processing unit 1001 (e.g., a central processing unit, a graphics processing unit, etc.), which can perform various appropriate actions and processes according to a program stored in read-only memory (ROM) 1002 or a program loaded from storage device 1003 into random access memory (RAM) 1004. The random access memory 1004 also stores various programs and data required for the operation of the sound source localization and recognition device based on multimodal fusion. The processing unit 1001, ROM 1002, and RAM 1004 are interconnected via a bus 1005. An input / output (I / O) interface 1006 is also connected to the bus. Typically, the following systems can be connected to I / O interface 1006: input devices 1007 including, for example, touchscreens, touchpads, keyboards, etc.; output devices 1008 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 1003 including, for example, magnetic tapes, hard disks, etc.; and communication devices 1009. Communication device 1009 allows the multimodal fusion-based sound source localization and identification device to communicate wirelessly or wiredly with other devices to exchange data. Although the figure shows a multimodal fusion-based sound source localization and identification device with various systems, it should be understood that it is not required to implement or possess all of the systems shown. More or fewer systems can be implemented alternatively.

[0152] Specifically, according to the embodiments disclosed in this application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments disclosed in this application include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device, or installed from storage device 1003, or installed from read-only memory 1002. When the computer program is executed by processing device 1001, it performs the functions defined in the methods of the embodiments disclosed in this application.

[0153] The sound source localization and recognition device based on multimodal fusion provided in this application, employing the sound source localization and recognition method based on multimodal fusion in the above embodiments, can solve the technical problem of difficulty in accurately associating sound sources with visual targets in complex scenes, resulting in the inability to accurately and reliably identify the speaker's identity. Compared with the prior art, the beneficial effects of the sound source localization and recognition device based on multimodal fusion provided in this application are the same as those of the sound source localization and recognition method based on multimodal fusion provided in the above embodiments, and other technical features in this sound source localization and recognition device based on multimodal fusion are the same as those disclosed in the previous embodiment method, and will not be repeated here.

[0154] It should be understood that the various parts disclosed in this application can be implemented using hardware, software, firmware, or a combination thereof. In the description of the above embodiments, specific features, structures, materials, or characteristics can be combined in any suitable manner in one or more embodiments or examples.

[0155] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

[0156] This application provides a computer-readable storage medium having computer-readable program instructions (i.e., a computer program) stored thereon, which are used to execute the sound source localization and recognition method based on multimodal fusion in the above embodiments.

[0157] The computer-readable storage medium provided in this application may be, for example, a USB flash drive, but is not limited to, electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to: electrical connections having one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof. In this embodiment, the computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, system, or device. The program code contained on the computer-readable storage medium may be transmitted using any suitable medium, including but not limited to: wires, optical cables, radio frequency (RF), etc., or any suitable combination thereof.

[0158] The aforementioned computer-readable storage medium may be included in a sound source localization and identification device based on multimodal fusion; or it may exist independently and not assembled into a sound source localization and identification device based on multimodal fusion.

[0159] The aforementioned computer-readable storage medium carries one or more programs. When these programs are executed by a sound source localization and recognition device based on multimodal fusion, the sound source localization and recognition device based on multimodal fusion performs the following actions: detects biological targets with lip movement characteristics and / or vocal cord vibration characteristics in a scene using millimeter-wave radar, and obtains the spatial coordinates of the biological targets; obtains the sound source direction angle of the sound source using a microphone array, and generates a corresponding sound source direction vector based on the sound source direction angle; determines whether the biological target is a valid sound-emitting target based on the spatial relationship between the spatial coordinates and the sound source direction vector; and if the biological target is a valid sound-emitting target, determines the identity information of the valid sound-emitting target.

[0160] Computer program code for performing the operations of this application can be written in one or more programming languages ​​or a combination thereof, including object-oriented programming languages ​​such as Java, Smalltalk, and C++, as well as conventional procedural programming languages ​​such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0161] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0162] The modules described in the embodiments of this application can be implemented in software or hardware. The names of the modules do not necessarily limit the functionality of the unit itself.

[0163] The readable storage medium provided in this application is a computer-readable storage medium that stores computer-readable program instructions (i.e., a computer program) for executing the aforementioned sound source localization and recognition method based on multimodal fusion. This solves the technical problem of difficulty in accurately associating sound sources with visual targets in complex scenes, leading to the inability to accurately and reliably identify the speaker. Compared with the prior art, the beneficial effects of the computer-readable storage medium provided in this application are the same as those of the sound source localization and recognition method based on multimodal fusion provided in the above embodiments, and will not be repeated here.

[0164] This application provides a computer program product, including a computer program that, when executed by a processor, implements the steps of the sound source localization and recognition method based on multimodal fusion as described above.

[0165] The computer program product provided in this application can solve the technical problem that it is difficult to accurately associate sound sources with visual targets in complex scenes, resulting in the inability to accurately and reliably identify the speaker's identity. Compared with the prior art, the beneficial effects of the computer program product provided in this application are the same as those of the sound source localization and recognition method based on multimodal fusion provided in the above embodiments, and will not be repeated here.

[0166] The above are merely preferred embodiments of this application and do not limit the patent scope of this application. Any equivalent structural or procedural transformations made using the content of this application's specification and drawings, or direct or indirect applications in other related technical fields, are similarly included within the patent scope of this application.

[0167] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or system that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or system. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or system that includes that element.

[0168] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method.

[0169] The above are merely preferred embodiments of this application and do not limit the patent scope of this application. Any equivalent structural or procedural transformations made using the content of this application's specification and drawings, or direct or indirect applications in other related technical fields, are similarly included within the patent protection scope of this application.

Claims

1. A sound source localization and identification method based on multimodal fusion, characterized in that, The sound source localization and identification method based on multimodal fusion includes: By using millimeter-wave radar to detect biological targets with lip movement characteristics and / or vocal cord vibration characteristics in a scene and to obtain the spatial coordinates of the biological targets, the high sensitivity of millimeter-wave radar to micro-movement physiological signals can be used to screen out candidate targets that are actually making sounds, effectively eliminating interference from silent people or non-speech sound sources. The sound source direction angle is obtained by using a microphone array, and a corresponding sound source direction vector is generated based on the sound source direction angle. Based on the spatial relationship between the spatial coordinates and the sound source direction vector, it is determined whether the biological target is a valid sound-emitting target; If the biological target is the effective sound-emitting target, obtain the horizontal and vertical angles of the effective sound-emitting target relative to the radar coordinate system; Based on the horizontal and vertical angles, the close-up gimbal is driven to adjust the shooting direction so that the close-up lens mounted on the close-up gimbal is aligned with the effective sound target; wherein, the millimeter-wave radar and the microphone array are integrated into the close-up gimbal in a fixed manner and do not move with the rotation of the close-up gimbal; when there is new sound source information that requires the close-up gimbal to adjust the shooting direction, a minimum turning threshold is set; when the deviation of the new effective sound target from the current pointing direction exceeds the minimum turning threshold, the close-up gimbal is triggered to turn. The visual image corresponding to the effective sound-emitting target is obtained through the close-up shot. Based on the visual image, extract the target facial features and target human shape features of the effective vocal target; The target facial features and target humanoid features are matched with detected targets or pre-registered identity information in the global image to determine the target identity information of the valid vocal target.

2. The sound source localization and identification method based on multimodal fusion as described in claim 1, characterized in that, The step of determining whether the biological target is a valid sound-emitting target based on the spatial relationship between the spatial coordinates and the sound source direction vector includes: Obtain the target distance of the biological target relative to the millimeter-wave radar; Based on the sound source direction vector and the target distance, determine the estimated location of the sound source corresponding to the sound source; Calculate the spatial deviation between the spatial coordinates and the estimated location of the sound source; If the spatial deviation is less than or equal to a preset threshold, the biological target is determined to be the effective vocal target.

3. The sound source localization and identification method based on multimodal fusion as described in claim 1, characterized in that, The step of acquiring the visual image corresponding to the effective sound-emitting target through the close-up shot includes: Determine the real-time imaging status of the effective sound-emitting target in the current imaging frame of the close-up shot; Based on the real-time imaging status, the zoom level of the close-up shot is adjusted so that the effective sound-emitting target is in the optimal imaging area; When the effective sound-emitting target is in the optimal imaging area, the close-up shot is controlled to zoom in and capture the visual image corresponding to the effective sound-emitting target.

4. The sound source localization and identification method based on multimodal fusion as described in claim 1, characterized in that, Before the step of matching the target facial features and the target humanoid features with detected targets or pre-registered identity information in the global image to determine the target identity information of the valid sound-emitting target, the sound source localization and recognition method based on multimodal fusion further includes: A global image of the target scene is obtained through the panoramic lens in a binocular camera; Target detection is performed on the global image to identify multiple detected targets; Visual features corresponding to each of the detection targets are extracted, and an association mapping between each detection target and the corresponding visual features is established, wherein the visual features include facial features and human-shaped features.

5. The sound source localization and identification method based on multimodal fusion as described in claim 1, characterized in that, Before the step of driving the close-up gimbal to adjust the shooting direction based on the horizontal and vertical angles so that the close-up lens mounted on the close-up gimbal is aligned with the effective sound target, the sound source localization and recognition method based on multimodal fusion further includes: Obtain the reference direction angle corresponding to the current pointing of the close-up gimbal; Calculate the angular deviation between the sound source direction angle of the effective sound-emitting target and the reference direction angle; If the angle deviation is greater than the preset angle threshold, then the step of adjusting the shooting direction of the driving close-up gimbal is executed; Otherwise, keep the current orientation of the gimbal unchanged.

6. The sound source localization and identification method based on multimodal fusion as described in claim 1, characterized in that, The step of detecting biological targets with lip movement characteristics and / or vocal cord vibration characteristics in a scene using millimeter-wave radar and obtaining the spatial coordinates of the biological targets includes: The millimeter-wave radar is controlled to transmit FMCW signals and receive reflected echoes from each target. Based on the reflected echo, the target distance, horizontal angle and vertical angle of each target relative to the millimeter-wave radar are extracted, and the micro-motion characteristics of each target are identified; Based on the micro-motion characteristics, biological targets possessing the lip movement characteristics and / or the vocal cord vibration characteristics are selected from among the targets; The spatial coordinates of the biological target are determined based on the target distance, the horizontal angle, and the vertical angle.

7. The sound source localization and identification method based on multimodal fusion as described in claim 1, characterized in that, The step of obtaining the sound source direction angle through a microphone array and generating a corresponding sound source direction vector based on the sound source direction angle includes: The sound source signal is acquired by the microphone array, and the arrival time difference of the sound source signal received by each microphone is obtained. Based on the time difference of arrival, determine the sound source direction angle relative to the microphone array; The sound source direction vector is determined based on the sound source direction angle.

8. A sound source localization and identification device based on multimodal fusion, characterized in that, The sound source localization and identification device based on multimodal fusion includes: a memory, a processor, and a computer program stored in the memory and executable on the processor, the computer program being configured to implement the steps of the sound source localization and identification method based on multimodal fusion as described in any one of claims 1 to 7.

9. A storage medium, characterized in that, The storage medium is a computer-readable storage medium, and a computer program is stored on the storage medium. When the computer program is executed by a processor, it implements the steps of the sound source localization and recognition method based on multimodal fusion as described in any one of claims 1 to 7.