Conference room spokesman positioning method, system and device and storage medium

By coordinating the processing of radar and microphone devices, and leveraging the stability of radar and the precision of microphones, the problem of poor sound source localization accuracy was solved, resulting in more accurate speaker localization and improved system robustness and sound pickup performance.

CN122017729APending Publication Date: 2026-05-12YEALINK (XIAMEN) NETWORK TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
YEALINK (XIAMEN) NETWORK TECHNOLOGY CO LTD
Filing Date
2025-12-31
Publication Date
2026-05-12

AI Technical Summary

Technical Problem

Existing sound source localization technologies have poor accuracy in complex scenarios such as conference rooms, and cannot reliably determine who the speaker is and where they are.

Method used

By combining the sensing data from radar and microphone devices, the first spatial information of candidate targets detected by radar is matched and correlated with the second spatial information of the sound-emitting targets determined by microphone devices. By utilizing the stability of radar and the accuracy of microphone, an optimized positioning result of the target speaker is generated.

Benefits of technology

It improves the speaker's positioning accuracy and system robustness, enabling more reliable and accurate speaker positioning in complex scenarios, laying the foundation for clear sound pickup and smooth interaction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122017729A_ABST
    Figure CN122017729A_ABST
Patent Text Reader

Abstract

According to the conference room spokesman positioning method, system and device and the storage medium, the first space information of the candidate target is collected through the radar device, the second space information of the sounding target is collected through the microphone device, the acoustic positioning error is corrected by means of the acoustic positioning compensation strategy, and the positioning accuracy is improved. And screening out a target spokesman based on a spatial matching relationship between the first spatial information and the second spatial information, and fusing the two types of spatial information to determine final spatial information of the target spokesman, so that high-precision positioning of the spokesman in the conference room can be realized. Therefore, the technical problem that the pickup performance is greatly reduced due to the fact that a sound source positioning scheme depending on a beam forming technology is easily influenced by position deviation and the positioning precision is insufficient in the prior art is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of sound source localization technology, specifically to a method, system, device, and storage medium for locating a speaker in a conference room. Background Technology

[0002] Sound source localization technology aims to determine the physical location of one or more sound sources in space, and it has broad application prospects in fields such as video conferencing, intelligent robots, security monitoring, and voice interaction devices. Currently, mainstream sound source localization schemes mainly rely on microphone array-based beamforming technology, which has poor sound source localization accuracy. Summary of the Invention

[0003] This application provides a method, system, device, and storage medium for locating a speaker in a conference room, aiming to solve the problem of poor sound source localization accuracy.

[0004] Firstly, a method for locating a speaker in a conference room is provided, the method comprising: Acquire the first spatial information of at least one candidate target output by the radar device; Acquire second spatial information of at least one sound-emitting target output by the microphone device; Based on the spatial matching relationship between the first spatial information and the second spatial information, the target speaker is determined; Based on the first spatial information and the second spatial information corresponding to the target speaker, the spatial information of the target speaker is determined.

[0005] Secondly, a method for locating a speaker in a conference room is provided, aiming to solve the problem of poor sound source localization accuracy. The method includes: Acquire the fourth spatial information of at least one fifth candidate target output by the radar device; By analyzing the captured images using a camera, facial motion features of at least one sixth candidate target can be obtained. Based on the matching degree between the facial motion features and the features in the preset vocal facial motion feature library, the camera determines at least one seventh candidate target from the at least one sixth candidate target; Obtain the fifth spatial information of at least one seventh candidate target output by the camera; Based on the fourth spatial information and the fifth spatial information, the first target speaker is determined; Based on the fourth and fifth spatial information corresponding to the first target speaker, the spatial information of the first target speaker is determined.

[0006] In some embodiments, the method further includes: The radar device transmits the fourth spatial information to the camera; The camera acquires the image based on the fourth spatial information.

[0007] Thirdly, a method for locating a speaker in a conference room is provided, aiming to solve the problem of poor sound source localization accuracy. The method includes: Acquire the first spatial information of at least one candidate target output by the radar device; Acquire second spatial information of at least one sound-emitting target output by the microphone device; Based on the spatial matching relationship between the first spatial information and the second spatial information, the first spatial information corresponding to each sound-emitting target is determined; Based on the first spatial information and the second spatial information corresponding to each sound-emitting target, the spatial information of each sound-emitting target is determined.

[0008] Fourthly, a conference room speaker positioning system is provided to solve the problem of poor sound source positioning accuracy. The system includes a processing device, a radar device, and a microphone device. The radar device is used to output first spatial information of at least one candidate target; The microphone device is used to output second spatial information of at least one sound-emitting target; The processing device is used to determine the target speaker based on the spatial matching relationship between the first spatial information and the second spatial information; The processing device is further configured to determine the spatial information of the target speaker based on the first spatial information and the second spatial information corresponding to the target speaker.

[0009] Fifthly, a conference room speaker positioning system is provided to solve the problem of poor sound source positioning accuracy. The system includes a processing device, a radar device, and a camera. The radar device is used to output fourth spatial information of at least one fifth candidate target; The camera is used to analyze the acquired images to obtain facial motion features of at least one sixth candidate target; The camera is also used to determine at least one seventh candidate target from the at least one sixth candidate target based on the matching degree between the facial motion features and the features in the preset vocal facial motion feature library; The processing device is used to acquire the fifth spatial information of at least one seventh candidate target output by the camera; The processing device is further configured to determine a first target speaker based on the fourth spatial information and the fifth spatial information; The processing device is further configured to determine the spatial information of the first target speaker based on the fourth and fifth spatial information corresponding to the first target speaker.

[0010] Sixthly, a conference room speaker positioning system is provided to solve the problem of poor sound source positioning accuracy. The system includes a processing device, a radar device, and a microphone device. The radar device is used to output first spatial information of at least one candidate target; The microphone device is used to output second spatial information of at least one sound-emitting target; The processing device is used to determine the first spatial information corresponding to each sound-emitting target based on the spatial matching relationship between the first spatial information and the second spatial information; The processing device is further configured to determine the spatial information of each sound-emitting target based on the first spatial information and the second spatial information corresponding to each sound-emitting target.

[0011] Seventhly, a conference room speaker positioning device is provided to solve the problem of poor sound source positioning accuracy. The device includes a first acquisition module and a first processing module. The first acquisition module is used to acquire first spatial information of at least one candidate target output by the radar device; The first acquisition module is further configured to acquire second spatial information of at least one sound-emitting target output by the microphone device; The first processing module is used to determine the target speaker based on the spatial matching relationship between the first spatial information and the second spatial information; The first processing module is further configured to determine the spatial information of the target speaker based on the first spatial information and the second spatial information corresponding to the target speaker.

[0012] Eighthly, a conference room speaker positioning device is provided to solve the problem of poor sound source positioning accuracy. The device includes a second acquisition module and a second processing module. The second acquisition module is used to acquire the fourth spatial information of at least one fifth candidate target output by the radar device; The second processing module is used to analyze the acquired images through the camera to obtain facial motion features of at least one sixth candidate target; The second processing module is further configured to determine at least one seventh candidate target from the at least one sixth candidate target based on the matching degree between the facial motion features and the features in the preset vocal facial motion feature library by the camera; The second acquisition module is further configured to acquire the fifth spatial information of at least one seventh candidate target output by the camera; The second processing module is further configured to determine the first target speaker based on the fourth spatial information and the fifth spatial information; The second processing module is further configured to determine the spatial information of the first target speaker based on the fourth and fifth spatial information corresponding to the first target speaker.

[0013] Ninthly, a conference room speaker positioning device is provided to solve the problem of poor sound source positioning accuracy. The device includes a third acquisition module and a third processing module. The third acquisition module is used to acquire first spatial information of at least one candidate target output by the radar device; The third acquisition module is also used to acquire second spatial information of at least one sound-emitting target output by the microphone device; The third processing module is used to determine the first spatial information corresponding to each sound-emitting target based on the spatial matching relationship between the first spatial information and the second spatial information. The third processing module is also used to determine the spatial information of each sound-emitting target based on the first spatial information and the second spatial information corresponding to each sound-emitting target.

[0014] Tenthly, a method for locating a speaker in a conference room is provided to address the problem of poor sound source localization accuracy. The method includes: Acquire the first spatial information of at least one candidate target output by the radar device; Acquire second spatial information of at least one sound-emitting target determined by the microphone device, wherein the radar device is integrated into the microphone device; Based on the spatial matching relationship between the first spatial information and the second spatial information, the target speaker is determined; Based on the first spatial information and the second spatial information corresponding to the target speaker, the spatial information of the target speaker is determined.

[0015] Eleventhly, a method for locating a speaker in a conference room is provided to solve the problem of poor sound source localization accuracy. The method includes: Acquire the first spatial information of at least one candidate target output by the radar device; Acquire second spatial information of at least one sound-emitting target determined by the microphone device, wherein the radar device is integrated into the microphone device; Based on the spatial matching relationship between the first spatial information and the second spatial information, the first spatial information corresponding to each sound-emitting target is determined; Based on the first spatial information and the second spatial information corresponding to each sound-emitting target, the spatial information of each sound-emitting target is determined.

[0016] In a twelfth aspect, a conference room speaker positioning system is provided to solve the problem of poor sound source positioning accuracy. The system includes a radar device and a microphone device; the radar device is integrated into the microphone device. The radar device is used to output first spatial information of at least one candidate target; The microphone device is used to determine second spatial information of at least one vocal target; determine a target speaker based on the spatial matching relationship between the first spatial information and the second spatial information; and determine the spatial information of the target speaker based on the first spatial information and the second spatial information corresponding to the target speaker.

[0017] In a thirteenth aspect, a conference room speaker positioning system is provided to solve the problem of poor sound source positioning accuracy. The system includes a radar device and a microphone device; the radar device is integrated into the microphone device. The radar device is used to output first spatial information of at least one candidate target; The microphone device is used to determine second spatial information of at least one sound-emitting target; determine first spatial information corresponding to each sound-emitting target based on the spatial matching relationship between the first spatial information and the second spatial information; and determine spatial information of each sound-emitting target based on the first spatial information and the second spatial information corresponding to each sound-emitting target.

[0018] In a fourteenth aspect, a computer-readable storage medium is provided, in which a computer program is stored, which, when executed by a processor, is used to implement the methods of the first, second, third, tenth, or eleventh aspects described above.

[0019] In a fifteenth aspect, a computer program product is provided, comprising a non-transitory computer-readable storage medium storing a computer program, wherein when the computer program is executed by a processor, it performs the methods described in the first, second, third, tenth, or eleventh aspects above.

[0020] This application provides a method, system, device, and storage medium for speaker localization in a conference room. By collaboratively processing the sensing data from a radar device and a microphone device, it improves the accuracy of speaker localization and the robustness of the system. The scheme first matches and correlates the first spatial information of candidate targets detected by the radar device with the second spatial information of the sound-emitting target determined by the microphone device, thereby accurately identifying the speaking target, i.e., the target speaker, and solving the correspondence problem between the sound source and the sound-emitting body. Based on this, this application comprehensively utilizes the two types of spatial information from the successfully matched radar device and microphone device sides to generate an optimized localization result for the target speaker, i.e., spatial information. This process, at its underlying logic, leverages the stability of the radar device in spatial detection and the accuracy of the microphone in sound source orientation. Through information complementarity, it overcomes the limitations of relying solely on acoustic localization, such as sensitivity to beam pointing deviation and susceptibility to environmental interference. Therefore, the system can achieve more reliable and accurate speaker localization in complex scenarios such as conference rooms, laying the foundation for clear sound pickup and smooth interaction. Attached Figure Description

[0021] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0022] Figure 1 A flowchart illustrating a conference room speaker location method provided in an embodiment of this application; Figure 2 A schematic diagram illustrating the switching of speaker facial tracking according to an embodiment of this application; Figure 3 A schematic flowchart illustrating another method for locating a speaker in a conference room, as provided in an embodiment of this application; Figure 4 A schematic flowchart illustrating another method for locating a speaker in a conference room, as provided in an embodiment of this application; Figure 5 A schematic diagram of a conference room speaker positioning system provided in an embodiment of this application; Figure 6 Another structural schematic diagram of the conference room speaker positioning system provided in the embodiments of this application; Figure 7 This is a schematic diagram illustrating the scenario of speaker positioning in a conference room, provided in an embodiment of this application. Figure 8 A schematic diagram of a conference room speaker positioning device provided in an embodiment of this application; Figure 9Another schematic diagram of the conference room speaker positioning device provided in the embodiments of this application; Figure 10 Another schematic diagram of the conference room speaker positioning device provided in the embodiments of this application; Figure 11 A schematic flowchart illustrating another method for locating a speaker in a conference room, as provided in an embodiment of this application; Figure 12 A schematic flowchart illustrating another method for locating a speaker in a conference room, as provided in an embodiment of this application; Figure 13 This is another schematic diagram illustrating the location of a speaker in a conference room, as provided in an embodiment of this application. Detailed Implementation

[0023] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0024] In the field of sound source localization technology, to achieve high-quality sound pickup (such as in remote conferencing and voice interaction), the industry has long followed a highly unified and continuously evolving technical approach: based on microphone arrays, improving localization accuracy and speech separation through optimized beamforming algorithms. The core assumption of this approach is that sound source localization is essentially and only a signal processing problem of "analyzing direction from acoustic signals."

[0025] Therefore, research and improvements by industry professionals naturally focus on: designing more complex array topologies, developing more powerful noise suppression and dereverberation algorithms, and utilizing deep learning to improve the robustness of azimuth estimation. Despite these efforts, a fundamental performance ceiling remains: all acoustic processing is extremely dependent on the estimation of the initial direction of the sound source. Even a slight angular deviation can lead to beam pointing errors, causing a precipitous drop in pickup performance.

[0026] Faced with this bottleneck, the mainstream approach still seeks breakthroughs within established paradigms, believing that the problem stems from "insufficient accuracy or robustness of acoustic algorithms." However, the inventors of this application, through in-depth analysis of numerous real-world scenarios (such as multi-person, mobile, and noisy conference rooms), have identified an unquestioned cognitive blind spot in existing technical approaches: the concept and processing of "sound source" and "physical entity emitting sound" have been equated.

[0027] The inventors realized that the key problem might not lie in the "insufficient accuracy of acoustic localization itself," but rather in the lack of an independent and reliable information dimension in pure acoustic solutions to determine a priori "who is making the sound" and "where they might be." Acoustic signals can tell the system "which direction the sound is coming from," but they cannot reliably answer the two preliminary questions, "Is that a person speaking?" and "Which specific person is speaking?" in complex scenarios. This makes acoustic processing like searching for a target in the dark with an uncertain location and unclear features, naturally making it extremely sensitive to initial aiming errors.

[0028] This cognitive shift is crucial, as it means moving beyond the single dimension of acoustic technology and redefining the problem from a more fundamental "speaker perception" task. The real technical challenge then emerges: how to reliably detect and locate the speaker at the moment of speaking, without relying on the sound itself, thereby providing a stable and reliable spatial prior for subsequent acoustic processing? Once this long-neglected fundamental problem was clearly identified, the path to finding solutions across technological fields became clear. The inventors turned their attention to radar sensing technology, especially its micro-Doppler detection capability, which is highly sensitive to micro-movements of living organisms (such as lip and chest vibrations), thus establishing a new technical route of "using radar to sense the physiological characteristics of the speaker to achieve pre-screening and coarse localization, and then coordinating with acoustic localization."

[0029] In view of this, embodiments of this application provide a method, system, device, and storage medium for locating a speaker in a conference room. First, the first spatial information of a candidate target detected by a radar device is matched and associated with the second spatial information of a sound-emitting target determined by a microphone device, thereby accurately identifying the speaking target, i.e., the target speaker, and solving the correspondence problem between sound sources and sound emitters. Based on this, this application comprehensively utilizes the two types of spatial information from the successfully matched radar device and microphone device sides to generate an optimized location result for the target speaker, i.e., spatial information. That is, by collaboratively processing the perception data from the radar device and microphone device, the accuracy of speaker location and system robustness can be improved. This process, at its underlying logic, utilizes the stability of the radar device in spatial detection and the accuracy of the microphone in sound source orientation. Through information complementarity, it overcomes the limitations of relying solely on acoustic positioning, which is sensitive to beam pointing deviations and susceptible to environmental interference. Therefore, the system can achieve more reliable and accurate speaker location in complex scenarios such as conference rooms, laying the foundation for clear sound pickup and smooth interaction.

[0030] The conference room speaker location method, system, and device provided in this application will be explained and described below with reference to specific embodiments: Firstly, such as Figure 1 As shown, Figure 1This is a flowchart illustrating a method for locating a speaker in a conference room, as provided in an embodiment of this application. Figure 1 The method of the embodiment includes, but is not limited to, steps 101-104: Step 101: Obtain the first spatial information of at least one candidate target output by the radar device.

[0031] Step 102: Obtain the second spatial information of at least one sound-emitting target output by the microphone device.

[0032] Step 103: Determine the target speaker based on the spatial matching relationship between the first spatial information and the second spatial information.

[0033] Step 104: Determine the spatial information of the target speaker based on the first spatial information and the second spatial information corresponding to the target speaker.

[0034] First of all, it should be noted that Figure 1 The conference room speaker positioning method provided in this embodiment can be applied to a conference room speaker positioning system, which may include a processing device, a radar device, and a microphone device. Some or all of the processing device, radar device, and microphone device may be independently configured devices. Optionally, the processing device may be integrated into the microphone device; alternatively, the radar device may also be integrated into the microphone device. For example, when the processing device is independent of the microphone device and radar device, steps 101-104 can be interpreted as being executed by the processing device; when the processing device is integrated into the microphone device, steps 101-104 can be interpreted as being executed by the microphone device (i.e., the processing device within the microphone device). Optionally, Figure 1 The conference room speaker location method provided in the embodiments can also be applied to a conference room speaker location device, wherein the device can be the aforementioned processing device independent of the microphone device and radar device, or it can be a microphone device integrated with the processing device. Furthermore, the above two configuration methods include various layout forms where the radar device and microphone device are integrated or not integrated, which will not be elaborated here, but will be specifically described in the embodiments below.

[0035] In the embodiments of this application, the radar device can be a device that uses wave signals to detect targets and obtain target spatial information. For example, the type of radar device can include at least one of microwave radar, millimeter-wave radar, ultra-wideband radar, lidar, ultrasonic radar, etc., and this application is not limited thereto; the number of radar devices can be one or more, and this application mainly uses one as an example for illustration. Candidate targets can be all possible physical entities or moving objects identified by the radar device within the detection area, including participants, tables and chairs, mobile devices, etc. in a conference room.

[0036] In optional embodiments of this application, the radar device may also determine candidate targets through the following embodiments, as detailed below: First, the radar device emits wave signals into the area (such as a conference room or classroom). When the wave signals propagate to the surface of objects (such as people or objects) within the area, they are reflected to form echo signals.

[0037] Then, the radar device analyzes the received echo signal to obtain the vocal micro-motion characteristics of at least one first candidate target. For example, the radar device can use a constant false alarm rate (CFAR) detection algorithm to perform preliminary detection of the echo signal and obtain a preliminary target list. For instance, for low signal-to-noise ratio (SNR) echo signals of static human targets, the radar device can dynamically adjust the number of reference units and the detection threshold factor (e.g., increasing the number of reference units from a first value to a second value, and adjusting the threshold factor from a third value to a fourth value, where the fourth value is less than the third value). This reduces the probability of missing detection of static human targets with low SNR, thus obtaining a preliminary target list that includes static potential targets (i.e., users who are static and may vocalize) and dynamic targets (including users who are dynamic and vocalizing and / or dynamic and not vocalizing). This ensures that static vocal targets are effectively acquired. The preliminary target list includes at least one first candidate target. Further, for each first candidate target, the radar device extracts features from the echo signal of each first candidate target to obtain the vocal micro-motion characteristics of each first candidate target.

[0038] For example, the radar device extracts the micro-Doppler features of the echo signal of each first candidate target. For instance, the radar device extracts the micro-Doppler features of the echo signal caused by lip micro-movements and / or chest cavity vibrations during vocalization. Then, the radar device performs time-frequency analysis on the micro-Doppler features, such as through short-time Fourier transform or wavelet transform, to extract the time-domain variation law of the micro-Doppler features, obtaining the vocal micro-motion feature spectrum characterizing the first candidate target, i.e., obtaining the vocal micro-motion features of the first candidate target. In an optional embodiment, when extracting the vocal micro-motion features of the first candidate target, the radar device can also use high-precision phase measurement to track the change of minute displacements on the target surface over time, extracting vibration modes corresponding to the speech frequency range, thus obtaining the vocal micro-motion features of the first candidate target. In an optional embodiment, when the target emits sound, its vocal cords, oral cavity, chest cavity, and other parts will generate minute vibrations, modulating the phase, frequency, or amplitude of the reflected echo signal. Accordingly, the radar device analyzes the echo signal of the first candidate target. For example, the radar device can extract periodic or non-periodic minute motion features related to sound emission through Doppler effect analysis, phase demodulation, or high-resolution range-Doppler spectrum analysis, thereby obtaining the sound emission micro-motion features of the first candidate target. It should be noted that this application does not limit the specific method by which the radar device determines the sound emission micro-motion features of the first candidate target.

[0039] Then, the radar device determines the at least one candidate target from the at least one first candidate target based on the matching degree between the acoustic micro-motion characteristics of the first candidate target and features in a preset acoustic micro-motion characteristic library. For example, the radar device determines the matching degree between the acoustic micro-motion characteristics of the first candidate target and each feature in the preset acoustic micro-motion characteristic template library, wherein the matching degree measures the similarity between the acoustic micro-motion characteristics extracted from the radar echo signal and the features in the acoustic micro-motion characteristic template library. By calculating the matching degree, the degree of conformity between the currently detected acoustic micro-motion characteristics and known features in the library can be quantified; a higher matching degree indicates a greater likelihood that the target (i.e., the user) is making a sound, and vice versa. The matching degree can be calculated using various methods. For example, it can be calculated by measuring the Euclidean distance, cosine similarity, or correlation coefficient between the vocal micro-motion features and features in the vocal micro-motion feature template library. The smaller the Euclidean distance or the higher the cosine similarity / correlation coefficient, the higher the matching degree. In practice, a mapping relationship between Euclidean distance, cosine similarity, or correlation coefficient and matching degree can be constructed to satisfy the relationship that the smaller the Euclidean distance or the higher the cosine similarity / correlation coefficient, the higher the matching degree. Alternatively, other methods can be used, which are not limited in this application. Alternatively, a trained classifier (such as a support vector machine or neural network) can be used to classify the extracted vocal micro-motion features. The confidence score output by the classifier can be used as the matching degree. The higher the confidence score, the higher the matching degree. In practice, a mapping relationship between confidence score and matching degree can be constructed to satisfy the relationship that the higher the confidence score, the higher the matching degree. Alternatively, other methods can be used, which are not limited in this application.

[0040] Therefore, when the vocal micro-motion features of the first candidate target match the features in the vocal micro-motion feature template library to a predetermined standard, the first candidate target can be identified as a candidate target for subsequent speaker location. For example, a predetermined threshold corresponding to the matching degree can be set. If the vocal micro-motion features of the first candidate target do not match the features in the vocal micro-motion feature template library to a predetermined threshold, the radar device determines the first candidate target as a non-vocal dynamic target (such as a waving or walking person) or irrelevant interference. Conversely, when the vocal micro-motion features match the features in the vocal micro-motion feature template library to a predetermined threshold, the radar device identifies the corresponding first candidate target as a candidate target. In this case, the candidate targets include dynamic and vocal targets and / or static and vocal targets. In an optional embodiment, the radar device can also sort the matching degrees of all first candidate targets that meet the predetermined threshold, and then select a predetermined number of first candidate targets as candidate targets from all first candidate targets that meet the predetermined threshold in descending order of matching degree.

[0041] Furthermore, the aforementioned pre-defined vocal micro-motion feature library can be a database or model storing micro-motion feature templates corresponding to known vocal behaviors. The features in the vocal micro-motion feature library can be pre-trained through experiments, simulations, or machine learning, serving as a benchmark for determining whether a target is vocalizing. The vocal micro-motion feature library can contain vocal micro-motion feature templates for different speech types, speech rates, and vocal intensities. The construction of the vocal micro-motion feature library can be based on collecting and analyzing a large amount of radar echo data from human vocalization, extracting common vocal micro-motion features, such as the vocal cord vibration frequency range, the amplitude of lip movements, and the frequency of lip movements, and quantifying these into feature vectors or temporal patterns for storage. Additionally, the vocal micro-motion feature library can be updated by continuously collecting new vocal data and using supervised or unsupervised learning algorithms to iteratively optimize the feature library, enabling it to adapt to more diverse vocal scenarios and individual differences.

[0042] It should be explained that this application effectively addresses the problem that radar devices may detect dynamic and non-vocal targets when acquiring candidate targets, leading to interference and affecting the accuracy of speaker location. By analyzing the vocal micro-motion characteristics of the radar echo signal and filtering based on its matching degree with a pre-set vocal micro-motion feature library, targets that are truly vocal (whether dynamic or static) can be accurately identified, thus excluding non-vocal interference targets (e.g., still objects or people with only slight body movements but no vocalization) from the candidate targets. This allows the subsequent speaker location process to be based on a purer and more accurate set of candidate targets, significantly improving the accuracy and reliability of speaker determination.

[0043] Furthermore, in an optional embodiment of this application, after determining the candidate target, the radar device can also determine the first spatial information of the candidate target through the following implementation methods, as follows: In the embodiments of this application, the first spatial information can be data output by the radar device for each candidate target, describing its position in space. This data may include, but is not limited to, part or all of the target's distance, azimuth, elevation, radial velocity, and three-dimensional coordinates in a preset coordinate system (such as polar coordinates, rectangular coordinates, Cartesian coordinates, etc., which are not limited in this application). Simply put, the radar device can directly output the position coordinates of each candidate target in the preset coordinate system, such as three-dimensional coordinates; or, the radar device can output the distance and angle information of each candidate target relative to the radar device, and then convert this distance and angle information into the first spatial information in the preset coordinate system through a coordinate transformation algorithm.

[0044] In optional embodiments of this application, the radar device may also determine the first spatial information of the candidate target through the following implementation: the radar device acquires the first distance and first direction of the at least one candidate target relative to the radar device; then, the radar device determines the first spatial information of each candidate target based on the first distance and first direction of each candidate target.

[0045] The first distance can be understood as the radial distance between the candidate target and the radar device. The radar device can calculate the first distance by measuring the time delay between the transmitted wave signal and the received echo signal. For example, the radar device can transmit a wave signal and record the total time from transmission to reception of the echo, and then calculate the radial distance based on the propagation speed of the wave signal. Alternatively, the radar device can also use Frequency Modulated Continuous Wave Radar (FMCW) technology to calculate the radial distance by measuring the frequency difference between the transmitted and received signals.

[0046] The first direction can be the direction of the candidate target relative to the radar device, such as the azimuth and / or elevation angle of the candidate target relative to the radar device. The azimuth angle represents the direction of the candidate target on the horizontal plane, and can be represented by the angle between the projection of the candidate target onto the horizontal detection plane of the radar device and the radar reference axis. The elevation angle represents the direction of the candidate target on the vertical plane, and can be represented by the angle between the projection of the candidate target onto the vertical detection plane of the radar and the radar reference axis. Alternatively, the radar device can receive the echo signal through an array antenna and use beamforming or angle-of-arrival estimation algorithms to determine the incident direction of the echo signal, thereby obtaining the first direction corresponding to the candidate target. Alternatively, multiple radar devices can be used for cross-location, and the first direction of the candidate target relative to the radar device can be deduced from the distance information measured by different radar devices.

[0047] Therefore, the radar device can use the first range and first direction of the candidate target as the first spatial information of the candidate target. That is, the first spatial information of each candidate target can include the first range (R) and first direction (such as azimuth (θ) and elevation (φ)) corresponding to the candidate target. Optionally, the first spatial information can also include the radial velocity (v) of the candidate target, which can be used to determine whether the candidate target is dynamic or static. For example, if the radial velocity is greater than or equal to a third preset threshold, the candidate target is determined to be dynamic; conversely, if the radial velocity is less than the third preset threshold, the candidate target is determined to be static.

[0048] In optional embodiments, the radar device can also determine the first spatial information of the candidate target based on the first distance and the first direction of the candidate target. For example, the radar device can output the position coordinates of the candidate target in a preset coordinate system, i.e., the first spatial information, based on the three-dimensional calculation results of the first distance and the first direction. For example, the coordinate values ​​(x, y, z) in the XYZ coordinate system in the scenario of the ceiling-mounted microphone device, or the coordinates (Rθ, φ) in the polar coordinate system, represent the spatial position of the candidate target.

[0049] In a multi-radar deployment scenario, where there are multiple radar devices (e.g., multiple radar devices integrated into a microphone device and evenly distributed), the multiple radar devices can send the four-dimensional parameters of the candidate target (first distance R, azimuth (θ), elevation (φ), and radial velocity v) calculated by themselves to the fusion processing unit (which can be a server, one of the multiple radar devices, or a processing device in the microphone device; this application does not limit this). Then, the fusion processing unit uses a fusion algorithm, such as triangulation or least squares fitting, to spatially register and associate the four-dimensional parameters of the candidate target under the multiple radar devices, and calculates the position coordinates of the candidate target in a preset coordinate system, i.e., the first spatial information of the candidate target. Accordingly, step 101 is to obtain the first spatial information of at least one candidate target output by the fusion processing unit.

[0050] It should be explained that this application avoids indirect measurement errors that may exist in related methods by directly measuring the first distance and first direction corresponding to the candidate target using a radar device, thus ensuring the accuracy and reliability of the first spatial information. Specifically, the radar device can penetrate non-metallic obstacles and is unaffected by the environment, and its echo signal has high stability, enabling stable acquisition of candidate target position data even in complex conference environments. Then, the first distance and first direction corresponding to the candidate target are used to obtain the first spatial information corresponding to the candidate target through precise coordinate transformation or multi-radar fusion algorithms, which greatly improves the accuracy of spatial information determination. This allows for more accurate identification of the target speaker when performing spatial matching based on the first spatial information and the second spatial information output by the microphone device, thereby significantly improving the accuracy and robustness of the overall conference room speaker positioning method.

[0051] In the embodiments of this application, the microphone device described above can be a device for receiving sound waves and converting them into electrical signals. In a conference room environment, the microphone device typically includes one or more microphones (i.e., a microphone array) for capturing the speaking voices of participants. This application does not limit the number of microphone devices, and mainly uses one as an example for illustration. The sound-emitting target can be a user identified by the microphone device who is emitting a sound, and may include a speaker in the conference room. The second spatial information can be data output by the microphone device for each sound-emitting target, describing the spatial position of its sound source, and may include, but is not limited to, some or all of the distance, azimuth angle, pitch angle, and three-dimensional coordinates in a preset coordinate system of the sound source relative to the microphone device. This application does not limit this information.

[0052] For example, a microphone device may consist of an array of one or more microphones capable of capturing sound signals within a meeting room. By processing the sound signals, the source of the sound, i.e., the target of the sound, can be identified. As an example, the microphone device may utilize acoustic localization techniques, such as methods based on time difference of arrival or beamforming, to estimate the position coordinates of each target, thus determining the position coordinates of the target in three-dimensional space as secondary spatial information of the target.

[0053] In an optional embodiment of this application, the microphone device may further determine the second spatial information of at least one sound-emitting target in step 102 through the following implementation method, as follows: First, the microphone acquires sound signals from the area. These raw sound signals are typically complex analog or digital waveforms that require processing to convert them into parameters usable for localization. The microphone device can employ a digital signal processor (DSP) or microcontroller (MCU) to perform real-time sampling, analog-to-digital conversion, filtering, and noise reduction on the received sound signals to improve signal quality and remove environmental interference. Alternatively, the microphone device can utilize an acoustic front-end module, which integrates a high-precision analog-to-digital converter and pre-amplification circuitry, to acquire the sound signals with high fidelity and perform preliminary frequency domain analysis, such as obtaining the signal's spectral information through a Fast Fourier Transform (FFT).

[0054] Then, the microphone device analyzes the preprocessed sound signal to obtain a second distance and a second direction relative to the at least one sound-emitting target. For example, the microphone device can employ microphone array technology. By calculating the time difference or phase difference between different microphones receiving the same sound source signal (i.e., the same sound-emitting target), and combining this with the array geometry, it can estimate the azimuth and elevation angles of the sound-emitting target relative to the microphone device using algorithms such as generalized cross-correlation, multiple signal classification, or rotation-invariant parameter estimation. In this case, the second direction can include the azimuth and / or elevation angles corresponding to the sound-emitting target, and the second distance corresponding to the sound-emitting target can be estimated using the speed of sound and the time difference. Alternatively, the microphone device can also combine acoustic models and machine learning methods, using a trained model to analyze the characteristics of the sound signal (e.g., sound spectrum, volume changes, etc.) to predict the second distance and second direction corresponding to the sound-emitting target.

[0055] Furthermore, the microphone device determines the second spatial information based on the second distance and the second direction. For example, the microphone device can use the second distance and the second direction as the second spatial information of the sound-emitting target. That is, the first spatial information of each candidate target can include the second distance and the second direction corresponding to the sound-emitting target. In an optional embodiment, the microphone device can convert the obtained second distance and second direction (e.g., azimuth and pitch angles) into corresponding three-dimensional spatial coordinates based on a preset coordinate system (e.g., a Cartesian coordinate system or a spherical coordinate system with the microphone device as the origin) as the second spatial information. Alternatively, the microphone device can also convert the second distance and second direction into relative position information relative to the reference point based on the layout of the conference room or a preset reference point as the second spatial information; or it can directly output spherical coordinates as the second spatial information based on the obtained second distance and second direction.

[0056] For example, the microphone device can be a microphone array, which employs a Uniform Linear Array (ULA) or Uniform Circular Array (UCA) structure. A high-resolution parameter estimation algorithm (Estimation of Signal Parameters via Rotational Invariance Techniques, ESPRIT) is used to decoherently process the multi-channel acoustic signals, calculating the azimuth and elevation angles of the sound-emitting target relative to the microphone device. Then, the microphone device combines the geometric dimensions of the microphone array with the sound propagation speed, using a spherical interpolation algorithm to calculate the radial distance r of the sound-emitting target, outputting the acoustic positioning coordinates Paudio(xa,ya,za) of the sound-emitting target, i.e., the second spatial information of the sound-emitting target. In an optional embodiment, the microphone device can also acquire a first spatial information dataset {Pradar(t)} collected by the radar device within the most recent preset time period. i Then, the microphone device establishes a historical state transition model of acoustic positioning coordinates and first spatial information through the Kalman filter algorithm, and performs smooth filtering on the second spatial information of the sound-emitting target to correct the distance estimation error caused by indoor reverberation.

[0057] It should be explained that this application uses a microphone device to analyze sound signals, ensuring the accuracy of extracting key spatial parameters from the raw data and avoiding the impact of inaccurate input on the positioning results. A second distance and a second direction of the sound-emitting target are obtained. Furthermore, based on the second distance and the second direction, second spatial information is determined, constructing complete and standardized spatial data, enhancing the accuracy and consistency of the information, thereby supporting the reliability of the overall positioning system.

[0058] In an optional embodiment of this application, before performing step 102, in addition to the microphone device being able to independently collect the sound signal of the sound-emitting target, the radar device can also guide the microphone to collect sound information. Specifically, the radar device sends first spatial information of at least one candidate target to the microphone device; then, the microphone device collects signals based on the first spatial information of the at least one candidate target to obtain the sound signal.

[0059] For example, after detecting a candidate target in an area (such as a conference room), the radar device transmits the first spatial information of the candidate target to the microphone device. This can be achieved in various ways. For instance, the radar device can transmit the first spatial information to the microphone device in real time via a wired communication interface (e.g., Ethernet, USB) or a wireless communication interface (e.g., Wi-Fi, Bluetooth). Alternatively, the radar device can encapsulate the first spatial information into a standard data packet and send it to the microphone device via a preset communication protocol (e.g., UDP or TCP) to ensure that the microphone device can accurately parse and utilize the first spatial information.

[0060] Then, after receiving the first spatial information of the candidate target provided by the radar device, the microphone device can adjust its acquisition strategy based on the first spatial information. For example, the microphone device can utilize the beamforming capability of its microphone array to calculate the target direction based on the received first spatial information, and adjust the weights and phases of the microphone array to form a sound beam pointing in that target direction, thereby enhancing the sound of the emitting target and effectively suppressing noise from other directions. Alternatively, the microphone device can also dynamically adjust gain control or selectively activate microphone units in specific directions based on the first spatial information to optimize the acquisition of sound from the target area corresponding to the first spatial information and reduce interference from sound from non-target areas.

[0061] It should be explained that the radar device of this application can provide the microphone device with accurate first spatial information of candidate targets, so that the microphone device can focus on the potential sound-emitting target area based on the first spatial information when collecting sound signals, which significantly improves the efficiency and accuracy of sound signal collection and reduces the interference of background noise and irrelevant signals.

[0062] In the embodiments of this application, the spatial matching relationship in step 103 can be the degree of correlation between the first spatial information and the second spatial information in terms of spatial location. For example, it can be determined by calculating the distance, overlapping area, or similarity between the two spatial information. The target speaker can include the participant who is speaking, whose candidate target detected by the radar is associated with the sound source identified by the microphone through the spatial matching relationship.

[0063] For example, after acquiring the first spatial information of candidate targets provided by the radar device and the second spatial information of the sound-emitting target provided by the microphone device, it is necessary to correlate the two. As one example, the Euclidean distance between the second spatial information of each sound-emitting target and the first spatial information of each candidate target is calculated. If the Euclidean distance between the second spatial information of the sound-emitting target and the first spatial information of the candidate target is less than a preset threshold, then the two spatial information sets are considered to have a matching relationship, and the sound-emitting target is identified as the target speaker. As another example, the first and second spatial information are mapped to regions in two-dimensional or three-dimensional space, respectively, and it is determined whether these regions in two-dimensional or three-dimensional space overlap or are adjacent; if they overlap or are adjacent, then the two spatial information sets are considered to have a matching relationship, and the sound-emitting target is identified as the target speaker.

[0064] In an optional embodiment of this application, step 103 can also be obtained through the following implementation, as detailed below: First, for each vocal target, a first matching degree is determined between the second spatial information of the vocal target and the first spatial information of each candidate target; then, based on the first matching degree corresponding to each vocal target, a target speaker is determined. For example, vocal targets with a first matching degree reaching a first preset threshold are determined as the target speaker, or a preset number of vocal targets with a first matching degree reaching the first preset threshold are selected as the target speaker in descending order from all vocal targets. This application does not impose any limitations on this.

[0065] In optional embodiments of this application, the first matching degree is used to quantify the degree of spatial correlation between the second spatial information of the sound-emitting target and the first spatial information of the candidate target. The higher the first matching degree, the higher the corresponding degree of spatial correlation. One implementation is to calculate the Euclidean distance between the first and second spatial information. The smaller the Euclidean distance, the higher the first matching degree. That is, in actual operation, a mapping relationship between Euclidean distance and first matching degree can be constructed to satisfy the relationship that the smaller the Euclidean distance, the higher the first matching degree. Alternatively, other methods can be adopted, which are not limited in this application. Another implementation method is to calculate the probability that one of the spatial information in the first spatial information and the second spatial information belongs to the region represented by the other spatial information based on a probability model. The higher the probability, the higher the first matching degree. Similarly, a mapping relationship between probability and first matching degree can be constructed to satisfy the relationship that the higher the probability, the higher the first matching degree. Alternatively, other methods can be adopted, which are not limited in this application. Another implementation method is to represent the first matching degree by calculating the cosine value of the angle between the two spatial information in a specific coordinate system. The smaller the angle, the higher the first matching degree. A mapping relationship between the angle and the first matching degree can be constructed to satisfy the relationship that the smaller the angle, the higher the first matching degree. Alternatively, other methods can be adopted, which are not limited in this application.

[0066] In optional embodiments of this application, a first preset threshold is used to determine whether the first matching degree has reached a sufficiently high degree of correlation. The first preset threshold can be empirically set according to the needs of the actual application scenario and system performance. For example, it can be set as an upper limit of distance; when the Euclidean distance is less than this upper limit, a match is considered successful. Alternatively, it can be set as a lower limit of probability; when the matching probability is higher than this lower limit, a match is considered successful. Or, it can be set as an upper limit of the included angle; when the included angle is less than this upper limit, a match is considered successful. Optionally, determining the target speaker can also involve selecting one or more speaking targets from multiple speaking targets that are spatially highly matched with the candidate targets detected by the radar device, and identifying them as the actual speaker in the conference room.

[0067] It should be explained that this application effectively solves the ambiguity and inaccuracy problems that may exist in spatial matching in related technical solutions by calculating the first matching degree between each sound source and the candidate targets, and identifying the qualified object as the target speaker. This avoids the risk of misjudgment caused by the lack of a specific matching mechanism. By performing a comprehensive matching degree calculation between each sound source and all candidate targets, and filtering based on a preset threshold, it ensures that only sound sources with high spatial consistency are identified as speakers, thereby significantly improving the accuracy and reliability of speaker location in the conference room.

[0068] In optional embodiments of this application, after determining the target speaker based on the first matching degree of the vocal target, the determination of the first spatial information corresponding to the target speaker may involve the following situations, as detailed below: In an optional embodiment, if there are multiple instances where the first matching degree of the vocal target reaches the first preset threshold, then the first spatial information corresponding to the highest first matching degree of the vocal target can be determined as the first spatial information corresponding to the target speaker.

[0069] In the initial matching stage, a sound-emitting target may have a matching degree with multiple different candidate targets that all reach a preset reliability standard. This can be achieved by setting a counter or list in the processing device. When traversing all candidate targets and calculating their first matching degree with the current sound-emitting target, if the first matching degree exceeds a first preset threshold, the candidate target and its corresponding first spatial information are recorded, and the number of first matching degrees is counted.

[0070] Among the multiple first matching degrees that meet the above conditions, the highest first matching degree is obtained by comparison and selection. Then, the first spatial information of the candidate target associated with this highest first matching degree is acquired. This can be achieved by executing a maximum value search algorithm in the processing device. For example, in a list recording all first matching degrees that meet the conditions and their corresponding first spatial information, the maximum matching degree value is iteratively searched, and its associated first spatial information is extracted. The first spatial information corresponding to the highest first matching degree is determined as the first spatial information corresponding to the vocal target. This can be done by storing the first spatial information in a data structure or assigning it to a variable representing the first spatial information of the target speaker.

[0071] It should be explained that this application effectively solves the problem of how to accurately select the first spatial information corresponding to the target speaker when there are multiple targets with a first matching degree reaching a first preset threshold. By prioritizing the selection of the first spatial information corresponding to the highest first matching degree, this solution can ensure that the most reliable and accurate radar positioning data is selected for the target speaker in multi-match scenarios, thereby significantly improving the accuracy and reliability of speaker positioning in the conference room and avoiding positioning ambiguity or errors caused by multiple matching.

[0072] In an optional embodiment, if there are multiple highest first matching degrees, the first spatial information corresponding to each highest first matching degree is used as the first spatial information corresponding to a target speaker, resulting in multiple first spatial information corresponding to multiple target speakers, wherein the multiple target speakers correspond one-to-one with the multiple first spatial information.

[0073] In this scenario, a single sound source may have a first matching degree calculated with multiple different candidate targets, all of which reach a first preset threshold and are equal. This typically occurs when multiple physical targets (e.g., multiple users) are spatially very close, or when the microphone device has some ambiguity in locating the sound source, causing it to achieve an equal degree of optimal matching with multiple candidate targets detected by the radar device. For example, when two users speak simultaneously and closely adjacent to each other, the microphone device may identify the two sound sources as a single sound source, while the radar device can distinguish between the two users and determine first spatial information for each of them. Furthermore, both of these first spatial information pieces have the highest matching degree with the same sound source detected by the microphone.

[0074] For the scenario of multiple highest first-match degrees, each first spatial information with the highest first-match degree is independently considered as the first spatial information corresponding to a potential target speaker, and each candidate target with the highest match degree with the vocal target is identified as an independent target speaker. For example, if two candidate targets both have the highest match degree with a vocal target, then the first spatial information of each candidate target will be considered as the first spatial information of a target speaker.

[0075] It should be explained that, when multiple candidate targets share the same highest first matching degree with the sound-emitting target, this application avoids mistakenly grouping multiple physically independent speakers into one, or confusing the spatial information of a speaker with multiple unrelated radar targets. By independently determining a target speaker for each highest first matching degree corresponding to the first spatial information, and establishing a one-to-one correspondence between multiple target speakers and multiple first spatial information, the application can fully utilize the fine spatial information provided by the radar device, effectively distinguish and identify multiple closely adjacent speakers, significantly improve the accuracy and reliability of speaker positioning in complex meeting scenarios, and solve the problem of reduced positioning accuracy caused by microphone positioning ambiguity or the proximity of multiple targets.

[0076] In optional embodiments of this application, the spatial information of the target speaker in step 104 can be obtained by fusing the first spatial information and the second spatial information corresponding to the target speaker; or it can be determined by selecting one of the first spatial information and the second spatial information based on the specific application scenario; or it can be determined by iterative calibration through the historical spatial data corresponding to the target speaker and the real-time first spatial information and the second spatial information. This application does not limit the method of determining the spatial information of the target speaker.

[0077] For example, after determining the target speaker, more accurate final spatial information is generated based on the first and second spatial information of the target speaker. As one example, a simple arithmetic average is performed on the first and second spatial information corresponding to the target speaker to obtain the target speaker's spatial information. As another example, one of the spatial information is selected as the target speaker's spatial information according to a preset priority rule. This application does not limit the method for determining the target speaker's spatial information.

[0078] In an optional embodiment of this application, step 104 can also be obtained by the following implementation: specifically, based on a preset fusion algorithm, the first spatial information and the second spatial information corresponding to the target speaker are fused to obtain the spatial information of the target speaker; wherein, the preset fusion algorithm includes, but is not limited to, one or more of the following: weighted fusion algorithm, tightly coupled fusion algorithm, Bayesian fusion algorithm, which are not limited in this application.

[0079] In this context, a pre-defined fusion algorithm can be understood as a data fusion method pre-determined and configured based on the characteristics of the devices (such as radar devices or microphone devices), environmental conditions, and application requirements before system design or operation. Its purpose is to effectively integrate spatial information from different devices. Pre-defined fusion algorithms can be implemented through pre-programming in the processing device; for example, by embedding specific fusion algorithm code in firmware or loading it as a configurable module. Alternatively, they can be trained and deployed using a machine learning model. This model learns during the training phase how to optimally fuse data from different devices and executes as a pre-defined algorithm during runtime.

[0080] It should be explained that fusing the first and second spatial information corresponding to the target speaker aims to combine the first spatial information from the radar device and the second spatial information from the microphone device to overcome the inherent limitations of a single positioning device. For example, the radar device may be insensitive to stationary targets, while the microphone device may be affected by ambient noise. Through fusion, the complementarity of the two types of information can be utilized to improve the accuracy and reliability of positioning.

[0081] In embodiments of this application, fusion can be performed at the data layer, i.e., by directly performing mathematical operations on the original or preprocessed spatial coordinate data, such as weighted averaging; or it can be performed at the feature layer, i.e., extracting features from the two types of spatial information (e.g., position, velocity, confidence level, etc.), fusing these features, and determining the final spatial information through the fused features. Spatial information can include position coordinates, representing the location of the target speaker within the region; or it can be a probability distribution containing position coordinates and an uncertain region (such as an ellipsoid), reflecting the confidence level of the location.

[0082] Weighted fusion algorithms can be understood as a method of combining different data sources by assigning weights to them. These weights typically reflect the reliability, accuracy, or importance of the data sources. The system dynamically adjusts the contribution of different spatial information based on sensor performance, environmental conditions, or target status, thereby optimizing the fusion result. For example, when the echo signal quality from the radar device is good, the weight of the first spatial information is increased; when the microphone positioning accuracy is high, the weight of the second spatial information is increased. Weighted fusion algorithms can be linear weighted averages or other methods.

[0083] Tightly coupled fusion algorithms can be understood as fusing data at a lower level (e.g., raw measurement data or feature data), fully considering the interdependence and correlation between different sensors. They can more deeply explore the intrinsic connections between sensor data and handle the interaction effects between sensors, thus providing more consistent and accurate positioning results in complex environments. This is particularly suitable for scenarios where sensor data exhibit strong correlation or complementarity. Tightly coupled fusion algorithms can be implemented using state estimation algorithms such as extended Kalman filters or unscented Kalman filters, directly using the raw measurements from radar and microphones as observation inputs to jointly estimate the target state. Alternatively, methods such as Factor Graph Optimization (FGO) can be used to construct a unified optimization problem from the measurement and motion models of different sensors for solution.

[0084] Bayesian fusion algorithms, based on Bayes' theorem, update the posterior probability distribution of a target state by combining prior knowledge and sensor observation data. This handles uncertainty in a probabilistic manner and provides a probabilistic estimate of the target's location. The algorithm effectively handles uncertainties and noise in sensor measurements, providing robust localization results and naturally integrating prior information, improving decision-making capabilities in situations with incomplete or ambiguous information. Bayesian fusion algorithms can be implemented using particle filtering, where a set of particles represents the probability distribution of the target state, each carrying a weight, and updating the particle weights and positions based on sensor observations. Alternatively, they can be implemented using probabilistic graphical models such as Gaussian Mixture Models (GMMs) or Bayesian Networks, modeling measurements from different sensors and fusing them using Bayesian inference.

[0085] In embodiments of this application, the preset fusion algorithm may further include a Kalman filter fusion algorithm, a least squares fitting fusion algorithm, etc. For example, the Kalman filter fusion algorithm can retrieve the historical location dataset of the radar device within the most recent preset time period, establish a state transition model between acoustic positioning coordinates and radar historical coordinates, and perform smoothing filtering on the fusion process to correct distance estimation errors caused by factors such as indoor reverberation. The least squares fitting fusion algorithm can solve for the optimal fused spatial information by minimizing the sum of squared residuals between the first spatial information and the second spatial information.

[0086] It should be explained that this application integrates the first and second spatial information corresponding to the target speaker by introducing a pre-defined fusion algorithm. This effectively solves the problems of insufficient accuracy and poor robustness in positioning results caused by the lack of an effective fusion mechanism in related positioning methods. Specifically, this scheme can fully utilize the respective advantages of radar and microphone devices, such as the radar's sensitivity to moving targets and the microphone's direct perception capability of sound-emitting targets. Through diverse fusion strategies such as weighted fusion algorithms, tightly coupled fusion algorithms, or Bayesian fusion algorithms, the system can dynamically select the most suitable fusion method according to the actual scene and sensor characteristics. This not only ensures the standardization and controllability of the fusion process and avoids errors that may be introduced by arbitrary data processing, but also significantly improves the accuracy and reliability of the target speaker's spatial information, enabling the positioning results to better adapt to different environmental changes and sensor characteristics.

[0087] The following explanation will use a weighted fusion algorithm as an example to illustrate how the first spatial information and the second spatial information corresponding to the target speaker are fused to obtain the spatial information. Specifically: the first weight corresponding to the first spatial information and the second weight corresponding to the second spatial information are obtained; based on the first weight and the second weight, the first spatial information and the second spatial information corresponding to the target speaker are weighted and fused to obtain the spatial information.

[0088] In this context, the first weight and the second weight represent the importance or confidence level of the first spatial information and the second spatial information in the fusion process, respectively. Obtaining the first and second weights aims to quantify the relative contributions of different sensor data sources (i.e., the radar device and microphone device of this application). As an example, weights can be preset based on the characteristics of the sensors themselves and environmental conditions. For instance, in an environment where radar signals are less interfered with, the first spatial information can be given a higher weight; in an environment where microphone reception is clear, the second spatial information can be given a higher weight. As another example, weights can be dynamically adjusted based on real-time data quality. For instance, the first weight can be determined by evaluating the signal-to-noise ratio or stability of the first spatial information output by the radar device, while the second weight can be determined by evaluating the clarity or positioning accuracy of the second spatial information output by the microphone device.

[0089] One possible implementation is to use a weighted average method, where the final spatial information is the weighted average of the first and second spatial information. Another implementation is to use a linear combination method, where the weights can be normalized or denormalized according to the specific application scenario to adapt to different fusion models. For example, the first spatial information is denoted as P. radar The second spatial information is denoted as P. audio The first weight is denoted as w. r The second weight is denoted as w. aThe spatial information of the target speaker is denoted as P. final The fusion formula is: P final =w r ×P radar +w a ×P audio .

[0090] It should be explained that this application limits the acquisition and application of weights in the weighted fusion algorithm, effectively solving the problem of lack of flexibility and accuracy in the fusion process of related schemes. By acquiring the first weight corresponding to the first spatial information and the second weight corresponding to the second spatial information, and performing weighted fusion of the two types of spatial information based on the first weight and the second weight, the system can dynamically adjust the fusion strategy according to the reliability of different sensor data or environmental conditions. This not only improves the accuracy of determining the spatial information of the target speaker, but also enhances the adaptability and robustness of the system in complex and ever-changing conference environments.

[0091] In optional embodiments of this application, the relationship between the first weight and the second weight may satisfy, but is not limited to, one of the following relationships: If the target speaker is in a dynamic and vocal state, the first weight is greater than the second weight; if the target speaker is in a static and vocal state, the first weight is less than the second weight; if the matching degree between the second spatial information of two vocal targets reaches a second preset threshold, the first weight is greater than the second weight; if the matching degree between the second spatial information of two vocal targets does not reach the second preset threshold, the first weight is less than or equal to the second weight.

[0092] The statement that the target speaker is in a dynamic and vocal state indicates that the speaker's spatial position or posture is changing significantly while emitting sound. Determining whether the target speaker is in a dynamic state can be done by analyzing the changes in a first spatial information over a continuous time period (e.g., calculating their movement speed (such as the radial velocity of a candidate target output by a radar device), velocity variance, or acceleration). If the changes in the first spatial information exceed a third preset threshold, the target speaker is determined to be in a dynamic state. In this state, because the radar device is more sensitive and accurate in detecting object motion, the first weight is assigned greater than the second weight to prioritize radar data, thereby more accurately capturing the speaker's real-time position.

[0093] The speaker being in a static, vocal state indicates that while emitting sound, their spatial position or posture remains relatively stable. Whether the speaker is static can be determined by analyzing the changes in their first spatial information over a continuous time period (e.g., calculating their velocity (such as the radial velocity of a candidate target output by a radar device), velocity variance, or acceleration). Similarly, if the changes in the first spatial information do not exceed a third preset threshold, the speaker is determined not to be static. In this state, the microphone device's localization of the sound source is generally more stable and accurate; therefore, a first weight is assigned less than a second weight to prioritize microphone data, thereby improving localization stability.

[0094] For example, taking the change in the first spatial information as the velocity variance as an example, the radar device outputs the velocity variance of the candidate target. If the velocity variance > the fifth value, it is determined that the target speaker is in a dynamic and vocal state; a first weight w is assigned. r =First preset weight, second weight w a =1-w r At this point, the first weight is greater than the second weight. If the velocity variance is less than the sixth value, the target speaker is determined to be in a static and vocal state; the second weight w is assigned. a =Second preset weight, first weight w r =1-w a At this point, the first weight is less than the second weight; optionally, the first preset weight is less than the second preset weight.

[0095] In a conference environment, there may be two or more sound sources emitting sounds that are spatially very close, making it difficult for microphones to effectively distinguish them. Therefore, it's helpful to determine the degree of proximity (e.g., matching degree) based on second spatial information (such as sound source direction and distance). For example, if the angle between the sound source directions of two sound sources is less than a certain threshold, or the distance between the sound sources is less than a certain threshold, the matching degree is considered high. If the matching degree is high, microphone data is prone to confusion, while radar devices are better able to identify and distinguish different physical entities. Therefore, a first weight is assigned greater than a second weight to enhance the radar data's ability to distinguish different speakers. For instance, if two participants are discussing at close range and the second spatial information of the two sound sources has a high matching degree, then the first weight is greater than the second weight, using the high-precision positioning of radar data to compensate for errors caused by acoustic interference.

[0096] Similarly, if the matching degree is low, there is sufficient spatial distance between the sound-emitting targets, or there is only one sound-emitting target, allowing the microphone device to clearly identify and locate each sound source. The microphone data has high reliability, therefore the first weight is assigned less than or equal to the second weight to fully utilize the advantages of microphone data in sound source localization. For example, in a single-person presentation, or when multiple speakers are seated far apart, the matching degree of the second spatial information between any two sound-emitting targets is low. In this case, the first weight is less than or equal to the second weight to adapt to the localization requirements of low-interference scenarios.

[0097] In some optional embodiments, the confidence level of the target speaker's spatial information can also be calculated, i.e., C=w r ×C r +w a ×C a C r For radar data confidence, C a For acoustic data confidence, when C > the sixth value, output the spatial information of the target speaker; otherwise, trigger multi-frame data fusion, that is, continue to perform the above steps based on the real-time acquired data until the confidence of the target speaker's spatial information meets the threshold (i.e., the sixth value) requirement.

[0098] It should be explained that this application effectively solves the problem of the lack of specific rules for weight setting in related weighted fusion algorithms, avoiding the decrease in positioning accuracy and reliability caused by improper weight allocation when the target speaker is in a dynamic or static state or when there are multiple vocal targets. This solution can intelligently adjust the fusion weights of radar and microphone data according to the speaker's actual movement state and the complexity of the acoustic environment, thereby significantly improving the accuracy and robustness of speaker positioning in conference rooms, enabling the system to provide stable and reliable positioning results in various complex conference scenarios.

[0099] In an optional embodiment of this application, a camera-assisted method for determining the target speaker can also be introduced. Accordingly, step 103 can also be obtained through the following implementation: specifically, acquiring the third spatial information of at least one second candidate target output by the camera; then, determining the target speaker based on the first spatial information, the second spatial information, and the third spatial information. Accordingly, step 104 corresponds to determining the spatial information of the target speaker based on the first spatial information, the second spatial information, and the third spatial information corresponding to the target speaker.

[0100] In the embodiments of this application, a camera can be used as a visual sensor to capture image or video information within the conference room. This image or video information is then processed to identify individuals in the conference room as second candidate targets. The third spatial information can be the spatial location data of the second candidate target, such as its three-dimensional coordinates in the camera coordinate system, two-dimensional image coordinates (combined with depth information), or its direction and distance relative to the camera. The camera can be a depth camera, directly acquiring a depth map of the scene. Combined with face or human detection results in the image, the three-dimensional spatial coordinates of the second candidate target can be calculated. Alternatively, a standard RGB camera can be used, employing image processing algorithms (e.g., target detection, pose estimation) to identify individuals in the image. Combined with pre-calibrated camera parameters and a conference room model, the three-dimensional spatial position of the second candidate target can be estimated using techniques such as triangulation or monocular ranging.

[0101] By introducing third spatial information, the first and second spatial information can be cross-validated and supplemented, thereby improving the accuracy and robustness of speaker identification. As an example, a multimodal data fusion model can be established, such as one based on a probabilistic graphical model (e.g., a Bayesian network) or a machine learning model. This model uses the first, second, and third spatial information as input features, outputs the probability of each candidate target as the target speaker, and selects the candidate with the highest probability as the target speaker. As another example, matching rules can be set. For instance, preliminary matching is performed based on the first and second spatial information to obtain at least one candidate speaker. For each candidate speaker, the matching degree between the third spatial information and the first and second spatial information is determined. Based on the matching degree, the target speaker is determined from the at least one candidate speaker, thus eliminating misjudgments caused by single sensor errors or interference. The process of determining the target speaker based on the matching degree can be referred to the foregoing embodiments and will not be repeated here.

[0102] After identifying the target speaker, all available spatial information (e.g., first, second, and third spatial information) corresponding to that speaker is fused to obtain a more accurate spatial location. The fusion of multi-source information can effectively reduce single-sensor errors and improve positioning accuracy and stability. A weighted average fusion algorithm can be used. Based on the accuracy and reliability of different sensors and current environmental conditions (e.g., dynamic / static, interference), different weights are assigned to the first, second, and third spatial information, and then a weighted average is calculated to obtain the final spatial information of the target speaker. Alternatively, more complex fusion algorithms, such as Kalman filtering or particle filtering, can be used. The first, second, and third spatial information are used as observations, combined with the speaker's motion model, to estimate the speaker's spatial location in real time, thereby obtaining a smoother and more accurate positioning result.

[0103] For example, a weighted fusion decision model is constructed, which dynamically allocates fusion weights for the data from each device based on the confidence levels of the first spatial information output by the radar device, the third spatial information output by the camera, and the VAD confidence levels of the second spatial information output by the microphone device. For example, the radar position data weight w r The first preset value, camera visual feature weight w c The second preset value, microphone acoustic feature weight w a The third preset value is used, for example, the second preset value is equal to the third preset value, and the second preset value is greater than the first preset value; then, based on the fusion weights and the first spatial information P output by the radar device... radar Third-space information P output by the camera camera The second spatial information P output by the microphone device audio Through formula P final =w r ×P radar +w c ×P camera +w a ×P audio The spatial information of the target speaker, i.e., the location coordinates P, is calculated. final Optionally, the overall confidence level C of the fusion result can be calculated. final =w r ×C r +w c ×C camera +w a ×C vad C r For the confidence level of the first spatial information, C camera For the confidence level of third-space information, C vad The confidence level of VAD for the second spatial information; if the overall confidence level Cfinal If the preset fusion confidence threshold is reached, the current location coordinates are output as the final speaker location coordinates, i.e., the spatial information of the target speaker; if not, multi-frame data iterative fusion is triggered until the overall confidence meets the preset fusion confidence threshold.

[0104] It should be explained that this application significantly improves the accuracy of speaker location in conference rooms by introducing third spatial information output by the camera and fusing it with first spatial information output by the radar device and second spatial information output by the microphone device in a multimodal manner. In related solutions, relying solely on radar and microphone data can easily lead to misjudgment or inaccurate location when there are multiple potential speakers, background noise interference, or dense crowds in the conference room. This application provides an additional, independent verification dimension for speaker identification by adding visual information from the camera. For example, when the microphone may be inaccurately positioned due to reverberation, or the radar cannot distinguish between stationary but vocal individuals, the camera can provide clear information about the person's position and face, effectively eliminating interference and more accurately locking onto the true speaker. This allows the system to more robustly and reliably determine the target speaker and their precise spatial location in complex and changing conference environments, thus effectively solving the technical problems of insufficient positioning accuracy and susceptibility to misjudgment in traditional solutions.

[0105] In optional embodiments of this application, at least one second candidate target output by the camera can also be determined in the following manner, as follows: The camera acquires images based on spatial indication information, which includes the first spatial information and / or the second spatial information; the camera analyzes the images to obtain facial motion features of at least one third candidate target; the camera determines at least one second candidate target from the at least one third candidate target based on the matching degree between the facial motion features and features in a preset vocal facial motion feature library.

[0106] The spatial indication information may include positioning data used to guide the camera to focus on a specific area or target. This spatial indication information may include first spatial information output by a radar device and / or second spatial information output by a microphone device. If the spatial information received by the camera includes the first spatial information received from the radar device and the second spatial information received from the microphone device, the camera can fuse these two types of spatial information (the specific fusion method is not limited) to allow the camera to acquire image data based on the fused spatial information. The first spatial information and / or the second spatial information work together to provide the camera with a preliminary, targeted observation range, thus avoiding blind scanning of the entire conference room. Image acquisition allows the camera to adjust its shooting parameters (e.g., focal length, field of view, shooting angle, etc.) according to the received spatial indication information to obtain high-quality visual data of a specific area or target. Image analysis can be understood as processing and parsing the image data acquired by the camera to extract meaningful information. This analysis process can utilize various image processing algorithms or artificial intelligence vision algorithms, such as target detection, face recognition, and key point tracking.

[0107] The third candidate target can be a potential target that may be the speaker, identified from the acquired images through image analysis. Facial motion features can be dynamic changes in the face related to human vocalization, including lip movements, subtle changes in facial muscles, eye movements, and even auxiliary head or hand movements. The pre-defined vocalization facial motion feature library can be a database or model that stores facial motion patterns of known vocalization behaviors. The vocalization facial motion feature library can be built based on a large amount of training data, for example, by learning and storing typical facial motion patterns of different individuals during vocalization through machine learning or deep learning models. The matching degree between facial motion features and features in the pre-defined vocalization facial motion feature library can be calculated using various algorithms, such as correlation coefficients, distance metrics, or probability values ​​output by classifiers, etc., which are not limited in this application.

[0108] It should be explained that this application effectively solves the problem of non-speaker interference targets that may be introduced when the camera acquires images. By utilizing the first and / or second spatial information provided by the radar and microphone devices as spatial indications, the camera can selectively focus on the potential speaker area for image acquisition, avoiding scanning of irrelevant areas, thereby improving the efficiency and quality of image acquisition. Simultaneously, by performing facial motion feature analysis on the third candidate target in the acquired image and matching it with a preset vocal facial motion feature database, the system can accurately identify second candidate targets with vocal behavior characteristics. This allows the system to effectively filter out static interference, non-speaking targets, or individuals who only overlap spatially but do not actually vocalize, significantly improving the accuracy and reliability of target speaker identification.

[0109] In an optional embodiment, the camera determines the at least one second candidate target from the at least one third candidate target based on the matching degree between the facial motion features and features in a preset vocal facial motion feature library. This determination can also be achieved through the following implementation methods: Based on the matching degree between the facial motion features and the features in the preset vocal facial motion feature library, the camera determines at least one fourth candidate target from the at least one third candidate target; based on the pose features of each fourth candidate target, the camera determines whether each fourth candidate target is an interfering target; and the camera determines the fourth candidate target that is not an interfering target as the at least one second candidate target.

[0110] For example, the camera determines at least one fourth candidate target from the at least one third candidate target based on the matching degree between facial motion features and a preset vocal facial motion feature library. For example, it determines a portion (e.g., a preset number selected in descending order of matching degree) or all of the third candidate targets whose matching degree reaches a fourth preset threshold as the fourth candidate target.

[0111] In embodiments of this application, the posture features may include the pose, actions, and relative positional relationships of various body parts of the fourth candidate target in space. For example, after the camera acquires an image, human pose estimation techniques, such as deep learning models like OpenPose and AlphaPose, are used to detect and track key points of the human skeleton in the image, thereby obtaining information such as head pose, body orientation, and gestures. Then, based on the head pose, body orientation, and gestures of the fourth candidate target, it is determined whether the fourth candidate target is in a normal speaking posture. For example, it analyzes whether the head is facing forward, the body is facing the center of the meeting, and whether there are large gestures. For instance, if the head is facing forward, the body is facing the center of the meeting, and there are large gestures, then it is determined that the fourth candidate target is in a normal speaking posture. Conversely, if a person is not facing forward, their body is facing the center of the meeting, and they are not making large gestures, then they are determined to be not in a normal speaking posture. Then, it is determined that the fourth candidate target in a normal speaking posture is not a distracting target, and those not in a normal speaking posture are considered distracting targets. A distracting target can be understood as someone whose facial movement features are similar to those in the vocal facial movement feature database, but whose posture features indicate that they are not a person giving a formal speech. For example, if a person's face shows signs of vocalization, but their body is turned to the side, their head is lowered, or they are whispering with someone next to them, they may be identified as a distracting target. Furthermore, the fourth candidate target that is not a distracting target is determined as at least one second candidate target to eliminate interference caused by non-speaking behaviors.

[0112] It should be explained that after initially identifying a third candidate target who may be speaking, this application further introduces posture feature analysis. This effectively identifies and eliminates interfering targets that, although showing signs of speaking on their faces, are not actually speakers. This avoids misjudgments that may result from relying solely on facial movement features, such as misidentifying private conversations or non-speaking body movements as speaking behavior. By combining facial movement features and posture features for dual verification, the accuracy and reliability of speaker identification are significantly improved.

[0113] In other optional embodiments, facial motion features may include inter-frame motion features of the lips; the vocal facial motion feature library may include lip action feature templates; alternatively, based on the acquired image data, the camera can extract the region of interest for the facial region of each third candidate target, and use a convolutional neural network combined with optical flow to analyze the inter-frame motion features of the lips in the facial region (e.g., lip opening and closing, peristalsis) to construct a lip action feature template library; the inter-frame motion features of the lips are matched with the preset lip action feature template library to output the lip action matching degree M. lip (For example, matching degree M) lip If the lip movement matching degree reaches the preset lip matching threshold, the corresponding third candidate target is determined as the suspected main speaker, i.e., the fourth candidate target. Then, the camera performs a second screening based on the suspected main speaker's facial orientation (e.g., whether it is facing the sound source), facial expression features (e.g., facial muscle changes when speaking), and posture features (e.g., whether it is the posture when talking to someone privately) to determine the candidate main speaker, i.e., the second candidate target. The camera can also output the third spatial information (i.e., position coordinates P) of the second candidate target. camera ) and the confidence level C corresponding to the third spatial information camera In other words, lip movements are re-detected to avoid the radar detecting "lip features" as isolated incidents.

[0114] Furthermore, in the embodiments of this application, the matching degree between any two of the first, second, and third spatial information corresponding to the target speaker reaches a first preset threshold. The matching degree can be calculated based on the Euclidean distance, Mahalanobis distance, angular difference, or the proportion of overlapping areas between the two spatial information pieces, and this application does not limit this calculation. The first preset threshold can be used to determine whether the matching degree between two spatial information pieces reaches an acceptable level of consistency, serving as a criterion for spatial information matching. The matching degree between the first spatial information identified by the radar device and the second spatial information located by the microphone device, the matching degree between the second spatial information located by the microphone device and the third spatial information of the second candidate target identified by the camera, and the matching degree between the first spatial information identified by the radar device and the third spatial information of the second candidate target identified by the camera are calculated. Only when the matching degree between any two spatial information pieces reaches the preset first preset threshold is it considered that the aforementioned first, second, and third spatial information all point to the same target speaker.

[0115] This application ensures that the spatial information acquired by the radar, microphone, and camera is highly consistent in space by determining that the matching degree between any two of the first, second, and third spatial information corresponding to the target speaker reaches a first preset threshold. This effectively avoids the risks caused by data deviation or mismatch from a single sensor, significantly improving the accuracy and robustness of target speaker identification. For example, when the radar detects a target, the microphone detects a sound source, and the camera detects facial movement, the same speaker will only be identified when these three are highly consistent in space. This effectively reduces false positives and false negatives, making the speaker location in the conference room more accurate and reliable.

[0116] In some alternative embodiments, a clock synchronization protocol can be used to timestamp (e.g., nanosecond level) the data frames of the images captured by the camera, the audio signals from the microphone, and the echo signals from the radar device, based on the clock source of the radar device, thereby achieving time synchronization of the radar device, camera, and microphone device. Then, spatial coordinate system calibration is performed, for example, by using Zhang's calibration method to complete the intrinsic parameter calibration of the camera, and combining hand-eye calibration to obtain the extrinsic parameter matrices (e.g., rotation matrix, translation vector) of the camera, radar device, and microphone device. The global three-dimensional coordinate system of the radar device, the imaging coordinate system of the camera, and the acoustic coordinate system of the microphone device are mapped to the system's global coordinate system (e.g., with the geometric center of the microphone array as the origin), thereby achieving spatial coordinate consistency of the radar device, camera, and microphone device, that is, achieving spatial dimensional consistency of the spatial information corresponding to the radar device, camera, and microphone device.

[0117] Therefore, further, in the optional embodiments of this application, after determining the target speaker based on the first spatial information, the second spatial information, and the third spatial information, the determination of the first spatial information and the third spatial information corresponding to the target speaker may involve the following situations, specifically: For each vocal target, if the number of vocal targets whose first matching degree reaches the first preset threshold is multiple, and the number of vocal targets whose second matching degree reaches the first preset threshold is 1, then the first spatial information corresponding to the highest first matching degree of the vocal target is determined as the first spatial information corresponding to the target speaker; if the number of vocal targets whose first matching degree reaches the first preset threshold is 1, and the number of vocal targets whose second matching degree reaches the first preset threshold is multiple, then the third spatial information corresponding to the highest second matching degree of the vocal target is determined as the third spatial information corresponding to the target speaker; wherein, the first matching degree represents the matching degree between the first spatial information and the second spatial information, and the second matching degree represents the matching degree between the second spatial information and the third spatial information.

[0118] In this embodiment, when multiple candidate targets have matching degrees with the emitting target that all reach a preset threshold, making it impossible to uniquely determine the spatial information of the speaker, a conditional branching logic is introduced to solve this problem. Based on acquiring the first spatial information output by the radar device, the second spatial information output by the microphone device, and the third spatial information output by the camera, this scheme first evaluates, for each emitting target, the number of times its first and second matching degrees with different candidate targets reach a first preset threshold.

[0119] Specifically, when the second spatial information of the same vocal target matches the first spatial information of multiple candidate targets, but only matches the third spatial information of one second candidate target, the system will prioritize selecting the first spatial information with the highest first matching degree and identify it as the first spatial information corresponding to the target speaker. When the second spatial information of the same vocal target matches the first spatial information of only one candidate target, but matches the third spatial information of multiple second candidate targets, the system will select the third spatial information with the highest second matching degree and identify it as the third spatial information corresponding to the target speaker.

[0120] This application filters the first spatial information corresponding to the highest first matching degree or the third spatial information corresponding to the highest second matching degree by different combinations of the number of first matching degrees and second matching degrees. At the same time, it adapts to the multi-target processing logic when the highest matching degrees are equal. In complex scenarios with multiple matching in one dimension and single matching in another dimension, it can accurately lock the spatial information most closely related to the speaking target, avoid the positioning confusion caused by multiple sets of valid spatial information, and further improve the accuracy and relevance of the target speaker's corresponding first and third spatial information. This improves the technical problem in related technologies where there is multiple valid information in a single dimension when matching multi-source data, resulting in fuzzy spatial information selection and unclear association logic, which leads to a decrease in the accuracy of subsequent multi-source fusion positioning.

[0121] In an optional embodiment, if the number of times the first matching degree of the vocal target reaches the first preset threshold is multiple, and the number of times the second matching degree of the vocal target reaches the first preset threshold is multiple, then a third matching degree is determined between the first spatial information corresponding to each first matching degree and the third spatial information corresponding to each second matching degree; the first spatial information and the third spatial information corresponding to the third matching degree that reach the first preset threshold are taken as the first spatial information and the third spatial information corresponding to the same target speaker respectively; the first spatial information corresponding to the third matching degree that does not reach the first preset threshold is taken as the first spatial information corresponding to one target speaker, and the third spatial information corresponding to the third matching degree that does not reach the first preset threshold is taken as the third spatial information corresponding to another target speaker.

[0122] Specifically, when a microphone device detects a sound-emitting target, the second spatial information of the sound-emitting target shows a high degree of matching with the first spatial information of candidate targets detected by multiple radar devices (the first matching degree reaches a first preset threshold). Simultaneously, the second spatial information of the sound-emitting target also shows a high degree of matching with the third spatial information of second candidate targets detected by multiple cameras (the second matching degree reaches a first preset threshold). The system calculates the spatial matching degree, i.e., the third matching degree, for each first spatial information whose first matching degree with the sound-emitting target reaches the first preset threshold, and for each third spatial information whose second matching degree with the sound-emitting target reaches the first preset threshold. The third matching degree can be calculated based on the geometric distance, angular difference, or other spatial feature similarity between the first and third spatial information, and this application does not limit this calculation. For example, the Euclidean distance between the first and third spatial information can be calculated; the smaller the Euclidean distance, the higher the matching degree. Alternatively, the angular difference between them in a specific coordinate system can be calculated; the smaller the angular difference, the higher the matching degree. Another approach is to use data fusion algorithms such as Kalman filtering or particle filtering to evaluate the correlation and consistency between the first and third spatial information over time, thereby obtaining the third matching degree.

[0123] Furthermore, if the third matching degree between a certain first spatial information and a certain third spatial information also reaches a preset first threshold, it is considered that the first spatial information and the third spatial information are highly likely to originate from the same actual speaker. Therefore, highly matched first spatial information and third spatial information are bound as the first spatial information and third spatial information of the target speaker. If a certain first spatial information fails to form a third matching degree with any third spatial information that reaches the first preset threshold, or if a certain third spatial information fails to form a third matching degree with any first spatial information that reaches the first preset threshold, it indicates that these unmatched spatial information may correspond to different, independent speakers, or they are false alarms or interference. The system will treat them as independent speaker candidates and process them separately. For example, unmatched first spatial information is treated as radar positioning information of a potential speaker, and unmatched third spatial information is treated as camera positioning information of a potential speaker, retaining all possible speaker information, avoiding information loss, and providing more comprehensive data for subsequent processing.

[0124] It should be explained that this application effectively solves the problem of accurately distinguishing and locating multiple speakers when multiple matches exist in multi-sensor data. By introducing a third matching degree between the first spatial information output by the radar device and the third spatial information output by the camera, the matching results between the second spatial information output by the microphone device and the spatial information output by the radar device and the camera are further verified and refined. This enables the system to more accurately assign spatial information from different sensors to the correct speaker in complex meeting environments, even when multiple potential speakers are close together or there is interference. This significantly improves the accuracy and robustness of speaker location, especially in scenarios where multiple speakers are speaking simultaneously, effectively avoiding misjudgment and confusion.

[0125] In optional embodiments of this application, the second spatial information can be acoustic positioning coordinates; the second candidate target can be understood as a candidate main speaker; the third spatial information can be understood as the position coordinates of the candidate main speaker; and matching the second spatial information with the third spatial information can be matching the acoustic positioning coordinates with the position coordinates of the candidate main speaker. To further ensure that the second candidate target output by the camera is a valid sound source, the microphone device can also perform acoustic verification on the second candidate target selected by the camera based on the acquired sound signal, and provide a precise location reference for the sound source. For example, by acquiring sound signals through the microphone device, a high-resolution algorithm based on subspace decomposition (Multiple Signal Classification, MUSIC) is used to perform high-resolution direction-of-arrival estimation on the sound signals, calculating the azimuth angle θ of the sound source corresponding to the sound target in a preset coordinate system (such as the global coordinate system). aunio Pitch angle φ aunio and distance r aunio The second spatial information P is obtained. aunio Then, speech activity detection is performed on the acoustic beam signal corresponding to the region of the second candidate target output by the camera, such as extracting short-time energy and zero-crossing rate features, and the VAD confidence C corresponding to the second candidate target is output. van If the VAD confidence level corresponding to the second candidate target reaches the first threshold, then the region where the second candidate target is located is verified to contain a valid sound signal, and the second spatial information P is calculated. aunio The third spatial information of the second candidate target output by the camera (denoted as P) came aThe system calculates the Euclidean distance between the second and third spatial information. If the Euclidean distance is less than or equal to a preset distance threshold, the second spatial information is considered to match the third spatial information. If the Euclidean distance is greater than the preset distance threshold, the camera is re-triggered to visually identify the area corresponding to the sound-emitting target. Conversely, if the VAD confidence level corresponding to the second candidate target does not reach the first threshold, the area of ​​the second candidate target is verified to lack a valid sound signal. The microphone device then removes the second candidate target output by the camera that corresponds to the VAD confidence level, retaining only the second candidate target output by the camera whose VAD confidence level reaches the first threshold.

[0126] In optional embodiments of this application, the radar device can be integrated into the microphone device; optionally, when there is only one radar device, the radar device can coincide with the geometric center of the microphone device; when there are multiple radar devices, the distribution pattern of the multiple radar devices on the microphone device includes, but is not limited to, any one of the following patterns: ring uniform distribution, linear uniform distribution, matrix uniform distribution, fan-shaped uniform distribution, etc.

[0127] For example, the radar device and the microphone device are physically designed, manufactured, or assembled into a single, indivisible integrated device. For instance, the radar device may be encapsulated inside the housing of the microphone device, or the radar device may be directly embedded in the circuit board of the microphone array; or the radar device may be tightly fixed to the outside of the microphone device, or the radar device may be secured to the housing of the microphone device, such as the microphone array, by structural components, forming a compact unit.

[0128] It's important to explain that when there's only one radar device, its installation position perfectly coincides with the geometric center of the microphone device. This ensures that the spatial reference for radar detection is consistent with the acoustic reference for microphone acquisition, reducing spatial calibration errors during subsequent data fusion. Similarly, when there are multiple radar devices, their distribution on the microphone device can be any of the following: ring-shaped uniform distribution, linear uniform distribution, matrix uniform distribution, or fan-shaped uniform distribution. When multiple radar devices are present, they are arranged according to a specific geometric pattern on the physical structure of the microphone device, and the distance or angle between each radar device is uniform. For example, a ring-shaped uniform distribution means that multiple radar devices are arranged around the center of the microphone device at equal angular intervals on a circle; a linear uniform distribution means that multiple radar devices are arranged equidistantly along a certain axis of the microphone device; a matrix uniform distribution means that multiple radar devices are arranged in rows and columns to form a regular grid on the plane of the microphone device; and a fan-shaped uniform distribution means that multiple radar devices are evenly distributed along the arc or radial direction of a fan-shaped area, with a certain reference point of the microphone device as the vertex.

[0129] This application achieves device integration by integrating the radar device into the microphone device and precisely configuring the geometry according to the number of radar devices. This significantly reduces external interference and installation complexity, ensures a close physical connection between the radar and microphone, and avoids signal mismatch between independent devices. When there is only one radar device, the geometric centers of the radar device and the microphone device coincide, eliminating positional deviations and making the acquisition of first and second spatial information more consistent. When there are multiple radar devices, various uniform distribution patterns optimize the radar coverage, reduce detection blind spots, and enhance multi-target detection capabilities.

[0130] In optional embodiments of this application, the microphone device may be, but is not limited to, any of the following forms: ceiling-mounted, wall-mounted, desktop, or front-mounted. A ceiling-mounted microphone device refers to a microphone device installed on the ceiling, which can utilize the overhead space of the conference room to achieve wide-area sound coverage of the conference area and reduce blind spots caused by obstacles. A wall-mounted microphone device refers to a microphone device installed on a wall, which can save desktop space and can be flexibly deployed according to the wall layout of the conference room, suitable for small and medium-sized conference rooms or specific areas requiring sound pickup. A desktop microphone device refers to a microphone device placed on a conference table or other flat surface, usually close to the speaker, enabling close-range, high-definition sound acquisition, effectively suppressing environmental noise, and improving the signal-to-noise ratio of the voice signal. A front-mounted microphone device refers to a microphone device placed at the front of the conference room (e.g., below the display screen or at the podium), which usually has strong directionality, can concentrate on picking up sound from the front area, reducing interference from the sides and rear, and is suitable for conference scenarios with a clearly defined speaker location.

[0131] This application provides microphone devices in various forms, such as ceiling-mounted, wall-mounted, desktop, and front-facing microphones. The system can be flexibly configured according to the actual environment and usage needs of the conference room, avoiding blind spots in sound acquisition and signal distortion.

[0132] Furthermore, in the embodiments of this application, after determining the spatial information of the target speaker, it can also be applied in many other scenarios. Several scenarios are listed below for illustrative purposes: In some embodiments, the microphone device may adjust the beam of the microphone device based on the spatial information of the target speaker. For example, the beamforming circuit of the microphone device may adjust the position of one or more fixed beams to cover the target speakers in the area based on the spatial information of the target speakers, such that each target speaker is covered by at least one fixed beam.

[0133] In some embodiments, the microphone device may further adjust one or more beams based on the spatial information of the target speaker to acquire the sound of the target speaker within the area; optionally, when there are multiple target speakers, the microphone device may further adjust one or more beams based on the spatial information of the target speakers to acquire the sound of one target speaker within the area (e.g., the target speaker with the loudest voice); then the microphone device may perform acoustic echo cancellation based on the corresponding beam signal received from the beamforming circuit and output the echo-corrected beam signal. Alternatively, the microphone device may further shield sound signals outside the area where the spatial information of the target speaker is located, based on the spatial information of the target speaker.

[0134] In some embodiments, the microphone device can also control the camera to perform real-time facial tracking of the target speaker based on the target speaker's spatial information. Correspondingly, the camera can send the tracked facial information of the target speaker to the terminal device for display on the terminal device's screen. Optionally, the microphone device can also simultaneously mute non-speaking individuals, preventing them from being displayed on the terminal device's screen. For ease of understanding facial tracking, such as... Figure 2 As shown, Figure 2 This is a schematic diagram of the speaker image tracking switching provided in the embodiment of this application. User A is displayed on the screen of the terminal device. When the speaker is identified as User B, the microphone device controls the camera to perform image tracking on User B based on the spatial information of User B, and sends the image of User B to the terminal device in real time. Then the terminal device displays the image of User B, realizing the effect of switching from User A to User B.

[0135] Finally, it should be noted that in the embodiments of this application, the execution order of steps 101 and 102 is not sequential and can be changed. For example, step 101 can be executed first and then step 102, or step 102 can be executed first and then step 101, or steps 101 and 102 can be executed simultaneously. The specific execution order can be set according to the actual situation and is not limited here.

[0136] Secondly, such as Figure 3 As shown, Figure 3 This is another flowchart illustrating the speaker location method in a conference room provided in this application embodiment, which provides a speaker location method in a conference room. Figure 3 The method of the embodiment includes, but is not limited to, steps 301-306: Step 301: Acquire the fourth spatial information of at least one fifth candidate target output by the radar device.

[0137] Step 302: Analyze the acquired images using the camera to obtain facial motion features of at least one sixth candidate target.

[0138] Step 303: Based on the matching degree between the facial motion features of the camera and the features in the preset vocal facial motion feature library, determine at least one seventh candidate target from at least one sixth candidate target.

[0139] Step 304: Obtain the fifth spatial information of at least one seventh candidate target output by the camera.

[0140] Step 305: Based on the fourth and fifth spatial information, determine the first target speaker.

[0141] Step 306: Determine the spatial information of the first target speaker based on the fourth and fifth spatial information corresponding to the first target speaker.

[0142] Similarly, it should first be noted that, Figure 3 The method provided in this embodiment can be applied to a conference room speaker positioning system, which may include a processing device, a camera, and a microphone device. Some or all of the processing device, camera, and microphone device may be independently configured devices; optionally, the processing device may be integrated into the microphone device. For example, when the processing device is independent of the microphone device and camera, steps 301-306 can be interpreted as steps executed by the processing device; when the processing device is integrated into the microphone device, steps 301-306 can be interpreted as steps executed by the microphone device (i.e., the processing device within the microphone device). Optionally, the conference room speaker positioning method can also be applied to a conference room speaker positioning device, wherein the device may be the aforementioned processing device independent of the microphone device and camera, or it may be a microphone device with an integrated processing device.

[0143] In the embodiments of this application, the fifth candidate target in step 301 can refer to the description of the candidate target in step 101, and the fourth spatial information can refer to the description of the first spatial information mentioned above, which will not be repeated here.

[0144] In the embodiments of this application, the sixth candidate target in step 302 can be all human targets identified by the camera from the acquired images through image segmentation and human detection algorithms. The camera acquires images of the area and analyzes them, for example, by identifying all human targets from the images through image segmentation and human detection algorithms, and determines them as the sixth candidate target. Optionally, the sixth candidate target can also refer to the description of the third candidate target in the above embodiments, which will not be repeated here.

[0145] In the embodiments of this application, the seventh candidate target in step 302 can be an object suspected of having vocalization behavior, selected by the camera based on the matching results of facial motion features and a preset vocalization facial motion feature library. The camera uses a convolutional neural network combined with optical flow to extract features from the face and related parts of each sixth candidate target to obtain facial motion features, such as lip movements, facial muscle movements, and eye changes. The extracted facial motion features are matched in real time with features in the preset vocalization facial motion feature library. If the matching degree reaches a preset threshold, the corresponding sixth candidate target is determined as the seventh candidate target. Optionally, the principle for determining the seventh candidate target can refer to the description of the principle for determining the second candidate target in the above embodiments, which will not be repeated here.

[0146] In the embodiments of this application, the description of obtaining the fifth spatial information of at least one seventh candidate target output by the camera in step 304 can refer to the description of obtaining the third spatial information of at least one second candidate target output by the camera above, and will not be repeated here.

[0147] In the embodiments of this application, the process of determining the first target speaker in step 305 can refer to the above description of determining the target speaker based on the spatial matching relationship between the first spatial information and the second spatial information, and will not be repeated here.

[0148] In the embodiments of this application, the process of determining the spatial information of the first target speaker in step 306 can refer to the description of determining the spatial information of the target speaker based on the first spatial information and the second spatial information corresponding to the target speaker, and will not be repeated here.

[0149] It should be explained that, even in complex conference room environments with background noise or reverberation interference, the system can still accurately identify the speaker's location. For example, when a participant speaks, their lip movement features match the feature database, and simultaneously, the fourth spatial information detected by radar matches the fifth spatial information located by the camera. The system fuses these two types of spatial information to output a high-precision positioning result, avoiding the limitations of a single sensor. Overall, this application significantly improves the accuracy and robustness of sound source localization, providing reliable technical support for video conferencing systems.

[0150] In optional embodiments of this application, the camera can acquire images either automatically or under the instruction of a radar device. Specifically, the radar device sends fourth-space information to the camera, and then the camera acquires images based on the fourth-space information.

[0151] The radar device can transmit fourth-space information to the camera via a wired communication interface, such as Ethernet or USB; alternatively, it can send the fourth-space information to the camera wirelessly, such as via Wi-Fi or Bluetooth. The camera adjusts its image acquisition behavior based on the received fourth-space information to obtain more targeted image data. The camera can, based on the target area indicated by the fourth-space information, control its pan / tilt and zoom functions to center and magnify the target area within its field of view; or, within its fixed field of view, the camera can crop the image or focus on specific areas based on the fourth-space information, thereby acquiring or focusing only on the image portion relevant to the fifth candidate target.

[0152] In this embodiment, acquiring images using a camera based on fourth spatial information may include: mapping the global coordinate system corresponding to the fourth spatial information to the camera's imaging coordinate system based on a preset sensor extrinsic calibration model, and calculating the pixel coordinates of the fifth candidate target on the camera's imaging plane; extracting target spatial range information from the fourth spatial information output by the radar device, and determining the target area and pixel percentage of the fifth candidate target in the camera's field of view; adjusting the azimuth and pitch axis rotation angles of the camera's gimbal drive module based on the azimuth and pitch angles of the fifth candidate target, so that the fifth candidate target is located at the center of the camera's field of view; and determining the target area and pixel percentage of the fifth candidate target based on the distance between the fifth candidate target and the camera. Based on the target size, the camera focal length is adjusted so that the pixel ratio of the fifth candidate target on the imaging plane reaches a preset threshold. The adjusted horizontal and vertical field of view angles are calculated in real time by the camera to verify whether the target area completely covers the camera's field of view. After completing the field of view matching, the camera triggers image acquisition according to the type of fourth spatial information output by the radar. Specifically, if the fourth spatial information is static target coordinates, a single-frame high-resolution image acquisition is triggered; if the fourth spatial information is a dynamic target trajectory, continuous frame acquisition is started at a preset frame rate, and the field of view and parameters are dynamically adjusted according to the fourth spatial information updated by the radar in real time. For the target area located by the radar, image data within the region of interest is acquired.

[0153] For example, after the camera receives the fourth spatial information transmitted by the radar device, it achieves accurate image acquisition of the radar positioning area through a closed-loop control process of spatial coordinate mapping, dynamic field-of-view matching, acquisition parameter optimization, and target area imaging. The specific technical implementation is as follows: Part 1: Transformation of the position coordinates of the vocalizing human target into the spatial coordinate system.

[0154] (1) The camera end receives the fourth spatial information of the fifth candidate target and completes the coordinate system transformation through the preset sensor extrinsic calibration model: Based on the pinhole camera model and the hand-eye calibration algorithm, the global coordinate system corresponding to the fourth spatial information generated by the radar device (e.g., with the geometric center of the microphone device / positioning system as the origin) is mapped to the imaging coordinate system of the camera (e.g., with the optical center of the camera as the origin). The calculation process of solving the pixel coordinates (u,v) of the fifth candidate target on the camera imaging plane is shown in formula (1): (1) Where K is the camera intrinsic parameter matrix (e.g., containing focal length and principal point coordinates), and R and T are the rotation matrix and translation vector between the radar and the camera, respectively (e.g., obtained through the coordinate calibration mentioned above, which will not be elaborated here).

[0155] (2) Extract the target spatial range information (e.g., the target bounding box size W×H×D) from the fourth spatial information output by the radar device, calculate the target area and pixel ratio of the fifth candidate target in the camera field of view, and provide a basis for subsequent field of view adjustment.

[0156] Part Two: Based on the coordinate system transformation results, the camera adjusts its attitude and focal length via the gimbal drive module, ensuring that the target area located by the radar falls completely within the camera's effective field of view. (1) Gimbal angle adjustment: Based on the azimuth and pitch angles of the fifth candidate target, drive the azimuth and pitch axes of the gimbal to rotate to the corresponding angles to ensure that the target area is in the center of the camera's field of view; the control accuracy of the rotation angle is determined by the resolution of the gimbal's servo motor. (2) Dynamic focal length adaptation: Based on the distance Z between the fifth candidate target and the camera r Given the target size, the camera focal length f is adjusted using an automatic zoom algorithm to ensure that the pixel percentage of the fifth candidate target on the imaging plane reaches a fifth preset threshold. The calculation formula is as follows: Where s is the camera pixel size, d is the target pixel width on the imaging plane, and S is the actual physical width of the target.

[0157] (3) Field of View (FOV) Verification: Real-time calculation of the adjusted horizontal / vertical field of view of the camera. Among them, W s / H s (The size is the camera size) to ensure that the target area located by the radar is completely covered within the field of view.

[0158] Part Three: Image Acquisition and Trigger Control of the Target Area. After completing field-of-view matching, the camera triggers image acquisition in the following ways: (1) Single acquisition trigger: If the position data output by the radar device, i.e. the fourth spatial information, is a static target, the camera will immediately trigger the acquisition of a single frame high-resolution image after receiving the position data.

[0159] (2) Continuous frame acquisition: If the position data output by the radar device, i.e. the fourth spatial information, is a dynamic target, the camera starts continuous acquisition at a preset frame rate and dynamically adjusts the field of view and parameters according to the position data updated in real time by the radar device to achieve target tracking and imaging.

[0160] (3) Region of Interest Acquisition: For the target area located by radar, the camera only acquires image data within the region of interest (such as the human image area of ​​the fifth candidate target) (instead of the full frame), reducing the data transmission bandwidth and subsequent image processing computing power consumption.

[0161] It should be explained that the radar device of this application can provide the camera with preliminary spatial indication information, enabling the camera to collect images in a targeted manner, avoiding blind scanning and processing of the entire scene, thereby effectively solving the problem of inaccurate image acquisition range, significantly reducing the data processing burden of the camera, and improving the efficiency and accuracy of image analysis, thereby enhancing the overall efficiency and accuracy of speaker positioning in the conference room.

[0162] It should be noted that, Figure 3 The radar device, microphone device, camera, and processing device in the embodiment can also directly or indirectly perform the above-mentioned functions. Figure 1 The steps performed by the processing device, radar device, and microphone device in the embodiments are the same and can achieve the same technical effect, and will not be described in detail here.

[0163] It should also be noted that in the embodiments of this application, the execution order of steps 301 and 302 is not sequential and can be changed. For example, step 301 can be executed first and then step 302, or step 302 can be executed first and then step 301, or steps 301 and 302 can be executed simultaneously. The specific execution order can be set according to the actual situation and is not limited here.

[0164] Thirdly, such as Figure 4 As shown, Figure 4 This is another flowchart illustrating the speaker location method in a conference room provided in this application embodiment, which provides a speaker location method in a conference room. Figure 4 The method of the embodiment includes, but is not limited to, steps 401-404: Step 401: Obtain the first spatial information of at least one candidate target output by the radar device.

[0165] Step 402: Obtain the second spatial information of at least one sound-emitting target output by the microphone device.

[0166] Step 403: Determine the first spatial information corresponding to each sound-emitting target based on the spatial matching relationship between the first spatial information and the second spatial information.

[0167] Step 404: Determine the spatial information of each sound-emitting target based on the first spatial information and the second spatial information corresponding to each sound-emitting target.

[0168] In the embodiments of this application, the principle of step 401 can be referred to the description of step 101, and will not be repeated here. The principle of step 402 can be referred to the description of step 102, and will not be repeated here. The principle of determining the first spatial information corresponding to each sound-emitting target in step 403 can be referred to the above. Figure 1 The principle of determining the first spatial information corresponding to the target speaker in the embodiment is described in detail here, and will not be repeated here. The principle of step 404 can be referred to the description of step 104, and will not be repeated here.

[0169] Similarly, it should be noted that, Figure 4 The conference room speaker positioning method provided in this embodiment can be applied to a conference room speaker positioning system, which may include a processing device, a radar device, and a microphone device. Some or all of the processing device, radar device, and microphone device may be independently configured devices. Optionally, the processing device may be integrated into the microphone device; alternatively, the radar device may also be integrated into the microphone device. For example, when the processing device is independent of the microphone device and radar device, steps 401-404 can be interpreted as steps executed by the processing device; when the processing device is integrated into the microphone device, steps 401-404 can be interpreted as steps executed by the microphone device (i.e., the processing device within the microphone device). Optionally, the conference room speaker positioning method can also be applied to a conference room speaker positioning device, wherein the device may be the aforementioned processing device independent of the microphone device and radar device, or it may be a microphone device integrated with a processing device. Furthermore, the above two configuration methods include various layout forms where the radar device and microphone device are integrated or not integrated, which will not be elaborated here.

[0170] In this embodiment, all valid sound-emitting targets are first identified using a microphone device. Then, the spatial positioning of each sound-emitting target is supplemented and calibrated using the first spatial information output by the radar device, ultimately outputting the accurate spatial information of all sound-emitting targets. The process of determining the first spatial information corresponding to each sound-emitting target based on the spatial matching relationship between the first and second spatial information can refer to the description of determining the spatial information of the target speaker based on the first and second spatial information corresponding to the target speaker, and will not be repeated here. It should be noted that the conference room speaker positioning method can be applied to a conference room speaker positioning system or a microphone device, and can be referred to the aforementioned embodiments, which will not be repeated here.

[0171] It should be noted that the third method differs from the first method in the speaker selection logic and the scope of the output objects. The third method primarily uses the microphone device to determine the target speaker. The first method only identifies the target speaker as an object that is simultaneously identified as a candidate by the radar device and as a vocal target by the microphone device, and whose spatial matching relationship meets the criteria. Essentially, it uses dual devices for dual verification to lock in high-confidence vocal targets, filtering out vocal targets identified by the microphone device but not by the radar device. The third method includes all vocal targets identified by the microphone device in the output scope, without additional main speaker selection, thoroughly covering the needs of scenarios with multiple sound sources emitting sound simultaneously.

[0172] The first approach filters target speakers based on spatial matching relationships, retaining only those that meet the matching criteria. The third approach supplements calibration based on spatial matching relationships. For each microphone-identified vocal target, it iterates through the first spatial information output by the radar to perform matching, assigning corresponding radar spatial data to correct acoustic positioning errors. Even if a vocal target does not meet the matching threshold with the first spatial information of all candidate targets, the vocal target will still be retained. Its second spatial information is used as a basis for iterative calibration based on historical data to supplement positioning, ensuring that no valid vocal targets are missed.

[0173] This application acquires first spatial information of candidate targets through a radar device and second spatial information of sound-emitting targets through a microphone device. For each sound-emitting target, it matches it with the first spatial information (ensuring no sound-emitting target is missed) and fuses the two types of spatial information to determine the final spatial information of the target speaker, ensuring the completeness of sound-emitting target identification and adapting to scenarios where multiple people speak at the same time; it also improves positioning accuracy and reliability, and enables the high-precision spatial information of the radar device to calibrate acoustic positioning.

[0174] It should be noted that, Figure 4 The radar device, microphone device, camera, and processing device in the embodiment can also directly or indirectly perform the above-mentioned functions. Figure 1The steps performed by the processing device, radar device, and microphone device in the embodiments are the same and can achieve the same technical effect, and will not be described in detail here.

[0175] It should also be noted that in the embodiments of this application, the execution order of steps 401 and 402 is not sequential and can be changed. For example, step 401 can be executed first and then step 402, or step 402 can be executed first and then step 401, or steps 401 and 402 can be executed simultaneously. The specific execution order can be set according to the actual situation and is not limited here.

[0176] Fourthly, such as Figure 5 As shown, Figure 5 This is a schematic diagram of a conference room speaker positioning system provided in an embodiment of this application. Figure 5 The system shown includes a processing unit 501, a radar unit 502, and a microphone unit 503; the relationship between the processing unit 501, radar unit 502, and microphone unit 503 will not be described in detail here, but can be referred to the above description. Figure 1 Corresponding description of the embodiments.

[0177] Radar device 502 is used to output first spatial information of at least one candidate target; Microphone device 503 is used to output second spatial information of at least one sound-emitting target; The processing device 501 is used to determine the target speaker based on the spatial matching relationship between the first spatial information and the second spatial information; The processing device 501 is also used to determine the spatial information of the target speaker based on the first spatial information and the second spatial information corresponding to the target speaker.

[0178] It should be explained that this application combines the first spatial information of the candidate target output by the radar device 502 with the second spatial information of the sound-emitting target output by the microphone device 503 in a spatial matching relationship, and fuses the two spatial information based on the matching result, thereby accurately associating physical targets with sound-emitting behavior in the conference room environment. This solves the problem that single sensor positioning is easily affected by environmental interference, and achieves the effect of significantly improving the positioning accuracy and reliability of the speaker.

[0179] Specifically, the radar device actively detects physical targets in the conference room, providing first spatial information less affected by the acoustic environment. This information describes the location of candidate targets in space. The microphone device captures sound signals in real time, providing second spatial information directly related to speaking behavior. The processing device analyzes the spatial distance or regional overlap between the first and second spatial information to establish a spatial matching relationship, thereby accurately mapping vocal behavior to specific physical targets and effectively eliminating misidentification caused by reverberation or background noise. After identifying the target speaker, the processing device further integrates the first and second spatial information, for example, through arithmetic averaging or selection based on preset priority rules, to generate more stable and accurate target speaker spatial information, using the stability of radar positioning and the real-time performance of microphone sound source identification for mutual correction.

[0180] It should be noted that, Figure 5 The principles of each step in the embodiment can be referred to the above. Figure 1 The description of the embodiments will not be repeated here; Figure 5 The processing device 501, radar device 502, and microphone device 503 in the embodiments can also directly or indirectly execute the steps executed by the processing device, radar device, and microphone device in the above embodiments, and can achieve the same technical effect, which will not be described in detail here.

[0181] Fifthly, such as Figure 6 As shown, Figure 6 This is another structural diagram of the conference room speaker positioning system provided in the embodiments of this application. Figure 6 The system shown includes a processing device 501, a radar device 502, and a camera 504; the relationship between the processing device 501, the radar device 502, and the camera 504 will not be described in detail here, but can be referred to the corresponding description in the above embodiments; A radar device for outputting fourth-space information of at least one fifth candidate target; A camera is used to analyze the acquired images to obtain facial motion features of at least one sixth candidate target; The camera is also used to determine at least one seventh candidate target from at least one sixth candidate target based on the matching degree between facial motion features and features in a preset vocal facial motion feature library. A processing device for acquiring fifth spatial information of at least one seventh candidate target output by a camera; The processing device is also used to determine the first target speaker based on fourth-space information and fifth-space information; The processing device is also used to determine the spatial information of the first target speaker based on the fourth and fifth spatial information corresponding to the first target speaker.

[0182] exist Figure 5 and Figure 6 Based on this, as an example, such as Figure 7 As shown, Figure 7 This is a schematic diagram of a conference room speaker positioning scenario provided in an embodiment of this application. For example, camera 504 can be positioned below the conference room display screen. Optionally, camera 504 can also be a desktop camera, a wall-mounted camera, etc. This application does not limit the type of camera, nor does it limit the number of cameras. This embodiment combines the fourth spatial information output by the radar device with the facial motion features and fifth spatial information acquired by the camera in a multimodal fusion manner, thereby solving the key problem of poor sound source positioning accuracy through spatial information cross-validation and visual confirmation of vocal behavior.

[0183] It should be explained that the radar device actively emits electromagnetic waves and receives reflected signals, outputting the fourth spatial information of the fifth candidate target, providing physical location data unaffected by the acoustic environment as a preliminary positioning basis; the camera performs real-time analysis of the acquired images, extracting facial motion features of the sixth candidate target, such as lip opening and closing, facial muscle movements, and other dynamic behaviors related to speech, and calculates the matching degree based on a preset facial motion feature library for speech, filtering out the seventh candidate target with a high probability of speech, effectively eliminating interference from non-speaking targets; the processing device integrates the fourth spatial information of the radar device and the fifth spatial information output by the camera, and determines the first target speaker by calculating the overlapping area of ​​spatial positions or distance thresholds, achieving accurate association between the sound source and the physical entity; finally, the processing device generates fused spatial information based on the fourth and fifth spatial information corresponding to the first target speaker, using a weighted average or priority selection strategy, significantly improving positioning accuracy.

[0184] It should be noted that, Figure 6 The principles of each step in the embodiment can be referred to the above. Figure 3 The description of the embodiments will not be repeated here; Figure 6 The processing device 501, radar device 502, and camera 504 in the embodiments can also directly or indirectly execute the steps executed by the processing device, radar device, and camera in the above embodiments, and can achieve the same technical effect, which will not be described in detail here.

[0185] Sixth aspect, such as Figure 5 As shown, a conference room speaker positioning system is provided. The system includes a processing device 501, a radar device 502, and a microphone device 503. The relationship between the processing device 501, the radar device 502, and the microphone device 503 will not be described in detail here, but can be referred to the corresponding description in the above embodiments.

[0186] Radar device 502 is used to output first spatial information of at least one candidate target; Microphone device 503 is used to output second spatial information of at least one sound-emitting target; The processing device 501 is used to determine the first spatial information corresponding to each sound-emitting target based on the spatial matching relationship between the first spatial information and the second spatial information; The processing device 501 is also used to determine the spatial information of each sound-emitting target based on the first spatial information and the second spatial information corresponding to each sound-emitting target.

[0187] This embodiment combines the first spatial information of the candidate target output by the radar device 502 with the second spatial information of the sound-emitting target output by the microphone device 503 in a spatial matching manner, and performs information fusion through a processing device, thereby effectively solving the positioning accuracy problem caused by the positioning error of a single sensor.

[0188] Specifically, radar device 502 can actively emit electromagnetic waves to detect physical targets in the environment, providing stable first spatial information, and its positioning results are less affected by the acoustic environment; microphone device 503 determines the second spatial information of the sound-emitting target through acoustic signal processing, ensuring a direct correlation with the sound-emitting behavior. Processing device 501 establishes a precise spatial matching relationship by calculating the spatial distance or regional overlap between the first and second spatial information, accurately associating the sound-emitting target with the physical target, and further integrating the advantages of the two information sources to obtain more robust final spatial information.

[0189] It should be noted that the principles of each step in the sixth aspect embodiment can be referred to the above. Figure 3 The description of the embodiments will not be repeated here; the processing device 501, radar device 502 and microphone device 503 in the sixth aspect embodiment can also directly / indirectly execute the steps executed by the processing device, radar device and microphone device in the above embodiments, and can achieve the same technical effect, which will not be described one by one here.

[0190] The seventh aspect, such as Figure 8 As shown, Figure 8 This is a schematic diagram of a conference room speaker positioning device provided in an embodiment of this application. The device includes a first acquisition module 801 and a first processing module 802. It should be noted that... Figure 8 The conference room speaker positioning device in the embodiment can be a device independent of the microphone device, radar device, and camera, or it can be a microphone device (or a processing device integrated in the microphone device). The first acquisition module 801 is used to acquire first spatial information of at least one candidate target output by the radar device; The first acquisition module 801 is also used to acquire second spatial information of at least one sound-emitting target output by the microphone device; The first processing module 802 is used to determine the target speaker based on the spatial matching relationship between the first spatial information and the second spatial information; The first processing module 802 is also used to determine the spatial information of the target speaker based on the first spatial information and the second spatial information corresponding to the target speaker.

[0191] It should be noted that, Figure 8 The principles of each step in the embodiment can be referred to the above. Figure 1 The description of the embodiments will not be repeated here; Figure 8 The conference room speaker positioning device in the embodiment can also directly or indirectly execute the steps executed by the processing device, radar device, microphone device and camera in the above embodiment, and can achieve the same technical effect, which will not be described in detail here.

[0192] The eighth aspect, such as Figure 9 As shown, Figure 9 This is another schematic diagram of the conference room speaker positioning device provided in the embodiments of this application. The device includes a second acquisition module 901 and a second processing module 902. It should be noted that... Figure 9 The conference room speaker positioning device in the embodiment can be a device independent of the microphone device, radar device, and camera, or it can be a microphone device (or a processing device integrated in the microphone device). The second acquisition module 901 is used to acquire the fourth spatial information of at least one fifth candidate target output by the radar device; The second processing module 902 is used to analyze the acquired images through the camera to obtain facial motion features of at least one sixth candidate target; The second processing module 902 is further configured to determine at least one seventh candidate target from at least one sixth candidate target by using the matching degree between the facial motion features of the camera and the features in the preset vocal facial motion feature library. The second acquisition module 901 is also used to acquire the fifth spatial information of at least one seventh candidate target output by the camera; The second processing module 902 is also used to determine the first target speaker based on the fourth spatial information and the fifth spatial information; The second processing module 902 is also used to determine the spatial information of the first target speaker based on the fourth spatial information and the fifth spatial information corresponding to the first target speaker.

[0193] It should be noted that, Figure 9The principles of each step in the embodiment can be referred to the above. Figure 1 The description of the embodiments will not be repeated here; Figure 9 The conference room speaker positioning device in the embodiment can also directly or indirectly execute the steps executed by the processing device, radar device, microphone device, and camera in the above embodiment, and can achieve the same technical effect, which will not be described in detail here.

[0194] Ninth aspect, such as Figure 10 As shown, Figure 10 This is another schematic diagram of the conference room speaker positioning device provided in the embodiments of this application. The device includes a third acquisition module 901 and a third processing module 902. It should be noted that... Figure 10 The conference room speaker positioning device in the embodiment can be a device independent of the microphone device, radar device, and camera, or it can be a microphone device (or a processing device integrated in the microphone device). The third acquisition module 1001 is used to acquire the first spatial information of at least one candidate target output by the radar device; The third acquisition module 1001 is also used to acquire second spatial information of at least one sound-emitting target output by the microphone device; The third processing module 1002 is used to determine the first spatial information corresponding to each sound-emitting target based on the spatial matching relationship between the first spatial information and the second spatial information. The third processing module 1002 is also used to determine the spatial information of each sound-emitting target based on the first spatial information and the second spatial information corresponding to each sound-emitting target.

[0195] It should be noted that, Figure 10 The principles of each step in the embodiment can be referred to the above. Figure 1 The description of the embodiments will not be repeated here; Figure 10 The conference room speaker positioning device in the embodiment can also directly or indirectly execute the steps executed by the processing device, radar device, microphone device, and camera in the above embodiment, and can achieve the same technical effect, which will not be described in detail here.

[0196] The tenth aspect, such as Figure 11 As shown, Figure 11 This is another flowchart illustrating the conference room speaker location method provided in this application. The method is applied to a conference room speaker location device, which is a microphone device. The relationship between the microphone device and the radar device will not be elaborated here. The method includes, but is not limited to, steps 1101-1104: Step 1101: Acquire the first spatial information of at least one candidate target output by the radar device; Step 1102: Determine the second spatial information of at least one sound-emitting target; That is to say, the radar device is integrated into the microphone device to acquire second spatial information of at least one sound-emitting target determined by the microphone device.

[0197] Step 1103: Determine the target speaker based on the spatial matching relationship between the first spatial information and the second spatial information; Step 1104: Determine the spatial information of the target speaker based on the first spatial information and the second spatial information corresponding to the target speaker.

[0198] In some embodiments, when there is only one radar device, the geometric center of the radar device coincides with that of the microphone device; when there are multiple radar devices, the distribution pattern of the multiple radar devices on the microphone device includes any one of the following: ring uniform distribution, linear uniform distribution, matrix uniform distribution, and fan-shaped uniform distribution.

[0199] In some embodiments, the microphone device may be in any of the following forms: ceiling-mounted, wall-mounted, desktop, or front-mounted.

[0200] It should be noted that the principles of each step in this embodiment can be referred to the corresponding description of the method in the first aspect, and will not be repeated here. Also, Figure 11 The conference room speaker positioning device in the embodiment can also directly or indirectly execute the steps executed by the processing device, radar device, microphone device, and camera in the above embodiment, and can achieve the same technical effect, which will not be described in detail here.

[0201] Eleventh aspect, such as Figure 12 As shown, Figure 12 This is another flowchart illustrating the conference room speaker location method provided in this application embodiment. The method is applied to a conference room speaker location device, which is a microphone device. The relationship between the microphone device and the radar device will not be elaborated here. The method includes, but is not limited to, steps 1201-1204: Step 1201: Obtain the first spatial information of at least one candidate target output by the radar device; Step 1202: Determine the second spatial information of at least one sound-emitting target; Step 1203: Determine the first spatial information corresponding to each sound-emitting target based on the spatial matching relationship between the first spatial information and the second spatial information; Step 1204: Determine the spatial information of each sound-emitting target based on the first spatial information and the second spatial information corresponding to each sound-emitting target.

[0202] In some embodiments, when there is only one radar device, the geometric center of the radar device coincides with that of the microphone device; when there are multiple radar devices, the distribution pattern of the multiple radar devices on the microphone device includes any one of the following: ring uniform distribution, linear uniform distribution, matrix uniform distribution, and fan-shaped uniform distribution.

[0203] In some embodiments, the microphone device may be in any of the following forms: ceiling-mounted, wall-mounted, desktop, or front-mounted.

[0204] It should be noted that the principles of each step in this embodiment can be referred to the corresponding description of the method in the first aspect, and will not be repeated here. Also, Figure 12 The conference room speaker positioning device in the embodiment can also directly or indirectly execute the steps executed by the processing device, radar device, microphone device, and camera in the above embodiment, and can achieve the same technical effect, which will not be described in detail here.

[0205] In a twelfth aspect, a conference room speaker positioning system is provided, the system including a radar device and a microphone device; the radar device is integrated into the microphone device; A radar device for outputting first spatial information of at least one candidate target; A microphone device is used to determine second spatial information of at least one vocal target; determine a target speaker based on a spatial matching relationship between first spatial information and second spatial information; and determine spatial information of the target speaker based on the first spatial information and second spatial information corresponding to the target speaker.

[0206] In some embodiments, when there is only one radar device, the geometric center of the radar device coincides with that of the microphone device; when there are multiple radar devices, the distribution pattern of the multiple radar devices on the microphone device includes any one of the following: ring uniform distribution, linear uniform distribution, matrix uniform distribution, and fan-shaped uniform distribution.

[0207] In some embodiments, the microphone device may be in any of the following forms: ceiling-mounted, wall-mounted, desktop, or front-mounted.

[0208] As an example, such as Figure 13 As shown, Figure 13This is another schematic diagram of a conference room speaker positioning scenario provided in the embodiments of this application. The microphone device 1301 can be a ceiling-mounted microphone, and the number of microphone devices 1301 can be determined according to the actual situation, and is not limited here. This embodiment can refer to the system description in the fourth aspect, which will not be repeated here.

[0209] In some embodiments, each radar device acquires location data (including location data of dynamic objects and static sound-emitting objects, such as first spatial information) of objects (e.g., candidate targets) within the area in real time. The location data may include location coordinates. Location data acquired by some or all radar devices are sent to each microphone device, including but not limited to the following scenarios: (1) location data acquired by each radar device is sent to each microphone device; (2) location data from at least one radar device among multiple radar devices is sent to each microphone device; (3) fused location data obtained by fusing location data acquired by multiple radar devices is sent to each microphone device. The microphone devices synchronously acquire sound signals within the area in real time. Based on the location data received from the radar devices and the sound signals synchronously acquired by the microphone devices within the area, the microphone devices output the location of the speaking object (e.g., spatial information).

[0210] It should be noted that the principles of each step in this embodiment can be referred to the corresponding descriptions in the above embodiments, and will not be repeated here. Furthermore, the radar device and microphone device in this embodiment can also directly / indirectly execute the steps executed by the processing device, radar device, microphone device, and camera in the above embodiments, and can achieve the same technical effect; these will not be described in detail here.

[0211] In a thirteenth aspect, a conference room speaker positioning system is provided, the system including a radar device and a microphone device; the radar device is integrated into the microphone device; A radar device for outputting first spatial information of at least one candidate target; A microphone device is used to determine second spatial information of at least one sound-emitting target; determine first spatial information corresponding to each sound-emitting target based on the spatial matching relationship between first spatial information and second spatial information; and determine spatial information of each sound-emitting target based on the first spatial information and second spatial information corresponding to each sound-emitting target.

[0212] In some embodiments, when there is only one radar device, the geometric center of the radar device coincides with that of the microphone device; when there are multiple radar devices, the distribution pattern of the multiple radar devices on the microphone device includes any one of the following: ring uniform distribution, linear uniform distribution, matrix uniform distribution, and fan-shaped uniform distribution.

[0213] In some embodiments, the microphone device may be in any of the following forms: ceiling-mounted, wall-mounted, desktop, or front-mounted.

[0214] It should be noted that this embodiment can refer to the system description in aspect six, and will not be repeated here. The principles of each step in this embodiment can refer to the corresponding descriptions in the above embodiments, and will not be repeated here. Furthermore, the radar device and microphone device in this embodiment can also directly / indirectly execute the steps executed by the processing device, radar device, microphone device, and camera in the above embodiments, and can achieve the same technical effect, which will not be described one by one here.

[0215] In a fourteenth aspect, embodiments of this application provide a computer-readable storage medium storing a computer program thereon, such as a memory storing the computer program, which can be executed by a processor to perform the aforementioned method steps. The computer-readable storage medium may be a memory such as FRAM, ROM, PROM, EPROM, EEPROM, Flash Memory, magnetic surface memory, optical disc, or CD-ROM.

[0216] In a fifteenth aspect, embodiments of this application also provide a computer program product, the computer program product including a non-transitory computer-readable storage medium storing a computer program operable to cause a computer to perform some or all of the steps of any of the conference room speaker positioning methods described in the above method embodiments.

[0217] In addition, each functional unit in the various embodiments of this application can be integrated into one processing unit, or each unit can be a separate unit, or two or more units can be integrated into one unit; the integrated unit can be implemented in hardware or in the form of hardware plus software functional units.

[0218] Those skilled in the art will understand that all or part of the steps of the above method embodiments can be implemented by hardware related to program instructions. The aforementioned program can be stored in a computer-readable storage medium. When the program is executed, it performs the steps of the above method embodiments. The aforementioned storage medium includes various media capable of storing program code, such as mobile storage devices, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0219] Alternatively, if the integrated units described above are implemented as software functional modules and sold or used as independent products, they can also be stored in a computer-readable storage medium. Based on this understanding, the technical solutions of the embodiments of this application, or the parts that contribute to related technologies, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the methods of the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as mobile storage devices, ROM, RAM, magnetic disks, or optical disks.

[0220] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

Claims

1. A method for locating a speaker in a conference room, characterized in that, The method includes: Acquire the first spatial information of at least one candidate target output by the radar device; Acquire second spatial information of at least one sound-emitting target output by the microphone device; Based on the spatial matching relationship between the first spatial information and the second spatial information, the target speaker is determined; Based on the first spatial information and the second spatial information corresponding to the target speaker, the spatial information of the target speaker is determined.

2. The method according to claim 1, characterized in that, The step of determining the target speaker based on the spatial matching relationship between the first spatial information and the second spatial information includes: For each sound-emitting target, a first matching degree is determined between the second spatial information of the sound-emitting target and the first spatial information of each candidate target; The vocal targets whose first matching degree reaches the first preset threshold are identified as the target speakers.

3. The method according to claim 2, characterized in that, The method further includes: If there are multiple instances where the first matching degree of the vocal target reaches the first preset threshold, then the first spatial information corresponding to the highest first matching degree of the vocal target is determined as the first spatial information corresponding to the target speaker.

4. The method according to claim 3, characterized in that, The method further includes: If there are multiple highest first matching degrees, then the first spatial information corresponding to each highest first matching degree is used as the first spatial information corresponding to a target speaker, resulting in multiple first spatial information corresponding to multiple target speakers, wherein the multiple target speakers correspond one-to-one with the multiple first spatial information.

5. The method according to any one of claims 1-4, characterized in that, Based on the first spatial information and the second spatial information corresponding to the target speaker, the spatial information of the target speaker is determined, including: Based on a preset fusion algorithm, the first spatial information and the second spatial information corresponding to the target speaker are fused to obtain the spatial information; The preset fusion algorithm includes one of the following: weighted fusion algorithm, tightly coupled fusion algorithm, and Bayesian fusion algorithm.

6. The method according to claim 5, characterized in that, Based on the weighted fusion algorithm, the first spatial information and the second spatial information corresponding to the target speaker are fused to obtain the spatial information, including: Obtain the first weight corresponding to the first spatial information and the second weight corresponding to the second spatial information; Based on the first weight and the second weight, the first spatial information and the second spatial information corresponding to the target speaker are weighted and fused to obtain the spatial information.

7. The method according to claim 6, characterized in that, The relationship between the first weight and the second weight satisfies one of the following: If the target speaker is in a dynamic and vocal state, then the first weight is greater than the second weight; If the target speaker is in a static and vocal state, then the first weight is less than the second weight; If the matching degree between the second spatial information of two sound-emitting targets reaches the second preset threshold, then the first weight is greater than the second weight; If there is no matching degree between the second spatial information of two sound-emitting targets that reaches the second preset threshold, then the first weight is less than or equal to the second weight.

8. The method according to any one of claims 1-4, 6, and 7, characterized in that, The step of determining the target speaker based on the spatial matching relationship between the first spatial information and the second spatial information includes: Acquire the third spatial information of at least one second candidate target output by the camera; The target speaker is determined based on the first spatial information, the second spatial information, and the third spatial information; The step of determining the spatial information of the target speaker based on the first spatial information and the second spatial information corresponding to the target speaker includes: Based on the first spatial information, second spatial information, and third spatial information corresponding to the target speaker, the spatial information of the target speaker is determined.

9. The method according to claim 8, characterized in that, The method further includes: The camera acquires images based on spatial indication information, wherein the spatial indication information includes the first spatial information and / or the second spatial information. The image is analyzed by the camera to obtain facial motion features of at least one third candidate target; The camera determines the at least one second candidate target from the at least one third candidate target based on the matching degree between the facial motion features and features in a preset vocal facial motion feature library.

10. The method according to claim 9, characterized in that, The step of determining the at least one second candidate target from the at least one third candidate target by means of the matching degree between the facial motion features and features in a preset vocal facial motion feature library using the camera includes: Based on the matching degree between the facial motion features and the features in the preset vocal facial motion feature library, the camera determines at least one fourth candidate target from the at least one third candidate target; The camera determines whether each fourth candidate target is an interfering target based on its pose characteristics. The camera identifies a fourth candidate target that is not an interfering target as the at least one second candidate target.

11. The method according to claim 10, characterized in that, The matching degree between any two of the first, second, and third spatial information corresponding to the target speaker reaches a first preset threshold.

12. The method according to claim 11, characterized in that, The method further includes: For each vocal target, if there are multiple instances where the first matching degree of the vocal target reaches the first preset threshold, and the number of instances where the second matching degree of the vocal target reaches the first preset threshold is 1, then the first spatial information corresponding to the highest first matching degree of the vocal target is determined as the first spatial information corresponding to the target speaker. If the number of the first matching degree of the vocal target reaching the first preset threshold is 1, and the number of the second matching degree of the vocal target reaching the first preset threshold is multiple, then the third spatial information corresponding to the highest second matching degree of the vocal target is determined as the third spatial information corresponding to the target speaker. Wherein, the first matching degree represents the matching degree between the first spatial information and the second spatial information, and the second matching degree represents the matching degree between the second spatial information and the third spatial information.

13. The method according to claim 12, characterized in that, The method further includes: If the number of times the first matching degree of the vocal target reaches the first preset threshold is multiple, and the number of times the second matching degree of the vocal target reaches the first preset threshold is multiple, then the third matching degree between the first spatial information corresponding to each first matching degree and the third spatial information corresponding to each second matching degree is determined. The first spatial information and the third spatial information corresponding to the third matching degree that reaches the first preset threshold are used as the first spatial information and the third spatial information corresponding to the same target speaker, respectively. The first spatial information corresponding to the third matching degree that does not reach the first preset threshold is taken as the first spatial information corresponding to one target speaker, and the third spatial information corresponding to the third matching degree that does not reach the first preset threshold is taken as the third spatial information corresponding to another target speaker.

14. The method according to any one of claims 1-4, 9, 10, 11, 12, and 13, characterized in that, The method further includes: By analyzing the received echo signals using the radar device, at least one acoustic micro-motion characteristic of a first candidate target can be obtained. The radar device determines at least one candidate target from at least one first candidate target based on the matching degree between the acoustic micro-motion characteristics and the characteristics in the preset acoustic micro-motion characteristic library.

15. The method according to claim 14, characterized in that, The method further includes: The radar device is used to obtain a first distance and a first direction of the at least one candidate target relative to the radar device; The radar device determines the first spatial information based on the first distance and the first direction.

16. The method according to claim 15, characterized in that, The method further includes: The received sound signal is analyzed by the microphone device to obtain a second distance and a second direction of the at least one sound-emitting target relative to the microphone device; The second spatial information is determined by the microphone device based on the second distance and the second direction.

17. The method according to claim 16, characterized in that, The method further includes: The radar device transmits first spatial information of at least one candidate target to the microphone device; The sound signal is obtained by acquiring the signal based on the first spatial information of the at least one candidate target through the microphone device.

18. The method according to claim 17, characterized in that, The radar device is integrated into the microphone device; When there is only one radar device, the geometric center of the radar device coincides with that of the microphone device; When there are multiple radar devices, the distribution pattern of the multiple radar devices on the microphone device includes any one of the following: ring uniform distribution, linear uniform distribution, matrix uniform distribution, and fan-shaped uniform distribution.

19. The method according to claim 18, characterized in that, The microphone device can be any of the following forms: ceiling-mounted, wall-mounted, desktop, or front-facing.

20. A method for locating a speaker in a conference room, characterized in that, The method includes: Acquire the fourth spatial information of at least one fifth candidate target output by the radar device; By analyzing the captured images using a camera, facial motion features of at least one sixth candidate target can be obtained. Based on the matching degree between the facial motion features and the features in the preset vocal facial motion feature library, the camera determines at least one seventh candidate target from the at least one sixth candidate target; Obtain the fifth spatial information of at least one seventh candidate target output by the camera; Based on the fourth spatial information and the fifth spatial information, the first target speaker is determined; Based on the fourth and fifth spatial information corresponding to the first target speaker, the spatial information of the first target speaker is determined.

21. The method according to claim 20, characterized in that, The method further includes: The radar device transmits the fourth spatial information to the camera; The camera acquires the image based on the fourth spatial information.

22. A method for locating a speaker in a conference room, characterized in that, The method includes: Acquire the first spatial information of at least one candidate target output by the radar device; Acquire second spatial information of at least one sound-emitting target output by the microphone device; Based on the spatial matching relationship between the first spatial information and the second spatial information, the first spatial information corresponding to each sound-emitting target is determined; Based on the first spatial information and the second spatial information corresponding to each sound-emitting target, the spatial information of each sound-emitting target is determined.

23. A conference room speaker positioning system, characterized in that, The system includes a processing unit, a radar unit, and a microphone unit; The radar device is used to output first spatial information of at least one candidate target; The microphone device is used to output second spatial information of at least one sound-emitting target; The processing device is used to determine the target speaker based on the spatial matching relationship between the first spatial information and the second spatial information; The processing device is further configured to determine the spatial information of the target speaker based on the first spatial information and the second spatial information corresponding to the target speaker.

24. A conference room speaker positioning system, characterized in that, The system includes a processing unit, a radar unit, and a camera; The radar device is used to output fourth spatial information of at least one fifth candidate target; The camera is used to analyze the acquired images to obtain facial motion features of at least one sixth candidate target; The camera is also used to determine at least one seventh candidate target from the at least one sixth candidate target based on the matching degree between the facial motion features and the features in the preset vocal facial motion feature library; The processing device is used to acquire the fifth spatial information of at least one seventh candidate target output by the camera; The processing device is further configured to determine a first target speaker based on the fourth spatial information and the fifth spatial information; The processing device is further configured to determine the spatial information of the first target speaker based on the fourth and fifth spatial information corresponding to the first target speaker.

25. A conference room speaker positioning system, characterized in that, The system includes a processing unit, a radar unit, and a microphone unit; The radar device is used to output first spatial information of at least one candidate target; The microphone device is used to output second spatial information of at least one sound-emitting target; The processing device is used to determine the first spatial information corresponding to each sound-emitting target based on the spatial matching relationship between the first spatial information and the second spatial information; The processing device is further configured to determine the spatial information of each sound-emitting target based on the first spatial information and the second spatial information corresponding to each sound-emitting target.

26. A speaker positioning device for a conference room, characterized in that, The device includes a first acquisition module and a first processing module; The first acquisition module is used to acquire first spatial information of at least one candidate target output by the radar device; The first acquisition module is further configured to acquire second spatial information of at least one sound-emitting target output by the microphone device; The first processing module is used to determine the target speaker based on the spatial matching relationship between the first spatial information and the second spatial information; The first processing module is further configured to determine the spatial information of the target speaker based on the first spatial information and the second spatial information corresponding to the target speaker.

27. A speaker positioning device for a conference room, characterized in that, The device includes a second acquisition module and a second processing module; The second acquisition module is used to acquire the fourth spatial information of at least one fifth candidate target output by the radar device; The second processing module is used to analyze the acquired images through the camera to obtain facial motion features of at least one sixth candidate target; The second processing module is further configured to determine at least one seventh candidate target from the at least one sixth candidate target based on the matching degree between the facial motion features and the features in the preset vocal facial motion feature library by the camera; The second acquisition module is further configured to acquire the fifth spatial information of at least one seventh candidate target output by the camera; The second processing module is further configured to determine the first target speaker based on the fourth spatial information and the fifth spatial information; The second processing module is further configured to determine the spatial information of the first target speaker based on the fourth and fifth spatial information corresponding to the first target speaker.

28. A speaker positioning device for a conference room, characterized in that, The device includes a third acquisition module and a third processing module; The third acquisition module is used to acquire first spatial information of at least one candidate target output by the radar device; The third acquisition module is also used to acquire second spatial information of at least one sound-emitting target output by the microphone device; The third processing module is used to determine the first spatial information corresponding to each sound-emitting target based on the spatial matching relationship between the first spatial information and the second spatial information. The third processing module is also used to determine the spatial information of each sound-emitting target based on the first spatial information and the second spatial information corresponding to each sound-emitting target.

29. A computer-readable storage medium, characterized in that, The computer-readable medium stores a computer program that, when executed by a processor, is used to implement the method according to any one of claims 1-19, 20, 21, or 22.