Precise radio system and method based on double microphone arrays and 3D visual identification
The precise sound pickup system using dual microphone arrays and 3D visual recognition solves the problems of linear microphone arrays being unable to distinguish sound sources at the same angle and low accuracy of visual reconstruction. It achieves precise sound source localization and dynamic tracking, improving the microphone array's sound pickup capability and real-time performance in complex environments.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- BEIJING DIZHIYUAN TECHNOLOGY CO LTD
- Filing Date
- 2026-01-16
- Publication Date
- 2026-04-14
AI Technical Summary
In existing technologies, linear microphone arrays cannot distinguish sound sources at the same angle, are easily affected by environmental interference, resulting in low sound pickup accuracy, and have low precision in distance measurement for monocular vision reconstruction, which cannot meet the requirements for accurate sound pickup.
A precise sound recording system based on dual-microphone array and 3D visual recognition is adopted. Through the sound source localization module, visual recognition module, audio processing module and sound recording control module, combined with beamforming and triangulation algorithms, the sound source localization and speech signal extraction are realized. The distance and angle of the sound source are calculated by using 3D visual recognition, and audio and video synchronous tracking is performed by combining multi-instance comparative learning.
It improves the sound pickup accuracy in multi-sound-source and background noise scenarios, realizes accurate sound source modeling and dynamic tracking, enhances the robustness and real-time performance of the microphone array in complex environments, and meets the needs of real-time interaction.
Smart Images

Figure CN121865166A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of microphone sound pickup, specifically to a precise sound pickup system and method based on dual-microphone array and 3D visual recognition. Background Technology
[0002] A dual-microphone array is an audio acquisition device composed of multiple microphones, incorporating functions such as beamforming, noise suppression, echo cancellation, and sound source localization. It is used for directional sound pickup and to enhance the quality of the audio signal. Microphone arrays are low-cost, easy to integrate, and capable of sound source angle localization. With the widespread adoption of smart terminals, they are widely used in products such as digital human assistants, smart screens in conference rooms, and smart speakers. The working principle of a microphone array is to calculate the sound source angle by measuring the time or phase difference of the voice signal across different microphones, thereby achieving directional sound pickup through beamforming.
[0003] Linear microphone arrays can only acquire the angle of the sound source, but cannot detect the actual distance between the sound source and the array. When multiple sound sources are at the same detection angle, the array cannot distinguish between valid and interfering sound sources, resulting in irrelevant speech signals being mixed in during recording and misplaying the voice of non-target users. In environments such as open conference rooms and shopping malls, where there is a lot of echo and background noise, the angle positioning function of the microphone array is easily affected by environmental interference, reducing the accuracy of sound pickup, increasing the error rate of downstream data processing, and further reducing the efficiency of the microphone array.
[0004] Furthermore, reconstructing a three-dimensional sound field in space using a monocular vision device requires scanning the depth of the environment, which involves a large amount of computation. Moreover, the distance can only be estimated based on the image size, resulting in low measurement accuracy. When the target sound source is in a moving state, the sound pickup adjustment is not timely enough, failing to meet the requirements for accurate sound pickup. The multi-target interference problem in the process of visual sound field reconstruction is also very serious. Summary of the Invention
[0005] The purpose of this invention is to provide a precise sound reception system and method based on dual-microphone array and 3D visual recognition, so as to solve the problems mentioned in the background art.
[0006] To solve the above-mentioned technical problems, the present invention provides the following technical solution: a precise sound recording system based on dual-microphone array and 3D visual recognition, comprising: a sound source localization module, a visual recognition module, an audio processing module, a sound recording control module, and a precise tracking module; The sound source localization module is used to deploy a dual-microphone array and a structured light camera to cover the sound field area, collect voice signals from the target area, expand the coverage area through the dual-microphone array layout, achieve directional focusing through beamforming, adjust the beamforming parameters of the microphone array according to the target listening angle, perform sound source localization based on time delay, obtain the voice signal of the effective sound source, and output the voice signal through the audio interface. The visual recognition module is used to control the camera to collect depth images and visual data of the target area in real time, obtain visual feature information from the captured video sequence, filter the face position information within the effective interaction distance as the effective sound source, calculate the angle information and distance information of the effective sound source relative to the camera based on the pixel depth of the image, and calculate the target listening angle of the linear microphone array through the triangulation algorithm. The audio processing module is used to decompose the received signal into sub-bands when different effective sound sources are present, perform cross-correlation operation on the sub-band signals, determine the position of each microphone by the sound field frequency, use the position of the base array microphone as the initial point and the remaining microphone positions as sound pressure interpolation points to perform spline interpolation, increase the sampling frequency, and perform cross-correlation operation again after fusing the sub-bands according to frequency weight, signal-to-noise ratio estimation and coherence analysis. The TDOA time difference of arrival characteristics of the speech signals of each microphone are obtained through dual threshold endpoint detection. The sound control module is used to arrange all TDOA features into feature vectors, with the TDOA features at each effective sound source as cluster centers. The feature vectors are clustered, and the distance between the TDOA vectors of each cluster center and the feature vectors is used as the correlation coefficient. The cluster with the highest correlation coefficient is selected, and the effective sound source corresponding to the selected cluster is output as the target sound source. The sound source localization is corrected by generating virtual proximity reference points. The precise tracking module generates a sound power heatmap based on the sound source location, time difference of arrival, and volume ratio of each microphone. It optimizes the sound power heatmap based on the adjacency matrix, maps the audio power to the scene visual space, trains a convolutional recurrent neural network using multi-instance contrastive learning, learns facial motion features and audio spectrum features, encodes the temporal information of the sound, performs weighted fusion of the camera image sequence and the temporal information of the sound, aligns the target sound source location with the camera angle, measures the camera pose, tracks the sound source based on the pose increment, updates the global coordinate system and the sound source location, and enables the camera to track the target sound source.
[0007] Furthermore, the sound source localization module includes: a dual-microphone array unit and a beamforming unit; The dual-microphone array unit is used to fix two microphones to the main body of the sound receiving device, calibrate the spatial offset between the optical center of the camera and the center of the microphone array, and determine the spatial coordinate system transformation matrix and time difference between the camera and the microphone. The beamforming unit is used to adjust the receiving direction of the microphone array using beamforming technology based on the direction information of the sound source, so that the microphone array is focused on the target sound source, and the receiving beam of the microphone array is determined according to the position of the target sound source and the positioning angle of the microphone array.
[0008] Furthermore, the visual recognition module includes: an image processing unit and a depth positioning unit; The image processing unit is used to synchronously acquire RGB-D video streams from the structured light camera and, based on a preset effective interaction distance range, filter out faces, speakers, or microphones within the range as sound sources for output. The depth positioning unit is used to calculate the pixel depth value of the center of each sound source, calculate the three-dimensional coordinates of the sound source relative to the optical center of the camera using camera intrinsic parameters and coordinate system transformation matrix, transform the angle and position of the sound source in the camera coordinate system to the microphone array coordinate system, and lock the horizontal azimuth angle of the sound source in the array plane as the target listening angle.
[0009] Furthermore, the audio processing module includes: a sub-band decomposition unit, a TDOA estimation unit, and a dynamic focusing unit; The subband decomposition unit is used to perform endpoint detection on the microphone's speech signal. Within adjacent endpoints, the audio signal received by each microphone is decomposed into subband signals through a filter bank, and the frequency of the subband is used to calibrate the geometric position of the microphone array. The TDOA estimation unit is used to select a pair of microphones for each sub-band to calculate the cross-correlation function, perform a weighted average based on the signal-to-noise ratio and coherence of each sub-band to obtain the time difference of arrival, and estimate the sound pressure signal at the virtual microphone position by interpolation to increase the feature dimension of TDOA. The dynamic focusing unit is used to weight and vectorize the TDOA values of all sub-bands according to frequency weight, signal-to-noise ratio estimation, and coherence to form a TDOA feature vector.
[0010] Furthermore, the sound reception control module includes: a sound source clustering unit and a localization correction unit; The sound source clustering unit is used to cluster the TDOA feature vectors extracted between all endpoints. The number of clusters is set according to the number of sound sources. The theoretical TDOA vector corresponding to each effective sound source is calculated from the spatial location of each effective sound source. The correlation coefficient between the TDOA vector of each cluster center and each theoretical TDOA vector is calculated. The positioning and correction unit is used to select the effective sound source with the highest correlation coefficient as the target sound source, take the TDOA vector of the cluster center corresponding to the target sound source as the sound source direction, use virtual reference points and geometric constraints to perform consistency verification on the sound source direction, correct the sound source direction, and use beamforming algorithm to enhance the audio in the sound source direction.
[0011] Furthermore, the precise tracking module includes: an environmental perception unit, an audio-visual fusion unit, and a posture adjustment unit; The environmental perception unit is used to generate a sound power heatmap using sound field simulation software based on the corrected sound source location, TDOA, and microphone volume ratio. The value of each pixel in the heatmap represents the intensity of the sound generated at that location. The audio-visual fusion unit is used to map the sound source position to the image coordinate system using the camera's calibration parameters, extract facial motion features and input them into CNN training, output the position of the face in the sounding state, align the face center with the sound power heatmap, and align the sound source position. The attitude adjustment unit is used to calculate the offset between the target sound source coordinates and the camera image center coordinates, determine the rotation angle of the camera's calculation gimbal, obtain the pose increment between each frame of the camera through feature point matching, update the position of the sound source in the global coordinate system according to the pose increment, and make the camera aligned with the sound source.
[0012] The precise sound pickup method based on dual-microphone array and 3D visual recognition includes the following steps: Step S1. Deploy a dual-microphone array and a camera to cover the sound field area. The dual-microphone array acquires voice signals from the environment, and the camera captures depth images of the target area to obtain a video sequence. Step S2. Obtain visual features from the video sequence, filter faces within the effective interaction distance as effective sound sources, calculate the angle and distance information of the effective sound sources relative to the camera, determine the target listening angle of the linear microphone array through triangulation, and adjust the beamforming parameters of the microphone array according to the target listening angle. Step S3. When different effective sound sources exist, the received signal is decomposed into sub-bands, cross-correlation is performed on the sub-band signals, and cross-correlation is performed again after the sub-bands are fused according to frequency weight, signal-to-noise ratio estimation and coherence analysis to obtain the time difference of arrival characteristics of the speech signals of each microphone. Step S4. Arrange all sound arrival time difference features into feature vectors, cluster the feature vectors, use the distance between the TDOA vector of each cluster center and the feature vector as the correlation coefficient, select the cluster with the highest correlation coefficient, and output the effective sound source corresponding to the selected cluster as the target sound source. Step S5. Generate a sound power heatmap based on the target sound source location, sound arrival time difference, and volume ratio of each microphone. Map the audio power to the scene visual space, encode the temporal information of the sound, align the target sound source location with the camera angle, measure the camera pose, update the sound source location, and enable the camera to track the target sound source.
[0013] Furthermore, step S1 includes: Step S11. Fix the two microphones to the main body of the recording device, calibrate the spatial offset between the optical center of the camera and the center of the microphone array, and determine the spatial coordinate system transformation matrix and time difference between the camera and the microphone; Step S12. Using the direction information of the sound source, the receiving direction of the microphone array is adjusted by beamforming technology so that the microphone array is focused on the target sound source. The receiving beam of the microphone array is determined according to the location of the target sound source and the positioning angle of the microphone array.
[0014] Furthermore, step S2 includes: Step S21. Synchronously acquire RGB-D video stream from structured light camera, and filter out faces, speakers or microphones within the preset effective interaction distance range as sound source outputs; Step S22. Calculate the pixel depth value of the center of each sound source, calculate the three-dimensional coordinates of the sound source relative to the optical center of the camera using the camera intrinsic parameters and coordinate system transformation matrix, transform the angle and position of the sound source in the camera coordinate system to the microphone array coordinate system, and lock the horizontal azimuth angle of the sound source in the array plane as the target listening angle.
[0015] Furthermore, step S3 includes: Step S31. Perform endpoint detection on the microphone's speech signal. Within adjacent endpoints, decompose the audio signal received by each microphone into sub-band signals through a filter bank, and use the frequency of the sub-bands to calibrate the geometric position of the microphone array. Step S32. For each sub-band, select a pair of microphones to calculate the cross-correlation function. Perform a weighted average based on the signal-to-noise ratio and coherence of each sub-band to obtain the time difference of arrival. Estimate the sound pressure signal at the virtual microphone position through interpolation to increase the feature dimension of TDOA. Based on the frequency weight, signal-to-noise ratio estimation and coherence, weight the TDOA values of all sub-bands and vectorize them to form the TDOA feature vector.
[0016] Furthermore, step S4 includes: Step S41. Cluster the TDOA feature vectors extracted between all endpoints. The number of clusters is set according to the number of sound sources. Calculate the theoretical TDOA vector corresponding to each effective sound source based on the spatial location of each effective sound source. Calculate the correlation coefficient between the TDOA vector of each cluster center and each theoretical TDOA vector. Step S42. Select the effective sound source with the highest correlation coefficient as the target sound source, take the TDOA vector of the cluster center corresponding to the target sound source as the sound source direction, use virtual reference points and geometric constraints to perform consistency verification of the sound source direction, correct the sound source direction, and use beamforming algorithm to enhance the audio in the sound source direction.
[0017] Furthermore, step S5 includes: Step S51. Based on the corrected sound source location, TDOA and microphone volume ratio, generate a sound power heatmap using sound field simulation software. The value of each pixel in the heatmap represents the intensity of the sound produced at that location. Step S52. Using the camera's calibration parameters, map the sound source location to the image coordinate system, extract facial motion features and input them into the CNN for training, output the position of the face in the sounding state, align the face center with the sound power heatmap, and align the sound source location. Step S53. Calculate the offset between the target sound source coordinates and the camera image center coordinates, determine the rotation angle of the camera's calculation gimbal, obtain the pose increment between each frame of the camera through feature point matching, update the position of the sound source in the global coordinate system according to the pose increment, and make the camera aligned with the sound source.
[0018] Compared with the prior art, the beneficial effects achieved by the present invention are: 1. This invention, by deploying a dual-microphone array and a structured light camera, calculates the listening angle of the effective sound source, performs effective sound source localization, and acquires the voice signal. It solves the problem that traditional linear microphone arrays cannot distinguish sound sources at the same angle, improves anti-interference capability, and increases the sound pickup accuracy in multi-sound source and background noise scenarios. It is easy to integrate and adapt to standardized hardware, has high precision and real-time performance, and meets the needs of real-time interactive scenarios.
[0019] 2. This invention can acquire the time difference of arrival features of microphone speech signals when different effective sound sources exist, select the effective sound source corresponding to the cluster with the highest correlation coefficient as the target sound source, realize accurate modeling of sound source localization scene, effectively improve the accuracy of feature matching, solve the problems of target loss and multi-target interference, improve the sound reception capability of microphone array in complex scene, and has good robustness and accuracy.
[0020] 3. This invention generates a sound power heatmap, maps the audio power to the scene visual space, aligns the target sound source position with the camera angle, enables the camera to track the target sound source, performs audio and video synchronous tracking, realizes the visual analysis of the sound source signal, can achieve good positioning accuracy, improve the sound source positioning accuracy and dynamic tracking capability, and enhance the environmental adaptability and real-time performance of the microphone. Attached Figure Description
[0021] The accompanying drawings are provided to further illustrate the invention and form part of the specification. They are used in conjunction with embodiments of the invention to explain the invention and do not constitute a limitation thereof. In the drawings: Figure 1 This is a schematic diagram of the structure of the precision sound reception system based on dual-microphone array and 3D visual recognition of the present invention; Figure 2 This is a schematic diagram illustrating the steps of the precise sound reception method based on dual-microphone array and 3D visual recognition of the present invention. Detailed Implementation
[0022] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0023] Please see Figures 1 to 2 The present invention provides a technical solution: a precise sound recording system based on dual-microphone array and 3D visual recognition, comprising: a sound source localization module, a visual recognition module, an audio processing module, a sound recording control module, and a precise tracking module; The sound source localization module is used to deploy a dual-microphone array and a structured light camera to cover the sound field area, collect voice signals from the target area, expand the coverage area through the dual-microphone array layout, achieve directional focusing through beamforming, adjust the beamforming parameters of the microphone array according to the target listening angle, perform sound source localization based on time delay, obtain the voice signal of the effective sound source, and output the voice signal through the audio interface. The sound source localization module includes: a dual-microphone array unit and a beamforming unit; The dual-microphone array unit is used to fix two microphones to the main body of the sound receiving device, calibrate the spatial offset between the optical center of the camera and the center of the microphone array, and determine the spatial coordinate system transformation matrix and time difference between the camera and the microphone. The beamforming unit is used to adjust the receiving direction of the microphone array using beamforming technology based on the direction information of the sound source, so that the microphone array is focused on the target sound source, and the receiving beam of the microphone array is determined according to the position of the target sound source and the positioning angle of the microphone array.
[0024] The visual recognition module is used to control the camera to collect depth images and visual data of the target area in real time, obtain visual feature information from the captured video sequence, filter the face position information within the effective interaction distance as the effective sound source, calculate the angle information and distance information of the effective sound source relative to the camera based on the pixel depth of the image, and calculate the target listening angle of the linear microphone array through the triangulation algorithm. The visual recognition module includes: an image processing unit and a depth positioning unit; The image processing unit is used to synchronously acquire RGB-D video streams from the structured light camera and, based on a preset effective interaction distance range, filter out faces, speakers, or microphones within the range as sound sources for output. The depth positioning unit is used to calculate the pixel depth value of the center of each sound source, calculate the three-dimensional coordinates of the sound source relative to the optical center of the camera using camera intrinsic parameters and coordinate system transformation matrix, transform the angle and position of the sound source in the camera coordinate system to the microphone array coordinate system, and lock the horizontal azimuth angle of the sound source in the array plane as the target listening angle.
[0025] The audio processing module is used to decompose the received signal into sub-bands when different effective sound sources are present, perform cross-correlation operation on the sub-band signals, determine the position of each microphone by the sound field frequency, use the position of the base array microphone as the initial point and the remaining microphone positions as sound pressure interpolation points to perform spline interpolation, increase the sampling frequency, and perform cross-correlation operation again after fusing the sub-bands according to frequency weight, signal-to-noise ratio estimation and coherence analysis. The TDOA time difference of arrival characteristics of the speech signals of each microphone are obtained through dual threshold endpoint detection. The audio processing module includes: a sub-band decomposition unit, a TDOA estimation unit, and a dynamic focusing unit; The subband decomposition unit is used to perform endpoint detection on the microphone's speech signal. Within adjacent endpoints, the audio signal received by each microphone is decomposed into subband signals through a filter bank, and the frequency of the subband is used to calibrate the geometric position of the microphone array. The TDOA estimation unit is used to select a pair of microphones for each sub-band to calculate the cross-correlation function, perform a weighted average based on the signal-to-noise ratio and coherence of each sub-band to obtain the time difference of arrival, and estimate the sound pressure signal at the virtual microphone position by interpolation to increase the feature dimension of TDOA. The dynamic focusing unit is used to weight and vectorize the TDOA values of all sub-bands according to frequency weight, signal-to-noise ratio estimation, and coherence to form a TDOA feature vector.
[0026] The sound control module is used to arrange all TDOA features into feature vectors, with the TDOA features at each effective sound source as cluster centers. The feature vectors are clustered, and the distance between the TDOA vectors of each cluster center and the feature vectors is used as the correlation coefficient. The cluster with the highest correlation coefficient is selected, and the effective sound source corresponding to the selected cluster is output as the target sound source. The sound source localization is corrected by generating virtual proximity reference points. The sound reception control module includes: a sound source clustering unit and a localization correction unit; The sound source clustering unit is used to cluster the TDOA feature vectors extracted between all endpoints. The number of clusters is set according to the number of sound sources. The theoretical TDOA vector corresponding to each effective sound source is calculated from the spatial location of each effective sound source. The correlation coefficient between the TDOA vector of each cluster center and each theoretical TDOA vector is calculated. The positioning and correction unit is used to select the effective sound source with the highest correlation coefficient as the target sound source, take the TDOA vector of the cluster center corresponding to the target sound source as the sound source direction, use virtual reference points and geometric constraints to perform consistency verification on the sound source direction, correct the sound source direction, and use beamforming algorithm to enhance the audio in the sound source direction.
[0027] The precise tracking module generates a sound power heatmap based on the sound source location, time difference of arrival, and volume ratio of each microphone. It optimizes the sound power heatmap based on the adjacency matrix, maps the audio power to the scene visual space, trains a convolutional recurrent neural network using multi-instance contrastive learning, learns facial motion features and audio spectrum features, encodes the temporal information of the sound, performs weighted fusion of the camera image sequence and the temporal information of the sound, aligns the target sound source location with the camera angle, measures the camera pose, tracks the sound source based on the pose increment, updates the global coordinate system and the sound source location, and enables the camera to track the target sound source.
[0028] The precise tracking module includes: an environmental perception unit, an audio-visual fusion unit, and a posture adjustment unit; The environmental perception unit is used to generate a sound power heatmap using sound field simulation software based on the corrected sound source location, TDOA, and microphone volume ratio. The value of each pixel in the heatmap represents the intensity of the sound generated at that location. The audio-visual fusion unit is used to map the sound source position to the image coordinate system using the camera's calibration parameters, extract facial motion features and input them into CNN training, output the position of the face in the sounding state, align the face center with the sound power heatmap, and align the sound source position. The attitude adjustment unit is used to calculate the offset between the target sound source coordinates and the camera image center coordinates, determine the rotation angle of the camera's calculation gimbal, obtain the pose increment between each frame of the camera through feature point matching, update the position of the sound source in the global coordinate system according to the pose increment, and make the camera aligned with the sound source.
[0029] The precise sound pickup method based on dual-microphone array and 3D visual recognition includes the following steps: Step S1. Deploy a dual-microphone array and a camera to cover the sound field area. The dual-microphone array acquires voice signals from the environment, and the camera captures depth images of the target area to obtain a video sequence. Step S1 includes: Step S11. Fix the two microphones to the main body of the recording device, calibrate the spatial offset between the optical center of the camera and the center of the microphone array, and determine the spatial coordinate system transformation matrix and time difference between the camera and the microphone; Step S12. Using the direction information of the sound source, the receiving direction of the microphone array is adjusted by beamforming technology so that the microphone array is focused on the target sound source. The receiving beam of the microphone array is determined according to the location of the target sound source and the positioning angle of the microphone array.
[0030] Step S2. Obtain visual features from the video sequence, filter faces within the effective interaction distance as effective sound sources, calculate the angle and distance information of the effective sound sources relative to the camera, determine the target listening angle of the linear microphone array through triangulation, and adjust the beamforming parameters of the microphone array according to the target listening angle. Step S2 includes: Step S21. Synchronously acquire RGB-D video stream from structured light camera, and filter out faces, speakers or microphones within the preset effective interaction distance range as sound source outputs; Step S22. Calculate the pixel depth value of the center of each sound source, calculate the three-dimensional coordinates of the sound source relative to the optical center of the camera using the camera intrinsic parameters and coordinate system transformation matrix, transform the angle and position of the sound source in the camera coordinate system to the microphone array coordinate system, and lock the horizontal azimuth angle of the sound source in the array plane as the target listening angle.
[0031] Step S3. When different effective sound sources exist, the received signal is decomposed into sub-bands, cross-correlation is performed on the sub-band signals, and cross-correlation is performed again after the sub-bands are fused according to frequency weight, signal-to-noise ratio estimation and coherence analysis to obtain the time difference of arrival characteristics of the speech signals of each microphone. Step S3 includes: Step S31. Perform endpoint detection on the microphone's speech signal. Within adjacent endpoints, decompose the audio signal received by each microphone into sub-band signals through a filter bank, and use the frequency of the sub-bands to calibrate the geometric position of the microphone array. Step S32. For each sub-band, select a pair of microphones to calculate the cross-correlation function. Perform a weighted average based on the signal-to-noise ratio and coherence of each sub-band to obtain the time difference of arrival. Estimate the sound pressure signal at the virtual microphone position through interpolation to increase the feature dimension of TDOA. Based on the frequency weight, signal-to-noise ratio estimation and coherence, weight the TDOA values of all sub-bands and vectorize them to form the TDOA feature vector.
[0032] Step S4 includes: Step S41. Cluster the TDOA feature vectors extracted between all endpoints. The number of clusters is set according to the number of sound sources. Calculate the theoretical TDOA vector corresponding to each effective sound source based on the spatial location of each effective sound source. Calculate the correlation coefficient between the TDOA vector of each cluster center and each theoretical TDOA vector. Step S42. Select the effective sound source with the highest correlation coefficient as the target sound source, take the TDOA vector of the cluster center corresponding to the target sound source as the sound source direction, use virtual reference points and geometric constraints to perform consistency verification of the sound source direction, correct the sound source direction, and use beamforming algorithm to enhance the audio in the sound source direction.
[0033] Step S4. Arrange all sound arrival time difference features into feature vectors, cluster the feature vectors, use the distance between the TDOA vector of each cluster center and the feature vector as the correlation coefficient, select the cluster with the highest correlation coefficient, and output the effective sound source corresponding to the selected cluster as the target sound source. Step S5. Generate a sound power heatmap based on the target sound source location, sound arrival time difference, and volume ratio of each microphone. Map the audio power to the scene visual space, encode the temporal information of the sound, align the target sound source location with the camera angle, measure the camera pose, update the sound source location, and enable the camera to track the target sound source.
[0034] Step S5 includes: Step S51. Based on the corrected sound source location, TDOA and microphone volume ratio, generate a sound power heatmap using sound field simulation software. The value of each pixel in the heatmap represents the intensity of the sound produced at that location. Step S52. Using the camera's calibration parameters, map the sound source location to the image coordinate system, extract facial motion features and input them into the CNN for training, output the position of the face in the sounding state, align the face center with the sound power heatmap, and align the sound source location. Step S53. Calculate the offset between the target sound source coordinates and the camera image center coordinates, determine the rotation angle of the camera's calculation gimbal, obtain the pose increment between each frame of the camera through feature point matching, update the position of the sound source in the global coordinate system according to the pose increment, and make the camera aligned with the sound source.
[0035] Example: After the device is powered on, the coordinate parameters of the dual-microphone array and the 3D camera are automatically calibrated. The 3D camera begins to acquire depth images in real time, and the dual-microphone array enters standby recording mode. When participant A speaks while sitting 3m directly in front of him, the 3D camera acquires the depth image of participant A, determines the candidate region through face detection, and the VAD algorithm detects the voice activity, determining that A is a valid sound source. Calculate the three-dimensional parameters of sound source A: distance d = 3m, horizontal angle θ = 0°, vertical angle φ = 0°. Using a triangulation algorithm, calculate the target listening angle α1 = 11.3° for the left microphone array A and the target listening angle α2 = -11.3° for the right microphone array B. Send control commands α1 = 11.3° and α2 = -11.3° to the sound control module to adjust the beamforming main lobe of the left microphone array to 11.3° and the right microphone array to -11.3°. Set the sidelobe suppression ratio to 28dB. Perform noise reduction processing on the acquired speech signal of A and output it to the playback system.
[0036] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus.
[0037] Finally, it should be noted that the above descriptions are merely preferred embodiments of the present invention and are not intended to limit the present invention. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing embodiments or make equivalent substitutions for some of the technical features. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.
Claims
1. A precise sound pickup method based on dual-microphone array and 3D visual recognition, characterized in that, The method includes the following steps: Step S1. Deploy a dual-microphone array and a camera to cover the sound field area. The dual-microphone array acquires voice signals from the environment, and the camera captures depth images of the target area to obtain a video sequence. Step S2. Obtain visual features from the video sequence, filter faces within the effective interaction distance as effective sound sources, calculate the angle and distance information of the effective sound sources relative to the camera, determine the target listening angle of the linear microphone array through triangulation, and adjust the beamforming parameters of the microphone array according to the target listening angle. Step S3. When different effective sound sources exist, the received signal is decomposed into sub-bands, cross-correlation is performed on the sub-band signals, and cross-correlation is performed again after the sub-bands are fused according to frequency weight, signal-to-noise ratio estimation and coherence analysis to obtain the time difference of arrival characteristics of the speech signals of each microphone. Step S4. Arrange all sound arrival time difference features into feature vectors, cluster the feature vectors, use the distance between the TDOA vector of each cluster center and the feature vector as the correlation coefficient, select the cluster with the highest correlation coefficient, and output the effective sound source corresponding to the selected cluster as the target sound source. Step S5. Generate a sound power heatmap based on the target sound source location, sound arrival time difference, and volume ratio of each microphone. Map the audio power to the scene visual space, encode the temporal information of the sound, align the target sound source location with the camera angle, measure the camera pose, update the sound source location, and enable the camera to track the target sound source.
2. The precise sound pickup method based on dual-microphone array and 3D visual recognition according to claim 1, characterized in that: Step S1 includes: Step S11. Fix the two microphones to the main body of the recording device, calibrate the spatial offset between the optical center of the camera and the center of the microphone array, and determine the spatial coordinate system transformation matrix and time difference between the camera and the microphone; Step S12. Using the direction information of the sound source, the receiving direction of the microphone array is adjusted by beamforming technology so that the microphone array is focused on the target sound source. The receiving beam of the microphone array is determined according to the position of the target sound source and the positioning angle of the microphone array. Step S2 includes: Step S21. Synchronously acquire RGB-D video stream from structured light camera, and filter out faces, speakers or microphones within the preset effective interaction distance range as sound source outputs; Step S22. Calculate the pixel depth value of the center of each sound source, calculate the three-dimensional coordinates of the sound source relative to the optical center of the camera using the camera intrinsic parameters and coordinate system transformation matrix, transform the angle and position of the sound source in the camera coordinate system to the microphone array coordinate system, and lock the horizontal azimuth angle of the sound source in the array plane as the target listening angle.
3. The precise sound recording method based on dual-microphone array and 3D visual recognition according to claim 2, characterized in that: Step S3 includes: Step S31. Perform endpoint detection on the microphone's speech signal. Within adjacent endpoints, decompose the audio signal received by each microphone into sub-band signals through a filter bank, and use the frequency of the sub-bands to calibrate the geometric position of the microphone array. Step S32. For each sub-band, select a pair of microphones to calculate the cross-correlation function. Perform a weighted average based on the signal-to-noise ratio and coherence of each sub-band to obtain the time difference of arrival. Estimate the sound pressure signal at the virtual microphone position through interpolation to increase the feature dimension of TDOA. Based on the frequency weight, signal-to-noise ratio estimation and coherence, weight the TDOA values of all sub-bands and vectorize them to form the TDOA feature vector.
4. The precise sound pickup method based on dual-microphone array and 3D visual recognition according to claim 3, characterized in that: Step S4 includes: Step S41. Cluster the TDOA feature vectors extracted between all endpoints. The number of clusters is set according to the number of sound sources. Calculate the theoretical TDOA vector corresponding to each effective sound source based on the spatial location of each effective sound source. Calculate the correlation coefficient between the TDOA vector of each cluster center and each theoretical TDOA vector. Step S42. Select the effective sound source with the highest correlation coefficient as the target sound source, take the TDOA vector of the cluster center corresponding to the target sound source as the sound source direction, use virtual reference points and geometric constraints to perform consistency verification of the sound source direction, correct the sound source direction, and use beamforming algorithm to enhance the audio in the sound source direction.
5. The precise sound pickup method based on dual-microphone array and 3D visual recognition according to claim 4, characterized in that: Step S5 includes: Step S51. Based on the corrected sound source location, TDOA and microphone volume ratio, generate a sound power heatmap using sound field simulation software. The value of each pixel in the heatmap represents the intensity of the sound produced at that location. Step S52. Using the camera's calibration parameters, map the sound source location to the image coordinate system, extract facial motion features and input them into the CNN for training, output the position of the face in the sounding state, align the face center with the sound power heatmap, and align the sound source location. Step S53. Calculate the offset between the target sound source coordinates and the camera image center coordinates, determine the rotation angle of the camera's calculation gimbal, obtain the pose increment between each frame of the camera through feature point matching, update the position of the sound source in the global coordinate system according to the pose increment, and make the camera aligned with the sound source.
6. A precise sound reception system based on dual-microphone array and 3D visual recognition, characterized in that, The system includes the following modules: Sound source localization module, visual recognition module, audio processing module, radio control module, and precision tracking module; The sound source localization module is used to deploy a dual-microphone array and a structured light camera to cover the sound field area, collect voice signals from the target area, expand the coverage area through the dual-microphone array layout, achieve directional focusing through beamforming, adjust the beamforming parameters of the microphone array according to the target listening angle, perform sound source localization based on time delay, obtain the voice signal of the effective sound source, and output the voice signal through the audio interface. The visual recognition module is used to control the camera to collect depth images and visual data of the target area in real time, obtain visual feature information from the captured video sequence, filter the face position information within the effective interaction distance as the effective sound source, calculate the angle information and distance information of the effective sound source relative to the camera based on the pixel depth of the image, and calculate the target listening angle of the linear microphone array through the triangulation algorithm. The audio processing module is used to decompose the received signal into sub-bands when different effective sound sources are present, perform cross-correlation operation on the sub-band signals, determine the position of each microphone by the sound field frequency, use the position of the base array microphone as the initial point and the remaining microphone positions as sound pressure interpolation points to perform spline interpolation, increase the sampling frequency, and perform cross-correlation operation again after fusing the sub-bands according to frequency weight, signal-to-noise ratio estimation and coherence analysis. The TDOA time difference of arrival characteristics of the speech signals of each microphone are obtained through dual threshold endpoint detection. The sound control module is used to arrange all TDOA features into feature vectors, with the TDOA features at each effective sound source as cluster centers. The feature vectors are clustered, and the distance between the TDOA vectors of each cluster center and the feature vectors is used as the correlation coefficient. The cluster with the highest correlation coefficient is selected, and the effective sound source corresponding to the selected cluster is output as the target sound source. The sound source localization is corrected by generating virtual proximity reference points. The precise tracking module generates a sound power heatmap based on the sound source location, time difference of arrival, and volume ratio of each microphone. It optimizes the sound power heatmap based on the adjacency matrix, maps the audio power to the scene visual space, trains a convolutional recurrent neural network using multi-instance contrastive learning, learns facial motion features and audio spectrum features, encodes the temporal information of the sound, performs weighted fusion of the camera image sequence and the temporal information of the sound, aligns the target sound source location with the camera angle, measures the camera pose, tracks the sound source based on the pose increment, updates the global coordinate system and the sound source location, and enables the camera to track the target sound source.
7. The precise sound reception system based on dual-microphone array and 3D visual recognition according to claim 6, characterized in that: The sound source localization module includes: a dual-microphone array unit and a beamforming unit; The dual-microphone array unit is used to fix two microphones to the main body of the sound receiving device, calibrate the spatial offset between the optical center of the camera and the center of the microphone array, and determine the spatial coordinate system transformation matrix and time difference between the camera and the microphone. The beamforming unit is used to adjust the receiving direction of the microphone array using beamforming technology based on the direction information of the sound source, so that the microphone array is focused on the target sound source, and the receiving beam of the microphone array is determined according to the position of the target sound source and the positioning angle of the microphone array. The visual recognition module includes: an image processing unit and a depth positioning unit; The image processing unit is used to synchronously acquire RGB-D video streams from the structured light camera and, based on a preset effective interaction distance range, filter out faces, speakers, or microphones within the range as sound sources for output. The depth positioning unit is used to calculate the pixel depth value of the center of each sound source, calculate the three-dimensional coordinates of the sound source relative to the optical center of the camera using camera intrinsic parameters and coordinate system transformation matrix, transform the angle and position of the sound source in the camera coordinate system to the microphone array coordinate system, and lock the horizontal azimuth angle of the sound source in the array plane as the target listening angle.
8. The precise sound reception system based on dual-microphone array and 3D visual recognition according to claim 7, characterized in that: The audio processing module includes: a sub-band decomposition unit, a TDOA estimation unit, and a dynamic focusing unit; The subband decomposition unit is used to perform endpoint detection on the microphone's speech signal. Within adjacent endpoints, the audio signal received by each microphone is decomposed into subband signals through a filter bank, and the frequency of the subband is used to calibrate the geometric position of the microphone array. The TDOA estimation unit is used to select a pair of microphones for each sub-band to calculate the cross-correlation function, perform a weighted average based on the signal-to-noise ratio and coherence of each sub-band to obtain the time difference of arrival, and estimate the sound pressure signal at the virtual microphone position by interpolation to increase the feature dimension of TDOA. The dynamic focusing unit is used to weight and vectorize the TDOA values of all sub-bands according to frequency weight, signal-to-noise ratio estimation, and coherence to form a TDOA feature vector.
9. The precise sound reception system based on dual-microphone array and 3D visual recognition according to claim 8, characterized in that: The sound reception control module includes: a sound source clustering unit and a localization correction unit; The sound source clustering unit is used to cluster the TDOA feature vectors extracted between all endpoints. The number of clusters is set according to the number of sound sources. The theoretical TDOA vector corresponding to each effective sound source is calculated from the spatial location of each effective sound source. The correlation coefficient between the TDOA vector of each cluster center and each theoretical TDOA vector is calculated. The positioning and correction unit is used to select the effective sound source with the highest correlation coefficient as the target sound source, take the TDOA vector of the cluster center corresponding to the target sound source as the sound source direction, use virtual reference points and geometric constraints to perform consistency verification on the sound source direction, correct the sound source direction, and use beamforming algorithm to enhance the audio in the sound source direction.
10. The precise sound reception system based on dual-microphone array and 3D visual recognition according to claim 9, characterized in that: The precise tracking module includes: an environmental perception unit, an audio-visual fusion unit, and a posture adjustment unit; The environmental perception unit is used to generate a sound power heatmap using sound field simulation software based on the corrected sound source location, TDOA, and microphone volume ratio. The value of each pixel in the heatmap represents the intensity of the sound generated at that location. The audio-visual fusion unit is used to map the sound source position to the image coordinate system using the camera's calibration parameters, extract facial motion features and input them into CNN training, output the position of the face in the sounding state, align the face center with the sound power heatmap, and align the sound source position. The attitude adjustment unit is used to calculate the offset between the target sound source coordinates and the camera image center coordinates, determine the rotation angle of the camera's calculation gimbal, obtain the pose increment between each frame of the camera through feature point matching, update the position of the sound source in the global coordinate system according to the pose increment, and make the camera aligned with the sound source.
Citation Information
Patent Citations
Signal-enhancing beamforming in augmented reality environment
CN104106267A
Method and apparatus for localizing target sound source
CN105467364A
Sound source positioning method based on dual-microphone array
CN109239667A
Association rule method based on KMeans algorithm
CN111222573A
Acoustic imaging method
CN112017688A
Cited By
Early warning method and device for disease infection risk
CN122091264A