Method for Sound Source Direction Localization and Directional Sound Enhancement Based on Microphone Array and Camera

By combining the multimodal data of the microphone array and camera, using depth detection and mouth motion analysis, dynamically adjusting the sound filter, solving the problems of inaccurate positioning and noise interference in complex environments by traditional technology, and achieving accurate sound enhancement and positioning.

CN119738779BActive Publication Date: 2025-07-22珠海金智维人工智能股份有限公司
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202411784897.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-06
Publication Date
2025-07-22
Estimated Expiration
2044-12-06

AI Technical Summary

Technical Problem

Traditional sound positioning and enhancement technologies are difficult to meet the accuracy requirements in complex acoustic environments, especially in multi-person voices and high-noise scenarios, where there are problems such as inaccurate positioning, serious noise interference, and insufficient filter adjustment.

Method used

Combining the microphone array and camera, through object detection, depth estimation and mouth motion analysis, the gain and frequency response of the sound filter is dynamically adjusted to achieve precise positioning and enhancement.

Benefits of technology

It realizes accurate screening of close-range vocalists and improves the clarity of sound signals in multiple people and high noise environments, reduces background noise interference, and improves the intelligence and real-timeness of sound processing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119738779B_ABST
    Figure CN119738779B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for sound source direction localization and directional sound enhancement based on a microphone array and a camera. All the people in the scene are detected by a target detection model (such as YOLO), and corresponding person bounding boxes are generated. Then, a depth detection module (Depth Anything model) is used to calculate the depth information of each pixel, and these depth information are matched with the detected person bounding boxes. The speaker closest to the camera is screened out through average depth calculation, and only the effective targets at close range are retained. The present invention can solve the problems of inaccurate sound source localization and poor sound enhancement effect in complex environments. By capturing images in real time through the camera and combining with the sound data of the microphone array, the system can accurately locate and enhance the target sound source, especially showing excellent performance in multi-source and high-noise environments.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of sound localization and enhancement, and particularly to a method for sound source direction localization and directional sound enhancement based on a microphone array and a camera. Background Art

[0002] With the rapid development of artificial intelligence (AI), the Internet of Things (IoT), and multimedia technologies, higher requirements for accurate sound capture and processing have been put forward in fields such as speech recognition, smart home, and video conferencing. Especially in complex acoustic environments, such as scenes with high background noise or multiple people speaking simultaneously, traditional sound localization and enhancement technologies are difficult to meet the actual needs. In this context, multi-modal fusion technologies that combine visual information (camera data) and audio information (microphone array data) have become a research hotspot.

[0003] Traditional methods mainly rely on a microphone array to locate the sound source, and judge the direction of the sound source by analyzing the phase difference and time delay of the sound signal. However, in the face of complex environments, this method has obvious limitations. Therefore, in recent years, the research trend has begun to attempt to combine cameras and microphone arrays to improve the accuracy of sound localization and enhancement through visual assistance technologies.

[0004] "CN112015364B - Method and Device for Adjusting Pickup Sensitivity" provides a technology for sound enhancement by combining a microphone array and a camera, which relies on the camera to track the movement of the lips to assist the microphone array in adjusting the pickup sensitivity. However, this technology has several obvious limitations:

[0005] Rough localization: This technology relies on the camera to identify the lip movement of the speaker as the basis for adjusting the microphone sensitivity. However, in a complex multi-person environment, especially when the background is relatively complex, it is difficult to accurately distinguish the position of the speaker, and it may mis-enhance irrelevant sound sources.

[0006] Lack of accurate depth information: This technology does not use depth detection to filter out irrelevant sound sources at a distance, but completely relies on the preliminary localization of the camera image and the microphone array. This method is prone to errors in a multi-person scenario and lacks effective discrimination of speakers at a distance.

[0007] Unable to dynamically adjust the filter: For sound enhancement, this patent only uses a fixed method to adjust the pickup sensitivity, lacking real-time analysis and adaptive adjustment of the sound characteristics.

[0008] "CN115862682B - Sound Detection Method and Related Equipment" adopts multi-modal fusion of audio and video features, comprehensively judges the sound source position through visual and audio data, and then enhances the target sound. However, this method also has limitations:

[0009] Imprecise positioning in a multi-source environment: This technology relies solely on the fusion of visual and audio features to determine the sound source. In a scenario where multiple people are speaking, it is difficult to accurately distinguish which are the main speakers, especially when there is confusion easily among multiple targets at relatively close distances.

[0010] Poor real-time performance: The process of multi-modal data fusion is complex and requires a large amount of computing resources, resulting in poor performance in application scenarios with high real-time requirements. Summary of the Invention

[0011] The object of the present invention is to overcome the shortcomings and deficiencies of the prior art and provide a sound source direction positioning and directional sound enhancement method based on a microphone array and a camera.

[0012] The object of the present invention is achieved by the following technical solutions:

[0013] A sound source direction positioning and directional sound enhancement method based on a microphone array and a camera, comprising the following steps:

[0014] S1. Install the camera at the central position of the microphone array and preprocess the images collected by the camera; detect all the people in the images through a target detection model and generate corresponding human bounding boxes;

[0015] S2. Use a depth estimation module to calculate the depth information of each pixel in the image and match the depth information of the pixels with the detected human bounding boxes;

[0016] By calculating the average depth of the pixels inside each human bounding box, judge the distance between the speaker and the camera, and filter out the targets at a relatively far distance by setting a depth threshold, retaining the effective human bounding boxes of those at a relatively close distance and likely to be speakers;

[0017] S3. Perform human key point analysis on the effective human bounding boxes and extract multiple human key points including the mouth, eyes, nose, etc.;

[0018] S4. Analyze the mouth movement state of the mouth key point: detect the degree of mouth opening and closing representing the sound intensity of the speaker, and detect the lip movement speed representing the speaking frequency and speed of the speaker;

[0019] The gain of the sound filter increases as the degree of mouth opening and closing of the speaker increases, and the sound signal intensity increases accordingly;

[0020] The frequency response of the sound filter improves as the mouth movement speed of the speaker becomes faster;

[0021] S5. Calculate the angle of the speaker's mouth relative to the center point of the camera, thereby identifying the orientation of the speaker, and adjust the gain of the microphone in the corresponding direction according to the layout of the microphone array.

[0022] The depth estimation module generates a depth map for each image. The depth map D(x, y) represents the distance between the pixel at each position (x, y) in the image and the camera; the depth map has the same coordinate system and resolution as the input image, and the depth information of the pixel is matched with the detected human bounding box through coordinate matching.

[0023] Match the depth information of the pixel with the detected human bounding box. The specific steps are as follows:

[0024] (1) Traverse each human bounding box: In the set of human bounding boxes B = {B1, B2,..., B n} detected by the object detection model, process each human bounding box B i one by one. The boundary coordinates of each human bounding box B i are (x i1 , x i2 , y i1 , y i2 ), indicating the position range of the target person in the image; where x i1 , y i1 are the abscissa and ordinate of the upper left corner of the human bounding box B i respectively, and x i2 , y i2 are the abscissa and ordinate of the lower right corner of the human bounding box B i respectively;

[0025] (2) Find the corresponding coordinate area from the depth map of the image: For each human bounding box B i , find the range of pixel points corresponding to the human bounding box from the depth map D(x, y) generated by the depth estimation model. The coordinates of this range are the same as the boundary coordinates of the human bounding box, that is, x i1 ≤ x ≤ x i2 and y i1 ≤ y ≤ y i2 , where x and y are the abscissa and ordinate of the pixel point respectively;

[0026] (3) Crop the corresponding depth map area: Crop the depth information part corresponding to the human bounding box B i from the depth map D(x, y). The cropped area contains the depth values of all pixel points within the human bounding box, that is, the depth information sub-map of the target person

[0027] The cropping formula is:

[0028] Among them, is the one corresponding to the human bounding box Bi The matching depth map sub-region, representing the corresponding region of the person in the depth map;

[0029] (4) Collect the depth information within the person's bounding box: Through the cropping operation, extract the set of depth values of all pixel points corresponding to the person's bounding box region from the depth map. This set represents the depth information of the person in the three-dimensional space and is used for subsequent average depth calculation.

[0030] The average depth of the pixels inside the person's bounding box is calculated as follows:

[0031]

[0032] where D avg (B i ) is the average depth value of the person's bounding box B, N is the total number of pixels inside the person's bounding box; D(x, y) is the depth value of each pixel point within the person's bounding box region, taken from the cropped depth map sub-region D i (x, y). Bi (x, y).

[0033] The degree of mouth opening and closing of the speaker is calculated by the following method:

[0034] Assume that the set of upper lip key points is {(x u1 , y u1 ), (x u2 , y u2 ),...(x un , y un )}, and the set of lower lip key points is {(x l1 , y l1 ), (x l2 , y l2 ),...(x ln , y ln )}. Each pair of upper and lower lip key points corresponds one by one;

[0035] Calculate the vertical distance between each pair of upper and lower lip key points, and then find the average of these distances;

[0036]

[0037] where n is the number of upper and lower lip key points; (x ui , y ui ) is the coordinate of the i-th upper lip key point;

[0038] (x li , y li ) is the coordinate of the i-th lower lip key point; d open is the average degree of mouth opening and closing.

[0039] The lip movement speed of the speaker is calculated as follows:

[0040]

[0041] where v(t) is the movement speed of the lips at time point t, d open (t), d open (t - 1) are the mouth opening degrees at time points t and t - 1 respectively, and Δt is the time interval.

[0042] Adjust the gain G(t) of the sound filter according to the mouth opening degree of the speaker:

[0043] G(t) = G0 + k·d open ;

[0044] where G0 is the base gain, k is the gain adjustment coefficient, and d open is the mouth opening degree.

[0045] Adjust the frequency response f(t) of the filter according to the lip movement speed of the speaker:

[0046] f(t) = f0 + k f ·v(t);

[0047] where f0 is the base frequency, k f is the frequency adjustment coefficient, and v(t) is the lip movement speed.

[0048] The specific implementation steps of step S5 are as follows:

[0049] (1) Precise azimuth calculation: Obtain the lip key point coordinates P lips = (x lips , y lips ) and the image center point (x center , y center ) through the camera, and calculate the angle of the lips relative to the camera center point, that is, the azimuth angle θ of the lips:

[0050]

[0051] (2) Adjust the gain of the microphone array based on the azimuth angle: For each microphone i, we set its center angle to θ, then the gain adjustment can be calculated by the following formula:

[0052] G i (t) = G0 + k·max(0, cos(θ - θ i ));

[0053] where G iG(i)(t) is the gain value of the i-th microphone at time t, G0 is the base gain, θ i is the central angle of the i-th microphone, and k is the adjustment coefficient of the gain;

[0054] By using the cosine function cos(θ - θ i ), to represent the relationship between the lip orientation and the central angle of each microphone, the system can dynamically adjust the gain of each microphone;

[0055] When the lip azimuth angle θ is close to the central angle θ i of a certain microphone, the gain G i will be maximized; other microphones that are farther away will receive smaller gains.

[0056] The microphone array will also perform distributed gain adjustment: the gain distribution based on the lip azimuth angle is smoothly transitioned among multiple microphones; when the lip position is in the middle of the coverage areas of two microphones, the gains of the two adjacent microphones will be adjusted accordingly to enhance the sound signal in this direction; for microphones that are far away from the lip azimuth, the gain will be reduced or even suppressed.

[0057] Meanwhile, the present invention provides:

[0058] A server, the server includes a processor and a memory, and at least one program is stored in the memory, and the program is loaded and executed by the processor to implement the above-mentioned sound source direction localization and directional sound enhancement method based on the microphone array and the camera.

[0059] A computer-readable storage medium, at least one program is stored in the storage medium, and the program is loaded and executed by a processor to implement the above-mentioned sound source direction localization and directional sound enhancement method based on the microphone array and the camera.

[0060] Compared with the prior art, the present invention has the following advantages and beneficial effects:

[0061] 1. The present invention can accurately screen the speaker: by combining depth detection, it can effectively filter out speakers who are far away or irrelevant. This process enables the system to focus on important sound sources at close range and greatly reduces the interference of distant noise.

[0062] 2. The present invention can improve the positioning accuracy: different from the prior art, the present invention uses depth information to accurately determine the distance between each person's frame and the camera, thereby accurately positioning each speaker and avoiding errors that may be caused by simply relying on images or microphone signals.

[0063] 3. The present invention can reduce background noise interference: Through in-depth detection, the system can accurately lock the speaker with a relatively close distance in a multi-person scenario. Even in a noisy environment, it can still effectively shield irrelevant noises.

[0064] 4. The mouth key point detection and dynamic filter adjustment of the present invention can improve the intelligence of sound processing. After completing the effective target screening, the present invention carefully analyzes the mouth shape of the speaker through the human key point detection model (DWPose) to obtain key information such as the degree of mouth opening and closing and the lip movement speed. Through this information, the gain and frequency response of the sound filter are adjusted in real time, thereby optimizing the enhancement effect of the sound signal.

[0065] 5. The present invention can achieve adaptive filter adjustment: Through the degree of mouth opening and closing and the lip movement speed, the system can dynamically adjust the sound filter. When the degree of opening is large, the gain is enhanced to ensure a higher sound signal intensity; when the degree of opening is small, the gain is weakened to suppress irrelevant noises. This adaptive adjustment method significantly improves the clarity of the sound signal.

[0066] 6. The present invention can accurately capture sound characteristics: The dynamic analysis of the mouth key points is not only used to judge whether a sound is being made, but also can precisely adjust the frequency response of the filter according to the subtle changes in the mouth shape, enabling the system to respond in real time to the changes in the speaker's speech rate and pitch, and improving the intelligence level of sound enhancement.

[0067] 7. The filter optimization of the present invention is consistent with the sound source direction: The change in the position of the corners of the mouth can help the system dynamically adjust the direction gain of the microphone array, so that the enhancement direction is consistent with the actual position of the speaker, further improving the accuracy of sound source localization. BRIEF DESCRIPTION OF THE DRAWINGS

[0068] Figure 1 It is a schematic structural diagram of a sound source direction localization and directional sound enhancement system based on a microphone array and a camera.

[0069] Figure 2 It is a flowchart of a sound source direction localization and directional sound enhancement method based on a microphone array and a camera. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0070] The following further describes the present invention in detail in conjunction with embodiments and the accompanying drawings, but the embodiments of the present invention are not limited thereto.

[0071] The present invention proposes a sound source direction localization and directional sound enhancement system based on a microphone array and a camera, aiming to solve the problems of inaccurate sound source localization and poor sound enhancement effect in complex environments. By capturing images in real time through the camera and combining with the sound data of the microphone array, the system can accurately locate and enhance the target sound source, especially showing excellent performance in multi-source and high-noise environments. To make the technical solution clearer, the following will describe each module of the present invention and its implementation principle in detail in combination with the system architecture diagram.

[0072] 3.1 System overall architecture

[0073] The overall architecture of the system is as Figure 1 shown. The system is mainly divided into three levels: the basic layer, the processing layer, and the output layer. Each level contains multiple key modules, and each module works together to ensure that the system can achieve efficient and accurate sound source localization and sound enhancement in complex environments.

[0074] 3.1.1 Basic layer

[0075] The basic layer is the underlying support part of the system, including two parts: basic devices and infrastructure:

[0076] 3.1.1.1 Basic devices

[0077] Camera: Used to capture real-time images in the scene. The position of the camera is precisely calculated and installed at the center of the microphone array to ensure the spatial consistency of the image and sound data.

[0078] Microphone array: The system uses an array composed of four directional microphones, arranged in a specific geometric layout (ring-shaped) to optimize the accuracy and coverage of sound source localization. The microphone array is not only used to capture sound signals in the environment but also determines the approximate direction of the sound source through a preliminary sound source localization algorithm.

[0079] Edge computing device: As the core processing platform of the system, the edge computing device is responsible for tasks such as image processing, sound signal processing, data synchronization, and fusion. The introduction of edge computing not only improves the real-time performance of the system but also reduces the dependence on cloud computing resources, meeting the deployment requirements in resource-constrained environments.

[0080] 3.1.1.2 Infrastructure

[0081] High-performance computing hardware, such as NVIDIA graphics cards, etc., is used to process image and sound data. The efficient operation of the computing part is the basis for ensuring the real-time performance and accuracy of the system.

[0082] The system accelerates the inference process of AI models through inference engines such as ONNX and TensorRT. Especially in image processing and speech recognition, the inference acceleration significantly reduces latency and improves the response speed of the system. To ensure the stability and scalability of the system, containerization technologies such as Docker are adopted in the deployment part. Containerization technologies enable each module of the system to be developed, tested, and deployed independently, making it easy to maintain and upgrade.

[0083] 3.1.2 Processing Layer

[0084] The processing layer is the core functional part of the system, responsible for processing the input image and sound data to achieve sound source localization and sound enhancement. The processing layer is divided into two main modules: image processing and audio processing:

[0085] 3.1.2.1 Image Processing

[0086] Object Detection: The image processing module uses the images captured by the camera for object detection. The system uses a deep learning model (YOLO) to quickly identify and mark people or objects in the image. The object detection module can accurately locate the possible sound source positions in the scene and generate corresponding bounding box data for subsequent processing. Depth Detection: The system uses the depth detection model of DepthAnything to calculate the distance information between the detected objects and the camera. The depth detection module can filter out distant background objects and only retain the objects that are closer and may be the sound source, so as to improve the processing efficiency and accuracy of the system.

[0087] Human Keypoint Detection: Based on object detection and depth detection, the system further performs human keypoint recognition on the detected objects, especially the recognition of important feature points such as the mouth, eyes, and nose.

[0088] By detecting the mouth shape, the system can determine whether the detected object is speaking, assisting the sound source localization module to improve the localization accuracy.

[0089] 3.1.2.2 Audio Processing

[0090] Directional Sound Pickup: The microphone array performs preliminary sound source localization by analyzing the phase difference and time delay of sound signals. The directional sound pickup module will enhance the sound in a specific direction according to the localization result, while suppressing the background noise in other directions. This process can effectively improve the signal-to-noise ratio of the sound, making the target sound source more clearly distinguishable.

[0091] Speech Recognition: Based on directional sound pickup, the system further processes the enhanced sound signal through speech recognition technology to recognize the speech content. This function is not only used to provide real-time feedback but also can provide data support for subsequent advanced applications (such as voice commands, dialogue systems).

[0092] 3.1.3 Output Layer

[0093] The output layer is responsible for the final output of the processed data and provides support for downstream applications or devices.

[0094] The main function of this layer is to ensure the quality and real-time performance of the processing results, specifically including:

[0095] Output of processing results: The edge computing device outputs the processed sound signal and image recognition results to the output module of the system in real time. The output module is responsible for transmitting the enhanced sound signal to the user side to ensure that the finally output sound has the characteristics of high quality and low noise. This module can also feedback the processing results to other application systems to meet the requirements of different scenarios.

[0096] 3.2 Image Processing Process

[0097] The image processing process is as Figure 2 shown and is a key part of the present invention for analyzing and processing the image data captured from the camera. Its main functions include object detection, human key point detection, depth detection, and lip detection and motion analysis. The implementation methods and technical details of each technical link are described in detail below.

[0098] 3.2.1 Image Capture and Preprocessing Image Capture: The camera module is installed at the center of the microphone array to ensure the spatial consistency of the image and sound data. The camera has high resolution and good low-light imaging ability, can capture the images in the environment in real time, and transmit the data to the image processing module at a high frame rate.

[0099] Image Preprocessing: Before the image enters the processing link, the image is first preprocessed, such as denoising, size scaling, etc.

[0100] 3.2.2 YOLO Object Detection

[0101] The YOLO (You Only Look Once) model is used to detect people in the image and generate corresponding bounding boxes for people. In the system of the present invention, the main task of YOLO is to identify people in the image through fast object detection and crop out each detected target person. The detection results will be combined with the results of depth estimation later to filter out the bounding boxes of the people close to the camera.

[0102] YOLO Detection Process: The system sends the input image into the YOLO model for object detection, and the output result is each detected bounding box for people. These bounding boxes define the area of each person in the image for subsequent processing steps. The YOLO object detection results are as follows:

[0103] Β={B1,B2,...,Bn};

[0104] Among them, Β is the set of all person bounding boxes detected by YOLO, and B i is the i-th detection bounding box, and each detection bounding box contains a person.

[0105] 3.2.3 Depth Detection

[0106] The depth estimation model (using DepthAnything) is used to calculate the depth information of each pixel point in the image, that is, the distance between each pixel point and the camera. The result of depth estimation is a depth map, where each pixel point is assigned a depth value.

[0107] Depth Map Generation: The depth estimation model calculates the depth map D(x,y) through the input image, where D(x,y) represents the distance between the pixel point at position (x,y) in the image and the camera. The depth map provides three-dimensional information in the scene and can help the system identify the distance of each pixel point from the camera.

[0108] Depth Estimation Formula:

[0109] D(x,y) = Depth(x,y);

[0110] Among them, D(x,y) is the depth value of each pixel point, and Depth(x,y) is the output result of the depth estimation model, representing the distance between each pixel point and the camera.

[0111] 3.2.4 Combining Depth Estimation with YOLO to Achieve Person Bounding Box Filtering

[0112] To determine the distance of each detected person bounding box B i from the camera, the system combines the person bounding boxes output by YOLO with the pixel depth information generated by the depth estimation model. By calculating the average depth of the pixels inside each person bounding box, the system can judge the distance of the target and filter out the targets that are too far away by setting a depth threshold.

[0113] A. Matching Person Bounding Boxes with the Depth Map

[0114] Person Bounding Boxes Detected by YOLO: What the YOLO model outputs are the person bounding boxes of each target person, defined as B i , and each person bounding box has its rectangular boundary, defined by the upper left corner coordinates (x1,y1) and the lower right corner coordinates (x2,y2). This means that the person bounding box contains the pixel points within the range from x1 to x2 and from y1 to y2.

[0115] Depth map generated by the depth estimation model: The depth estimation model generates depth information for each pixel in the entire image. This depth map D(x, y) represents the distance between the pixel at each position (x, y) in the image and the camera. This depth map has the same resolution and coordinate system as the input image of YOLO, so they can be directly matched through coordinates.

[0116] The matching steps are as follows:

[0117] (1) Traverse each person bounding box: In the set of person bounding boxes B = {B1, B2,..., B n}, process each person bounding box B i one by one. The boundary coordinates of each person bounding box B i are (x i1 , x i2 , y i1 , y i2 ), which represents the position range of the target person in the image. Among them, x i1 , y i1 are the abscissa and ordinate of the upper left corner of the person bounding box B i respectively, and x i2 , y i2 are the abscissa and ordinate of the lower right corner of the person bounding box B i respectively;

[0118] (2) Find the corresponding coordinate region from the depth map of the image: For each person bounding box B i , find the range of pixel points corresponding to the person bounding box from the depth map D(x, y) generated by the depth estimation model. The coordinates of this range are the same as the boundary coordinates of the person bounding box, that is, x i1 ≤ x ≤ x i2 and y i1 ≤ y ≤ y i2 , where x and y are the abscissa and ordinate of the pixel point respectively;

[0119] (3) Crop the corresponding depth map region: Crop the depth information part corresponding to the person bounding box B i from the depth map D(x, y). The cropped region contains the depth values of all pixel points within the person bounding box, that is, the depth information sub-map of the target person

[0120] The cropping formula is: for x i1 ≤ x ≤ x i2 , y i1 ≤ y ≤ y i2 ;

[0121] Among them, is the one corresponding to the person bounding box B iThe matching depth map sub-region, representing the corresponding region of the person in the depth map;

[0122] (4) Collect the depth information within the person's bounding box: Through a cropping operation, extract the set of depth values of all pixel points corresponding to the person's bounding box region from the depth map. This set represents the depth information of the person in three-dimensional space and is used for subsequent average depth calculation.

[0123] B. Calculation of the average depth within the person's bounding box

[0124] After obtaining the depth map sub-region corresponding to the person's bounding box B i The system can calculate the average depth value of this region to determine the distance of the target person from the camera.

[0125] Average depth calculation formula:

[0126]

[0127] Where:

[0128] D avg (B i ) is the average depth value of the person's bounding box B i

[0129] N is the total number of pixel points within the person's bounding box, and the calculation formula is N = (x2 - x1 + 1) * (y2 - y1 + 1). D(x, y) is the depth value of each pixel point within the person's bounding box region, taken from the cropped depth map sub-region

[0130] C. Distance filtering

[0131] After obtaining the average depth value D i (B avg ) for each person's bounding box B i , the system uses a preset depth threshold D thresh for distance filtering. If the average depth of the target person exceeds the threshold, it is considered that the target is too far from the camera, and the system filters it out.

[0132] Distance filtering formula:

[0133]

[0134] D avg (B i ) is the average depth value within the person's bounding box of the target person.

[0135] D thresh is the depth threshold set by the system, representing the maximum acceptable distance.

[0136] ​​ Indicates a valid human bounding box after depth filtering; if the person is too far away, the system will no longer process the target (i.e., ).

[0137] 3.2.4 Human Keypoint Detection

[0138] After depth estimation and distance filtering are completed, the system only retains those valid human bounding boxes that are relatively close and may be the speaker. Next, these valid human bounding boxes are fed into the human keypoint detection model one by one to further extract the keypoint information of the speaker, especially the keypoints in the mouth area. This keypoint information will provide an important basis for tasks such as voice enhancement and sound source localization.

[0139] The specific steps are as follows:

[0140] (1) The system traverses each filtered human bounding box Crops it into a separate image region and feeds it into the DWPose keypoint detection model.

[0141] (2) The DWPose model analyzes the human keypoints in this region and extracts multiple keypoints including the mouth, eyes, nose, etc.

[0142] (3) The set of detected keypoints K i Includes the coordinates of the upper and lower lips and the corners of the mouth of the mouth, providing basic data for subsequent analysis of mouth shape changes.

[0143] Keypoint set formula:

[0144] K i ={P lips ,P eyes ,P nose ...};

[0145] Among them, P lips Is the keypoint in the mouth area, including the position coordinates of the upper and lower lips and the corners of the mouth. The keypoints in the mouth area are an important basis for the system to subsequently analyze the mouth shape movement of the speaker.

[0146] 3.2.5 Mouth Shape Movement Analysis

[0147] After detecting the mouth keypoints of the speaker, the system starts a more in-depth analysis. By observing the changes in the mouth shape movement, the system can monitor the speaking state of the speaker in real time and generate dynamic parameters for voice enhancement and filtering. The following are several key analysis steps:

[0148] (1) Mouth opening and closing degree detection: To determine the vocalization state of the speaker, the system estimates the mouth opening and closing degree by calculating the vertical distance between the key points of the upper and lower lips. The mouth opening and closing degree is directly related to the sound intensity. When the mouth is opened wide, it usually indicates a high sound intensity, while a closed mouth indicates a weak or stopped sound.

[0149] Assume the set of key points of the upper lip is {(x u1 ,y u1 ),(x u2 ,y u2 ),...(x un ,y un )}, and the set of key points of the lower lip is {(x l1 ,y l1 ),(x l2 ,y l2 ),...(x ln ,y ln )}. Each pair of upper and lower lip key points corresponds one by one.

[0150] To calculate the overall mouth opening and closing degree, the vertical distance between each pair of upper and lower lip key points can be calculated, and then the average of these distances can be obtained.

[0151]

[0152] Where:

[0153] n is the number of upper and lower lip key points.

[0154] (x ui ,y ui ) is the coordinate of the i-th upper lip key point.

[0155] (x li ,y li ) is the coordinate of the i-th lower lip key point.

[0156] d open is the average mouth opening and closing degree.

[0157] (2) Lip movement speed detection: The system not only cares about the mouth opening and closing degree but also tracks the lip movement speed. By calculating the change rate of the mouth opening and closing degree between consecutive time points, the system can infer the vocalization frequency and speed of the speaker. Fast lip movement usually means a high vocalization frequency, and vice versa.

[0158] Lip movement speed formula:

[0159]

[0160] Where, v(t) is the lip movement speed at time point t, dopen (t) is the degree of mouth opening at that time point, and Δt is the time interval. The system infers the current vocal rhythm of the speaker based on the speed of lip movement.

[0161] 3.3 Interaction with the microphone array

[0162] The interaction with the microphone array is a core part of the sound source localization and directional sound enhancement system of the present invention. Through precise interaction, the directivity and gain settings of the microphone array can be controlled in real time to ensure accurate capture of the target sound source and effective suppression of background noise. The following details each step and technical details of the interaction with the microphone array.

[0163] 3.3.1 Receiving the azimuth information from the image processing module

[0164] Input of azimuth information: First, receive the azimuth information about the sound source target from the image processing module. This information includes:

[0165] Spatial position of the target: The position of the human target detected from the image in the image.

[0166] Position information of the mouth shape key points: The position information of the mouth shape key points based on human key point detection.

[0167] Mouth shape motion state: Based on the results of mouth shape detection and motion analysis, determine whether the target is making a sound.

[0168] This information provides the precise position of the sound source and the vocal state of the target, which is the key basis for subsequent directivity adjustment of the microphone array.

[0169] 3.3.2 Multimodal fusion of image and audio data

[0170] After the system captures the image data, this mouth shape information is transmitted to the microphone array in real time to further optimize the sound filtering and enhancement process. The microphone array initially determines the approximate direction of the sound source through acoustic algorithms (such as the phase difference method or the time delay method), while the image data provides more accurate information about the position of the speaker and the mouth shape movement. Through the fusion of multimodal data, the system can improve the positioning accuracy and dynamically adjust the sound processing parameters.

[0171] (1) Generating dynamic filter parameters from mouth shape information

[0172] The system adjusts the gain and frequency response of the sound filter through the dynamic parameters generated by the mouth key points to ensure that the sound enhancement process can respond to the dynamic changes of the speaker in real time. Specifically:

[0173] When a large opening and closing degree of the mouth is detected, the system will correspondingly increase the gain of the filter to enhance the intensity of the sound signal; while when the mouth is closed, the gain will be reduced to suppress environmental noise.

[0174] The lip movement speed is used to adjust the frequency response of the filter. Faster lip movements correspond to higher frequency components, and slower movements correspond to lower frequency components. Through this adaptive adjustment, the system can dynamically optimize the sound enhancement effect.

[0175] (2) Specific implementation of dynamic filter adjustment

[0176] Based on the above mouth key point information, the system constructs an adaptive dynamic filtering function. The adjustment of the filter is divided into the following aspects:

[0177] Gain adjustment: Adjust the gain of the filter according to the opening and closing degree of the mouth. The gain formula is:

[0178] G(t) = G0 + k·d open ;

[0179] Where G0 is the base gain, k is the gain adjustment coefficient, and d open is the opening and closing degree of the mouth.

[0180] Frequency response adjustment: Adjust the frequency response of the filter according to the lip movement speed. The frequency adjustment formula is:

[0181] f(t) = f0 + k f ·v(t);

[0182] Where f0 is the base frequency, k f is the frequency adjustment coefficient, and v(t) is the lip movement speed.

[0183] 3.3.3 Dynamic adjustment of the directivity of the microphone array

[0184] The microphone array consists of 6 uniformly distributed microphones. The system can calculate the corresponding angle based on the position of the lips relative to the center of the camera and dynamically adjust the gain of each microphone according to this angle.

[0185] (1) Precise azimuth calculation

[0186] By obtaining the lip key point coordinates P lips =(x lips , y lips ) of the sound emitter and the image center point (x center , y center ) through the camera, the angle of the lips relative to the center can be calculated. We can use the following formula to calculate the azimuth angle θ of the lips:

[0187]

[0188] This azimuth angle θ represents the angle of the lips relative to the center point of the camera, with a range from -180° to 180°. Through this angle information, the system can accurately identify the azimuth of the speaker and adjust the gain in the corresponding direction according to the layout of the microphone array.

[0189] (2) Adjust the gain of the 6-array microphone based on the azimuth angle

[0190] The 6-microphone arrays are usually distributed at 60-degree intervals. Assuming that each microphone covers an area of 60 degrees, the system can dynamically adjust the gain of the corresponding microphone according to the calculated azimuth angle θ.

[0191] A. Determine the microphone gain distribution based on the azimuth angle

[0192] If the azimuth angle θ of the lip key point falls within the area covered by a certain microphone (for example, θ is between -30° and 30°), the microphone in this direction will receive more gain; the gain of the adjacent microphones can be interpolated according to the distance from this azimuth. For example, the closer the microphone is to this angle, the greater its gain, and the farther the microphone is from this angle, the smaller its gain.

[0193] For each microphone i, we set its central angle as θ, then the gain adjustment can be calculated by the following formula:

[0194] G i (t) = G0 + k·max(0, cos(θ - θ i ));

[0195] Where:

[0196] G i (t) is the gain value of the i-th microphone at time t.

[0197] G0 is the base gain.

[0198] θ is the azimuth angle of the current speaker calculated through the lip azimuth.

[0199] θ i is the central angle of the i-th microphone (such as 0 degrees, 60 degrees, 120 degrees, etc.).

[0200] k is the gain adjustment coefficient.

[0201] By using the cosine function cos(θ - θ i ) to represent the relationship between the lip azimuth and the central angle of each microphone, the system can dynamically adjust the gain of each microphone.

[0202] When the lip azimuth angle θ approaches the central angle θ of a certain microphone i , the gain G i will be maximized. Other microphones that are farther away will receive smaller gains.

[0203] B. Distributed gain adjustment

[0204] In addition, we also perform distributed gain adjustment. The gain distribution based on the lip azimuth angle can achieve smooth transition among multiple microphones. This means that when the lip position is in the middle of the coverage areas of two microphones, the gains of two adjacent microphones will be adjusted accordingly, thereby enhancing the sound signal in this direction. On the contrary, for microphones that are far away from the lip azimuth, the gain will be reduced or even suppressed.

[0205] Compared with the prior art, the present invention has significant advantages in sound source localization and directional sound enhancement. First of all, the present invention combines the multi-modal data of the microphone array and the camera. Through object detection, depth detection, lip movement analysis, etc. of the image processing module, accurate localization of the sound source target is achieved. This multi-modal fusion effectively overcomes the problem of insufficient accuracy of traditional single acoustic localization methods in complex environments, especially performing excellently in multi-person scenarios and high-noise environments.

[0206] Secondly, the in-depth interaction between the system and the microphone array enables the present invention to not only accurately adjust the directivity of the microphone, but also suppress environmental noise through a dynamic adaptive algorithm, thereby significantly improving the clarity of the target sound source. This real-time dynamic adjustment ability enables the present invention to maintain stable performance in different application scenarios, meeting the higher actual requirements for sound processing accuracy and noise suppression. Generally speaking, the present invention surpasses the prior art in terms of the accuracy of sound source localization, the effect of sound enhancement, and the ability to adapt to complex environments, and has broad application prospects.

[0207] The above embodiments are the preferred embodiments of the present invention, but the embodiments of the present invention are not limited by the above embodiments. Any other changes, modifications, substitutions, combinations, and simplifications made without departing from the spirit and principle of the present invention shall be equivalent replacement methods and are all included in the protection scope of the present invention.

Claims

1. A method for sound source direction localization and directional sound enhancement based on a microphone array and a camera, characterized in that The following steps are involved: S1, installing the camera at the center of the microphone array and preprocessing the image collected by the camera; Detect all people in the image through the object detection model and generate corresponding person frames; S2. Calculate the depth information of each pixel in the image using the depth estimation module, and match the depth information of the pixel with the detected human frame; By calculating the average depth of the pixels inside each person's frame, the distance between the speaker and the camera is determined. By setting a depth threshold, targets that are far away are filtered out, and valid frames of those that are close and may be the speaker are retained. S3, analyzing the key points of the human body on the valid human frame, and extracting multiple key points of the human body including the mouth, eyes, and nose; S4, analyzing the mouth shape movement state of the key points of the mouth: detecting the degree of mouth opening and closing representing the intensity of the sound of the speaker, and detecting the lip movement speed representing the frequency and speed of the sound of the speaker; The gain of the sound filter increases as the speaker's mouth opens and closes more, and the strength of the sound signal increases accordingly; The frequency response of the sound filter increases as the speed of the speaker's mouth movement becomes faster; S5, calculating the angle of the speaker's mouth relative to the center point of the camera, thereby identifying the speaker's position, and adjusting the gain of the microphone in the corresponding direction according to the layout of the microphone array; The specific implementation steps of step S5 are as follows: (1) Precise azimuth calculation: Obtain the coordinates P of the lip key points of the sound emitter through the camera lips =(x lips , y lips ) and the center point of the image (x center , y center ), calculate the angle of the lips relative to the center point of the camera, that is, the azimuth angle θ of the lips: (2) Adjust the gain of the microphone array based on the azimuth angle: For each microphone i, we set its center angle to θ, then the gain adjustment can be calculated by the following formula: G i (t) = G0 + k·max(0, cos(θ - θ i )); Among them, G i (t) is the gain value of the i-th microphone at time t, G0 is the base gain, and θ i is the central angle of the i-th microphone, and k is the adjustment coefficient of the gain; By using the cosine function cos(θ - θ i ), to represent the relationship between the lip orientation and the angle of each microphone center, the system can dynamically adjust the gain of each microphone; When the lip azimuth angle θ approaches the center angle θ of a certain microphone i the gain G i will be maximized; other microphones at a greater distance will receive a small gain.

2. The method for sound source direction localization and directional sound enhancement based on a microphone array and a camera according to claim 1, wherein The depth information of the pixel is matched with the detected human frame, and the specific steps are as follows: (1) Traverse each person bounding box: In the set of person bounding boxes B = {B1, B2,..., B n} detected by the object detection model, process each person bounding box B i one by one. The boundary coordinates of each person bounding box B i are (x i1 , x i2 , y i1 , y i2 ), representing the position range of the target person in the image. Among them, x i1 , y i1 are the abscissa and ordinate of the upper left corner of the person bounding box B i respectively, and x i2 , y i2 are the abscissa and ordinate of the lower right corner of the person bounding box B i respectively; (2) Find the corresponding coordinate region from the depth map of the image: For each person bounding box B i , find the range of pixel points corresponding to the person bounding box from the depth map D(x, y) generated by the depth estimation model. The coordinates of this range are the same as the boundary coordinates of the person bounding box, that is, x i1 ≤x≤x i2 and y i1 ≤y≤y i2 , where x and y are the abscissa and ordinate of the pixel point respectively; (3) Crop the corresponding depth map region: Crop out the depth information part corresponding to the human bounding box B in the depth map D(x, y). The cropped region contains the depth values of all pixels within the human bounding box, that is, the depth information sub-map of the target person. i The cropping formula is: Among them, is the depth map sub-region matching the person box B i indicating the corresponding region of the person in the depth map; (4) Collecting depth information within the human frame: Through cropping operations, a set of depth values of all pixels corresponding to the human frame area of the person is extracted from the depth map. This set represents the depth information of the person in three-dimensional space and is used for subsequent average depth calculations.

3. The method for sound source direction localization and directional sound enhancement based on a microphone array and a camera according to claim 1, wherein The average depth of the pixels inside the human frame is calculated as follows: Among them, D avg (B i ) is the average depth value of the human bounding box B i , N is the total number of pixel points within the human bounding box; D(x, y) is the depth value of each pixel point within the human bounding box area, taken from the cropped depth map sub-region 4. The method for sound source direction localization and directional sound enhancement based on a microphone array and a camera according to claim 1, wherein The degree of opening and closing of the mouth of the speaker is calculated as follows: Suppose the set of upper lip key points is \(\{(x u1 , y u1 ), (x u2 , y u2 ),... (x un , y un )\}, and the set of lower lip key points is \(\{(x l1 , y l1 ), (x l2 , y l2 ),... (x ln , y ln )\}, and each pair of upper and lower lip key points corresponds one-to-one; Calculate the vertical distance between each pair of upper and lower lip key points, and then find the average of these distances; where n is the number of upper and lower lip key points; (x ui , y ui ) are the coordinates of the i-th upper lip key point; (x li , y li ) are the coordinates of the i-th lower lip key point; d open is the average opening degree of the mouth.

5. The method for sound source direction localization and directional sound enhancement based on a microphone array and a camera according to claim 1, wherein The lip movement speed of the speaker is calculated as follows: where v(t) is the movement speed of the lips at time point t, and d open (t), d open (t - 1) are the mouth opening degrees at time points t and t - 1 respectively, and Δt is the time interval.

6. The method for sound source direction localization and directional sound enhancement based on a microphone array and a camera according to claim 1, wherein The gain G(t) of the sound filter is adjusted according to the degree of opening and closing of the speaker's mouth: G(t) = G0 + k·d open ; where G0 is the base gain, k is the gain adjustment coefficient, and d open is the degree of mouth opening and closing; The frequency response f(t) of the filter is adjusted according to the lip movement speed of the speaker: f(t) = f0 + k f ·v(t); where f0 is the base frequency, k f is the frequency adjustment coefficient, and v(t) is the lip movement speed.

7. The method for sound source direction localization and directional sound enhancement based on a microphone array and a camera according to claim 1, characterized in that, The microphone array also performs distributed gain adjustment: the gain distribution based on the lip azimuth angle makes a smooth transition between multiple microphones; when the lip position is in the middle of the coverage area of two microphones, the gains of the two adjacent microphones will be adjusted accordingly to enhance the sound signal in this direction; for microphones that are far away from the lip position, the gain will be reduced or even suppressed.

8. A server, the server comprising a processor and a memory, characterized in that, At least one program is stored in the memory, and the program is loaded and executed by the processor to implement the method for sound source direction positioning and directional sound enhancement based on a microphone array and a camera as described in any one of claims 1 to 7.

9. A computer-readable storage medium storing at least one program, characterized in that, The program is loaded and executed by a processor to implement the method for sound source direction localization and directional sound enhancement based on a microphone array and a camera according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Method and device for adjusting sound pickup sensitivity

    CN112015364B

  • Sound detection methods and related equipment

    CN115862682B

  • Face and voiceprint authentication system and method based on deep transfer learning

    CN111723679A

  • Customer service robot and dialogue processing system and method based on deep learning

    CN117995187A