Speaker target detection and tracking methods, systems, and media
By combining visual two-dimensional partitioning and audio three-dimensional localization techniques, the accuracy and cost issues of speaker detection in complex scenarios are solved, achieving efficient and low-cost speaker target detection and tracking.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SHANGHAI GOLDEN BRIDGE INFOTECH CO LTD
- Filing Date
- 2026-04-15
- Publication Date
- 2026-07-21
AI Technical Summary
Existing speaker detection technologies struggle to balance positioning accuracy and hardware cost in complex scenarios, making it impossible to achieve real-time, accurate speaker detection and dynamic tracking.
By combining visual two-dimensional partitioning with audio three-dimensional localization, and through the collaborative work of detection cameras and microphones, the speaker's coordinates are determined using visual partitioning and sound source localization, and the position of the tracking camera is adjusted to achieve precise tracking.
It improves the accuracy and reliability of speaker detection, reduces hardware costs and deployment complexity, and adapts to the actual needs of different scenarios.
Smart Images

Figure CN122435671A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of computer vision technology, and more specifically to a method, system, and medium for speaker target detection and tracking. Background Technology
[0002] In scenarios such as meetings, teaching, and live streaming, accurate real-time speaker capture and tracking are core requirements for improving content recording quality and viewing experience. Current mainstream speaker detection technologies are mainly divided into two categories: audio-based and vision-based solutions, both of which have significant limitations.
[0003] Among these, audio-based solutions often employ a deployment of multiple fixed microphones on a desktop, identifying the speaker by recognizing the microphone position corresponding to the sound source. This approach lacks flexibility, supporting only fixed-seat speaker detection and failing to track speaker movement or identify areas without seats. Furthermore, the hardware costs are high, requiring a separate microphone for each seat, resulting in complex wiring and making it unsuitable for large-scale or flexible layout scenarios.
[0004] Vision-based solutions capture images using one or more cameras and employ object detection algorithms to identify people and determine the speaker's location. However, these solutions only partition the camera image into two-dimensional coordinate planes. In scenarios where multiple people are present, moving simultaneously, or obstructing each other's view, they struggle to accurately distinguish speakers, limiting their location accuracy.
[0005] In summary, current technical solutions cannot balance positioning accuracy and hardware cost control, and are insufficient to meet the actual needs for real-time speaker detection, precise positioning, and dynamic tracking in complex scenarios.
[0006] In view of this, the present invention is hereby proposed. Summary of the Invention
[0007] The present invention is proposed in view of the above-mentioned problems. According to one aspect of the present invention, a speaker target detection method is provided for a speaker target detection system, the system including a detection camera and a microphone; the method includes: The detection camera continuously acquires real-time images of the scene; Target detection is performed on the real-time footage of each scene to identify human targets in the real-time footage of each scene. Based on the mouth changes of each human target in the real-time scene of a continuously preset number of frames, determine the target speaker and the target visual partition in which it is located. The sound source is located based on the sound signal collected in real time by the microphone, and the final sound source coordinates are determined. Based on the coordinates of the target visual partition and the final sound source coordinates, the speaker coordinates are determined, where the speaker coordinates are the three-dimensional coordinates of the target speaker's location.
[0008] For example, determining the target speaker and its target visual partition based on the mouth changes of each human target in a continuously preset number of frames of real-time scene footage includes: For any of the aforementioned human targets Determine the aspect ratio of the human target's mouth in the real-time frame of each of the aforementioned scenes; Based on the magnitude and frequency of the change in the aspect ratio of the human target's mouth in the real-time scene of the scene in the continuous preset number of frames, it is determined whether the human target is speaking. When the human target exhibits speaking behavior, that human target is identified as the target speaker; Preferably, determining the aspect ratio of the human target's mouth in each real-time scene includes: For any of the aforementioned real-time scenes, extract the key points of the human target's mouth and determine the coordinates of the key points of the human target's mouth. The key points of the mouth include the key points of the upper lip, the key points of the lower lip, the key points of the left corner of the mouth, and the key points of the right corner of the mouth. Based on the coordinates of the key points of the human target's mouth, determine the aspect ratio of the human target's mouth in the real-time image of the current scene.
[0009] For example, the step of locating the sound source and determining the final sound source coordinates based on the sound signal collected in real time by the microphone includes: The sound signal continuously collected by the microphone is subjected to temporal convolution calculation to obtain a clean sound signal; The pure sound signal is calculated using signal time-of-flight technology and time difference positioning technology to obtain the initial sound source coordinates; Calculate the distance between the initial sound source coordinates and different preset sound anchor points; The coordinates of the preset sound anchor point that is closest to the initial sound source coordinates are taken as the final sound source coordinates; The preset sound anchor points are obtained by: establishing a three-dimensional sound coordinate system with the location of the microphone as the origin, and measuring the coordinates of the sound anchor points at each preset speaking position in the space to obtain the coordinates of each preset sound anchor point.
[0010] For example, determining the speaker coordinates based on the coordinates of the target visual partition and the final sound source coordinates includes: Determine whether the final sound source coordinates match the coordinates of the target visual partition; When a match is found, the final sound source coordinates are determined to be the speaker coordinates.
[0011] According to another aspect of the present invention, a speaker target tracking method is provided. This method is used in a speaker target tracking system, the speaker target tracking system including a speaker target detection system and a tracking camera, the speaker target detection system including a detection camera and a microphone; the method includes: determining speaker coordinates using the aforementioned speaker target detection method; Based on the speaker's coordinates, the horizontal position, vertical position, and focal length of the tracking camera are adjusted so that the speaker is located in the center of the tracking camera's image.
[0012] For example, before adjusting the horizontal position, vertical position, and focal length of the tracking camera, the method further includes: Calculate the distance change between the speaker's coordinates and the historical coordinates; The step of adjusting the horizontal position, vertical position, and focal length of the tracking camera is performed when the distance change value exceeds a distance threshold.
[0013] Exemplarily, the method further includes: When no speaker is detected using the speaker target detection method, the image captured by the detection camera is displayed; When a speaker is detected using the speaker target detection method, the image captured by the tracking camera is displayed.
[0014] According to another aspect of the present invention, a speaker target detection system is provided, comprising: a detection camera, a microphone, and a target detection controller, wherein the detection camera and the microphone are both connected to the target detection controller; the target detection controller is used to perform the speaker target detection method described above.
[0015] According to another aspect of the present invention, a speaker target tracking system is provided, comprising: a speaker target detection system, a tracking camera, and a tracking controller, wherein the speaker target detection system includes a detection camera and a microphone; the detection camera, the microphone, and the tracking camera are all connected to the tracking controller; and the tracking controller is used to perform the speaker target tracking method described above.
[0016] According to another aspect of the present invention, a computer-readable storage medium is provided, which stores a computer program / instructions that, when executed by a processor, implement the method described above.
[0017] In the above technical solution, by combining visual two-dimensional partitioning with audio three-dimensional localization, the shortcomings of single visual solutions in ignoring distance differences can be overcome. This effectively avoids localization errors in complex scenarios with multiple people and occlusion, achieving multimodal information complementarity and significantly improving the accuracy and reliability of speaker target detection. The localization accuracy is significantly superior to existing single visual or single audio localization solutions. Moreover, this technical solution only requires a single matrix microphone to achieve full-scene sound source acquisition and localization, eliminating the need for a separate microphone for each seat, simplifying wiring and deployment processes, and significantly reducing hardware procurement and construction costs. In summary, this solution can reduce implementation costs while ensuring localization accuracy and can adapt to the actual application needs of different scenarios.
[0018] The above description is merely an overview of the technical solution of the present invention. In order to better understand the technical means of the present invention and to implement it in accordance with the contents of the specification, and in order to make the above and other objects, features and advantages of the present invention more apparent and understandable, specific embodiments of the present invention are described below. Attached Figure Description
[0019] The above and other objects, features, and advantages of the present invention will become more apparent from the more detailed description of the embodiments of the invention in conjunction with the accompanying drawings. The drawings are provided to further illustrate the embodiments of the invention and form part of the specification. They are used together with the embodiments of the invention to explain the invention and do not constitute a limitation thereof. In the drawings, the same reference numerals generally represent the same parts or steps.
[0020] Figure 1 A schematic flowchart of a speaker target detection method according to an embodiment of the present invention is shown; Figure 2 A schematic diagram of the adjustment of a tracking camera according to an embodiment of the present invention is shown; Figure 3 A schematic diagram of a speaker target tracking system according to an embodiment of the present invention is shown. Detailed Implementation
[0021] To make the objectives, technical solutions, and advantages of the present invention more apparent, exemplary embodiments according to the present invention will be described in detail below with reference to the accompanying drawings. Obviously, the described embodiments are merely a part of the embodiments of the present invention, and not all of the embodiments of the present invention. It should be understood that the present invention is not limited to the exemplary embodiments described herein. Based on the embodiments of the present invention described herein, all other embodiments obtained by those skilled in the art without inventive effort should fall within the protection scope of the present invention.
[0022] According to one aspect of the present invention, a speaker target detection method is provided. This method is used in a speaker target detection system, which includes a detection camera and a microphone. In this document, the detection camera is used to acquire panoramic images, and the microphone is used to acquire spatial sound signals. The microphone includes multiple pickup units that can work collaboratively to perform sound source localization.
[0023] Figure 1 A schematic flowchart of a speaker target detection method according to an embodiment of the present invention is shown. Figure 1 As shown, the method may include the following steps: S110, S120, S130, S140, and S150.
[0024] In step S110, the detection camera continuously acquires real-time images of the scene.
[0025] In some solutions of this example, visual point presets can be performed before step S110. Specifically, the detection camera can be controlled to capture panoramic images first. Then, a two-dimensional visual coordinate system (U, V) is established, where U is the horizontal coordinate value and V is the vertical coordinate value, dividing the image into several uniform partitions, each visual partition corresponding to a unique two-dimensional coordinate range.
[0026] In this example, the number of detection cameras can be one or more, with the goal of covering the entire indoor space. When there are multiple detection cameras, the real-time scene view obtained at each moment includes the image acquired by each detection camera at the corresponding moment.
[0027] In step S120, target detection is performed on each scene's real-time frame to determine human targets in each scene's real-time frame.
[0028] Object detection methods include, but are not limited to, human object detection using pre-trained neural network models or object detection algorithms. In a specific embodiment, the YOLO series object detection algorithms (such as YOLO11s-pose, with algorithm weights using officially released pre-trained model weights) can be invoked to perform object detection on each scene's real-time image. This model can accurately detect 17 key points of the human body (including nose, left eye, right eye, left ear, right ear, left shoulder, right shoulder, left elbow, right elbow, left wrist, right wrist, left hip, right hip, left knee, right knee, left ankle, and right ankle), and determine whether a human object exists in the image based on the existence and distribution characteristics of the key points. If a human object is detected, the model outputs the two-dimensional visual coordinate range corresponding to the human object and the coordinates of each key point.
[0029] In step S130, the target speaker and the target visual partition in which they are located are determined based on the mouth changes of each human target in the real-time scene of a continuously preset number of frames.
[0030] After identifying the human target, the changes in the mouth of the same human target in a series of preset number of real-time frames of the scene can be analyzed. For example, the preset number of real-time frames of the scene can be directly input into a pre-trained neural network model, which can then identify the target speaker and the visual region in which they are located. Alternatively, the amplitude and / or frequency of mouth changes can be calculated and compared with preset thresholds to determine whether a specific human target is a speaker.
[0031] The preset number can be selected according to actual needs, such as 14 frames, 15 frames, 16 frames, etc.
[0032] In some embodiments, the same human target in a predetermined number of consecutive frames of real-time scene footage can be pre-determined for subsequent analysis of mouth changes of specific human targets. For example, an IOU-based tracking strategy can be used, measuring the overlap of human bounding boxes in adjacent frames using the Intersection over Union (IOU). If the overlap exceeds a threshold (e.g., 0.5), the target is identified as the same and assigned the same ID. Another example is the use of 17 key points of the human body output by YOLO11s-pose to calculate the similarity and positional distance of the key points for tracking. Of course, direct face recognition can also be used for confirmation, which will not be elaborated upon here.
[0033] As described above, there can be multiple detection cameras. When there are multiple detection cameras, step S130 can be performed on a preset number of consecutive real-time frames of scene footage detected by each detection camera to identify the final target speaker.
[0034] In some embodiments, before performing step S130, the method may further include the following step: cropping the face region of each frame of the real-time scene. When subsequently performing mouth change recognition, the face image can be analyzed directly, which helps reduce computational load and minimizes interference from other areas besides the face. During face region cropping, if the head turn angle is excessive, the shoulders are detected; in this case, the mouthPosition is masked. This ensures the continuity of multi-frame tracking and target recognition.
[0035] In one specific embodiment, the same human target in a consecutive preset number of real-time scene frames can be identified first, with the same human target identified by the same ID. Then, the face region is cropped from each frame of the real-time scene, and faces with the same ID are arranged into a face sequence according to time sequence. Finally, mouth change analysis is performed on each face sequence (for example, when determining whether the human target is speaking based on the magnitude and frequency of changes in the aspect ratio of the mouth in the consecutive preset number of real-time scene frames, the judgment can be made based on the corresponding face region images of the human target in each of the consecutive preset number of real-time scene frames), thereby determining whether the human target corresponding to the face sequence is speaking.
[0036] In some embodiments, a frontal face quality score can also be calculated to obtain the best avatar that may be needed in subsequent business processes (such as sending it to a server, displaying it on a conference interface, associating it with a speaker ID, etc.).
[0037] After identifying the target speaker, the visual partition containing the key points on the speaker's head can be used as the target visual partition. For example, in the embodiment for identifying 17 key points of the human body described above, for a human target, the visual partition containing the most head key points (nose, left eye, right eye, left ear, right ear) is selected as the target visual partition.
[0038] In step S140, the sound source is located based on the sound signal collected in real time by the microphone, and the final sound source coordinates are determined.
[0039] In some embodiments, a three-dimensional sound coordinate system (X, Y, Z) can be established in advance with the microphone installation position as the origin, where the X-axis is the horizontal direction, the Y-axis is the vertical direction, and the Z-axis is the front-to-back distance direction; possible speaking positions in the space are measured, and several sound anchor points are set, each anchor point corresponding to a unique three-dimensional coordinate (X, Y, Z). n Y n Z n (n is the anchor point number, n=1, 2, ..., N), the anchor points cover the entire speaking area, ensuring that any speaking position corresponds to the nearest anchor point.
[0040] If the microphone used adopts a spherical coordinate system Devices for determining spatial coordinates can be used to... The distance from the coordinate point to the origin. The angle between the coordinate point and the Z-axis, ranging from... , Let be the angle between the projection of the coordinate point onto the XY plane and the positive Z-axis, ranging from . Then it needs to be converted to rectangular coordinates using the following formula: .
[0041] In step S150, the speaker coordinates are determined based on the coordinates of the target visual partition and the final sound source coordinates. The speaker coordinates are the three-dimensional coordinates of the target speaker's location.
[0042] In some embodiments, a mapping relationship between the visual two-dimensional coordinate system and the audio three-dimensional coordinate system can be established in advance. Then, the final sound source coordinates are projected onto the visual two-dimensional coordinate system. Next, it is determined whether the projection point is within the target visual partition. If the projection point is within the target visual partition, the final sound source coordinates are determined to be the speaker coordinates. Of course, other methods can also be used to constrain and verify the final sound source coordinates using the coordinates of the target visual partition to determine the speaker coordinates.
[0043] The aforementioned technical solution combines visual two-dimensional partitioning with audio three-dimensional localization, overcoming the limitation of single-vision solutions that ignore distance differences. It effectively avoids localization errors in complex scenarios with multiple people and occlusion, achieving multimodal information complementarity and significantly improving the accuracy and reliability of speaker target detection. Its localization accuracy is significantly superior to existing single-vision or single-audio localization solutions. Furthermore, this solution requires only a single matrix microphone to achieve full-scene sound source acquisition and localization, eliminating the need for a separate microphone for each seat, simplifying wiring and deployment processes, and significantly reducing hardware procurement and construction costs. In summary, this solution can reduce implementation costs while ensuring localization accuracy and can adapt to the practical application needs of different scenarios.
[0044] For example, determining the target speaker and its target visual partition based on the mouth changes of each human target in a continuously preset number of real-time scene frames includes: for any human target, determining the mouth aspect ratio of the human target in each real-time scene frame; determining whether the human target has speaking behavior based on the magnitude and frequency of the mouth aspect ratio change of the human target in a continuously preset number of real-time scene frames; and determining the human target as the target speaker when the human target has speaking behavior.
[0045] The aspect ratio of the mouth can be directly identified using a pre-trained neural network model (such as ResNet-18 / MobileNet-v2), or it can be calculated by first determining the key points of the mouth and then using the coordinates of the key points.
[0046] In some embodiments, determining whether a human target is speaking based on the magnitude and frequency of changes in the mouth aspect ratio in a consecutive preset number of real-time frames of a scene may include the following steps: calculating the magnitude of the change in the mouth aspect ratio in every two adjacent real-time frames of the scene; summing the magnitudes; calculating the number of times the magnitude exceeds a minimum magnitude threshold to obtain the frequency; determining that the human target is speaking when the sum of the magnitudes exceeds the magnitude threshold and the frequency exceeds the frequency threshold; otherwise, determining that the human target is not speaking. The minimum magnitude threshold, the magnitude threshold, and the frequency threshold can all be determined based on actual working conditions or experience, and will not be elaborated further.
[0047] In other embodiments, determining whether a human target is speaking based on the magnitude and frequency of changes in the aspect ratio of its mouth in a continuous preset number of real-time frames can include the following steps: dividing the continuous preset number of real-time frames into multiple sliding windows, wherein the step size of each sliding window is a preset step size, the first frame of the first sliding window is the first frame in the continuous preset number of real-time frames, and the last frame of the last sliding window is the last frame in the continuous preset number of real-time frames; for each sliding window, calculating the average magnitude of change in the aspect ratio of the human target's mouth in that sliding window, the average magnitude of change being the ratio of the sum of the differences in the aspect ratio of the human target's mouth in two adjacent frames of the real-time scene within the sliding window to the length of the sliding window; determining that the human target is speaking when the average magnitude of change in at least some of the multiple sliding windows exceeds a magnitude threshold, or when the average magnitude of change in the last sliding window exceeds the magnitude threshold and the average magnitude of change in at least one of the multiple sliding windows (excluding the last sliding window) exceeds the magnitude threshold. For example, in this embodiment... The preset number of consecutive sliding windows is 14, the sliding window length is 8 frames, the preset step size is 2 frames, and the number of sliding windows in at least some parts is 3. By dividing the window into sliding windows [0,7], [2,9], [4,11], and [6,13], if the average variation of [0,7], [2,9], and [4,11] all exceeds the variation threshold, it can be determined that the human target is exhibiting speech behavior. If the average variation of any one of [0,7], [2,9], and [4,11] exceeds the variation threshold, and the average variation of [6,13] also exceeds the variation threshold, it can also be determined that the human target is exhibiting speech behavior. Otherwise, it can be determined that the human target is not exhibiting speech behavior.
[0048] The above technical solution calculates and analyzes the mouth aspect ratio of each human target in a real-time scene of a continuously preset number of frames. Based on the change range and frequency of the mouth aspect ratio, it accurately determines whether the target is speaking, and then determines the target speaker and the target visual partition in which it is located. It can quickly and accurately identify the person who is speaking in complex scenes, effectively improving the real-time performance and accuracy of speaker detection, while reducing the probability of misjudgment. It can provide a more accurate basis for subsequent speaker location detection.
[0049] For example, determining the aspect ratio of the human target's mouth in each real-time scene includes: for any real-time scene, extracting key points of the human target's mouth, determining the coordinates of the key points of the human target's mouth, including key points of the upper lip, lower lip, left corner of the mouth, and right corner of the mouth; and determining the aspect ratio of the human target's mouth in the current real-time scene based on the coordinates of the key points of the human target's mouth.
[0050] Mouth landmarks can be extracted using any existing facial landmark detection model, such as YoLo11L-pose. When using a model to detect mouth landmarks, multiple (e.g., four) real-time scene images (which can be the original scene images or cropped images of the face region) can be pre-stitched together into a single image and then input into the model for mouth landmark detection. This helps improve detection efficiency. Taking the example above where faces with the same ID are arranged in chronological order to form a face sequence, every four face region images in each face sequence can be stitched together into one image. Then, mouth landmark detection can be performed on this image to identify the mouth landmarks of each face in the image.
[0051] After obtaining the key points of the mouth, the aspect ratio of the corresponding human mouth can be determined based on the coordinates of the key points. Specifically, the distance between the key points of the upper lip and the lower lip can be used as the height of the mouth region, and the distance between the key points of the left and right corners of the mouth can be used as the width of the mouth region. The aspect ratio is obtained by the ratio of the width to the height.
[0052] The above scheme is simple to calculate and can quickly and accurately determine the mouth aspect ratio of each human target in a single frame image, thus providing a more accurate basis for speaker recognition in subsequent processes.
[0053] For example, the method of locating the sound source and determining the final sound source coordinates based on the sound signal collected in real time by the microphone includes: performing temporal convolution calculation on the sound signal continuously collected by the microphone to obtain a clean sound signal; calculating the clean sound signal using signal time-of-flight technology and time difference positioning technology to obtain the initial sound source coordinates; calculating the distance between the initial sound source coordinates and different preset sound anchor points; and taking the coordinates of the preset sound anchor point closest to the initial sound source coordinates as the final sound source coordinates. The preset sound anchor points are obtained by: establishing a three-dimensional sound coordinate system with the microphone location as the origin, and measuring the coordinates of the sound anchor points at each preset speaking position in the space to obtain the coordinates of each preset sound anchor point.
[0054] In this example, noise filtering and sound source signal enhancement can be achieved through continuous sampling, which is implemented using the following formula: ; in, The raw sound signal (including ambient noise) continuously collected by the matrix microphone. The preset Gaussian filter convolution kernel is used to filter environmental noise and enhance the characteristics of human voice signals. This is the clean human voice signal output after the convolution operation. For the signal time-domain integral variable, This is the convolution operator.
[0055] Then, the purified human voice signal can be used in conjunction with Time-of-Flight (ToF) technology to obtain the sound propagation time. t Combined with the speed of sound v (Default 343m / s), calculate the straight-line distance from the sound source to the microphone: d = v × t The time difference of sound arrival at different pickup units is calculated using Time Difference Occurrence Addressing (TDOA) technology to determine the horizontal and vertical angles of the sound source and obtain the initial sound source coordinates. X s , Y s , Z s ).
[0056] Next, the spatial distance between the initial sound source coordinates and each preset sound anchor point is calculated using the following formula, and the coordinates of the preset sound anchor point with the smallest distance are selected as the final sound source coordinates. The calculation formula is: ; in, D n The spatial distance between the initial sound source coordinates and the nth preset sound anchor point, (X n , Y n , Z n () represents the coordinates of the nth preset sound anchor point.
[0057] The above technical solution acquires sound signals in real time using a microphone. First, it performs temporal convolution calculations on the continuously acquired signals to remove interference and obtain a clean sound signal. Then, it combines signal time-of-flight technology and time difference positioning technology to calculate the initial sound source coordinates. Subsequently, it constructs a three-dimensional sound coordinate system with the microphone position as the origin. The initial sound source coordinates are compared with the coordinates of the sound anchor points corresponding to the preset speaking positions in the space. The coordinates of the closest sound anchor points are selected as the final sound source coordinates. This can effectively improve the accuracy and stability of sound source positioning, reduce the impact of environmental noise and signal errors on the positioning results, and achieve rapid and accurate identification and determination of the sound source position within the preset speaking area.
[0058] For example, determining the speaker coordinates based on the coordinates of the target visual partition and the final sound source coordinates includes: determining whether the final sound source coordinates match the coordinates of the target visual partition; and if they match, determining the final sound source coordinates as the speaker coordinates.
[0059] In this example, a pre-established correspondence between preset sound anchor points and visual zones can be created. Specifically, each preset sound anchor point can be projected onto a two-dimensional visual coordinate system, and the visual zone where the projection point is located can be associated with the corresponding preset sound anchor point. During matching, it can be determined whether the final sound source coordinates are sound anchor points associated with the target visual zone. If so, the match is confirmed, and the final sound source coordinates are determined to be the speaker coordinates.
[0060] The above technical solution matches the final sound source coordinates with the target visual partition coordinates. When the two coordinates match, the final sound source coordinates are directly determined as the speaker coordinates. This can effectively improve the accuracy and reliability of speaker positioning, reduce positioning errors caused by deviations in single sound source positioning information, and achieve effective fusion and verification of sound source positioning information and visual partition information. This allows for accurate determination of the real speaker's position coordinates and improves the credibility and stability of the overall positioning results.
[0061] According to another aspect of the present invention, a speaker target tracking method is provided. The method is used in a speaker target tracking system, which includes a speaker target detection system and a tracking camera. The speaker target detection system includes a detection camera and a microphone. The method includes: determining the speaker coordinates using the above-described speaker target detection method; and adjusting the horizontal position, vertical position, and focal length of the tracking camera based on the speaker coordinates so that the speaker is located at the center of the tracking camera's image.
[0062] In this example, the speaker's coordinates can be determined first using the speaker target detection method described above. Then, the horizontal and vertical positions of the tracking camera can be fine-tuned based on the X and Y coordinates of the speaker's coordinates, and the focal length can be adjusted based on the Z coordinate to zoom in and out of the image. Ultimately, the speaker is positioned in the center of the image, and the aspect ratio meets the requirements for a close-up (the upper body of the person accounts for 70%-80%). Figure 2 A schematic diagram of the adjustment of a tracking camera according to an embodiment of the present invention is shown. Figure 2 As shown in this example, preset positions for the tracking camera and each sound anchor point can be pre-set. After determining the speaker's coordinates, the tracking camera invokes the preset positions, then adjusts the horizontal (Pan) and vertical (Tilt) directions according to the X and Y coordinates, and zooms according to the Z coordinate.
[0063] The aforementioned technical solution combines a detection camera and a microphone to collaboratively determine the speaker's coordinates. Based on these coordinates, it precisely adjusts the horizontal and vertical positions and focal length of the tracking camera, enabling real-time and stable tracking of the speaker target. This ensures the speaker remains centered within the tracking camera's view, effectively improving the accuracy of speaker tracking and image presentation. This provides reliable technical support for target tracking and speaker close-up needs in scenarios such as video conferencing and intelligent monitoring. Furthermore, the tracking camera in this solution can be adjusted in real-time according to the speaker's location, supporting real-time dynamic tracking during speaker movement. It eliminates the need for a fixed position, adapting to scenarios where the speaker moves freely. Compared to fixed microphone solutions, this significantly improves flexibility and broadens the range of applications.
[0064] For example, before adjusting the horizontal position, vertical position, and focal length of the tracking camera, the method further includes: calculating the distance change value between the speaker's coordinates and historical coordinates; wherein the step of adjusting the horizontal position, vertical position, and focal length of the tracking camera is performed when the distance change value exceeds a distance threshold. The distance threshold can be selected according to actual needs, for example, it can be 5cm.
[0065] The above solution pre-calculates the distance change between the speaker's coordinates and historical coordinates before adjusting the horizontal position, vertical position, and focal length of the tracking camera. The camera position and focal length are only adjusted when the distance change exceeds a preset distance threshold. This effectively avoids frequent and unnecessary camera adjustments caused by slight fluctuations in the speaker's position, significantly reduces the frequency of camera adjustments and system resource consumption, improves tracking stability and image continuity, and extends the lifespan of the equipment.
[0066] For example, the method further includes: displaying the image captured by the detection camera when no speaker is detected using the speaker target detection method; and displaying the image captured by the tracking camera when a speaker is detected using the speaker target detection method.
[0067] In this example, when there is no speaking activity, a panoramic view is displayed; when a speaker is detected, the view switches to a close-up of the speaker; when the speaker changes, the view quickly switches to a close-up of the new speaker. The switching process is smooth and achieves a recording effect that intelligently follows the current speaker.
[0068] The above solution can intelligently switch between panoramic and close-up views. When no one is speaking, it displays a panoramic view, and when someone is speaking, it automatically switches to a close-up view and follows them. It can also quickly adapt when the speaker changes, thus improving the professionalism of video recording and display and the viewing experience.
[0069] In this example, detection cameras and tracking cameras can be grouped together; that is, one detection camera and one tracking camera can be installed at each location. This ensures that every area can be captured and tracked, especially in large scenes. After the speaker is identified, the tracking camera corresponding to the detection camera that shows the speaker's face in the frame can be adjusted. Figure 3 A schematic diagram of a speaker target tracking system according to an embodiment of the present invention is shown. Figure 3 As shown, the application scenario is a large conference room. Three groups of detection and tracking cameras are deployed in different locations. A microphone is positioned in the center of the scene. The panoramic view captured by the detection cameras is used to determine the target visual zone where the speaker is located, while the sound signal collected by the microphone is used for sound source localization to determine the final sound source coordinates. The visual zones determined by both cameras are then matched with the sound source coordinates to obtain the speaker's coordinates. The tracking camera adjusts itself based on these speaker coordinates to center the speaker within its field of view.
[0070] According to another aspect of the present invention, a speaker target detection system is provided, the system comprising: a detection camera, a microphone, and a target detection controller, wherein the detection camera and the microphone are both connected to the target detection controller; the target detection controller is used to perform the speaker target detection method described above.
[0071] According to another aspect of the present invention, a speaker target tracking system is provided, comprising: a speaker target detection system, a tracking camera, and a tracking controller; the speaker target detection system includes a detection camera and a microphone; the detection camera, the microphone, and the tracking camera are all connected to the tracking controller; the tracking controller is used to execute the speaker target tracking method described above. In some embodiments, the tracking controller may be the same controller as the target detection controller, or it may be controlled by two separate controllers.
[0072] According to another aspect of the present invention, a computer-readable storage medium is also provided. The storage medium stores a computer program / instructions that, when executed by a processor, implement the method described above. The storage medium may, for example, include a read-only memory (ROM), an erasable programmable read-only memory (EPROM), a portable compact disc read-only memory (CD-ROM), a USB memory, or any combination of the above storage media. The computer-readable storage medium may be any combination of one or more computer-readable storage media.
[0073] Those skilled in the art will readily understand the implementation structure, working principle, and beneficial effects of electronic devices and computer-readable storage media by reading the above methods. For the sake of brevity, further details will not be elaborated here.
[0074] Although exemplary embodiments have been described herein with reference to the accompanying drawings, it should be understood that the above exemplary embodiments are merely illustrative and are not intended to limit the scope of the invention. Various changes and modifications can be made therein by those skilled in the art without departing from the scope and spirit of the invention. All such changes and modifications are intended to be included within the scope of the invention as claimed in the appended claims.
[0075] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementations should not be considered beyond the scope of this invention.
[0076] In the several embodiments provided by this invention, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another device, or some features may be ignored or not executed.
[0077] Numerous specific details are set forth in the specification provided herein. However, it will be understood that embodiments of the invention may be practiced without these specific details. In some instances, well-known methods, structures, and techniques have not been shown in detail so as not to obscure the understanding of this specification.
[0078] Similarly, it should be understood that, in order to streamline the invention and aid in understanding one or more of the various aspects of the invention, features of the invention are sometimes grouped together in a single embodiment, figure, or description thereof in the description of exemplary embodiments of the invention. However, this approach should not be construed as reflecting an intention that the claimed invention requires more features than are expressly recited in each claim. Rather, as reflected in the corresponding claims, its inventive point lies in solving the corresponding technical problem with fewer features than all of those in a single disclosed embodiment. Therefore, the claims following the detailed description are hereby expressly incorporated into that detailed description, wherein each claim itself is a separate embodiment of the invention.
[0079] Those skilled in the art will understand that, apart from the mutual exclusion of features, all features disclosed in this specification (including the accompanying claims, abstract, and drawings) and all processes or elements of any method or apparatus so disclosed may be combined in any combination. Unless otherwise expressly stated, each feature disclosed in this specification (including the accompanying claims, abstract, and drawings) may be replaced by an alternative feature that serves the same, equivalent, or similar purpose.
[0080] Furthermore, those skilled in the art will understand that although some embodiments described herein include certain features but not others included in other embodiments, combinations of features from different embodiments are intended to be within the scope of the invention and form different embodiments. For example, in the claims, any of the claimed embodiments can be used in any combination.
[0081] The various component embodiments of the present invention can be implemented in hardware, or as software modules running on one or more processors, or a combination thereof. Those skilled in the art will understand that microprocessors or digital signal processors (DSPs) can be used in practice to implement some or all of the functions of some modules in the electronic device according to embodiments of the present invention. The present invention can also be implemented as an apparatus program (e.g., a computer program and computer program product) for performing some or all of the methods described herein. Such programs implementing the present invention can be stored on a computer-readable medium or can be in the form of one or more signals. Such signals can be downloaded from an Internet website, provided on a carrier signal, or provided in any other form.
[0082] It should be noted that the above embodiments are illustrative of the invention and not restrictive, and that those skilled in the art can devise alternative embodiments without departing from the scope of the appended claims. In the claims, any reference signs placed between parentheses should not be construed as limiting the claims. The word "comprising" does not exclude the presence of elements or steps not listed in the claims. The word "a" or "an" preceding an element does not exclude the presence of a plurality of such elements. The invention can be implemented by means of hardware comprising several different elements and by means of a suitably programmed computer. In the unit claims enumerating several means, several of these means may be embodied by the same item of hardware. The use of the words first, second, and third, etc., does not indicate any order. These words can be interpreted as names.
[0083] The above description is merely a specific embodiment of the present invention or an explanation of that embodiment. The scope of protection of the present invention is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in the present invention should be included within the scope of protection of the present invention. The scope of protection of the present invention should be determined by the scope of the claims.
Claims
1. A speaker target detection method, characterized in that, A speaker target detection system, the system comprising a detection camera and a microphone; the method comprising: The detection camera continuously acquires real-time images of the scene; Target detection is performed on the real-time footage of each scene to identify human targets in the real-time footage of each scene. Based on the mouth changes of each human target in the real-time scene of a continuously preset number of frames, determine the target speaker and the target visual partition in which it is located. The sound source is located based on the sound signal collected in real time by the microphone, and the final sound source coordinates are determined. Based on the coordinates of the target visual partition and the final sound source coordinates, the speaker coordinates are determined, where the speaker coordinates are the three-dimensional coordinates of the target speaker's location.
2. The method according to claim 1, characterized in that, The step of determining the target speaker and its corresponding visual partition based on the mouth changes of each human target in a continuously preset number of frames of real-time scene footage includes: For any of the aforementioned human targets Determine the aspect ratio of the human target's mouth in the real-time frame of each of the aforementioned scenes; Based on the magnitude and frequency of the change in the aspect ratio of the human target's mouth in the real-time scene of the scene in the continuous preset number of frames, it is determined whether the human target is speaking. When the human target exhibits speaking behavior, that human target is identified as the target speaker; Preferably, determining the aspect ratio of the human target's mouth in each real-time scene includes: For any of the aforementioned real-time scenes, extract the key points of the human target's mouth and determine the coordinates of the key points of the human target's mouth. The key points of the mouth include the key points of the upper lip, the key points of the lower lip, the key points of the left corner of the mouth, and the key points of the right corner of the mouth. Based on the coordinates of the key points of the human target's mouth, determine the aspect ratio of the human target's mouth in the real-time image of the current scene.
3. The method according to claim 1, characterized in that, The step of locating the sound source and determining the final sound source coordinates based on the sound signal collected in real time by the microphone includes: The sound signal continuously collected by the microphone is subjected to temporal convolution calculation to obtain a clean sound signal; The pure sound signal is calculated using signal time-of-flight technology and time difference positioning technology to obtain the initial sound source coordinates; Calculate the distance between the initial sound source coordinates and different preset sound anchor points; The coordinates of the preset sound anchor point that is closest to the initial sound source coordinates are taken as the final sound source coordinates; The preset sound anchor points are obtained by: establishing a three-dimensional sound coordinate system with the location of the microphone as the origin, and measuring the coordinates of the sound anchor points at each preset speaking position in the space to obtain the coordinates of each preset sound anchor point.
4. The method according to claim 1, characterized in that, Determining the speaker coordinates based on the coordinates of the target visual partition and the final sound source coordinates includes: Determine whether the final sound source coordinates match the coordinates of the target visual partition; When a match is found, the final sound source coordinates are determined to be the speaker coordinates.
5. A speaker target tracking method, characterized in that, A speaker target tracking system, the speaker target tracking system including a speaker target detection system and a tracking camera, the speaker target detection system including a detection camera and a microphone; the method includes: The speaker coordinates are determined using the speaker target detection method according to any one of claims 1-4; Based on the speaker's coordinates, the horizontal position, vertical position, and focal length of the tracking camera are adjusted so that the speaker is located in the center of the tracking camera's image.
6. The method according to claim 5, characterized in that, Before adjusting the horizontal position, vertical position, and focal length of the tracking camera, the method further includes: Calculate the distance change between the speaker's coordinates and the historical coordinates; The step of adjusting the horizontal position, vertical position, and focal length of the tracking camera is performed when the distance change value exceeds a distance threshold.
7. The method according to claim 5, characterized in that, The method further includes: When no speaker is detected using the speaker target detection method, the image captured by the detection camera is displayed; When a speaker is detected using the speaker target detection method, the image captured by the tracking camera is displayed.
8. A speaker target detection system, characterized in that, include: The system includes a detection camera, a microphone, and a target detection controller, wherein the detection camera and the microphone are both connected to the target detection controller. The target detection controller is used to perform the method according to any one of claims 1-4.
9. A speaker target tracking system, characterized in that, include: The method includes a speaker target detection system, a tracking camera, and a tracking controller. The speaker target detection system includes a detection camera and a microphone. The detection camera, the microphone, and the tracking camera are all connected to the tracking controller. The tracking controller is used to perform the method according to any one of claims 5-7.
10. A computer-readable storage medium, characterized in that, The system contains a computer program / instructions that, when executed by a processor, implement the method as described in any one of claims 1-4 or the method as described in any one of claims 5-7.