Automatic conference spokesman positioning method, system and device and medium
By using a microphone array and a pan-tilt camera in tandem and leveraging a neural network to generate a coordinate transformation matrix, the system enables automatic positioning and close-up shooting of speakers in the smart conference room. This solves the maintenance problems caused by equipment relocation and renovation, and enables unattended operation and maintenance of the smart conference room.
Patent Information
- Application Number
- CN202610187041.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-02-09
- Publication Date
- 2026-05-15
AI Technical Summary
In existing technologies, the speaker close-up function of smart conference rooms is easily affected by equipment movement, system expansion and renovation, resulting in a high burden of manual secondary calibration and maintenance, and failing to meet the needs of unattended operation.
It employs a microphone array and a pan-tilt camera working together to collect training sound and image data, and uses a neural network to autonomously generate a coordinate transformation matrix to achieve automatic positioning and close-up shooting, including audio-visual synchronization verification and the application of a lightweight network model.
It enables unattended operation and maintenance of intelligent conference rooms, automatically capturing speakers accurately and switching between close-ups, ensuring smooth remote consultations and meeting recordings, and reducing manual intervention and maintenance costs.
Smart Images

Figure CN122053783A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of intelligent positioning technology, and in particular to a method, system, device and medium for automatic positioning of conference speakers. Background Technology
[0002] As hybrid office practices become the mainstream for businesses, smart meeting rooms, as the core platform, are increasingly crucial for their unattended speaker close-up automatic switching function. They need to accurately capture speakers and quickly switch close-ups to ensure the efficiency of remote collaboration and meeting debriefing.
[0003] Current mainstream solutions rely on the collaboration between ceiling microphone arrays and PTZ cameras. Commonly used commercial microphones, such as the Sennheiser TeamConnect Ceiling 2 and Shure MXA920, can collect voice and output the azimuth angle (DOA) or coordinates of the sound source. Commands are then issued from the central control system to drive the PTZ camera to complete close-up switching, seemingly achieving automation. The core premise of this solution is that the relative coordinates of the microphone array and the PTZ camera are manually calibrated in advance to establish a coordinate mapping relationship. However, in actual use, scenarios such as equipment movement, system expansion, and conference room renovations can easily disrupt the calibration relationship, causing the close-up function to malfunction, requiring manual recalibration by professionals.
[0004] Manual secondary calibration requires additional manpower and time costs, interrupts normal use of meeting rooms, and significantly increases the maintenance burden in multi-meeting scenarios. It may also result in accuracy deviations, affecting the meeting experience. It cannot meet the long-term stable and unattended requirements of smart meeting rooms in a hybrid office environment, and a solution is urgently needed. Summary of the Invention
[0005] This invention provides a method, system, device, and medium for automatically locating conference speakers, in order to solve the problem of high maintenance burden caused by manual secondary calibration.
[0006] This invention discloses an automatic speaker location method for a conference, which is applied to an intelligent conference system. The intelligent conference system includes: at least one microphone array set in a conference room, and at least one PTZ camera.
[0007] The automatic speaker location method for the conference includes: Drive the at least one set of microphone arrays to collect training sound data; drive any one of the pan-tilt cameras to perform a spiral scan covering the speaking area of the conference room to collect training image data; If training voice data is detected in the training sound data, and lip micro-movement data is detected in the training image data simultaneously, then the three-dimensional coordinates of the sound source are obtained based on the training voice data, and the current shooting parameters of the PTZ camera are recorded. The current shooting parameters include at least one of horizontal rotation angle, pitch angle and focal length. The corresponding three-dimensional coordinates of the sound source and the current shooting parameters are used as a set of training samples. Collect N sets of training samples, N≥30, input the N sets of training samples into the neural network for training, and obtain the coordinate transformation matrix of the pan-tilt camera relative to the at least one set of microphone arrays; When there is a speaker in the conference room, the real-time sound source coordinates are obtained through the at least one set of microphone arrays, and the real-time sound source coordinates are transformed through the coordinate transformation matrix to generate the coordinates of the visual search center. A close-up shot of the speaker is taken based on the coordinates of the visual search center.
[0008] Optionally, the step of taking a close-up shot of the speaker based on the coordinates of the visual search center includes: Centered on the coordinates of the visual search center, the image captured by the pan-tilt camera is cropped according to the target detection window to obtain the cropped image; Feature extraction is performed on the captured image to obtain visual features; The visual features are analyzed using a network model to detect the precise bounding box of the corresponding inventor region in the captured image. Obtain the center of the precise bounding box, and drive the target camera to adjust the horizontal rotation angle and / or pitch angle so that the center of the pan-tilt camera's image overlaps with the center of the bounding box.
[0009] Optionally, the step of taking a close-up shot of the speaker based on the coordinates of the visual search center further includes: When multiple PTZ cameras are able to capture the speaker, one is selected as the best camera based on the shooting distance and screen occlusion of each PTZ camera. The speaker was filmed in close-up using the best camera.
[0010] Optionally, the steps of driving the at least one set of microphone arrays to acquire training sound data and driving any one of the pan-tilt cameras to perform a spiral scan covering the speaking area of the conference room to acquire training image data include: The training sound data is used to perform speech detection using a human voice detection model to obtain audio confidence. Determine whether the audio confidence exceeds a preset audio threshold. If so, perform lip micro-movement detection on the central region of the training image data using a lip detection model with a temporal receptive field of 5-7 frames and a spatial receptive field of 3×3 pixels, and obtain the lip confidence. The audio confidence score and the lip confidence score are fused to obtain a comprehensive confidence score; When the overall confidence level exceeds the preset overall threshold, it is determined that the current training sound data includes training human voice data, and simultaneously, the training image data includes lip micro-movement data.
[0011] Optionally, the step of inputting the N sets of training samples into the neural network for training includes: Obtain the initial transformation matrix of the PTZ camera; obtain the theoretical imaging coordinates based on the initial transformation matrix and the three-dimensional coordinates of the sound source; obtain the theoretical shooting parameters based on the theoretical imaging coordinates. The loss value of the theoretical shooting parameters and the current shooting parameters is calculated based on the loss function, which includes a reprojection error term and a regularization constraint term. The neural network is then trained in reverse based on the loss value; When the loss value is stable and less than a preset standard value, and the decrease in value for n consecutive iterations is less than a preset change threshold, n>100, the neural network training is determined to be complete.
[0012] Optionally, the automatic speaker location method further includes: Drive the at least one group of microphone arrays to collect real-time sound data according to a preset cycle, and obtain the real-time prediction parameters corresponding to the real-time sound data according to the coordinate transformation matrix; Obtain the current real-time shooting parameters of the PTZ camera, and calculate the real-time prediction parameters and the real-time reprojection error value of the real-time shooting parameters; If the real-time reprojection error value exceeds the preset error threshold multiple times in a row, it is determined that the PTZ camera has shifted. Reacquire the coordinate transformation matrix of the PTZ camera.
[0013] Optionally, the step of re-acquiring the coordinate transformation matrix of the PTZ camera includes: Collect M sets of training samples, where 20 ≤ M ≤ 30; Freeze the trained neural network, retaining only the output layer for training; The frozen neural network is trained using the training samples. Training is stopped when the reprojection error is ≤2cm, and the update transformation matrix corresponding to the neural network at this time is used as the coordinate transformation matrix.
[0014] The present invention also discloses an intelligent meeting system, comprising: At least one microphone array and at least one pan-tilt camera should be installed in the conference room; A controller connected to each of the microphone arrays and each of the pan-tilt-zoom cameras, the controller being used to implement the automatic speaker positioning method described above.
[0015] The present invention also discloses a storage medium storing a computer program, which, when executed by a processor, causes the processor to perform the steps of the method described above.
[0016] The present invention also discloses an intelligent conferencing device, including a memory and a processor, wherein the memory stores a computer program, and when the computer program is executed by the processor, the processor performs the steps of the method described above.
[0017] The beneficial effects of the automatic speaker positioning method, system, device, and medium provided in this invention are as follows: It automatically drives a microphone array to collect training voice data and a PTZ camera to collect training image data through helical scanning. Valid samples are screened through audio-visual synchronization verification. The system autonomously completes neural network training and generates a coordinate transformation matrix, thus solving the problem of manual recalibration required after equipment movement, expansion, or renovation in existing solutions. This truly achieves unattended operation and maintenance of intelligent conference rooms. From training sample collection and transformation matrix generation to real-time sound source coordinate acquisition, visual search center coordinate transformation, and speaker close-up shooting, the entire process is fully automated. Accurate speaker capture and close-up switching can be completed without manual intervention, ensuring smooth operation in scenarios such as remote consultations and meeting recordings. Attached Figure Description
[0018] The technical solution of the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. In the accompanying drawings: Figure 1 This is a flowchart illustrating an embodiment of the automatic speaker location method provided by the present invention; Figure 2 This is a schematic diagram of the structure of an embodiment of the intelligent meeting system provided by the present invention; Figure 3 This is a schematic diagram of the internal structure of an intelligent conference device in one embodiment of the present invention.
[0019] The labels for the attached figures are as follows: 10. Intelligent conference system; 11. Microphone array; 12. Pan-tilt camera; 13. Controller. Detailed Implementation
[0020] It should be noted that, unless otherwise specified, the embodiments and features described in this application can be combined with each other. The preferred embodiments of the present invention will now be described in detail with reference to the accompanying drawings.
[0021] Please refer to the following: Figure 1 and Figure 2 , Figure 1 This is a flowchart illustrating an embodiment of the automatic speaker location method provided by the present invention. Figure 2 This is a schematic diagram of an embodiment of the intelligent conference system provided by the present invention. The intelligent conference system 10 includes at least one microphone array 11 and at least one PTZ camera 12, installed in a conference room. The microphone array 11 can be ceiling-mounted, without occupying conference table space or obstructing the view, while achieving full conference room voice coverage. Each microphone array consists of multiple high-sensitivity pickup units and has the ability to output real-time three-dimensional coordinates of the sound source. Based on the DOA (Direction of Arrival) algorithm, combined with the time difference and phase difference of the array pickup units, the three-dimensional coordinates of the speaker's sound source can be calculated. The PTZ camera 12 is a PTZ (Pan, Tilt, Zoom) conference camera, which can be flexibly deployed on the walls, ceiling, or edge of the table in the conference room according to the layout and area of the conference room, prioritizing unobstructed locations with a wide field of view to ensure coverage of all areas where people are active in the conference room. Each PTZ camera 12 has high-definition imaging capabilities, capable of capturing clear facial details and lip movements of the speaker, providing high-quality image data for visual detection of lip micro-movements. The intelligent conferencing system 10 also includes a controller 13 that is connected (including wired and wireless connections) to each microphone array 11 and each pan-tilt camera 12. The controller 13 can be an AI-BOX (edge computing box) that integrates multiple technologies such as audio and video decoding, data transmission, and intelligent algorithms.
[0022] The microphone array 11 works in conjunction with the pan-tilt camera 12 and the controller 13 to achieve unattended operation, enabling the speaker to be located as they speak and for the location to be captured in close-up.
[0023] The automatic speaker location method for meetings provided by this invention includes the following steps: S101: Drive at least one set of microphone arrays to collect training sound data; drive any pan-tilt camera to perform spiral scanning to cover the speaking area of the conference room and collect training image data.
[0024] In a specific implementation scenario, a speaker can speak from any area of the conference room. At this time, at least one microphone array is driven to collect training sound data. The pickup units in each microphone array synchronously collect the raw audio stream of the conference room, and perform real-time noise reduction, echo suppression, and dereverberation on the raw audio stream to obtain training sound data.
[0025] Simultaneously, based on the room dimensions and the installation position of the PTZ camera, the scanning parameters for the PTZ camera's spiral scan were calculated in advance. These parameters included the scanning range, spiral step length, scanning speed, and number of spiral rotations. The PTZ camera initiated the scan based on these parameters to acquire raw image data. It first rotated to the initial angle corresponding to the center of the speaking area (e.g., pan=0°, tilt=0°), serving as the starting point for the spiral scan. Following the preset step length and speed, the horizontal rotation angle (pan) and the tilt angle (tilt) were simultaneously adjusted. The pan angle increased / decreased at a uniform speed, while the tilt angle changed sinusoidally with the pan angle, forming a continuous spiral trajectory to ensure the PTZ camera lens could scan every position where a speaker was present. During the scanning process to acquire raw image data, the PTZ camera's built-in high-precision encoder recorded the current horizontal rotation angle (pan) and tilt angle (tilt) in real time, while the autofocus module simultaneously recorded the focal length (zoom). Each frame of raw image data corresponded to a set of (pan, tilt, zoom) shooting parameters. Exposure correction, white balance adjustment, and noise reduction were performed on the acquired raw image data to improve image clarity and acquire training image data. The training image data and shooting parameters are mapped one-to-one.
[0026] S102: If training voice data is detected in the training sound data, and lip micro-movement data is detected in the training image data simultaneously, then the three-dimensional coordinates of the sound source are obtained based on the training voice data, and the current shooting parameters of the PTZ camera are recorded. The current shooting parameters include at least one of encoder data, horizontal rotation angle, pitch angle and focal length. The corresponding three-dimensional coordinates of the sound source and the current shooting parameters are used as a set of training samples.
[0027] In a specific implementation scenario, VAD (Voice Activity Detection) is performed on the training audio data to detect whether human voice data exists in the training audio data. If so, the microphone array obtains the three-dimensional coordinates of the sound source corresponding to the human voice data based on the DOA (Directional Azimuth) algorithm and beamforming technology.
[0028] Specifically, by utilizing the time difference of arrival (TDOA) of sound waves between the array pickup units, the horizontal azimuth (0~360°) and vertical elevation (-45°~+45°) of the human voice signal are calculated. Combined with the energy attenuation model of the human voice signal and the acoustic parameters of the conference room (such as reverberation time), the straight-line distance from the speaker to the microphone array is estimated. The azimuth, elevation, and distance are converted into three-dimensional coordinates in the microphone coordinate system, which are the three-dimensional coordinates of the sound source.
[0029] At this point, it is checked whether the training image data acquired at the same time as the training speech data includes lip micro-movement data. If so, the current shooting parameters (pan, tilt, zoom) corresponding to the training image data are obtained, and the three-dimensional coordinates of the sound source corresponding to the training speech data are obtained. The current shooting parameters and the corresponding three-dimensional coordinates of the sound source are used as a set of training samples.
[0030] In one embodiment, a human voice detection model is used to perform speech detection on the training speech data. This model can accurately extract the core features of human speech (such as Mel spectrum and fundamental frequency) and filter out all non-speech sounds (air conditioner noise, furniture movement, equipment noise, background noise, etc.). The audio confidence score output by the human voice detection model is then obtained. , 0≤ ≤1. Preset audio threshold (e.g., 0.8). If the audio threshold is ≥, then lip movement detection is performed. If the audio threshold is exceeded, it indicates that the probability of a real person speaking is low, and the process terminates.
[0031] like To detect lip movements, the training image data is used to detect lip micro-movements in the central region of the image. This central region is designed to exclude lip movements irrelevant to the edges (such as distant individuals pursing their lips or drinking water), while simultaneously reducing the number of pixels processed and improving detection speed. A spatiotemporal separation convolutional kernel is used for lip micro-movement detection, designed for scenarios where speakers in conference rooms speak softly and lip movements are minimal. The spatiotemporal separation convolutional kernel has a temporal receptive field of 5-7 frames and a spatial receptive field of 3×3 pixels.
[0032] The pan-tilt camera captures images at a frame rate of 30fps, with a temporal window of 200-300ms corresponding to 5-7 frames. This 200-300ms range is the shortest time period for a single lip movement during human speech, allowing for precise capture of continuous, minute lip deformations during pronunciation and preventing missed actions due to single-frame detection. Convolutional feature extraction is performed only on a small 3×3 pixel area of the lip region, representing the optimal spatial scale for capturing lip movements. This approach accurately locates minute pixel changes in the lips (such as pixel shifts from lip opening and closing) while significantly reducing computational load, making it suitable for edge computing.
[0033] Specifically, after training image data is input into the lip detection model, the central region image is acquired. A spatiotemporal separation convolution kernel is used to extract features from the central region image, separating temporal and spatial features. Convolution is performed on 5-7 consecutive frames of the central region image to extract continuous dynamic deformation features of the lips (distinguishing between continuous micro-movements of the lips during speech and static / silent lip movements, such as pursing or licking the lips). A 3×3 pixel convolution is performed on the lip region of each frame to extract subtle spatial morphological changes of the lips (such as pixel shifts from closed to open lips, adapting to the subtle movements of whispering). The lip detection model outputs lip confidence scores. , 0≤ ≤1.
[0034] Ordinary convolution mixes temporal and spatial features simultaneously, making it difficult to capture low-amplitude, high-frequency lip movements. Spatiotemporal separation, on the other hand, can accurately extract dynamic changes in continuous frames and minute deformations in local areas, effectively improving detection accuracy and meeting the needs of meeting room lip movement scenarios.
[0035] The audio confidence score and lip confidence score are fused together, for example by addition, multiplication, or exponential operation, to obtain the overall confidence score. A preset comprehensive threshold (e.g., 0.7) can be set in advance. If the value is ≥0.7, it is considered that the training audio data and training image data collected at the same time contain valid human voice data and lip micro-movement data. The three-dimensional coordinates of the sound source corresponding to the training audio data and the current shooting parameters corresponding to the training image data are obtained as a set of training samples.
[0036] S103: Collect N sets of training samples, N≥30, input the N sets of training samples into the neural network for training, and obtain the coordinate transformation matrix of the pan-tilt camera relative to at least one set of microphone arrays.
[0037] In a specific implementation scenario, N sets of training samples are collected, where N≥30. 30 sets represent the minimum effective sample size, covering different positions (near / far, left / right, front / back) in the conference room speaking area. This ensures the trained transformation matrix generalizes to the entire conference room area and avoids the matrix only fitting a local region due to insufficient samples (e.g., <20 sets). The coordinate transformation matrix is not a simple square matrix, but rather a combination of a rotation matrix and a translation vector representing the spatial transformation relationship between the microphone coordinate system of the microphone array and the camera coordinate system of the PTZ camera. It is the only target to be optimized in neural network training.
[0038] Traditional geometric methods (such as directly solving spatial equations) are extremely sensitive to sample errors. In a conference room environment, there are slight environmental interference errors (such as voice reflection and slight lens shake) in microphone sound source localization and pan-tilt camera angle acquisition. Even a small error can lead to a significant decrease in the accuracy of the calculated coordinate transformation matrix. Neural networks, on the other hand, have a strong nonlinear fitting capability, can automatically learn the small error patterns in the samples, are robust to interference data, and can achieve accurate optimization through self-supervised training (using the actual shooting parameters of the pan-tilt camera as labels), all without human intervention.
[0039] Furthermore, only in neural networks and For trainable parameters, the weights of the remaining networks (such as the coding layer) are lightweight and fixed, and can be set according to engineering experience values to further reduce computing power consumption.
[0040] The neural network is trained using N sets of training samples. After training converges, the final data is extracted from the neural network. and The combination of these two elements forms the coordinate transformation matrix of the PTZ camera relative to the microphone array.
[0041] Furthermore, the extracted coordinate transformation matrix is physically verified to ensure it conforms to the actual conditions of the conference room. It satisfies orthogonality (an inherent physical property of rotation matrices). If the module length is less than or equal to the maximum physical size of the conference room (e.g., if the conference room is 10 meters long, then the module length is less than or equal to 10 meters), then the coordinate transformation matrix can be considered to have passed the verification.
[0042] In one implementation scenario, the initial transformation matrix of the PTZ camera is obtained. The theoretical shooting coordinates in the camera coordinate system are calculated based on the initial transformation matrix and the three-dimensional coordinates of the sound source. The camera intrinsic parameters of the PTZ camera are obtained, and based on these parameters, the two-dimensional pixel plane perspective projection of the theoretical shooting coordinates onto the PTZ camera's image is obtained, yielding the theoretical imaging coordinates. The pixel coordinates of the PTZ camera's image are linearly related to the pan / tilt angle; therefore, the theoretical shooting parameters can be calculated based on the theoretical imaging coordinates. Through the above transformation, the theoretically correct pixel position of the speaker in the image is converted into the angle at which the PTZ camera needs to rotate to capture the speaker at that pixel position.
[0043] The loss function calculates the loss value based on the theoretical and current shooting parameters. It includes a reprojection error term and a regularization constraint term, representing a customized design that balances accuracy requirements with actual physical constraints.
[0044] The loss function can be calculated using the following formula:
[0045] in, This is the reprojection error term. This is a regularization constraint term. These are the current shooting parameters. The three-dimensional coordinates of the sound source and Let be the coordinate transformation matrix to be optimized. These are the theoretical shooting parameters.
[0046] To minimize the loss value, a new set of more accurate transformation matrix parameters is obtained. This is a precise fine-tuning of the initial transformation matrix, which better matches the spatial relationship between microphone and camera coordinates than the previous transformation matrix. The new transformation matrix serves as the initial matrix for the next round, and the above training steps are repeated cyclically. Each iteration yields a better transformation matrix, and the loss value L continuously decreases and gradually stabilizes until the dual convergence criteria are met. The purpose of the dual criteria is to ensure both the actual accuracy of the transformation matrix meets the target and that the training fully converges without fluctuations, avoiding problems such as loss values meeting the target but still oscillating, or loss values stabilizing but with insufficient accuracy.
[0047] The dual convergence criteria are: the loss value is stable and less than a preset standard value, and the decrease in the loss value is less than a preset threshold value for n (n>100) consecutive iterations. A stable loss value less than the preset standard value ensures that the actual accuracy of the transformation matrix meets engineering requirements; that is, the deviation between the theoretical and actual shooting parameters is small enough for the PTZ camera to accurately capture the speaker. The decrease in the loss value for n (n>100) consecutive iterations being less than the preset threshold value ensures that the training has fully converged; the loss value hardly decreases anymore, and the parameters of the transformation matrix hardly need adjustment. Continuing to iterate would only increase computational cost and could even lead to overfitting.
[0048] S104: When there is a speaker in the conference room, obtain the real-time sound source coordinates through at least one microphone array, and transform the real-time sound source coordinates through a coordinate transformation matrix to generate the coordinates of the visual search center.
[0049] In a specific implementation scenario, once the coordinate transformation matrix is obtained, the speaker can be located automatically in real time. When there is a speaker in the conference room, the microphone array captures the speaker's voice signal and calculates the real-time sound source coordinates of the voice signal in the microphone coordinate system. The specific acquisition method is basically the same as the method for obtaining the three-dimensional coordinates of the sound source described above, and will not be repeated here.
[0050] The coordinate transformation matrix obtained through the preceding steps is used to perform a rigid body transformation on the real-time sound source coordinates, mapping them from the microphone coordinate system to the camera coordinate system, thus generating the visual search center coordinates. It is understood that as the speaker moves within the conference room, the real-time sound source coordinates will update in real time, and similarly, the visual search center will dynamically update as the speaker's position changes, ensuring that the visual search center always closely follows the speaker.
[0051] S105: Take close-up shots of the speaker based on the coordinates of the visual search center.
[0052] In a specific implementation scenario, by combining the intrinsic parameters of the PTZ camera and obtaining real-time shooting parameters based on the coordinates of the visual search center, the PTZ camera is driven to shoot based on the real-time shooting parameters, thereby ensuring that the speaker is always in the shooting frame of the PTZ camera.
[0053] The visual search center is defined by its three-dimensional coordinates in camera coordinates. Based on the intrinsic parameter matrix of the PTZ camera, the visual two-dimensional coordinates in the PTZ camera's captured image corresponding to the visual search center are obtained. A preset target detection window size (e.g., 64×64) is used. The image captured from the PTZ camera is cropped according to the target detection window, centered on the visual two-dimensional coordinates, to obtain the cropped image.
[0054] The system extracts visual features from the captured image to uniquely distinguish the speaker from the background and irrelevant individuals. A lightweight keypoint detection algorithm is used to extract 16 core feature points on the speaker's face, including the corners of the eyes, mouth, nose, and jawline. These feature points accurately locate facial areas. Furthermore, specifically for scenarios where the speaker is speaking, dynamic features of the lips are extracted, including the lip line outline, changes in the distance between the upper and lower edges of the lips, and textural details of the lip area. These features further confirm that the face in the image is that of the person speaking, rather than a seated attendee or other irrelevant individual.
[0055] A specially trained lightweight network model analyzes the extracted visual features to accurately identify the speaker's location within the captured image and marks it with a bounding box. The midpoint between the top-left and bottom-right corners of this bounding box is taken as its center within the captured image. The bounding box center coordinates are then compared with the center coordinates of the image currently captured by the PTZ camera to obtain the horizontal and vertical pixel deviations.
[0056] By combining the pixel-to-angle conversion coefficients set at the factory of the PTZ camera, the horizontal and vertical pixel deviations are converted into the PTZ camera's horizontal rotation angle deviation and pitch angle deviation, respectively. The PTZ camera's horizontal rotation angle and pitch angle are then adjusted according to the horizontal rotation angle deviation and pitch angle deviation to ensure that the center of the PTZ camera's image overlaps with the center of the bounding box.
[0057] It should be noted that the above describes the automatic positioning method for a single PTZ camera. When multiple PTZ cameras are installed in a conference room, the above steps need to be performed on each PTZ camera. When a new PTZ camera needs to be installed in the conference room, driving the PTZ camera and microphone array to perform the above steps will automatically complete the calibration of the PTZ camera.
[0058] In other implementation scenarios, after the PTZ camera obtains the coordinate transformation matrix, its position or angle may be slightly offset due to maintenance or cleaning. This will cause errors in the actual shooting. Therefore, in subsequent use, the current coordinate transformation matrix of the PTZ camera can be checked at preset intervals (e.g., one week or one month). If a deviation is found, the coordinate transformation matrix will be automatically corrected to ensure that the PTZ camera can always accurately and automatically position the speaker.
[0059] Drive at least one microphone array to acquire real-time sound data, and execute subsequent steps only if the real-time sound data includes human voice data. If the real-time sound data includes human voice data, obtain the real-time three-dimensional coordinates of the sound source based on the human voice data, transform the three-dimensional coordinates of the sound source to the camera coordinate system based on the already effective coordinate transformation matrix, and obtain the actual prediction parameters of the pan-tilt camera by combining the intrinsic parameter matrix of the pan-tilt camera. Obtain the current actual shooting parameters of the pan-tilt camera.
[0060] Calculate the real-time reprojection error value of the real-time prediction parameters and the real-time shooting parameters. The reprojection error value is the physical distance deviation between the center of the image captured according to the theoretical parameters and the center of the image captured according to the actual parameters on the plane where the speaker is located. The larger the deviation, the lower the matching degree between the coordinate transformation matrix and the current actual position of the PTZ camera.
[0061] To avoid single-shot error exceeding the limit due to sudden environmental interference (such as rapid speaker movement, strong voice reflection, or brief equipment noise), this embodiment adopts a continuous multiple-judgment principle. If the real-time reprojection error value exceeds the preset error threshold multiple times (e.g., 3-5 times), it is determined that the PTZ camera has shifted.
[0062] Specifically, a sliding error queue (e.g., a queue containing 5 elements) can be maintained. Each time a reprojection error value is obtained, it is added to the sliding error queue. If all reprojection error values in the queue exceed a preset error threshold, the current coordinate transformation matrix is considered invalid. If any reprojection error value in the queue exceeds the preset error threshold, the queue is cleared and accumulation begins again.
[0063] If the current coordinate transformation matrix is determined to be invalid, a new coordinate transformation matrix is obtained. Based on the sample collection steps similar to those described above, M sets of training samples are collected, where 20 ≤ M ≤ 30, to avoid redundant data consuming computing power.
[0064] The trained neural network is frozen, retaining only the output layer as trainable, while setting the weights of all hidden layers to non-trainable. During training, the weights of the hidden layers are no longer updated, fully reusing the general features learned from the initial calibration. Since only the parameters of the output layer need to be optimized, the training speed and efficiency are greatly improved. By reusing the general features from the previous calibration, even with only 20-30 small samples, the updated matrix can be guaranteed to have good generalization ability across the entire conference room area, rather than only adapting to a local region.
[0065] The same loss function as described above is used for model training. During training, the average reprojection error of the entire sample is calculated every m epochs (e.g., m=10). Training stops when the average reprojection error of the entire sample is less than or equal to 2 cm. The final update transformation matrix is extracted from the output layer and used to replace the original coordinate transformation matrix as the new coordinate transformation matrix.
[0066] As can be seen from the above description, in this embodiment, each PTZ camera can independently perform displacement monitoring and incremental calibration without interfering with each other.
[0067] As described above, this embodiment automatically drives the microphone array to collect training voice data and drives the PTZ camera to perform helical scanning to collect training image data. Valid samples are screened through audio-visual synchronization verification. The system autonomously completes neural network training and generates a coordinate transformation matrix, eliminating the need for manual recalibration after equipment movement, expansion, or renovation, as in previous solutions. This truly achieves unattended operation and maintenance of the smart conference room. From training sample collection and transformation matrix generation to real-time sound source coordinate acquisition, visual search center coordinate transformation, and speaker close-up shooting, the entire process is fully automated. Accurate speaker capture and close-up switching can be achieved without manual intervention, ensuring smooth operation in scenarios such as remote consultations and meeting recordings.
[0068] Figure 3 A schematic diagram of the internal structure of a smart conferencing device in one embodiment is shown. Figure 3As shown, the intelligent conferencing device includes a processor, a memory, and a network interface connected via a system bus. The memory includes a non-volatile storage medium and internal memory. The non-volatile storage medium stores an operating system and may also store a computer program. When executed by the processor, this computer program enables the processor to implement a method for automatically locating conference speakers. The internal memory may also store a computer program, which, when executed by the processor, enables the processor to implement the method for automatically locating conference speakers. Those skilled in the art will understand that… Figure 3 The structure shown is merely a block diagram of a portion of the structure related to the solution of this application and does not constitute a limitation on the intelligent interactive device to which the solution of this application is applied. Specific intelligent conferencing devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.
[0069] In one embodiment, an intelligent interactive device is provided, including a memory and a processor, wherein the memory stores a computer program, and when the computer program is executed by the processor, the processor performs the steps described above.
[0070] In one embodiment, a computer-readable storage medium is provided storing a computer program that, when executed by a processor, causes the processor to perform the steps described above.
[0071] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), RAMbus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and RAMbus dynamic RAM (RDRAM), etc.
[0072] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0073] It should be understood that the above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit them. Those skilled in the art can modify the technical solutions described in the above embodiments, or make equivalent substitutions for some of the technical features; and all such modifications and substitutions should fall within the protection scope of the appended claims of the present invention.
Claims
1. A method for automatically locating a speaker at a conference, characterized in that, The system is applied to an intelligent conference system, which includes at least one microphone array installed in a conference room and at least one pan-tilt camera. The automatic speaker location method for the conference includes: Drive the at least one set of microphone arrays to collect training sound data; drive any one of the pan-tilt cameras to perform a spiral scan covering the speaking area of the conference room to collect training image data; If training voice data is detected in the training sound data, and lip micro-movement data is detected in the training image data simultaneously, then the three-dimensional coordinates of the sound source are obtained based on the training voice data, and the current shooting parameters of the PTZ camera are recorded. The current shooting parameters include at least one of horizontal rotation angle, pitch angle and focal length. The corresponding three-dimensional coordinates of the sound source and the current shooting parameters are used as a set of training samples. Collect N sets of training samples, N≥30, input the N sets of training samples into the neural network for training, and obtain the coordinate transformation matrix of the pan-tilt camera relative to the at least one set of microphone arrays; When there is a speaker in the conference room, the real-time sound source coordinates are obtained through the at least one set of microphone arrays, and the real-time sound source coordinates are transformed through the coordinate transformation matrix to generate the coordinates of the visual search center. A close-up shot of the speaker is taken based on the coordinates of the visual search center.
2. The automatic speaker location method according to claim 1, characterized in that, The step of taking a close-up shot of the speaker based on the coordinates of the visual search center includes: Centered on the coordinates of the visual search center, the image captured by the pan-tilt camera is cropped according to the target detection window to obtain the cropped image; Feature extraction is performed on the captured image to obtain visual features; The visual features are analyzed using a network model to detect the precise bounding box of the corresponding inventor region in the captured image. Obtain the center of the precise bounding box, and drive the target camera to adjust the horizontal rotation angle and / or pitch angle so that the center of the pan-tilt camera's image overlaps with the center of the bounding box.
3. The automatic speaker positioning method according to claim 2, characterized in that, The step of taking a close-up shot of the speaker based on the coordinates of the visual search center further includes: When multiple PTZ cameras are able to capture the speaker, one is selected as the best camera based on the shooting distance and screen occlusion of each PTZ camera. The speaker was filmed in close-up using the best camera.
4. The automatic speaker positioning method according to claim 1, characterized in that, The driver is used to collect training sound data from at least one set of microphone arrays; The steps of driving any one of the pan-tilt cameras to perform spiral scanning to acquire training image data covering the speaking area of the conference room include: The training sound data is used to perform speech detection using a human voice detection model to obtain audio confidence. Determine whether the audio confidence exceeds a preset audio threshold. If so, perform lip micro-movement detection on the central region of the training image data using a lip detection model with a temporal receptive field of 5-7 frames and a spatial receptive field of 3×3 pixels, and obtain the lip confidence. The audio confidence score and the lip confidence score are fused to obtain a comprehensive confidence score; When the overall confidence level exceeds the preset overall threshold, it is determined that the current training sound data includes training human voice data, and simultaneously, the training image data includes lip micro-movement data.
5. The automatic speaker positioning method according to claim 1, characterized in that, The step of inputting the N sets of training samples into the neural network for training includes: Obtain the initial transformation matrix of the PTZ camera; obtain the theoretical imaging coordinates based on the initial transformation matrix and the three-dimensional coordinates of the sound source; obtain the theoretical shooting parameters based on the theoretical imaging coordinates. The loss value of the theoretical shooting parameters and the current shooting parameters is calculated based on the loss function, which includes a reprojection error term and a regularization constraint term. The neural network is then trained in reverse based on the loss value; When the loss value is stable and less than a preset standard value, and the decrease in value for n consecutive iterations is less than a preset change threshold, n>100, the neural network training is determined to be complete.
6. The automatic speaker positioning method according to claim 5, characterized in that, The automatic speaker location method for the meeting also includes: Drive the at least one group of microphone arrays to collect real-time sound data according to a preset cycle, and obtain the real-time prediction parameters corresponding to the real-time sound data according to the coordinate transformation matrix; Obtain the current real-time shooting parameters of the PTZ camera, and calculate the real-time prediction parameters and the real-time reprojection error value of the real-time shooting parameters; If the real-time reprojection error value exceeds the preset error threshold multiple times in a row, it is determined that the PTZ camera has shifted. Reacquire the coordinate transformation matrix of the PTZ camera.
7. The automatic speaker positioning method according to claim 6, characterized in that, The step of re-acquiring the coordinate transformation matrix of the PTZ camera includes: Collect M sets of training samples, where 20 ≤ M ≤ 30; Freeze the trained neural network, retaining only the output layer for training; The frozen neural network is trained using the training samples. Training is stopped when the reprojection error is ≤2cm, and the update transformation matrix corresponding to the neural network at this time is used as the coordinate transformation matrix.
8. An intelligent conference system, characterized in that, include: At least one microphone array and at least one pan-tilt camera should be installed in the conference room; A controller connected to each of the microphone arrays and each of the pan-tilt cameras, the controller being configured to implement the automatic speaker positioning method as described in any one of claims 1-7.
9. A storage medium, characterized in that, The device stores a computer program that, when executed by a processor, causes the processor to perform the steps of the method as described in any one of claims 1 to 7.
10. A smart conferencing device, characterized in that, It includes a memory and a processor, the memory storing a computer program that, when executed by the processor, causes the processor to perform the steps of the method as described in any one of claims 1 to 7.