Conference speaker tracking method and device, electronic equipment and storage medium
By combining audio and video signals, the location of the sound source and the state of lip movement are determined, and the speaker in the video conference is accurately located. This solves the problems of camera perspective switching and inaccurate focus caused by multiple participants speaking alternately or environmental interference, and improves the continuity of the video conference.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- RUISHI (GUANGZHOU) INFORMATION TECH CO LTD
- Filing Date
- 2026-02-16
- Publication Date
- 2026-05-15
AI Technical Summary
In existing video conferences, alternating speeches by multiple participants or environmental reverberation can cause changes or drifts in sound source localization results. Cameras frequently switch perspectives or focus on locations other than the actual speaker, affecting the continuity of the video feed.
The location information of the sound source is determined by the audio acquisition device, and the initial set of face regions is extracted by the video acquisition device. Based on the lip movement information and the location information of the sound source, the time synchronization is analyzed to accurately lock the effective speaker target and control the imaging center of the video acquisition device to be aligned with the speaker.
It ensures video conferencing continuity even when multiple participants are speaking alternately or when there is ambient reverberation interference, avoiding frequent camera angle switching and focus deviation.
Smart Images

Figure CN122053784A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of computer technology, and in particular to a method, apparatus, electronic device, and storage medium for tracking conference speakers. Background Technology
[0002] In existing conference scenarios, a common approach to achieve real-time speaker tracking is to use a scheme that links fixed-area sound source localization with a preset camera viewing angle. This method detects the direction of the current sound source using a microphone array and drives the PTZ camera to turn in that direction to capture the image.
[0003] However, existing methods have a significant problem: when multiple participants speak alternately in a short period of time or when there is environmental reverberation interference, the sound source localization results are prone to jumps or drifts. This causes the camera to frequently switch perspectives or focus on the location of a non-speaker, making it impossible to stably and continuously track the actual speaker. This instability seriously affects the continuity of the video conference. Summary of the Invention
[0004] This invention provides a method, apparatus, electronic device, and storage medium for tracking conference speakers, in order to improve the continuity of video conferencing.
[0005] In a first aspect, the present invention provides a method for tracking conference speakers, including: The sound source location information is determined based on the multi-channel audio signals collected by the audio acquisition device, and the initial face region set under each viewpoint is extracted based on the multi-view video frame sequence collected by the video acquisition device. Based on the sound source location information and the initial set of face regions, a spatial mapping is performed to obtain a set of target face regions located within the sound source location coverage area. Based on each face region in the target face region set, lip region features are extracted from the video frames of the corresponding viewpoint to obtain the lip movement state information of each speaking target. Based on the lip movement state information and the sound source location information, a time synchronization analysis is performed to obtain the effective speaker target. Based on the spatial position of the effective speaker target in the current video frame, the imaging center of at least one video acquisition device is controlled to align with the effective speaker target.
[0006] Secondly, the present invention also provides a conference speaker tracking device, applied to the conference speaker tracking method as described in the first aspect; the conference speaker tracking device includes: The audio and video analysis module is used to determine the location information of the sound source based on the multi-channel audio signals collected by the audio acquisition device, and to extract the initial set of face regions from each viewpoint based on the multi-view video frame sequence collected by the video acquisition device. The spatial mapping module is used to perform spatial mapping based on the sound source orientation information and the initial set of face regions to obtain a set of target face regions located within the sound source orientation coverage area; The lip feature recognition module is used to extract lip region features from video frames of corresponding viewpoints based on each face region in the target face region set to obtain lip movement state information of each speaking target. The target tracking and positioning module is used to perform time synchronization analysis based on the lip movement state information and the sound source orientation information to obtain the effective speaker target, and to control the imaging center of at least one video acquisition device to align with the effective speaker target based on the spatial position of the effective speaker target in the current video frame.
[0007] Thirdly, the present invention also provides an electronic device, comprising: a memory for storing computer software programs; and a processor for reading and executing the computer software programs, thereby implementing the conference speaker tracking method described above.
[0008] Fourthly, the present invention also provides a non-transitory computer-readable storage medium storing a computer software program that, when executed by a processor, implements the conference speaker tracking method described above.
[0009] Fifthly, the present invention also provides a computer program product, including a computer program that, when executed by a processor, implements the conference speaker tracking method described above.
[0010] The speaker tracking method provided in this invention performs spatial mapping between sound source location information and an initial set of face regions to obtain a set of target face regions located within the sound source location coverage area. Therefore, by establishing a correlation between audio location and video faces through spatial mapping, candidate faces related to the sound source location are filtered from a massive initial face region set, effectively eliminating interference from faces unrelated to the current sound source and avoiding the problem of misselection of irrelevant targets caused by slight sound source drift. Based on the temporal synchronization analysis of lip movement state information and sound source location information, valid speaker targets are obtained. Therefore, synchronization verification distinguishes the real speaker from interfering sound sources. When the sound source location information and lip movement state information are consistent in time, the speaker can be identified as valid; otherwise, it is determined to be an interfering target caused by sound source jumps or drifts, accurately locking onto the actual speaker target. Based on the spatial position of the effective speaker target in the current video frame, the imaging center of at least one video acquisition device is controlled to align with the effective speaker target. This avoids problems such as frequent camera perspective switching and focus deviation caused by abnormal sound source localization. It solves the problems of sound source localization jumps and drifts, frequent camera perspective switching and inaccurate focus caused by multiple participants speaking alternately or environmental reverberation interference, thereby improving the continuity of video conference footage. Attached Figure Description
[0011] Figure 1 This is a flowchart illustrating the conference speaker tracking method provided in an embodiment of the present invention; Figure 2 This is a schematic diagram of the conference speaker tracking device provided in an embodiment of the present invention; Figure 3 An embodiment diagram of the electronic device provided in this invention; Figure 4 An embodiment diagram of a computer-readable storage medium provided in accordance with the present invention. Detailed Implementation
[0012] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0013] Optionally, see Figure 1 , Figure 1 This is a flowchart illustrating the speaker tracking method provided by the present invention. In this embodiment, the executing entity of the speaker tracking method is a meeting management device. Therefore, the speaker tracking method includes: Step 10: Determine the sound source location information based on the multi-channel audio signals acquired by the audio acquisition device, and extract the initial face region set under each viewpoint based on the multi-view video frame sequence acquired by the video acquisition device.
[0014] Optionally, the conference management device controls the audio acquisition device to start and acquire multi-channel audio signals. The audio acquisition device is a device with multiple pickup units. The multiple pickup units are arranged according to a preset rule, which is that the spacing between the pickup units is uniform and not less than 5 centimeters, to ensure that the differences in audio signals emitted by sound sources from different directions can be captured. The multi-channel audio signal refers to the audio signal corresponding to the same sound event acquired by each pickup unit in the audio acquisition device. Each pickup unit corresponds to an independent audio channel, and the audio signal of each channel carries information such as the intensity and propagation time of the sound emitted by the sound source.
[0015] Furthermore, the conference management device preprocesses the audio signal of each channel. The preprocessing process includes noise reduction and signal amplification. The noise reduction adopts an environmental noise cancellation method. The conference management device first collects the environmental noise signal when the audio acquisition device does not capture the sound source signal. This environmental noise signal is used as the reference noise. Then, the reference noise is subtracted from the multi-channel audio signal collected by each channel to remove environmental interference. The signal amplification adopts a fixed gain amplification method to amplify the audio signal of each channel after noise reduction to a preset signal strength range, which is 0.5 volts to 2 volts.
[0016] Furthermore, the conference management device determines the sound source's azimuth information based on multi-channel audio signals. The specific process is as follows: by comparing the arrival time difference and intensity difference of the same sound signal in each channel's audio signal, and combining this with the preset arrangement of each pickup unit in the audio acquisition device, the horizontal and vertical azimuth angles of the sound source relative to the audio acquisition device are calculated, thereby determining the specific azimuth information of the sound source. Here, the arrival time difference refers to the time difference between the arrival of the same sound signal at different pickup units, and the intensity difference refers to the difference in signal intensity acquired by the same sound signal at different pickup units. The sound source's azimuth information refers to the specific coordinates of the sound source in the conference scene space. These coordinates are based on a preset origin point of the conference scene, which is the spatial coordinate origin preset by the conference management device, typically set at the center of the conference scene. The azimuth information specifically includes the horizontal distance, horizontal azimuth angle, and vertical azimuth angle of the sound source relative to the preset origin point, clearly indicating the sound source's specific location within the conference scene.
[0017] Furthermore, the conference management device controls the video acquisition device to start and acquire multi-view video frame sequences. The video acquisition device consists of multiple cameras with image acquisition capabilities. These cameras are arranged in different positions in the conference scene according to preset angles, with each camera corresponding to an independent acquisition angle, ensuring full coverage of the entire conference scene without any blind spots. The multi-view video frame sequence refers to the set of continuous image frames acquired by each camera at a preset frame rate of 25 to 30 frames per second. The continuous image frames acquired by each camera are arranged in chronological order of acquisition time to form the video frame sequence under that angle. The multi-view video frame sequence is the set of video frame sequences acquired by all cameras.
[0018] Furthermore, the conference management device extracts the initial face region set frame by frame for each video frame sequence from each viewpoint. The initial face region set refers to the set of all regions containing facial features in a single video frame from a single viewpoint, where facial features include the contour features of facial organs such as eyebrows, eyes, nose, and mouth. Optionally, the extraction process in this embodiment of the invention is as follows: The conference management device preprocesses single-frame video frames under each viewpoint. The preprocessing includes grayscale processing and image enhancement processing. Grayscale processing converts color video frames into grayscale images, reducing the amount of data processed by the image. Image enhancement processing uses contrast enhancement to improve the contrast between the face region and the background region in the grayscale image, facilitating subsequent facial feature recognition. After preprocessing, a face detection algorithm is used to scan the grayscale image. The face detection algorithm identifies regions in the image that conform to facial contour features and facial organ distribution features, and determines these regions as initial face regions. All detected initial face regions in a single-frame video frame are summarized to form the initial face region set for that video frame under that viewpoint. The above extraction process is performed on all video frames under that viewpoint to obtain the initial face region set corresponding to all video frames under that viewpoint, thus completing the extraction of the initial face region set under each viewpoint.
[0019] In one embodiment, the preset origin is set at the center of the conference room. The audio acquisition device is a microphone array with four pickup units, evenly arranged on a circular bracket with a diameter of 10 cm and a spacing of 7.85 cm, conforming to the preset arrangement rules. The video acquisition device consists of three cameras, installed at the front, left, and right sides of the conference room respectively. The three cameras complement each other, fully covering the entire conference room without blind spots, and the preset frame rate is set to 30 frames per second. After the meeting starts, the audio and video acquisition devices are activated simultaneously. The four pickup units of the audio acquisition device simultaneously acquire sound signals from the conference scene, forming four-channel audio signals. The conference management device first acquires the ambient noise signal (such as air conditioning noise and slight external interference noise) when no one is speaking in the conference room, as the reference noise. The reference noise is subtracted from the four-channel audio signals to complete the noise reduction process. Then, the noise-reduced audio signals of each channel are amplified to 1 volt, within the preset signal strength range (0.5 volts to 2 volts).
[0020] Subsequently, the conference management device compares the arrival time difference and intensity difference of the same speaking sound (such as the voice of a participant speaking) in the four audio signals. For example, the speaking sound arrives at the front-end pickup unit 0.002 seconds earlier than it arrives at the rear-end pickup unit, and the signal strength arriving at the left pickup unit is 0.3 volts higher than the signal strength arriving at the right pickup unit. Combining the arrangement of the four pickup units, the device calculates that the horizontal distance of the sound source relative to the preset origin (center of the conference room) is 3 meters, the horizontal azimuth angle is 30 degrees, and the vertical azimuth angle is 0 degrees, thus determining the sound source's directional information.
[0021] Simultaneously, three cameras capture video frame sequences at a frame rate of 30 frames per second. The front-end camera captures video frame sequences from the front area of the conference room, the left-side camera captures video frame sequences from the left side of the conference room, and the right-side camera captures video frame sequences from the right side of the conference room. For each video frame captured by the front-end camera, the conference management device first converts the color video frame into a grayscale image, then enhances the contrast of the grayscale image, and then scans the grayscale image using a face detection algorithm. Two regions containing facial features are detected, namely the face regions of attendee A and attendee B. These two regions are combined to form the initial face region set for that video frame from the front-end perspective. The same method is used to process the single-frame video frames captured by the left and right cameras. One face region (attendee C) is detected in the video frame from the left camera, and one face region (attendee D) is detected in the video frame from the right camera, forming the initial face region sets for that video frame from the left and right perspectives, respectively. The above operations are performed on all video frames, ultimately obtaining the initial face region sets for each of the three perspectives.
[0022] Step 20: Spatial mapping is performed based on the sound source orientation information and the initial set of face regions to obtain the set of target face regions located within the sound source orientation coverage area.
[0023] Optionally, based on preset spatial mapping rules, the initial face region sets from various perspectives are mapped to the three-dimensional space of the conference scene to obtain the position coordinates of each initial face region in the three-dimensional space. Then, the conference management device determines whether the position coordinates of each initial face region in the three-dimensional space are within the spatial area covered by the sound source orientation information. The sound source orientation coverage area refers to a spherical area with a preset radius centered on the position coordinates corresponding to the sound source orientation information. The preset radius is determined according to the size of the conference scene, typically set to 1 to 1.5 meters, to ensure that a reasonable range around the sound source location is covered. All initial face regions within the sound source orientation coverage area are aggregated to form a target face region set. The target face region set is a subset of the initial face region set, containing only face regions related to the sound source orientation, as described in steps 201 to 204.
[0024] Step 30: Extract lip region features from the video frames of the corresponding viewpoints for each face region in the target face region set to obtain lip movement state information for each speaking target.
[0025] Optionally, for each face region in the target face region set, the conference management device determines the acquisition viewpoint corresponding to the face region and the video frame under the corresponding viewpoint. The corresponding acquisition viewpoint refers to the camera viewpoint corresponding to the initial face region set where the face region is located, and the corresponding video frame refers to the single video frame where the face region is located, that is, the original video frame when the face region is extracted.
[0026] Furthermore, the conference management device extracts lip region features from the video frames of the corresponding viewpoint for the face region. The lip region features refer to the contour features of the lips, the opening and closing state features of the lips, and the movement trajectory features of the lips. The lip contour features include the width, height, and contour curve of the lips; the lip opening and closing state features include different states such as closed lips, slightly open lips, moderately open lips, and fully open lips; and the lip movement trajectory features refer to the trajectory of the lips' position changes in adjacent video frames.
[0027] Optionally, the specific process of extracting lip region features in this embodiment of the invention is as follows: The conference management device performs region positioning on the video frame containing the target face region, and segments the target face region from the background region of the video frame. The segmentation method adopts the threshold segmentation method. By setting a preset grayscale threshold, the face region (grayscale value within the preset threshold range) is separated from the background region (grayscale value exceeding the preset threshold range). The preset grayscale threshold is determined according to the skin color features of the face to ensure that the segmented face region is complete and free from background interference.
[0028] Furthermore, within the segmented face region, the lip region is located. By identifying the color features (lip color differs significantly from the skin tone of other parts of the face, typically red, pink, etc.) and contour features (lip contour is a closed or semi-closed curve) of the lips in the face region, the specific range of the lip region is determined. This range is then segmented from the face region to obtain an independent lip region. Finally, feature extraction is performed on the segmented lip region, extracting contour features such as the lip contour curve, width, and height. By comparing the distance between the upper and lower edges of the lip region, the opening and closing state features of the lips are determined. By tracking the positional changes of the lip region in adjacent video frames, the motion trajectory features of the lips are extracted. All extracted lip region features are summarized to form the lip region features corresponding to the target face region.
[0029] For each face region in the target face region set, the above-mentioned lip region feature extraction process is performed. Simultaneously, for adjacent video frames corresponding to the same target face region, lip region features are continuously extracted. Combined with the acquisition time of each video frame, lip motion state information corresponding to that target face region is formed. Lip motion state information refers to the changes in lip movement of the target face region's lips in a continuous video frame sequence, including changes in lip opening and closing states and movement trajectory.
[0030] In one embodiment, the target face region set contains only the face region of participant A. From the video frame captured by the front-end camera, the meeting management device uses a threshold segmentation method for the video frame, setting the grayscale threshold to 80 to 180 (determined according to the skin color characteristics of participant A), to segment the face region of participant A from the background of the video frame. The segmented face region is complete and free from background interference. Subsequently, in the segmented face region, the lip region is located by recognizing the pink features and closed curve contour of the lips, and the lip region is segmented from the face region to obtain an independent lip region.
[0031] Subsequently, the meeting management device extracted the contour features of the lip area, determining that the lip width was 2 cm, the height was 1 cm, and the contour curve was a closed curve. By comparing the distance between the upper and lower edges of the lips (0.1 cm), it was determined that the lips were in a closed state. At the same time, the lip area of participant A was extracted from three adjacent video frames of the front-end camera (with an acquisition time interval of 1 / 30 second). The comparison revealed that the position of the lip area did not change significantly, and the movement trajectory was stationary. These features were summarized to form the lip area features of participant A.
[0032] The system continuously tracks the continuous video frames captured by the front-end camera. For each frame, it extracts the lip region features of participant A. Combined with the acquisition time of each frame, it forms the lip movement state information of participant A. For example: in the first to fifth video frames, the lips are closed with no obvious movement; in the sixth to fifteenth video frames, the lips gradually open to a moderately open state with an upward movement trajectory; in the sixteenth to twenty-fifth video frames, the lips remain moderately open with slight up-and-down movement; in the twenty-sixth to thirtyth video frames, the lips gradually close with a downward movement trajectory. This part of the information is the lip movement state information of participant A.
[0033] Step 40: Perform time synchronization analysis based on lip movement state information and sound source orientation information to obtain the effective speaker target, and control the imaging center of at least one video acquisition device to align with the effective speaker target based on the spatial position of the effective speaker target in the current video frame.
[0034] Optionally, the conference management device performs time synchronization analysis based on lip movement state information and sound source location information to obtain effective speaker targets, as described in steps 401 to 404.
[0035] Furthermore, the conference management device obtains the spatial position of the valid speaker target in the current video frame. The current video frame refers to the latest video frame acquired during the time synchronization analysis. The spatial position of the valid speaker target in the current video frame is the position coordinate of the target face region corresponding to the valid speaker target in the three-dimensional space of the conference scene.
[0036] Subsequently, based on this spatial location, the conference management device controls at least one video capture device to align its imaging center with the valid speaker target. Here, the imaging center refers to the center point of the lens image of the video capture device (camera). Aligning with the valid speaker target means adjusting the camera's shooting angle and focal length so that the valid speaker target's face area is centered in the camera's field of view and clearly visible, facilitating subsequent tracking, shooting, and image acquisition of the valid speaker target. Optionally, the control process in this embodiment of the invention is as follows: The conference management device first calculates the angle difference and distance difference between the spatial location of the valid speaker target and the spatial location corresponding to the current imaging center of the video capture device to be controlled. The angle difference includes a horizontal angle difference and a vertical angle difference. The horizontal angle difference refers to the difference between the horizontal azimuth angle of the valid speaker target's spatial location and the horizontal azimuth angle of the camera's current imaging center; the vertical angle difference refers to the difference between the vertical azimuth angle of the valid speaker target's spatial location and the vertical azimuth angle of the camera's current imaging center; the distance difference refers to the straight-line distance between the valid speaker target's spatial location and the camera.
[0037] Subsequently, the conference management device sends control commands to the corresponding video acquisition device. The control commands include angle adjustment parameters and focal length adjustment parameters. The angle adjustment parameters are determined based on the calculated horizontal and vertical angle differences and are used to adjust the horizontal and vertical shooting angles of the camera so that the imaging center of the camera is aligned with the spatial position of the valid speaker target.
[0038] The focal length adjustment parameter is determined based on the calculated distance difference and is used to adjust the camera's shooting focal length so that the face area of the effective speaker target reaches a preset clarity in the shooting image. The preset clarity is such that the facial organs of the face can be clearly identified, ensuring the effectiveness of the shooting image.
[0039] The video acquisition device adjusts its own shooting angle and shooting focal length according to the angle adjustment parameters and focal length adjustment parameters in the control command, and completes the operation of aligning the imaging center with the valid speaker target. If there are multiple valid speaker targets, one of the highest priority valid speaker targets can be selected according to the preset priority rules, and at least one video capture device can be controlled to be aimed at that valid speaker target; or multiple video capture devices can be controlled to be aimed at different valid speaker targets respectively, to ensure that each valid speaker target can be clearly captured. The specific control method is determined according to the preset settings of the meeting management device.
[0040] In one embodiment, through time synchronization analysis, participant A is determined to be a valid speaker target. The meeting management device obtains the spatial position of participant A in the current video frame. The three-dimensional coordinates corresponding to this spatial position are: a horizontal distance of 3 meters, a horizontal azimuth angle of 30 degrees, and a vertical azimuth angle of 0 degrees relative to the preset origin of the meeting scene. The meeting management device selects and controls the imaging center of the front-end camera (the camera in the view of participant A) to be aligned with participant A. First, it calculates the angular difference and distance difference between the spatial position of participant A and the current imaging center of the front-end camera: the calculated horizontal angular difference is 5 degrees, the vertical angular difference is 0 degrees, and the distance difference is 3 meters.
[0041] Subsequently, the conference management device sends control commands to the front-end camera. The angle adjustment parameters in the control commands are: increase the horizontal shooting angle by 5 degrees, and keep the vertical shooting angle unchanged; the focal length adjustment parameters are: adjust the shooting focal length to the focal length corresponding to 3 meters (determined according to the focal length range of the front-end camera. Assuming that the focal length range of the front-end camera is 1 meter to 10 meters, the focal length corresponding to a distance of 3 meters is 50 millimeters).
[0042] After receiving the control command, the front-end camera increases the horizontal shooting angle by 5 degrees while keeping the vertical shooting angle unchanged. At the same time, the shooting focal length is adjusted to 50 mm. After the adjustment, the face area of participant A is in the center of the front-end camera's shooting image, and the face area is clearly visible, and the facial organs can be clearly identified, thus completing the image center alignment operation.
[0043] The embodiments of the present invention solve the problems of sound source localization jumps and drifts caused by multiple participants speaking alternately or environmental reverberation interference, as well as frequent camera perspective switching and inaccurate focus, thereby improving the continuity of video conference footage.
[0044] Optionally, the processes of steps 201 to 204 include: Step 201: Based on the sound source horizontal angle range and sound source vertical angle range in the sound source orientation information, determine the cone-shaped coverage area of the sound source in three-dimensional space.
[0045] Optionally, the meeting management device acquires sound source location information including the sound source horizontal azimuth angle, the sound source vertical azimuth angle, and the horizontal distance of the sound source relative to the preset origin of the meeting scene. The preset origin is the spatial coordinate origin of the meeting scene preset by the meeting management device, which is usually set at the center of the meeting scene. The sound source horizontal azimuth angle refers to the angle between the projection of the line connecting the sound source and the preset origin on the horizontal plane and the preset horizontal baseline. The sound source vertical azimuth angle refers to the angle between the line connecting the sound source and the preset origin and the preset horizontal baseline.
[0046] Furthermore, the conference management device determines the cone-shaped coverage area of the sound source in three-dimensional space based on the horizontal and vertical angle ranges of the sound source location information. The horizontal angle range refers to the angular interval formed by extending a preset horizontal angle to both sides of the horizontal azimuth angle of the sound source. The preset horizontal angle is determined according to the sound pickup accuracy requirements of the conference scenario and is typically set to 5 to 10 degrees to ensure coverage of any slight horizontal offset that the sound source may have. The vertical angle range refers to the angular interval formed by extending a preset vertical angle upwards and downwards of the vertical azimuth angle of the sound source. The preset vertical angle is typically set to 3 to 5 degrees to cover any slight vertical offset that the sound source may have.
[0047] Optionally, the process of determining the cone-shaped coverage area in this embodiment of the invention is as follows: taking the actual position of the sound source in three-dimensional space as the vertex, taking the horizontal angle range and the vertical angle range of the sound source as the vertex angle parameters of the cone, and taking a preset length as the generatrix length of the cone, a three-dimensional cone-shaped spatial area is formed. The generatrix length is determined according to the size of the meeting scene, and is usually set to half the length of the longest diagonal of the meeting scene to ensure that the cone-shaped coverage area can completely cover the location of the sound source and the surrounding possible speaking target areas, while avoiding the coverage area being too large, which would reduce the subsequent screening accuracy. This cone-shaped coverage area is the three-dimensional spatial coverage area corresponding to the sound source location.
[0048] Step 202: Based on the geometric correspondence between the cone-shaped coverage area and the imaging field of view of each video acquisition device, determine the two-dimensional projection coverage area corresponding to the viewpoint of each video acquisition device.
[0049] Optionally, the conference management device acquires the installation location coordinates, imaging field of view, lens focal length, imaging direction, etc. of each video acquisition device. The installation location coordinates refer to the specific position of each video acquisition device in the three-dimensional space of the conference scene, with a preset origin as the reference.
[0050] The imaging field of view refers to the range of spatial angles that a video capture device can capture, including the horizontal and vertical imaging field of view, reflecting the shooting coverage of the video capture device; the lens focal length refers to the focal length parameter of the lens of the video capture device, which affects the size of the imaging field of view and the shooting clarity; the imaging direction refers to the orientation of the lens of the video capture device, that is, the spatial orientation corresponding to the imaging center.
[0051] Furthermore, the conference management device determines the corresponding two-dimensional projection coverage area from the perspective of each video acquisition device based on the geometric correspondence between the cone-shaped coverage area and the imaging field of view of each video acquisition device. Here, the geometric correspondence refers to the geometric relationship formed by mapping the cone-shaped coverage area in three-dimensional space onto the imaging plane of each video acquisition device through the principle of perspective projection. The principle of perspective projection is the principle by which points, lines, and surfaces in three-dimensional space are imaged by the lens of the video acquisition device and projected onto a two-dimensional imaging plane to form corresponding two-dimensional image elements. The specific real-time process of this invention is as follows: The conference management device first calculates the projection coordinates of all boundary points of the cone-shaped coverage area on the imaging plane of each video acquisition device. Boundary points refer to points on the contour edge of the cone-shaped coverage area, including the points on the cone's apex and base contour. Subsequently, the projection coordinates of all boundary points are connected to form a closed two-dimensional geometric area, which is the two-dimensional projection coverage area from the perspective of the corresponding video acquisition device. Each video acquisition device corresponds to an independent two-dimensional projection coverage area, which matches the imaging field of view of the video acquisition device and only contains the projection portion of the cone-shaped coverage area from that perspective.
[0052] Step 203: Based on the spatial coordinates of each face region in the initial face region set under the corresponding viewpoint and the two-dimensional projection coverage area, determine the first face region subset located within the two-dimensional projection coverage area.
[0053] Optionally, for each video capture device's perspective, the conference management device maps the spatial coordinates of each face region in the initial set of face regions under that perspective onto the imaging plane of the corresponding two-dimensional projection coverage area using the principle of perspective projection, thus obtaining the two-dimensional projection coordinates of each face region on the imaging plane. Subsequently, it determines whether the two-dimensional projection coordinates of each face region are within the two-dimensional projection coverage area corresponding to that perspective. The criterion is whether the center point of the two-dimensional projection coordinates of the face region falls within the closed boundary of the two-dimensional projection coverage area. If it falls within the boundary, the face region is determined to be related to the sound source location; if it falls outside the boundary, the face region is determined to be unrelated to the sound source location.
[0054] Furthermore, the conference management device aggregates all face regions within the corresponding two-dimensional projection coverage area from the perspective of each video acquisition device, forming the first subset of face regions from that perspective.
[0055] The first face region subset contains only the face regions related to the sound source location from that viewpoint. The above operation is performed for all video acquisition device viewpoints to obtain the corresponding first face region subsets for each viewpoint, thus completing the initial screening of the initial face region set.
[0056] Step 204: Determine the target face region set based on the bounding box coordinates of each face region in the first face region subset and the boundary coordinates of the two-dimensional projection coverage area.
[0057] Optionally, the bounding box coordinates refer to the two-dimensional boundary coordinates of the face region on the imaging plane of the corresponding viewpoint video frame, used to define the range of the face region in the video frame; the boundary coordinates refer to the two-dimensional coordinates of the closed boundary of the two-dimensional projection coverage area on the imaging plane, used to define the range of the two-dimensional projection coverage area. The conference management device determines the target face region set based on the bounding box coordinates of each face region in the first face region subset and the boundary coordinates of the two-dimensional projection coverage area, as described in steps 2041 to 2044.
[0058] This invention achieves precise correlation between audio location and video face. Through step-by-step mapping from three-dimensional to two-dimensional, it improves the accuracy of face region selection, effectively eliminates interference from faces unrelated to the current sound source, and avoids the problem of misselection of irrelevant targets caused by slight drift of the sound source by defining the cone-shaped coverage area. Thus, it can accurately lock the actual speaking target, solve the problem caused by multiple participants speaking alternately or environmental reverberation interference, and improve the continuity of video conference.
[0059] Optionally, the process of steps 2041 to 2044 includes: Step 2041: Based on the bounding box coordinates of each face region in the first face region subset and the boundary coordinates of the two-dimensional projection coverage area, perform inclusion relationship analysis to determine the inclusion relationship between each face region and the two-dimensional projection coverage area.
[0060] Optionally, the conference management device performs inclusion relationship analysis on each face region in the first face region subset under each video capture device's viewpoint. Inclusion relationship analysis refers to determining the spatial overlap between the bounding box range of the face region and the boundary range of the two-dimensional projection coverage area. The specific analysis process of this embodiment is as follows: extract the coordinates of the four bounding box vertices of a single face region in the first face region subset one by one, and simultaneously extract all boundary coordinates of the two-dimensional projection coverage area under that viewpoint; compare the coordinates of the four bounding box vertices of the face region with the boundary coordinates of the two-dimensional projection coverage area to determine whether each vertex coordinate is within the boundary range of the two-dimensional projection coverage area, and simultaneously determine whether the edge of the face region bounding box overlaps with the boundary of the two-dimensional projection coverage area.
[0061] Furthermore, based on the comparison results, the inclusion relationship between each face region and the 2D projection coverage area is determined. The inclusion relationship determination results are divided into three types: complete inclusion, partial overlap, and complete non-overlap. Complete inclusion means that all four vertices of the face region's bounding box are within the boundary of the 2D projection coverage area, and the entire bounding box of the face region is inside the 2D projection coverage area without any part exceeding it. Partial overlap means that at least one vertex of the face region's bounding box is within the boundary of the 2D projection coverage area, or the edge of the face region's bounding box intersects with the boundary of the 2D projection coverage area, and the face region is not completely inside the 2D projection coverage area. Complete non-overlap means that all four vertices of the face region's bounding box are outside the boundary of the 2D projection coverage area, and the edge of the face region's bounding box does not intersect with the boundary of the 2D projection coverage area. The conference management device performs the above analysis process on each face region in the first face region subset and records the inclusion relationship determination results for each face region.
[0062] Step 2042: Based on the face regions determined to be completely contained or partially overlapping in the inclusion relationship determination results, determine the candidate face region set.
[0063] Optionally, the conference management device checks the inclusion relationship determination results of each face region in the first face region subset one by one, retaining face regions whose inclusion relationship determination results are completely inclusive or partially overlapping, and removing face regions whose inclusion relationship determination results are completely non-overlapping. Among them, the completely inclusive face regions have the highest correlation with the two-dimensional projection coverage area, and the corresponding face targets have the strongest correlation with the sound source location; the partially overlapping face regions, although not completely inside the two-dimensional projection coverage area, have some areas overlapping with the two-dimensional projection coverage area, and may still be face targets related to the sound source location, so they are retained; the completely non-overlapping face regions have no correlation with the two-dimensional projection coverage area, and the corresponding face targets are unrelated to the current sound source location, so removing them can further reduce the interference of irrelevant faces.
[0064] Furthermore, the meeting management device aggregates all face regions whose inclusion relationship determination results are complete inclusion or partial overlap, forming a candidate face region set. The face regions included in the candidate face region set are all related to the two-dimensional projection coverage area, that is, they have a potential correlation with the sound source location.
[0065] Step 2043: Based on the center point coordinates of each first candidate face region in the candidate face region set in the corresponding video frame, and the center direction vector of the two-dimensional projection coverage area, obtain the angular deviation value of each first candidate face region relative to the principal axis direction of the sound source.
[0066] Optionally, the center direction vector of the two-dimensional projection coverage area refers to the unit vector pointing from the center point of the two-dimensional projection coverage area to the position of the sound source in the three-dimensional space of the conference scene. It is used to characterize the center orientation of the two-dimensional projection coverage area, and its determination process is based on the coordinates of the center point of the two-dimensional projection coverage area and the position coordinates of the sound source in the three-dimensional space. The principal axis direction of the sound source refers to the core propagation direction of the sound source in the three-dimensional space, that is, the extension direction of the line connecting the sound source position and the preset origin of the conference scene. It is consistent with the direction of the central axis of the cone-shaped coverage area. The preset origin is the origin of the spatial coordinates of the conference scene, which is usually set at the center of the conference scene.
[0067] Optionally, the conference management device determines the center point coordinates of each first candidate face region in the candidate face region set within the corresponding video frame. A first candidate face region refers to a single face region in the candidate face region set. Its center point coordinates in the corresponding video frame are obtained by calculating the coordinates of the intersection of the diagonals of the face region's bounding box. Specifically, the calculation process involves extracting the coordinates of the four vertices of the first candidate face region's bounding box, calculating the average x-coordinate and average y-coordinate of the upper left and lower right vertices, and simultaneously calculating the average x-coordinate and average y-coordinate of the upper right and lower left vertices. The intersection of these two average value sets is the center point coordinate of the first candidate face region in the corresponding video frame. This center point coordinate is used to characterize the core position of the first candidate face region on the imaging plane.
[0068] Subsequently, the conference management device calculates the angular deviation value of each first candidate face region relative to the principal axis of the sound source, based on the center point coordinates of the first candidate face region and the center direction vector of the two-dimensional projection coverage area under the corresponding viewpoint. The angular deviation value refers to the angle between the direction from the center point of the first candidate face region to the sound source position and the principal axis of the sound source. It is used to characterize the degree of deviation of the face target corresponding to the first candidate face region from the principal axis of the sound source. The smaller the angular deviation value, the stronger the correlation between the face target and the sound source location, and the more likely it is to be the speaking target corresponding to the current sound source; the larger the angular deviation value, the weaker the correlation between the face target and the sound source location, and the more likely it is to be an interference target.
[0069] Optionally, the specific calculation process of the angle deviation value in the embodiment of the present invention is as follows: taking the center point of the two-dimensional projection coverage area as a reference, a coordinate system on the two-dimensional imaging plane is established, the position of the center point of the first candidate face region in the coordinate system is determined, and the projection direction of the center direction vector of the two-dimensional projection coverage area in the coordinate system is determined; the angle between the direction vector of the center point of the first candidate face region pointing to the center point of the two-dimensional projection coverage area and the projection direction of the center direction vector on the imaging plane is calculated, and the angle is the angle deviation value of the first candidate face region relative to the main axis direction of the sound source.
[0070] Step 2044: Based on the changing trend of the angle deviation values of each first candidate face region in a continuous preset number of video frames, determine the target face region set.
[0071] Optionally, the preset number refers to the number of video frames used to analyze the changing trend of the angle deviation value. The preset number is determined based on the frame rate of the video acquisition device and the stability requirements of the sound source localization, and is usually set to 10 to 20 frames to ensure that the changing pattern of the angle deviation value can be accurately reflected and to avoid inaccurate analysis results due to accidental deviations in a single video frame. Therefore, the conference management device determines the target face region set based on the changing trend of the angle deviation value of each first candidate face region in a continuous preset number of video frames, as described in steps 20441 to 20444.
[0072] The embodiments of the present invention effectively eliminate interference from faces unrelated to the current sound source, avoid the problem of misselection of irrelevant targets caused by slight drift of the sound source and accidental deviation of a single video frame, ensure that the selected target face regions are highly correlated with the sound source location, guarantee the accuracy of the viewing angle control of the video acquisition device, and improve the continuity of the video conference.
[0073] Optionally, the process of steps 20441 to 20444 includes: Step 20441: Based on the changing trend of the angle deviation values of each first candidate face region in a consecutive preset number of video frames, determine the dynamic fluctuation amplitude of the angle deviation, and remove the first candidate face regions whose dynamic fluctuation amplitude of the angle deviation exceeds the preset stability upper limit to obtain the second candidate face regions.
[0074] Optionally, the conference management device extracts the angle deviation values corresponding to each first candidate face region in a predetermined number of consecutive video frames. Based on the changing trends of these angle deviation values, it determines the dynamic fluctuation amplitude of the angle deviation. The dynamic fluctuation amplitude of the angle deviation refers to the difference between the maximum and minimum values of all angle deviation values for a single first candidate face region in a predetermined number of consecutive video frames. It characterizes the degree of fluctuation of the angle deviation value of the first candidate face region in consecutive frames. The smaller the fluctuation amplitude, the more stable the facial target position corresponding to the first candidate face region, and the more reliable the correlation with the sound source location. The larger the fluctuation amplitude, the more unstable the facial target position corresponding to the first candidate face region, and the more likely it is an interfering target or a facial target unrelated to the sound source. The specific calculation process of the dynamic fluctuation amplitude of the angle deviation is as follows: all angle deviation values for a single first candidate face region in a predetermined number of consecutive video frames are statistically analyzed, the maximum and minimum values are selected, and the difference between the maximum and minimum values is the dynamic fluctuation amplitude of the angle deviation for the first candidate face region.
[0075] Optionally, the preset stability upper limit refers to a threshold value preset by the conference management device to determine whether the fluctuation of the angle deviation value is stable. This threshold value is determined according to the stability requirements of the conference scene and the sound source drift range, and is usually set to 2 to 5 degrees. First candidate face regions whose dynamic fluctuation amplitude of angle deviation does not exceed the threshold value are judged as candidates with stable positions, while those exceeding the threshold value are judged as candidates with unstable positions. The conference management device compares the dynamic fluctuation amplitude of the angle deviation of each first candidate face region with the preset stability upper limit, eliminates first candidate face regions whose dynamic fluctuation amplitude of angle deviation exceeds the preset stability upper limit, and retains first candidate face regions whose dynamic fluctuation amplitude of angle deviation does not exceed the preset stability upper limit. The retained first candidate face regions are defined as second candidate face regions.
[0076] Step 20442: Based on the angular deviation values of each second candidate face region under at least two different viewpoints, spatial pointing is reversed to obtain the three-dimensional spatial pointing, and based on whether the three-dimensional spatial pointing intersects within the cone-shaped coverage area of the sound source, the subset of the second face region is determined.
[0077] Optionally, the conference management device determines at least two different viewing angles for each second candidate face region. These at least two different viewing angles refer to the acquisition angles of two or more video acquisition devices at different positions that can simultaneously capture the second candidate face region, ensuring the accuracy and reliability of spatial pointing inversion and avoiding inversion deviations caused by a single viewing angle. Subsequently, the angular deviation value of the second candidate face region under each corresponding viewing angle is extracted. Combined with the installation position coordinates of the corresponding viewing angle video acquisition device, imaging field parameters, and the center direction vector of the two-dimensional projection coverage area, spatial pointing inversion is performed to obtain the three-dimensional spatial pointing of the second candidate face region. Spatial pointing inversion refers to calculating the pointing direction of the face target corresponding to the second candidate face region in the three-dimensional space of the conference scene based on the angular deviation information on the two-dimensional imaging plane, combined with the spatial position and imaging parameters of the video acquisition device. This three-dimensional spatial pointing is the spatial direction vector pointing from the installation position of the corresponding video acquisition device to the face target, which can characterize the approximate position and direction of the face target in three-dimensional space. The meeting management device performs a reverse calculation of the three-dimensional spatial orientation for each second candidate face region based on its angular deviation values at least two different viewpoints, thereby obtaining the three-dimensional spatial orientation corresponding to each second candidate face region.
[0078] Subsequently, the conference management device determines whether the three-dimensional spatial pointers corresponding to each second candidate face region intersect within the cone-shaped coverage area of the sound source. The determination process is as follows: The intersection point of at least two three-dimensional spatial pointers corresponding to each second candidate face region is calculated. The coordinate position of this intersection point in the three-dimensional space of the conference scene is determined. This coordinate position is compared with the spatial range of the cone-shaped coverage area of the sound source to determine whether the intersection point is within the boundary range of the cone-shaped coverage area. If the intersection point of the three-dimensional spatial pointers is within the cone-shaped coverage area, the second candidate face region is determined to be highly correlated with the sound source orientation; if the intersection point is not within the cone-shaped coverage area, the second candidate face region is determined to be unrelated to the sound source orientation.
[0079] The meeting management device will aggregate the second candidate face regions identified as conical coverage areas where the three-dimensional spatial directions converge at the sound source, forming a subset of the second face regions.
[0080] Step 20443: Based on the angle between the imaging optical axis of the video acquisition device and the main axis of the sound source for each face region in the second face region subset, determine whether each face region is within the main lobe range of the effective radiation of the sound source energy, and obtain the third face region subset.
[0081] Optionally, the imaging optical axis direction of the video acquisition device refers to the direction of the central axis of the lens of the video acquisition device, that is, the spatial pointing direction corresponding to the imaging center, which is used to characterize the shooting orientation of the video acquisition device; the main axis direction of the sound source refers to the core propagation direction of the sound source in three-dimensional space, that is, the extension direction of the line connecting the sound source position and the preset origin of the conference scene, which is consistent with the direction of the central axis of the cone-shaped coverage area; the main lobe range of the effective radiation of the sound source energy refers to the spatial angle range in which the energy of the sound emitted by the sound source is mainly concentrated. This range is formed by extending to the sides and up and down by a preset angle with the main axis direction of the sound source as the center. The preset angle is determined according to the sound pickup characteristics of the sound source and the needs of the conference scene, and is usually set to 8 degrees to 12 degrees. The sound energy is strongest within the main lobe range, and the corresponding face target is most likely to be the speaking target. The sound energy outside the main lobe range is weaker, and the corresponding face target is likely to be the interference target.
[0082] Optionally, the conference management device calculates the angle between the imaging optical axis of the video acquisition device and the main axis of the sound source for each face region in the second face region subset. The specific calculation process for this angle is as follows: taking the preset origin of the conference scene as a reference, determine the spatial vector corresponding to the direction of the imaging optical axis of the video acquisition device and the spatial vector corresponding to the direction of the main axis of the sound source, and calculate the angle between the two spatial vectors. This angle is the angle between the imaging optical axis of the video acquisition device and the main axis of the sound source for the face region, and is used to characterize the degree to which the face target corresponding to the face region is within the effective radiation range of the sound source energy. Subsequently, the conference management device determines whether the included angle is within the main lobe range of the effective radiation of the sound source energy. The judgment criterion is: whether the calculated included angle is less than or equal to half the angle of the main lobe range of the effective radiation of the sound source energy (i.e., the angle at which the main lobe range expands to both sides with the main axis of the sound source as the center). If the included angle is less than or equal to half the angle, it is determined that the face area is within the main lobe range of the effective radiation of the sound source energy; if the included angle is greater than half the angle, it is determined that the face area is not within the main lobe range of the effective radiation of the sound source energy.
[0083] The conference management device aggregates the face regions within the main lobe range of the effective radiation of the sound source energy from the second face region subset to form the third face region subset.
[0084] Step 20444: Determine the target face region set based on whether the angle deviation values of each face region in the third face region subset remain consistent in adjacent video frames.
[0085] Optionally, the sign of the angle deviation value refers to the positive or negative attribute of the angle deviation value of each third face region in the corresponding video frame. If the signs are consistent, it means that the positive or negative attribute of the angle deviation value of the same third face region remains unchanged in adjacent video frames. This is used to characterize that the offset direction of the face target corresponding to the third face region relative to the main axis of the sound source remains stable, further verifying its correlation with the sound source location. If the signs are inconsistent, it means that the offset direction of the face target is unstable, and it is likely to be an interfering target.
[0086] Optionally, the conference management device tracks the sign of the angle deviation value of each face region in the subset of third face regions in consecutive adjacent video frames, and determines whether the face region maintains a consistent sign of the angle deviation value in adjacent video frames. The determination process in this embodiment is as follows: extract the angle deviation value of the same third face region in multiple consecutive sets of adjacent video frames, check the positive or negative attribute of the angle deviation value in each set of adjacent video frames. If the sign of the angle deviation value of the face region does not change in all adjacent video frames, it is determined that the sign remains consistent; if the sign of the angle deviation value changes in any set of adjacent video frames, it is determined that the sign does not remain consistent. The number of consecutive sets of adjacent video frames is consistent with the preset number in the aforementioned sub-step.
[0087] The conference management device summarizes the face regions that maintain a consistent angular deviation value in adjacent video frames from the third face region subset, and finally determines them as the target face region set.
[0088] The embodiments of the present invention effectively eliminate various interfering facial targets with unstable positions, inconsistent spatial orientations, weak energy radiation, and variable offset directions. It avoids the problem of misselection of irrelevant targets caused by slight sound source drift and environmental interference, ensuring that the final determined target facial areas are highly correlated with the sound source location. This guarantees the accuracy of the subsequent video acquisition device's perspective control, solves the problem caused by multiple participants speaking alternately or environmental reverberation interference, and improves the continuity of video conference footage.
[0089] Optionally, the processes of steps 401 to 404 include: Step 401: Based on the first time point in the lip opening and closing time sequence corresponding to each speaking target in the lip movement state information where the lip opening area increases between adjacent frames, and the second time point in the sound source energy intensity time sequence in the sound source orientation information where the sound source energy intensity increases between adjacent sampling times, a joint time alignment sequence is obtained.
[0090] Optionally, the lip movement state information includes the lip opening and closing time sequence corresponding to each speaking target. The speaking target refers to the participant corresponding to each target face region in the target face region set. The lip opening and closing time sequence refers to the lip opening and closing state data set of each speaking target in consecutive video frames arranged in chronological order of video acquisition time. Each data point in this data set corresponds to the lip opening area of the speaking target in one video frame. The lip opening area refers to the actual area of the opening and closing part in the lip region, which is used to characterize the degree of lip opening and closing.
[0091] Optionally, the sound source location information includes a time sequence of sound source energy intensity. The time sequence of sound source energy intensity refers to a set of sound source energy intensity data at consecutive sampling times arranged in chronological order of audio acquisition time. Sound source energy intensity refers to the energy level of the sound emitted by the sound source as acquired by the audio acquisition device, which is used to characterize the activity level of the sound source. The sampling time refers to a fixed time point at which the audio acquisition device acquires sound signals. The interval between sampling times is determined according to the sampling frequency of the audio acquisition device. The sampling frequency is usually set to 16,000 to 48,000 times per second, that is, the interval between adjacent sampling times is 1 / 16,000 to 1 / 48,000 seconds.
[0092] Optionally, the conference management device analyzes the corresponding lip opening and closing timing sequence for each speaking target and extracts the first time point at which the lip opening area increases between adjacent frames.
[0093] In this context, adjacent frames refer to two consecutive video frames arranged in chronological order of video capture time within the lip opening and closing sequence. An increase in lip opening area between adjacent frames means that the lip opening area of the speaking target in the later video frame is greater than that in the earlier video frame, and the area difference is greater than a preset area threshold. The preset area threshold is a critical value set by the conference management device to determine whether the lip opening area has effectively increased. It is usually set to 5% to 10% of the maximum lip opening area to avoid misjudgment due to slight vibrations. The first time point refers to the capture time of the later video frame, which is the time node when the lip opening area begins to increase.
[0094] Optionally, the conference management device analyzes the time series of sound source energy intensity and extracts the second time point in which the sound source energy intensity increases between adjacent sampling times.
[0095] Among them, adjacent sampling time refers to two consecutive sampling times arranged in the order of audio acquisition time in the time sequence of sound source energy intensity. An increase in sound source energy intensity between adjacent sampling times means that the sound source energy intensity of the later sampling time is greater than that of the earlier sampling time, and the intensity difference is greater than a preset intensity threshold. The preset intensity threshold is a critical value preset by the conference management device to determine whether the sound source energy intensity has effectively increased. It is usually set to 3% to 5% of the maximum sound source energy intensity to avoid misjudgment caused by fluctuations in environmental noise. The second time point refers to the time point of the later sampling time, which is the time node when the sound source energy intensity begins to increase.
[0096] Furthermore, the conference management device integrates all first time points corresponding to each speaking target with all second time points in the sound source energy intensity time sequence in chronological order to form a joint time sequence alignment sequence. The joint time sequence alignment sequence refers to a unified time sequence containing all first and second time points, where each time point is labeled with its corresponding type (first time point or second time point) and corresponding speaking target (only the first time point is labeled).
[0097] Step 402: Based on the time interval between the first time point and the second time point in the joint time alignment sequence, select the paired time points whose time interval does not exceed the preset synchronization time tolerance, and obtain the synchronization event set corresponding to each speaking target.
[0098] Optionally, the preset synchronization time tolerance refers to the maximum time interval threshold used to determine whether the first time point and the second time point have a synchronous correlation. This threshold is determined based on the matching relationship between the video capture frame rate and the audio sampling frequency, and is usually set to 10 milliseconds to 30 milliseconds to ensure that it can cover the small time deviation caused by the asynchrony between audio capture and video capture, while avoiding mismatch caused by an excessively large tolerance.
[0099] For each first time point in the joint timing alignment sequence, the conference management device searches for all second time points within a preset synchronization time tolerance range before and after that first time point, and calculates the time interval between each first time point and each corresponding second time point. The time interval refers to the time difference between two time points. The calculation process is as follows: subtract the earlier time point from the later time point, and the difference is the time interval between the two time points.
[0100] Subsequently, the conference management device compares the calculated time interval of each time point with the preset synchronization time tolerance, and filters out paired time points whose time interval does not exceed the preset synchronization time tolerance. A paired time point refers to a pair of time points consisting of a first time point and a second time point, and the time interval between the two time points does not exceed the preset synchronization time tolerance.
[0101] If multiple second time points corresponding to a first time point all meet the time point interval requirement, then the second time point with the smallest time point interval is selected and paired with the first time point to avoid subsequent statistical bias caused by multiple pairs corresponding to a first time point.
[0102] If a second time point corresponds to multiple first time points and all meet the time point interval requirement, then each second time point is paired with a first time point to form a pairing time point, corresponding to different speaking targets.
[0103] The conference management device summarizes all the paired time points corresponding to each speaking target to form a set of synchronous events corresponding to each speaking target. The set of synchronous events refers to the set of all paired time points with synchronous correlation corresponding to a single speaking target, and each paired time point is a synchronous event.
[0104] Step 403: Based on the set of synchronous events of each speaking target, count the number of synchronous events contained in each speaking target within a consecutive preset time window to obtain the synchronous event count of each speaking target.
[0105] Optionally, a time window refers to a fixed time interval used to count the number of synchronous events. The duration of the time window is determined based on the average speaking speed in the meeting scenario, and is usually set to 1 to 3 seconds to ensure that it can fully cover the synchronous events in a short speech. A set of consecutive time windows refers to a set number of consecutive time windows arranged in chronological order. The set number is usually set to 3 to 5. By counting multiple consecutive time windows, misjudgments caused by statistical bias of a single time window can be avoided.
[0106] The conference management device, for each speaking target, counts the number of paired time points corresponding to each synchronous event in the target's synchronous event set that fall within each consecutive preset time window. The counting process is as follows: first, the time range (start time point and end time point) of each consecutive time window is determined; then, each paired time point in the target's synchronous event set is checked one by one to determine whether the time of the paired time point is within the time range of the current time window. If it is within the range, the synchronous event is counted in the number of synchronous events in the current time window; otherwise, it is not counted.
[0107] Furthermore, the conference management device summarizes the number of synchronization events for each speaking target within a consecutive preset time window to obtain the synchronization event count for that speaking target. The synchronization event count refers to the total number of all synchronization events for a single speaking target within a consecutive preset time window. It is used to characterize the frequency of synchronization between the changes in lip movement and the changes in sound source energy of that speaking target. The higher the synchronization event count, the stronger the synchronization between the lip movement and the changes in sound source energy of that speaking target, and the more likely it is a genuine speaking target; the lower the synchronization event count, the weaker the synchronization, and the more likely it is an interfering target.
[0108] During the statistical process, if there are no synchronization events for the speaking target within a certain time window, the number of synchronization events in that time window is counted as 0. This ensures that the synchronization event count can fully reflect the synchronization status of the speaking target in all consecutive time windows and avoids counting deviations caused by missing a certain time window.
[0109] Step 404: Perform synchronization analysis based on the synchronization event count of each speaking target to obtain the effective speaker targets.
[0110] Optionally, the conference management device performs synchronization analysis based on the synchronization event count of each speaking target to obtain the valid speaker targets, as described in steps 4041 to 4044.
[0111] This invention enables time synchronization analysis of lip movement state information and sound source location information, effectively distinguishing between the real speaker and interfering sound sources. It avoids the problem of misjudging the interference target caused by sound source jumps, drifts, and environmental noise fluctuations, ensuring that the locked valid speaker target is highly matched with the current sound source, and guaranteeing the continuity of the video conference.
[0112] Optionally, the processes of steps 4041 to 4044 include: Step 4041: Based on the relationship between the synchronization event count of each speaking target and the preset minimum synchronization event number threshold, speaking targets with synchronization event counts less than the preset minimum synchronization event number threshold are eliminated, thus forming the first candidate speaking target set.
[0113] Optionally, the preset minimum synchronous event threshold is a critical value used to determine whether a speaking target has preliminary synchronous correlation. This critical value is determined based on the total duration of several consecutive preset time windows, the average speaking speed, and the frequency of synchronous events. It is usually set to 5 to 10. Speaking targets with a synchronous event count not lower than this threshold have preliminary synchronous correlation, while speaking targets with a count lower than this threshold have extremely weak synchronous correlation and can be identified as interference targets. For each speaking target, the conference management device compares the synchronous event count of that speaking target with the preset minimum synchronous event threshold to determine the relationship between the two. The comparison process is as follows: directly compare the specific value of the synchronous event count with the specific value of the preset minimum synchronous event threshold. If the value of the synchronous event count is greater than or equal to the preset minimum synchronous event threshold, the speaking target is determined to have preliminary synchronous correlation and is retained; if the value of the synchronous event count is less than the preset minimum synchronous event threshold, the speaking target is determined to have extremely weak synchronous correlation and does not meet the conditions to be a valid speaking target, and is eliminated.
[0114] Furthermore, the conference management device aggregates all speaking targets whose synchronization event counts are greater than or equal to the preset minimum synchronization event count threshold and have not been eliminated, forming a first candidate speaking target set.
[0115] Step 4042: Determine the second set of candidate speaking targets based on whether each candidate speaking target in the first set of candidate speaking targets is earlier than or equal to the second time point in each synchronization event at the first time point.
[0116] Optionally, the conference management device analyzes each paired time point in the synchronization event set of each candidate speaking target in the first candidate speaking target set, determining the time sequence of the first and second time points in each paired time point. The core judgment is whether the first time point of the candidate speaking target in each synchronization event is earlier than or equal to the second time point. Specifically, if the first time point is earlier than the second time point, it indicates that the speaking target's lips begin to open first (opening area increases), followed by an increase in sound source energy intensity, consistent with the physiological laws and sound propagation logic of real speech. If the first time point is equal to the second time point, it means that the specific time values of the first and second time points are completely consistent, indicating that the increase in the opening area of the speaking target's lips and the increase in sound source energy intensity occur synchronously. Due to slight synchronization deviations between audio and video acquisition, this situation also conforms to the synchronization logic of real speech. If the first time point is later than the second time point, it indicates that the sound source energy intensity increases first, followed by the speaking target's lips opening, which does not conform to the physiological laws and sound propagation logic of real speech. This type of synchronization event can be identified as an abnormal synchronization event, and the corresponding candidate speaking target is likely an interference target.
[0117] Optionally, the specific judgment process in this embodiment of the invention is as follows: The conference management device extracts each paired time point from the synchronous event set of candidate speaking targets one by one, obtains the specific time values of the first time point and the second time point in the paired time point, compares the size of the two time values, and records the judgment result of each paired time point (the first time point is earlier than, equal to, or later than the second time point). Subsequently, it judges whether all paired time points in the synchronous event set of the candidate speaking target satisfy the condition that the first time point is earlier than or equal to the second time point. If all paired time points satisfy this condition, it is determined that the time sequence of the synchronous events of the candidate speaking target conforms to the actual speaking logic and is retained; if any paired time point satisfies the condition that the first time point is later than the second time point, it is determined that the synchronous event of the candidate speaking target is abnormal and the synchronization correlation is unreliable, and is removed. The conference management device summarizes all candidate speaking targets in the synchronous event set whose paired time points satisfy the condition that the first time point is earlier than or equal to the second time point and have not been removed, forming a second candidate speaking target set.
[0118] Step 4043: Determine the third candidate speaking target set based on whether the variation range of the time interval of the synchronization events of each candidate speaking target in the second candidate speaking target set within a consecutive preset number of time windows does not exceed the preset time jitter tolerance.
[0119] Optionally, the preset time jitter tolerance refers to the critical value used to determine whether the change in the time interval of the synchronization event is stable. This critical value is determined based on the synchronization accuracy of audio and video acquisition and the range of environmental noise fluctuation. It is usually set to 5 milliseconds to 15 milliseconds. Synchronization events whose time interval change does not exceed this critical value are judged to have stable time jitter, while those that exceed this critical value are judged to have abnormal time jitter.
[0120] The conference management device analyzes the synchronization events in the synchronization event set of each candidate speaking target in the second candidate speaking target set, and measures the variation range of the time interval within a consecutive preset number of time windows. The time interval of a synchronization event refers to the time difference between the first and second time points within a single synchronization event. The variation range of the synchronization event time interval refers to the difference between the maximum and minimum time intervals of all synchronization events of the candidate speaking target within a consecutive preset number of time windows. This variation range characterizes the stability of the synchronization event time intervals; the smaller the variation range, the more stable the synchronization event time intervals, and the more reliable the synchronization correlation of the candidate speaking target; the larger the variation range, the more severe the jitter in the synchronization event time intervals, and the more likely it is an interfering target.
[0121] Optionally, the specific analysis process of this embodiment of the invention is as follows: The conference management device first extracts all synchronous events in the synchronous event set of the candidate speaking target, and calculates the time interval of each synchronous event one by one (subtracting the first time point from the second time point to obtain the time difference); then, it filters out the maximum and minimum values among these time intervals, calculates the difference between the maximum and minimum values, and this difference is the variation range of the time interval of the synchronous events of the candidate speaking target within a consecutive preset number of time windows; finally, it compares the calculated time interval variation range with the preset time jitter tolerance to determine whether the variation range does not exceed the preset time jitter tolerance.
[0122] If the variation range of the synchronization event time interval of a candidate speaking target does not exceed the preset time jitter tolerance, the candidate speaking target is determined to have a stable synchronization event time interval and reliable synchronization correlation, and is retained; if the variation range exceeds the preset time jitter tolerance, the candidate speaking target is determined to have severe synchronization event time interval jitter and unreliable synchronization correlation, and is eliminated. The conference management device summarizes all candidate speaking targets whose synchronization event time interval variation range does not exceed the preset time jitter tolerance and have not been eliminated, forming a third set of candidate speaking targets.
[0123] Step 4044: Based on whether each candidate speaker in the third candidate speaker target set has a lip opening area increase event corresponding to its spatial position in the video frame sequence acquired by different video acquisition devices, and the time interval between the lip opening area increase event and the sound source energy intensity increase event does not exceed the preset synchronization time tolerance, determine the effective speaker target.
[0124] Optionally, the video frame sequence acquired by each video acquisition device refers to a continuous set of image frames acquired by multiple video acquisition devices at a preset frame rate, with each video acquisition device corresponding to an independent acquisition perspective. The conference management device performs multi-view verification on each candidate speaking target in the third candidate speaking target set, focusing on whether the candidate speaking target exhibits an increase in lip opening area corresponding to its spatial location in all video frame sequences acquired by different video acquisition devices, and whether the time interval between the increase in lip opening area event and the increase in sound source energy intensity event does not exceed a preset synchronization time tolerance.
[0125] Optionally, the specific judgment process in this embodiment of the invention is as follows: For a single candidate speaking target, its spatial coordinates are extracted, and the video frame sequences captured by each video acquisition device are matched one by one. In the video frame sequence of each video acquisition device, based on the spatial coordinates, the corresponding face region is searched, and it is determined whether the face region is the face region of the candidate speaking target (confirmed by face feature matching). If it is the face region of the candidate speaking target, it is further determined whether there is a lip opening area increase event in the face region (i.e., the lip opening area increases between adjacent frames). If a lip opening area increase event corresponding to the spatial position of the candidate speaking target is found in the video frame sequences of all different video acquisition devices, the first time point corresponding to the lip opening area increase event in each video frame sequence is extracted, and the second time point corresponding to the sound source energy intensity increase event is extracted. The time interval between each first time point and the second time point is calculated. If all calculated time intervals do not exceed the preset synchronization time tolerance, the candidate speaking target is determined to have passed multi-view verification and the synchronization correlation is real and reliable, and is retained. If in any video frame sequence of a video acquisition device, no lip opening area increase event corresponding to its spatial position is found, or the time interval between the found lip opening area increase event and the sound source energy intensity increase event exceeds the preset synchronization time tolerance, the candidate speaking target is determined to have failed multi-view verification and is likely an interference target, and is eliminated.
[0126] Furthermore, the meeting management device aggregates all candidate speakers that have passed multi-view verification and have not been eliminated, and identifies them as valid speaker targets, thus achieving accurate differentiation between real speakers and interference sources.
[0127] The embodiments of the present invention effectively distinguish between real speakers and interfering sound sources, avoiding the problems of misselection of interference targets due to sound source jumps, drifts, environmental noise fluctuations, and misjudgments from a single perspective. This ensures that the locked valid speaker target is highly matched with the current sound source and that the synchronization correlation is real and reliable, thus guaranteeing the continuity of the video conference.
[0128] Optionally, the process of steps 50 to 80 includes: Step 50: Determine the spatial offset vector based on the spatial position coordinates of the effective speaker target in the current video frame and the current imaging center coordinates of at least one video acquisition device.
[0129] Optionally, the current video frame refers to the latest video frame containing the valid speaker target when determining the spatial position; the current imaging center coordinates of the video acquisition device refer to the specific position coordinates of the current lens imaging center point of the video acquisition device in the three-dimensional space of the meeting scene. Similarly, based on the preset origin of the meeting scene, the imaging center point refers to the spatial position corresponding to the optical center of the lens of the video acquisition device, and its coordinates are jointly determined by the installation position coordinates of the video acquisition device and the current shooting angle.
[0130] Optionally, the conference management device determines a spatial offset vector based on the spatial position coordinates of the valid speaker target in the current video frame and the current imaging center coordinates of at least one video acquisition device. The spatial offset vector is a three-dimensional vector characterizing the spatial position of the valid speaker target and the degree and direction of offset between it and the current imaging center position of the video acquisition device. This vector not only reflects the offset distance between the two but also specifies the horizontal and vertical directions of the offset.
[0131] Optionally, the specific process for determining the spatial offset vector in this embodiment of the invention is as follows: Extract the horizontal, vertical, and depth components of the effective speaker's target spatial position coordinates, as well as the horizontal, vertical, and depth components of the current imaging center coordinates of the video acquisition device; subtract the horizontal component of the current imaging center coordinates of the video acquisition device from the horizontal component of the effective speaker's target spatial position coordinates to obtain the horizontal component of the spatial offset vector; subtract the vertical component of the current imaging center coordinates of the video acquisition device from the vertical component of the effective speaker's target spatial position coordinates to obtain the vertical component of the spatial offset vector; subtract the depth component of the current imaging center coordinates of the video acquisition device from the depth component of the effective speaker's target spatial position coordinates to obtain the depth component of the spatial offset vector; integrate the horizontal, vertical, and depth components to form a complete spatial offset vector, thus completing the determination of the spatial offset vector.
[0132] Step 60: Based on the horizontal and vertical components of the spatial offset vector, determine whether both are less than the corresponding preset center alignment tolerance to determine the spatial position judgment result. The spatial position judgment result indicates whether the valid speaker target is within the stable area corresponding to the imaging center of the video acquisition device.
[0133] Optionally, the horizontal preset center alignment tolerance refers to the critical value used to determine whether the horizontal offset is within the allowable range, and the vertical preset center alignment tolerance refers to the critical value used to determine whether the vertical offset is within the allowable range. The conference management device compares the horizontal and vertical components of the spatial offset vector with the corresponding preset center alignment tolerances, determines whether both are less than the corresponding preset center alignment tolerances, and thus determines the spatial position judgment result.
[0134] The judgment process in this embodiment of the invention is divided into: comparing the absolute value of the horizontal component of the spatial offset vector with the horizontal preset center alignment tolerance, and determining whether the absolute value of the horizontal component is less than the horizontal preset center alignment tolerance; comparing the absolute value of the vertical component of the spatial offset vector with the vertical preset center alignment tolerance, and determining whether the absolute value of the vertical component is less than the vertical preset center alignment tolerance.
[0135] The spatial location determination result is a binary determination result, which includes only two cases, and can clearly indicate whether the effective speaker target is within the stable area corresponding to the imaging center of the video acquisition device.
[0136] The stable region refers to a two-dimensional area centered on the current imaging center of the video capture device, bounded by horizontal and vertical preset center alignment tolerances. This region is located at the center of the video capture device's image and is the core area ensuring a stable and clear image of the effective speaker. If the absolute value of the horizontal component of the spatial offset vector is less than the horizontal preset center alignment tolerance, and the absolute value of the vertical component is less than the vertical preset center alignment tolerance, the spatial position determination result indicates that the effective speaker target is within the stable region. If the absolute value of the horizontal component is greater than or equal to the horizontal preset center alignment tolerance, or the absolute value of the vertical component is greater than or equal to the vertical preset center alignment tolerance, or both conditions are met, the spatial position determination result indicates that the effective speaker target is not within the stable region. Step 70: If the spatial position judgment result indicates that it is not in the stable area, then control the imaging center to move towards the spatial position of the effective speaker target, and analyze the displacement vector between adjacent frames based on the spatial position coordinate sequence of the effective speaker target in a series of preset video frames to obtain the direction of motion trend.
[0137] Optionally, if the spatial location determination result indicates that the valid speaker target is within a stable area, there is no need to adjust the imaging center of the video acquisition device, and the current shooting angle and imaging state of the video acquisition device are maintained; if the spatial location determination result indicates that the valid speaker target is not within a stable area, the imaging center adjustment process is initiated, and the movement trend direction of the valid speaker target is analyzed.
[0138] Optionally, the imaging center adjustment process in this embodiment of the invention is as follows: Based on the spatial offset vector determined in step 50, an imaging center movement control command is sent to the corresponding video acquisition device. The control command includes movement direction and movement distance parameters, wherein the movement direction is determined by the positive and negative attributes of the horizontal and vertical components of the spatial offset vector, and the movement distance is determined by the absolute values of the horizontal and vertical components of the spatial offset vector. Specifically, if the horizontal component of the spatial offset vector is positive, the imaging center is controlled to move in the positive horizontal direction, and the movement distance is equal to the absolute value of the horizontal component; if the horizontal component is negative, the imaging center is controlled to move in the negative horizontal direction, and the movement distance is equal to the absolute value of the horizontal component; if the vertical component is positive, the imaging center is controlled to move in the positive vertical direction, and the movement distance is equal to the absolute value of the vertical component; if the vertical component is negative, the imaging center is controlled to move in the negative vertical direction, and the movement distance is equal to the absolute value of the vertical component. Through this movement control, the imaging center gradually moves closer to the spatial position of the effective speaker target.
[0139] While controlling the movement of the imaging center, the conference management device acquires the spatial position coordinate sequence of the effective speaker target in a series of preset video frames.
[0140] Among them, the consecutive preset video frames refer to a preset number of consecutive image frames arranged in chronological order of video acquisition time. The preset number is determined based on the video acquisition frame rate and the movement speed of the effective speaker target, and is usually set to 10 to 20 frames to ensure that the movement trajectory of the effective speaker target can be accurately reflected. The spatial position coordinate sequence refers to the sequence formed by arranging the spatial position coordinates of the effective speaker target in the consecutive preset video frames in chronological order of acquisition time.
[0141] Furthermore, based on this spatial coordinate sequence, the conference management device analyzes the displacement vectors between adjacent frames to obtain the motion trend direction. The displacement vector between adjacent frames refers to the vector obtained by subtracting the spatial coordinates of the effective speaker target from the spatial coordinates of the effective speaker target in the previous video frame from the spatial coordinates of the effective speaker target in the next video frame within a consecutive preset number of video frames; each adjacent frame pair corresponds to one displacement vector. The motion trend direction refers to the overall motion direction of the effective speaker target within a consecutive preset number of video frames, used to predict the subsequent motion trajectory of the effective speaker target.
[0142] Optionally, the specific analysis process of the motion trend direction in this embodiment of the invention is as follows: The conference management device calculates the displacement vectors between all adjacent frames in a consecutive preset number of video frames; it statistically analyzes the positive and negative attributes and numerical values of the horizontal and vertical components of these displacement vectors; if the horizontal components of most displacement vectors have the same positive and negative attributes and the numerical values tend to be stable, it is determined that the effective speaker target has a corresponding motion trend in the horizontal direction; if the vertical components of most displacement vectors have the same positive and negative attributes and the numerical values tend to be stable, it is determined that the effective speaker target has a corresponding motion trend in the vertical direction; the horizontal and vertical motion trends are integrated to form the overall motion trend direction of the effective speaker target, thus completing the analysis of the motion trend direction.
[0143] Step 80: Control the video acquisition device based on the direction of motion trend.
[0144] Optionally, the conference management device controls the video acquisition device according to the direction of the movement trend, as in steps 801 to 804.
[0145] This invention combines real-time offset calibration with motion trend prediction to ensure that the effective speaker target is always within the stable imaging center area of the video acquisition device. This avoids problems such as frequent camera perspective switching and focus deviation caused by abnormal sound source localization, movement of the effective speaker target, etc. It solves the problem of unstable picture caused by multiple participants speaking alternately or environmental reverberation interference, improves the picture continuity and clarity of video conferences, and ensures the smooth conduct of video conferences.
[0146] Optionally, the processes of steps 801 to 804 include: Step 801: Based on the relative relationship between the direction of motion trend and the field of view boundary of the video acquisition device, predict whether the effective speaker target will move out of the current field of view within a preset number of future frames, and obtain a field of view loss warning signal.
[0147] Optionally, the field-of-view boundary parameter refers to the coordinate range of the boundary of the current imaging field of view of the video acquisition device in the three-dimensional space of the conference scene, including the spatial coordinates corresponding to the horizontal left boundary, horizontal right boundary, vertical upper boundary, and vertical lower boundary of the field of view, used to define the spatial range that the video acquisition device can currently capture. The preset future frame count refers to the number of future video frames preset by the conference management device to predict whether a valid speaker target has moved out of the current field of view. The preset number is determined based on the video acquisition frame rate and the movement speed of the valid speaker target, and is usually set to 15 to 30 frames to ensure the accuracy of the prediction results and avoid false or missed warnings due to too few predicted frames.
[0148] Optionally, the conference management device predicts whether the target speaker will move out of the current field of view within a preset number of future frames based on the relative relationship between the direction of motion and the field of view boundary of the video acquisition device, thereby obtaining a field of view loss warning signal. Here, the relative relationship refers to the spatial relationship between the direction of motion of the target speaker and the spatial position of the field of view boundary of the video acquisition device; that is, whether the direction of motion is towards the outside of the field of view boundary, and whether the target speaker will exceed the range of the field of view boundary if it moves according to this motion trend.
[0149] The specific prediction process in this embodiment of the invention is as follows: Based on the current spatial coordinates and motion trend direction of the effective speaker target, and the time length corresponding to a preset number of future frames (calculated from the video capture frame rate, i.e., the preset number of future frames divided by the video capture frame rate), the predicted spatial coordinates of the effective speaker target at the end of the preset number of future frames are calculated. The calculation method is as follows: using the current spatial coordinates as a reference, the motion direction is determined according to the motion trend direction, and combined with the average motion speed of the effective speaker target in a consecutive preset number of video frames (obtained by dividing the average length of the displacement vector between adjacent frames by the frame interval time), the predicted displacement distance within the preset number of future frames is calculated, thereby obtaining the predicted spatial coordinates. The calculated predicted spatial coordinates are compared with the field of view boundary parameters of the currently controlled video capture device to determine whether the predicted spatial coordinates exceed the spatial coordinate range corresponding to the field of view boundary. If the predicted spatial coordinates exceed the field of view boundary, it is determined that the valid speaker target will move out of the current field of view within a preset number of future frames, generating a field-of-view loss warning signal indicating that the target has moved out of the current field of view. If the predicted spatial coordinates do not exceed the field of view boundary, it is determined that the valid speaker target will not move out of the current field of view within a preset number of future frames, generating a field-of-view loss warning signal indicating that the target has not moved out of the current field of view. The field-of-view loss warning signal is used to indicate whether the valid speaker target will subsequently leave the shooting range of the current video acquisition device.
[0150] Step 802: If the field of view loss warning signal indicates that the target has moved out of the current field of view, then based on the field of view coverage of each video acquisition device, candidate video acquisition devices that contain the future location of the effective speaker target are selected.
[0151] Optionally, if the field-of-view loss warning signal indicates that the valid speaker target will not move out of the current field of view within a preset number of future frames, then there is no need to screen candidate video acquisition devices, and the current video acquisition device will maintain its tracking and shooting status of the valid speaker target; if the field-of-view loss warning signal indicates that the valid speaker target will move out of the current field of view within a preset number of future frames, then the candidate video acquisition device screening process will be initiated.
[0152] The conference management device checks each video capture device's field of view individually to determine whether the future location coordinates of the valid speaker are included. The criterion is whether the future location coordinates of the valid speaker are located within the horizontal left boundary and horizontal line of the video capture device's field of view. If a video capture device falls within the range between the right boundary and the vertical upper and lower boundaries, its field of view is deemed sufficient to cover the future location of the valid speaker target, making it a candidate video capture device. If it falls outside this range, its field of view is deemed insufficient to cover the future location of the valid speaker target, and it is excluded. The conference management device aggregates all video capture devices whose field of view covers the future location of the valid speaker target, forming a candidate video capture device set. This candidate video capture device set refers to the set of video capture devices capable of covering the future location of the valid speaker target and suitable for taking over tracking and filming from the current video capture device. This set must contain at least one video capture device to ensure that the valid speaker target does not leave the shooting range of any video capture device.
[0153] Step 803: Based on the spatial distance between each candidate video acquisition device and the effective speaker target, select the candidate video acquisition device with the smallest spatial distance as the main imaging device, and align its imaging center with the effective speaker target.
[0154] Optionally, based on the future location coordinates of the effective speaker target, i.e., the predicted spatial location coordinates at the end of a preset future frame number, the spatial distance between each candidate video acquisition device and the effective speaker target is determined. Here, spatial distance refers to the straight-line distance between the installation location coordinates of the candidate video acquisition device and the future location coordinates of the effective speaker target, used to characterize the proximity between the candidate video acquisition device and the future location of the effective speaker target. The smaller the spatial distance, the clearer the image of the effective speaker target captured by the candidate video acquisition device, the more accurate the tracking, and the more suitable it is as the main imaging device.
[0155] Optionally, the specific calculation process of spatial distance in this embodiment of the invention is as follows: The conference management device extracts the installation position coordinates of each candidate video acquisition device in the candidate video acquisition device set one by one. These coordinates refer to the specific position coordinates of the candidate video acquisition device in the three-dimensional space of the conference scene, with the preset origin of the conference scene as the reference, and are consistent with the reference of all the aforementioned spatial coordinates; then, the horizontal component, vertical component, and depth component of the installation position coordinates of the candidate video acquisition device, as well as the horizontal component, vertical component, and depth component of the future position coordinates of the effective speaker target, are extracted respectively; the difference between the horizontal component, the difference between the vertical component, and the difference between the depth component of the two coordinates are calculated, and the three differences are squared respectively. The results of the three squared operations are added together to obtain a sum, and then the square root operation is performed on the sum. The result is the spatial distance between the candidate video acquisition device and the effective speaker target.
[0156] After calculating the spatial distance between all candidate video capture devices and the target speaker, the conference management device compares all spatial distances and selects the candidate video capture device with the smallest spatial distance, which is then designated as the primary imaging device. The primary imaging device is the video capture device that takes over from the current video capture device and is primarily responsible for tracking and capturing the target speaker. Its shooting angle is closest to the future position of the target speaker, ensuring the clarity and stability of the captured image.
[0157] The conference management device sends an imaging center adjustment control command to the main imaging device to control its imaging center to align with the future position coordinates of the valid speaker target. The specific control process is consistent with the imaging center movement control logic in step 70 above, that is, based on the spatial offset vector between the current imaging center coordinates of the main imaging device and the future position coordinates of the valid speaker target, the shooting angle and focal length of the main imaging device are adjusted so that the future position of the valid speaker target is within the stable area corresponding to the imaging center of the main imaging device.
[0158] Step 804: If no target synchronization event is detected within a consecutive preset time window, the tracking of the valid speaker target is terminated, and the corresponding video acquisition device control authority is released.
[0159] Optionally, the determination condition for the target synchronization event is: the time interval between the point in time when the lip opening area corresponding to the effective speaker target increases between adjacent video frames and the point in time when the sound source energy intensity increases between adjacent audio sampling times does not exceed a preset synchronization time tolerance. The point in time when the lip opening area increases between adjacent video frames is the aforementioned first time point, and the point in time when the sound source energy intensity increases between adjacent audio sampling times is the aforementioned second time point. This determination condition is used to verify whether the effective speaker target is still speaking. If no target synchronization event is detected for a long time, it indicates that the effective speaker target has stopped speaking and no further tracking is needed.
[0160] The conference management device initiates the target synchronization event detection process. The specific detection process is as follows: For each time window in a series of preset time windows, the lip movement state information corresponding to the effective speaker target is checked one by one. The first time point when the lip opening area increases between adjacent video frames is extracted. At the same time, the time sequence of sound source energy intensity is checked, and the second time point when the sound source energy intensity increases between adjacent audio sampling times is extracted. The time interval between each first time point and the corresponding second time point is calculated. It is determined whether the time interval does not exceed the preset synchronization time tolerance. If there is at least one time interval that meets the condition, it is determined that a target synchronization event has been detected in the time window. If no first time point or second time point is extracted in the time window, or if the time interval between all extracted first time points and second time points exceeds the preset synchronization time tolerance, it is determined that no target synchronization event has been detected in the time window.
[0161] The meeting management device statistically analyzes the target synchronization event detection results within a consecutive preset time window. If no target synchronization event is detected within any consecutive preset time window, it is determined that the valid speaker target has stopped speaking, and the tracking process for the valid speaker target is terminated. At the same time, a control permission release command is sent to the corresponding video acquisition device (main imaging device or currently controlled video acquisition device) to release the control permission of the video acquisition device, so that the video acquisition device can be restored to the default shooting state and can be used for subsequent tracking and shooting of other valid speaker targets. If a target synchronization event is detected in at least one time window within a consecutive preset time window, it is determined that the valid speaker target is still speaking, and the tracking control and target synchronization event detection of the valid speaker target are maintained.
[0162] The embodiments of the present invention realize the intelligent control of dynamic tracking of valid speaker targets, prediction of field of view switching, and tracking termination. It effectively avoids problems such as screen interruption caused by valid speaker targets moving out of the field of view, unstable switching of multiple video acquisition devices, and invalid tracking occupying device resources. It ensures that only valid speaker targets that are continuously speaking are tracked, thus guaranteeing the continuity and smoothness of video conferences.
[0163] Furthermore, the conference speaker tracking device provided by the present invention will be described below. The conference speaker tracking device described below can be referred to in correspondence with the conference speaker tracking method described above.
[0164] Optional, refer to Figure 2 , Figure 2 This is a schematic diagram of the conference speaker tracking device provided by the present invention. The conference speaker tracking device includes: The audio and video analysis module 210 is used to determine the sound source location information based on the multi-channel audio signals acquired by the audio acquisition device, and to extract the initial face region set under each viewpoint based on the multi-view video frame sequence acquired by the video acquisition device. The spatial mapping module 220 is used to perform spatial mapping based on the sound source orientation information and the initial set of face regions to obtain a set of target face regions located within the sound source orientation coverage area. The lip feature recognition module 230 is used to extract lip region features from video frames of corresponding viewpoints based on each face region in the target face region set to obtain lip movement state information of each speaking target. The target tracking and positioning module 240 is used to perform time synchronization analysis based on lip movement state information and sound source orientation information to obtain the effective speaker target, and to control the imaging center of at least one video acquisition device to align with the effective speaker target based on the spatial position of the effective speaker target in the current video frame.
[0165] The embodiments of the present invention solve the problems of sound source localization jumps and drifts caused by multiple participants speaking alternately or environmental reverberation interference, as well as frequent camera perspective switching and inaccurate focus, thereby improving the continuity of video conference footage.
[0166] Please see Figure 3 , Figure 3 An embodiment diagram of an electronic device provided in accordance with the present invention. For example... Figure 3 As shown, an embodiment of the present invention provides an electronic device 300, including a memory 310, a processor 320, and a computer program 311 stored in the memory 310 and executable on the processor 320. When the processor 320 executes the computer program 311, it implements the processes of steps 10 to 40.
[0167] Please see Figure 4 , Figure 4 An embodiment diagram of a computer-readable storage medium provided in accordance with an embodiment of the present invention is shown. Figure 4 As shown, this embodiment provides a computer-readable storage medium 400 on which a computer program 311 is stored. When the computer program 311 is executed by a processor, it implements the processes of steps 10 to 40.
[0168] On the other hand, the present invention also provides a computer program product, which includes a computer program that can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer is able to execute the conference speaker tracking method provided by the above methods, which includes the process of steps 10 to 40.
[0169] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for tracking conference speakers, characterized in that, include: The sound source location information is determined based on the multi-channel audio signals collected by the audio acquisition device, and the initial face region set under each viewpoint is extracted based on the multi-view video frame sequence collected by the video acquisition device. Based on the sound source location information and the initial set of face regions, a spatial mapping is performed to obtain a set of target face regions located within the sound source location coverage area. Based on each face region in the target face region set, lip region features are extracted from the video frames of the corresponding viewpoint to obtain the lip movement state information of each speaking target. Based on the lip movement state information and the sound source location information, a time synchronization analysis is performed to obtain the effective speaker target. Based on the spatial position of the effective speaker target in the current video frame, the imaging center of at least one video acquisition device is controlled to align with the effective speaker target.
2. The method for tracking conference speakers according to claim 1, characterized in that, The step of spatially mapping the sound source location information with the initial set of face regions to obtain a set of target face regions located within the sound source location coverage area includes: Based on the sound source's horizontal and vertical angle ranges in the sound source's location information, the cone-shaped coverage area of the sound source in three-dimensional space is determined. Based on the geometric correspondence between the cone-shaped coverage area and the imaging field of view of each video acquisition device, the two-dimensional projection coverage area corresponding to the viewpoint of each video acquisition device is determined. Based on the spatial coordinates of each face region in the initial face region set under the corresponding viewpoint within the two-dimensional projection coverage area, a first subset of face regions located within the two-dimensional projection coverage area is determined. The target face region set is determined based on the bounding box coordinates of each face region in the first face region subset and the boundary coordinates of the two-dimensional projection coverage area.
3. The method for tracking conference speakers according to claim 2, characterized in that, The step of determining the target face region set based on the bounding box coordinates of each face region in the first face region subset and the boundary coordinates of the two-dimensional projection coverage area includes: Based on the bounding box coordinates of each face region in the first face region subset and the boundary coordinates of the two-dimensional projection coverage area, an inclusion relationship analysis is performed to determine the inclusion relationship between each face region and the two-dimensional projection coverage area. Based on the face regions determined to be completely contained or partially overlapping in the inclusion relationship determination results, a set of candidate face regions is determined. Based on the center point coordinates of each first candidate face region in the candidate face region set in the corresponding video frame, and the center direction vector of the two-dimensional projection coverage area, the angular deviation value of each first candidate face region relative to the main axis direction of the sound source is obtained. The target face region set is determined based on the changing trend of the angle deviation values of each first candidate face region in a continuous preset number of video frames.
4. The method for tracking conference speakers according to claim 3, characterized in that, The determination of the target face region set based on the changing trend of the angle deviation values of each first candidate face region in a consecutive preset number of video frames includes: Based on the changing trend of the angle deviation values of each first candidate face region in a consecutive preset number of video frames, the dynamic fluctuation amplitude of the angle deviation is determined, and the first candidate face regions whose dynamic fluctuation amplitude of the angle deviation exceeds the preset stability upper limit are eliminated to obtain the second candidate face regions. Based on the angular deviation values of each second candidate face region under at least two different viewpoints, the spatial pointing is reversed to obtain the three-dimensional spatial pointing, and based on whether the three-dimensional spatial pointing intersects within the cone-shaped coverage area of the sound source, the subset of the second face region is determined. Based on the angle between the imaging optical axis of the video acquisition device and the main axis of the sound source for each face region in the second face region subset, it is determined whether each face region is within the main lobe range of the effective radiation of the sound source energy, and a third face region subset is obtained. The target face region set is determined based on whether the angle deviation values of each face region in the third face region subset remain consistent in adjacent video frames.
5. The method for tracking conference speakers according to claim 1, characterized in that, The steps for obtaining the effective speaker target through time synchronization analysis include: Based on the first time point in the lip opening and closing time sequence corresponding to each speaking target in the lip movement state information where the lip opening area increases between adjacent frames, and the second time point in the sound source energy intensity time sequence in the sound source orientation information where the sound source energy intensity increases between adjacent sampling times, a joint time alignment sequence is obtained. Based on the time interval between the first and second time points in the joint timing alignment sequence, paired time points with a time interval not exceeding a preset synchronization time tolerance are selected to obtain the set of synchronization events corresponding to each speaking target. Based on the set of synchronous events for each speaking target, the number of synchronous events contained in each speaking target within a consecutive preset time window is counted to obtain the synchronous event count for each speaking target; Synchronization analysis is performed based on the synchronization event count of each speaking target to obtain the effective speaker targets.
6. The method for tracking conference speakers according to claim 5, characterized in that, The synchronization analysis based on the synchronization event count of each speaking target yields valid speaker targets, including: Based on the relationship between the synchronization event count of each speaking target and the preset minimum synchronization event number threshold, speaking targets whose synchronization event count is less than the preset minimum synchronization event number threshold are eliminated, resulting in the first candidate speaking target set. The second set of candidate speaking targets is determined based on whether each candidate speaking target in the first set of candidate speaking targets is earlier than or equal to the second time point in each synchronization event at the first time point. The third set of candidate speaking targets is determined based on whether the variation range of the time interval of the synchronization events of each candidate speaking target in the second set of candidate speaking targets within a consecutive preset number of time windows does not exceed the preset time jitter tolerance. The effective speaker target is determined based on whether each candidate speaker target in the third candidate speaker target set has a lip opening area increase event corresponding to its spatial position in the video frame sequence acquired by different video acquisition devices, and whether the time interval between the lip opening area increase event and the sound source energy intensity increase event does not exceed the preset synchronization time tolerance.
7. The method for tracking conference speakers according to any one of claims 1 to 6, characterized in that, The method for tracking conference speakers also includes: Based on the spatial position coordinates of the effective speaker target in the current video frame and the current imaging center coordinates of at least one video acquisition device, a spatial offset vector is determined. Based on whether the horizontal and vertical components of the spatial offset vector are both less than the corresponding preset center alignment tolerance, the spatial position judgment result is determined; the spatial position judgment result indicates whether the effective speaker target is within the stable area corresponding to the imaging center of the video acquisition device. If the spatial position determination result indicates that it is not in a stable area, the imaging center is controlled to move towards the spatial position of the effective speaker target, and the displacement vector between adjacent frames is analyzed based on the spatial position coordinate sequence of the effective speaker target in a series of preset video frames to obtain the direction of motion trend. The video acquisition device is controlled based on the direction of the motion trend.
8. The method for tracking conference speakers according to claim 7, characterized in that, The video acquisition device controlled based on the motion trend direction includes: Based on the relative relationship between the direction of motion trend and the field of view boundary of the video acquisition device, it is predicted whether the effective speaker target will move out of the current field of view within a preset number of future frames, and a field of view loss warning signal is obtained. If the field of view loss warning signal indicates that the target has moved out of the current field of view, then based on the field of view coverage of each video acquisition device, candidate video acquisition devices that contain the future location of the effective speaker target are selected. Based on the spatial distance between each candidate video acquisition device and the effective speaker target, the candidate video acquisition device with the smallest spatial distance is selected as the main imaging device, and its imaging center is aligned with the effective speaker target. If no target synchronization event is detected within a consecutive preset time window, the tracking of the effective speaker target is terminated, and the corresponding video acquisition device control permissions are released.
9. A conference speaker tracking device, characterized in that, Applied to the conference speaker tracking method as described in any one of claims 1 to 8; The conference speaker tracking device includes: The audio and video analysis module is used to determine the location information of the sound source based on the multi-channel audio signals collected by the audio acquisition device, and to extract the initial set of face regions from each viewpoint based on the multi-view video frame sequence collected by the video acquisition device. The spatial mapping module is used to perform spatial mapping based on the sound source orientation information and the initial set of face regions to obtain a set of target face regions located within the sound source orientation coverage area; The lip feature recognition module is used to extract lip region features from video frames of corresponding viewpoints based on each face region in the target face region set to obtain lip movement state information of each speaking target. The target tracking and positioning module is used to perform time synchronization analysis based on the lip movement state information and the sound source orientation information to obtain the effective speaker target, and to control the imaging center of at least one video acquisition device to align with the effective speaker target based on the spatial position of the effective speaker target in the current video frame.
10. An electronic device, comprising: Memory, used to store computer software programs; A processor for reading and executing the computer software program, characterized in that, when the processor executes the computer software program, it implements the conference speaker tracking method as described in any one of claims 1 to 8.