A conference room audio-video synchronization visualization control method and system

CN122601899APending Publication Date: 2026-08-18MAINTENANCE COMPANY OF STATE GRID XINJIANG ELECTRIC POWER COMPANY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610827522.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-09
Publication Date
2026-08-18

AI Technical Summary

Technical Problem

[0003]但是现有技术在实际应用中仍存在一定的局限性:在中大型会议室场景下,音画同步效果往往难以达到预期,会出现轻微但持续的时空漂移现象;具体表现为发言者的口型与声音不同步,或者发言区可视化标注与实际发言位置存在偏差

Benefits of technology

[0022] The beneficial effects of this invention are as follows: By uniformly converting audio and video media clocks to wall clock time, a unified time reference is established, providing a foundation for subsequent physical sound path correction and helping to improve the accuracy of time conversion. By calculating the reflected sound path difference using the mirror sound source method and incorporating frame raster coefficients based on the inter-frame optical axis changes of the camera, dynamic synchronization compensation based on the physical structure of the conference room and the video frame state is achieved, correcting structural synchronization errors. By correcting the audio acquisition time before associating audio and video events, the reflected sound path offset is avoided from being mixed into network latency, improving the accuracy of audio and video event matching. By using a continuous non-negative reference delay calculation method and locking the presentation time to the display refresh raster, smooth distribution of audio and video delays is achieved, eliminating playback jitter and secondary desynchronization problems. By converting the compensation time shift into the width of the visual bar and generating joint control commands, the visualization of the synchronization compensation process and unified control of audio and video are achieved, improving the system's debuggability and user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122601899A_ABST
    Figure CN122601899A_ABST
Patent Text Reader

Abstract

The application relates to the technical field of audio-video synchronization control, and discloses a conference room audio-video synchronization visual control method and system, wherein the conference room audio-video synchronization visual control method comprises the following steps: step S1, calculating audio-video collection time; step S2, calculating compensation time shift; step S3, extracting a target audio packet; step S4, calculating audio playing delay and video waiting delay; step S5, calculating a presentation moment; step S6, calculating a graph strip pixel width; and step S7, generating a joint control instruction. The application calculates the reflection sound path difference through a mirror sound source method, introduces a frame grid coefficient in combination with the interframe optical axis change of a camera, realizes dynamic synchronization compensation based on the physical structure of a conference room and the video frame state, can correct structural synchronization errors, locks the presentation moment to a display refresh grid, realizes the visualization of the synchronization compensation process and the unified control of audio-video and visualization.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of audio and video synchronization control technology, and more specifically, to a visual control method and system for audio and video synchronization in a conference room. Background Technology

[0002] Audio and video synchronization is one of the core functions of a conferencing system, directly impacting communication efficiency and user experience. Current mainstream audio and video synchronization technologies are based on the RTP / RTCP protocol. This establishes a time correspondence between media streams by comparing the wall clock time reported by the RTCP sender with the RTP timestamp. Combined with buffering and presentation control at the receiver, it compensates for latency differences caused by network transmission and encoding processing. Some systems also incorporate lip-sync and speech correlation analysis, dynamic buffer adjustment, and other technologies to further improve synchronization accuracy.

[0003] However, existing technologies still have certain limitations in practical applications: in medium to large conference room scenarios, the audio-visual synchronization effect often fails to meet expectations, resulting in a slight but persistent spatiotemporal drift phenomenon; specifically, the speaker's lip movements and voice are out of sync, or there is a discrepancy between the visual annotation of the speaking area and the actual speaking position. This phenomenon is particularly noticeable in speaking areas near highly reflective surfaces such as glass walls and whiteboards, and it is exacerbated by camera rotation or electronic cropping.

[0004] In summary, existing technologies only address synchronization issues at the digital media stream level, neglecting the impact of the conference room's physical environment on audio propagation. Hard reflective boundaries in the conference room generate reflected sound, causing the audio events received by the ceiling array microphones to lag behind the actual sound emission time. Furthermore, video signals are acquired and displayed in discrete frames; when camera status changes, the same acoustic time shift will appear as different offsets on the screen. Existing technologies struggle to distinguish between this structural synchronization error and network latency, thus hindering the achievement of true physical time synchronization. Summary of the Invention

[0005] This invention provides a method and system for synchronized visual control of audio and video in a conference room, which solves the technical problems mentioned in the background.

[0006] This invention provides a method for visually controlling synchronized audio and video in a conference room, comprising the following steps:

[0007] Step S1: Convert the audio timestamp and video timestamp into audio capture time and video capture time respectively based on the reference timestamp and clock frequency;

[0008] Step S2: Extract the mirror source coordinates based on the coordinates of the sound-emitting area, the boundary normal vector, and the coordinates of the boundary points. Calculate the direct sound path and the equivalent sound path based on the array coordinates, combined with the coordinates of the sound-emitting area and the mirror source coordinates. Calculate the acoustic time difference accordingly. Combine the acoustic time difference, the video frame period, and the optical axis angular velocity to calculate the compensation time shift.

[0009] Step S3: Subtract the compensation time shift from the audio acquisition time to obtain the corrected audio time, and extract the target audio packet based on the principle of minimizing the difference between the video acquisition time and the corrected audio time. Subtract the audio arrival time corresponding to the target audio packet from the video arrival time to obtain the arrival time difference.

[0010] Step S4: Calculate the audio playback delay and video waiting delay based on the arrival time difference and display refresh cycle;

[0011] Step S5: Calculate the presentation time based on the display refresh cycle, the video arrival time, and the video waiting delay;

[0012] Step S6: Use the mapping matrix to convert the sound area coordinates into screen pixel coordinates, and combine the compensation time shift, video frame period, basic pixel width and pixel width increment to calculate the strip pixel width;

[0013] Step S7: Combine the target audio package, audio playback delay, video waiting delay, presentation time, screen pixel coordinates, and bar pixel width to generate a joint control command.

[0014] This invention provides a conference room audio-video synchronization and visualization control system, comprising:

[0015] The acquisition time conversion module converts the audio timestamp and video timestamp into audio acquisition time and video acquisition time respectively, based on the reference timestamp and clock frequency.

[0016] The time-shift compensation module extracts the mirror source coordinates based on the coordinates of the sound-emitting area, the boundary normal vector, and the boundary point coordinates. It calculates the direct sound path and the equivalent sound path based on the array coordinates, the coordinates of the sound-emitting area, and the coordinates of the mirror source, respectively. Based on this, it calculates the acoustic time difference and calculates the time-shift compensation by combining the acoustic time difference, the video frame period, and the optical axis angular velocity.

[0017] The target audio packet extraction module subtracts the compensation time shift from the audio acquisition time to obtain the corrected audio time, and extracts the target audio packet based on the principle of minimizing the difference between the video acquisition time and the corrected audio time. The arrival time difference is obtained by subtracting the arrival time of the audio corresponding to the target audio packet from the arrival time of the video.

[0018] The audio and video latency calculation module calculates audio playback latency and video waiting latency based on the arrival time difference and display refresh cycle.

[0019] The presentation time calculation module calculates the presentation time based on the display refresh cycle, the video arrival time, and the video waiting delay.

[0020] The pixel width calculation module uses a mapping matrix to convert the sound area coordinates into screen pixel coordinates, and combines compensation time shift, video frame period, basic pixel width and pixel width increment to calculate the pixel width of the bar.

[0021] The control instruction generation module combines the target audio package, audio playback delay, video waiting delay, presentation time, screen pixel coordinates, and bar pixel width to generate a joint control instruction.

[0022] The beneficial effects of this invention are as follows: By uniformly converting audio and video media clocks to wall clock time, a unified time reference is established, providing a foundation for subsequent physical sound path correction and helping to improve the accuracy of time conversion. By calculating the reflected sound path difference using the mirror sound source method and incorporating frame raster coefficients based on the inter-frame optical axis changes of the camera, dynamic synchronization compensation based on the physical structure of the conference room and the video frame state is achieved, correcting structural synchronization errors. By correcting the audio acquisition time before associating audio and video events, the reflected sound path offset is avoided from being mixed into network latency, improving the accuracy of audio and video event matching. By using a continuous non-negative reference delay calculation method and locking the presentation time to the display refresh raster, smooth distribution of audio and video delays is achieved, eliminating playback jitter and secondary desynchronization problems. By converting the compensation time shift into the width of the visual bar and generating joint control commands, the visualization of the synchronization compensation process and unified control of audio and video are achieved, improving the system's debuggability and user experience. Attached Figure Description

[0023] Figure 1 This is a calculation flowchart of a conference room audio and video synchronization visualization control method according to the present invention. Detailed Implementation

[0024] The subject matter described herein will now be discussed with reference to exemplary embodiments. It should be understood that these embodiments are discussed only to enable those skilled in the art to better understand and implement the subject matter described herein, and changes may be made to the function and arrangement of the elements discussed without departing from the scope of this specification. Various processes or components may be omitted, substituted, or added as needed in the examples. Furthermore, features described in some examples may be combined in other examples.

[0025] It should be noted that, unless otherwise defined, the technical or scientific terms used in one or more embodiments of the present invention should have the ordinary meaning understood by one of ordinary skill in the art to which this invention pertains. The terms "first," "second," and similar terms used in one or more embodiments of the present invention do not indicate any order, quantity, or importance, but are merely used to distinguish different components. Terms such as "comprising" or "including" indicate that the element or object preceding the term encompasses the elements or objects listed following the term and their equivalents, without excluding other elements or objects. Terms such as "connected" or "linked" are not limited to physical or mechanical connections, but can include electrical connections, whether direct or indirect. Terms such as "upper," "lower," "left," and "right" are used only to indicate relative positional relationships; when the absolute position of the described object changes, the relative positional relationship may also change accordingly.

[0026] like Figure 1 As shown, a method for synchronized and visualized control of audio and video in a conference room includes the following steps:

[0027] Step S1: Convert the audio timestamp and video timestamp into audio capture time and video capture time respectively based on the reference timestamp and clock frequency;

[0028] Step S2: Extract the mirror source coordinates based on the coordinates of the sound-emitting area, the boundary normal vector, and the coordinates of the boundary points. Calculate the direct sound path and the equivalent sound path based on the array coordinates, combined with the coordinates of the sound-emitting area and the mirror source coordinates. Calculate the acoustic time difference accordingly. Combine the acoustic time difference, the video frame period, and the optical axis angular velocity to calculate the compensation time shift.

[0029] Step S3: Subtract the compensation time shift from the audio acquisition time to obtain the corrected audio time, and extract the target audio packet based on the principle of minimizing the difference between the video acquisition time and the corrected audio time. Subtract the audio arrival time corresponding to the target audio packet from the video arrival time to obtain the arrival time difference.

[0030] Step S4: Calculate the audio playback delay and video waiting delay based on the arrival time difference and display refresh cycle;

[0031] Step S5: Calculate the presentation time based on the display refresh cycle, the video arrival time, and the video waiting delay;

[0032] Step S6: Use the mapping matrix to convert the sound area coordinates into screen pixel coordinates, and combine the compensation time shift, video frame period, basic pixel width and pixel width increment to calculate the strip pixel width;

[0033] Step S7: Combine the target audio package, audio playback delay, video waiting delay, presentation time, screen pixel coordinates, and bar pixel width to generate a joint control command.

[0034] In one embodiment of the present invention, the specific calculation process of step S1 includes:

[0035] The formula for calculating the audio acquisition time based on the audio reference time, audio timestamp, audio reference stamp, and audio clock frequency is as follows:

[0036]

[0037] in, Indicates the first Audio capture time for each audio packet Indicates the audio package number. When representing an audio reference, Indicates the first Audio timestamps of each audio packet Indicates an audio reference stamp. Indicates the audio clock frequency;

[0038] The formula for calculating the video capture time, based on the video reference time, video timestamp, video reference stamp, and video clock frequency, is as follows:

[0039]

[0040] in, Indicates the first The video capture time for each video frame Indicates the video frame number. When referring to a video, Indicates the first The video timestamp of each video frame Indicates a video reference stamp. Indicates the video clock frequency.

[0041] It should be noted that the audio reference time is the wall clock time corresponding to the audio RTCP (Real-Time Transport Control Protocol) reference point, which can be obtained from the RTCP sender report sent by the receiving audio and video devices. The video reference time is the wall clock time corresponding to the video RTCP reference point, and its acquisition method is the same as that of the audio reference time. The audio reference stamp is the audio reference RTP (Real-Time Transport Protocol) timestamp, extracted from the RTCP sender report. The video reference stamp is the video reference RTP timestamp, and its extraction method is the same as that of the audio reference stamp.

[0042] The audio clock frequency is the audio RTP clock frequency, which is a preset parameter or session negotiation parameter. Its value is determined by the clock frequency in the audio RTP payload type or session description, and is usually consistent with the audio sampling period frequency; for example, it can be 8000 ticks per second, 16000 ticks per second, 48000 ticks per second, or 96000 ticks per second. For audio streams using a specific RTP payload format, the value can also be determined according to the specifications of that payload format (for example, the RTP payload format corresponding to G.711 audio encoding specifies a clock frequency of 8000 ticks per second, while the RTP payload format corresponding to AAC audio encoding can support clock frequencies of 48000 ticks per second or 96000 ticks per second). The video clock frequency is the video RTP clock frequency, which is a preset parameter or session negotiation parameter. Its value is determined by the clock frequency in the video RTP load type or session description. In common RTP audio and video conferencing loads, the video RTP timestamp frequency is preferably 90,000 ticks per second. This video clock frequency is different from the video frame rate. The tick is the smallest counting unit of the RTP timestamp and is not a time unit itself. The actual time it represents depends on the clock frequency.

[0043] The audio acquisition time is the unified wall clock acquisition time converted from the audio packet, representing the actual physical moment when the sound wave corresponding to the audio packet was captured by the audio acquisition device. The video acquisition time is the unified wall clock acquisition time converted from the video frame, representing the actual physical moment when the image corresponding to the video frame was captured by the video acquisition device.

[0044] Specifically, the RTCP reference point is obtained by receiving RTCP sender reports periodically sent by the audio and video devices. These reports contain the current device's wall clock time and the corresponding RTP timestamp. The sending interval of the RTCP sender reports can be determined based on the RTCP bandwidth, the number of session participants, and the interval calculation rules specified in the protocol, and can be randomized. In practice, the reference point update period can be configured to approximately 5 seconds or other periods that meet synchronization accuracy requirements (e.g., when synchronization accuracy is required to be better than 0.5 milliseconds, the reference point update period can be configured to 2 seconds; when synchronization accuracy is required to be 2 milliseconds, it can be configured to 10 seconds). After receiving the latest RTCP sender report, the controller automatically updates the corresponding audio reference time, audio reference stamp, video reference time, and video reference stamp to ensure the accuracy of time conversion.

[0045] The RTP timestamps of audio packets and video frames are extracted from the RTP header fields of the corresponding media streams. The second 32-bit field of the RTP header is the RTP timestamp field. When the controller receives audio and video media streams, it parses each RTP packet, extracts the RTP timestamp, and associates it with the corresponding packet sequence number for storage.

[0046] It should be noted that when the audio / video device restarts or the network is restored after an interruption, the controller needs to reacquire the RTCP reference point and reset the time conversion base to avoid time conversion errors caused by device clock reset. Specifically, after the controller detects that the audio / video stream is interrupted for more than 10 seconds, it automatically clears the currently stored RTCP reference point. After the audio / video stream is restored and a new RTCP sender report is received, the time conversion base is re-established.

[0047] It should be further explained that the embodiment corresponding to step S1 uniformly converts the audio and video media clocks with different sampling frequencies and timestamp scales into wall clock time, providing a unified time reference for the subsequent introduction of physical path correction, and avoiding misjudging the sound path error reflected in the conference room as simple network latency. Traditional audio and video synchronization methods usually align directly based on RTP timestamps, without considering the influence of the conference room's physical environment on audio propagation time, resulting in structural path differences mixed into the synchronization error. This embodiment establishes a unified wall clock time axis, mapping audio and video events to the same real time coordinate, so that the subsequent physical path correction can be accurately superimposed on the time reference, thereby eliminating the time reference inconsistency problem caused by the clock differences of different media streams.

[0048] In one embodiment of the present invention, the specific calculation process of step S2 includes:

[0049] The formula for calculating the coordinates of the mirror source, based on the coordinates of the sound-emitting area, the boundary normal vector, and the coordinates of the boundary points, is as follows:

[0050]

[0051] in, Indicates the first The mirror source coordinates corresponding to each video frame Indicates the first The coordinates of the sound area corresponding to each video frame Represents the boundary normal vector. Represents the coordinates of the boundary points. Represents the transpose of the boundary normal vector;

[0052] The formula for calculating the direct sound path is as follows, based on the coordinates of the sound-generating area and the array coordinates:

[0053]

[0054] in, Indicates the first The direct path of sound corresponding to each video frame. Indicates the first The coordinates of the sound area corresponding to each video frame Indicates array coordinates, Represents the Euclidean norm;

[0055] The formula for calculating the equivalent sound path is as follows, based on the coordinates of the mirror source and the array:

[0056]

[0057] in, Indicates the first The equivalent sound path corresponding to each video frame Indicates the first The mirror source coordinates corresponding to each video frame Indicates array coordinates, Represents the Euclidean norm;

[0058] The steps to calculate the acoustic time difference based on the equivalent sound path, the direct sound path, and the sound speed are as follows: subtract the direct sound path from the equivalent sound path to obtain the sound path difference, and divide the sound path difference by the sound speed to obtain the acoustic time difference.

[0059] The formula for calculating the optical axis angular velocity based on the current optical axis vector, the previous frame's optical axis vector, and the video frame period is as follows:

[0060]

[0061] in, Indicates the first The optical axis angular velocity corresponding to each video frame Indicates the first The current optical axis vector of each video frame. Indicates the first The optical axis vector of the previous frame corresponding to each video frame. Indicates the video frame period. Represents the inverse cosine function. Indicates the transpose of the current optical axis quantity;

[0062] Based on the equivalent sound path, direct sound path, sound speed, optical axis angular velocity, and video frame period, the formula for calculating the time shift compensation is as follows:

[0063]

[0064] in, Indicates the first Compensated time shift corresponding to each video frame Indicates the first The equivalent sound path corresponding to each video frame Indicates the first The direct path of sound corresponding to each video frame. Indicates the speed of sound. Indicates the first The optical axis angular velocity corresponding to each video frame Indicates the video frame period. This represents the cosine function.

[0065] It should be noted that the sound area coordinates are the coordinates of the representative point of the sound area corresponding to the video frame, which can be obtained through seat calibration, sound area geometry of the ceiling array microphone output, or video area geometry analysis. The boundary normal vector is the unit normal vector of the dominant reflection boundary, which is a user-defined parameter and is obtained through pre-calibration of the conference room structure. The boundary point coordinates are the coordinates of known calibration points on the dominant reflection boundary, which are user-defined parameters and are obtained through pre-calibration of the conference room structure.

[0066] The array coordinates are the coordinates of the array center of the ceiling array microphone. These are user-defined parameters obtained during device installation and calibration. The mirror source coordinates are the coordinates of the mirror source point formed by the sound-emitting area coordinates with respect to the dominant reflection boundary, representing the spatial location of the equivalent reflected sound source. The direct path is the Euclidean distance from the representative point of the sound-emitting area to the array center, representing the path length of the sound wave traveling directly from the sound-emitting area to the array center.

[0067] The equivalent sound path is the Euclidean distance from the mirrored sound source point to the center of the array, representing the equivalent path length of a sound wave propagating from the sound-generating area, after reflection at the dominant reflection boundary, to the center of the array. Acoustic time difference is the fundamental time offset caused by the difference in reflected sound path, representing the time difference in propagation of reflected sound relative to direct sound. Optical axis angular velocity is the camera's line-of-sight angular velocity or electronically clipped equivalent angular velocity between adjacent frames, representing the rate of change of the camera's line-of-sight between adjacent frames.

[0068] The frame raster coefficient is a dimensionless coefficient reflecting the amplification or compression effect of the video frame raster on acoustic time shift; its value changes with the rate of change of the camera's line of sight. Compensation time shift is the mirror path video frame raster time shift, representing the length of time required to correct the audio acquisition time.

[0069] Specifically, the dominant reflective boundary is pre-calibrated based on the conference room structure, prioritizing highly reflective surfaces such as glass walls, whiteboards, and video walls. The calibration process involves: after the conference room is deployed, using a laser rangefinder to measure the spatial coordinates of at least three non-collinear points on the reflective plane; calculating the equation of the reflective plane through plane fitting; and then obtaining the boundary normal vector and boundary point coordinates. For a conference room with multiple strong reflective surfaces (such as glass walls, whiteboards, and video walls), the parameters of each reflective surface are calibrated. During actual operation, the dominant reflective boundary is determined based on at least one of the following: the location of the sound-emitting area, the reflective boundary material, the reflective path length, and the array received energy. When multiple reflective surfaces meet the dominant reflective condition, the compensation time shift corresponding to each reflective surface can be calculated separately, and the final compensation time shift is determined according to a preset selection rule. The preset selection rule includes selecting the compensation time shift with the largest reflective energy, selecting the compensation time shift with the largest contribution of the reflective path, or performing weighted fusion of multiple compensation time shifts (for example, the compensation time shift corresponding to the glass wall has a weight of 0.6, the compensation time shift corresponding to the whiteboard has a weight of 0.4, and the final compensation time shift is equal to the compensation time shift of the glass wall multiplied by 0.6 plus the compensation time shift of the whiteboard multiplied by 0.4).

[0070] The coordinates of the sound-emitting zone can be obtained in three ways: The first is seat calibration, where the spatial coordinates of each seat are pre-measured and stored in the controller after the conference room is deployed. When a microphone corresponding to a seat is activated, the seat's coordinates are directly used as the sound-emitting zone coordinates. The second is acoustic zone geometry, which calculates the spatial position of the sound source using the beamforming output of the ceiling array microphones. The third is video region geometry, which detects the speaker's facial position through video image analysis and maps this position to the conference room's spatial coordinates as the sound-emitting zone coordinates. In actual operation, the controller prioritizes using acoustic zone geometry to obtain the sound-emitting zone coordinates. When acoustic zone geometry fails to obtain valid results (e.g., the acoustic zone energy value output by the ceiling array microphones is below a preset threshold of -40 dB, or the acoustic zone position deviation detected for three consecutive frames exceeds 1 meter), it is determined that no valid result can be obtained, and it automatically switches to seat calibration or video region geometry.

[0071] The camera's optical axis vector is read from the attitude parameters or electronic cropping parameters of a PTZ (Horizontal Rotation-Vertical Pitch-Zoom) camera. For a PTZ camera, the optical axis vector can be calculated from the camera's horizontal rotation angle, vertical pitch angle, and zoom magnification. Before calculation, the camera's intrinsic and extrinsic parameters need to be pre-calibrated. Intrinsic parameters include focal length, principal point coordinates, and distortion coefficients, while extrinsic parameters include the camera's position and attitude. The calibration process uses the Zhang Zhengyou calibration method and is completed using a checkerboard calibration board. For cameras with electronic cropping capabilities, the optical axis vector can be calculated from the center position and field of view of the electronic cropping area. Each time the controller receives a video frame, it simultaneously reads the corresponding camera attitude parameters or electronic cropping parameters, calculates the optical axis vector corresponding to that frame, and stores it.

[0072] It should be noted that when the conference room structure changes, such as moving the whiteboard or glass wall, the parameters of the dominant reflection boundary need to be recalibrated; when the camera position or installation angle changes, the camera's intrinsic and extrinsic parameters need to be recalibrated to ensure the accuracy of the optical axis vector calculation.

[0073] It should be further explained that the embodiment corresponding to step S2 equates the hard reflective boundary of the conference room to a mirror boundary, uses the mirror sound source method to calculate the reflected sound path difference, and introduces a frame grating coefficient in combination with the inter-frame optical axis change of the camera to solve the position-related and frame-related synchronization drift problems caused by the reflective structure of the conference room and the discrete sampling of video. In medium and large conference rooms, the sound received by the ceiling array microphone usually contains direct sound and multiple reflected sounds. The intensity of strong reflected sound may be close to or even exceed that of direct sound (for example, when the speaker is 0.5 meters away from the glass wall and 3 meters away from the ceiling array microphone, the intensity of the reflected sound from the glass wall can reach 0.8 to 1.2 times the intensity of the direct sound), causing the audio event captured by the array microphone to lag behind the actual sound time. At the same time, the video signal is acquired and displayed in the form of discrete frames. When the camera rotates or is electronically cropped, the same acoustic time shift will manifest as different image shifts in different video frames. This embodiment combines acoustic reflection characteristics with video frame grating characteristics to calculate dynamic compensation time shifts related to position and frame state, which can accurately correct this structural synchronization error and thus improve audio and video synchronization accuracy.

[0074] In one embodiment of the present invention, the specific calculation process of step S3 includes:

[0075] The formula for calculating the target audio packet index is as follows, based on the audio acquisition time, compensation time shift, and video acquisition time:

[0076]

[0077] in, Indicates the first The target audio packet index corresponding to each video frame. Indicates the candidate audio package number. Indicates the first Audio acquisition time for each candidate audio packet Indicates the first Compensated time shift corresponding to each video frame Indicates the first The video capture time for each video frame This indicates that the sequence number that minimizes the absolute value is selected from the candidate audio packet sequence numbers;

[0078] The formula for calculating the time difference of arrival (TDOA) based on the arrival time of the video and the arrival time of the corresponding audio packet is as follows:

[0079]

[0080] in, Indicates the first The time difference of arrival for each video frame Indicates the first The arrival time of each video frame Indicates the first The arrival time of the target audio packet corresponding to each video frame. Indicates the first The target audio packet index corresponding to each video frame.

[0081] It should be noted that the corrected audio time is the audio acquisition time minus the compensated time shift, resulting in a time approximating the actual sound emission time. This represents the physical moment of the actual sound event corresponding to the audio packet. The target audio packet is the optimal audio packet corresponding to the video frame; it is the audio packet with the smallest difference between the corrected audio time and the video acquisition time among all candidate audio packets.

[0082] The arrival time difference (OTD) is the difference between the arrival time of the video frame and the arrival time of the corresponding target audio packet, representing the time difference between when the video frame and the corresponding audio packet enter the controller. When the OTD is positive, it means that the video frame arrives at the controller later than the corresponding audio packet; when the OTD is negative, it means that the video frame arrives at the controller earlier than the corresponding audio packet.

[0083] Candidate audio packets are limited to all audio packets within three video frame periods before and after the current video frame capture time, avoiding a global search. Specifically, the controller maintains an audio packet cache queue with a length of seven video frame periods. When processing video frames, candidate audio packets are searched only in this cache queue. This approach effectively reduces computation and improves the real-time performance of event correlation. It should be noted that the length of the audio packet cache queue can be adjusted based on network jitter. When network jitter is significant, the cache queue length can be appropriately increased to ensure that the corresponding target audio packet can be found. For example, when the peak network jitter exceeds 20 milliseconds, the audio packet cache queue length can be increased from 7 video frame periods to 10 video frame periods; when the peak network jitter exceeds 50 milliseconds, it can be increased to 15 video frame periods.

[0084] The arrival times of audio packets and video frames are recorded by the same local system clock of the controller. Upon receiving each audio packet and video frame, the controller immediately reads the current value of its local system clock as the arrival time of that packet or frame and associates it with the corresponding packet or frame sequence number for storage. It should be noted that since both audio packet and video frame arrival times are recorded by the same local system clock of the controller, their arrival time difference does not depend on the clock synchronization of the transmitting device. When audio playback, video display, and visualization rendering are performed by different execution devices, the controller's local system clock and the output clocks of each execution device can be synchronized via the IEEE 1588 Precision Time Protocol or other clock synchronization mechanisms (such as the NTP protocol), preferably with a synchronization accuracy better than 1 millisecond.

[0085] When multiple audio packets have the same difference between their corrected audio time and video capture time, the audio packet with the largest sequence number is selected as the target audio packet. This is because audio packets with larger sequence numbers usually correspond to later capture times and are more likely to correspond to speech actions in the current video frame.

[0086] It should be further explained that the embodiment corresponding to step S3 first performs physical path correction on the audio acquisition time to restore the actual sound event time, and then associates it with the video acquisition time to ensure that the audio and video events are matched based on the same physical sound event moment. Traditional audio and video event association usually directly compares the audio acquisition time and the video acquisition time, without considering the path difference during audio propagation. This leads to the time lag caused by reflected sound being mistakenly identified as network latency or encoding latency, thus associating it with the wrong audio packet. This embodiment, by first correcting the audio acquisition time to obtain a corrected audio time that is closer to the actual sound event moment, and then associating it with the video acquisition time, can accurately match the audio and video corresponding to the same physical sound event, providing a correct basis for subsequent delay allocation.

[0087] In one embodiment of the present invention, the specific calculation process of step S4 includes:

[0088] The formula for calculating the baseline latency, based on the arrival time difference and the display refresh cycle, is as follows:

[0089]

[0090] in, Indicates the first The baseline delay corresponding to each video frame Indicates the first The time difference of arrival for each video frame Indicates the display refresh cycle;

[0091] Based on the baseline delay, the formula for calculating audio playback delay is as follows:

[0092]

[0093] in, Indicates the first Each video frame corresponds to the audio playback delay of the target audio packet. Indicates the first The baseline delay corresponding to each video frame;

[0094] The formula for calculating video waiting delay based on the baseline delay and the time difference of arrival is as follows:

[0095]

[0096] in, Indicates the first Video delay per video frame Indicates the first The baseline delay corresponding to each video frame Indicates the first The time difference of arrival for each video frame.

[0097] It should be noted that the baseline latency is a non-negative audio playback latency baseline, used to uniformly allocate the waiting time for audio and video. Audio playback latency is the playback waiting time for the target audio packet corresponding to the video frame, representing the length of time it takes for the target audio packet to arrive at the controller before playback begins. Video waiting latency is the display waiting time for the video frame, representing the length of time it takes for the video frame to arrive at the controller before display begins.

[0098] It should be noted that the display refresh rate is read from the extended display identifier data of the display device or set through system configuration parameters. Specifically, when the controller establishes a connection with the display device, it automatically reads the extended display identifier data of the display device, extracts the supported display refresh rates, and sets the display refresh rate to the reciprocal of the display refresh rate. When the display device supports multiple display refresh rates, the controller prioritizes the highest display refresh rate to obtain a smoother display effect. For example, a display device with a 60Hz refresh rate can update 60 frames per second, resulting in more continuous image movement and no noticeable stuttering compared to a 30Hz refresh rate device. If the display refresh rate cannot be read from the extended display identifier data, the controller uses the system-configured default display refresh rate, which has a default value of 16.67 milliseconds, corresponding to a 60Hz display refresh rate.

[0099] It should be noted that when the display refresh rate of the display device changes, the controller needs to reread the extended display identifier data and update the display refresh cycle to ensure that the delay allocation is consistent with the actual refresh capability of the display device.

[0100] It should be further explained that the embodiment corresponding to step S4 adopts a continuous non-negative baseline delay calculation method, using the display refresh cycle as the smooth dimension, to avoid the boundary jitter problem caused by the traditional hard threshold judgment of whether to delay audio or video. Traditional delay allocation methods usually use hard threshold judgment, delaying video when the arrival time difference is greater than the threshold, and delaying audio when the arrival time difference is less than the threshold. This method will frequently switch the delay object when the arrival time difference is close to the threshold, resulting in playback stuttering and audio-video synchronization jitter. The baseline delay calculation method used in this embodiment can continuously allocate the waiting time of audio and video according to the arrival time difference. Regardless of whether the arrival time difference is positive or negative, it can ensure that the audio playback delay and video waiting delay are both non-negative and change continuously and smoothly; that is, through the continuous delay allocation method, the controller can adaptively adjust the waiting time of audio and video without causing playback jitter, adapting to changes in network jitter and device processing latency.

[0101] In one embodiment of the present invention, the specific calculation process of step S5 includes:

[0102] The formula for calculating the presentation time, based on the display refresh cycle, video arrival time, and video waiting delay, is as follows:

[0103]

[0104] in, Indicates the first The presentation time of each video frame and the target audio packet Indicates the display refresh cycle. Indicates the first The arrival time of each video frame Indicates the first Video delay per video frame This represents the function for rounding up.

[0105] It should be noted that the display base time is the sum of the video frame arrival time and the video waiting delay, i.e., the theoretical presentation reference time, representing the earliest theoretical time when the video frame should be displayed. The refresh sequence number is the most recent display refresh grid sequence number no earlier than the display base time, representing the refresh cycle sequence number at which the video frame should be displayed. The presentation time is the unified output presentation time of the video frame and its corresponding audio packet, representing the actual physical time of video frame display and audio packet playback. The presentation time is aligned with the refresh grid of the display device to ensure that the video frame can be displayed at the stable refresh time of the display device.

[0106] The rounding up operation rounds up towards positive infinity to ensure that the rendering time is no earlier than the theoretical display base time. Specifically, when the display base time is exactly equal to the time of a certain display refresh grid, the refresh number is equal to the quotient of the display base time divided by the display refresh period. When the display base time is between two adjacent display refresh grids, the refresh number is equal to the integer part of the quotient of the display base time divided by the display refresh period plus 1. For example, when the display refresh period is 16.67 milliseconds and the display base time is 30 milliseconds, the quotient of the display base time divided by the display refresh period is approximately 1.8. After rounding up, the refresh number is 2, and the corresponding rendering time is 33.34 milliseconds.

[0107] It should be noted that the controller maintains a global display refresh clock, which is synchronized with the display device's refresh clock. Based on the current value of the display refresh clock and the calculated presentation time, the controller adds video frames and audio packets to the corresponding display queue and playback queue, respectively, waiting for the presentation time to arrive before outputting them. The length of both the display queue and the playback queue does not exceed 5 elements to avoid excessive latency due to overly long queues.

[0108] It should be further explained that the embodiment corresponding to step S5 locks the theoretical rendering reference to the discrete refresh grid of the display device, adapting to the discrete refresh characteristics of video display and reproduction, and avoiding secondary asynchrony caused by display refresh. Even if the controller calculates the precise theoretical rendering time, the display device can only update the image at a fixed refresh time. If the display is performed directly according to the theoretical rendering time, video frames may be output between two refreshes of the display device, resulting in screen tearing or unstable display delay. This embodiment ensures that video frames can be displayed at the stable refresh time of the display device by locking the rendering time to the refresh grid of the display device, and at the same time, the playback time of the audio package is also adjusted to the same refresh grid, avoiding secondary asynchrony of audio and video caused by display refresh.

[0109] In one embodiment of the present invention, the specific calculation process of step S6 includes:

[0110] The formula for calculating screen pixel coordinates based on the mapping matrix and the coordinates of the sound-emitting area is as follows:

[0111]

[0112] in, Indicates the first The screen x-coordinate corresponding to each video frame Indicates the first The screen vertical coordinate corresponding to each video frame Represents the mapping matrix. Indicates the first The horizontal coordinate of the sound-producing area corresponding to each video frame Indicates the first The vertical coordinate of the sound-producing area corresponding to each video frame Represents the third basis vector. This represents the transpose of the third basis vector;

[0113] The formula for calculating the pixel width of the bar, based on the base pixel width, pixel width increment, compensation time shift, and video frame period, is as follows:

[0114]

[0115] in, Indicates the first The width of the bar corresponding to each video frame (in pixels). Indicates the base pixel width. Indicates the pixel width increment. Indicates the first Compensated time shift corresponding to each video frame Indicates the video frame period.

[0116] It should be noted that homogeneous coordinates are three-dimensional coordinates composed of the coordinates of the sound-emitting area plane and a constant 1, used for planar projection transformation. Screen homogeneous coordinates are coordinates obtained by multiplying the mapping matrix by the homogeneous coordinates of the sound-emitting area, representing the homogeneous projection coordinates of the sound-emitting area coordinates on the screen pixel plane.

[0117] The normalized component is the third component of the screen homogeneous coordinates, used for homogeneous coordinate normalization. The screen pixel coordinates are the pixel coordinates of the representative point of the sound-emitting area mapped onto the display screen, representing the corresponding position of the sound-emitting area on the display screen.

[0118] The frame-time ratio is the dimensionless ratio of the compensated time shift to the video frame period, representing how many video frame periods the compensated time shift is equivalent to. The pixel increment is the increase in the width of the visualization bar calculated from the frame-time ratio and the pixel width increment, representing the increase in the width of the visualization bar due to the compensated time shift.

[0119] The pixel width of the image bar is the pixel width of the synchronized visualization bar corresponding to the video frame, representing the actual width of the synchronized visualization bar displayed on the screen. The base pixel width is the basic pixel length of the visualization bar when there is no compensation or minimal compensation. It is a custom parameter, with a preferred value of 10 pixels and a value range of 5 to 20 pixels. The value is adjusted according to the screen resolution; the higher the screen resolution, the larger the base pixel width.

[0120] The pixel width increment is the screen width calibrated value corresponding to one video frame cycle. It is a custom parameter, with a preferred value of 50 pixels per frame cycle. The value range is from 20 pixels per frame cycle to 100 pixels per frame cycle. The value is adjusted according to the screen size and viewing distance. The larger the screen size or the farther the viewing distance, the larger the pixel width increment.

[0121] It should be noted that the mapping matrix is ​​calculated based on the correspondence between at least four non-collinear calibration points on the conference room plane and their corresponding screen pixels. The specific calibration process is as follows: Select four non-collinear calibration points on the conference room plane, such as the four corners of the conference table, and measure the spatial coordinates of each calibration point using a laser rangefinder. Then, assign the corresponding screen pixels to these conference room plane calibration points in the visual control interface, or obtain the screen pixel coordinates corresponding to each calibration point through manual selection, interface calibration, or camera-assisted calibration. Finally, based on the correspondence between at least four sets of conference room plane coordinates and screen pixel coordinates, calculate the mapping matrix from the conference room plane coordinates to the screen pixel coordinates using a mapping matrix solving method (such as the direct linear transformation method). During calculation, the spatial coordinates of each calibration point are converted to homogeneous coordinates, and the corresponding screen pixel coordinates are also converted to homogeneous coordinates. Then, a system of linear equations is constructed, and the elements of the mapping matrix are obtained by solving the system of linear equations. It should be noted that when the position or angle of the display device changes, the mapping matrix needs to be recalibrated to ensure the accuracy of the coordinate mapping.

[0122] The preferred base pixel width is 10 pixels, and the preferred pixel width increment is 50 pixels per frame. For a 1920×1080 resolution display screen, the base pixel width can be set to 10 pixels, and the pixel width increment can be set to 50 pixels per frame. For a 3840×2160 resolution display screen, the base pixel width can be set to 20 pixels, and the pixel width increment can be set to 100 pixels per frame. In actual use, users can adjust these two parameters according to their own visual preferences to obtain the best visualization effect.

[0123] It should be further explained that the embodiment corresponding to step S6 converts the time-dimension compensation time shift into the width of a spatial-dimension visualization bar, making the structural synchronization compensation amount in the conference room directly visible in the video overlay. In traditional audio-visual synchronization systems, the compensation process is usually performed in the background, and users cannot intuitively understand the degree and source of the compensation. This embodiment, by converting the compensation time shift into the width of a visualization bar, allows users to intuitively see the amount of compensation being applied to each speaking area. When the compensation time shift is large, the visualization bar widens; when the compensation time shift is small, the visualization bar narrows. This visualization method helps users quickly identify the source of synchronization problems, such as a speaking area near a glass wall having a large compensation amount, thus allowing them to take appropriate adjustment measures.

[0124] In one embodiment of the present invention, the specific calculation process of step S7 includes:

[0125] Based on the target audio package, audio playback delay, video waiting delay, presentation time, screen pixel coordinates, and bar pixel width, the generative formula for generating the joint control command is as follows:

[0126]

[0127] in, Indicates the first The joint control instructions corresponding to each video frame, and the audio control instructions are written to the fields of the target audio packet index, presentation time, and audio playback delay. Indicates the first The target audio packet index corresponding to each video frame. Indicates the first The presentation time of each video frame and the target audio packet Indicates the first Each video frame corresponds to the audio playback delay of the target audio packet. The video control indicates that the fields for video frame sequence number, presentation time, and video wait delay are written. Indicates the video frame number. Indicates the first The video latency for each video frame is controlled by a bar that indicates the field containing the coordinates of the two endpoints of the bar. Indicates the first The screen x-coordinate corresponding to each video frame Indicates the first The screen vertical coordinate corresponding to each video frame Indicates the first The width of the bar corresponding to each video frame (in pixels). This indicates the horizontal coordinate of the left end of the graph. This indicates the horizontal coordinate at the right end of the graph bar.

[0128] It should be noted that the audio control field contains the target audio package index, rendering time, and audio playback delay, and is used to control the playback operation of the audio device. The video control field contains the video frame number, rendering time, and video wait delay, and is used to control the display operation of the video device. The bar control field contains the coordinates of the two endpoints of the synchronized visualization bar, and is used to control the rendering operation of the visualization interface. The combined control command is a structured control command that integrates audio, video, and visualization control information, and is the final control signal output by the controller.

[0129] It should be noted that the joint control commands are encoded in JSON format and sent to the audio / video playback controller and the visualization interface renderer via the TCP protocol. Each time the controller processes a video frame, it immediately generates and sends the corresponding joint control command. Upon receiving the joint control command, the audio / video playback controller and the visualization interface renderer parse the corresponding control fields and execute the appropriate operations. Based on the audio and video control fields, the audio / video playback controller plays the corresponding audio packet and displays the corresponding video frame at the specified presentation time. Based on the bar control field, the visualization interface renderer draws a synchronized visual bar at a specified position on the video screen.

[0130] The synchronization visualization bar is drawn on the top layer of the video image. Its coordinates are determined by the bar control field, calculated from the screen pixel coordinates and the bar's pixel width. When it's necessary to avoid the speaker's face area, a preset display offset can be applied to the vertical coordinates in the bar control field without changing the synchronization compensation calculation logic; for example, offsetting downwards by approximately 20 pixels. The color of the synchronization visualization bar can be set to a color with high contrast to the video image (referring to color combinations with a brightness difference greater than 120, such as using a blue with a brightness of 50 on a white background with a brightness of 180, or using a green with a brightness of 200 on a dark background with a brightness of 60), such as green or blue, to ensure clear visibility under various backgrounds. The height of the synchronization visualization bar is fixed at 2 pixels, while its length is dynamically adjusted based on the calculated bar pixel width.

[0131] It should be further explained that the embodiment corresponding to step S7 integrates the three independent control objectives of audio playback, video display, and visual overlay into a single structured instruction, ensuring that all three are executed at the same presentation moment. Traditional audio and video control systems typically separate audio control, video control, and visual control, sending different control instructions separately. This approach can easily lead to differences in transmission delays between different control instructions, resulting in asynchrony between audio, video, and visual overlay. This embodiment integrates the three control objectives into a single joint control instruction, ensuring that the three control operations are executed at the same presentation moment, avoiding synchronization errors caused by separate control.

[0132] In one embodiment of the present invention, a conference room audio-video synchronization and visualization control system includes:

[0133] The acquisition time conversion module converts the audio timestamp and video timestamp into audio acquisition time and video acquisition time respectively, based on the reference timestamp and clock frequency.

[0134] The time-shift compensation module extracts the mirror source coordinates based on the coordinates of the sound-emitting area, the boundary normal vector, and the boundary point coordinates. It calculates the direct sound path and the equivalent sound path based on the array coordinates, the coordinates of the sound-emitting area, and the coordinates of the mirror source, respectively. Based on this, it calculates the acoustic time difference and calculates the time-shift compensation by combining the acoustic time difference, the video frame period, and the optical axis angular velocity.

[0135] The target audio packet extraction module subtracts the compensation time shift from the audio acquisition time to obtain the corrected audio time, and extracts the target audio packet based on the principle of minimizing the difference between the video acquisition time and the corrected audio time. The arrival time difference is obtained by subtracting the arrival time of the audio corresponding to the target audio packet from the arrival time of the video.

[0136] The audio and video latency calculation module calculates audio playback latency and video waiting latency based on the arrival time difference and display refresh cycle.

[0137] The presentation time calculation module calculates the presentation time based on the display refresh cycle, the video arrival time, and the video waiting delay.

[0138] The pixel width calculation module uses a mapping matrix to convert the sound area coordinates into screen pixel coordinates, and combines compensation time shift, video frame period, basic pixel width and pixel width increment to calculate the pixel width of the bar.

[0139] The control instruction generation module combines the target audio package, audio playback delay, video waiting delay, presentation time, screen pixel coordinates, and bar pixel width to generate a joint control instruction.

[0140] Specifically, this invention is applicable to medium to large conference rooms equipped with multiple PTZ cameras, ceiling array microphones, and large-screen display systems. During deployment, the equipment installation and system cabling are completed first, connecting the cameras, microphones, display devices, and controllers to the same local area network. Then, system calibration is performed, including calibrating the array center coordinates of the ceiling array microphones, the boundary normal vectors and boundary point coordinates of each strongly reflective plane in the conference room, the intrinsic and extrinsic parameters of the cameras, and the mapping matrix from the conference room plane to the screen pixel coordinates. After calibration, system parameters are configured, including audio clock frequency, video clock frequency, base pixel width, and pixel width increment.

[0141] During system operation, the controller receives video streams from cameras and audio streams from ceiling array microphones in real time. For each video frame, the controller first establishes a unified wall clock timeline, calculates the mirror sound path video frame raster time shift, then performs audio-video event association and delay allocation, generates unified presentation times and synchronized visualization bar parameters, and finally outputs joint control commands. The audio-video playback controller and the visualization interface renderer, based on the joint control commands, play audio, display video, and draw synchronized visualization bars at the specified times.

[0142] For example, in a conference room measuring 10 meters long, 8 meters wide, and 3 meters high, when a speaker near the glass wall speaks, if the compensated time shift calculated based on the calibrated direct path, equivalent path, air velocity, and frame raster coefficient is approximately 15 milliseconds, and the video frame period is 33.33 milliseconds, then the frame-to-time ratio is approximately 0.45. With a base pixel width of 10 pixels and a pixel width increment of 50 pixels per frame period, the width of the synchronization visualization bar is 32.5 pixels. The final output joint control command will instruct the playback of the corresponding audio packet, the display of the corresponding video frame, and the drawing of a 32.5-pixel-long green synchronization visualization bar below the speaker's face at the presentation moment.

[0143] It should be noted that the interval and threshold sizes are set for ease of comparison. The size of the threshold depends on the amount of sample data and the base number set by those skilled in the art for each set of sample data, as long as it does not affect the proportional relationship between the parameter and the quantized value. Furthermore, the above formulas are all dimensionless calculations, and the formulas are derived from software simulations using a large amount of collected data to obtain a formula that is closest to the real situation. The preset parameters in the formulas are set by those skilled in the art according to the actual situation.

[0144] The content of this embodiment has been described above, but this embodiment is not limited to the specific implementation methods described above. The specific implementation methods described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of this embodiment, all of which are within the protection scope of this embodiment.

Claims

1. A method for synchronized and visualized control of audio and video in a conference room, characterized in that, Includes the following steps: Step S1: Convert the audio timestamp and video timestamp into audio capture time and video capture time respectively based on the reference timestamp and clock frequency; Step S2: Extract the mirror source coordinates based on the coordinates of the sound-emitting area, the boundary normal vector, and the coordinates of the boundary points. Calculate the direct sound path and the equivalent sound path based on the array coordinates, combined with the coordinates of the sound-emitting area and the mirror source coordinates. Calculate the acoustic time difference accordingly. Combine the acoustic time difference, the video frame period, and the optical axis angular velocity to calculate the compensation time shift. Step S3: Subtract the compensation time shift from the audio acquisition time to obtain the corrected audio time, and extract the target audio packet based on the principle of minimizing the difference between the video acquisition time and the corrected audio time. Subtract the audio arrival time corresponding to the target audio packet from the video arrival time to obtain the arrival time difference. Step S4: Calculate the audio playback delay and video waiting delay based on the arrival time difference and display refresh cycle; Step S5: Calculate the presentation time based on the display refresh cycle, the video arrival time, and the video waiting delay; Step S6: Use the mapping matrix to convert the sound area coordinates into screen pixel coordinates, and combine the compensation time shift, video frame period, basic pixel width and pixel width increment to calculate the strip pixel width; Step S7: Combine the target audio package, audio playback delay, video waiting delay, presentation time, screen pixel coordinates, and bar pixel width to generate a joint control command.

2. The method for synchronized and visualized audio and video control in a conference room according to claim 1, characterized in that, Step S1 specifically includes the following steps: Step S101: Read the audio reference time, video reference time, audio reference stamp, video reference stamp, audio clock frequency, and video clock frequency; Step S102: Subtract the audio timestamp from the audio reference timestamp to obtain the audio timestamp difference; divide the audio timestamp difference by the audio clock frequency to obtain the audio time difference; add the audio time difference to the audio reference time to obtain the audio acquisition time. Step S103: Subtract the video timestamp from the video reference timestamp to obtain the video timestamp difference; divide the video timestamp difference by the video clock frequency to obtain the video time difference; add the video time difference to the video reference time to obtain the video acquisition time.

3. The method for synchronized and visualized audio and video control in a conference room according to claim 1, characterized in that, Step S2 specifically includes the following steps: Step S201: Subtract the coordinates of the sound-emitting area from the coordinates of the boundary point to obtain the boundary offset; take the inner product of the boundary normal vector and the boundary offset to obtain the boundary inner product; multiply the boundary normal vector, the boundary inner product, and the numerical value to obtain the mirror offset; subtract the mirror offset from the coordinates of the sound-emitting area to obtain the coordinates of the mirror source. Step S202: The Euclidean distance between the sound-emitting area coordinates and the array coordinates is taken as the direct sound path, and the Euclidean distance between the mirror source coordinates and the array coordinates is taken as the equivalent sound path. Step S203: Subtract the direct sound path from the equivalent sound path to obtain the sound path difference, and divide the sound path difference by the sound speed to obtain the acoustic time difference; Step S204: Take the inner product of the current optical axis vector and the optical axis vector of the previous frame, take the inverse cosine of the inner product value, and then divide it by the video frame period to obtain the optical axis angular velocity. Step S205: Multiply the optical axis angular velocity by the video frame period to obtain the angle product, take the cosine of the angle product, subtract the cosine value from the first value and divide it by the second value, and then add it to the first value to obtain the frame raster coefficient. Step S206: Multiply the acoustic time difference by the frame grating coefficient to obtain the compensated time shift.

4. The method for synchronized and visualized audio and video control in a conference room according to claim 1, characterized in that, Step S3 specifically includes the following steps: Step S301: For the candidate audio packet, subtract the compensation time shift from the audio acquisition time of the candidate audio packet to obtain the corrected audio time. Step S302: Subtract the corrected audio time from the video acquisition time to obtain the acquisition time difference, and take the absolute value of the acquisition time difference; Step S303: Select the candidate audio packet with the smallest absolute value from the candidate audio packets as the target audio packet; Step S304: Subtract the arrival time of the target audio packet from the arrival time of the video to obtain the arrival time difference.

5. The method for synchronized and visualized audio and video control in a conference room according to claim 1, characterized in that, Step S4 specifically includes the following steps: Step S401: Add the square of the arrival time difference to the square of the display refresh cycle to obtain the sum of squares, and take the square root of the sum of squares to obtain the root value; Step S402: Add the arrival time difference to the square root value to obtain the delay sum value, and divide the delay sum value by the numerical value two to obtain the reference delay; Step S403: Use the reference delay as the audio playback delay; Step S404: Subtract the arrival time difference from the baseline delay to obtain the video waiting delay.

6. The method for synchronized and visualized audio and video control in a conference room according to claim 1, characterized in that, Step S5 specifically includes the following steps: Step S501: Add the video arrival time to the video waiting delay to obtain the display base time; Step S502: Divide the display base time by the display refresh cycle to obtain the refresh quotient; Step S503: Round up the refresh quotient to obtain the refresh sequence number; Step S504: Multiply the refresh sequence number by the display refresh cycle to obtain the presentation time.

7. The method for synchronized and visualized audio and video control in a conference room according to claim 1, characterized in that, Step S6 specifically includes the following steps: Step S601: Combine the plane abscissa, plane ordinate, and constant in the sound-emitting area coordinates to form homogeneous coordinates; Step S602: Multiply the mapping matrix by the homogeneous coordinates to obtain the screen homogeneous coordinates; Step S603: Divide the horizontal component in the screen homogeneous coordinates by the normalized component to obtain the screen horizontal coordinate, divide the vertical component in the screen homogeneous coordinates by the normalized component to obtain the screen vertical coordinate, and combine the screen horizontal coordinate and the screen vertical coordinate to form the screen pixel coordinate. Step S604: Remove the frame time ratio obtained by the video frame period during compensation, multiply the frame time ratio by the pixel width increment to obtain the pixel increment, and add the pixel increment to the base pixel width to obtain the bar pixel width.

8. The method for synchronized and visualized audio and video control in a conference room according to claim 1, characterized in that, Step S7 specifically includes the following steps: Step S701: Write the target audio package, presentation time, and audio playback delay into the audio control field; Step S702: Write the video frame number, presentation time, and video waiting delay into the video control field; Step S703: Write the screen pixel coordinates and the bar pixel width into the bar control field. Subtract half of the bar pixel width from the screen horizontal coordinate to obtain the left horizontal coordinate. Add half of the bar pixel width to the screen horizontal coordinate to obtain the right horizontal coordinate. Use the screen vertical coordinate as the left and right vertical coordinates. Step S704: Merge the audio control field, video control field, and graphic bar control field to generate a joint control instruction.

9. A conference room audio-video synchronized visual control system, characterized in that, The method for synchronous visual control of audio and video in a conference room as described in any one of claims 1 to 8 includes: The acquisition time conversion module converts the audio timestamp and video timestamp into audio acquisition time and video acquisition time respectively, based on the reference timestamp and clock frequency. The time-shift compensation module extracts the mirror source coordinates based on the coordinates of the sound-emitting area, the boundary normal vector, and the boundary point coordinates. It calculates the direct sound path and the equivalent sound path based on the array coordinates, the coordinates of the sound-emitting area, and the coordinates of the mirror source, respectively. Based on this, it calculates the acoustic time difference and calculates the time-shift compensation by combining the acoustic time difference, the video frame period, and the optical axis angular velocity. The target audio packet extraction module subtracts the compensation time shift from the audio acquisition time to obtain the corrected audio time, and extracts the target audio packet based on the principle of minimizing the difference between the video acquisition time and the corrected audio time. The arrival time difference is obtained by subtracting the arrival time of the audio corresponding to the target audio packet from the arrival time of the video. The audio and video latency calculation module calculates audio playback latency and video waiting latency based on the arrival time difference and display refresh cycle. The presentation time calculation module calculates the presentation time based on the display refresh cycle, the video arrival time, and the video waiting delay. The pixel width calculation module uses a mapping matrix to convert the sound area coordinates into screen pixel coordinates, and combines compensation time shift, video frame period, basic pixel width and pixel width increment to calculate the pixel width of the bar. The control instruction generation module combines the target audio package, audio playback delay, video waiting delay, presentation time, screen pixel coordinates, and bar pixel width to generate a joint control instruction.