Audio and video synchronous control system for multimedia movie hall
By constructing a three-dimensional coordinate system and extracting the coordinates and motion vectors of three-dimensional sound sources, and combining mapping units and calibration matrices, the problems of insufficient three-dimensional spatial sense, sound source motion accuracy and video synchronization precision in traditional audio and video synchronization systems are solved, achieving higher quality audio and video synchronization control and enhancing the audience's immersive experience.
Patent Information
- Application Number
- CN202510917730.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-03
- Publication Date
- 2025-11-04
AI Technical Summary
Traditional multimedia cinema audio-visual synchronization systems are deficient in terms of three-dimensional spatial perception and sound source motion accuracy, audio signal mapping precision, video synchronization accuracy, and projection system redundancy, resulting in a poor audience experience.
A three-dimensional coordinate system is constructed, and the coordinates and motion vectors of the three-dimensional sound source are extracted. Accurate sound source localization and dynamic tracking are achieved through mapping units and calibration matrices. Video signals are synchronously output to the main and backup projectors. Combined with mechanical vibration compensation and frequency offset correction, audio and video synchronization is ensured.
It achieves precise sound source localization and dynamic tracking, enhances the spatial sense and immersion of the audio system, reduces the risk of video interruption, and improves the stability of video transmission and the immersive experience for the audience.
Smart Images

Figure CN120897080A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of audio and video transmission, in particular to a multimedia theater audio and video synchronization control system. BACKGROUND
[0002] In the development process of multimedia theaters, audio and video synchronization control technology has always been one of the key factors affecting the audience's immersive experience. The audio and video synchronization systems of traditional theaters are mostly based on early two-dimensional audio formats and relatively simple video synchronization mechanisms.
[0003] However, with the growing demand for high-quality audio-visual experience, the shortcomings of traditional audio and video synchronization control technology have gradually become apparent. On the one hand, in terms of audio processing, two-dimensional audio formats cannot meet the requirements of modern theaters for three-dimensional spatial sense and accurate sound source movement. The construction of sound field lacks depth and stereoscopic effect, making it difficult to achieve precise sound source positioning and dynamic tracking. On the other hand, in terms of video synchronization, traditional systems fail to fully consider the physical characteristics of film projection windows and the impact of mechanical vibrations, resulting in insufficient synchronization accuracy between video frames and audio signals, and easy occurrence of phenomena such as out-of-sync between picture and sound, which is particularly evident in fast action scenes or special effects scenes requiring high precision synchronization.
[0004] In addition, the audio system of traditional theaters also has obvious shortcomings in sound field optimization. Due to the lack of effective means to counteract reflected sound and individualized compensation mechanisms for seat sound fields, the audio effects received by audiences at different positions differ greatly, making it difficult to achieve uniform and consistent immersive experience. At the same time, existing technologies lack real-time and precision in handling the Doppler effect caused by sound source movement, resulting in insufficient compensation for audio frequency shifts and affecting the realism and naturalness of sound source movement.
[0005] In terms of projection systems, traditional theaters usually do not have effective redundancy mechanisms. Once the main projector fails, the switching process is often accompanied by obvious picture interruption and audio distortion, seriously affecting the audience's viewing experience. SUMMARY
[0006] To overcome the shortcomings of the prior art, the present application provides a multimedia theater audio and video synchronization control system.
[0007] A multimedia theater audio and video synchronization control system, comprising:
[0008] The synchronization control module constructs a three-dimensional coordinate system in the multimedia theater, with the center of the multimedia theater as the origin, the horizontal direction as the x-axis, the vertical direction as the y-axis, and the vertical direction as the z-axis.
[0009] An audio processing module includes a sound image coordinate extraction unit that receives audio and video object data in a multimedia theater and extracts three-dimensional sound source coordinates (x, y, z) and motion vectors (vx, vy, vz) in a three-dimensional coordinate system;
[0010] It should be noted that the specific steps are as follows: the sound image coordinate extractor receives input audio object metadata, which is packaged in Dolby Atmos / Auro 3D format and contains various information of the audio object in three-dimensional space. It is the original data source for extracting three-dimensional sound source coordinates and motion vectors.
[0011] The Dolby Atmos / Auro 3D format audio object metadata is deeply analyzed, and the three-dimensional sound source coordinates (x, y, z) are extracted. Here, x, y, and z represent the position coordinates of the sound source in different directions in three-dimensional space, which are used to determine the specific position of the audio object in space, so that the audio system can accurately position the audio according to these coordinate information.
[0012] In addition to the three-dimensional sound source coordinates, the metadata is further analyzed to extract the motion vectors (vx, vy, vz) related to the motion of the audio object. The three motion vectors represent the motion velocity components of the sound source in the x, y, and z axis directions. Through analysis and processing of the motion vectors,
[0013] It also includes a mapping unit that stores the mapping relationship between the screen pixel coordinates in the theater and the surround sound array, and converts the sound source coordinates (x, y, z) into a target sound number set through a calibration matrix.
[0014] According to the target sound number set, a driving signal is generated to form a directional sound beam in the theater space.
[0015] A video processing module includes a projection redundancy unit that synchronously outputs two video signals to the main / backup projector.
[0016] By constructing a three-dimensional coordinate system and extracting three-dimensional sound source coordinates and motion vectors, the problem of traditional two-dimensional audio format not meeting the requirements of modern theaters for three-dimensional space and sound source motion accuracy is solved. Precise sound source positioning and dynamic tracking are achieved, and the audio system has stronger spatial sense and immersion;
[0017] The mapping unit and the use of the calibration matrix effectively solve the problem of inaccurate mapping relationship between audio signals and specific sound equipment in traditional systems, and the sound field distribution effect is not good. The sound source coordinates can be more accurately converted into a target sound number set, and a more ideal three-dimensional sound field can be constructed.
[0018] The video processing module synchronously outputs two video signals to the main / backup projectors, compared with the traditional projection system, the risk of picture interruption caused by single point failure is greatly reduced, the stability and reliability of video transmission are improved, and the continuity of video display is ensured.
[0019] Preferably, the specific working steps of the mapping unit are as follows:
[0020] Nine calibration points are set on the screen of the theater, specifically, four corner points, a center point and edge middle points of the screen;
[0021] According to the ultrasonic transmitter and the sound array microphone at each calibration point, the sound wave transmission time difference Δt1-Δt n is measured.
[0022] Through the sound speed V, and by using the measured Δt1-Δt n , the sound source coordinates (xs, ys, zs) corresponding to each calibration point are calculated, the coordinates of each calibration point are obtained, and the Euclidean distance between the current sound source coordinates and the coordinates is calculated.
[0023] The calculated Euclidean distance is compared with the recorded minimum distance, if the Euclidean distance is smaller, the minimum distance variable is updated, and the sound group number set corresponding to the calibration point is recorded, after traversing all the calibration points, the sound group number set corresponding to the calibration point closest to the sound source coordinates is found.
[0024] According to the sound group number set found in the above matching process, the target sound number set corresponding to the current sound source coordinates is directly determined.
[0025] Preferably, the mapping unit generates a driving signal according to the target sound number set, so that the separated sound source forms a directional sound beam in the theater space, and the specific working steps are as follows:
[0026] The sound source coordinates (x, y, z) and the target sound number set are obtained, and the time delay that each sound needs to add is calculated according to the sound source coordinates (x, y, z) and the target sound number set, the unit is second;
[0027] Then the contribution weight W k of each sound to the sound beam is calculated, the driving signal is generated according to the contribution weight W k and the time delay, so that the separated sound source forms a directional sound beam in the theater space.
[0028] Preferably, the audio processing module further comprises a correction unit, which calculates a frequency offset in real time according to the sound source motion vector (vx, vy, vz), and performs frequency compensation on each sound and the initial audio stream of video transmission, and the specific steps are as follows:
[0029] The three-dimensional coordinates of each seat in the multimedia theater are obtained, and then the vector (dx, dy, dz) from the sound source to the seat is obtained, and the relative speed V from the sound source to the seat is calculated rel ;
[0030] The frequency offset Δf is calculated according to the formula , wherein f0 is the original frequency, which is the frequency in the audio stream that is not affected by the Doppler effect. According to the frequency offset, the initial audio stream of each sound and video transmission is frequency compensated.
[0031] Preferably, the audio processing module further comprises a compensation unit, which generates an anti-phase sound wave to cancel the non-direct sound when the beamformer generates a directional sound beam to direct the sound source sound to the seat area. The specific steps are as follows:
[0032] Based on the HRTF function library of each seat, the direct sound is filtered;
[0033] The direct sound filtered by the HRTF is played through the seat headrest speaker;
[0034] At each seat in the theater, a measuring device is used to obtain the impulse response R(t) of the early reflection sound;
[0035] A microphone is placed at each seat in the theater, a specific sound source signal is played, and the reflected sound signal is recorded to extract the impulse response;
[0036] The measured impulse response R(t) is convolved with the original audio signal, and the anti-phase sound wave C(t) is generated by taking the inverse. The generated anti-phase sound wave C(t) is emitted through the loudspeaker array.
[0037] Preferably, it further comprises a synchronization control module, which is used to establish a video frame position model based on the physical characteristics of the film projection window, and outputs a frame position signal (Fn, Xn, Yn). The specific steps further comprise the following:
[0038] According to the film movement speed, the time when the center of the gth frame passes through the projection window is calculated;
[0039] Align the center of the digital video frame with the calculated time when it passes through the projection window, and ensure that the alignment error is less than the preset frame period.
[0040] Preferably, the synchronization control module further comprises a mechanical vibration compensation unit, and the specific steps are as follows:
[0041] A three-axis MEMS accelerometer is installed on the projector to detect the vibration acceleration information of the projector in three axes in real time;
[0042] According to the detected vibration acceleration information, a vibration displacement model is established, and displacement of the projector in horizontal direction δx(t) is calculated;
[0043] The vibration acceleration information is fused with vibration information and frame position prediction information by using Kalman filtering algorithm, a dynamic adjustment of the frame position prediction model is performed, a time offset ΔT(a) is introduced on the basis of the original frame position prediction model, and an adjusted prediction model is obtained.
[0044] According to the calculation results of the adjusted frame position prediction model and the vibration displacement model, corresponding correction signals are generated, and the signals are output to the projector.
[0045] Preferably, the specific steps of the video processing module are as follows:
[0046] The screen brightness is monitored by the light sensor, and when the screen brightness is monitored to drop by more than a set threshold, the switching mechanism is triggered.
[0047] During the working process of the main projector, the standby projector preloads the next frame of content, and switches the signal source during the video vertical blanking period.
[0048] The audio clock source is locked during the switching period, and the audio buffer area is inhibited from refreshing during the switching period.
[0049] Preferably, the specific working steps of the video processing module further include the following:
[0050] The lateral deviation Δx and the longitudinal deviation Δz are determined by comparing the sound source coordinates (x, y, z) and the video object screen coordinates (Xm, Ym):
[0051] When the absolute value |Δx| of the lateral deviation exceeds a set value of the screen width, the mapping unit is driven to adjust the sound beam pointing direction to compensate for the lateral deviation.
[0052] Longitudinal correction: when the absolute value |Δz| of the longitudinal deviation exceeds a set value, the mapping unit is driven to adjust the sound beam pointing direction to compensate for the lateral deviation.
[0053] The present application provides a multimedia theater audio and video synchronization control system. It has the following beneficial effects:
[0054] 1. By constructing a three-dimensional coordinate system and extracting three-dimensional sound source coordinates and their motion vectors, the problem that the traditional two-dimensional audio format cannot meet the requirements of modern theaters for three-dimensional spatial sense and sound source motion accuracy is solved, accurate sound source positioning and dynamic tracking are achieved, the audio system has stronger spatial sense and immersion, the setting of the mapping unit and the use of the calibration matrix effectively solve the problems of insufficient accuracy of the mapping relationship between the audio signal and the specific sound equipment in the traditional system and poor sound field distribution effect. BRIEF DESCRIPTION OF DRAWINGS
[0055] Figure 1 The flow chart is for the present application. DETAILED DESCRIPTION
[0056] The technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative work fall within the scope of protection of the present application.
[0057] As Figure 1 shown, the present application provides a multimedia theater audio and video synchronization control system, comprising:
[0058] A synchronization control module constructs a three-dimensional coordinate system in the multimedia theater, taking the center of the multimedia theater as the origin, the horizontal direction as the x-axis, the vertical direction as the y-axis, and the vertical direction as the z-axis.
[0059] An audio processing module, comprising a sound image coordinate extraction unit, receives audio and video object data in the multimedia theater, and extracts three-dimensional sound source coordinates (x, y, z) and their motion vectors (vx, vy, vz) in the three-dimensional coordinate system.
[0060] It should be noted that the specific steps are as follows: the sound image coordinate extractor receives the input audio object metadata, which is packaged in Dolby Atmos / Auro 3D format and contains various information of the audio object in three-dimensional space. It is the original data source for extracting three-dimensional sound source coordinates and their motion vectors.
[0061] The Dolby Atmos / Auro 3D format audio object metadata is deeply analyzed, and the three-dimensional sound source coordinates (x, y, z) are extracted. Here, x, y, and z represent the position coordinates of the sound source in different directions in three-dimensional space, which are used to determine the specific position of the audio object in space, so that the audio system can accurately position the audio according to these coordinate information.
[0062] In addition to the three-dimensional sound source coordinates, the metadata is further analyzed to extract the motion vectors (vx, vy, vz) related to the motion of the audio object. The three motion vectors represent the motion velocity components of the sound source in the x, y, and z axis directions. Through analysis and processing of the motion vectors,
[0063] It also includes a mapping unit that stores the mapping relationship between the screen pixel coordinates in the theater and the surround sound array, and converts the sound source coordinates (x, y, z) into a target sound number set through a calibration matrix.
[0064] According to the target sound number set, a driving signal is generated to make the separated sound source form a directional sound beam in the hall space.
[0065] The video processing module comprises a projection redundancy unit, which synchronously outputs two video signals to the main and backup projectors.
[0066] By constructing a three-dimensional coordinate system and extracting the three-dimensional sound source coordinates and their motion vectors, the problem that the traditional two-dimensional audio format cannot meet the requirements of modern theaters for three-dimensional spatial sense and sound source motion accuracy is solved, precise sound source positioning and dynamic tracking are achieved, and the audio system has stronger spatial sense and immersion;
[0067] The mapping unit and the use of the calibration matrix effectively solve the problem of inaccurate mapping relationship between audio signals and specific sound equipment and poor sound field distribution effect in traditional systems, and can more accurately convert sound source coordinates into a target sound number set, thereby achieving more ideal three-dimensional sound field construction.
[0068] The video processing module synchronously outputs two video signals to the main and backup projectors, which greatly reduces the risk of picture interruption caused by single point failure compared with traditional projection systems, improves the stability and reliability of video transmission, and ensures the continuity of video display.
[0069] As an optional embodiment, the specific working steps of the mapping unit are as follows:
[0070] Nine calibration points are set on the screen of the theater, specifically the four corner points, the center point and the edge midpoint of the screen.
[0071] According to the ultrasonic transmitters and the sound array microphones at each calibration point, the sound wave transmission time difference Δt1-Δt n is measured.
[0072] Through the sound speed V, and the measured Δt1-Δt n , the sound source coordinates (xs, ys, zs) corresponding to each calibration point are calculated, the coordinates of each calibration point are obtained, and the Euclidean distance between the current sound source coordinates is calculated.
[0073] It should be noted that the specific calculation formula is:
[0074]
[0075] Where d1 represents the position coordinates of the first microphone, d n n represents the position coordinates of the nth microphone. These position coordinates of the microphones, together with the measured sound wave transmission time difference and the sound speed, are used to solve the coordinates of the sound source.
[0076] Compare the calculated Euclidean distance with the previously recorded minimum distance. If the Euclidean distance is smaller, update the minimum distance variable and record the corresponding sound group number set of the calibration point. After traversing all calibration points, find the sound group number set corresponding to the calibration point closest to the sound source coordinates.
[0077] According to the sound group number set found by the above matching process, directly determine it as the target sound number set corresponding to the current sound source coordinates.
[0078] It should be noted that, for example, if the sound group number set corresponding to the calibration point closest to the sound source coordinates in the mapping lookup table is [1, 3, 5], then the target sound number set is [1, 3, 5];
[0079] It should also be noted that in this embodiment, the target sound number set can be adjusted appropriately according to the direction and size of the motion vector, for example, if the motion vector is larger in the x-axis direction and towards a certain side, the number of the sound on that side can be appropriately increased in the set or the number of adjacent sound can be added to enhance the spatial transition effect during the movement of the audio object;
[0080] Output the finally determined target sound number set in the specified data format for subsequent audio signal allocation and playback control, for example, the number set can be encapsulated into a data packet and sent to the signal allocation module in the audio playback system through a specific communication protocol, so that the audio signal can be correctly allocated to the corresponding numbered sound equipment for playback;
[0081] It should be noted that the scheme of using nine calibration points combined with ultrasonic transmitters and sound array microphones to measure the time difference of sound wave transmission, and then calculating the sound source coordinates using the speed of sound, effectively solves the problem of inaccurate mapping of sound and sound source positions in traditional audio systems, and can more accurately determine the position of the sound source, thereby realizing more accurate sound positioning and sound field distribution, and improving the accuracy and spatial sense of the audio;
[0082] The introduction of Euclidean distance comparison and minimum distance updating mechanism optimizes the matching process of sound source coordinates and sound group number set from the algorithm level, avoids the sound positioning deviation or uneven sound field distribution that may occur in traditional systems, ensures that the audio signal can be correctly allocated to the optimal sound combination, and enhances the consistency and stability of the audio effect;
[0083] The strategy of adjusting the target sound number set according to the direction and size of the motion vector further compensates for the defects of traditional systems in handling the spatial transition effect of audio object movement, making the spatial sense of audio more coherent and natural during movement, and improving the auditory experience of the audience.
[0084] As an optional embodiment: the mapping unit generates the driving signal according to the target sound number set, so that the separated sound sources form directional sound beams in the theater space. The specific working steps are as follows:
[0085] The sound source coordinates (x, y, z) and the target sound number set are obtained, and the time delay that each sound needs to add is calculated according to the sound source coordinates (x, y, z) and the target sound number set, which is in seconds;
[0086] The specific calculation formula is:
[0087]
[0088] Where (x k ,y k ,z k ) is the coordinate of the kth sound, which is obtained through the target sound number set;
[0089] T k is the time delay of the kth sound;
[0090] Then the contribution weight W k of each sound to the sound beam is calculated, and the driving signal is generated according to the contribution weight W k and the time delay, so that the separated sound sources form directional sound beams in the theater space.
[0091] It should be noted that in this embodiment, the specific calculation formula is:
[0092]
[0093] Where θ k is the azimuth angle of the sound relative to the sound source, and θ0 is the best radiation angle;
[0094] It should be noted that specifically, a corresponding time delay is applied to each sound, and then the sound wave amplitude of each sound is adjusted according to the contribution weight W k . Specifically, the main lobe is enhanced: when the azimuth angle of the sound is close to the best radiation angle, the weight is larger, the sound wave amplitude in that direction is enhanced, the main lobe of the sound beam is more concentrated, the energy is stronger, the directivity and transmission efficiency are improved;
[0095] The side lobe is suppressed: when the azimuth angle of the sound deviates greatly from the best radiation angle, the weight decreases, the sound wave amplitude in that direction decreases, the side lobe is suppressed, the interference is reduced, and the purity of the sound beam is improved;
[0096] By calculating the time delay that each sound needs to add and generating the driving signal combined with the contribution weight, the problem of inaccurate sound beam formation and weak directivity in traditional audio systems is solved, so that the sound beam can be more accurately pointed to the target area, and the concentration and clarity of the audio are improved;
[0097] The strategy of enhancing the main lobe and suppressing the side lobe effectively overcomes the problems of energy dispersion and large interference of the sound beam in the traditional system, makes the main lobe of the sound beam more concentrated and the energy stronger, reduces the interference caused by the side lobe, improves the purity and quality of the sound beam, and thus improves the auditory experience of the audience.
[0098] As an optional embodiment, the audio processing module further comprises a correction unit, which calculates a frequency offset in real time according to the sound source motion vector (vx, vy, vz), and compensates the frequency of the initial audio stream transmitted by each sound and video for each sound and video, and the specific steps are as follows: It should be noted that the sound source motion vector will cause the Doppler effect of the sound wave frequency. The dynamic Doppler corrector calculates the frequency offset in real time according to the sound source motion vector, which provides a basis for subsequent audio stream frequency shift compensation.
[0099] The three-dimensional coordinates of each seat in the multimedia hall are obtained, and then the vector (dx, dy, dz) from the sound source to the seat is obtained, and the relative speed V of the sound source to the seat is calculated. rel
[0100] It should be noted that in this embodiment, the specific calculation formula is
[0101] According to the formula The frequency offset Δf is calculated and obtained, where f0 is the original frequency, which is the frequency of the audio stream without the influence of the Doppler effect. According to the frequency offset, the frequency of the initial audio stream transmitted by each sound and video is compensated.
[0102] It should also be noted that the audio stream contains the original sound information of the audio object and is the signal source of the directional sound beam. The multi-channel beamformer processes the audio stream according to the target sound number set to generate a driving signal and then forms a directional sound beam.
[0103] The frequency shift compensation of the dynamic Doppler corrector ensures the frequency accuracy of the audio stream, so that the frequency characteristics of the directional sound beam match the audio content, improve the accuracy of the sound source position and motion, and enhance the spatial sense and immersion of the audio.
[0104] By considering the sound source motion vector and calculating the frequency offset in real time, the frequency of the audio stream is compensated, which solves the problems of insufficient Doppler effect processing and insufficient audio frequency offset compensation in the traditional audio system, ensures the frequency accuracy of the audio stream, makes the expression of the sound source position and motion more accurate, enhances the spatial sense and immersion of the audio, and enables the audience to obtain a more real and natural auditory experience.
[0105] As an optional embodiment: the audio processing module further comprises a compensation unit, which generates anti-phase sound waves to offset non-direct sound when the beamformer generates directional sound beams to direct sound source sound to the seat area, and the specific steps are as follows:
[0106] Based on the HRTF function library of each seat, the direct sound is filtered; the HRTF function describes the frequency and phase changes of sound waves propagating from the sound source to the human ear, and can give the sound direction and spatial sense, specifically, according to the direction and position of the sound source, select the appropriate HRTF function and apply it to the direct sound signal, so that it has the spatial characteristics unique to the direction and position;
[0107] The direct sound filtered by the HRTF is played through the seat headrest speaker; the headrest speaker is close to the audience's ear and can accurately deliver sound to the audience, enhancing the direction and immersion of the audio, making the audience feel as if they are in the sound environment;
[0108] At each seat in the theater, use a measuring device to obtain the impulse response R(t) of the early reflection sound; the impulse response describes the characteristics of the sound wave reflection in the theater, reflecting the propagation characteristics of the sound wave after reflection by walls, ceilings, etc. in this time range;
[0109] Place a microphone at each seat in the theater, play a specific sound source signal, and record the reflected sound signal to extract the impulse response; existing sound systems can also be used to play test signals and measure them;
[0110] Convolve the measured impulse response R(t) with the original audio signal, and take the negative to generate anti-phase sound waves C(t), which are emitted by the loudspeaker array.
[0111] R(t) represents the impulse response, which reflects the reflection intensity and characteristics of sound waves at different time points over time;
[0112] Generate anti-phase sound waves using convolution integral to calculate anti-phase sound waves;
[0113] Where S(t) is the original audio signal, this formula convolves the impulse response with the original audio signal and takes the negative to generate a signal opposite to the reflected sound wave.
[0114] C(t) represents the anti-phase sound wave, which matches the reflected sound wave in time and amplitude, but the phase is opposite, and is used to offset the reflected sound.
[0115] Convolve the measured impulse response R(t) with the original audio signal S(t), and take the negative to generate anti-phase sound waves C(t), which can be implemented using a digital signal processor or specialized audio processing software;
[0116] The side wall distributed speaker array is installed on the side wall of the theater, covering different directions and areas to ensure that the reflected sound at each seat can be effectively canceled by the counter-phase sound waves;
[0117] The generated counter-phase sound waves C(t) are emitted through the speaker array, and the volume and direction of the speaker are adjusted according to the seat position and the characteristics of the reflected sound to achieve the best cancellation effect;
[0118] The volume and direction of the speaker are adjusted according to the intensity and direction of the reflected sound to ensure that the counter-phase sound waves and the reflected sound are effectively canceled at the seat;
[0119] The direct sound is filtered based on the HRTF function library of each seat and played through the headrest speaker, solving the problem of poor direct sound processing effect, weak direction and space feeling in traditional systems, making the direct sound have more realistic direction and space characteristics, and enhancing the audience's sense of immersion;
[0120] The counter-phase sound wave is introduced to cancel the non-direct sound, effectively solving the problem of large reflected sound interference in traditional audio systems, and the counter-phase sound wave with opposite phase to the reflected sound is generated to cancel it, optimizing the sound field environment at the seat, so that the audience can hear the direct sound more clearly, and the purity and quality of the audio are improved.
[0121] As an optional embodiment, it further includes a synchronization control module for establishing a video frame position model based on the physical characteristics of the film projection window and outputting a frame position signal (Fn, Xn, Yn), and the specific steps further include the following:
[0122] According to the film movement speed, the time when the center of the gth frame passes through the projection window is calculated. It should be noted that in this embodiment, the specific calculation formula is:
[0123]
[0124] Where T0 is the initial time when the center of the 0th frame passes through the projection window, which can be determined by triggering a time marker signal at the beginning of projection; g is the frame number; D is the frame spacing, representing the physical distance between adjacent two frames; V p is the film movement speed, with a unit of meters per second (m / s);
[0125] Align the center of the digital video frame with the calculated time when it passes through the projection window to ensure that the alignment error is less than the preset frame period. The specific implementation is to fine-tune the sending time of the digital video signal through a digital signal processor (DSP) or a video synchronization circuit to accurately align the center time of each video frame to T g . For example, if the current video frame sending time is ahead of T g , delay sending the video frame to achieve alignment, and in this embodiment, the preset frame period is 0.5 frames;
[0126] The video frame position model is established based on the physical characteristics of the film projection window, and a frame position signal is output, which solves the problem of insufficient synchronization accuracy caused by the fact that the traditional video synchronization technology does not fully consider the physical characteristics of the film projection window, improves the accuracy of video frame position prediction, and provides a more reliable basis for accurate synchronization of video frames and audio signals.
[0127] The transmission time of the digital video signal is fine-tuned by the digital signal processor or the video synchronization circuit, so as to ensure that the center of the digital video frame is aligned with the time calculated to pass through the projection window, and the alignment error is controlled within the preset frame period, thereby effectively avoiding the phenomenon of picture and sound being out of synchronization caused by large alignment error in the traditional system, and improving the audio-video synchronization effect.
[0128] As an optional embodiment, the synchronization control module further comprises a mechanical vibration compensation unit, and the specific steps are as follows:
[0129] A three-axis MEMS accelerometer is installed on the projector to detect the vibration acceleration information of the projector in three axial directions in real time; the accelerometer has the characteristics of high precision, small size and low power consumption, and can accurately capture the slight vibration of the projector during operation.
[0130] According to the detected vibration acceleration information, a vibration displacement model is established to calculate the displacement δx(t) of the projector in the horizontal direction;
[0131] It should be noted that the specific calculation formula is
[0132] Wherein, α x (S) represents the vibration acceleration of the projector in the horizontal direction, s and τ are integral variables, and through the calculation of the model, the displacement change of the projector caused by vibration can be accurately obtained;
[0133] The vibration acceleration information is fused with the vibration information and the frame position prediction information by using the Kalman filtering algorithm, the frame position prediction model is dynamically adjusted, the time offset ΔT(a) is introduced on the basis of the original frame position prediction model, and the adjusted prediction model is obtained; so as to compensate the influence of vibration on frame position, and ensure the stable display of video frames;
[0134] The actual steps are as follows: on the basis of the original frame position prediction model , the adjusted prediction model ΔT(a) is obtained according to δx(t) divided by the running speed V p of the film;
[0135] According to the calculation results of the adjusted frame position prediction model and the vibration displacement model, corresponding correction signals are generated and output to the projector, and after receiving the correction signals, displacement compensation is performed. The display position of each pixel is accurately adjusted to offset the image jitter and displacement error caused by the vibration of the projector, thereby ensuring the clarity and stability of the projected image and improving the overall display effect and user experience.
[0136] As an optional embodiment, the specific steps of the video processing module are as follows:
[0137] The light sensor is installed near the screen to monitor the brightness change of the screen in real time, and the brightness audio and video are transmitted to the fault detection module for analysis. In this embodiment, the set threshold is 30%;
[0138] During the operation of the main projector, the standby projector preloads the next frame of content, and switches the signal source during the video vertical blanking period; ensure that the standby projector is always in standby state and can quickly take over the projection task;
[0139] Switch the signal source during the video vertical blanking period to avoid visual interference during the switching process. The video vertical blanking period refers to the interval period for synchronization in the video signal;
[0140] Lock the audio clock source during switching, and prohibit audio buffer area refresh during switching. Specifically, ensure that the switching time is controlled within 8.3 milliseconds (1 / 120 seconds) to achieve seamless switching and ensure that the audience cannot perceive the switching process. Lock the audio clock source during switching to ensure the stability of the audio signal and avoid audio distortion or interruption caused by switching. During switching, the audio buffer area refresh is prohibited to prevent audio data loss or audio interruption phenomenon, and to ensure the continuity of the audio. It should be noted that when the main projector fails, it can quickly and seamlessly switch to the standby projector, while ensuring the continuous and stable output of the audio, and ensuring that the audience can obtain uninterrupted high-quality audio-visual experience.
[0141] As an optional embodiment, the specific working steps of the video processing module further include the following:
[0142] By comparing the sound source coordinates (x, y, z) and the video object screen coordinates (Xm, Ym), the horizontal deviation Δx and the vertical deviation Δz are determined:
[0143] When the absolute value of lateral deviation |Δx| exceeds a set value of screen width, the driving mapping unit adjusts the sound beam direction to compensate for the lateral deviation; the specific operation is to control the multi-channel beamformer to change the propagation direction of the sound beam according to the deviation direction and size, so that the sound beam is directed closer to the screen coordinates of the video object; in this embodiment, the set value is 3%, and the actual value is adjusted according to the use requirement;
[0144] Longitudinal correction: when the absolute value of longitudinal deviation |Δz| exceeds a set value, the driving mapping unit adjusts the sound beam direction to compensate for the lateral deviation. The dynamic Doppler corrector adjusts the frequency characteristics of the audio signal according to the longitudinal deviation, simulates the position change of the sound source in the height direction, and matches the screen coordinates of the video object; in this embodiment, the set value is 3%, and the actual value is adjusted according to the use requirement;
[0145] Through the above correction strategy, the sound and picture position deviation is ensured to be controlled within 2% of the screen width, so as to realize the accurate coupling of sound and image with the video object.
[0146] Although the embodiments of the present application have been shown and described, it can be understood by those skilled in the art that various changes, modifications, replacements and variations can be made to the embodiments without departing from the principles and spirits of the present application, and the scope of the present application is defined by the appended claims and their equivalents.
Claims
1. A multimedia theater audio-video synchronization control system, characterized in that, The application relates to a multimedia theater synchronous control system, which comprises the following modules: A synchronous control module, which establishes a three-dimensional coordinate system in the multimedia theater, takes the center of the multimedia theater as an origin, takes a horizontal direction as an x axis, takes a vertical direction as a y axis, and takes a vertical direction as a z axis; An audio processing module, which comprises a sound image coordinate extraction unit, receives audio and video object data in the multimedia theater, extracts three-dimensional sound source coordinates (x, y, z) and motion vectors (vx, vy, vz) in the three-dimensional coordinate system, and further comprises a mapping unit which stores a mapping relationship between screen pixel coordinates in the theater and a surround sound array, converts the sound source coordinates (x, y, z) into a target sound number set through a calibration matrix, generates a driving signal according to the target sound number set, and makes the separated sound source form a directional sound beam in the theater space; A video processing module, which comprises a projection redundancy unit and synchronously outputs two video signals to main and standby projectors. The specific working steps of the mapping unit are as follows: Nine calibration points are arranged on the screen of the theater, specifically four corner points, a center point and middle points of the edges of the screen; 2. The multimedia theater audio-video synchronization control system of claim 1, wherein: If the Euclidean distance is smaller, the minimum distance variable is updated, and the sound group number set corresponding to the calibration point is recorded; after all the calibration points are traversed, the sound group number set corresponding to the calibration point closest to the sound source coordinates is found; According to the sound group number set found in the above matching process, the target sound number set corresponding to the current sound source coordinates is directly determined. At each calibration point, the difference in sound wave transmission time Δti - Δt is measured between the ultrasonic transmitter and the acoustic array microphone n ; By the sound velocity V, and by the measured Δt1-Δt n The sound source coordinates (xs, ys, zs) corresponding to each calibration point are calculated, the coordinates of each calibration point are obtained, and the Euclidean distance between the coordinates and the current sound source coordinates is calculated. The specific working steps of the mapping unit for generating a driving signal according to the target sound number set, so that the separated sound source forms a directional sound beam in the theater space are as follows: The sound source coordinates (x, y, z) and the target sound number set are obtained, the time delay that needs to be added to each sound is calculated according to the sound source coordinates (x, y, z) and the target sound number set, and the unit is second; 3. The multimedia theater audio-video synchronization control system of claim 2, wherein: The audio processing module further comprises a correction unit which calculates a frequency offset in real time according to the sound source motion vector (vx, vy, vz), and compensates the initial audio stream of each sound and video transmission by frequency conversion, and the specific steps are as follows: The audio processing module further comprises a compensation unit which generates an inverted sound wave to offset non-direct sound when the beamformer generates a directional sound beam to direct the sound source sound to the seat area, and the specific steps are as follows: The contribution weight W of each sound to the sound beam is then calculated k According to the contribution weight W k And the time delay, the driving signal is generated to make the separated sound sources form directional sound beams in the theater space.
4. The multimedia theater audio-video synchronization control system of claim 3, wherein: Based on the HRTF function library of each seat, the direct sound is filtered; The three-dimensional coordinates of each seat in the multimedia theater are obtained, and then the vector (dx, dy, dz) from the sound source to the seat is calculated, and the relative speed V from the sound source to the seat is calculated rel ; According to the formula The frequency offset Δf is calculated, where f0 is the original frequency, which is the frequency in the audio stream that is not affected by the Doppler effect, and the initial audio stream for each audio and video transmission is frequency compensated according to the frequency offset.
5. The multimedia theater audio-video synchronization control system of claim 3, wherein: The direct sound filtered through the HRTF is played through the seat headrest loudspeaker; A measuring device is used to obtain the impulse response R(t) of the early reflection sound at each seat in the theater; A microphone is placed at each seat in the theater, a specific sound source signal is played, and a reflection sound signal is recorded to extract the impulse response; The measured impulse response R(t) is convoluted with the original audio signal, and an inverted sound wave C(t) is generated by taking the negative, and the generated inverted sound wave C(t) is emitted through the loudspeaker array. The synchronous control module is used for establishing a video frame position model based on the physical characteristics of a film projection window, and outputting a frame position signal (Fn, Xn, Yn), and the specific steps further comprise the following: The time when the center of the gth frame passes through the projection window is calculated according to the film movement speed.
6. The multimedia theater audio-video synchronization control system of claim 1, wherein: Aligning the center of the digital video frame with the calculated time alignment through the projection window ensures that the alignment error is less than the preset frame period.
7. The multimedia theater audio-video synchronization control system of claim 6, wherein: The synchronization control module further comprises a mechanical vibration compensation unit, and the specific steps are as follows: A three-axis MEMS accelerometer is installed on the projector to detect the vibration acceleration information of the projector in three axial directions in real time. According to the detected vibration acceleration information, a vibration displacement model is established to calculate the displacement of the projector in the horizontal direction. The vibration acceleration information is fused with vibration information and frame position prediction information by using a Kalman filtering algorithm to dynamically adjust the frame position prediction model. According to the calculation results of the adjusted frame position prediction model and the vibration displacement model, corresponding correction signals are generated and output to the projector.
8. The multimedia theater audio-video synchronization control system of claim 1, wherein: The specific steps of the video processing module are as follows: The screen brightness is monitored by an optical sensor, and when the monitored screen brightness drops below a set threshold, the switching mechanism is triggered. During the operation of the main projector, the standby projector preloads the next frame of content, and switches the signal source during the vertical blanking period of the video. The audio clock source is locked during the switching period, and the audio buffer area is prohibited from refreshing during the switching period.
9. The multimedia theater audio-video synchronization control system of claim 8, wherein: The specific working steps of the video processing module further include the following: By comparing the sound source coordinates (x, y, z) with the video object screen coordinates (Xm, Ym), the horizontal deviation Δx and the vertical deviation Δz are determined: When the absolute value |Δx| of the horizontal deviation exceeds the set value of the screen width, the mapping unit is driven to adjust the sound beam pointing direction to compensate for the horizontal deviation. Vertical correction: when the absolute value |Δz| of the vertical deviation exceeds the set value, the mapping unit is driven to adjust the sound beam pointing direction to compensate for the horizontal deviation.