Intelligent safety helmet audio and video data synchronization method and system

By constructing a predicted drift curve and using a multi-round data synchronization method, the problem of clock drift accumulation error in traditional audio and video synchronization is solved, achieving highly stable and real-time audio and video data synchronization.

CN121888022APending Publication Date: 2026-04-17山东港源管道物流有限公司
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
山东港源管道物流有限公司
Filing Date
2025-12-23
Publication Date
2026-04-17

AI Technical Summary

Technical Problem

Traditional audio and video data synchronization methods cannot effectively suppress the cumulative errors caused by hardware clock drift, resulting in poor audio and video data synchronization stability.

Method used

By acquiring clock drift reference parameters and operating characteristic parameters, a predicted drift curve is constructed, and multiple rounds of data synchronization are performed, including the first round of predictive alignment, the second round of real-time clock drift compensation, and the third round of synchronization deviation statistical analysis, to generate enhanced synchronized audio and video data streams and ensure synchronization stability.

Benefits of technology

It effectively suppresses clock drift caused by individual hardware differences and changes in operating conditions, significantly improving the real-time performance and stability of audio and video synchronization, and is suitable for high-reliability audio and video synchronization under complex operating conditions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121888022A_ABST
    Figure CN121888022A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of audio and video data processing, and discloses an intelligent safety helmet audio and video data synchronization method and system. According to the invention, the expected drift curve is constructed according to the clock drift reference parameter, and the first round of data synchronization is carried out on the current processing frame group based on the expected drift curve, so that long-term and stable main body clock drift caused by hardware individual difference is compensated; according to the method, decoding, separation and statistical feature extraction are carried out on an initial synchronous audio and video data stream to obtain a main drift vector, and real-time compensation of frequency fluctuation caused by operation condition changes is realized through a method of establishing a mapping relation between operation characteristic parameters and clock frequency drift. And the second round of data synchronization is carried out on the to-be-processed new frame group based on the real-time clock drift compensation coefficient, so that accumulation of synchronization errors is suppressed from a signal source, the limitation that a traditional scheme can only correct a timestamp at the tail end of data is broken through, and the real-time performance and stability of audio and video synchronization are remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of audio and video data processing technology, and in particular to a method and system for synchronizing audio and video data in a smart safety helmet. Background Technology

[0002] The smart safety helmet is a modern personal protective equipment integrating computer, communication, sensing, and artificial intelligence technologies. It transcends the single passive protection function of traditional safety helmets, utilizing built-in high-definition cameras, noise-canceling microphones, environmental sensors, and positioning modules to achieve comprehensive perception, recording, and remote interaction of the work environment, personnel status, and operational processes while ensuring head safety. Its core applications are concentrated in high-risk or precision-driven work scenarios such as power line inspection, construction, petrochemicals, emergency rescue, and industrial manufacturing. Through real-time audio and video transmission, first-person perspective collaboration, and hazardous behavior identification and location management, it effectively improves on-site operational safety, collaborative efficiency, and intelligent management levels.

[0003] Traditional audio and video data synchronization methods are essentially post-synchronization strategies based on software timestamps. This method uses the system clock to timestamp audio and video data and encapsulates, stores, or transmits them over the network based on these timestamps in order to achieve synchronization at the playback end. However, this method cannot overcome clock drift at the hardware level. Since the video capture module and the audio codec have their own independent physical clock sources, their frequencies have slight deviations. These deviations accumulate as the device runs, causing the initial slight asynchrony to evolve into significant audio-visual delays at the second level after several hours, forming cumulative drift. Therefore, traditional synchronization methods have the fundamental flaw of being unable to suppress the accumulation of errors, resulting in poor stability of audio and video data synchronization. Summary of the Invention

[0004] The main objective of this invention is to provide a method and system for synchronizing audio and video data in a smart safety helmet, aiming to solve the technical problems in the prior art.

[0005] This invention proposes a method for synchronizing audio and video data in a smart safety helmet, applied to the receiving and processing end of the smart safety helmet, comprising: Acquire first-in-first-out audio and video buffer data transmitted from the smart helmet, and obtain clock drift reference parameters and operating characteristic parameters; Based on the clock drift reference parameters, a expected drift curve is constructed, and the audio and video buffer data is synchronized for the first time based on the expected drift curve to generate and output the initial synchronized audio and video data stream in real time. The initial synchronized audio and video data stream is analyzed to obtain the main drift vector, and the main drift vector is corrected based on the operating characteristic parameters to obtain the real-time clock drift compensation coefficient. Extract the new frame group to be processed sequentially from the audio and video cache data, and perform a second round of data synchronization on the new frame group to be processed based on the real-time clock drift compensation coefficient to generate and output the tracking synchronized audio and video data stream in real time. The tracking and synchronized audio and video data stream is analyzed to obtain the short-term time-series deviation accumulation. Based on the short-term time-series deviation accumulation, the audio and video buffer data is synchronized for the third time to generate and output the enhanced synchronized audio and video data stream in real time. End-to-end synchronization verification is performed on the enhanced synchronized audio and video data stream to obtain the synchronization stability index; Determine whether the synchronization stability index is greater than a preset threshold; If the synchronization stability index is greater than the preset threshold, then the audio and video data synchronization is deemed to meet the requirements. If the synchronization stability index is not greater than a preset threshold, the real-time clock drift compensation coefficient is optimized based on the enhanced synchronized audio and video data stream to obtain an optimized clock drift compensation coefficient. The optimized clock drift compensation coefficient is then used as the real-time clock drift compensation coefficient to return to the step of performing a second round of data synchronization on the new frame group to be processed based on the real-time clock drift compensation coefficient.

[0006] This application also provides an intelligent safety helmet audio and video data synchronization system, including multiple modules, which are used to implement the steps of the above-described intelligent safety helmet audio and video data synchronization method.

[0007] Preferably, the module includes multiple units, which are used to implement the steps of the above-described method for synchronizing audio and video data of a smart safety helmet.

[0008] The present invention also provides a computer device, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps of the above-described smart safety helmet audio and video data synchronization method.

[0009] The present invention also provides a computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the steps of the above-described method for synchronizing audio and video data of a smart safety helmet.

[0010] The beneficial effects of this invention are as follows: This invention constructs a predicted drift curve based on clock drift reference parameters and extracts the current processing frame group sequentially from the audio and video buffer data. Based on the predicted drift curve, it performs a first round of data synchronization on the current processing frame group to achieve predictive alignment of audio and video data, compensating for long-term and stable main clock drift caused by individual hardware differences. This invention obtains the main drift vector by decoding, separating, and extracting statistical features from the initial synchronized audio and video data stream, and then corrects the main drift vector based on operating characteristic parameters to obtain the real-time clock drift compensation coefficient. This method, by establishing a mapping relationship between operating characteristic parameters and clock frequency drift, allows for the real-time monitoring of chip junction temperature, power supply voltage, etc. The parameter changes are transformed into dynamic correction of the main drift vector, thereby achieving real-time compensation for frequency fluctuations caused by changes in operating conditions and effectively suppressing cumulative drift caused by crystal oscillator temperature characteristics and power supply fluctuations. Based on the real-time clock drift compensation coefficient, a second round of data synchronization is performed on the new frame group to be processed, thereby suppressing the accumulation of synchronization errors from the signal source and breaking through the limitation of traditional solutions that can only correct timestamps at the end of the data. This invention performs statistical analysis of synchronization deviation within a fixed time window on the tracking and synchronizing audio and video data streams, and initiates enhanced adjustment based on the analysis results, thereby further offsetting sudden timing deviations caused by factors such as instantaneous changes in operating conditions and signal transmission delays, significantly improving the real-time performance and stability of audio and video synchronization. Attached Figure Description

[0011] Figure 1 This is a schematic diagram of a method flow according to an embodiment of the present invention.

[0012] Figure 2 This is a schematic diagram of the system structure according to an embodiment of the present invention.

[0013] Figure 3 This is a schematic diagram of the internal structure of a computer device according to an embodiment of this application.

[0014] The realization of the objective, functional features and advantages of the present invention will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation

[0015] It should be understood that the specific embodiments described herein are merely illustrative of the invention and are not intended to limit the invention.

[0016] like Figure 1 As shown, this application provides a method for synchronizing audio and video data in a smart safety helmet, applied to the receiving and processing end of a smart safety helmet, including: S1. Obtain a first-in-first-out audio and video buffer data transmitted from the smart safety helmet, and obtain clock drift reference parameters and operating characteristic parameters; S2. Construct an expected drift curve based on the clock drift reference parameters, and perform the first round of data synchronization on the audio and video buffer data based on the expected drift curve to generate and output the initial synchronized audio and video data stream in real time; S3. Analyze the initial synchronized audio and video data stream to obtain the main drift vector, and correct the main drift vector based on the operating characteristic parameters to obtain the real-time clock drift compensation coefficient. S4. Extract the new frame group to be processed from the audio and video cache data in sequence, and perform a second round of data synchronization on the new frame group to be processed based on the real-time clock drift compensation coefficient to generate and output the tracking and synchronization audio and video data stream in real time. S5. Analyze the tracking and synchronization audio and video data stream to obtain the short-term timing deviation accumulation, and perform a third round of data synchronization on the audio and video cache data based on the short-term timing deviation accumulation to generate and output the enhanced synchronization audio and video data stream in real time. S6. Perform end-to-end synchronization verification on the enhanced synchronous audio and video data stream to obtain the synchronization stability index; S7. Determine whether the synchronization stability index is greater than a preset threshold. If the synchronization stability index is greater than the preset threshold, then the audio and video data synchronization is deemed to meet the requirements. If the synchronization stability index is not greater than a preset threshold, the real-time clock drift compensation coefficient is optimized based on the enhanced synchronized audio and video data stream to obtain an optimized clock drift compensation coefficient. The optimized clock drift compensation coefficient is then used as the real-time clock drift compensation coefficient to return to the step of performing a second round of data synchronization on the new frame group to be processed based on the real-time clock drift compensation coefficient.

[0017] As described in steps S1-S7 above, this invention acquires a first-in-first-out audio and video buffer data transmitted from the smart safety helmet. This audio and video buffer data refers to a queue set of unsynchronized original video frames and audio data blocks stored in the order of acquisition time. During electrical initialization or periodic calibration of the smart safety helmet, a hardware synchronization signal is triggered by the system's underlying driver. This signal simultaneously acts on the video acquisition module and the audio codec. By reading the difference between the internal clock counter values ​​of both at this time, the initial offset of the video and audio clocks is obtained. The nominal frequencies of the operating clocks are read from the registers of the audio codec and the video sensor chip, respectively, to calculate the nominal frequency ratio between the two. The initial offset of the video and audio clocks and the nominal frequency ratio constitute the clock drift reference parameters. Through the chip sensor interface provided by the embedded system kernel, the frequency stabilization characteristic parameters reflecting the physical state of the clock source, including the core power supply voltage and chip junction temperature, are read in real-time via polling or interruption. Aging characteristic parameters calculated based on the manufacturing time and cumulative operating power consumption are read from the lifetime model preset in the device's non-volatile memory. The real-time operating time of the device is obtained from the system running daemon process. Operating characteristic parameters, aging characteristic parameters, and real-time running time together constitute the operating characteristic parameters. By acquiring multi-dimensional data, a comprehensive understanding of the physical characteristics of the clock source and the system operating status can be established. Traditional synchronization methods only add software timestamps to the collected audio and video data at the operating system application layer and use this to synchronize the audio and video. However, this method has inherent defects: it cannot perceive and compensate for the inherent frequency deviation between the video and audio hardware clock sources. This micro-frequency difference will accumulate into macro-level audio and video delay over time. In contrast, this invention constructs a predicted drift curve based on clock drift reference parameters. The predicted drift curve refers to the linear relationship trajectory reflecting the theoretical deviation of the audio and video clock over time. The current processing frame group is extracted sequentially from the audio and video buffer data. Based on the predicted drift curve, the current processing frame group is synchronized in the first round, generating and outputting the initial synchronized audio and video data stream in real time. This invention achieves predictive alignment of audio and video data through the first round of data synchronization, which is used to compensate for the long-term and stable main clock drift caused by individual hardware differences. This step corrects the system error introduced by hardware trigger delay differences from the source, thereby establishing a high-precision reference for the entire synchronization system.

[0018] Since the chip junction temperature and core power supply voltage of the clock source are key variables affecting the physical characteristics of the clock circuit, they indirectly induce frequency drift by changing the operating state of the oscillation circuit. This leads to a relative frequency deviation between the audio clock and the video reference clock, which accumulates over time and ultimately manifests as significant asynchrony between audio and video data. To address the dynamic clock drift caused by changes in operating conditions during system operation, this invention decodes and separates the initial synchronized audio and video data streams to obtain a timing deviation sequence. Statistical feature extraction based on a sliding window is performed on the timing deviation sequence to obtain a trend feature vector. Regression analysis is then used to analyze the trend feature vector to obtain the dominant drift vector. The dominant drift vector refers to the dominant frequency drift feature extracted from the timing deviation using signal decomposition techniques. The dominant drift vector is then corrected based on operating characteristic parameters to obtain the real-time clock drift compensation coefficient. This method establishes a mapping relationship between operating characteristic parameters and clock frequency drift. The changes in parameters such as chip junction temperature and power supply voltage monitored in real time are converted into dynamic corrections of the main drift vector, thereby achieving real-time compensation for frequency fluctuations caused by changes in operating conditions and effectively suppressing cumulative drift caused by crystal oscillator temperature characteristics and power supply fluctuations. New frame groups to be processed are extracted sequentially from the audio and video buffer data. These new frame groups are continuous audio and video data arriving in the buffer queue after the currently processed frame group that has completed initial synchronization. A second round of data synchronization is performed on the new frame groups to be processed based on the real-time clock drift compensation coefficient, generating and outputting a real-time tracking synchronized audio and video data stream. Through the above processing, a synchronization paradigm shift from passive timestamp correction to active frequency tracking is achieved: the audio resampling rate is dynamically controlled by the real-time clock drift compensation coefficient, and the audio acquisition frequency is fine-tuned at the software level, so that the generation rate of the audio data stream actively matches the video reference clock, thereby suppressing the accumulation of synchronization errors from the signal source and overcoming the limitation of traditional solutions that can only correct timestamps at the end of the data.

[0019] This invention performs statistical analysis of synchronization deviations within a fixed time window on the tracked synchronized audio and video data stream to obtain the cumulative amount of short-term timing deviations. Trend analysis of this cumulative amount is then performed to determine whether the aforementioned synchronization method is sufficient to maintain stable synchronization. If the analysis results indicate that the timing deviations are accelerating and the cumulative amount exceeds the dynamic safety threshold, enhanced regulation is initiated: batch timestamp recalibration is performed on the audio data backed up in the buffer queue, while severely out-of-synchronization and irreparable data frames are selectively removed, and an enhanced synchronized audio and video data stream is output. This method achieves rapid synchronization of the audio and video streams. Step-by-step reset effectively prevents the continuous expansion of errors, ensuring that the system can maintain a usable synchronized state even under extreme conditions such as sudden load fluctuations. By decoding and extracting the display timestamps of multiple consecutive video frames and their corresponding audio blocks from the enhanced synchronized audio and video data stream, the difference between each pair of audio and video timestamps is calculated to obtain an end-to-end deviation sequence. The standard deviation of this end-to-end deviation sequence is then calculated, and its reciprocal is used as the synchronization stability index. The synchronization stability index is compared with a preset threshold. If the synchronization stability index is greater than the preset threshold, the audio and video data synchronization is deemed to meet the requirements; otherwise, the synchronization stability index is deemed to be less than the preset threshold. If the value is not greater than a preset threshold, the real-time clock drift compensation coefficient is differentially corrected based on the enhanced synchronized audio and video data stream to obtain an optimized clock drift compensation coefficient. This optimized clock drift compensation coefficient is then used as the real-time clock drift compensation coefficient to return to the step of performing a second round of data synchronization on the new frame group to be processed based on the real-time clock drift compensation coefficient. This on-demand correction mechanism enables the system to have continuous optimization capabilities, fundamentally solving the problem of synchronization accuracy attenuation caused by clock aging and changes in operating conditions during long-term operation of traditional methods. The execution entity of this application is a receiving and processing end independent of the front-end acquisition equipment, in typical... In application scenarios, the front-end smart safety helmet, as a lightweight data acquisition and transmission unit, is responsible for capturing and uploading raw audio and video data in complex environments such as construction sites. The back-end receiving and processing end, as an edge server or command center platform with powerful computing capabilities, is dedicated to running the high-complexity synchronization algorithm described in this application. From the system design level, resource-intensive computing tasks are decoupled from lightweight acquisition tasks, which fundamentally ensures that in industrial scenarios with extremely high reliability requirements, such as power inspection and construction monitoring, the audio and video synchronization reliability and long-term stability of the smart safety helmet under complex working conditions are improved.

[0020] In one embodiment, step S2, which involves constructing a predicted drift curve based on the clock drift reference parameters, performing a first round of data synchronization on the audio and video buffer data based on the predicted drift curve, and generating and outputting the initial synchronized audio and video data stream in real time, includes: S21. Obtain the nominal frequency ratio and reference offset based on the clock drift reference parameters, and construct the expected drift curve based on the nominal frequency ratio and the reference offset; S22. Extract the first video data frame and the corresponding first audio data block from the audio and video cache data in sequence, and obtain the original video timestamp and the original audio timestamp according to the first video data frame and the first audio data block respectively. S23. Perform salient region detection on the first video data frame to obtain multiple salient visual region segments and their corresponding binarized visual hash sequences, and perform acoustic event analysis on the first audio data block to obtain multiple burst acoustic segments and their corresponding binarized acoustic hash sequences. S24. Calculate the Hamming distance based on the binarized visual hash sequence and the binarized acoustic hash sequence, and match the salient visual region segment with the sudden acoustic segment according to the Hamming distance to obtain multiple effective sampling points; S25. Obtain the original difference sequence based on multiple effective sampling points, and fit the original difference sequence and the expected drift curve using least squares linear fitting to obtain the measured drift calibration function. S26. Obtain the audio calibration timestamp based on the original audio timestamp and the measured drift calibration function, and encapsulate the first audio data block and the first video data frame with timestamp alignment based on the audio calibration timestamp and the original video timestamp to generate and output the initial synchronized audio and video data stream in real time.

[0021] As described in steps S21-S26 above, this invention obtains the nominal frequency ratio and reference offset based on clock drift reference parameters. The reference offset characterizes the absolute difference between the time counts of the video hardware clock and the audio hardware clock at the initial moment. Since there is an inherent frequency deviation between the two hardware clock sources, their respective time counts will generate a linear cumulative error proportional to the running time. Therefore, by using the system running time as the independent variable and the theoretically predicted cumulative time deviation as the dependent variable, a linear prediction function is constructed. The reference offset is used as the intercept of the linear prediction function, and the difference between the nominal frequency ratio and "1" is calculated. This difference is used as the slope of the linear prediction function, which characterizes the theoretical difference in speed between the video clock and the audio clock per second. Based on this linear prediction function, an expected drift curve can be constructed. The expected drift curve can predict the dynamic evolution trend of clock deviation with running time in a forward-looking manner, thereby providing a theoretical reference for subsequent real-time synchronization compensation.

[0022] Traditional methods assume that audio and video data frames aligned in chronological order serve as the synchronization benchmark. This assumption of temporal proximity has a fundamental flaw: it presupposes that video frames and audio data blocks under the same time tag must also occur synchronously in the physical world. However, since multiple independent physical events may be captured within a single acquisition cycle, and video and audio acquisition hardware has inherent and inconsistent triggering and processing delays, this method directly incorporates the temporal misalignment and hardware latency differences between different events into the initial synchronization parameters, leading to an inherent lack of benchmark accuracy. This invention, however, extracts the currently processed frame group sequentially from the audio and video buffer data header. The currently processed frame group refers to the frames retrieved from the audio and video buffer data header within the current synchronization cycle that are undergoing initial alignment processing. The system generates a first video data frame and its corresponding first audio data block, and obtains the original timestamps of the video and audio data frames based on the first video data frame and the first audio data block, respectively. A lightweight video decoding library is used to decode the compressed first video data frame into a pixel matrix, which is then uniformly scaled to a low resolution to reduce computational complexity. A salient region detection algorithm is used to locate multiple independent salient visual region segments in the preprocessed first video data frame, and a fixed-length, compact binary visual hash sequence is generated for each salient visual region segment. Simultaneously, a lightweight audio decoding library is used to decode the compressed first audio data block into PCM waveform data, and pre-emphasis and frame-by-frame windowing processing are performed. For each windowed audio frame... The system calculates the Mel-frequency cepstral coefficients of the high-frequency signal, resulting in a feature matrix consisting of multiple frames, each containing multiple Mel-frequency cepstral coefficient dimensions. The first-dimensional coefficient sequence is extracted from this feature matrix, and the short-time variance of all dimensions across consecutive frames is calculated. By detecting the peak values ​​of the first-dimensional coefficients and supplementing this with variance abrupt change points, the system locates time segments of sudden acoustic events (such as impact sounds or sudden command sounds), obtaining multiple sudden acoustic segments. A fixed-length binary acoustic hash sequence is generated for each sudden acoustic segment. The system iterates through each pair of binary visual hash sequences and binary acoustic hash sequences, calculating the Hamming distance between them. The Hamming distance is compared with a preset matching threshold. If a certain pair of hash sequences... If the Hamming distance is less than the matching degree threshold, it is determined that the salient visual region segment represented by the binarized visual hash sequence and the bursty acoustic segment represented by the binarized acoustic hash sequence originate from the same physical event, and this pairing is marked as a valid sampling point. Through the above-mentioned cross-modal content hashing process, this invention generates comparable identity identifiers for visual and acoustic events that must occur simultaneously in the physical world. By distinguishing different events within a frame, accurate cross-modal association is achieved, thereby ensuring that each valid sampling point carries a real synchronization relationship. For each valid sampling point, the system reads its corresponding video frame hardware timestamp and audio block hardware timestamp, and calculates the difference between these two timestamps to obtain the original time difference value of the sampling point.The original time difference sequence is constructed by arranging the original time differences corresponding to all valid sampling points obtained within a synchronization period in chronological order.

[0023] Using system runtime as the independent variable and the original difference sequence as the dependent variable, a least-squares linear fit is performed with the expected drift curve to obtain the measured drift calibration function. This initial data-driven approach evolves the synchronization benchmark from a fixed model dependent on hardware nominal parameters to a dynamic model that fits the actual operating state of the current device, thus fundamentally overcoming the problem of individual clock drift characteristics differences caused by hardware manufacturing tolerances. The original audio timestamp is substituted into the measured drift calibration function for calculation to obtain a forward-compensated audio calibration timestamp aligned with the video clock domain. The first audio data block carrying the audio calibration timestamp is then timestamped and encapsulated with the first video data frame carrying the original video timestamp to generate and output the initial synchronized audio and video data stream in real time.

[0024] In one embodiment, step S3, which involves analyzing the initial synchronized audio and video data stream to obtain the main drift vector and performing correlation analysis between the operating characteristic parameters and the main drift vector to obtain the real-time clock drift compensation coefficient, includes: S31. Decode and separate the initial synchronized audio and video data stream to obtain the initial synchronized video frame sequence and the corresponding initial synchronized audio data block, and obtain the timing deviation sequence based on the initial synchronized video frame sequence and the initial synchronized audio data block through timestamp alignment operation. S32. Perform statistical feature extraction based on a sliding window on the time-series deviation sequence to obtain a trend feature vector, and use regression analysis to analyze the trend feature vector to obtain the main drift vector; S33. Obtain the historical clock drift calibration database, and construct the operating condition drift mapping function based on the historical clock drift calibration database; S34. Obtain the frequency stabilization characteristic parameters, aging characteristic parameters and real-time running time according to the operating characteristic parameters, and obtain the operating condition drift vector according to the frequency stabilization characteristic parameters and the operating condition drift mapping function. S35. Obtain the aging-induced drift vector based on the aging characteristic parameters and the real-time running time, and fuse the amplitude information of the operating condition-induced drift vector, the aging-induced drift vector and the main drift vector to obtain the real-time clock drift compensation coefficient.

[0025] As described in steps S31-S35 above, this invention decodes and separates the initial synchronized audio and video data stream to obtain an initial synchronized video frame sequence with an initial synchronized timestamp and a corresponding initial synchronized audio data block. Using the initial synchronized timestamp of the Nth frame in the initial synchronized video frame sequence as an alignment reference, and extending from this initial synchronized timestamp as the center point forward and backward by twice the reciprocal of the nominal video frame rate, an audio extraction window with a total duration equal to the reciprocal of the nominal video frame rate is determined within the initial synchronized audio data block. All audio sampling points whose sampling times fall within this window are extracted from the initial synchronized audio data block to form an audio analysis segment corresponding to the Nth frame. The arithmetic mean of the absolute timestamps of the first and last sampling points of this segment is calculated to obtain the audio center timestamp. The difference between the audio center timestamp corresponding to the Nth frame and the initial synchronized timestamp is calculated to obtain the single-frame deviation value corresponding to that frame. A series of obtained single-frame deviation values ​​are arranged in chronological order to form an initial timing deviation sequence.

[0026] Statistical features of the time-series deviation sequence are extracted using a sliding window approach. The drift trend is characterized by calculating the moving average of the deviation sequence within a continuous time window, and the variance is calculated to characterize the fluctuation intensity. The trend feature vector composed of these statistical features is then input into a pre-trained drift prediction model. This model establishes a nonlinear mapping relationship from multidimensional features to frequency drift through a multivariate regression method, and outputs a dominant drift vector. The direction of this dominant drift vector characterizes the type of dominant drift frequency, and its magnitude characterizes the rate of the dominant drift.

[0027] Given that video acquisition clocks typically have higher frequency stability and their fixed frame rates are difficult to adjust via software, while audio clock sources are more sensitive to load changes and their data streams can be fine-tuned through techniques such as resampling, this invention chooses the video acquisition clock as the synchronization benchmark. By constructing a composite drift prediction model for the audio clock source, it achieves advanced synchronization compensation for the audio data. This invention obtains a historical clock drift calibration database, which refers to the video clock frequency and audio clock frequency corresponding to different historical chip junction temperatures and historical core power supply voltages obtained through previous calibration experiments. Based on a large number of historical chip junction temperatures and historical core power supply voltages forming operating condition groups and their corresponding drift observation vectors, regression analysis is used to fit and obtain an operating condition drift mapping function describing the quantitative relationship between clock operating conditions and drift vectors. This function indicates that the chip junction temperature and core power supply voltage are related to the operating conditions of the clock. At the same operating point, a two-dimensional vector field is formed that determines the drift direction and rate. The element values ​​in the vector, as specific implementations of the mapping function, describe how the chip junction temperature and core power supply voltage jointly determine the component values ​​of the drift vector in the X and Y axes through linear or nonlinear combinations. Based on the operating characteristic parameters, the frequency stabilization characteristic parameters, aging characteristic parameters, and real-time operating time of the audio clock source are obtained. The frequency stabilization characteristic parameters are substituted into the operating condition drift mapping function, and the operating condition-induced drift vector is calculated. The operating condition-induced drift vector is a two-dimensional vector that characterizes the expected frequency drift trend of the audio clock source under the current operating condition. This invention reveals the coupling law of the chip junction temperature and core power supply voltage acting together on clock drift, unifying two originally independent physical quantities into a vector output with directional characteristics. This accurately describes the nonlinear relationship between the operating condition and drift characteristics of the clock source, laying a theoretical foundation for subsequent intelligent compensation.

[0028] Through "V" x =δ*cos(θ)*sin(ω*t+φ)” calculates the long-term common drift, where V x Let δ represent the long-term common drift, θ represent the aging oscillation amplitude, θ represent the aging oscillation direction angle, ω represent the angular frequency, t represent the real-time running time, and φ represent the initial phase. In the formula, "sin(ω*t+φ)" represents the time-varying amplitude of the virtual aging. By multiplying this amplitude by the cosine of the aging oscillation direction angle, it is projected onto the X-axis representing the common drift. The resulting projected component is used to quantify the common influence of the aging effect on the absolute frequency of the audio and video clocks. Through "V..." y =δ*sin(θ)*sin(ω*t+φ)” calculates the long-term differential drift, where V yLet δ represent the long-term differential drift, θ represent the aging oscillation amplitude, θ represent the aging oscillation direction angle, ω represent the angular frequency, t represent the real-time running time, and φ represent the initial phase. By multiplying the time-varying amplitude of the same virtual aging by the sine of the aging oscillation direction angle and projecting it onto the Y-axis representing the differential drift, the resulting projected component is used to directly quantify the relative frequency deviation between the audio and video clocks caused by the aging effect. The long-term common drift and the long-term differential drift are combined into a two-dimensional vector to obtain the aging-induced drift vector.

[0029] The dynamic drift correction vector is obtained by superimposing the operating condition-induced drift vector and the aging-induced drift vector. Based on the reference direction established by the dynamic drift correction vector, the main drift vector is decomposed to separate the projection component representing the explainable drift and the vertical component representing the abnormal disturbance. The abnormal drift is determined by calculating the magnitude of the vertical component and comparing it with a preset threshold. When the threshold is exceeded, the projection component is weighted and suppressed by an attenuation factor to reduce the impact of observation noise. Finally, the optimized projection component and the dynamic drift correction vector are synthesized, and the magnitude of the synthesized vector is used as the real-time clock drift compensation coefficient. Through the anomaly detection and suppression mechanism based on vector projection, unreliable components in the observation data are corrected. Finally, the corrected observation values ​​are fused with the predicted values, thereby significantly improving the accuracy of the system in complex real-world environments.

[0030] In one embodiment, step S4, which involves sequentially extracting new frame groups to be processed from the audio and video buffer data, performing a second round of data synchronization on the new frame groups to be processed based on the real-time clock drift compensation coefficient, and generating and outputting the tracking and synchronized audio and video data stream in real time, includes: S41. Obtain the second video data frame and the corresponding second audio data block according to the new frame group to be processed; S42. Decode the second audio data block to obtain a pulse code modulated audio signal, and use speech activity detection technology to analyze the pulse code modulated audio signal to obtain multiple second audio segments and their corresponding speech existence probabilities; S43. Based on the probability of the existence of the speech, multiple second audio segments are segmented to obtain candidate speech segments and candidate environmental sound segments, and the candidate speech segments and candidate environmental sound segments are analyzed to obtain key audio segments and non-key audio segments. S44. Based on the real-time clock drift compensation coefficient, perform differentiated synchronization processing on the key audio segments and the non-key audio segments respectively to obtain the first drift compensation audio sequence and the second drift compensation audio sequence; S45. The first drift-compensated audio sequence and the second drift-compensated audio sequence are spliced ​​together to obtain a mixed audio sequence. The mixed audio sequence and the second video data frame are then timestamped and encapsulated to generate and output a tracking and synchronized audio and video data stream in real time.

[0031] As described in step S41- As described in S45, this invention obtains a second video data frame and a corresponding second audio data block based on the new frame group to be processed, and decodes the second audio data block to obtain a pulse code modulation audio signal. It then uses speech activity detection technology to analyze the pulse code modulation audio signal, obtaining multiple second audio segments and their corresponding speech presence probabilities. When the speech presence probability of multiple consecutive second audio segments exceeds a preset threshold, this series of consecutive second audio segments is determined as a candidate speech segment. Consecutive second audio segments that do not meet the above conditions are classified as candidate ambient sound segments. A fundamental frequency extraction algorithm is applied to the candidate speech segments to analyze their harmonic structure. The presence of a stable fundamental frequency trajectory is detected to confirm and mark them as key speech segments and non-key speech segments. For candidate ambient sound segments, the spectral flatness and zero-crossing rate of the frame signal are calculated for auxiliary determination: when the spectral flatness is high and the zero-crossing rate is within a specific range, it is determined as a non-key ambient sound segment; otherwise, it is determined as a key ambient sound segment. Key speech segments and key ambient sound segments are considered key audio segments, while non-key speech segments and non-key ambient sound segments are considered non-key audio segments.

[0032] For key audio segments, they are converted into a unified PCM audio sampling point sequence as processing input, and processed using a phase vocoder-based time-domain stretching technique to obtain the resampling ratio. The resampling ratio refers to the scaling factor used to stretch or compress the audio time axis to compensate for relative clock drift. The original phase trajectory corresponding to the key audio segment is obtained from the second audio data block. Based on the resampling ratio, the original phase trajectory is mapped onto a new time axis to form the target phase trajectory. The target phase trajectory is then discretely sampled according to the adjusted time axis, generating a new phase spectrum corresponding to each frequency domain frame. The original amplitude spectrum is combined with the scaled new phase spectrum, and then resynthesized into a time-domain signal using an inverse short-time Fourier transform to obtain the first drift-compensated audio sequence. Through this process of precisely controlling the evolution of phase continuity, while changing the audio duration to compensate for clock drift, it is beneficial to obtain a high-fidelity audio sequence that is strictly consistent with the video clock domain. Aligned audio data is obtained while fully preserving the formant structure and fundamental frequency characteristics of the speech. For non-critical environmental audio segments, they are converted into a unified pulse-code modulation audio sampling point sequence as processing input, and a linear interpolation algorithm is used for repositioning to obtain a theoretical position index, thereby establishing the sampling point position mapping relationship required for clock drift compensation. By locating the two nearest original sampling points before and after the theoretical position index, i.e., the sampling points corresponding to the floor index and the sampling points corresponding to the floor index, and calculating the final sampling value through linear weighting, the second drift-compensated audio sequence is obtained by performing the above operation on all current output indices. This method can complete the synchronization processing of non-critical environmental audio segments with extremely low computational complexity while ensuring basic synchronization requirements. This invention constructs two heterogeneous processing channels with high fidelity and high efficiency for audio segments of different importance, thereby achieving accurate compensation for clock drift and optimized allocation of system computing resources.

[0033] By reconstructing the first audio sequence and the second audio sequence into a complete mixed audio sequence, and using timestamp-aligned encapsulation technology, the mixed audio sequence is recombined with the second video data frame. Then, through the edit box structure in the MP4 encapsulation standard, the matching audio and video data frames are atomically bound to generate a tracking synchronized audio and video data stream.

[0034] In one embodiment, step S5, which involves analyzing the tracking synchronized audio and video data stream to obtain the short-term timing deviation accumulation, performing a third round of data synchronization on the audio and video buffer data based on the short-term timing deviation accumulation, and generating and outputting the enhanced synchronized audio and video data stream in real time, includes: S51. The tracking synchronous audio and video data stream is analyzed using fixed time window technology to obtain the short-term time series deviation accumulation, and the short-term time series deviation accumulation is analyzed to obtain trend characteristic parameters. S52. Obtain the cumulative threshold value of time series deviation, and adjust the cumulative threshold value of time series deviation based on the trend feature parameters to obtain the cumulative threshold value of time series deviation; S53. Determine whether the accumulated short-term time series deviation is greater than the accumulated time series deviation threshold; If the cumulative amount of the short-term timing deviation corresponding to any fixed time window is greater than the cumulative timing deviation threshold, then the third video data frame and the corresponding third audio data block are extracted sequentially from the audio and video cache data. S54. Analyze the third audio data block to obtain audio synchronization quality characteristics, and divide the third audio data block based on the audio synchronization quality characteristics to obtain correctable audio segments, mergeable audio segments, and pass-through audio segments. S55. Process the correctable audio segment and the mergeable audio segment to obtain a timestamp-calibrated audio segment and a merged audio segment. Then, splice the timestamp-calibrated audio segment, the merged audio segment, and the pass-through audio segment to obtain an enhanced and corrected audio block. S56. The enhanced correction audio block and the third video data frame are timestamped and encapsulated to generate and output the enhanced synchronous audio and video data stream in real time.

[0035] As described in steps S51-S56 above, this invention performs real-time quality monitoring of the tracking synchronous audio and video data stream in fixed time windows. It decodes the data stream and extracts the presentation timestamps of each video frame and its corresponding audio block, calculating the difference in presentation timestamps for each pair of matching audio and video frames as the instantaneous timing deviation. A time window contains multiple such instantaneous timing deviations. The absolute values ​​of all instantaneous timing deviations within the current time window are summed to obtain a single short-term timing deviation accumulation. The system continuously records and saves the short-term timing deviation accumulations of the most recent N time windows in chronological order, thus forming a deviation accumulation sequence. Trend analysis is performed on this deviation accumulation sequence: it is fitted to an exponential function model, and the coefficient of the exponential term is calculated using the least squares method for curve fitting. The coefficient of determination of the fitted value is then calculated. This coefficient measures the degree of fit between the accumulation sequence and the exponential model; a value closer to 1 indicates a higher degree of fit. The coefficients and the coefficient of determination of the fit together constitute the trend feature parameters; the CPU utilization rate and the remaining battery percentage are obtained; the timing deviation cumulative threshold is obtained and dynamically adjusted: if the exponential coefficient is greater than zero and the coefficient of determination of the fit is greater than 0.9, it is determined that the change gradient is exponentially increasing, and the CPU utilization rate exceeds the high load threshold of 70%, then it is determined that the system is in a vicious cycle of decreased clock stability due to high load, which in turn exacerbates synchronization drift. At this time, the timing deviation cumulative threshold is tightened to 80% of the timing deviation cumulative threshold to reduce the fault tolerance rate, trigger the correction mechanism in advance, and prevent the error from getting out of control. If the remaining battery percentage is lower than the low battery threshold of 20%, considering that the synchronization correction operation will significantly increase power consumption, in order to ensure the continuous operation capability of the device, the timing deviation cumulative threshold is relaxed to 150% of the timing deviation cumulative threshold to reduce the triggering frequency of high-intensity correction operations and achieve intelligent balance between power consumption and synchronization accuracy.

[0036] The accumulated short-term timing deviation for each monitoring time window is compared with the accumulated timing deviation threshold. If the accumulated short-term timing deviation for any monitoring time window exceeds the accumulated timing deviation threshold, forced correction is initiated. After forced correction is initiated, the latest batch of collected but not yet synchronized frames to be strengthened is retrieved from the first-in-first-out audio and video buffer data. The frame group to be strengthened refers to a batch of data to be processed, defined sequentially from the header of the audio and video buffer data when forced correction is triggered due to the excessive accumulated short-term timing deviation. This batch consists of subsequent third video data frames and their corresponding third audio data blocks. The third audio data blocks are divided into multiple third audio segments, and the original timestamp of each third audio segment is retrieved sequentially and compared with the current timestamp of the system master clock. The theoretical time deviation caused by clock drift in the third audio segment is calculated, which is the theoretical deviation value. Each third audio segment is decoded to obtain a pulse-code modulation signal, and the root mean square value of its sampling points is calculated to obtain the short-time energy of the data block. Based on the audio synchronization quality characteristics constituted by the theoretical deviation value and the short-time energy, the third audio segments are classified as follows: third audio segments with an absolute theoretical deviation value of less than 50 milliseconds are classified as correctable audio segments; third audio segments with an absolute theoretical deviation value greater than 50 milliseconds but not greater than 200 milliseconds and a short-time energy below the -60 dBFS silence threshold are classified as mergeable audio segments; third audio segments with an absolute theoretical deviation value greater than 200 milliseconds are classified as pass-through audio segments. This invention's segment classification method based on audio feature analysis achieves accurate identification and classification of abnormal audio data.

[0037] Using the original timestamp and real-time clock drift compensation coefficient of the calibrable audio segment as input parameters, the measured drift calibration function is substituted into the forward calculation to obtain the calibrated audio timestamp aligned with the video clock domain. The original timestamp in the calibrable audio segment is then replaced with the calibrated audio timestamp, thus outputting the timestamp-calibrated audio segment. For audio segments marked as mergingable, duration scaling is performed: the target video frame sequence corresponding to the mergingable audio segment in time is determined, and the total duration of this video frame sequence is calculated. The original total duration of the mergingable audio segment is then calculated. Based on the ratio of these two durations, the required time scaling factor is calculated. A resampling-based audio time scaling algorithm is used to stretch or compress the mergingable audio segment on the time axis according to this scaling factor, ensuring that its total duration precisely matches the total duration of the target video frame sequence. The processed audio segment is the merged audio segment.

[0038] The continuity detection of timestamp-calibrated audio segments, merged audio segments, and pass-through audio segments is performed. Energy breakpoints are located by calculating the difference between the first and last sampling points of adjacent audio segments. For each identified breakpoint, a 10-millisecond waveform segment is extracted from the audio segment before the breakpoint, and a segment of the same length is extracted from the audio segment after the breakpoint. The cross-fade function of the two segments is calculated, and they are weighted and superimposed based on waveform similarity to generate a smooth 20-millisecond transition audio segment. This transition audio segment is inserted into the breakpoint position to output an enhanced correction audio block, ensuring that the final output audio data has strict continuity in both timing and sound, eliminating audio stuttering or jumps that may be caused by batch audio segment processing. The enhanced correction audio block and the third video data frame are timestamp-aligned and encapsulated to generate and output an enhanced synchronous audio and video data stream in real time, thus completing this round of synchronous enhancement adjustment.

[0039] In one embodiment, step S7, which optimizes the real-time clock drift compensation coefficient based on the enhanced synchronous audio and video data stream to obtain the optimized clock drift compensation coefficient, includes: S71. Based on the enhanced synchronous audio and video data stream, trace the source to obtain the original data acquisition time period; S72. Obtain synchronous processing cache data, and extract frequency stability characteristic parameter sequence and timing deviation sequence fragment from the synchronous processing cache data based on the original data acquisition period; S73. Obtain the failure mode rule base, and match the frequency stability characteristic parameter sequence and the timing deviation sequence segment based on the failure mode rule base to obtain the synchronization failure type; S74. Based on the synchronization failure type, frequency stability parameter sequence, and timing deviation sequence segment, the real-time clock drift compensation coefficient is adaptively mapped and corrected to obtain the optimized clock drift compensation coefficient.

[0040] As described in steps S71-S74 above, this invention obtains the enhancement correction timestamp based on the enhanced synchronous audio and video data stream, and uses the traceability identifier carried in the enhanced synchronous audio and video data stream to reverse query its original state before the synchronous enhancement processing, thereby locating the original data acquisition period corresponding to the enhanced synchronous audio and video data stream when performing synchronous processing on the new frame group to be processed; and extracts the frequency stability characteristic parameter sequence and timing deviation sequence fragments collected and stored in the original data acquisition period from the synchronous processing cache data. The synchronous processing cache data refers to a first-in-first-out circular buffer that only stores lightweight parameters of the recent synchronization process and does not store the audio and video data itself. This application achieves full-link tracking of the audio and video data processing process by introducing the traceability identifier and the synchronous processing cache data, thereby solving the problem of difficulty in locating the source of error in traditional methods.

[0041] By acquiring a pre-set failure mode rule base and matching the frequency stability parameter sequence and the timing deviation sequence segment based on the failure mode rule base, synchronous failure types are obtained. These synchronous failure types include operational synchronous failure and systemic synchronous failure. The failure mode rule base pre-sets feature criteria for different failure types. The first-order difference between the real-time junction temperature sequence and the core power supply voltage sequence in the frequency stability parameter sequence is calculated. The number of times the absolute value of the real-time junction temperature difference sequence exceeds its corresponding preset threshold is counted, as is the number of times the absolute value of the core power supply voltage difference sequence exceeds its corresponding preset threshold. If the percentage of times any parameter sequence exceeds the limit per unit time exceeds a first preset ratio... For example, if the parameters show drastic fluctuations, it indicates that the audio and video data synchronization failure is mainly caused by sudden changes in external operating conditions, and is therefore classified as an operating condition-related synchronization failure. Least squares linear fitting is performed on the timing deviation sequence segments to obtain the slope of the fitted line and the statistical significance probability value. If the absolute value of the slope of the fitted line is greater than zero and the statistical significance probability value is less than 0.05, it indicates that the timing deviation shows a continuous linear growth trend in the same direction, indicating that the audio and video data synchronization failure is mainly caused by the inherent frequency deviation of the clock source, and is therefore classified as a systemic synchronization failure. This invention achieves accurate identification of the root cause of synchronization failure through a dual-channel diagnostic mechanism based on physical parameter mutation detection and timing trend statistical verification, using multi-dimensional signal feature analysis.

[0042] Adaptive mapping correction is performed based on the type of synchronization failure: a two-dimensional correction data table corresponding to different synchronization failure types is obtained from the system. This data table uses key diagnostic feature values ​​as the horizontal axis and precise correction amounts as the vertical axis. The corresponding data table is selected according to the synchronization failure type, and the aforementioned feature values ​​are used as input to query the selected data table. If the feature value lies between two preset nodes in the data table, a linear interpolation algorithm is used to calculate the correction value. The correction value is then arithmetically superimposed with the real-time clock drift compensation coefficient to obtain the optimized clock drift compensation coefficient. Through this targeted correction mechanism for different failure root causes, the self-optimization capability of the synchronization system is realized.

[0043] This application also provides an intelligent safety helmet audio and video data synchronization system, including: The data acquisition module is used to acquire first-in-first-out audio and video buffer data transmitted from the smart safety helmet, and to acquire clock drift reference parameters and operating characteristic parameters; The first synchronization module is used to construct an expected drift curve based on the clock drift reference parameters, and perform the first round of data synchronization on the audio and video buffer data based on the expected drift curve, generating and outputting the initial synchronized audio and video data stream in real time; The drift correction module is used to analyze the initial synchronized audio and video data stream to obtain the main drift vector, and correct the main drift vector based on the operating characteristic parameters to obtain the real-time clock drift compensation coefficient. The second synchronization module is used to extract the new frame group to be processed sequentially from the audio and video cache data, and perform a second round of data synchronization on the new frame group to be processed based on the real-time clock drift compensation coefficient, thereby generating and outputting the tracking and synchronization audio and video data stream in real time. The enhanced synchronization module is used to analyze the tracking and synchronization audio and video data stream, obtain the short-term time series deviation accumulation, and perform a third round of data synchronization on the audio and video cache data based on the short-term time series deviation accumulation, thereby generating and outputting the enhanced synchronization audio and video data stream in real time. The synchronization verification module is used to perform end-to-end synchronization verification on the enhanced synchronized audio and video data stream to obtain the synchronization stability index. The result judgment module is used to determine whether the synchronization stability index is greater than a preset threshold. If the synchronization stability index is greater than the preset threshold, then the audio and video data synchronization is deemed to meet the requirements. If the synchronization stability index is not greater than a preset threshold, the real-time clock drift compensation coefficient is optimized based on the enhanced synchronized audio and video data stream to obtain an optimized clock drift compensation coefficient. The optimized clock drift compensation coefficient is then used as the real-time clock drift compensation coefficient to return to the step of performing a second round of data synchronization on the new frame group to be processed based on the real-time clock drift compensation coefficient.

[0044] In one embodiment, the drift correction module includes: The decoding and separation unit is used to decode and separate the initial synchronized audio and video data stream to obtain the initial synchronized video frame sequence and the corresponding initial synchronized audio data block, and to obtain the timing deviation sequence based on the initial synchronized video frame sequence and the initial synchronized audio data block through a timestamp alignment operation. The drift analysis unit is used to extract statistical features of the time-series deviation sequence based on a sliding window to obtain a trend feature vector, and to analyze the trend feature vector using a regression analysis method to obtain the main drift vector; A mapping construction unit is used to obtain a historical clock drift calibration database and construct a working condition drift mapping function based on the historical clock drift calibration database. The operation analysis unit is used to obtain frequency stability characteristic parameters, aging characteristic parameters and real-time running time based on the operation characteristic parameters, and to obtain the operating condition drift vector based on the frequency stability characteristic parameters and the operating condition drift mapping function. The compensation fusion unit is used to obtain the aging-induced drift vector based on the aging characteristic parameters and the real-time running time, and to fuse the amplitude information of the operating condition-induced drift vector, the aging-induced drift vector and the main drift vector to obtain the real-time clock drift compensation coefficient.

[0045] The present invention also provides a computer device, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps of the above-described smart safety helmet audio and video data synchronization method.

[0046] The present invention also provides a computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the steps of the above-described method for synchronizing audio and video data of a smart safety helmet.

[0047] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in this application and in the embodiments can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in a variety of forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual-speed SDRAM (SSRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), RAMbus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM).

[0048] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, apparatus, article, or method that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, apparatus, article, or method. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, apparatus, article, or method that includes that element.

[0049] The above description is merely a preferred embodiment of the present invention and does not limit the patent scope of the present invention. Any equivalent structural or procedural transformations made based on the content of the present invention's specification and drawings, or direct or indirect applications in other related technical fields, are similarly included within the patent protection scope of the present invention.

Claims

1. A method for audio-video data synchronization of a smart safety helmet, applied to a receiving and processing end of a smart safety helmet, characterized in that, include: Acquire first-in-first-out audio and video buffer data transmitted from the smart helmet, and obtain clock drift reference parameters and operating characteristic parameters; Based on the clock drift reference parameters, a expected drift curve is constructed, and the audio and video buffer data is synchronized for the first time based on the expected drift curve to generate and output the initial synchronized audio and video data stream in real time. The initial synchronized audio and video data stream is analyzed to obtain the main drift vector, and the main drift vector is corrected based on the operating characteristic parameters to obtain the real-time clock drift compensation coefficient. Extract the new frame group to be processed sequentially from the audio and video cache data, and perform a second round of data synchronization on the new frame group to be processed based on the real-time clock drift compensation coefficient to generate and output the tracking synchronized audio and video data stream in real time. The tracking and synchronized audio and video data stream is analyzed to obtain the short-term time-series deviation accumulation. Based on the short-term time-series deviation accumulation, the audio and video cache data is synchronized for the third time to generate and output the enhanced synchronized audio and video data stream in real time. End-to-end synchronization verification is performed on the enhanced synchronized audio and video data stream to obtain the synchronization stability index; Determine whether the synchronization stability index is greater than a preset threshold; If the synchronization stability index is greater than the preset threshold, then the audio and video data synchronization is deemed to meet the requirements. If the synchronization stability index is not greater than a preset threshold, the real-time clock drift compensation coefficient is optimized based on the enhanced synchronized audio and video data stream to obtain an optimized clock drift compensation coefficient. The optimized clock drift compensation coefficient is then used as the real-time clock drift compensation coefficient to return to the step of performing a second round of data synchronization on the new frame group to be processed based on the real-time clock drift compensation coefficient.

2. The method for synchronizing audio and video data in a smart safety helmet according to claim 1, characterized in that, The steps of constructing a predicted drift curve based on the clock drift reference parameters, performing a first round of data synchronization on the audio and video buffer data based on the predicted drift curve, and generating and outputting the initial synchronized audio and video data stream in real time include: The nominal frequency ratio and reference offset are obtained based on the clock drift reference parameters, and the expected drift curve is constructed based on the nominal frequency ratio and the reference offset. The first video data frame and the corresponding first audio data block are extracted sequentially from the audio and video cache data, and the original video timestamp and the original audio timestamp are obtained according to the first video data frame and the first audio data block, respectively. The first video data frame is subjected to salient region detection to obtain multiple salient visual region segments and their corresponding binarized visual hash sequences. The first audio data block is subjected to acoustic event analysis to obtain multiple burst acoustic segments and their corresponding binarized acoustic hash sequences. Hamming distance is calculated based on the binarized visual hash sequence and the binarized acoustic hash sequence, and the salient visual region segment is matched with the sudden acoustic segment according to the Hamming distance to obtain multiple effective sampling points; The original difference sequence is obtained based on multiple valid sampling points, and the original difference sequence and the expected drift curve are fitted by least squares linear fitting to obtain the measured drift calibration function. The audio calibration timestamp is obtained based on the original audio timestamp and the measured drift calibration function. The first audio data block and the first video data frame are then timestamped and encapsulated based on the original audio timestamp and the original video timestamp to generate and output the initial synchronized audio and video data stream in real time.

3. The method for synchronizing audio and video data in a smart safety helmet according to claim 1, characterized in that, The steps of analyzing the initial synchronized audio and video data stream to obtain the main drift vector, and correlating the operating characteristic parameters with the main drift vector to obtain the real-time clock drift compensation coefficient, include: The initial synchronized audio and video data stream is decoded and separated to obtain the initial synchronized video frame sequence and the corresponding initial synchronized audio data block. The timing deviation sequence is obtained based on the initial synchronized video frame sequence and the initial synchronized audio data block through a timestamp alignment operation. The time-series deviation sequence is subjected to statistical feature extraction based on a sliding window to obtain a trend feature vector, and the trend feature vector is analyzed using a regression analysis method to obtain the main drift vector; Obtain a historical clock drift calibration database and construct a working condition drift mapping function based on the historical clock drift calibration database; Based on the operating characteristic parameters, obtain the frequency stabilization characteristic parameters, aging characteristic parameters, and real-time operating time; and based on the frequency stabilization characteristic parameters and the operating condition drift mapping function, obtain the operating condition drift vector. The aging-induced drift vector is obtained based on the aging characteristic parameters and the real-time running time. The amplitude information of the operating condition-induced drift vector, the aging-induced drift vector, and the main drift vector are then fused to obtain the real-time clock drift compensation coefficient.

4. The method for synchronizing audio and video data in a smart safety helmet according to claim 1, characterized in that, The steps of sequentially extracting new frame groups to be processed from the audio and video buffer data, performing a second round of data synchronization on the new frame groups to be processed based on the real-time clock drift compensation coefficient, and generating and outputting the tracking and synchronized audio and video data stream in real time include: The second video data frame and the corresponding second audio data block are obtained according to the new frame group to be processed; The second audio data block is decoded to obtain a pulse code modulated audio signal, and the pulse code modulated audio signal is analyzed using speech activity detection technology to obtain multiple second audio segments and their corresponding speech existence probabilities; Based on the probability of the speech presence, multiple second audio segments are segmented to obtain candidate speech segments and candidate ambient sound segments. The candidate speech segments and candidate ambient sound segments are then analyzed to obtain key audio segments and non-key audio segments. Based on the real-time clock drift compensation coefficient, the key audio segments and the non-key audio segments are subjected to differentiated synchronization processing to obtain the first drift compensation audio sequence and the second drift compensation audio sequence. The first drift-compensated audio sequence and the second drift-compensated audio sequence are concatenated to obtain a mixed audio sequence. The mixed audio sequence and the second video data frame are then timestamped and encapsulated to generate and output a tracking synchronized audio and video data stream in real time.

5. The method for synchronizing audio and video data in a smart safety helmet according to claim 1, characterized in that, The steps of analyzing the tracking synchronized audio and video data stream to obtain the short-term timing deviation accumulation, and performing a third round of data synchronization on the audio and video buffer data based on the short-term timing deviation accumulation to generate and output the enhanced synchronized audio and video data stream in real time include: The tracking synchronous audio and video data stream is analyzed using a fixed time window technique to obtain the short-term time series deviation accumulation, and the short-term time series deviation accumulation is analyzed to obtain trend characteristic parameters. Obtain the cumulative threshold value of time series deviation, and adjust the cumulative threshold value of time series deviation based on the trend feature parameters to obtain the cumulative threshold value of time series deviation; Determine whether the accumulated short-term time series deviation is greater than the accumulated time series deviation threshold; If the cumulative amount of the short-term timing deviation corresponding to any fixed time window is greater than the cumulative timing deviation threshold, then the third video data frame and the corresponding third audio data block are extracted sequentially from the audio and video cache data. The third audio data block is analyzed to obtain audio synchronization quality characteristics, and the third audio data block is divided based on the audio synchronization quality characteristics to obtain correctable audio segments, mergeable audio segments, and pass-through audio segments. The correctable audio segment and the mergeable audio segment are processed to obtain a timestamp-calibrated audio segment and a merged audio segment. The timestamp-calibrated audio segment, the merged audio segment, and the pass-through audio segment are spliced ​​together to obtain an enhanced and corrected audio block. The enhanced and corrected audio block and the third video data frame are timestamped and encapsulated to generate and output an enhanced synchronized audio and video data stream in real time.

6. The method for synchronizing audio and video data in a smart safety helmet according to claim 1, characterized in that, The step of optimizing the real-time clock drift compensation coefficient based on the enhanced synchronous audio and video data stream to obtain the optimized clock drift compensation coefficient includes: The original data acquisition time period was obtained by tracing the source of the enhanced synchronous audio and video data stream. Acquire synchronous processing cache data, and extract frequency stability characteristic parameter sequence and timing deviation sequence fragment from the synchronous processing cache data based on the original data acquisition period; Obtain the failure mode rule base, and match the frequency stability characteristic parameter sequence and the timing deviation sequence segment based on the failure mode rule base to obtain the synchronization failure type; Based on the synchronization failure type, the frequency stability parameter sequence, and the timing deviation sequence segment, the real-time clock drift compensation coefficient is adaptively mapped and corrected to obtain the optimized clock drift compensation coefficient.

7. A smart safety helmet audio and video data synchronization system, characterized in that, It includes multiple modules for implementing the steps of the method according to any one of claims 1 to 6.

8. The intelligent safety helmet audio and video data synchronization system according to claim 7, characterized in that, It includes multiple units, which are used to implement the steps of the method according to any one of claims 1 to 6.

9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 6.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 6.