Multi-terminal audio synchronization test method and system for vehicle-mounted KTV (Karaoke Television)
By embedding short pulses and coded markers in the vehicle-mounted KTV terminal, and utilizing high-precision microphone probe acquisition and weighted fusion technology, the problem of multi-terminal audio synchronization testing in the vehicle environment was solved, achieving efficient and reliable synchronous measurement in complex environments.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-30
- Publication Date
- 2026-03-24
AI Technical Summary
Existing technologies lack a dedicated solution for directly, reliably, quantitatively, and efficiently testing the audio playback synchronization accuracy of multiple KTV terminals in real, complex, and mobile in-vehicle environments. Especially in the context of vehicle networking, it is impossible to accurately measure the hardware and driver layer latency between the issuance of rendering commands and the actual output of the sound card/graphics card.
A dual-marking strategy of short pulse marking and coded marking is adopted. By embedding short pulses and coded markings at each beat point, audio signals are collected using a high-precision microphone probe. Through weighted fusion and unified time reference correction, the audio synchronization deviation of each terminal is calculated to generate audio synchronization calibration parameters.
It realizes a measurement paradigm shift from data transmission synchronization to user experience synchronization, and can obtain reliable and accurate beat point arrival times in complex acoustic environments. It solves the problem of clock asynchrony in distributed measurement and ensures the accuracy of multi-terminal synchronous testing.
Smart Images

Figure CN121728409A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of audio signal processing technology, and more specifically, to a multi-terminal audio synchronization testing method and system for in-vehicle KTVs. Background Technology
[0002] With the rapid development of in-vehicle infotainment systems, in-vehicle karaoke functionality has become a key feature for enhancing the driving experience. Especially in the context of connected vehicles, in-vehicle karaoke systems supporting multi-vehicle interaction and online group singing have begun to emerge. The core experience of these systems lies in the ability to synchronously play the accompaniment of the same song on different in-vehicle terminals, and to achieve real-time or near-real-time mixing and sharing of user vocals. Therefore, ensuring a high degree of synchronization in audio playback across all participating terminals is a crucial technical challenge for maintaining a good user experience.
[0003] Currently, the main methods for testing the synchronization performance of networked multimedia systems are as follows: 1. Network transmission latency-based testing: This type of method focuses on the time delay (network latency, jitter) of data packets being transmitted from the server to the client, typically using tools such as Ping and network performance analysis. Its fundamental flaw is that it measures the time at the data transmission layer, not the actual time when the audio / video signal is rendered and output to the speaker or screen by the playback device. Delays introduced by decoding, buffering, digital-to-analog conversion, and operating system scheduling in the playback chain are not taken into account, therefore failing to reflect the true synchronization of the user experience.
[0004] 2. Software log timestamp-based comparison method: Insert logs into the playback application to record the timestamps of key events such as decoding start and rendering calls, and compare them on the server side. The problems with this method are: (a) the log timestamps depend on the system clock of each terminal, and the clocks of multiple terminals are usually not synchronized, introducing comparison errors; (b) it still stays at the software level and cannot capture the hardware and driver layer latency between the rendering command being issued and the actual output of the sound card / graphics card.
[0005] 3. Single-point detection method based on external sensors (such as microphones): For example, a high-precision microphone is used to record the played audio, and the playback time is determined by detecting a specific marker (such as a short "beep") in the audio. Although this method touches on the actual physical playback output, it is usually used for single-point testing. When it is extended to multi-terminal group testing, the following problems are faced: (a) Inconsistent time reference: Multiple microphone probes are scattered across various terminals. Without a strictly synchronized clock, the recorded timestamps themselves are biased, leading to the accumulation of group testing errors and making it impossible to distinguish whether the playback is out of sync or the measurement clock is out of sync. (b) Insufficient detection robustness: In the complex acoustic environment of a vehicle (with engine noise, road noise, wind noise, and in-vehicle acoustic reflections and reverberation), a single marker (especially a simple short pulse) is easily submerged by noise or the detection peak position is shifted due to multipath effects, making the detection results unreliable.
[0006] In summary, existing technologies lack a dedicated solution for directly, reliably, quantitatively, and efficiently testing the audio playback synchronization accuracy of multiple KTV terminals in a real, complex, and mobile in-vehicle environment. The industry urgently needs a group testing technology that uses the actual audio beats as a unified anchor point, is resistant to environmental interference, and can unify the time reference of each measurement point. Summary of the Invention
[0007] This invention provides a method and system for testing the audio synchronization of multiple terminals in a vehicle-mounted KTV, in order to address the lack of a dedicated solution in the prior art that can directly, reliably, quantitatively, and efficiently test the audio playback synchronization accuracy of multiple KTV terminals in a real, complex, and mobile vehicle environment.
[0008] To achieve the above objectives, on the one hand, this invention provides a multi-terminal audio synchronization testing method for in-vehicle KTVs. The method includes: S1, simultaneously embedding short pulse markers and encoded markers at each beat point of the test audio to generate marked test audio and sending it to multiple in-vehicle KTV terminals under test for synchronized playback; S2, synchronously acquiring the audio signals actually played by each in-vehicle KTV terminal using microphone probes deployed in the playback environment of each terminal; S3, performing short pulse marker detection and encoded marker detection on each acquired audio signal to obtain the pulse detection timestamp and encoded detection timestamp corresponding to each beat point, as well as pulse detection indicators and encoded detection indicators; S4, generating pulse clarity weights and encoded clarity weights based on the pulse detection indicators and encoded detection indicators; and for... For each detected beat point, the pulse detection timestamp and the code detection timestamp are weighted and fused using the pulse clarity weight and the code clarity weight to obtain the fused timestamp of each beat point; S5, a unified time reference is obtained for each microphone probe, and the fused timestamps of all beat points output by each microphone probe are time-aligned and corrected using the unified time reference; S6, based on the time-aligned and corrected set of fused timestamps from all microphone probes, the time deviation of each non-reference terminal relative to the reference terminal for each beat point is calculated, and audio synchronization calibration parameters for each non-reference terminal are generated based on the time deviation; the reference terminal is a terminal specified from the multiple in-vehicle KTV terminals, and the non-reference terminals are other terminals in the multiple in-vehicle KTV terminals besides the reference terminal.
[0009] Optionally, the pulse detection indicators include: pulse peak sharpness, pulse local signal-to-noise ratio, and pulse peak-side ratio; the coding detection indicators include: coding peak sharpness, coding local signal-to-noise ratio, and coding peak-side ratio; generating pulse sharpness weights and coding sharpness weights based on the pulse detection indicators and coding detection indicators includes: normalizing the pulse peak sharpness, pulse local signal-to-noise ratio, pulse peak-side ratio, coding peak sharpness, coding local signal-to-noise ratio, and coding peak-side ratio; taking a weighted average of the normalized pulse peak sharpness, pulse local signal-to-noise ratio, and pulse peak-side ratio to obtain a pulse sharpness score; taking a weighted average of the normalized coding peak sharpness, coding local signal-to-noise ratio, and coding peak-side ratio to obtain a coding sharpness score; and calculating the pulse sharpness weights and coding sharpness weights based on the pulse sharpness score and coding sharpness score.
[0010] Optionally, S3 includes: removing the DC component from each acquired audio signal to obtain all preprocessed audio signals; applying a bandpass filter to each preprocessed audio signal and retaining the first frequency band to obtain all pulse-filtered signals; applying a bandpass filter to each preprocessed audio signal and retaining the second frequency band to obtain all encoded-filtered signals; performing short-pulse marker detection on each pulse-filtered signal to obtain the pulse detection timestamp corresponding to each beat point, and the pulse cross-correlation data during the detection process; extracting indicators from the pulse cross-correlation data to obtain pulse detection indicators; performing encoding marker detection on each encoded-filtered signal to obtain the encoding detection timestamp corresponding to each beat point, and the encoding cross-correlation data during the detection process; extracting indicators from the encoding cross-correlation data to obtain encoding detection indicators.
[0011] Optionally, the unified time reference is provided in one of the following ways: the second pulse (PPS) signal provided by the GPS module of the central server is distributed to each microphone probe; or the central server distributes time synchronization data packets periodically broadcast over the network to each microphone probe.
[0012] Optionally, generating audio synchronization calibration parameters for each non-reference terminal based on the time deviation includes: calculating the average time deviation, 95th percentile time deviation, and standard deviation of the time deviation of the current non-reference terminal based on the reference terminal at each beat point; and generating calibration parameters for the current non-reference terminal based on the average time deviation, 95th percentile time deviation, and standard deviation of the time deviation of the current non-reference terminal based on the reference terminal.
[0013] On the other hand, the present invention provides a multi-terminal audio synchronization testing system for in-vehicle KTVs. The system includes: a marker generation and playback unit, used to simultaneously embed short pulse markers and coded markers at each beat point of the test audio, generating marked test audio and sending it to multiple in-vehicle KTV terminals under test for synchronized playback; a microphone probe acquisition unit, used to synchronously acquire the audio signals actually played by each in-vehicle KTV terminal through microphone probes deployed in the playback environment of each in-vehicle KTV terminal; a dual-marker detection unit, used to perform short pulse marker detection and coded marker detection for each acquired audio signal, obtaining the pulse detection timestamp and coded detection timestamp corresponding to each beat point, as well as obtaining the pulse detection index and coded detection index; and a fusion unit, used to generate pulse clarity weights and coded clarity weights based on the pulse detection index and coded detection index. The system comprises the following components: a pulse detection timestamp and a code detection timestamp; a time alignment correction unit; a time alignment correction unit; a calibration suggestion generation unit; a time deviation of each non-reference terminal relative to the reference terminal; and an audio synchronization calibration parameter for each non-reference terminal. The reference terminal is a terminal designated from among the multiple in-vehicle KTV terminals, and the non-reference terminals are other terminals among the multiple in-vehicle KTV terminals besides the reference terminal.
[0014] Optionally, the pulse detection indicators include: pulse peak sharpness, pulse local signal-to-noise ratio, and pulse peak-side ratio; the coding detection indicators include: coding peak sharpness, coding local signal-to-noise ratio, and coding peak-side ratio; the generation of pulse sharpness weights and coding sharpness weights based on the pulse detection indicators and coding detection indicators includes: a normalization subunit for normalizing the pulse peak sharpness, pulse local signal-to-noise ratio, pulse peak-side ratio, coding peak sharpness, coding local signal-to-noise ratio, and coding peak-side ratio; a first weighted averaging subunit for weighted averaging the normalized pulse peak sharpness, pulse local signal-to-noise ratio, and pulse peak-side ratio to obtain a pulse sharpness score; a second weighted averaging subunit for weighted averaging the normalized coding peak sharpness, coding local signal-to-noise ratio, and coding peak-side ratio to obtain a coding sharpness score; and a weight calculation subunit for calculating the pulse sharpness weights and coding sharpness weights based on the pulse sharpness score and the coding sharpness score.
[0015] Optionally, the dual-marker detection unit includes: a removal subunit, used to remove the DC component from each acquired audio signal to obtain all preprocessed audio signals; a filtering subunit, used to apply a bandpass filter to each preprocessed audio signal and retain the first frequency band to obtain all pulse-filtered signals; apply a bandpass filter to each preprocessed audio signal and retain the second frequency band to obtain all encoded-filtered signals; a pulse mark detection subunit, used to perform short-pulse mark detection on each pulse-filtered signal to obtain the pulse detection timestamp corresponding to each beat point, and the pulse cross-correlation data during the detection process; extracting an index from the pulse cross-correlation data to obtain a pulse detection index; and an encoded mark detection subunit, used to perform encoded mark detection on each encoded-filtered signal to obtain the encoded detection timestamp corresponding to each beat point, and the encoded cross-correlation data during the detection process; extracting an index from the encoded cross-correlation data to obtain an encoded detection index.
[0016] Optionally, the unified time reference is provided in one of the following ways: the second pulse (PPS) signal provided by the GPS module of the central server is distributed to each microphone probe; or the central server distributes time synchronization data packets periodically broadcast over the network to each microphone probe.
[0017] Optionally, the calibration suggestion generation unit includes: a time deviation calculation subunit, used to calculate the time deviation of each non-reference terminal based on the reference terminal at each beat point based on the fused timestamp set from all microphone probes after time alignment correction; and an audio synchronization calibration subunit, used to perform audio synchronization calibration on each non-reference terminal in the following manner: calculating the average time deviation, 95th percentile time deviation, and standard deviation of the time deviation of the current non-reference terminal based on the reference terminal based on the time deviation of each beat point of the current non-reference terminal based on the reference terminal; and generating calibration parameters for the current non-reference terminal based on the average time deviation, 95th percentile time deviation, and standard deviation of the time deviation of the current non-reference terminal based on the reference terminal.
[0018] The beneficial effects of this invention are: This invention provides a multi-terminal audio synchronization testing method and system for in-vehicle KTVs. This method abandons the traditional approach of using network or software events as a benchmark, pioneering the use of music beats—a direct carrier of user experience—as the ultimate reference for synchronization testing, thus realizing a paradigm shift from "data transmission synchronization" to "user experience synchronization." It proposes a dual-marking strategy of "short pulse + coded marker." Short pulse markers are easy to locate at the sub-sampling level in good environments; coded markers, utilizing their excellent autocorrelation properties, have strong anti-interference capabilities and accurate time resolution in noisy and reverberant environments. The two complement each other, and subsequent weighted fusion ensures reliable and accurate beat arrival times in various complex acoustic environments. By introducing a unified time reference, it solves the inherent clock asynchrony problem in distributed measurements, enabling fair comparison of measurement data from dozens or even more dispersed probes on an absolutely unified time scale, thereby accurately identifying the true playback synchronization deviation. Attached Figure Description
[0019] Figure 1 This is a flowchart of a multi-terminal audio synchronization testing method for in-vehicle KTV provided in an embodiment of the present invention; Figure 2 This is a schematic diagram of the structure of a multi-terminal audio synchronization testing system for in-vehicle KTV provided in an embodiment of the present invention. Detailed Implementation
[0020] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this invention, and not all of them. Based on the embodiments of this invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this invention.
[0021] Figure 1 This is a flowchart of a multi-terminal audio synchronization testing method for in-vehicle KTV provided by an embodiment of the present invention; as follows: Figure 1 As shown, the method includes: S1. Simultaneously embed short pulse markers and encoded markers at each beat point of the test audio to generate marked test audio and send it to multiple in-vehicle KTV terminals under test for synchronous playback; The operator configures the system at the test control center. First, they select or design the music tracks for the test (i.e., the test audio). For a comprehensive evaluation, three types of test tracks are typically prepared: Type A (Steady-state accuracy test): Select songs with a stable rhythm and slow tempo (e.g., ballads at 60-80 BPM). The longer beat intervals facilitate observation of the synchronization capabilities of multiple vehicle-mounted KTV terminals under steady-state conditions.
[0022] Type B (Dynamic Response Test): Select songs with fast tempos and dense beats (e.g., dance music at 120-140 BPM). Dense beat events are used to test the response accuracy and data processing capabilities of multiple in-vehicle KTV terminals under high event rates.
[0023] Type C (Adaptability Test): Select songs with varying tempos (such as crescendo, crescendo, or free tempo) to test the synchronization algorithm's ability to track and adapt to tempo changes.
[0024] After selecting the test audio, the music beat analysis tool is used to automatically identify the beat times of the music; for example: beat 1 time: 0.000 seconds, beat 2 time: 0.500 seconds, beat 3 time: 1.000 seconds, ...; then, short pulse markers and coded markers are embedded in each music beat point in the test audio waveform. Short pulse marker: An audio pulse with extremely short duration (typically ≤2 milliseconds) and concentrated energy. For example, it could be a sine wave segment truncated by a window function (such as the Hanning window), with a center frequency selectable within the 4kHz range (this frequency band is within the range where the human ear is sensitive but not harsh). Its characteristics include a sharp time-domain waveform, which can be located with sub-sampling interval accuracy using algorithms such as cross-correlation in quiet environments; its energy is 10-15dB higher than background music (ensuring it can be detected without affecting the listening experience).
[0025] Encoded marker: A pseudo-random coded audio segment with good autocorrelation properties. Golay complementary sequence pairs are preferred because their autocorrelation function exhibits perfect single-peak characteristics and zero sidelobes. The duration of the encoded marker is relatively long; for example, at a 48kHz sampling rate, the sequence length can be selected as 512, 1024, or 2048 samples (corresponding to approximately 10.7ms, 21.3ms, or 42.7ms, respectively). Its modulation method involves modulating a binary sequence onto a 6kHz carrier (staggered from the pulse frequency to avoid interference). It features good noise and multipath (reverberation) immunity; even at low signal-to-noise ratios, the location of the correlation peak can be accurately found through matched filtering; its energy is comparable to that of a short pulse.
[0026] The insertion strategy is as follows: at each beat point, short pulse markers and coded markers are inserted temporally adjacently or with slight overlap, together forming dual evidence of a "beat event." Marked test audio is then generated.
[0027] At the start of the test, the playback control module sends instructions via the network to all the in-vehicle KTV terminals under test, requiring them to begin playing the specified test audio at the same absolute time (e.g., a full second after receiving the instruction). Each in-vehicle KTV terminal attempts to start playback simultaneously based on its own synchronization mechanism.
[0028] S2. By deploying microphone probes in the playback environment of each vehicle-mounted KTV terminal, the audio signals actually played by each vehicle-mounted KTV terminal are collected synchronously. In an optional implementation, microphone probes deployed in each vehicle (placed near the driver's headrest to simulate the user's listening position) are already activated and ready for use. Each probe is an independent hardware device, whose core components include: A high-precision analog-to-digital converter (ADC) that continuously samples at a fixed sampling rate (e.g., 48 kHz).
[0029] A real-time clock (RTC) with a highly stable crystal oscillator, synchronized and continuously calibrated using a uniform time reference source. Internally, each acquired audio sample block is timestamped with the local clock.
[0030] A processor and memory are used to run simple pre-processed firmware and cache data.
[0031] A wireless communication module (such as a Wi-Fi or cellular module) is used to upload the collected audio data and timestamp information to the data analysis server.
[0032] When the microphone probe detects that the car speakers have started playing test audio, it continuously records the audio stream and transmits the timestamped audio data to the data analysis server in real-time or near real-time via a wireless network. When bandwidth is limited, the probe can also cache the data locally and upload it all at once after the test is complete.
[0033] After receiving audio data from all microphone probes, the data analysis server activates the dual-marker detection unit for offline or online analysis. The detection process is performed independently for each audio signal from a specific probe.
[0034] S3. For each acquired audio signal, perform short pulse marker detection and code marker detection respectively to obtain the pulse detection timestamp and code detection timestamp corresponding to each beat point, as well as the pulse detection index and code detection index. In an optional implementation, S3 includes: S31. Remove the DC component from each acquired audio signal to obtain all preprocessed audio signals. For each acquired audio signal, the average value of the entire audio segment is calculated. For each acquired audio signal, it is sampled at a fixed sampling rate. The average value is subtracted from each sampling point to eliminate the DC offset that may be introduced by the microphone. All preprocessed audio signals are obtained one by one.
[0035] S32. Apply a bandpass filter to each preprocessed audio signal and retain the first frequency band to obtain all pulse filtered signals; apply a bandpass filter to each preprocessed audio signal and retain the second frequency band to obtain all encoded filtered signals. The first frequency band is the 2kHz-6kHz band; the second frequency band is the 5kHz-8kHz band; a bandpass filter is applied to each preprocessed audio signal to filter and retain the 2kHz-6kHz band in order to retain the main energy of the short pulses and filter out low-frequency engine noise and high-frequency wind noise, thus obtaining all pulse filtered signals; a bandpass filter is applied to each preprocessed audio signal to filter and retain the 5kHz-8kHz band in order to retain the main energy of the encoded markers, thus obtaining all encoded filtered signals.
[0036] S33. Perform short pulse marker detection on each pulse filter signal to obtain the pulse detection timestamp corresponding to each beat point, as well as the pulse cross-correlation data during the detection process; extract the index from the pulse cross-correlation data to obtain the pulse detection index; The precise timing of the short pulse markers is detected from each pulse-filtered signal. Specifically, taking a single pulse-filtered signal as an example: The pulse-filtered signal is divided into small segments, and the energy of each segment is calculated. ,in, The energy of the current segment, The sampling points contained in this segment are N consecutive sampling points. The energy of each segment is arranged in time sequence to obtain the total energy envelope. The energy envelope is cross-correlated with a pre-stored pulse energy template (the energy distribution sequence corresponding to a typical pulse). Specifically, the pre-stored pulse energy template is slid across the energy envelope in time sequence, and the similarity between the pulse energy template and the corresponding segment of the energy envelope is calculated at each sliding position. When the similarity reaches the maximum, the corresponding time position is determined as the coarse time position of the short pulse marker in the energy envelope. The peak position of the cross-correlation result is the coarse time of the short pulse marker. The accuracy of the peak position of the cross-correlation result is limited by the sampling interval (approximately 20.8 microseconds at 48kHz). To obtain higher accuracy: take the peak point and its two adjacent points, a total of 3 points, and fit a parabola with these 3 points. The vertex of the parabola is the more accurate peak time, which can improve the accuracy to about 5 microseconds; not all peaks are true short pulse markers and need to be verified. The verification conditions are: (1) peak height > background mean (the background mean is the background energy baseline obtained by statistically averaging the remaining energy values after excluding short pulse candidate peak segments in the energy envelope) 5 times (to ensure it is not noise); (2) the deviation between the peak and the expected beat interval (the expected beat interval is the time interval between adjacent beat points determined in advance based on the beat analysis results of the test audio) < 50ms (to exclude abnormal peaks); (3) peak half-width < 5ms (to ensure the peak is sharp enough); only the peaks that pass the verification are recorded as valid pulse times. The final output pulse detection timestamp { , , ,……}, where each value is a pulse detection timestamp for a beat.
[0037] The cross-correlation results during the short pulse marker detection process, which are also intermediate data in the detection process, are used to extract indices to obtain pulse detection indices.
[0038] S34. Perform encoding mark detection on each encoded filter signal to obtain the encoding detection timestamp corresponding to each beat point, as well as the encoding cross-correlation data during the detection process; extract the index from the encoding cross-correlation data to obtain the encoding detection index.
[0039] The precise time of the coded marker is detected from each coded filter signal. Specifically, taking a single coded filter signal as an example: The signal is cross-correlated with a pre-stored Golay sequence template. The peak position of the cross-correlation result is the time of the encoded marker. A special property of the Golay sequence is that the sidelobes (peaks other than the main peak) of its autocorrelation function cancel each other out, leaving only a sharp main peak. This makes the detection highly reliable and unaffected by false peaks caused by noise. Furthermore, due to potential slight sampling rate deviations in in-vehicle systems, the playback speed of the encoded marker's time series may differ slightly from the expected speed. Within the range, the Golay sequence template is time-scaled, for example: generating a Golay sequence template at 0.98x speed, generating a Golay sequence template at 0.99x speed, generating a Golay sequence template at 1.00x speed, generating a Golay sequence template at 1.01x speed, generating a Golay sequence template at 1.02x speed, performing cross-correlation operation between the signal and each Golay sequence template, and taking the cross-correlation result with the highest peak. Not all peaks are true coding markers and need to be verified. The verification conditions are: (1) Cross-correlation peak value > noise floor (in this embodiment, noise floor is used to characterize the correlation baseline level generated by background noise and non-coded signals in the cross-correlation result, which can be obtained by excluding the main peak segment of the coding marker in the cross-correlation result and then statistically analyzing the remaining correlation values); (2) Peak-to-side lobe ratio (main peak height / highest side lobe height) > 3; Only the peaks that pass the verification are recorded as valid coding time. The final output coding detection timestamp { , , Each value is a time stamp encoded as a beat.
[0040] The cross-correlation results during the coding and tag detection process, which are also intermediate data in the detection process, are used to extract indicators to obtain coding detection indicators.
[0041] S4. For each detected beat point, generate a pulse sharpness weight and a code sharpness weight based on the pulse detection index and the code detection index; use the pulse sharpness weight and the code sharpness weight to perform a weighted fusion of the pulse detection timestamp and the code detection timestamp to obtain the fused timestamp of each beat point; In one optional implementation, the pulse detection metrics include: pulse peak sharpness, pulse local signal-to-noise ratio, and pulse peak-side ratio; the coding detection metrics include: coding peak sharpness, coding local signal-to-noise ratio, and coding peak-side ratio. Peak sharpness is generated as follows: Find the position of the maximum amplitude in the cross-correlation result and record it as the peak height H; scan from the peak to the left and right sides and find the position where the height drops to H / 2 as the half height point; calculate the distance between the two half height points and record it as the half height width W (unit: number of sampling points), and calculate the peak sharpness S=H / W.
[0042] The local signal-to-noise ratio is generated as follows: The peak region is defined as the area within 50 sampling points to the left and right of the peak position. The average signal energy within the peak region is calculated and denoted as E_signal. The background region is defined as the area beyond 200 sampling points from the peak. The average signal energy within the background region is calculated and denoted as E_noise. The local signal-to-noise ratio (SNR) is calculated as: SNR = 10. log10 (E_signal / E_noise), in dB.
[0043] Peak-side ratios are generated as follows: Find the maximum peak value in the cross-correlation results and denote it as the main peak height H_main. Set the main peak region (20 sampling points on each side of the main peak) to zero. Find the maximum value in the remaining data and denote it as the highest side lobe height H_side. Calculate the peak-side ratio PSR = H_main / H_side.
[0044] The generation of pulse sharpness weights and coded sharpness weights based on pulse detection indicators and coded detection indicators includes: Normalize the pulse peak sharpness, pulse local signal-to-noise ratio, pulse peak-side ratio, coded peak sharpness, coded local signal-to-noise ratio, and coded peak-side ratio; The pulse sharpness score is obtained by weighting the normalized pulse peak sharpness, pulse local signal-to-noise ratio, and pulse peak-side ratio. The coding sharpness score is obtained by weighting the normalized coding peak sharpness, coding local signal-to-noise ratio, and coding peak-side ratio. The pulse clarity weight and the code clarity weight are calculated based on the pulse clarity score and the code clarity score.
[0045] The following is an example to illustrate this: The pulse peak sharpness, pulse local signal-to-noise ratio, and pulse peak-side ratio are respectively: S_A=20, SNR_A=20dB, PSR_A=5; The coding peak sharpness, coding local signal-to-noise ratio, and coding peak-side ratio are respectively: S_B=8, SNR_B=10dB, PSR_B=1.33; For each indicator type, normalization is performed based on an empirical range: Normalization formula: Normalized value = (original value - minimum value) / (maximum value - minimum value); if the normalized value < 0, set it to 0, and if the normalized value > 1, set it to 1.
[0046] Empirical ranges for each indicator: Peak sharpness S: The empirical minimum value is 2, and the empirical maximum value is 50; Local signal-to-noise ratio (SNR): The empirical minimum is 5dB, and the empirical maximum is 30dB; Peak-side ratio (PSR): The empirical minimum is 1.5, and the empirical maximum is 10.
[0047] Normalization of pulse detection: S'_A=(20-2) / (50-2)=18 / 48=0.375; SNR'_A=(20-5) / (30-5)=15 / 25=0.6; PSR'_A=(5-1.5) / (10-1.5)=3.5 / 8.5=0.412; Normalization of encoding detection: S'_B=(8-2) / (50-2)=6 / 48=0.125; SNR'_B=(10-5) / (30-5)=5 / 25=0.2; PSR'_B=(1.33-1.5) / (10-1.5)=-0.17 / 8.5=-0.02 → Set to 0; Pulse sharpness score = α×S'_A + β×SNR'_A + γ×PSR'_A Encoding clarity score = α×S'_B + β×SNR'_B + γ×PSR'_B Where α = weighting coefficient, representing the importance of sharpness; β = weighting coefficient, representing the importance of signal-to-noise ratio; γ = weighting coefficient, representing the importance of peak-side ratio; default setting: α=β=γ=1 / 3 (all three indicators are equally important); Pulse clarity score: Score_A=1 / 3×0.375+1 / 3×0.6+1 / 3×0.412=0.462 Encoding clarity score: Score_B=1 / 3×0.125+1 / 3×0.2+1 / 3×0=0.109 Total score = Sum of pulse clarity score and coded clarity score; The weight of each test result = the sharpness score of that result / the total score; Total score = Score_A + Score_B = 0.462 + 0.109 = 0.571 Pulse sharpness weight: w_A = 0.462 / 0.571 = 0.809; Encoding clarity weight: w_B = 0.109 / 0.571 = 0.191; Verification: w_A + w_B = 0.809 + 0.191 = 1.0 For each detected beat point, the pulse detection timestamp and the code detection timestamp are weighted and fused using the pulse sharpness weight and the code sharpness weight to obtain the fused timestamp of each beat point; In one optional implementation, taking one of the beat points as an example: the pulse detection timestamp is =10.052 seconds, =10.058 seconds, w_A=0.809, w_B=0.191; The fusion timestamp captured by this node = w_A× + w_B× =0.809×10.052+0.191×10.058=10.053 seconds.
[0048] The simple average result is (10.052 + 10.058) / 2 = 10.055 seconds; the fusion result of this invention is 10.053 seconds. Because the sharpness of impulse detection is much higher than that of coded detection (0.462 vs 0.109), the fusion result (10.053 seconds) is closer to the result of impulse detection (10.052 seconds) than the median value of the simple average (10.055 seconds). This is precisely the effect of "high-quality results receiving greater weight".
[0049] S5. Obtain a unified time reference for each microphone probe, and use this unified time reference to perform time alignment correction on the fused timestamps of all beat points output by each microphone probe. All microphone probes use the same time reference to record data, eliminating discrepancies in the internal clocks of each microphone probe.
[0050] In one alternative implementation, the unified time base is provided in one of the following ways: The GPS module of the central server provides a pulse-per-second (PPS) signal, which is distributed to each microphone probe. (1) The GPS module of the central server receives satellite signals; (2) The GPS module outputs a pulse signal (PPS) every second, and the rising edge of the pulse corresponds precisely to the UTC integer second. (3) Distribute the PPS signal to all microphone probes (via wired or wireless means). (4) When each microphone probe receives a PPS pulse, it immediately sets its local clock to the current UTC whole second; (5) Repeat the above process and continue to calibrate.
[0051] The central server distributes time synchronization data packets, which are periodically broadcast over the network, to each microphone probe.
[0052] (1) The central server records the current precise time T1 and sends a correction packet to the microphone probe; (2) The microphone probe receives the calibration packet and records the time T2 displayed on the local clock; (3) The microphone probe immediately sends an acknowledgment packet back to the central server; (4) The central server receives the confirmation packet and records the current time T3; (5) Calculate the network round-trip delay: RTT = T3 - T1; (6) Calculate one-way delay: One-way delay = RTT / 2; (7) Calculate the microphone probe clock deviation: Deviation = T2 - (T1 + one-way delay); (8) The microphone probe corrects the local clock based on the deviation value.
[0053] Furthermore, clock calibration is not a one-time fix. The crystal oscillator used by the microphone probe will drift due to temperature changes (typically 0.1-1 microseconds per second).
[0054] The procedure is as follows: A time synchronization operation is performed every 10 seconds throughout the entire test, and the clock deviation value is recorded each time. If the deviation exceeds a threshold (e.g., >0.5ms), it is corrected immediately. All microphone probe clocks are calibrated to a unified time reference and remain synchronized throughout the test.
[0055] Time series of each probe's beat: Probe 1 (Car 1): {t1_1, t1_2, t1_3, ...} Probe 2 (Car 2): {t2_1, t2_2, t2_3, ...} Probe 3 (Car 3): {t3_1, t3_2, t3_3, ...} (Assuming 3 vehicles are tested, and 3 cycles are monitored): Probe 1: {10:00:00.000, 10:00:00.500, 10:00:01.000} Probe 2: {10:00:00.006, 10:00:00.507, 10:00:01.005} Probe 3: {09:59:59.997, 10:00:00.498, 10:00:00.996} S6. Based on the fused timestamp set from all microphone probes after time alignment correction, calculate the time deviation of each non-reference terminal relative to the reference terminal for each beat point, and generate audio synchronization calibration parameters for each non-reference terminal based on the time deviation; the reference terminal is a terminal designated from the plurality of vehicle-mounted KTV terminals, and the non-reference terminals are other terminals in the plurality of vehicle-mounted KTV terminals besides the reference terminal.
[0056] In an optional implementation, generating audio synchronization calibration parameters for each non-reference terminal based on the time deviation includes: The average time deviation, 95th percentile time deviation, and standard deviation of time deviation of the current non-reference terminal based on the reference terminal are calculated based on the time deviation of each beat point of the current non-reference terminal based on the reference terminal. Calibration parameters for the current non-reference terminal are generated based on the average time deviation, 95th percentile time deviation, and standard deviation of the time deviation of the reference terminal.
[0057] Select a vehicle-mounted KTV terminal as a reference (usually the one with the smallest number, i.e., vehicle-mounted KTV terminal 1). For each beat point i and each non-reference vehicle-mounted KTV terminal j, calculate: Time deviation (non-reference vehicle-mounted KTV terminal j, beat i) = beat i time of vehicle-mounted KTV terminal j - beat i time of vehicle-mounted KTV terminal 1 Time deviation of beat 1: Time deviation of vehicle-mounted KTV terminal 2 = 10:00:00.006 - 10:00:00.000 = +6ms (vehicle-mounted KTV terminal 2 is 6ms slower than vehicle-mounted KTV terminal 1) Time deviation of vehicle-mounted KTV terminal 3 = 09:59:59.997 - 10:00:00.000 = -3ms (Vehicle-mounted KTV terminal 3 is 3ms faster than vehicle-mounted KTV terminal 1) Time deviation of beat 2: Time deviation of vehicle-mounted KTV terminal 2 = 10:00:00.507 - 10:00:00.500 = +7ms Time deviation of vehicle-mounted KTV terminal 3 = 10:00:00.498 - 10:00:00.500 = -2ms Time deviation of beat 3: Time deviation of vehicle-mounted KTV terminal 2 = 10:00:01.005 - 10:00:01.000 = +5ms Time deviation of vehicle-mounted KTV terminal 3 = 10:00:80.996 - 10:00:01.000 = -4ms For each non-reference terminal, calculate the following statistics: Mean deviation = average deviation of all beat deviations; 95th percentile deviation = the value at the 95th percentile after sorting all deviations by absolute value; Jitter = Standard deviation of the deviation series (reflecting the degree of fluctuation of the deviation) = sqrt(Σ(deviation i - mean deviation)) 2 / N); A positive deviation (+) indicates that the playback time of the in-vehicle KTV terminal is later than that of the reference terminal, i.e., "playback is too slow"; a negative deviation (-) indicates that the playback time of the in-vehicle KTV terminal is earlier than that of the reference terminal, i.e., "playback is too fast".
[0058] The final result is: Time deviation sequence of vehicle-mounted KTV terminal 2: {+6ms, +7ms, +5ms} Mean deviation = (6 + 7 + 5) / 3 = +6.0 ms Deviation 95th percentile ≈ +7ms Jitter = sqrt(((6-6)) 2 +(7-6) 2 +(5-6) 2 ) / 3)=sqrt((0+1+1) / 3)= 0.82ms Time deviation sequence of vehicle-mounted KTV terminal 3: {-3ms, -2ms, -4ms} Mean deviation = (-3 - 2 - 4) / 3 = -3.0 ms Deviation 95th percentile = -4ms jitter = sqrt(((-3-(-3)) 2 +(-2-(-3)) 2 +(-4-(-3)) 2 ) / 3)= sqrt((0+1+1) / 3)=0.82ms Group measurement deviation vector = [ (Vehicle-mounted KTV terminal 2, mean = +6.0ms, 95th percentile = +7ms, jitter = 0.82ms) (Vehicle-mounted KTV terminal 3, mean = -3.0ms, 95th percentile = -4ms, jitter = 0.82ms) Based on the group measurement deviation vector, generate audio synchronization calibration parameters for each non-reference vehicle: Vehicle-mounted KTV terminal 2: Average deviation +6ms (playback is slow), it is recommended to reduce the audio buffer by 6ms; The 95th percentile deviation (+7ms) is close to the mean (+6ms) and has very little jitter (0.82ms), indicating stable latency without extreme anomalies. Maintain the existing audio buffer size; Jitter: Extremely low jitter level (0.82ms), no need to enable or adjust anti-jitter algorithm.
[0059] Vehicle-mounted KTV terminal 3: Average deviation -3ms (playback is too fast), it is recommended to add an audio buffer of 3ms; The deviation at the 95th percentile (-4ms) is close to the mean (-3ms) and the jitter is very small (0.82ms), indicating that the playback is advanced and stable, and the existing audio buffer size can be maintained. Jitter: Extremely low jitter level (0.82ms), no need to enable or adjust anti-jitter algorithm.
[0060] This invention uses the beat points of the music itself as natural and perceptible synchronization event anchors. Through innovative "audio dual-labeling" technology and "event fusion" algorithms, it robustly captures the actual playback times of these anchors in complex acoustic environments. It then uses a unified time reference to align the dispersed measurement points in time. Finally, through group measurement statistical analysis, it accurately quantifies the synchronization deviation between multiple terminals in a mobile karaoke system and provides specific calibration guidance. This invention achieves multi-terminal audio synchronization in mobile karaoke systems.
[0061] Figure 2 This is a schematic diagram of the structure of a multi-terminal audio synchronization testing system for in-vehicle KTV provided in an embodiment of the present invention, as shown below. Figure 2 As shown, the system includes: The marker generation and playback unit 201 is used to simultaneously embed short pulse markers and coded markers at each beat point of the test audio, generate tagged test audio, and send it to multiple in-vehicle KTV terminals under test for synchronous playback. The microphone probe acquisition unit 202 is used to synchronously acquire the audio signals actually played by each vehicle-mounted KTV terminal through the microphone probes deployed in the playback environment of each vehicle-mounted KTV terminal. The dual-marker detection unit 203 is used to perform short pulse mark detection and code mark detection for each acquired audio signal, to obtain the pulse detection timestamp and code detection timestamp corresponding to each beat point, as well as the pulse detection index and code detection index. The fusion unit 204 is used to generate pulse sharpness weight and code sharpness weight based on pulse detection index and code detection index; for each detected beat point, the pulse detection timestamp and code detection timestamp are weighted and fused using the pulse sharpness weight and code sharpness weight to obtain the fused timestamp of each beat point; Alignment correction unit 205 is used to obtain a unified time reference for each microphone probe and use the unified time reference to perform time alignment correction on the fused timestamps of all beat points output by each microphone probe. The calibration suggestion generation unit 206 is used to calculate the time deviation of each non-reference terminal relative to the reference terminal based on the fused timestamp set from all microphone probes after time alignment correction, and to generate audio synchronization calibration parameters for each non-reference terminal based on the time deviation; the reference terminal is a terminal designated from the plurality of vehicle-mounted KTV terminals, and the non-reference terminals are other terminals in the plurality of vehicle-mounted KTV terminals other than the reference terminal.
[0062] In one optional implementation, the pulse detection metrics include: pulse peak sharpness, pulse local signal-to-noise ratio, and pulse peak-side ratio; the coding detection metrics include: coding peak sharpness, coding local signal-to-noise ratio, and coding peak-side ratio. The generation of pulse sharpness weights and coded sharpness weights based on pulse detection indicators and coded detection indicators includes: The normalization subunit 2041 is used to normalize the pulse peak sharpness, pulse local signal-to-noise ratio, pulse peak-side ratio, coded peak sharpness, coded local signal-to-noise ratio, and coded peak-side ratio. The first weighted average subunit 2042 is used to perform a weighted average of the normalized pulse peak sharpness, pulse local signal-to-noise ratio, and pulse peak-side ratio to obtain the pulse sharpness score. The second weighted average subunit 2043 is used to perform a weighted average of the normalized coding peak sharpness, coding local signal-to-noise ratio, and coding peak-side ratio to obtain the coding sharpness score. The weight calculation subunit 2044 is used to calculate the pulse sharpness weight and the coding sharpness weight based on the pulse sharpness score and the coding sharpness score.
[0063] In an optional implementation, the dual-label detection unit 203 includes: The removal subunit 2031 is used to remove the DC component from each acquired audio signal, so that all preprocessed audio signals are obtained one by one. The filtering subunit 2032 is used to apply a bandpass filter to each preprocessed audio signal and retain the first frequency band, so as to obtain all pulse filtered signals in a one-to-one correspondence; and to apply a bandpass filter to each preprocessed audio signal and retain the second frequency band, so as to obtain all coded filtered signals in a one-to-one correspondence. The pulse marker detection subunit 2033 is used to perform short pulse marker detection on each pulse filtered signal to obtain the pulse detection timestamp corresponding to each beat point, as well as the pulse cross-correlation data during the detection process; and to extract the index from the pulse cross-correlation data to obtain the pulse detection index. The encoding marker detection subunit 2034 is used to perform encoding marker detection on each encoded filter signal, obtain the encoding detection timestamp corresponding to each beat point, and the encoding cross-correlation data during the detection process; extract the index from the encoding cross-correlation data to obtain the encoding detection index.
[0064] In one alternative implementation, the unified time base is provided in one of the following ways: The GPS module of the central server provides a pulse-per-second (PPS) signal, which is distributed to each microphone probe. The central server distributes time synchronization data packets, which are periodically broadcast over the network, to each microphone probe.
[0065] In an optional implementation, the calibration recommendation generation unit 206 includes: The time deviation calculation subunit 2061 is used to calculate the time deviation of each non-reference terminal based on the reference terminal based on the fused timestamp set from all microphone probes after time alignment correction. The audio synchronization calibration subunit 2062 is used to perform audio synchronization calibration on each non-reference terminal in the following manner: The average time deviation, 95th percentile time deviation, and standard deviation of time deviation of the current non-reference terminal based on the reference terminal are calculated based on the time deviation of each beat point of the current non-reference terminal based on the reference terminal. Calibration parameters for the current non-reference terminal are generated based on the average time deviation, 95th percentile time deviation, and standard deviation of the time deviation of the reference terminal.
[0066] The system of the present invention corresponds to the method described above, and the specific implementation of the system will not be repeated here.
[0067] The beneficial effects of this invention are: This invention provides a multi-terminal audio synchronization testing method and system for in-vehicle KTVs. This method abandons the traditional approach of using network or software events as a benchmark, pioneering the use of music beats—a direct carrier of user experience—as the ultimate reference for synchronization testing, thus realizing a paradigm shift from "data transmission synchronization" to "user experience synchronization." It proposes a dual-marking strategy of "short pulse + coded marker." Short pulse markers are easy to locate at the sub-sampling level in good environments; coded markers, utilizing their excellent autocorrelation properties, have strong anti-interference capabilities and accurate time resolution in noisy and reverberant environments. The two complement each other, and subsequent weighted fusion ensures reliable and accurate beat arrival times in various complex acoustic environments. By introducing a unified time reference, it solves the inherent clock asynchrony problem in distributed measurements, enabling fair comparison of measurement data from dozens or even more dispersed probes on an absolutely unified time scale, thereby accurately identifying the true playback synchronization deviation.
[0068] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for testing multi-terminal audio synchronization in a vehicle-mounted karaoke system, characterized in that, include: S1. Simultaneously embed short pulse markers and encoded markers at each beat point of the test audio to generate marked test audio and send it to multiple in-vehicle KTV terminals under test for synchronous playback; S2. By deploying microphone probes in the playback environment of each vehicle-mounted KTV terminal, the audio signals actually played by each vehicle-mounted KTV terminal are collected synchronously. S3. For each acquired audio signal, perform short pulse marker detection and code marker detection respectively to obtain the pulse detection timestamp and code detection timestamp corresponding to each beat point, as well as the pulse detection index and code detection index. S4. Generate pulse sharpness weights and code sharpness weights based on pulse detection indicators and code detection indicators; For each detected beat point, the pulse detection timestamp and the code detection timestamp are weighted and fused using the pulse sharpness weight and the code sharpness weight to obtain the fused timestamp of each beat point; S5. Obtain a unified time reference for each microphone probe, and use this unified time reference to perform time alignment correction on the fused timestamps of all beat points output by each microphone probe. S6. Based on the fused timestamp set from all microphone probes after time alignment correction, calculate the time deviation of each non-reference terminal relative to the reference terminal for each beat point, and generate audio synchronization calibration parameters for each non-reference terminal based on the time deviation; the reference terminal is a terminal designated from the plurality of vehicle-mounted KTV terminals, and the non-reference terminals are other terminals in the plurality of vehicle-mounted KTV terminals besides the reference terminal.
2. The method according to claim 1, characterized in that: The pulse detection metrics include: pulse peak sharpness, pulse local signal-to-noise ratio, and pulse peak-side ratio; the encoding detection metrics include: encoding peak sharpness, encoding local signal-to-noise ratio, and encoding peak-side ratio. The generation of pulse sharpness weights and coded sharpness weights based on pulse detection indicators and coded detection indicators includes: Normalize the pulse peak sharpness, pulse local signal-to-noise ratio, pulse peak-side ratio, coded peak sharpness, coded local signal-to-noise ratio, and coded peak-side ratio; The pulse sharpness score is obtained by weighting the normalized pulse peak sharpness, pulse local signal-to-noise ratio, and pulse peak-side ratio. The coding sharpness score is obtained by weighting the normalized coding peak sharpness, coding local signal-to-noise ratio, and coding peak-side ratio. The pulse clarity weight and the code clarity weight are calculated based on the pulse clarity score and the code clarity score.
3. The method according to claim 1, characterized in that, S3 includes: The DC component of each acquired audio signal is removed, and all preprocessed audio signals are obtained one by one. Each preprocessed audio signal is filtered by a bandpass filter and the first frequency band is retained, resulting in a one-to-one correspondence of all pulse filtered signals; each preprocessed audio signal is filtered by a bandpass filter and the second frequency band is retained, resulting in a one-to-one correspondence of all coded filtered signals. Short pulse marker detection is performed on each pulse filtered signal to obtain the pulse detection timestamp corresponding to each beat point, as well as the pulse cross-correlation data during the detection process; the index is extracted from the pulse cross-correlation data to obtain the pulse detection index; For each encoded filtered signal, perform encoded marker detection to obtain the encoded detection timestamp corresponding to each beat point, as well as the encoded cross-correlation data during the detection process; extract the index from the encoded cross-correlation data to obtain the encoded detection index.
4. The method according to claim 1, characterized in that, The unified time base is provided in one of the following ways: The GPS module of the central server provides a pulse-per-second (PPS) signal, which is distributed to each microphone probe. The central server distributes time synchronization data packets, which are periodically broadcast over the network, to each microphone probe.
5. The method according to claim 1, characterized in that, The generation of audio synchronization calibration parameters for each non-reference terminal based on the time deviation includes: The average time deviation, 95th percentile time deviation, and standard deviation of time deviation of the current non-reference terminal based on the reference terminal are calculated based on the time deviation of each beat point of the current non-reference terminal based on the reference terminal. Calibration parameters for the current non-reference terminal are generated based on the average time deviation, 95th percentile time deviation, and standard deviation of the time deviation of the reference terminal.
6. A multi-terminal audio synchronization testing system for in-vehicle KTV, characterized in that, include: The marker generation and playback unit is used to simultaneously embed short pulse markers and coded markers at each beat point of the test audio, generate tagged test audio, and send it to multiple in-vehicle KTV terminals under test for synchronous playback. The microphone probe acquisition unit is used to synchronously acquire the audio signals actually played by each vehicle-mounted KTV terminal through microphone probes deployed in the playback environment of each vehicle-mounted KTV terminal. The dual-marker detection unit is used to perform short pulse mark detection and code mark detection for each acquired audio signal, to obtain the pulse detection timestamp and code detection timestamp corresponding to each beat point, as well as the pulse detection index and code detection index. The fusion unit is used to generate pulse sharpness weights and coded sharpness weights based on pulse detection metrics and coded detection metrics. For each detected beat point, the pulse detection timestamp and the code detection timestamp are weighted and fused using the pulse sharpness weight and the code sharpness weight to obtain the fused timestamp of each beat point; The alignment correction unit is used to obtain a unified time reference for each microphone probe and use this unified time reference to perform time alignment correction on the fused timestamps of all beat points output by each microphone probe. The calibration suggestion generation unit is used to calculate the time deviation of each non-reference terminal relative to the reference terminal based on the fused timestamp set from all microphone probes after time alignment correction, and to generate audio synchronization calibration parameters for each non-reference terminal based on the time deviation; the reference terminal is a terminal designated from the plurality of vehicle-mounted KTV terminals, and the non-reference terminals are other terminals among the plurality of vehicle-mounted KTV terminals besides the reference terminal.
7. The system according to claim 6, characterized in that: The pulse detection metrics include: pulse peak sharpness, pulse local signal-to-noise ratio, and pulse peak-side ratio; the encoding detection metrics include: encoding peak sharpness, encoding local signal-to-noise ratio, and encoding peak-side ratio. The generation of pulse sharpness weights and coded sharpness weights based on pulse detection indicators and coded detection indicators includes: The normalization subunit is used to normalize the pulse peak sharpness, pulse local signal-to-noise ratio, pulse peak-side ratio, coded peak sharpness, coded local signal-to-noise ratio, and coded peak-side ratio. The first weighted average subunit is used to calculate the pulse sharpness score by weighting the normalized pulse peak sharpness, pulse local signal-to-noise ratio, and pulse peak-side ratio. The second weighted average subunit is used to calculate the coding sharpness score by weighted averaging the normalized coding peak sharpness, coding local signal-to-noise ratio, and coding peak-side ratio. The weight calculation subunit is used to calculate the pulse clarity weight and the code clarity weight based on the pulse clarity score and the code clarity score.
8. The system according to claim 6, characterized in that, The dual-label detection unit includes: The removal subunit is used to remove the DC component from each acquired audio signal, so that all preprocessed audio signals can be obtained one by one. The filtering subunit is used to apply a bandpass filter to each preprocessed audio signal and retain the first frequency band, thus obtaining all pulse filtered signals in a one-to-one correspondence; and to apply a bandpass filter to each preprocessed audio signal and retain the second frequency band, thus obtaining all coded filtered signals in a one-to-one correspondence. The pulse marker detection subunit is used to perform short pulse marker detection on each pulse filtered signal to obtain the pulse detection timestamp corresponding to each beat point, as well as the pulse cross-correlation data during the detection process; and to extract the index from the pulse cross-correlation data to obtain the pulse detection index. The coding marker detection subunit is used to perform coding marker detection on each coding filter signal, obtain the coding detection timestamp corresponding to each beat point, and the coding cross-correlation data during the detection process; extract the index from the coding cross-correlation data to obtain the coding detection index.
9. The system according to claim 6, characterized in that, The unified time base is provided in one of the following ways: The GPS module of the central server provides a pulse-per-second (PPS) signal, which is distributed to each microphone probe. The central server distributes time synchronization data packets, which are periodically broadcast over the network, to each microphone probe.
10. The system according to claim 6, characterized in that, The calibration recommendation generation unit includes: The time deviation calculation subunit is used to calculate the time deviation of each non-reference terminal based on the reference terminal for each beat point, based on the fused timestamp set from all microphone probes after time alignment correction. The audio synchronization calibration subunit is used to perform audio synchronization calibration on each non-reference terminal in the following manner: The average time deviation, 95th percentile time deviation, and standard deviation of time deviation of the current non-reference terminal based on the reference terminal are calculated based on the time deviation of each beat point of the current non-reference terminal based on the reference terminal. Calibration parameters for the current non-reference terminal are generated based on the average time deviation, 95th percentile time deviation, and standard deviation of the time deviation of the reference terminal.