Intercom channel quality assurance method and system based on bidirectional audio stream delay jitter
Patent Information
- Application Number
- CN202611085546.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-21
- Publication Date
- 2026-09-29
AI Technical Summary
[0003]但门店现场的对讲终端长期处于7×24小时连续运行状态,拾音硬件会随服役时间增长出现物理性能衰减,叠加门店内制冷设备启停、外界传入等多变的环境噪声,现有系统无法准确区分音质下降是源于网络传输抖动,还是源于硬件状态变化或现场环境干扰,频繁出现误判--不需要切换时盲目切换、需要切换时响应滞后,导致对讲服务异常中断,依赖人工排查又耗时较长,直接影响无人门店的服务连续性与运行安全性
本申请通过获取目标终端的采集设备的阻抗衰减特征对目标环境的第一噪声数据进行处理得到第二噪声数据,再结合目标终端对应音频数据的抖动特征生成目标处理指令,从而能够基于硬件物理状态校正环境噪声的检测结果,区分硬件衰减、环境噪声与网络抖动对对讲音质的不同影响,解决了现有技术中对讲通道质量判定易受硬件老化与现场噪声干扰、导致自愈动作误触发或响应滞后的问题,有效提升了对讲通道质量判定的准确性,减少不必要的通道切换,保障对讲服务的连续性与运行稳定性。
Smart Images

Figure CN122845573A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of remote audio and video intercom technology, specifically to a method and system for ensuring the quality of intercom channels based on bidirectional audio stream delay jitter. Background Technology
[0002] Cloud-based unmanned stores generally use remote audio and video intercom systems to enable real-time communication between customer service and customers, handling of abnormalities and emergency guidance. Existing systems typically assess channel quality by monitoring network latency jitter of the two-way audio stream and trigger self-healing actions such as channel switching.
[0003] However, the intercom terminals in stores operate continuously 24 / 7. The physical performance of the audio pickup hardware degrades over time. Coupled with the changing environmental noise from the start and stop of the cooling equipment and external sources, the existing system cannot accurately distinguish whether the sound quality degradation is due to network transmission jitter, changes in hardware status, or interference from the environment. This often leads to misjudgments—blindly switching when it is not necessary and responding slowly when a switch is needed, resulting in abnormal interruptions to the intercom service. Manual troubleshooting is time-consuming and directly affects the service continuity and operational safety of unmanned stores. Summary of the Invention
[0004] To address the shortcomings of existing technologies, this invention provides a method and system for ensuring the quality of intercom channels based on bidirectional audio stream delay jitter.
[0005] To solve the above-mentioned technical problems, the present invention provides the following technical solution: This invention provides a method for ensuring the quality of intercom channels based on bidirectional audio stream delay jitter, comprising: Acquire the impedance attenuation characteristics of the target terminal's acquisition device, as well as the first noise data of the target environment; The first noise data is processed according to the impedance attenuation characteristics to obtain the second noise data; Obtain the jitter characteristics of the audio data corresponding to the target terminal; Based on the second noise data and the jitter characteristics, a target processing instruction for the audio data is generated.
[0006] As a preferred embodiment of the present invention, the step of processing the first noise data according to the impedance attenuation characteristics to obtain the second noise data includes: If the impedance attenuation characteristic is less than the target impedance data, the first noise data is determined as the second noise data; When the impedance attenuation characteristic is greater than or equal to the target impedance data, noise correction data is obtained based on the impedance attenuation characteristic, the preset correction data, and the first noise data. The second noise data is obtained based on the first noise data and the noise correction data.
[0007] As a preferred embodiment of the present invention, the step of generating a target processing instruction for the audio data based on the second noise data and the jitter characteristics includes: When the second noise data is greater than or equal to the target noise data, a weighted feature is obtained based on the impedance attenuation feature and the second noise data, and the preset first threshold data is adjusted based on the weighted feature to obtain the second threshold data; If the jitter feature is less than the second threshold data, the allocation ratio of the audio data between the first processing module and the second processing module is determined based on the weighted feature.
[0008] As a preferred embodiment of the present invention, the step of obtaining the weighted features based on the impedance attenuation characteristics and the second noise data includes: The weighted feature is obtained by multiplying the impedance attenuation feature with the second noise data. The step of adjusting the preset first threshold data based on the weighted features to obtain the second threshold data includes: Based on the weighted features, the maximum weighted features, and the lifting coefficient data, the lifting ratio data is obtained; The second threshold data is obtained based on the lifting ratio data and the first threshold data.
[0009] As a preferred embodiment of the present invention, the method further includes: If the jitter characteristic is greater than or equal to the second threshold data, a switching instruction is generated; Acquire the first displacement data of the acquisition device, and determine the target time information based on the first displacement data; At the target time information, the transmission path of the audio data is switched from the first transmission module to the second transmission module.
[0010] As a preferred embodiment of the present invention, determining the target time information based on the first displacement data includes: Interpolation processing is performed on the first displacement data of adjacent sampling periods to obtain the target time information, which corresponds to the time when the displacement value is zero; The step of switching the transmission path of the audio data from the first transmission module to the second transmission module at the target time information includes: Within the first time information containing the target time information, the audio data sending terminal is redirected to the second transmission module, and the synchronous reception state of the first transmission module and the second transmission module is maintained. When the first time information ends, a disconnection command is generated to disconnect the receiving state of the first transmission module.
[0011] As a preferred embodiment of the present invention, before acquiring the impedance attenuation characteristics of the acquisition device of the target terminal, the method further includes: Generate a first control instruction, which is used to control the playback module of the target terminal to output a first audio signal; Acquire the second displacement data of the acquisition device collected by the sensing module; The first amplitude feature corresponding to the target frequency information is determined based on the second displacement data; The impedance attenuation characteristic is obtained based on the first amplitude characteristic and the second amplitude characteristic.
[0012] As a preferred embodiment of the present invention, before acquiring the first noise data of the target environment, the method further includes: Acquire the second audio data collected by the target terminal; If the second audio data satisfies the first condition data, the second audio data is subjected to frequency domain transformation to obtain spectrum data; Based on the spectral data, the full-band energy characteristics are determined to obtain the first noise data.
[0013] As a preferred embodiment of the present invention, the step of obtaining the jitter characteristics of the audio data corresponding to the target terminal includes: Acquire the first jitter data and the second jitter data of the audio data; The jitter characteristics of the audio data are determined based on the maximum value between the first jitter data and the second jitter data.
[0014] The present invention also provides a communication channel quality assurance system, comprising: The first acquisition module is used to acquire the impedance attenuation characteristics of the acquisition device of the target terminal and the first noise data of the target environment; The first correction module is used to process the first noise data according to the impedance attenuation characteristics to obtain the second noise data; The second acquisition module is used to acquire the jitter characteristics of the audio data corresponding to the target terminal; The target generation module is used to generate target processing instructions for the audio data based on the second noise data and the jitter characteristics.
[0015] The beneficial effects of this invention are: This application obtains second noise data by processing the first noise data of the target environment through the impedance attenuation characteristics of the target terminal's acquisition device, and then generates target processing instructions by combining the jitter characteristics of the corresponding audio data of the target terminal. This enables the correction of the detection results of environmental noise based on the physical state of the hardware, distinguishing the different effects of hardware attenuation, environmental noise and network jitter on the intercom audio quality. It solves the problem in the prior art that the intercom channel quality judgment is easily affected by hardware aging and on-site noise interference, leading to false triggering of self-healing actions or delayed response. It effectively improves the accuracy of intercom channel quality judgment, reduces unnecessary channel switching, and ensures the continuity and operational stability of intercom services. Attached Figure Description
[0016] The accompanying drawings are provided to further illustrate the invention and form part of the specification. They are used in conjunction with embodiments of the invention to explain the invention and do not constitute a limitation thereof. In the drawings:
[0017] Figure 1 This is a schematic diagram of the workflow of the intercom channel quality assurance method based on bidirectional audio stream delay jitter according to the present invention. Detailed Implementation
[0018] The technical solutions of this application will now be clearly and completely described with reference to the accompanying drawings. Obviously, the described embodiments are merely some embodiments of this application, and not all embodiments. The components of this application described and shown in the accompanying drawings can generally be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of this application provided in the accompanying drawings is not intended to limit the scope of the claimed application, but merely to illustrate selected embodiments of this application. All other embodiments obtained by those skilled in the art based on the embodiments of this application without inventive effort are within the scope of protection of this application.
[0019] It should be noted that similar reference numerals and letters in the following figures indicate similar items; therefore, once an item is defined in one figure, it does not need to be further defined and explained in subsequent figures. Furthermore, in the description of this application, terms such as "first," "second," etc., are used only to distinguish descriptions and should not be construed as indicating or implying relative importance.
[0020] like Figure 1 As shown, the present invention provides a method for ensuring the quality of intercom channels based on bidirectional audio stream delay jitter. Specifically, the store edge gateway obtains the diaphragm acoustic impedance attenuation characteristics of the microphone in the store intercom terminal through the diaphragm status monitoring channel.
[0021] The acoustic impedance attenuation characteristic of the diaphragm is obtained by processing the diaphragm amplitude response signal output by the diaphragm displacement sensing unit, and is used to represent the physical aging degree of the pickup diaphragm at the current moment relative to the factory calibration state.
[0022] At the same time, the store edge gateway obtains the first noise data of the store's on-site environment through the main audio output path of the microphone. This first noise data represents the initial measurement value of the noise energy level of the store's acoustic environment at the current moment.
[0023] After acquiring the impedance attenuation characteristics and the first noise data, the store edge gateway performs a noise baseline inverse correction operation.
[0024] Specifically, the edge gateway corrects the first noise data based on the impedance attenuation characteristics. When the impedance attenuation characteristics indicate that the microphone diaphragm is in good condition, the first noise data is directly used as the corrected second noise data. When the impedance attenuation characteristics indicate that the microphone diaphragm has undergone significant physical aging, the first noise data is subjected to a reverse stripping operation with the impedance attenuation characteristics as a weighting factor to eliminate the artificially high background noise component introduced by the aging of the microphone hardware, thereby obtaining the corrected second noise data.
[0025] The second noise data represents the actual acoustic environment noise level at the store.
[0026] The store edge gateway also collects the transmission layer quality parameters of the bidirectional audio stream between the store and the cloud in real time through the RTP / RTCP protocol stack, and extracts the latency jitter statistics as the jitter feature of the audio data. This jitter feature indicates the latency fluctuation of the current intercom link at the network transmission layer.
[0027] After obtaining the corrected second noise data and the real-time jitter features, the store edge gateway inputs both into the joint gating decision module. This module generates a target processing instruction for the audio data based on different state combinations between the real environmental noise energy level represented by the second noise data and the network transmission jitter level represented by the jitter features.
[0028] The target processing instruction is used to control the store edge gateway or store intercom terminal to perform corresponding channel quality assurance actions. When the ambient noise is high but the network transmission is still acceptable, the target processing instruction is used to adjust the allocation ratio of audio encoding bandwidth between anti-jitter processing and noise reduction processing. When both ambient noise and network jitter are high, the target processing instruction is used to trigger cross-service provider channel switching of audio data transmission path.
[0029] Furthermore, in obtaining the impedance attenuation characteristics A of the pickup diaphragm... dAfter obtaining the first noise data N1 of the store environment, the store edge gateway performs a noise baseline reverse correction operation to eliminate the artificially high noise floor component introduced by the aging of the microphone hardware.
[0030] Edge gateways will have impedance attenuation characteristics A d Compared with the preset target impedance data T a The comparisons are made, and different correction strategies are implemented based on the comparison results.
[0031] The first scenario (the diaphragm is in good condition).
[0032] When A d Less than T a At this time, the edge gateway determines that the physical state of the microphone diaphragm is still in the healthy range, and the deviation between its acoustic-electric conversion characteristics and the factory calibration state is within an acceptable range. At this time, the first noise data N1 can truly reflect the acoustic background of the store environment. The edge gateway directly determines N1 as the corrected second noise data N2, that is, N2=N1.
[0033] The second scenario (the diaphragm has undergone significant physical aging).
[0034] When A d Greater than or equal to T a At that time, the edge gateway determined that the microphone diaphragm had undergone significant physical aging, with the elastic modulus of the diaphragm material deteriorating and the compliance of the suspension system decreasing, resulting in a systematic decrease in the acoustic-to-electrical conversion efficiency.
[0035] To compensate for this attenuation, the front-end automatic gain control module applies a larger amplification factor. During this process, the thermal noise and diaphragm nonlinear distortion components inside the microphone are amplified simultaneously, resulting in the first noise data N1 calculated based on the signal containing a portion of artificially high background noise that does not originate from the actual acoustic environment of the store, but rather from the physical state of the hardware itself.
[0036] Regarding the second scenario described above, the edge gateway uses impedance attenuation characteristic A d Using the weighting factor α, noise correction data Δ is generated as follows: N : ΔN=min(A d ×α×N1,N1×50%); Among them, A d The impedance attenuation characteristic (value from 0% to 100%) represents the degree of attenuation of the acoustic-electric conversion efficiency of the diaphragm relative to its factory condition; α is the correction coefficient (valued at 0.6 in this embodiment), which represents the noise floor increase coefficient introduced by the pickup model under unit impedance attenuation rate. This value is pre-calibrated based on the statistical regression relationship between the noise floor increase and impedance attenuation rate measured in accelerated aging tests of the same pickup model. N1 is the first noise data; ΔN is the noise correction data, representing the artificially high background noise component that needs to be removed from the first noise data; the second term N1×50% in the min operation represents the upper limit of the maximum removal ratio, that is, the removal amount does not exceed 50% of the first noise data.
[0037] The purpose of setting this upper limit is to prevent damage under extreme aging conditions (A... d Excessive stripping (approaching 100%) leads to distortion of the corrected ambient noise energy level, ensuring that the lowest perceptible component of the store's true acoustic environment is preserved.
[0038] The edge gateway obtains the corrected second noise data N2 based on the first noise data N1 and the noise correction data ΔN, as follows: N2 = N1 - ΔN; That is, the second noise data N2 is equal to the first noise data N1 minus the artificial high noise component ΔN stripped from it.
[0039] When A d When it reaches 100%, N2 = N1 × (1 - 0.6) = 0.4 × N1, which means that 40% of the real acoustic environment component is retained, ensuring that the system can still perceive the changing trend of the store's acoustic environment under extreme aging conditions.
[0040] The above correction coefficient α=0.6 is set based on the following: Based on the accelerated aging test of the same model of microphone under 85℃ and 85% relative humidity conditions with continuous application of 100dBSPL pink noise excitation, the background noise output level of the microphone during the silent period was measured at different stages of the aging process (impedance attenuation rate reached 5%, 10%, 15%, 20%, 25%, and 30%, respectively). Linear regression analysis was performed with the impedance attenuation rate of the microphone as the independent variable and the background noise increase as the dependent variable. The regression coefficients of the two were approximately 0.58 to 0.62, and the median value of 0.6 was taken as the preset correction coefficient.
[0041] The actual effect of this coefficient is that for every 1% increase in diaphragm impedance attenuation rate, the proportion of hardware overestimation in the background noise detection value increases by about 0.6 percentage points.
[0042] The reasonable range for the correction coefficient α is 0.5 to 0.7. This range is determined based on the following: Accelerated aging test data shows that the regression coefficient between the noise floor increase and impedance attenuation rate of different MEMS microphone models is distributed in the range of 0.5 to 0.7. The lower limit of 0.5 corresponds to the anti-aging model with a thicker diaphragm design, and the upper limit of 0.7 corresponds to the high-sensitivity model with a thinner diaphragm design. In this embodiment, the middle value of 0.6 is taken.
[0043] The reasonable range for the maximum stripping ratio of 50% is 40% to 60%. This range ensures that the system can still retain no less than 40% of the real acoustic environment information under extreme aging conditions, and avoids the complete loss of the ability of the noise level to follow changes in the acoustic environment of the store.
[0044] Furthermore, after obtaining the corrected second noise data N2, the store edge gateway compares N2 with the latency jitter sampling value D. j Input the joint gating decision module.
[0045] The joint gating decision module first determines whether the second noise data N2 reaches the target noise data T. n2 .
[0046] The target noise data T n2 (In this embodiment, the preset value is 65 dBSPL equivalent) The actual meaning is the critical sound pressure level at which the ambient noise in the store begins to significantly affect speech intelligibility. When the ambient noise exceeds this threshold, if no noise reduction measures are taken, the voice signal received by the remote customer service will be partially masked by the ambient noise, resulting in a decrease in speech intelligibility.
[0047] This threshold is determined based on the ITU-T P.800 series of subjective listening test standards: when the average subjective opinion score (MOS) of a speech signal in a background of 65 dBSPL pink noise drops below 3.0, it reaches the critical state of being partially intelligible but difficult to hear.
[0048] First-level condition judgment (noise threshold trigger).
[0049] When N2 <T n2 If the joint gate control determination module determines that the current acoustic environment of the store has not reached the level that requires joint gate control adjustment, it will not perform weighted feature calculation and threshold adjustment actions, and will continue to maintain the baseline processing flow.
[0050] When N2≥T n2 At that time, the joint gate control determination module determines that the current store has entered a high-noise operating condition, and then initiates the following joint gate control calculation.
[0051] The joint gating decision module acquires the impedance attenuation characteristic A from the diaphragm state monitoring channel. d And the corrected second noise data N2, the weighted feature W is generated as follows: W=A d ×N2; Among them, A dN1 represents the impedance attenuation characteristic (values from 0% to 100%), indicating the current physical aging degree of the pickup diaphragm; N2 is the second noise data (unit: dBSPL equivalent), representing the actual acoustic environment noise level at the store; the weighting characteristic W is in dB·%, which means the product of the diaphragm aging degree and the ambient noise intensity, representing the combined compression effect of the two on the effective dynamic range of speech.
[0052] The design principle of this weighted feature W is as follows: In store intercom scenarios, diaphragm aging (A) d The effects of N2 elevation and environmental noise on uplink speech quality are not independent, but rather exhibit a product-coupled relationship. Due to the decrease in the compliance of the suspension system and the increase in material internal friction, the aging diaphragm exhibits more significant nonlinear distortion (the linear relationship between diaphragm displacement and sound pressure deteriorates) under high sound pressure levels, which further amplifies the masking effect of environmental noise on the speech signal.
[0053] If the two are treated as independent variables separately (such as by additive combination), the coupling amplification effect cannot be reflected. Therefore, this implementation uses multiplicative combination to construct joint weighted feature W to accurately reflect the degree of synergistic compression of the effective dynamic range of speech by diaphragm aging and environmental noise.
[0054] After obtaining the weighted feature W, the joint gating decision module applies the preset first threshold data T. j1 Dynamic adjustments are made to generate a second threshold data T suitable for the current operating conditions. j2 .
[0055] First threshold data T j1 (In this embodiment, it is preset to 50 milliseconds) is the upper limit of the latency jitter tolerance of the store intercom link under standard working conditions. Its actual meaning is the latency jitter threshold at which the human ear begins to perceive audio discontinuity.
[0056] This value is determined based on statistical experimental data on latency jitter perception in the ITU-T G.114 standard: when latency jitter exceeds 50 milliseconds, more than 85% of ordinary listeners can perceive the discontinuity in audio playback.
[0057] First threshold data T j1 Stored in the local configuration cache of the edge gateway, serving as the baseline value for all subsequent dynamic adjustments.
[0058] Second threshold data T j2 Generate in the following manner: R boos t=γ×(W / W m ); T j2 =min(T j1×(1+R boost ),T j1 ×1.5); Among them, R boost The upscaling percentage represents the relative increase in the jitter tolerance threshold under the current weighted feature W; W is the aforementioned weighted feature. W m The maximum possible value of the weighted feature (corresponding to the W value when Ad=100% and N2=120dBSPL equivalent, i.e., W). m =100%×120dB=120dB·%), the upper limit of this maximum value is: the maximum value of impedance attenuation characteristic is 100% (the diaphragm completely loses its ability to convert sound to electricity), and the maximum value of environmental noise is 120dBSPL (human auditory pain threshold). γ is a preset lift coefficient (taken as 0.5 in this embodiment), representing the system's response gain to improved jitter tolerance in noisy environments; The second term T in the min operation j1 ×1.5 indicates that the second threshold data does not exceed 150% of the first threshold data (i.e., the maximum is no more than 75 milliseconds). The purpose of setting this upper limit is to prevent the threshold from being raised too much under high noise extreme conditions, which would cause the system to delay the channel switching due to the excessively high threshold when the network transmission has been significantly degraded, resulting in voice interruption.
[0059] Second threshold data T j2 The actual meaning is: under the coupled state of the current store noise environment and the aging of the microphone, the upper limit of the system's tolerance for network latency jitter is dynamically expanded to the value of T. j2 ≥T j1 This means that the system's tolerance for network jitter is moderately relaxed in high-noise environments. Even if there is a certain degree of network jitter, the enhanced acoustic noise reduction processing at the store end can still ensure the intelligibility of customer service voice and avoid misjudging network failure due to degraded noise perception.
[0060] The above-mentioned lift coefficient γ=0.5 is based on the following: subjective listening test conducted in a simulated store noise environment (65dBSPL to 75dBSPL pink noise). The test subjects were 20 adult subjects with normal hearing. The test content was to score the standard speech samples with MOS at different jitter levels (20ms to 100ms, step size 10ms).
[0061] Test results show that, in noisy environments, the increase in jitter tolerance is approximately 10% to 12% for every 10 dB increase in the equivalent noise sound pressure level.
[0062] In this embodiment, γ is set to 0.5, that is, the weighted feature W is relative to W mFor every 10% increase, the jitter threshold is raised by 5% to ensure that the increase does not exceed the actual tolerance limit of the human ear in a noisy environment.
[0063] The second threshold data T is obtained through the above calculation. j2 Subsequently, the joint gating determination module further analyzes the current jitter feature D. j With the second threshold data T j2 The comparison is performed, and the corresponding control action is executed based on the comparison result.
[0064] When D j <T j2 This indicates that although the current network latency jitter has triggered the joint gating logic under noisy conditions, the jitter level has not yet reached the raised switching threshold, and the system does not need to perform channel-level switching.
[0065] At this point, the joint gating decision module, based on the aforementioned weighted feature W, determines the allocation ratio of audio data between the first processing module and the second processing module in the following manner: ΔB=min(β×(W / W m )×B t ×R m B t ×L m ); Wherein, ΔB is the amount of bandwidth allocated from the first processing module to the second processing module, and β is the preset bandwidth allocation coefficient (taken as 0.8 in this embodiment), representing the driving gain of the weighted feature on the amount of bandwidth allocation. B t The total available bandwidth for audio encoding (determined by current network bandwidth conditions and encoding configuration); R m The maximum allocation ratio coefficient is preset (in this embodiment, the value is 0.3). L m The maximum single allocation amount is preset to a coefficient (in this embodiment, the value is 0.25). The min operation ensures that the allocation amount is limited by both the maximum proportional coefficient and the maximum allocation amount per transaction, preventing excessive adjustments in a single transaction from causing insufficient resources in the anti-jitter redundancy module and resulting in playback stuttering.
[0066] The allocation ratio after the transfer is as follows: the bandwidth share of the first processing module (anti-jitter redundancy module, responsible for anti-jitter processing such as forward error correction (FEC) coding and packet loss retransmission requests) is reduced from the baseline ratio (e.g., 70%) to 70% - ΔB / B. t ; The bandwidth allocation of the second processing module (acoustic noise reduction module, responsible for noise reduction processing such as adaptive noise suppression and speech enhancement) increases from the baseline (e.g., 30%) to 30% + ΔB / B.t .
[0067] The meaning of this bandwidth allocation action is: under high noise conditions (N2≥T) n2 The acoustic environment of the store has become a major bottleneck restricting the quality of uplink voice calls. However, network jitter has not yet deteriorated to an unacceptable level. Therefore, the system has transferred some of the coding resources originally used to resist network jitter to enhance the acoustic noise reduction processing at the store end.
[0068] The bandwidth allocation factor β=0.8 is set based on the following: Based on different bandwidth configurations (B... t The coding efficiency test was conducted under the combined conditions of 32kbps to 128kbps (step size 16kbps) and different noise levels (N2 is 65dB, 70dB, and 75dBSPL). Five sets of values were used in the test: β=0.5, 0.6, 0.7, 0.8, and 0.9. MOS score was used as the evaluation index.
[0069] Test results show that when β=0.8, the average MOS score reaches 3.45, higher than the other four groups (MOS=3.02 when β=0.5, MOS=3.28 when β=0.9), indicating that β=0.8 achieves a balance between noise reduction gain and jitter attenuation. Maximum single transfer amount L m The setting of 0.25 (i.e. 25%) is based on the following constraints: the anti-jitter redundancy module retains at least 75% of the baseline bandwidth, which is sufficient to maintain the minimum configuration requirement of 50% redundancy of forward error correction (FEC) coding, and avoids the anti-jitter capability from dropping to an unacceptable level due to excessive bandwidth allocation.
[0070] When D j ≥T j2 When this occurs, it indicates that the network latency jitter under noisy conditions has further deteriorated to the raised switching threshold. The system determines that the current coding parameter adjustment cannot effectively guarantee call quality and a channel switching operation needs to be performed.
[0071] By triggering weighted feature calculation and threshold raising through the first-level condition judgment (comparison of the second noise data and the target noise data), and driving the determination of bandwidth allocation ratio through the second-level condition judgment (comparison of jitter features and the second threshold data), the joint gating decision module realizes a two-level linkage decision mechanism of noise-triggered raising and jitter-determined branch. If the noise is high, the threshold is raised and the bandwidth is reconfigured. If the jitter exceeds the raised threshold, the channel is switched. If the jitter does not exceed the raised threshold, the encoding strategy after bandwidth reconfiguration is maintained.
[0072] This two-level linkage mechanism ensures that the system performs differentiated control under different operating conditions, and the reasonable range of the lifting coefficient γ is 0.3 to 0.7. The lower limit of 0.3 corresponds to a relatively conservative system configuration that improves the jitter tolerance response in noisy environments, and is suitable for scenarios where the acoustic environment of the store is relatively stable and the noise fluctuation is small; The upper limit of 0.7 corresponds to a system configuration that responds more positively to increased jitter tolerance in noisy environments, suitable for scenarios where the acoustic environment of the store fluctuates greatly and high noise impacts occur frequently. This implementation takes the middle value of 0.5 to balance the two.
[0073] The reasonable range of bandwidth allocation coefficient β is 0.6 to 0.9. The lower limit of 0.6 corresponds to the system configuration with high requirements for maintaining anti-jitter redundancy capability, which is suitable for stores with large network environment fluctuations. The upper limit of 0.9 corresponds to the system configuration with high requirements for noise reduction capability, which is suitable for scenarios where the store environment noise is continuously high and the network is relatively stable. In this implementation, 0.8 is taken to prioritize the noise reduction requirements under high noise conditions.
[0074] Maximum allocation ratio coefficient R m The reasonable range for this value is 0.2 to 0.4. The lower limit of this range is determined based on the constraint that the anti-jitter redundancy module must retain at least 60% of the baseline bandwidth to ensure basic anti-jitter capability; the upper limit is determined based on engineering test results showing that the marginal benefit of the acoustic noise reduction module tends to saturate after obtaining more than 40% additional bandwidth.
[0075] Furthermore, a second threshold data T is generated. j2 Subsequently, the joint gating judgment module continuously monitors the current jitter characteristic D. j With the second threshold data T j2 The size relationship.
[0076] When D j ≥T j2 This indicates that, under the coupled conditions of the current store noise environment and the aging degree of the microphone diaphragm, the network latency jitter has deteriorated to the raised switching threshold.
[0077] At this point, the joint gate control judgment module determines that the conventional encoding parameter adjustment (i.e., the bandwidth allocation ratio adjustment between the first processing module and the second processing module) can no longer effectively guarantee the call quality of the intercom channel. The acoustic environment interference and network transmission degradation in the store have formed a double superposition, and the system must perform channel-level self-healing actions to maintain the continuity of the intercom link.
[0078] The edge gateway then generates a switching instruction. This switching instruction contains the target channel identifier information (i.e., the channel identifier of the second transmission module) and execution time window constraint information. After the switching instruction is generated, it is sent to the instruction queue of the channel switching control module to wait for execution, but it is not immediately sent to the transmission module to perform physical switching. Instead, it enters a waiting phase constrained by the physical state.
[0079] While generating the switching command, the edge gateway initiates real-time acquisition and caching of the first displacement data.
[0080] The first displacement data comes from the diaphragm amplitude response signal collected by the sensing module (i.e., the diaphragm displacement sensing unit integrated inside the microphone of the store intercom terminal). This sensing unit shares the same diaphragm physical structure with the acquisition device (microphone), but the output channel is independent of the main audio acquisition path. It can continuously output a signal representing the real-time mechanical displacement state of the diaphragm without interfering with the normal intercom audio acquisition.
[0081] The first displacement data and the second displacement data used to calculate impedance attenuation characteristics both originate from the same sensing unit. The difference is that the second displacement data is displacement data collected within the calibration time window for calculating impedance attenuation characteristics, while the first displacement data is displacement data collected after channel switching is triggered for zero-crossing detection. Both have the same sampling rate (1000 times per second in this embodiment), but the starting time and purpose of the acquisition are different.
[0082] The zero-crossing detection module in the edge gateway takes the real-time displacement value output by the diaphragm displacement sensing unit as input and performs the operation of determining the target time information in the first displacement data.
[0083] The target moment information is defined as the moment when the instantaneous displacement value of the diaphragm crosses its static equilibrium position (i.e., the displacement value is zero). At this moment, the instantaneous velocity of the diaphragm is at its maximum value but the acceleration is zero. The diaphragm is in a state of dynamic equilibrium and has the lowest sensitivity to external disturbances.
[0084] The zero-crossing detection module starts running after receiving the switching trigger signal. It takes the real-time displacement data stream as input and continuously detects the change in the sign of the displacement value of adjacent sampling points. When the signs of the displacement values of two consecutive sampling points are opposite (i.e., one is positive and the other is negative), it is determined that there is a zero-crossing point between the two sampling points.
[0085] Based on this determination, the module marks the first time that meets the zero-crossing condition after the switching trigger signal and identifies it as the target time information.
[0086] After determining the target time information, the edge gateway performs an audio data transmission path switching operation. Specifically, at the time point corresponding to the target time information (the error window is ±1 millisecond in this embodiment), the edge gateway switches the output data stream of the store-side audio encoder from the sending buffer corresponding to the first transmission module to the sending buffer corresponding to the second transmission module, and at the same time switches the receiving data source of the cloud downlink audio stream from the receiving port corresponding to the first transmission module to the receiving port corresponding to the second transmission module.
[0087] After the switch is completed, the audio data is switched from the transmission path carried by the first transmission module to the transmission path carried by the second transmission module.
[0088] When the channel switching control module performs a switching operation, its control command directly acts on the audio data channel routing unit in the store edge gateway. This routing unit is located at the end of the audio data processing pipeline of the edge gateway and is connected to the network interface driver layer of the first transmission module and the second transmission module respectively.
[0089] After receiving the switching execution command, the routing unit performs the redirection of the transmit buffer pointer (writing the audio encoder output data into the transmit queue of the second transmission module instead of the transmit queue of the first transmission module) and the redirection of the receive data source (reading downlink audio frames from the receive queue of the second transmission module instead of the receive queue of the first transmission module) within a ±1 millisecond window corresponding to the target time information, thereby completing the physical switching of the data transmission path.
[0090] Furthermore, after acquiring the first displacement data from the acquisition device and determining the switching trigger condition, the edge gateway needs to locate the target time information from the first displacement data and perform a switching action across transmission modules at the time sequence position corresponding to the target time information.
[0091] The zero-crossing detection module in the edge gateway takes the real-time displacement value output by the diaphragm displacement sensing unit as input and performs interpolation processing on the first displacement data of adjacent sampling periods. The diaphragm displacement sensing unit continuously outputs the instantaneous displacement value of the diaphragm according to the first preset sampling period (1 millisecond in this embodiment), and the zero-crossing detection module monitors the sign of the displacement value between two adjacent sampling points.
[0092] When two consecutive sampling points are detected to have opposite signs of displacement values (i.e., the displacement value of one sampling point is positive and the displacement value of the next sampling point is negative, or vice versa), it indicates that the diaphragm has crossed its static equilibrium position between these two sampling times, that is, there is a zero-crossing point.
[0093] Due to the existence of the sampling period, the precise time of the zero crossing may not fall exactly on the sampling time, but rather at a certain position between two sampling times. The zero crossing detection module uses linear interpolation to estimate the precise time when the displacement value is zero based on the displacement values of the two adjacent sampling points and their respective timestamps, and determines this precise time as the target time information.
[0094] This target time information corresponds to the moment when the instantaneous displacement of the diaphragm is zero, that is, the moment when the diaphragm vibration waveform crosses the zero-level line. At this moment, the instantaneous velocity of the diaphragm reaches its maximum value but the acceleration is zero, and the diaphragm is in a state of dynamic equilibrium. This characteristic makes the transient impact on the diaphragm caused by the switching of audio data streams when a switching action is performed at this moment smaller. The edge gateway uses this target time information as the timing reference for subsequent switching actions.
[0095] After determining the target time information, the edge gateway defines a first time information around the time point corresponding to the target time information. The first time information is a time window centered on the target time information, and its boundary range is set to extend a preset time length before and after the target time information (±1 millisecond in this embodiment, that is, the total length of the first time information is 2 milliseconds). Within the first time information, the edge gateway performs the switching operation of the audio data transmission path.
[0096] At the start of the first time information, the edge gateway initiates the operation of switching the audio data sending terminal to the second transmission module. This switching operation includes switching the output data stream of the store-side audio encoder from the sending buffer corresponding to the first transmission module to the sending buffer corresponding to the second transmission module. At the same time, within the first time information, the edge gateway maintains the synchronous receiving state of the first and second transmission modules. The edge gateway simultaneously receives downlink audio data frames from both transmission modules and buffers the two data streams in their respective receiving buffers.
[0097] This synchronous reception state ensures that even if there is a slight timing offset in the data stream of the second transmission module at the moment of switching, the edge gateway can still maintain the continuity of downlink audio output. The duration of the first time information of the synchronous reception state is preset to 2 milliseconds (based on a window of 1 millisecond before and after the target time).
[0098] When the first information ends, the edge gateway generates a disconnect command to disconnect the receiving state of the first transmission module. This disconnect command instructs the edge gateway to stop reading downlink audio data frames from the receiving buffer of the first transmission module and completely switch the data source for downlink audio playback to the receiving buffer of the second transmission module. At this time, the first transmission module no longer undertakes the task of transmitting audio data, and the second transmission module becomes the only active transmission channel.
[0099] Furthermore, during the establishment and continuation of the intercom session in the store, the edge gateway needs to acquire the impedance attenuation characteristics of the acquisition device (i.e., the microphone of the intercom terminal in the store). This characteristic is the benchmark for subsequent noise baseline inverse correction and joint gating determination.
[0100] The edge gateway generates a first control command according to a preset calibration cycle (once every 30 seconds in this embodiment), and the first control command is sent to the playback module of the target terminal (i.e., the speaker unit of the store intercom terminal module).
[0101] After receiving the first control command, the playback module outputs a first audio signal. The first audio signal is a calibration audio excitation signal with known spectral characteristics. The calibration signal is a composite signal containing multiple single-frequency components, and its frequency covers the audible frequency range of the human ear (20Hz to 20kHz). The amplitude of each frequency component is known from the factory calibration.
[0102] After the first audio signal output by the playback module propagates within the store space, it is received by the acquisition device (i.e., the microphone). The microphone diaphragm undergoes forced vibration under the excitation of this sound pressure. The amplitude of this forced vibration is directly related to the current mechanical compliance, elastic modulus, and other characteristics of the diaphragm. A diaphragm in a healthy state produces the expected displacement amplitude under the same sound pressure level excitation. However, a diaphragm that has undergone physical aging will produce a displacement amplitude that is less than the factory calibration value under the same excitation condition due to the decrease in the compliance of the suspension system and the deterioration of the elastic modulus of the material.
[0103] The edge gateway collects the second displacement data of the acquisition device through a sensing module, which is the diaphragm displacement sensing unit integrated inside the microphone of the store intercom terminal. It shares the same diaphragm physical structure with the acquisition device, but its output channel is independent of the main audio acquisition channel. During the period when the playback module outputs the first audio signal, the diaphragm displacement sensing unit continuously outputs the amplitude response signal of the diaphragm under the sound pressure excitation. The edge gateway performs low-pass filtering on the raw signal output by the sensing module (the cutoff frequency is set to more than 3 times the mechanical resonance frequency of the diaphragm, for example, 60kHz) to eliminate high-frequency electrical noise and obtain a smooth diaphragm displacement curve as the second displacement data. The sampling rate of the second displacement data is consistent with the preset sampling rate of the diaphragm state monitoring channel (1000 times per second in this embodiment). Its time window covers the entire calibration signal playback period plus a margin of 50 milliseconds before and after, so as to fully capture the steady-state response of the diaphragm.
[0104] The edge gateway performs time-domain to frequency-domain transformation on the second displacement data (e.g., through Fast Fourier Transform) to extract the first amplitude feature corresponding to the target frequency information. In this embodiment, the target frequency information is a 1kHz frequency point, which is the frequency band where the energy of the speech signal is most concentrated, and is also a typical frequency point representing the acoustic-to-electrical conversion efficiency of the microphone.
[0105] The first amplitude characteristic is the displacement response amplitude of the diaphragm at a frequency of 1kHz, which represents the actual mechanical displacement capability of the diaphragm under the current physical aging state to the excitation of 1kHz sound pressure.
[0106] The edge gateway calculates the impedance attenuation characteristics based on the first amplitude characteristic and the second amplitude characteristic.
[0107] The second amplitude characteristic is the reference displacement amplitude stored in the local configuration memory of the edge gateway during the factory calibration stage. It means the theoretical displacement amplitude that the same model of microphone should produce under the same sound pressure level and the same frequency calibration audio excitation in the factory calibration state (i.e., the diaphragm material has not undergone any aging or fatigue). The edge gateway compares the first amplitude characteristic with the second amplitude characteristic and calculates the percentage of the difference between the two to the second amplitude characteristic as the impedance attenuation characteristic. The calculation method is: impedance attenuation characteristic = (second amplitude characteristic - first amplitude characteristic) / second amplitude characteristic × 100%.
[0108] The impedance attenuation characteristic indicates the degree of attenuation of the diaphragm's acoustic-electric conversion efficiency relative to its factory condition at the current moment. When this value is less than the target impedance data (preset to 15% in this embodiment), the diaphragm is considered to be in a healthy state. When this value reaches or exceeds the target impedance data, it indicates that the diaphragm has undergone significant physical aging and noise baseline reverse correction needs to be initiated.
[0109] Through the generation of the first control command, the acquisition of the second displacement data, the extraction of the first amplitude feature corresponding to the target frequency, and the calculation of impedance attenuation features based on the first and second amplitude features, the edge gateway completes the real-time monitoring and representation of the physical status of the acquisition device, providing a benchmark for subsequent noise correction and joint gating determination.
[0110] Furthermore, during the establishment and continuation of the intercom session at the store, the edge gateway needs to acquire the first noise data of the target environment, which is one of the inputs for subsequent noise baseline inverse correction.
[0111] During the intercom session, the store edge gateway continuously monitors the real-time audio signal output by the store intercom terminal acquisition device (i.e., microphone) and extracts second audio data for environmental noise assessment.
[0112] The acquisition of the second audio data is subject to a prerequisite, namely the first condition data.
[0113] In this embodiment, the first condition data is defined as a silent period in which the store is inactive and there are no effective customer voice activities, and this silent period lasts for more than a preset stable time window.
[0114] The determination of the silent period state is completed by the voice activity detection module built into the store edge gateway. This module performs a comprehensive statistical analysis of the short-time energy and zero-crossing rate of the digital audio signal output by the microphone main channel, and outputs a binary status indicator of whether there is voice activity.
[0115] When the voice activity detection module outputs a status indicator indicating that there is no voice activity and this status is maintained for a duration exceeding a preset stable time window (e.g., 200 milliseconds), the edge gateway determines that the current data meets the first condition and then triggers the environmental noise acquisition process.
[0116] The purpose of this prerequisite is to ensure that the signal collected by the microphone does not contain customer voice components, so that the signal can truly reflect the acoustic background state of the store environment and avoid the customer voice energy being mistakenly included in the environmental noise level.
[0117] If the second audio data satisfies the first condition data, the edge gateway performs frequency domain transformation processing on the second audio data. The edge gateway takes the second audio data as input, extracts a frame of environmental audio sample according to the preset sampling window length (e.g., 1 second), performs fast Fourier transform on the sample, and converts it from time domain representation to frequency domain representation. Through this frequency domain transformation processing, the energy distribution of each frequency component in the second audio data can be analyzed.
[0118] After obtaining the frequency domain representation, the edge gateway determines the full-band energy characteristics based on the spectrum data obtained from the frequency domain transformation. The edge gateway calculates the total energy value of the spectrum data in the full-band range and uses this total energy value as the first noise data of the target environment. This first noise data is expressed as the equivalent value of dBSPL (decibels of sound pressure level), representing the initial measurement value of the store's environmental noise. This value is used as the input of the subsequent noise baseline reverse correction module to perform hardware aging effect stripping processing.
[0119] Through the above frequency domain conversion and full-band energy feature extraction operations, the edge gateway completed the real-time measurement of the ambient noise level at the store site under the silent period trigger condition, which served as the initial measurement value for subsequent noise correction based on the diaphragm physical state.
[0120] Furthermore, during the intercom conversation in the store, the edge gateway needs to obtain the jitter characteristics of the audio data. These jitter characteristics are one of the input parameters used for comparison with the jitter threshold in the subsequent joint gating decision.
[0121] The store edge gateway collects the transport layer quality parameters of the bidirectional audio stream in real time based on the RTP / RTCP protocol stack, and obtains the first jitter data and the second jitter data of the audio data respectively.
[0122] The first jitter data represents the statistical variance of the time interval between the arrival of the uplink audio stream data packets transmitted from the store to the cloud. Specifically, it is the jitter field value extracted from the RTCP receive report block periodically returned by the store edge gateway from the cloud intercom service cluster, expressed in milliseconds.
[0123] The second jitter data represents the statistical variance of the time interval between the arrival of downlink audio stream data packets from the cloud to the store edge gateway. Specifically, it is the jitter field value extracted from the RTCP receive report block returned by the cloud intercom service cluster from the store edge gateway, also expressed in milliseconds. Both jitter data are real-time quality measures reflecting the stability of audio stream transmission in the corresponding direction. The larger the value, the more uneven the arrival time of audio stream data packets in that direction, and the higher the risk of continuity of audio playback at the receiving end.
[0124] After acquiring the first jitter data and the second jitter data, the edge gateway compares the two and determines the jitter characteristics of the audio data based on the maximum value between them.
[0125] The edge gateway selects the larger of the first jitter data and the second jitter data as the jitter feature of the current intercom link. This jitter feature represents the degree of latency fluctuation of the entire two-way intercom link at the current moment.
[0126] The purpose of taking the larger value of the two rather than the average or smaller value is that severe jitter in either direction of the bidirectional audio stream will directly affect the call experience at the far end or near end. Therefore, the system uses the side with more severe jitter in the two directions as the basis for the quality assessment of the entire link, so that the quality assurance decision can cover the worse transmission direction.
[0127] In the subsequent joint gating decision, the edge gateway compares the jitter feature with the raised second threshold data to determine whether to perform bandwidth allocation adjustment or channel switching.
[0128] Through the acquisition, comparison, and maximum value taking of the first and second jitter data, the edge gateway completes a comprehensive representation of the current intercom link latency jitter level, providing jitter input parameters for subsequent joint gating judgment.
[0129] To implement the above method, the present invention also provides an information processing system. In one specific embodiment, the system can be deployed in a store edge gateway and communicate with the store intercom terminal module to execute the aforementioned information processing method.
[0130] The system includes a first acquisition module, a first correction module, a second acquisition module, and a target generation module.
[0131] The first acquisition module is configured to acquire the impedance attenuation characteristics of the acquisition device of the target terminal, as well as the first noise data of the target environment.
[0132] This module communicates with the built-in diaphragm displacement sensing unit of the acquisition device (i.e., the microphone of the store intercom terminal) through a diaphragm status monitoring channel independent of the main audio acquisition link, and continuously acquires the acoustic impedance attenuation characteristics of the diaphragm. This characteristic is expressed as a percentage of the degree of attenuation of the acoustic-electric conversion efficiency of the diaphragm relative to its factory calibration state.
[0133] Meanwhile, this module communicates with the main audio output path of the microphone, triggering the collection of environmental audio samples and frequency domain energy extraction when the customer is in a period of silence at the store, thereby obtaining the first noise data of the target environment.
[0134] The first correction module is configured to process the first noise data according to the impedance attenuation characteristics to obtain the second noise data.
[0135] This module receives impedance attenuation characteristics and first noise data from the first acquisition module, compares the impedance attenuation characteristics with the target impedance data preset by the system, and directly outputs the first noise data as the second noise data when the impedance attenuation characteristics are less than the target impedance data. When the impedance attenuation characteristic is greater than or equal to the target impedance data, the impedance attenuation characteristic is used as a weighting factor in combination with the preset correction data to perform an inverse correction operation on the first noise data, removing the artificially high noise floor component introduced by the aging of the acquisition equipment hardware, and generating the corrected second noise data.
[0136] The second acquisition module is configured to acquire the jitter characteristics of the audio data corresponding to the target terminal. This module collects the transport layer quality parameters of the bidirectional audio stream in real time based on the RTP / RTCP protocol stack, acquires the first jitter data in the uplink direction and the second jitter data in the downlink direction respectively, and determines the larger value of the two as the latency jitter characteristics of the current link.
[0137] The target generation module is configured to generate target processing instructions for the audio data based on the second noise data and the jitter characteristics.
[0138] This module receives the corrected second noise data from the first correction module and the jitter feature from the second acquisition module. It inputs the two into the joint gating judgment logic and maintains the baseline processing strategy when the second noise data is less than the target noise data. When the second noise data is greater than or equal to the target noise data, a weighted feature is further generated based on the product of the impedance attenuation feature and the second noise data. The preset first threshold data is dynamically adjusted based on the weighted feature to obtain the second threshold data. The corresponding target processing instruction is generated based on the comparison result between the jitter feature and the second threshold data. When the jitter feature is less than the second threshold data, a bandwidth allocation ratio adjustment instruction is generated. When the jitter feature is greater than or equal to the second threshold data, a channel switching instruction is generated.
[0139] Through the coordinated work of the above four functional modules, the system is used to adaptively ensure the quality of the intercom channel. Its deployment in the store edge gateway enables physical state perception, noise correction and joint gate control determination to be completed on the edge side close to the acquisition device, reducing the delay and deviation of cloud decision-making on the physical state perception of the store site.
[0140] Finally, it should be noted that the above descriptions are merely preferred embodiments of the present invention and are not intended to limit the present invention. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing embodiments or make equivalent substitutions for some of the technical features. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.
Claims
1. A method for ensuring the quality of intercom channels based on bidirectional audio stream delay jitter, characterized in that, include: Acquire the impedance attenuation characteristics of the target terminal's acquisition device, as well as the first noise data of the target environment; The first noise data is processed according to the impedance attenuation characteristics to obtain the second noise data; Obtain the jitter characteristics of the audio data corresponding to the target terminal; Based on the second noise data and the jitter characteristics, a target processing instruction for the audio data is generated.
2. The method for ensuring the quality of intercom channels based on bidirectional audio stream delay jitter according to claim 1, characterized in that, The step of processing the first noise data according to the impedance attenuation characteristics to obtain the second noise data includes: If the impedance attenuation characteristic is less than the target impedance data, the first noise data is determined as the second noise data; When the impedance attenuation characteristic is greater than or equal to the target impedance data, noise correction data is obtained based on the impedance attenuation characteristic, the preset correction data, and the first noise data. The second noise data is obtained based on the first noise data and the noise correction data.
3. The method for ensuring the quality of intercom channels based on bidirectional audio stream delay jitter according to claim 1, characterized in that, The step of generating target processing instructions for the audio data based on the second noise data and the jitter characteristics includes: When the second noise data is greater than or equal to the target noise data, a weighted feature is obtained based on the impedance attenuation feature and the second noise data, and the preset first threshold data is adjusted based on the weighted feature to obtain the second threshold data; If the jitter feature is less than the second threshold data, the allocation ratio of the audio data between the first processing module and the second processing module is determined based on the weighted feature.
4. The method for ensuring the quality of intercom channels based on bidirectional audio stream delay jitter according to claim 3, characterized in that, The step of obtaining the weighted features based on the impedance attenuation characteristics and the second noise data includes: The weighted feature is obtained by multiplying the impedance attenuation feature with the second noise data. The step of adjusting the preset first threshold data based on the weighted features to obtain the second threshold data includes: Based on the weighted features, the maximum weighted features, and the lifting coefficient data, the lifting ratio data is obtained; The second threshold data is obtained based on the lifting ratio data and the first threshold data.
5. The method for ensuring the quality of intercom channels based on bidirectional audio stream delay jitter according to claim 3, characterized in that, The method further includes: If the jitter characteristic is greater than or equal to the second threshold data, a switching instruction is generated; Acquire the first displacement data of the acquisition device, and determine the target time information based on the first displacement data; At the target time information, the transmission path of the audio data is switched from the first transmission module to the second transmission module.
6. The method for ensuring the quality of intercom channels based on bidirectional audio stream delay jitter according to claim 5, characterized in that, Determining the target time information based on the first displacement data includes: Interpolation processing is performed on the first displacement data of adjacent sampling periods to obtain the target time information, which corresponds to the time when the displacement value is zero; The step of switching the transmission path of the audio data from the first transmission module to the second transmission module at the target time information includes: Within the first time information containing the target time information, the audio data sending terminal is redirected to the second transmission module, and the synchronous reception state of the first transmission module and the second transmission module is maintained. When the first time information ends, a disconnection command is generated to disconnect the receiving state of the first transmission module.
7. The method for ensuring the quality of intercom channels based on bidirectional audio stream delay jitter according to claim 1, characterized in that, Before acquiring the impedance attenuation characteristics of the target terminal's acquisition device, the method further includes: Generate a first control instruction, which is used to control the playback module of the target terminal to output a first audio signal; Acquire the second displacement data of the acquisition device collected by the sensing module; The first amplitude feature corresponding to the target frequency information is determined based on the second displacement data; The impedance attenuation characteristic is obtained based on the first amplitude characteristic and the second amplitude characteristic.
8. The method for ensuring the quality of intercom channels based on bidirectional audio stream delay jitter according to claim 1, characterized in that, Before acquiring the first noise data of the target environment, the method further includes: Acquire the second audio data collected by the target terminal; If the second audio data satisfies the first condition data, the second audio data is subjected to frequency domain transformation to obtain spectrum data; Based on the spectral data, the full-band energy characteristics are determined to obtain the first noise data.
9. The method for ensuring the quality of intercom channels based on bidirectional audio stream delay jitter according to claim 1, characterized in that, The step of acquiring the jitter features of the audio data corresponding to the target terminal includes: Acquire the first jitter data and the second jitter data of the audio data; The jitter characteristics of the audio data are determined based on the maximum value between the first jitter data and the second jitter data.
10. A two-way audio channel quality assurance system, used to implement the two-way audio stream delay jitter-based two-way audio channel quality assurance method according to any one of claims 1-9, characterized in that, include: The first acquisition module is used to acquire the impedance attenuation characteristics of the acquisition device of the target terminal and the first noise data of the target environment; The first correction module is used to process the first noise data according to the impedance attenuation characteristics to obtain the second noise data; The second acquisition module is used to acquire the jitter characteristics of the audio data corresponding to the target terminal; The target generation module is used to generate target processing instructions for the audio data based on the second noise data and the jitter characteristics.