A petrochemical IP broadcast audio packet loss compensation method based on multi-modal perception
Patent Information
- Application Number
- CN202611298395.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-08-26
- Publication Date
- 2026-09-22
AI Technical Summary
[0007]本发明的目的在于提供一种基于多模态感知的石化IP广播音频丢包补偿方法,以解决背景技术中提出的现有技术中,石化IP广播传输面临防护滞后、环境误判、干扰盲区、环网中断、识别低、决策单一、硬切换故障及低算力冲突的问题
本发明突破传统事后补救模式,依托物理层链路劣化特征实现故障前置预判,使断音故障大幅降低;通过温湿度基线补偿消除宽温高湿环境漂移,应急语音识别准确率显著提升,并能精准区分变频器周期性干扰与雷击、分合闸等瞬时突发干扰,覆盖厂区主要电磁干扰场景;针对MRP环网,提前预判切换并启用冗余传输,将音频中断时长由数百毫秒压缩至数十毫秒以内,有效规避环网自愈断音问题;双重音频校验机制可过滤工业噪声,确保应急语音优先处理;多因子加权分级决策按需匹配防护策略,大幅提升带宽资源利用率;四段式时序可控平滑切换机制有效降低策略切换引起的卡顿爆音;整套方案仅需PHY硬件原生数据运算,无需海量样本与离线训练,适用于石化防爆现场的低算力嵌入式网关,落地成本低。
Smart Images

Figure CN122802487A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of broadcast communication technology, specifically a method for compensating for packet loss in petrochemical IP broadcast audio based on multimodal perception. Background Technology
[0002] Petrochemical production plants deploy numerous high-power frequency converters and high-voltage electrical equipment, resulting in complex electromagnetic interference from frequent motor starts and stops, circuit breaker opening and closing, and lightning strikes. On-site equipment operates in a wide temperature range (-40℃ to 70℃) and high-humidity corrosive environment, making the transmission baseline prone to drift. The plant's industrial network generally employs the Media Redundancy Protocol (MRP) as specified in IEC 62439-2 to construct redundant ring networks. The IP broadcast system handles production scheduling, safety warnings, and emergency incident broadcasting, placing high demands on the continuity and reliability of audio transmission.
[0003] In existing technologies, solutions for ensuring audio transmission in industrial IP broadcasts mainly fall into the following categories: The first category is receiver-side voice packet loss compensation schemes based on deep learning. For example, Chinese invention patent application CN115171705A discloses a method for voice packet loss compensation, a method and apparatus for voice communication, which predicts the waveform of lost packets by using parameters of normally received packets and reconstructs lost frames at the receiver. Chinese invention patent application CN118136026A discloses a voice packet loss compensation method based on neural networks, which uses deep neural networks to predict features and reconstruct signals from lost audio frames. The above schemes have the following drawbacks: they all perform passive reconstruction at the receiver after packet loss occurs, which is a post-event remediation mode, making it difficult to take preventive measures before packet loss occurs; they rely on deep learning models and a large number of training samples, resulting in a large number of model parameters and computational load, making it difficult to adapt to the deployment conditions of low-computing-power embedded gateways in petrochemical explosion-proof sites.
[0004] The second category is narrowband audio transmission schemes based on audio coding compression and redundant transmission. For example, Chinese invention patent application CN121585717A discloses a real-time multi-channel audio transmission and synchronization method based on Ethernet, which realizes Ethernet transmission of multiple audio channels through device discovery and clock synchronization mechanisms; the combined scheme of Forward Error Correction (FEC) and Automatic Repeat-reQuest (ARQ) resists packet loss by adding redundant check packets or requesting retransmission of lost packets at the sending end. The above scheme reduces bandwidth consumption through compression coding and improves anti-packet loss capability through redundancy mechanisms, but it has the following defects: it is not specifically adapted for the extreme working conditions of wide temperature and humidity and strong electromagnetic interference in petrochemical plants, it does not involve the pre-judgment of physical layer underlying data, and it does not integrate MRP network status for collaborative protection, and it cannot dynamically adjust the protection strategy according to the on-site electromagnetic environment and network status.
[0005] The third type is the MRP ring network protocol redundancy scheme. The MRP protocol, as specified in the IEC 62439-2 standard, monitors link status through a ring network manager and automatically switches to a backup link when the primary link fails, achieving redundancy protection for industrial Ethernet. However, this scheme has the following drawbacks: it only focuses on link connectivity and self-healing switching at the network layer, without addressing the linkage mechanism with IP broadcast audio packet loss compensation. Furthermore, there is a communication interruption period of tens to hundreds of milliseconds during ring network switching, making it difficult to actively fill the audio transmission interruption during ring network switching.
[0006] In summary, existing industrial IP broadcast audio transmission protection solutions generally suffer from technical problems such as lagging protection, misjudgment or missed judgment caused by environmental drift, blind spots in interference type adaptation, audio interruption caused by ring network switching, low noise recognition rate of emergency voice, secondary faults caused by hard switching of strategies, and conflicts between deep learning solutions and explosion-proof low computing power deployment environments. Summary of the Invention
[0007] The purpose of this invention is to provide a multimodal perception-based audio packet loss compensation method for petrochemical IP broadcasting, in order to solve the problems faced by petrochemical IP broadcasting transmission in the prior art, such as protection lag, environmental misjudgment, interference blind spots, ring network interruption, low recognition, single decision-making, hard handover failure and low computing power conflict.
[0008] To solve the above-mentioned technical problems, the technical solution adopted by the present invention is as follows: A method for compensating for audio packet loss in petrochemical IP broadcasts based on multimodal perception includes the following steps: Link and environment multimodal perception steps: Real-time acquisition of raw bit error data from network physical layer chips and on-site temperature and humidity environmental data; baseline drift compensation of raw bit error data based on temperature and humidity environmental data; removal of environmental drift components; acquisition of electromagnetic interference characteristic data after environmental compensation. Packet loss risk prediction steps: Based on the pure electromagnetic interference characteristic data, match the corresponding interference frequency band characteristics of the frequency converter, evaluate and output the short-term packet loss prediction risk level; Audio modal feature recognition steps: parse IP broadcast multicast audio streams, jointly verify the oscillation rate of audio spectrum energy and audio amplitude, distinguish emergency broadcast voice from industrial noise, and output emergency voice confidence score; Network modality awareness steps: Interact with the protocol stack of the industrial redundant ring network to read the ring network operation status and link failure precursor data in real time, quantify and output network vulnerability scores, and predict network link switching actions in advance. Multimodal information fusion integrated decision-making steps: Integrate three types of heterogeneous data: short-term packet loss prediction risk level, emergency voice confidence level, and network vulnerability score, and adaptively match the corresponding hierarchical packet loss compensation strategy through weighted calculation; Smooth switching of strategy execution steps: When a new packet loss compensation strategy is generated in the multimodal information fusion and comprehensive decision-making step, the audio transmission unit is controlled to perform a smooth switching of the packet loss compensation strategy by executing the new and old strategies in parallel.
[0009] According to the above technical solution, the specific process of removing the environmental drift component in the link and environment multimodal sensing step is as follows: Calculate the original electromagnetic interference intensity index; The temperature drift compensation component is calculated based on the difference between the real-time temperature and the reference calibration temperature, combined with the preset temperature drift compensation coefficient. The humidity drift compensation component is calculated based on the real-time relative humidity on site and the preset humidity drift compensation coefficient. The temperature drift compensation component and the humidity drift compensation component are subtracted from the original electromagnetic interference intensity index to obtain the electromagnetic interference intensity index after temperature and humidity compensation.
[0010] According to the above technical solution, the specific process of assessing and outputting the short-term packet loss risk level in the packet loss risk prediction step is as follows: Determine whether the interference frequency in the pure electromagnetic interference characteristic data is within the preset upper and lower limits of the corresponding interference frequency band of the frequency converter. If it is within the range, calculate the frequency converter interference characteristic matching degree by combining the ratio of the interference amplitude at the current frequency point to the amplitude of the interference-free steady-state baseline. The original electromagnetic interference intensity index and the frequency converter interference characteristic matching degree are weighted and summed to construct an electromagnetic interference intensity index specific to the petrochemical scenario. When the dedicated electromagnetic interference intensity index exceeds the preset interference intensity threshold, the predicted packet loss probability within the future prediction window is obtained based on the exponential decay logic calculation, and this probability is used as the short-term packet loss prediction risk level.
[0011] According to the above technical solution, the specific process of joint verification in the audio modality feature recognition step is as follows: Extract the frequency and real-time amplitude of the current audio sampling point, determine whether the frequency is within the preset emergency voice frequency band, and at the same time determine whether the oscillation rate of the real-time audio amplitude is greater than the preset oscillation threshold. Only when the above two frequency band and oscillation rate verification conditions are met simultaneously, the ratio of the current real-time audio amplitude to the average amplitude of the audio baseline is calculated, and combined with the preset dual verification weight parameters, a standardized emergency voice confidence parameter with a value range of zero to one is output.
[0012] According to the above technical solution, the specific process of predicting the switching action of network links in advance in the network modality perception step is as follows: Real-time monitoring of current ring network link quality indicators; when the ring network link quality indicators are determined to be lower than the warning threshold, the handover prediction logic is triggered. The ring network handover prediction probability is calculated by combining the ratio of the difference between the real-time ring network handover delay and the maximum allowable handover delay. When it is confirmed that a link switching action is about to occur, a dual-path multicast redundant transmission command is issued in advance to fill the communication interruption period during the switching process before the network protocol actually performs the physical switch.
[0013] According to the above technical solution, the specific process of weighted calculation in the multimodal information fusion and comprehensive decision-making step is as follows: The emergency voice confidence level, the short-term packet loss prediction risk level, and the network status information derived from the network vulnerability score are multiplied by their respective independent weight coefficients and then weighted and summed to calculate a comprehensive protection decision score in the range of zero to one. Based on the different numerical ranges of the comprehensive protection decision score, different levels of combined compensation strategies are invoked, including forward error correction coding strength adjustment, dual-path multicast redundancy activation, and automatic retransmission request switch control functions.
[0014] Based on the above technical solution, the corresponding hierarchical packet loss compensation strategy is specifically divided into three levels: high, medium, and low adaptive scheduling. When the comprehensive protection decision score is greater than or equal to the set high-risk score boundary, or when the system receives a fire emergency broadcast trigger signal, the first-level high protection strategy is executed, dual-path multicast redundant transmission is enabled and high-intensity forward error correction coding is used. When the comprehensive protection decision score is in the set medium risk range, the level 2 medium protection strategy is implemented, using adaptive forward error correction coding in conjunction with the backup automatic retransmission request mechanism. When the comprehensive protection decision score is less than the set low-risk score boundary, a level-three conventional protection strategy is implemented, and basic compensation is performed through the audio error hiding mechanism at the receiving end.
[0015] According to the above technical solution, the strategy smooth switching execution step specifically adopts a four-stage parallel transition execution mechanism: The steps, including the parallel execution of the old and new packet loss compensation strategies within a preset time period, are executed sequentially. Perform the audio cache quality comparison step; Perform a linear crossover transition step between the old and new audio amplitudes within a preset time period; Finally, a smooth shutdown procedure using the old compensation strategy is executed to ensure continuous output of the IP broadcast audio stream throughout the entire strategy switching cycle.
[0016] According to the above technical solution, the specific calculation source for the original electromagnetic interference intensity index is as follows: Obtain the ratio of the link bit error rate collected in real time at the network physical layer to the pre-calibrated interference-free baseline bit error rate; Synchronously acquire the ratio of the number of cyclic redundancy check error frames within a unit statistics window to the total number of transmitted frames within that window; The two ratios mentioned above are assigned to set evaluation weight coefficients and then summed using a weighted average to serve as the basic data source for the original electromagnetic interference intensity index.
[0017] According to the above technical solution, the specific parameters for quantifying and outputting the network vulnerability score in the network modality sensing step are based on: Obtain the ratio of normally online nodes to the total number of nodes in the redundant ring network through the protocol stack; Obtain a parameter that characterizes the available margin of the current ring network handover delay. This parameter is the difference between the real-time ring network handover delay and the proportion of the maximum allowable handover delay. Obtain the ring network state correction coefficient corresponding to the current network operating condition, multiply the ratio data, available margin parameters and ring network state correction coefficient to obtain the ring network redundancy reliability score, and output the complementary value of the reliability score as the network vulnerability score.
[0018] Compared with the prior art, the present invention has the following beneficial effects: This invention breaks through the traditional post-event remediation mode, relying on the physical layer link degradation characteristics to achieve pre-fault prediction, significantly reducing audio interruption faults; by eliminating drift in wide temperature and high humidity environments through temperature and humidity baseline compensation, the accuracy of emergency voice recognition is significantly improved, and it can accurately distinguish between periodic interference from frequency converters and instantaneous sudden interference such as lightning strikes and circuit breaker opening and closing, covering the main electromagnetic interference scenarios in the plant area; for MRP ring networks, it predicts switching in advance and enables redundant transmission, compressing the audio interruption time from hundreds of milliseconds to less than tens of milliseconds, effectively avoiding the ring network self-healing audio interruption problem; a dual audio verification mechanism can filter industrial noise to ensure priority processing of emergency voices; multi-factor weighted hierarchical decision-making matches protection strategies on demand, greatly improving bandwidth resource utilization; a four-stage time-controlled smooth switching mechanism effectively reduces stuttering and popping sounds caused by strategy switching; the entire solution only requires native data computation of PHY hardware, without the need for massive samples and offline training, and is suitable for low-computing-power embedded gateways in petrochemical explosion-proof sites, with low implementation costs. Attached Figure Description
[0019] Figure 1 This is an overall flowchart of the compensation method of the present invention; Figure 2 This is a flowchart illustrating the specific implementation of the compensation method of the present invention. Detailed Implementation
[0020] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0021] Example 1:
[0022] Figure 1 and Figure 2As shown, in one specific embodiment, the present invention provides a method for compensating for audio packet loss in petrochemical IP broadcasts based on multimodal perception. The hardware architecture of this method includes an explosion-proof IP broadcast gateway (with built-in PHY chip, MRP protocol stack, and temperature and humidity acquisition module), a multimodal perception fusion unit, an audio encoding and decoding unit, a multimodal fusion decision operation unit, and a dual-channel multicast transmission unit. Specifically, the system and method can be deployed in two typical petrochemical industrial conditions: one is a large-scale petrochemical refining main unit, which is deployed in an explosion-proof cabinet, with an operating environment temperature range of -40℃ to 70℃ and a relative humidity of 30% to 95%RH, adopting an MRP ring network architecture of IEC 62439-2 standard (24 ring network nodes), with an audio sampling rate preferably of 16kHz and a single frame audio data packet length of 128 bytes; the other is a small and medium-sized petrochemical tank area, which adopts a lightweight, low-computing-power embedded gateway and a simplified MRP ring network topology, with audio parameters consistent with the main unit, and only minor adjustments made to some operation parameters.
[0023] Furthermore, in order to ensure that all baseline parameters conform to the interference-free steady-state operating conditions on site, this embodiment sets standardized parameter calibration rules: after the system has been running stably and continuously for 72 hours after power-on, steady-state data of each dimension are collected and the statistical average value is taken to complete the on-site baseline calibration.
[0024] Preferably, to avoid long-term baseline drift caused by device aging or fiber optic connector contamination, the system introduces a slow adaptive baseline tracking mechanism. Interference-free baseline bit error rate. With audio baseline average amplitude It is not fixed, but rather updated according to the following formula during the time window when the system does not trigger any level 1 or level 2 protection strategies (i.e., the silent steady-state period):
[0025]
[0026] In the formula, For the updated baseline parameters, For the current real-time baseline parameters, These are the historical baseline parameters from the previous moment. This is the slow update time constant. Wherein, the time constant... The value is set to 0.001, corresponding to an update cycle of seconds, to ensure that the baseline parameters can change slowly with the environment without responding to sudden disturbances.
[0027] like Figure 2 As shown, based on the above hardware and environment architecture, this invention iteratively executes the following multimodal fusion closed-loop protection process: Step 1: Link-Environment Dual-Modal Sensing and Environmental Drift Compensation. Considering that the wide temperature and high humidity environment in petrochemical sites easily causes baseline drift in the underlying semiconductor layer, this step overcomes the limitations of monitoring only a single communication level. Specifically, it reads the bit error rate and verification data of the PHY chip in real time and incorporates environmental variables for fusion.
[0028] First, calculate the original electromagnetic interference intensity index. The mathematical formula is as follows:
[0029]
[0030] In the above formula: This is the real-time time variable for the system's current continuous operation; In order to be in The original electromagnetic interference intensity index at any given moment, without environmental correction; The evaluation weight coefficient based on link error characteristics is preferably set to 0.5 in the refining and tank farm scenarios; In order to be in Real-time link bit error rate read from the underlying hardware at all times; The preferred calibration value is the interference-free baseline bit error rate obtained through 72-hour calibration. ; The evaluation weighting coefficient based on the cyclic redundancy check error characteristics is preferably set to 0.5; In order to be in The number of cyclic redundancy check error frames generated within a time unit statistics window (fixed at 100ms); The total number of data frames transmitted within the statistical window for this unit; To prevent the compensation term of a very small constant in the denominator or calculation term from being zero, it is preferable to take a value of [value missing]. .
[0031] Specifically, it refers to the total number of data frames actually received by the physical layer chip within the unit statistical window, rather than the expected number of frames received. If the actual number of frames received within the window is zero, the CRC error rate calculation for that window is skipped, and the statistical value of the previous valid window is maintained to avoid calculation anomalies caused by a denominator of zero.
[0032] Furthermore, the duration of the unit statistical window is fixed at 100ms and strictly synchronized with the interrupt sampling period of the PHY chip. This 100ms window duration is set based on the integer multiple of the power frequency interference (50Hz and its harmonics) in the petrochemical plant area (100ms is 5 complete cycles of 50Hz), ensuring that the number of cyclic redundancy check error frames can fully cover the periodic fluctuations of the power frequency interference and avoid statistical deviations caused by truncating the interference cycle.
[0033] Furthermore, to effectively isolate the environmental drift component, a temperature and humidity compensation algorithm is introduced. The mathematical formula for calculating the pure electromagnetic interference intensity after temperature and humidity compensation is as follows:
[0034]
[0035] In the above formula, The pure electromagnetic interference intensity index after eliminating the influence of wide temperature and high humidity environments; To match the temperature drift compensation coefficient of the underlying hardware characteristics, a value of 0.01 / ℃ is taken in large-scale refining and chemical plants, and a value of 0.008 / ℃ is taken in lightweight tank areas; The real-time temperature parameters collected by the sensor at the site; The reference ambient temperature during the calibration process is fixed at 25℃. The humidity drift compensation coefficient is set at 0.002 / %RH in large-scale refining and chemical plants and 0.0015 / %RH in lightweight tank areas. This refers to the real-time relative humidity parameters collected synchronously by the sensor.
[0036] Temperature drift compensation coefficient Humidity drift compensation coefficient The value is calculated based on the temperature frequency stability parameter (typically ±50ppm / ℃) and humidity influence coefficient provided in the datasheet of the selected physical layer chip (PHY). Furthermore, the calibration dataset is fitted and optimized using the gradient descent method under corresponding petrochemical conditions, with the goal of minimizing the sum of squared residuals after temperature and humidity compensation to determine the final value. For different chip models or operating conditions, those skilled in the art can refit the data using the above method. and The specific value.
[0037] Preferably, to prevent compensation overshoot in extreme temperature and humidity scenarios (such as a refining unit with a high temperature of 70°C and 95%RH) from causing the electromagnetic interference intensity index to drop to zero or become negative, thus distorting its physical meaning, this embodiment applies dynamic lower limit clamping protection to the compensated pure electromagnetic interference intensity index. The processing logic is as follows: If... Then force take To preserve the inherent minimum noise floor of the underlying physical layer, ensuring that subsequent packet loss prediction logic does not fail due to the value returning to zero.
[0038] Step Two: Electromagnetic Interference Feature Matching and Packet Loss Risk Prediction. To accurately distinguish between the differential interference caused by the periodic operation of large frequency converters and transient line impacts, this embodiment performs spectral feature mapping on the aforementioned compensated clean data.
[0039] Specifically, the formula for calculating the inverter interference characteristic matching degree is as follows:
[0040]
[0041] In the above formula: The characteristic matching degree of the frequency converter is calculated in real time; This represents the total number of sampling points for the spectrum or time-domain sequence; here, it is uniformly set to 15. This refers to the index variable of discrete sampling points in the corresponding statistical sequence; This is a logical indicator function. The value of this item is 1 when the logical condition in parentheses is true, and 0 when it is false. For the first Interference frequency data captured at each sampling frequency point; and These are the lower and upper limits of the frequency band corresponding to the interference frequency band of the petrochemical inverter, respectively, with fixed values of 4kHz and 10kHz. For the first The interference amplitude measured at each frequency point; This represents the amplitude of the uninterrupted steady-state baseline obtained through on-site calibration.
[0042] Preferably, by integrating the original indicators of the link with the above-mentioned matching features, a petrochemical-specific electromagnetic interference intensity assessment model is constructed:
[0043]
[0044] In the above formula, This is a petrochemical-specific electromagnetic interference intensity index characterized by comprehensive three-dimensional underlying data; it should be noted that in this three-dimensional fusion model, , With new parameters Together they constitute the three-dimensional weight coefficients, with a total weight of 1. These are the weighted components of the frequency conversion interference matching characteristics. In the refining and tank farm scenarios, the optimal combination of all three is [missing information]. , , .
[0045] In one specific implementation, the exponential decay calculation formula for predicting the short-term packet loss probability based on the aforementioned dedicated index is as follows:
[0046]
[0047] In the above formula, To predict the risk level of short-term packet loss within the advance prediction window using quantitative output; The lead time for advance prediction is uniformly fixed at 50ms. It is an exponential function with the natural constant as its base; The threshold value for interference intensity used to trigger the calculation of packet loss probability is 0.6 for refining and chemical units and 0.55 for tank farms.
[0048] Given that embedded gateways in petrochemical explosion-proof sites generally use ARM Cortex-M series processors with a main frequency of less than 500MHz, in order to reduce the load and power consumption caused by directly performing exponential floating-point operations, this embodiment replaces the above-mentioned exponential decay logic with a linear piecewise approximation algorithm based on lookup table (LUT). That is, the ratio... exist The numerical range is divided into 16 equal segments, and each segment has a pre-stored corresponding exponential decay value. The runtime system only needs to perform one division (or equivalent shift operation) table lookup and one multiplication operation to complete the packet loss probability mapping, so that the computation time of a single prediction decision is controlled within 50 microseconds, which can be adapted to low computing power embedded environments.
[0049] Step 3: Audio Modal Feature Verification and Emergency Voice Confidence Recognition. Addressing the technical challenge of misidentification caused by low-frequency noise generated by on-site pumps and other equipment, this step combines audio frequency band energy and vibration rate as dual perspectives for multi-dimensional verification.
[0050] Specifically, the formula for calculating the confidence level of emergency voice messages is as follows:
[0051]
[0052] In the above formula, Standardize the confidence level of emergency voice messages in the final output range between 0 and 1; The total number of audio temporal feature sampling points is uniformly set to 20; This is the first verification weighting coefficient for the frequency band energy dimension, with a preferred value of 0.6; For the first The instantaneous frequency of each audio sampling point; and These are the lower and upper limits of the dedicated core frequency band for emergency voice communication, fixed at 500Hz and 2000Hz respectively. This represents the real-time audio amplitude at the corresponding sampling point; The average amplitude of the audio baseline maintained by the system; The second verification weighting coefficient for the oscillation rate dimension is preferably set to 0.4; It is the first derivative of the real-time audio amplitude with respect to time, i.e., the audio oscillation rate. To identify the oscillation judgment threshold for voice change characteristics, a value of 50ms was set for the refining unit and a value of 45ms was set for the tank area.
[0053] In the embedded digital signal processing implementation of this embodiment, the start-up rate of the audio amplitude Instead of continuous differentiation in the analog domain, the calculation is performed using the forward finite difference method. The specific digital calculation formula is as follows:
[0054]
[0055] in, The audio amplitude at the current discrete sampling point. The audio amplitude of the previous adjacent sampling point. The sampling period of the system's audio analog-to-digital converter (corresponding to a sampling rate of 16kHz, i.e.) To enhance noise immunity and reduce the computational load on low-performance gateways, the system does not perform calculations at each raw sampling point. Instead, it uses the average amplitude within a fixed time window (preferably 2ms) as the basis for calculation. The input is filtered out to eliminate the interference of high-frequency random noise on the oscillation rate judgment.
[0056] Audio baseline average amplitude It is not fixed, but updated using the same slow adaptive tracking mechanism as the aforementioned bit error rate baseline. Specifically, during the time window when the system does not detect emergency voice features (i.e., the silent period), it is updated according to the formula:
[0057]
[0058] In the formula, The updated audio baseline average amplitude, The current audio baseline amplitude, The measured amplitude of the audio obtained through real-time sampling. Update the weighting coefficients for the audio baseline.
[0059] Update, in which Values This corresponds to a minute-level update speed, ensuring that the baseline can follow the slow changes in ambient background noise without responding to sudden audio events.
[0060] The selection of the dedicated core frequency band for emergency voice communication, 500Hz~2000Hz, is based on the following: the fundamental frequency and its lower harmonics of the main industrial noise sources (pumps, compressors, fans, etc.) in petrochemical plants are mostly concentrated below 500Hz, while the energy of human voice is most concentrated in the 500Hz~2000Hz resonant peak range. Therefore, the two are separable in the frequency domain. Based on statistical analysis of the spectral characteristics of these two types of sound sources, this embodiment determines that 500Hz is the lower limit and 2000Hz is the upper limit to constitute the dedicated discrimination frequency band for emergency voice communication.
[0061] Oscillation Judgment Threshold The different values for the refining unit and tank farm scenarios are based on the differences in background noise characteristics between the two types of scenarios: In the refining unit area, rotating equipment such as pumps and compressors are densely packed, resulting in a higher background noise onset rate, requiring a higher threshold (50ms) to avoid misinterpreting equipment startup noise as speech; In the tank farm area, background noise is relatively sparse, so the threshold is appropriately narrowed to 45ms to more sensitively capture the leading edge of sudden emergency speech. The above values are all based on the statistical optimization results of no less than 1000 sets of on-site measured samples in the corresponding scenarios.
[0062] Step 4: Network Modal State Awareness and MRP Network Vulnerability Scoring. To avoid communication interruptions caused by physical switching of network links, this step makes the ring network vulnerability data explicit by interfacing with the underlying protocol stack.
[0063] Specifically, the formula for scoring the vulnerability of a ring network is as follows:
[0064]
[0065] In the above formula, The network vulnerability score, which characterizes the current MRP ring network's resilience, ranges from 0 to 1. This refers to the number of currently active online ring network nodes reported by the protocol stack. This represents the total number of registered nodes in the network topology. This refers to the real-time ring network self-healing handover delay time reported by the underlying protocol stack. The maximum allowable switching delay boundary to be tolerated by the system is 100ms, based on the typical switching performance of the MRP standard protocol (IEC 62439-2) and the worst-case test results under harsh petrochemical conditions. The ring network state correction coefficient is set according to different network health levels, and its specific value is determined by the real-time mapping of the ring network protocol stack state machine.
[0066] Specifically, the ring network state correction coefficient The real-time mapping logic is as follows: The system obtains the current network health by parsing the standardized status messages of the industrial redundant ring network protocol stack. When the protocol stack reports the current ring network status as "Closed / RingOK" and there are no fault alarms, The value is 1.0; when the protocol stack reports "Link Loss" or "Ring Open", The value is 0.6; when the protocol stack reports "MultipleFailures" or "ManagerFailed", The value is 0.2; when the protocol stack reports "ring network test frame timeout (TestTimeout)", if it lasts for more than 5ms, it is classified as a single link failure.
[0067] Furthermore, the prediction formula for quantifying the switching trigger probability is as follows:
[0068]
[0069] In the above formula, This represents the predicted probability of a link switching action occurring at a future time. To monitor and extract the comprehensive quality attenuation index of the ring network links in real time; This is the threshold for triggering link quality early warnings obtained through on-site steady-state calibration.
[0070] When a link switching action is confirmed to be imminent (i.e., when the predicted probability value is greater than 0.5), the system issues a dual-path multicast redundant transmission command 20ms before the actual physical switching of the network protocol. This advance is based on the typical inherent delay of MRP ring network switching (about 10~15ms) plus a 5ms margin, ensuring that the dual-path redundancy command takes effect before the physical switching, effectively filling the communication interruption period during the switching process.
[0071] Comprehensive quality degradation index of ring network links The following sub-indicators are weighted and combined: Ring network port receive bit error rate Ring network management frame timeout rate Ring network port signal strength attenuation value The specific calculation formula is as follows:
[0072]
[0073] In the formula, This is a real-time comprehensive quality attenuation index for ring network links; Real-time ring network port bit error rate; The steady-state baseline bit error rate of the ring network; Real-time timeout rate for ring network management frames; This represents the real-time signal strength attenuation value at the port. This is a reference value for the port's rated signal strength. , , These are the weighted coefficients for the three sub-indicators.
[0074] in , , These are the weighting coefficients for each sub-indicator. The steady-state baseline bit error rate of the ring network port. This is a reference value for the port signal strength. When Below the calibrated threshold When (preferred value 0.3), the switching prediction logic is triggered.
[0075] Step 5: Multimodal Data Weighted Fusion and Hierarchical Protection Decision. Based on the heterogeneous quantified data collected and processed in the aforementioned serial steps, this step serves as the decision-making center, projecting all states onto the same decision evaluation space. Specifically, the comprehensive decision score calculation formula is as follows:
[0076]
[0077] In the above formula, The final comprehensive protection decision score used to guide system behavior ranges from 0 to 1; To assign a multi-factor fusion weight to prioritize emergency voice messages, the value is 0.35 for refining and chemical plants and 0.40 for tank farms. To assign a multi-factor fusion weight to predict the risk level of short-term packet loss, the weight is 0.35 for refining and chemical processing and 0.30 for tank farms. To impart vulnerability to network topology (i.e.) The multi-factor fusion weight is set to 0.30 for refining and 0.30 for tank farms. And under the three configurations... It is always equal to 1.
[0078] In one specific implementation, based on this score, the system adaptively matches a three-level strategy: when Or, when a fire emergency broadcast signal is captured, the first-level strategy (dual-channel multicast redundancy + high-intensity FEC) is executed; when When, execute the secondary strategy (adaptive FEC + standby ARQ); when At that time, the three-level conventional strategy (basic receiver error concealment) is executed.
[0079] In this embodiment, the high-intensity forward error correction coding specifically adopts a redundant packet generation mechanism based on Reed-Solomon codes (RS(256,224)), that is, adding 32 check redundancy packets to every 224 original audio payload packets, so that the redundancy reaches 12.5%; the adaptive forward error correction coding adopts RS(256,240), and the redundancy is reduced to 6.25%. The specific implementation of dual-channel multicast redundant transmission is as follows: the same IP broadcast audio stream is encapsulated into two multicast streams with the same SSRC but different destination IP addresses. One stream is sent through the primary ring network port, and the other stream is sent through the backup ring network port. When constructing redundant multicast packets, the transmitting unit uniformly writes a 64-bit hardware timestamp synchronized based on the IEEE 1588 precision clock synchronization protocol into the RTP extension header. The receiving end uses this timestamp as the sole comparison basis and discards old redundant packets whose timestamps lag behind the current system time by more than 50ms.
[0080] The FEC coding strength switching between the Level 1 high protection strategy and the Level 2 medium protection strategy is based on the comprehensive protection decision score. Dynamic changes trigger: when When the risk level rises from the medium-risk range to the high-risk range, the FEC encoding switches from RS(256,240) to RS(256,224); conversely, it switches from high intensity to adaptive intensity. During the switching process, the old and new FEC encoding parameters use the same four-segment parallel transition logic as in step six to ensure that no additional audio interruptions or decoding errors are introduced during the switching.
[0081] The receiver audio error concealment mechanism in the three-level conventional protection strategy specifically adopts the pitch waveform replication method based on adjacent correct packets: when the current audio packet is detected to be lost, the receiver calculates the pitch period based on the two most recently received correct audio packets, copies the audio waveform of the previous correct packet, stretches it in the time domain according to the pitch period, and fills it into the position of the lost packet. At the same time, a cosine attenuation window (with a window length of 10ms) is applied to the amplitude of the filled waveform to smooth the splicing boundary and avoid introducing abrupt noise.
[0082] Step Six: Four-Stage Parallel Transition Execution Mechanism. To effectively eliminate secondary faults such as audio popping and stuttering caused by hard switching of strategies, this embodiment abandons the traditional instantaneous switching mode. Specifically, a four-stage parallel transition execution mechanism is adopted. In the implementation scenario of a refining unit, the system executes the following sequentially: a 20ms parallel operation buffering phase for the old and new strategies, a 10ms concurrent audio stream quality comparison phase, a 30ms linear crossover transition phase for audio amplitude, and finally, a smooth shutdown phase for the old strategy. Preferably, in the lightweight tank area scenario, the timing parameters of the above four stages are adjusted to: 15ms parallel buffering, 8ms quality comparison, 25ms linear transition, and shutdown action. The timing buffering effectively ensures the continuous output of the audio stream and effectively completes the adaptive protection closed loop.
[0083] During the audio buffer quality comparison step, the system simultaneously monitors the buffer levels and continuous packet loss rates of both the old and new processing links. The specific decision logic is as follows: only when the buffer level of the new strategy link reaches more than 80% of the old strategy link's buffer level, and the instantaneous packet loss rate of the new link is lower than that of the old link, is the subsequent linear crossover transition phase allowed. If the new link's buffer is insufficient, the system forcibly extends the parallel operation time of the old and new strategies (up to a maximum of 40ms) until the transition admission conditions are met, thereby avoiding audio gaps caused by a forced switch before the new strategy has established a link.
[0084] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus.
[0085] Finally, it should be noted that the above descriptions are merely preferred embodiments of the present invention and are not intended to limit the present invention. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing embodiments or make equivalent substitutions for some of the technical features. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.
Claims
1. A method for compensating for audio packet loss in petrochemical IP broadcasting based on multimodal perception, characterized in that: Includes the following steps: Link and environment multimodal perception steps: Real-time acquisition of raw bit error data from network physical layer chips and on-site temperature and humidity environmental data; baseline drift compensation of raw bit error data based on temperature and humidity environmental data; removal of environmental drift components; acquisition of electromagnetic interference characteristic data after environmental compensation. Packet loss risk prediction steps: Based on the pure electromagnetic interference characteristic data, match the corresponding interference frequency band characteristics of the frequency converter, evaluate and output the short-term packet loss prediction risk level; Audio modal feature recognition steps: parse IP broadcast multicast audio streams, jointly verify the oscillation rate of audio spectrum energy and audio amplitude, distinguish emergency broadcast voice from industrial noise, and output emergency voice confidence score; Network modality awareness steps: Interact with the protocol stack of the industrial redundant ring network to read the ring network operation status and link failure precursor data in real time, quantify and output network vulnerability scores, and predict network link switching actions in advance. Multimodal information fusion integrated decision-making steps: Integrate three types of heterogeneous data: short-term packet loss prediction risk level, emergency voice confidence level, and network vulnerability score, and adaptively match the corresponding hierarchical packet loss compensation strategy through weighted calculation; Smooth switching of strategy execution steps: When a new packet loss compensation strategy is generated in the multimodal information fusion and comprehensive decision-making step, the audio transmission unit is controlled to perform a smooth switching of the packet loss compensation strategy by executing the new and old strategies in parallel.
2. The method for compensating for audio packet loss in petrochemical IP broadcasting based on multimodal perception according to claim 1, characterized in that: In the link and environment multimodal sensing steps, the specific process of removing the environment drift component is as follows: Calculate the original electromagnetic interference intensity index; The temperature drift compensation component is calculated based on the difference between the real-time temperature and the reference calibration temperature, combined with the preset temperature drift compensation coefficient. The humidity drift compensation component is calculated based on the real-time relative humidity on site and the preset humidity drift compensation coefficient. The temperature drift compensation component and the humidity drift compensation component are subtracted from the original electromagnetic interference intensity index to obtain the electromagnetic interference intensity index after temperature and humidity compensation.
3. The method for compensating for audio packet loss in petrochemical IP broadcasting based on multimodal perception according to claim 2, characterized in that: The specific process for assessing and outputting the short-term packet loss risk level in the packet loss risk prediction step is as follows: Determine whether the interference frequency in the pure electromagnetic interference characteristic data is within the preset upper and lower limits of the corresponding interference frequency band of the frequency converter. If it is within the range, calculate the frequency converter interference characteristic matching degree by combining the ratio of the interference amplitude at the current frequency point to the amplitude of the interference-free steady-state baseline. The original electromagnetic interference intensity index and the frequency converter interference characteristic matching degree are weighted and summed to construct an electromagnetic interference intensity index specific to the petrochemical scenario. When the dedicated electromagnetic interference intensity index exceeds the preset interference intensity threshold, the predicted packet loss probability within the future prediction window is obtained based on the exponential decay logic calculation, and this probability is used as the short-term packet loss prediction risk level.
4. The method for compensating for audio packet loss in petrochemical IP broadcasting based on multimodal perception according to claim 1, characterized in that: In the audio modal feature recognition step, the specific process of joint verification is as follows: Extract the frequency and real-time amplitude of the current audio sampling point, determine whether the frequency is within the preset emergency voice frequency band, and at the same time determine whether the oscillation rate of the real-time audio amplitude is greater than the preset oscillation threshold. Only when the above two frequency band and oscillation rate verification conditions are met simultaneously, the ratio of the current real-time audio amplitude to the average amplitude of the audio baseline is calculated, and combined with the preset dual verification weight parameters, a standardized emergency voice confidence parameter with a value range of zero to one is output.
5. The method for compensating for audio packet loss in petrochemical IP broadcasting based on multimodal perception according to claim 1, characterized in that: In the network modality sensing step, the specific process of predicting network link switching actions in advance is as follows: Real-time monitoring of current ring network link quality indicators; when the ring network link quality indicators are determined to be lower than the warning threshold, the handover prediction logic is triggered. The ring network handover prediction probability is calculated by combining the ratio of the difference between the real-time ring network handover delay and the maximum allowable handover delay. When it is confirmed that a link switching action is about to occur, a dual-path multicast redundant transmission command is issued in advance to fill the communication interruption period during the switching process before the network protocol actually performs the physical switch.
6. The method for compensating for audio packet loss in petrochemical IP broadcasting based on multimodal perception according to claim 1, characterized in that: In the multimodal information fusion and integrated decision-making process, the specific process of weighted calculation is as follows: The emergency voice confidence level, the short-term packet loss prediction risk level, and the network status information derived from the network vulnerability score are multiplied by their respective independent weight coefficients and then weighted and summed to calculate a comprehensive protection decision score in the range of zero to one. Based on the different numerical ranges of the comprehensive protection decision score, different levels of combined compensation strategies are invoked, including forward error correction coding strength adjustment, dual-path multicast redundancy activation, and automatic retransmission request switch control functions.
7. The method for compensating for audio packet loss in petrochemical IP broadcasting based on multimodal perception according to claim 6, characterized in that: The corresponding hierarchical packet loss compensation strategy is specifically divided into three levels: high, medium, and low adaptive scheduling. When the comprehensive protection decision score is greater than or equal to the set high-risk score boundary, or when the system receives a fire emergency broadcast trigger signal, the first-level high protection strategy is executed, dual-path multicast redundant transmission is enabled and high-intensity forward error correction coding is used. When the comprehensive protection decision score is in the set medium risk range, the level 2 medium protection strategy is implemented, using adaptive forward error correction coding in conjunction with the backup automatic retransmission request mechanism. When the comprehensive protection decision score is less than the set low-risk score boundary, a level-three conventional protection strategy is implemented, and basic compensation is performed through the audio error hiding mechanism at the receiving end.
8. The method for compensating for audio packet loss in petrochemical IP broadcasting based on multimodal perception according to claim 1, characterized in that: The strategy smooth switching execution step specifically adopts a four-stage parallel transition execution mechanism: The steps, including the parallel execution of the old and new packet loss compensation strategies within a preset time period, are executed sequentially. Perform the audio cache quality comparison step; Perform a linear crossover transition step between the old and new audio amplitudes within a preset time period; Finally, a smooth shutdown procedure using the old compensation strategy is executed to ensure continuous output of the IP broadcast audio stream throughout the entire strategy switching cycle.
9. A method for compensating for audio packet loss in petrochemical IP broadcasting based on multimodal perception according to claim 2, characterized in that: The specific calculation source for the original electromagnetic interference intensity index is: Obtain the ratio of the link bit error rate collected in real time at the network physical layer to the pre-calibrated interference-free baseline bit error rate; Synchronously acquire the ratio of the number of cyclic redundancy check error frames within a unit statistics window to the total number of transmitted frames within that window; The two ratios mentioned above are assigned to set evaluation weight coefficients and then summed using a weighted average to serve as the basic data source for the original electromagnetic interference intensity index.
10. A method for compensating for audio packet loss in petrochemical IP broadcasting based on multimodal perception according to claim 1, characterized in that: In the network modality sensing step, the specific parameters for quantifying the network vulnerability score are based on: Obtain the ratio of normally online nodes to the total number of nodes in the redundant ring network through the protocol stack; Obtain a parameter that characterizes the available margin of the current ring network handover delay. This parameter is the difference between the real-time ring network handover delay and the proportion of the maximum allowable handover delay. Obtain the ring network state correction coefficient corresponding to the current network operating condition, multiply the ratio data, available margin parameters and ring network state correction coefficient to obtain the ring network redundancy reliability score, and output the complementary value of the reliability score as the network vulnerability score.
Citation Information
Patent Citations
Voice packet loss compensation method, voice communication method and device
CN115171705A
Voice packet loss compensation method and device based on neural network
CN118136026A
Real-time multi-channel audio transmission and synchronization method and system based on Ethernet
CN121585717A