A method and system for IP broadcast audio processing in a converged broadcast system
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-13
- Publication Date
- 2026-08-11
AI Technical Summary
然而,该技术方案中,虽然涉及音频采集和处理,但未涉及多终端同步控制、丢包补偿、降噪处理等关键音频处理技术;且仅针对对讲场景设计,未考虑应急广播、消防广播、公共广播等多元广播类型的融合需求;同时,32K固定采样率无法根据终端算力、广播类型及场景动态调整,适配性不足
[0043] The IP broadcast audio processing method and system of the integrated broadcast system (1) achieves seamless integration of multiple types of broadcasts and solves the problem of collaborative management: For the first time, emergency broadcasts, fire broadcasts, public broadcasts and loudspeaker intercoms are included in a unified processing framework. Through priority scheduling, mode switching, intercom and broadcast linkage and other collaborative logic, unified management of all types of broadcasts is achieved, solving the problems of independent deployment of multiple systems, poor coordination and switching lag in the existing technology; it supports precise management of multiple areas, and different types of broadcasts can be sent out simultaneously in different areas, adapting to the needs of multiple scenarios such as campuses, parks, and factories;
Smart Images

Figure CN122554256A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of IP broadcasting technology, and in particular to an IP broadcasting audio processing method and system for a converged broadcasting system. Background Technology
[0002] With the rapid development of IP network technology, IP broadcasting, with its advantages of flexible deployment, wide coverage, and controllable cost, has gradually become the core carrier for the convergence of multiple broadcast types. However, existing IP broadcasting systems still have some shortcomings in audio processing, synchronization control, and packet loss compensation when integrating multiple broadcast types, affecting the overall performance and reliability of the converged broadcasting system.
[0003] A search revealed a cloud broadcasting system and method with publication number CN102291244A. This patent constructs an IP broadcasting system by setting up an application layer, a network service layer, and a terminal layer. The application layer provides various audio applications, the network service layer manages session objects and broadcast terminals, and the terminal layer consists of multiple broadcast terminals to complete the digital-to-analog conversion of audio signals. This cloud broadcasting system breaks through the performance bottleneck of a single server and can support large-scale IP broadcasting networks. However, this technical solution only provides a basic broadcasting system architecture and does not involve specific implementation schemes for multi-terminal synchronization accuracy, does not have targeted designs for packet loss compensation, and does not consider the differentiated needs of different broadcast types such as emergency broadcasting, fire broadcasting, public broadcasting, and intercom. At the same time, the solution adopts a fixed architecture design, which cannot adapt to the computing power limitations of embedded terminals, and lacks a dynamic parameter linkage mechanism between modules, failing to meet the unified management and control requirements of a converged broadcasting system.
[0004] A search revealed a patent for an IP network-based intercom system, publication number CN104468146A. This patent includes an IP network intercom host, IP network intercom extensions, and an IP network address box that enable communication via a local area network (LAN) or wide area network (WAN). The host and extensions each include a voice acquisition module and a core processor integrating a full-duplex echo cancellation circuit. The audio sampling rate is 32kHz. This invention improves the sound quality of full-duplex intercoms. However, while this technical solution involves audio acquisition and processing, it lacks key audio processing technologies such as multi-terminal synchronous control, packet loss compensation, and noise reduction. Furthermore, it is designed only for intercom scenarios and does not consider the integration needs of diverse broadcast types such as emergency broadcasts, fire broadcasts, and public broadcasts. Additionally, the fixed 32kHz sampling rate cannot be dynamically adjusted according to terminal computing power, broadcast type, and scenario, resulting in insufficient adaptability.
[0005] The aforementioned problems indicate that existing IP broadcasting systems still have certain shortcomings in terms of multi-terminal synchronization accuracy, packet loss compensation effectiveness, noise reduction algorithm adaptation, multi-broadcast type integration, and unified management. Therefore, this invention provides an IP broadcasting audio processing method and system for integrated broadcasting systems, aiming to address the deficiencies of existing technologies, improve the overall performance and reliability of integrated broadcasting systems, and meet the needs of integrated systems for emergency response, fire protection, public address systems, and intercom systems. Summary of the Invention
[0006] The purpose of this invention is to provide an IP broadcast audio processing method and system for a converged broadcast system, so as to solve the problems mentioned in the background art.
[0007] To achieve the above objectives, the present invention provides the following technical solution: an IP broadcast audio processing method for a converged broadcast system, wherein the method achieves seamless integration of four major types of broadcasts: emergency broadcast, fire broadcast, public broadcast, and intercom, and is compatible with various embedded terminals ranging from low-end MCUs to high-end ARM / Linux, comprising the following steps:
[0008] S1. Audio Preprocessing and Encoding: The broadcast server collects multiple types of audio signals and adopts a three-dimensional decision model of broadcast type + terminal computing power + scene noise to dynamically adjust the sampling rate of 16000Hz / 32000Hz / 48000Hz. The terminal completes seamless switching and feedback according to the sampling rate instructions issued by the server; the encoding method is dynamically switched according to the broadcast type, and preprocessing noise reduction and intercom-specific echo cancellation are performed.
[0009] S2. Audio Interleaving and FEC Redundancy Generation: The server constructs an adjustable interleaving matrix of order 4-8, shuffles and rearranges continuous audio frames and encapsulates them into RTP packets. Every 4-8 consecutive RTP service frames are grouped together, and FEC redundancy frames are generated through XOR operation. Differentiated FEC redundancy is set according to broadcast type. The FEC redundancy frames and service frames are sent down via UDP multicast with the same identifier.
[0010] S3. Multi-terminal clock synchronization: The server deploys the chronyNTP clock service. The terminal sends synchronization requests and reports its own computing power level according to a dynamic period. It calculates the link delay and clock deviation, performs smoothing correction and long-term drift correction, and triggers secondary synchronization when the deviation of the fire / emergency terminal exceeds 1ms.
[0011] S4. Terminal packet loss detection and hierarchical compensation: A four-layer protection mechanism is adopted, namely "interleaving to prevent sudden bursts → FEC priority error correction → unicast resume transmission → local hierarchical compensation". Error correction and resume transmission are processed according to the priority of broadcast type, and differentiated local compensation is performed according to the packet loss level, broadcast type and sampling rate.
[0012] S5. Terminal audio noise reduction and synchronous playback: The terminal performs three-level lightweight noise reduction on various audio frames, dynamically adjusts the noise reduction parameters to adapt to broadcast type, scene and sampling rate; restores the timing according to the interleaved inverse matrix, dynamically adjusts the adaptive buffer delay, and realizes synchronous playback of multiple terminals and smooth switching of multiple broadcast types.
[0013] S6. Iterative optimization and multi-broadcast type collaborative management: The server processes terminal resume requests according to priority, and the terminal executes the processing flow in a loop; it realizes priority scheduling of multiple broadcast types, mode switching, intercom and broadcast linkage and regional management, and dynamically optimizes the parameters of the entire link.
[0014] In a preferred embodiment of this scheme, the adaptive sampling described in step S1 includes: during the NTP synchronization or login phase, the terminal reports its computing power level through the TCP control channel, and the server determines the target sampling rate according to the priority of "broadcast type > terminal computing power > scene noise";
[0015] Fire / emergency broadcasts should prioritize 32000Hz, with low-end terminals reduced to 16000Hz and high-end terminals maintaining 32000Hz; public broadcasts / intercoms should use 48000Hz for high-end terminals, 32000Hz for mid-range terminals, and 16000Hz for low-end terminals; in noisy and complex scenarios, the sampling rate should be increased by one level, while in normal scenarios it should be decreased by one level, but not exceeding 16000Hz. 48000Hz range.
[0016] In this preferred embodiment, the encoding method switching logic in step S1 is as follows: fire / emergency broadcasts use G.711 encoding to prioritize real-time performance.
[0017] Public address / loudspeaker intercom uses AAC-LD encoding, balancing sound quality and latency;
[0018] Preprocessing noise reduction includes 32-point median filtering and 300Hz-3400Hz bandpass FIR filtering. For industrial plant scenes, the median filtering window is increased to 64 points.
[0019] The intercom terminal uses the NLMS adaptive echo cancellation algorithm, with a computing power consumption of ≤30MIPS; high-end terminals can optionally integrate the RNNoise lightweight AI noise reduction module, with a computing power consumption of ≤50MIPS.
[0020] In this preferred embodiment, the adjustment logic of the interleaving matrix in step S2 is as follows: an 8×8 interleaving matrix is used for fire / emergency broadcasting, and a 4-6×4-6 interleaving matrix is used for public broadcasting / intercom; the interleaving mapping relationship is as follows:
[0021]
[0022] Where M is the order of the interleaving matrix, i and j are the row and column indices of the interleaving matrix, respectively, and SendIdx(i,j) is the sending index of the shuffled frame; Algorithm implementation process: ① Matrix initialization: Determine the order M of the interleaving matrix according to the broadcast type, and construct an M×M empty matrix; ② Frame padding: Patch the continuous frames... ① The audio frames are filled into the matrix in row-major order, where i and j are integers from 0 to M-1; ② Frame shuffling: According to the above mapping formula, the send index SendIdx(i,j) corresponding to each matrix position (i,j) is calculated. The audio frames in the matrix are extracted in ascending order of SendIdx(i,j) to complete the shuffling and rearrangement of the frames; ③ Encapsulation and distribution: The shuffled audio frames are encapsulated into RTP packets, carrying the sampling rate identifier, and distributed via UDP multicast to realize the shuffling of sudden packet loss and reduce the difficulty of subsequent packet loss compensation.
[0023] FEC redundancy is set as follows: 20% for fire / emergency broadcasting, and 10% for public broadcasting / intercom. The FEC redundancy frame calculation formula is as follows: 15%
[0024]
[0025] in For k service frames in the same group, where k is 4 8. Adjustable and matched with the interleaving matrix order; Algorithm implementation process: ① Frame grouping: Divide k consecutive RTP service frames into a group (k matches the interleaving matrix order M), and assign a unique frame group identifier to each group; ② Redundant frame generation: Generate redundant frames within the group. All service frames are XORed sequentially to obtain the corresponding FEC redundant frames. ③ Identification Carrying: Add the same frame group identifier, broadcast type identifier, and sampling rate identifier as the corresponding service frame to the FEC redundant frame to ensure accurate terminal matching; ④ Distribution and Transmission: Distribute the FEC redundant frame and the corresponding service frame together via UDP multicast. The redundancy is preset according to the broadcast type to ensure the error correction reliability of high-priority broadcasts.
[0026] In this preferred embodiment, the specific implementation of multi-terminal clock synchronization in step S3 includes: the synchronization period of the fire / emergency terminal is forcibly set to 1 second, while that of the low-end terminal can be adjusted to 2 seconds; the calculation formulas for link delay d and clock deviation θ are as follows: , ,in When the terminal sends a request, When the server receives a request, The time when the server sends a response. The time when the terminal receives the response; the smoothing correction formula is: Long-term drift correction is performed every 10 seconds, based on drift error. Adjust the local clock.
[0027] In this preferred embodiment, the packet loss detection and hierarchical compensation in step S4 includes: determining packet loss by frame sequence number difference, and the number of lost frames. ,in The current frame number. The largest frame number in history;
[0028] FEC priority correction is performed according to the priority order: fire > emergency > public > intercom. When the number of lost frames is ≤1, the FEC restoration formula is used. Restore the lost frames;
[0029] The conditions for triggering a resume request are L≥2 and FEC restoration failure or keyframe packet loss, and the fire / emergency broadcast timeout period. Number of retries Next, public address / intercom , Secondly, unicast continuation is used to avoid network storms.
[0030] In this preferred embodiment, the specific logic of the local PLC layered compensation in step S4 is as follows: ① Single frame packet loss: Low sampling rate frames use linear interpolation + fade-in / fade-out, high sampling rate frames use quadratic interpolation + fade-in / fade-out, and fire / emergency frames receive an additional 1.2 times gain; ② Short-term continuous packet loss (2≤L≤3): The previous frame's pitch period waveform is reused, low sampling rate frames are combined with linear interpolation, high sampling rate frames are combined with cubic interpolation, and fade-in / fade-out is used throughout; ③ Long-term packet loss (L>3): Low-end terminals use gradual mute, high-end terminals use third-order polynomial fitting compensation, fire / emergency broadcasts trigger local alarms, and intercoms trigger interruption prompts.
[0031] In this preferred embodiment, the specific implementation of the three-level lightweight noise reduction in step S5 includes: ① Basic noise reduction: 8 A 32-point moving average filter is used for high sampling rates and noisy, complex scenarios, while an 8-point filter is used for low sampling rates and normal scenarios. 16 points; ② Adaptive noise reduction: Improved spectral subtraction combined with minimum statistical noise estimation, fire / emergency broadcasting Public Broadcasting intercom High sampling rate frames The value is slightly lower; ③ Post-noise reduction: VAD detection distinguishes between voice / silent segments, fire / emergency segments. Public address / intercom High-end terminals can add sub-band noise reduction for 8 sub-bands; intercom terminals can add dynamic adjustment of noise threshold.
[0032] In this preferred embodiment, the adaptive buffer delay mentioned in step S5... The adjustment logic is as follows: fire / emergency broadcasts should prioritize 40ms, while public broadcasts / intercoms should be adjusted to 60ms based on jitter. 120ms, high sampling rate frame The value is slightly higher;
[0033] The jitter estimation formula is as follows , ,in Let i be the time of receiving the i-th frame. The time when the i-th frame is sent. This is the frame interval deviation. This is the real-time jitter value; 50 is applied when switching between multiple broadcast types. Smooth fade-in and fade-out processing for 60ms.
[0034] In this preferred embodiment, the specific logic of the multi-broadcast type collaborative management in step S6 includes: ① Priority scheduling: Broadcast requests are processed in order of fire > emergency > public > intercom; ② Mode switching: When a fire / emergency alarm is triggered, the mode is automatically switched and other broadcasts are paused, and the sampling rate is adjusted to the appropriate range; ③ Intercom and broadcast linkage: When intercom is used, the public broadcast in the corresponding area is paused, and the sampling rate is matched synchronously. Intercom voice can be converted to public broadcast with one click; ④ Area control: Different types of broadcasts can be accurately distributed to designated areas, and the sampling rate is matched according to the computing power of the area terminals.
[0035] An IP broadcast audio processing system for a converged broadcast system, used to implement the method, includes a broadcast server, at least one IP broadcast terminal and at least one public address intercom terminal, which are connected to each other via an IP network based on the TCP / IP protocol;
[0036] The broadcast server includes an audio acquisition and encoding module, an interleaving encoding module, an FEC redundancy generation module, an NTP clock service module, a resume response module, a parameter optimization module, and a multi-broadcast collaborative management and control module.
[0037] The IP broadcast terminal includes a clock synchronization module, a receiver parsing module, a packet loss detection module, an FEC error correction module, a resume request module, a local compensation module, a multi-level noise reduction module, a synchronous playback module, and a terminal parameter optimization module.
[0038] The loudspeaker intercom terminal integrates the core functions of an IP broadcast terminal and adds an intercom audio acquisition module, an intercom transmission module, an intercom reception module, and an intercom linkage module to achieve seamless linkage between intercom and broadcast.
[0039] In this preferred embodiment, the audio acquisition and encoding module of the broadcast server integrates an adaptive sampling unit, which can dynamically adjust the sampling rate to three levels: 16000Hz / 32000Hz / 48000Hz, dynamically switch the encoding method, and integrate an echo cancellation module and an optional lightweight AI noise reduction module; the multi-broadcast collaborative management module is used to realize priority scheduling of multiple broadcast types, mode switching, intercom and broadcast linkage, and area management, and coordinate the linkage of parameters of each module.
[0040] In this preferred embodiment, the clock synchronization module of the IP broadcast terminal can synchronize with the NTP server according to a dynamic period, report its own computing power level, perform clock smoothing correction and long-term drift suppression, and support early warning of clock deviation for fire / emergency terminals; the local compensation module can perform differentiated compensation according to packet loss level, broadcast type and sampling rate; the multi-level noise reduction module can dynamically adjust noise reduction parameters to adapt to different scenarios and broadcast type requirements.
[0041] In this preferred embodiment, the intercom audio acquisition module of the loudspeaker terminal integrates an adaptive sampling unit, synchronously matches the sampling rate of the broadcast terminal, and performs echo cancellation preprocessing; the intercom linkage module can receive server instructions to realize the pause, resumption, and voice conversion of intercom and broadcast, and synchronously match the sampling rate parameters.
[0042] Compared with the prior art, the technical effects and advantages of the present invention are as follows:
[0043] The IP broadcast audio processing method and system of the integrated broadcast system (1) achieves seamless integration of multiple types of broadcasts and solves the problem of collaborative management: For the first time, emergency broadcasts, fire broadcasts, public broadcasts and loudspeaker intercoms are included in a unified processing framework. Through priority scheduling, mode switching, intercom and broadcast linkage and other collaborative logic, unified management of all types of broadcasts is achieved, solving the problems of independent deployment of multiple systems, poor coordination and switching lag in the existing technology; it supports precise management of multiple areas, and different types of broadcasts can be sent out simultaneously in different areas, adapting to the needs of multiple scenarios such as campuses, parks, and factories;
[0044] (2) The synchronization accuracy of multiple terminals is greatly improved to meet the needs of emergency / firefighting coordination: The combination of "chrony NTP synchronization + smooth correction + long-term drift suppression" is adopted, combined with the secondary synchronization mechanism of firefighting / emergency terminals, to control the synchronization error of multiple terminals to ≤2ms, which is far superior to the existing technology. 10ms synchronization accuracy; no need to use high-cost PTP synchronization, compatible with low-end embedded terminals, balancing accuracy, cost and versatility, ensuring synchronized alarms from multiple terminals in fire / emergency broadcasting, and improving emergency response efficiency.
[0045] (3) Excellent packet loss compensation effect, avoid network storms, and adapt to multiple broadcast priorities: The four-layer protection mechanism of "interleaving anti-burst → FEC priority error correction → unicast continuation → local hierarchical compensation" is adopted. Combined with the differentiated compensation strategy of packet loss level, broadcast type and sampling rate, the packet loss compensation success rate is increased to more than 95%, and the sound quality loss is controlled within an acceptable range when there is continuous packet loss. The unicast continuation + retry number limit is adopted to effectively avoid network storms caused by multicast continuation. The error correction and continuation of high priority broadcasts (fire / emergency) are given priority to ensure that alarm instructions and emergency instructions are clearly transmitted.
[0046] (4) Through a combination of innovative technologies such as adaptive sampling strategy, four-layer packet loss protection mechanism, three-level lightweight noise reduction, multi-terminal clock synchronization optimization and full-link collaborative linkage, the system achieves seamless integration and unified management of four types of broadcasts: emergency broadcast, fire broadcast, public broadcast and loudspeaker intercom. It is also compatible with various embedded terminals from low-end MCU to high-end ARM / Linux, taking into account both engineering practicality and technological innovation, and significantly improving the playback quality, robustness and emergency response efficiency of the integrated broadcast system. Attached Figure Description
[0047] To more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the drawings used in the description of the specific embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.
[0048] Figure 1 This is a diagram showing the overall architecture of the IP broadcast audio processing system of the converged broadcast system of the present invention.
[0049] Figure 2 This is an overall flowchart of the IP broadcast audio processing method of the present invention;
[0050] Figure 3 This is a flowchart illustrating the decision-making process of the adaptive sampling strategy of the present invention.
[0051] Figure 4 This is a flowchart illustrating the packet loss detection and hierarchical compensation mechanism of the present invention.
[0052] Figure 5 This is a flowchart illustrating the implementation of the three-stage lightweight noise reduction method of the present invention.
[0053] Figure 6 This is a logic diagram of the multi-broadcast type collaborative management of the present invention. Detailed Implementation
[0054] In the following description, numerous specific details are set forth in order to provide a more thorough understanding of the invention. However, it will be apparent to those skilled in the art that the invention can be practiced without one or more of these details. In other instances, certain technical features well-known in the art have not been described in order to avoid obscuring the invention.
[0055] Example 1
[0056] This embodiment provides, for example Figures 1 to 6 The IP broadcast audio processing method for a converged broadcast system, as shown, includes the following steps:
[0057] S1. Audio Preprocessing and Encoding (Lightweight design, adaptable to multiple broadcast types, using adaptive sampling)
[0058] The broadcast server collects various audio signals (background music / announcements from public broadcasts, command voices from emergency broadcasts, alarm voices from fire broadcasts, and two-way voices from public address systems). It employs an adaptive sampling strategy, with the sampling rate decision centrally determined by the broadcast server; the terminal only executes and provides feedback. The server dynamically determines the target sampling rate based on a three-dimensional model of "broadcast type + terminal computing power + scene noise." (Only 16000Hz / 32000Hz / 48000Hz are allowed). The decision-making and control process is as follows:
[0059] 1) Computing power reporting: During the NTP synchronization or login phase, the terminal actively reports its own computing power level (low-end / mid-end / high-end) and device model through an independent TCP control channel, and the server establishes a terminal computing power table;
[0060] 2) Sampling rate calculation: The server determines the target sampling rate according to priority (broadcast type > terminal computing power > scene noise):
[0061] Firefighting / Emergency: Prioritize 32000Hz (mid-range), reduce to 16000Hz for low-end, and maintain 32000Hz for high-end;
[0062] Public / intercom: High-end 48000Hz, mid-range 32000Hz, low-end 16000Hz;
[0063] In complex noise scenarios (factory area / rescue), increase the noise level by one gear; in normal scenarios, decrease the noise level by one gear (not lower than 16k and not higher than 48k).
[0064] 3) Command Issuance: The server issues a "sampling rate switching command" (including target) via the TCP control channel. The RTP extension header (1 byte) also carries a sampling rate identifier: 0=16k, 1=32k, 2=48k, as a data plane redundancy check.
[0065] 4) Terminal handover: After receiving the control command, the terminal aligns with the RTP timestamp boundary (20ms frame end) to complete the seamless resampling handover to avoid half-frame distortion; within 100ms after the handover, it sends a "handover successful / failed" message to the server.
[0066] 5) Dynamic closed loop: The server collects terminal feedback, network packet loss rate and noise estimation every 5 seconds. If the packet loss is >5% for 3 consecutive seconds, the sampling rate can be temporarily reduced. After the network recovers, the sampling rate is smoothly increased in increments of 1 level / 2 seconds to prevent jitter.
[0067] Abandoning a fixed sampling frequency, the sampling parameters are dynamically adjusted based on the broadcast type, terminal computing power, and scenario requirements, as follows: Sampling rate Available at 16000Hz Adaptive switching within the 48000Hz range, single frame duration Keeping 20ms constant, the number of sampling points per frame (Dynamically matching sampling rate, balancing sound quality and computing power); Sampling rate switching logic:
[0068] 1) Adaptation by broadcast type: Fire / emergency broadcasts prioritize real-time performance, with a sampling rate adaptation of 16000Hz. 32000Hz (16000Hz for low-end terminals, 32000Hz for mid-range terminals); Public address / loudspeaker systems balance sound quality and latency, with a sampling rate adapted to 32000Hz. 48000Hz (48000Hz for high-end terminals, 32000Hz for mid-range terminals);
[0069] 2) Adaptation based on terminal computing power: For low-end MCUs (such as the STM32F1 series, computing power ≤ 100 MIPS), the sampling rate is capped at 16000Hz; for mid-range MCUs (such as the STM32F4 series, computing power ≤ 100 MIPS), the sampling rate is capped at 16000Hz. The sampling rate limit for 200MIPS terminals is 32000Hz; the sampling rate limit for high-end ARM / Linux terminals (computing power ≥ 200MIPS) is 48000Hz.
[0070] 3) Adaptable to various scenarios: Sampling rate increased to 32000Hz for complex noise environments such as factory areas and emergency rescue sites. 48000Hz improves noise recognition and suppression accuracy; in typical scenarios such as office buildings, control rooms, and laboratories, the sampling rate can be reduced to 16000Hz. 32000Hz, reducing computing power consumption.
[0071] In this embodiment, the low-latency encoding method is dynamically switched according to the broadcast type (fire / emergency broadcasts use G.711 encoding to prioritize real-time performance; public broadcasts / loudspeaker intercoms can use AAC-LD encoding to balance sound quality and latency), while pre-processing noise reduction is performed: impulse noise is suppressed sequentially through a 32-point median filter, and 300Hz... A 3400Hz bandpass FIR filter removes power frequency and ultra-high frequency noise, achieving initial noise suppression. For different types of broadcast scenarios, the filter parameters are dynamically adjusted (in industrial plant scenarios, the median filter window is increased to 64 points to suppress equipment impulse noise; in emergency scenarios, the bandpass filter attenuation is enhanced to filter out ambient noise). For high-end terminal scenarios, a lightweight RNNoise AI noise reduction module can be selectively added, connected in series in the preprocessing flow, to further suppress complex environmental noise (such as wind noise, vehicle noise, and equipment noise). The AI noise reduction module's computing power is controlled within 50 MIPS, making it compatible with mid-range embedded terminals.
[0072] In this embodiment, a new audio processing adaptation for amplified intercom is added: the two-way voice collected by the intercom terminal is preprocessed with echo cancellation (NLMS adaptive echo cancellation algorithm, computing power ≤30MIPS) to eliminate echo interference during intercom and ensure clear two-way communication; when switching between intercom voice and broadcast voice, a smooth transition processing (fade-in and fade-out duration 50ms) is performed to avoid popping sounds during switching; the sampling rate of the intercom terminal is synchronously adaptively adjusted to match the sampling rate of the broadcast terminal to avoid audio quality distortion during switching.
[0073] S2. Audio interleaving and FEC redundancy generation (anti-burst packet loss, adaptable to multiple broadcast priorities)
[0074] S21. Broadcast Server Construction Interleaving matrix (M is 4) 8 adjustable, dynamically selecting based on network environment and broadcast type: Fire / emergency broadcasts prioritize M=8 to improve resilience against sudden packet loss; public broadcasts / intercoms use M=4 to balance bandwidth and efficiency. After the frame audio is filled into the matrix in row and column order, it is then processed according to the mapping relationship. Shuffle and rearrange, among which To shuffle the transmission index of subsequent frames, i and j are the row and column indices of the interleaving matrix, respectively (i and j are both 0). The scrambled audio frames are encapsulated into RTP packets in a new order. The interleaving algorithm breaks down sudden packet loss into random packet loss, reducing the difficulty of subsequent packet loss compensation. This design is different from the existing "direct transmission without interleaving" or "simple sequential interleaving" schemes and is adapted to the priority requirements of multiple broadcast types. The RTP packets carry a sampling rate identifier, which is convenient for the terminal to synchronously adapt to decoding.
[0075] S22. Each group consists of k consecutive RTP service frames (k is 4). 8 (adjustable, matching the interleaving matrix order M), generates one FEC redundant frame corresponding to the group through XOR operation. The FEC redundant frame calculation formula is:
[0076] ,in For k service frames within the same group, This is the FEC redundancy frame corresponding to this group; the FEC redundancy is set according to the broadcast priority (fire / emergency broadcast redundancy is increased to 20%, public broadcast / intercom redundancy is 10%). 15%), ensuring the reliability of error correction for high-priority broadcasts; the FEC redundant frames and corresponding service frames are carried with the same frame group identifier, broadcast type identifier (distinguishing between fire, emergency, public, and intercom) and sampling rate identifier, and are sent together to all IP broadcast terminals via UDP multicast. The frame group identifier is used by the terminal to accurately match the FEC redundant frames with the service frames, the broadcast type identifier is used by the terminal to identify the broadcast type and execute the corresponding processing logic, and the sampling rate identifier is used by the terminal to synchronously adjust the decoding sampling parameters to avoid false error correction, type confusion and audio quality distortion. This design solves the problems of poor FEC error correction targeting, insufficient adaptation to multiple broadcast types and uncoordinated sampling rates in the existing technology.
[0077] S3. Multi-terminal clock synchronization (dynamically optimized, adaptable to multiple scenarios)
[0078] S31. The broadcast server deploys the chronyNTP clock service (which improves clock stability by 30% compared to ordinary NTP services). After powering on, IP broadcast terminals (including intercom terminals) send time requests to the NTP server at 1-second intervals. The synchronization period can be dynamically adjusted according to the terminal's computing power and broadcast type (low-end terminals can be adjusted to 2 seconds to reduce computing power consumption; fire / emergency broadcast terminals are forced to use a 1-second interval to ensure synchronization accuracy). Unlike the high-cost PTP synchronization solutions in existing technologies, this solution uses NTP synchronization + subsequent optimization to balance accuracy, cost, and the needs of multiple broadcast types. When synchronizing, the terminal synchronously reports its own computing power level, which facilitates the server to issue matching sampling rate parameters.
[0079] S32. Terminal records the time of sending request. Receive server response time The server records the time when the request is received. Sending response time The link delay d and clock skew θ are calculated using the following formulas:
[0080] The formula for calculating link delay d is:
[0081]
[0082] Algorithm implementation process: When the terminal sends a time synchronization request, it records the sending time. And include it in the request packet; after receiving the request packet, the server records the time of receipt. Immediately generate a response packet and record the sending time. ,Will , The response packet is included in the data and sent out; after receiving the response packet, the terminal records the time of reception. The terminal calculates the link delay d according to the above formula, eliminating the impact of network round-trip transmission delay on clock synchronization, providing a basis for subsequent clock deviation calculation. The calculation process is lightweight, involving only simple arithmetic operations, and is compatible with the computing power of various embedded terminals.
[0083] Formula for calculating clock skew θ:
[0084]
[0085] Algorithm implementation process: Based on the above steps, the following steps were obtained. , , , Four time values are substituted into the formula to calculate the deviation θ between the terminal's local clock and the server's reference clock. If θ is positive, it means that the terminal's local clock is ahead of the server's clock. If θ is negative, it means that the terminal's local clock is behind the server's clock. After the calculation is completed, θ is passed to the smoothing correction algorithm to avoid clock jumps, ensure synchronization accuracy, and reduce computing power consumption.
[0086] Terminal clock smoothing correction formula:
[0087] .
[0088] Algorithm implementation process: First, define the Clamp function to limit the amplitude of a single clock correction. When θ > 20ms, When θ < -20ms, When -20ms ≤ θ ≤ 20ms, The terminal obtains the local current clock. Substitute into the above formula to calculate the corrected local clock. This avoids audio stuttering caused by excessively large single calibration amplitude; after calibration is completed, the local clock is updated, and the next synchronization cycle begins.
[0089] in, : The original local clock time currently running on the terminal before a single smooth correction is executed;
[0090] The updated local clock time of the terminal after amplitude limiting and smoothing correction provides a clock reference for subsequent audio timing alignment and multi-terminal synchronous playback;
[0091] : Dedicated limiting and truncation function, used to constrain the amplitude of a single clock correction, avoiding audio stuttering, dropouts, and timing errors caused by large clock jumps;
[0092] The system's preset maximum upper and lower limits for single clock correction are fixed engineering parameters.
[0093] Function execution logic: When When, the function outputs ;when When, the function outputs ;when At that time, the function directly outputs the original offset. .
[0094] Long-term clock drift correction: through drift error Adjust the local clock using the following formula:
[0095]
[0096] The global standard reference time output by the broadcast server chronyNTP clock service serves as a unified clock reference for all terminals across the network.
[0097] : The terminal's local real-time clock before long-term drift correction is performed;
[0098] The clock drift error caused by the long-term operation of the terminal hardware crystal oscillator is measured in milliseconds. The system performs drift correction every 10 seconds, and uses this error to fine-tune the local clock, suppressing long-term clock offset and ensuring stable synchronization accuracy in long-term broadcast scenarios.
[0099]
[0100] Algorithm implementation process: Long-term drift correction is performed every 10 seconds, and the terminal obtains the server's NTP reference time. Calculate the current local clock and drift error ;Will Divide by 64 to obtain the step size for each drift correction, avoiding clock instability caused by excessively rapid drift correction; add this step size to the current local clock. It completes drift correction, continuously suppresses long-term clock drift, and ensures that the terminal clock and the server reference clock remain consistent over a long period of time, especially suitable for the high-precision synchronization requirements of fire / emergency terminals.
[0101]
[0102] S33. The terminal performs a smooth correction on the local clock. The correction formula is as follows: The Clamp function limits the magnitude of a single correction to prevent audio stuttering caused by time jumps; it also compares the local time with the NTP reference time every 10 seconds, suppressing long-term clock drift through a drift correction formula. , ,in For drift error, The NTP reference time is used; for fire / emergency broadcast terminals, a new clock deviation warning mechanism is added. When the deviation exceeds 1ms, secondary synchronization is immediately triggered to ensure that the synchronization error of multiple terminals is ≤2ms. This solves the problems of existing technologies that only use single NTP synchronization, have weak anti-drift capabilities, and cannot meet the needs of emergency / fire coordination. During the synchronization process, the terminal adjusts its local sampling and decoding rate in real time according to the sampling rate parameters sent by the server to ensure multi-terminal sampling synchronization.
[0103] S4. Terminal packet loss detection and layered compensation (four-layer protection, protection against network storms, and adaptation to multiple broadcast types)
[0104] This step is the core innovation, employing a four-layer packet loss protection mechanism: "interleaving for burst damage prevention → FEC priority error correction → unicast continuation → local hierarchical compensation." Combined with a multi-broadcast type priority optimization compensation strategy, it differs from existing "single compensation" or "no hierarchical compensation" solutions, balancing sound quality, playback continuity, network stability, and the needs of multiple broadcast types. Furthermore, it incorporates adaptive sampling parameters to dynamically adjust the accuracy of the compensation algorithm, adapting to packet loss compensation requirements at different sampling rates.
[0105] S41. The terminal receives RTP packets and FEC redundant frames, parses the broadcast type identifier and sampling rate identifier, and buffers the maximum sequence number of historical received frames. and the corresponding sampling rate parameters, each time a new frame is received The following formula is used to determine whether packets are lost and to calculate the number of lost frames L:
[0106] If Loss=true, if Otherwise, Loss = false.
[0107] Formula for calculating the number of lost frames L:
[0108]
[0109] Algorithm implementation process: The terminal maintains the maximum sequence number of historically received frames. Each time a new RTP packet is received, its frame sequence number is parsed. ;Will and Perform a comparison, if If packet loss is detected, the number of lost frames L is calculated using the formula above; simultaneously, a set of lost frame sequence numbers is generated. It can accurately locate the specific sequence number of lost frames, providing accurate packet loss information for subsequent FEC error correction, retransmission application and local compensation. The algorithm logic is simple, the execution efficiency is high, and it is compatible with various embedded terminals.
[0110] Set of lost frame sequence numbers:
[0111]
[0112] Algorithm implementation process: Based on the number of lost frames L calculated above, using The starting number, To determine the terminating sequence number, enumerate the sequence numbers of all lost frames sequentially to form a set of lost frame sequence numbers. The terminal will It is stored in association with the corresponding frame group identifier, broadcast type identifier, and sampling rate identifier. This information can be used later during FEC error correction. Match the FEC redundant frames of the corresponding frame group, and in the retransmission request, Include it in the request packet to ensure that the server accurately sends the resume frame.
[0113] FEC lost frame restoration formula:
[0114]
[0115] Algorithm implementation process: When the number of lost frames is ≤1, the terminal determines the sequence number of the lost frames. Match the FEC redundant frames in its frame group and within that group, except Other service frames (excluding) );right and (excluding) Perform XOR operations sequentially to obtain the lost frames. The data is restored; the XOR operation logic is: 0⊕0=0, 0⊕1=1, 1⊕0=1, 1⊕1=0. The operation process is lightweight, requiring no complex computing power, and is compatible with low-end embedded terminals; after restoration, the data is restored... Decoding verification is performed to ensure that the restored audio quality meets the requirements. If restoration fails, a resume request is triggered.
[0116] ;
[0117] S42. FEC Priority Error Correction: Adjust the error correction priority according to the broadcast type priority (fire > emergency > public > intercom), and prioritize FEC error correction for lost frames of high-priority broadcasts; if there are FEC redundant frames in the same group as the lost frame, and the number of lost frames is ≤1, the lost frame is recovered using the FEC restoration formula: ,in For lost frames, For the FEC redundant frames corresponding to the group containing the lost frames, For the group excluding Other service frames besides those mentioned above; if restoration is successful, proceed to step S5; if restoration fails, trigger a resume request judgment; prioritize FEC error correction to maximize the audio quality of high-priority broadcasts and avoid unnecessary resume requests; during error correction, match the error correction precision according to the sampling rate identifier, high sampling rate frames (32000Hz) The 48000Hz frame uses a high-precision error correction algorithm, while the low sampling rate frame (16000Hz) uses a lightweight error correction algorithm, balancing computing power and sound quality.
[0118] S43. Resumption Request Judgment and Execution: Resumption priority is set according to broadcast type priority. Lost frames in fire / emergency broadcasts trigger a resumption request first. The terminal sends a unicast resumption request to the broadcast server when any of the following conditions are met. (including the set of lost frame sequence numbers) (Corresponding frame group identifier, broadcast type identifier, and its own sampling rate parameters): ① L≥2 and FEC restoration fails; ② Key frames (first frame, encoding switching frame, fire / emergency alarm frame) are lost; The terminal sets the timeout and retry count according to the broadcast type (fire / emergency broadcast: , Next; Public address / intercom: , If no response is received within a timeout period, a retry is performed. If the retry limit is reached, the local PLC packet loss compensation is initiated. If a unicast continuation frame is received from the server, it is decoded and proceeds to step S5. Unicast continuation is used instead of multicast continuation, and the number of retries is limited to effectively avoid network storms caused by continuation requests. At the same time, priority settings ensure the reliability of continuation for high-priority broadcasts, solving the problems of excessive bandwidth consumption for multicast continuation and lack of priority for continuation of multiple broadcast types in existing technologies. The server sends continuation frames that match the sampling rate based on the sampling rate parameters reported by the terminal to avoid decoding distortion.
[0119] S44. Local PLC layered compensation (differentiated processing based on packet loss level + broadcast type + sampling rate, balancing audio quality and computing power):
[0120] ① Single-frame packet loss (L=1): Low sampling rate (16000Hz) frames use linear interpolation + fade-in / fade-out smoothing; high sampling rate (32000Hz) frames use linear interpolation + fade-in / fade-out smoothing. The 48000Hz frame uses double interpolation and smooth fade-in / fade-out to ensure sound quality; for fire / emergency broadcasts, additional amplitude enhancement processing (gain increased by 1.2 times) is added to ensure that alarm voices are clearly identifiable;
[0121] ② Short-term continuous packet loss (2≤L≤3): The fundamental frequency waveform of the previous frame before packet loss is reused, and the waveform of the lost frame in the middle is fitted by linear interpolation (low sampling rate) or cubic interpolation (high sampling rate). The entire process is smoothly transitioned by fade-in and fade-out to avoid pitch distortion. Compared with the existing technology of "simply repeating the previous frame", the sound quality is significantly improved.
[0122] ③Long-term packet loss (L>3): Low-end terminals use a gradual mute formula to achieve amplitude attenuation; high-end terminals use a third-order polynomial fitting compensation: using the 10 most recent consecutive effective PCM frames before packet loss as the fitting window, a third-order polynomial is constructed. Boundary conditions: Minimize the sum of squared errors for the first three consecutive frames; ensure the first derivative of the last frame matches the slope of the final frame. Solve for the coefficients using the least squares method and extrapolate the compensation data for L frames. Perform a 50ms fade-in / fade-out transition after extrapolation to avoid abrupt distortion. For fire / emergency broadcasts, trigger a local alarm (playing a preset emergency alert tone) during prolonged packet loss to prevent alarm interruption due to packet loss. For intercom systems, trigger an intercom interruption alert during prolonged packet loss to remind the user to re-initiate the intercom. Simultaneously, adjust the compensation transition time according to the sampling rate: extend the transition time to 60ms for high sampling rate frames and maintain 50ms for low sampling rate frames to adapt to the differentiated needs of multiple broadcast types and different sampling rates.
[0123] S5. Terminal audio noise reduction and synchronized playback (multi-level lightweight design, end-to-end collaboration, adaptable to multiple broadcast scenarios)
[0124] S51. The terminal performs three-level lightweight noise reduction processing on normal frames, FEC restored frames, resumed frames, and compensation frames. It dynamically adjusts the noise reduction parameters based on the broadcast type, scene, and sampling rate. This differs from existing technologies that use "complex AI noise reduction" or "single filtering" solutions, and balances noise suppression, voice fidelity, computing power requirements, and multi-scene adaptation.
[0125] ①Basic noise reduction: through 8 A 32-point moving average filter (adjusted according to the scene and sampling rate: 32 points for industrial plants and high sampling rate scenes, and 16 or 8 points for parks / commercial buildings and low sampling rate scenes, balancing smoothing effect and computing power) further smooths the white noise. The formula is:
[0126]
[0127] in For sampling after filtering, Let the window length be denoted by 'window'. Algorithm implementation process: The terminal initializes the sliding window length. (Based on scenario and sampling rate presets), cache the most recent sampling points (k=0 to Each time a new sampling point is received Discard the earliest sampling point Add the new sampling point to the buffer, and calculate the current filtered sample value using the formula above. It outputs filtered audio signals in real time, ensuring lightweight and low latency, and adapting to the computing power of embedded terminals.
[0128] ② Adaptive noise reduction: An improved spectral subtraction method is used to suppress residual noise. The formula is as follows:
[0129] ,like ;otherwise
[0130] in This is a noisy spectrum. For noise spectrum, The lower limit suppression coefficient (adjusted according to broadcast type and sampling rate: fire / emergency broadcast) Ensure clear audio; public address system. Balancing sound quality and noise reduction; intercom Balancing call clarity with noise reduction; high sampling rate frames A slightly lower value improves audio fidelity and reduces the sampling rate frame rate. (Take a slightly higher value to enhance noise suppression); Algorithm implementation process: First, process the noisy frequency signal... Perform a short-time Fourier transform (STFT) to obtain the noisy spectrum. The noise spectrum is updated in real time using a minimum statistical noise estimation algorithm. Specifically, a noise spectrum history buffer is maintained for 1 second (20 frames) with an update period of 50ms. During each update, the minimum value of the history buffer is taken as the current noise spectrum estimate for each frequency point. The noise update is triggered by 20ms of continuous silence detected by VAD. The noise smooth update uses a time constant. To avoid introducing musical noise through sudden changes; then calculate the denoised spectrum using the formula above. Finally, the spectrum is converted back to the time domain audio signal through inverse short-time Fourier transform (ISTFT) to complete the adaptive noise reduction processing; for different sampling rate frames, the STFT frame length is dynamically adjusted (1024-point frame length for low sampling rate 16000Hz, and 2048-point frame length for high sampling rate 32000Hz / 48000Hz) to balance noise reduction accuracy and computing power consumption, and to adapt to various embedded terminals.
[0131] ③ Post-noise reduction: Based on VAD (Voice Activity Detection) technology, it distinguishes between voice segments and silent segments, dynamically adjusting the noise reduction intensity to avoid distortion caused by excessive noise reduction in voice segments and residual noise in silent segments affecting the user experience; VAD detection uses dual thresholds of short-time energy and zero-crossing rate for judgment, and the threshold for voice segments is adjusted according to the broadcast type (the short-time energy threshold for fire / emergency broadcasts is reduced by 15% to ensure that weak alarm voices are not misjudged as silent; the standard threshold is used for public broadcasts / intercoms to balance detection accuracy and false positive rate); deep noise reduction is performed on silent segments, while more voice details are preserved in voice segments. Specific parameters are as follows: fire / emergency broadcasts (Minimum gain to avoid residual noise in silent sections), Public address / intercom (Dynamically adjusted according to scene noise intensity); High-end terminals can add 8 sub-band noise reduction modules to perform differentiated suppression for different frequency band noise (such as low-frequency equipment noise in industrial plants and high-frequency environmental noise in parks), further improving the noise reduction effect. The computing power of the sub-band noise reduction module is controlled within 40MIPS, so as not to affect the operation of other modules of the terminal; The intercom terminal adds additional dynamic noise threshold adjustment logic, which adjusts the VAD threshold and noise reduction gain in real time according to the intercom voice intensity to avoid voice distortion and background noise interference during intercom, and ensure clear two-way communication.
[0132] S52. Audio Timing Restoration and Synchronous Playback: The audio frames received by the terminal are interleaved and scrambled. The timing needs to be restored according to the inverse interleaving matrix. The inverse mapping relationship is as follows: Where RecvIdx(i,j) is the received index of the restored frame, M is the order of the interleaving matrix, and i and j are the row and column indices of the matrix, respectively. This forms a closed loop with the interleaving mapping in step S2, ensuring correct audio frame timing and avoiding playback misalignment. After restoring the timing, the terminal uses the real-time jitter value... Dynamically adjust adaptive buffer delay This ensures synchronized playback across multiple devices while avoiding excessive stuttering and latency.
[0133] Adaptive buffer delay Supplementary implementation of dynamic adjustment logic: The terminal calculates the jitter value in real time. ,when When fire / emergency broadcasts maintain a 40ms buffer delay, public broadcasts / intercoms are adjusted to 60ms; when 10ms ≤ When the buffer delay is ≤20ms, the buffer delay is uniformly increased by 20ms; when When buffer delay increases by 30ms, but not exceeding 120ms (to avoid excessive delay affecting emergency response and intercom experience); the average jitter value is recalculated every 500ms, and the buffer delay is smoothly adjusted according to the average jitter value, with an adjustment step of no more than 10ms / time, to avoid audio stuttering caused by sudden changes in buffer delay.
[0134] When switching between multiple broadcast types, execute 50. The 60ms fade-in / fade-out smoothing is implemented as follows: 5ms before the transition, the gain of the current broadcast audio is gradually reduced to 0; 5ms after the transition, the gain of the remaining audio is reduced to 0. The target broadcast audio gain is gradually increased to the normal level over 60ms. High sampling rate frames use a 60ms transition duration, while low sampling rate frames use a 50ms transition duration to avoid issues such as popping or dropping sounds during switching. At the same time, the sampling rate, noise reduction parameters, and buffer delay are adjusted synchronously to ensure smooth audio playback and stable sound quality after switching, achieving seamless connection between four broadcast types: fire, emergency, public, and intercom.
[0135] S6. Iterative optimization and collaborative management of multiple broadcast types (end-to-end linkage, adapting to management needs in multiple scenarios).
[0136] S61. Iterative Optimization: The server collects feedback data from all terminals in real time, including terminal computing power utilization, sampling rate switching status, packet loss rate, noise reduction effect, clock synchronization deviation, etc., and performs full-link parameter optimization every 10 seconds. If the terminal computing power utilization rate continues to exceed 80%, the server automatically lowers the corresponding terminal sampling rate by one level to reduce computing power consumption. If the packet loss rate continues to exceed 5%, the FEC redundancy of the corresponding broadcast type is appropriately increased (not exceeding 25% to avoid excessive bandwidth consumption). If the clock synchronization deviation continues to exceed 1ms, the terminal is triggered to perform secondary synchronization and adjust the synchronization period. The terminal reports its own operating status to the server every 30 seconds. The server dynamically adjusts the parameters sent based on the reported data, forming a closed loop of "collection-feedback-optimization-execution" to ensure system adaptability and stability.
[0137] S62. Multi-broadcast Type Collaborative Management: Based on the priority of "Fire > Emergency > Public > Intercom", unified scheduling and linkage of all types of broadcasts are achieved. The specific logic is supplemented as follows:
[0138] ① Priority scheduling: The server maintains a broadcast request queue and processes them in order of priority. High-priority broadcasts (fire / emergency) can preempt low-priority broadcasts (public / intercom). When preempting, a smooth transition is performed to avoid poor user experience caused by sudden interruption of low-priority broadcasts. If multiple broadcast requests of the same priority are triggered at the same time, they are processed in the order of their initiation time. Furthermore, broadcasts of the same priority in different regions can be distributed separately according to regional needs to achieve precise control.
[0139] ② Mode Switching: The system has three preset working modes—normal mode, emergency mode, and fire mode. The default mode is normal (public address / intercom is running normally). When a fire alarm is triggered, it automatically switches to fire mode, pauses all other types of broadcasts, adjusts the sampling rate to 32000Hz (16000Hz for low-end terminals), increases FEC redundancy to 20%, and enhances clock synchronization accuracy to ensure clear and synchronized playback of fire alarm voice. When an emergency command is triggered, it switches to emergency mode, pauses public address / intercom, prioritizes the transmission of emergency command voice, and the parameter configuration is the same as fire mode. After the emergency command ends, it automatically returns to normal mode. During mode switching, the terminal synchronously adjusts encoding, noise reduction, buffering, and other parameters to ensure smooth and lag-free switching.
[0140] ③ Intercom and Broadcast Integration: When an intercom terminal initiates an intercom call, the server automatically pauses the public broadcast in the corresponding area and simultaneously adjusts the sampling rate of the intercom terminal and the corresponding area's broadcast terminal to be consistent, avoiding interference between intercom and broadcast voice. During the intercom call, if a fire / emergency broadcast is triggered, the intercom call will automatically pause, prioritizing the playback of the fire / emergency voice, and will automatically resume when the intercom call is restored. It supports one-click conversion of intercom voice to public broadcast. After the user initiates the conversion command, the server will encode and process the intercom voice according to public broadcast parameters and distribute it to the designated area, achieving seamless integration of intercom and broadcast, adapting to the emergency communication and daily notification needs of scenarios such as parks and factories.
[0141] ④ Regional Control: The server supports dividing terminals into groups by region, which can accurately send different types of broadcasts to designated areas (such as office areas in a park, production areas in a factory, and teaching buildings in a campus). When sending broadcasts, the server matches the corresponding sampling rate and processing parameters according to the average computing power of the terminals in the region to ensure that all terminals in the region can operate stably. The server supports sending different types of broadcasts to multiple regions at the same time (such as sending public broadcasts from office areas and sending emergency instructions from production areas) without interference, thus achieving precise control in multiple scenarios and regions.
[0142] In this embodiment, to implement the IP broadcast audio processing method of the above-mentioned integrated broadcast system, the present invention also provides a corresponding processing system. The system architecture adopts a distributed "server-terminal" design, realizes data transmission and command interaction based on IP network, adapts to embedded terminals with different computing power, and realizes seamless integration and unified management of emergency broadcast, fire broadcast, public broadcast, and intercom. The specific structure and functions are supplemented as follows:
[0143] 1. System Overall Architecture: The system includes a broadcast server, at least one IP broadcast terminal, and at least one public address intercom terminal. The three are connected via an IP network (LAN or WAN) based on the TCP / IP protocol. The broadcast server, as the core management node, is responsible for audio acquisition and encoding, parameter decision-making, clock synchronization, data transmission, resume response, and collaborative management. The IP broadcast terminal is responsible for audio reception, packet loss compensation, noise reduction, synchronized playback, and status feedback. The public address intercom terminal, based on the core functions of the IP broadcast terminal, adds a dedicated intercom module to achieve two-way intercom and broadcast linkage. The three work together to form a complete converged broadcast processing system.
[0144] 2. Supplementary information on the functions of each module of the broadcast server:
[0145] ① Audio Acquisition and Encoding Module: Integrates an adaptive sampling unit, encoding switching unit, preprocessing noise reduction unit, and echo cancellation unit; the adaptive sampling unit dynamically decides on three sampling rates (16000Hz / 32000Hz / 48000Hz) based on a three-dimensional model of "broadcast type + terminal computing power + scene noise," and sends switching commands through the TCP control channel; the encoding switching unit dynamically switches between G.711 / AAC-LD encoding methods according to the broadcast type; the preprocessing noise reduction unit performs 32-point median filtering and 300Hz... 3400Hz bandpass FIR filtering with dynamic adjustment of filtering parameters; echo cancellation unit for intercom audio, performing NLMS adaptive echo cancellation to ensure intercom audio quality; high-end terminal adapter module can optionally integrate RNNoise lightweight AI noise reduction module to improve noise reduction effect in complex scenarios.
[0146] ② Interleaving coding module: Dynamically adjusts the interleaving matrix order according to the broadcast type (fire / emergency M=8, public / intercom M=4) 6) Perform audio frame interleaving and shuffling processing according to the mapping relationship. The audio frames are rearranged, encapsulated into RTP packets, and carry identifiers such as sampling rate and broadcast type, providing a foundation for resisting sudden packet loss. At the same time, it supports dynamic optimization of interleaving parameters, adjusting the matrix order according to the network packet loss rate to balance the anti-packet loss effect and bandwidth consumption.
[0147] ③ FEC redundancy generation module: Press 4 Eight service frames are grouped together, and FEC redundant frames are generated through XOR operation. Differential redundancy is set according to broadcast priority (20% for fire / emergency, 10% for public / intercom). (15%), redundant frames are sent via UDP multicast with the same identifier as service frames; dynamic adjustment of redundancy is supported, and the redundancy of the corresponding broadcast type is automatically increased when the network packet loss rate increases to ensure error correction reliability.
[0148] ④NTP Clock Service Module: Deploys chronyNTP clock service to provide a high-precision reference clock, receives terminal synchronization requests, records request and response times, and assists terminals in calculating link latency and clock deviation; supports dynamic synchronization period adjustment, with a mandatory 1s synchronization period for fire / emergency terminals and a 2s period for low-end terminals, balancing synchronization accuracy and terminal computing power consumption; integrates a clock drift monitoring unit to monitor terminal clock deviation in real time and trigger a secondary synchronization mechanism.
[0149] ⑤ Resumption Response Module: Receives unicast resumption requests from terminals, processes them according to broadcast priority, and prioritizes responses to resumption requests from fire / emergency terminals; parses parameters such as lost frame sequence number and sampling rate in the request, sends out matching resumption frames, controls the resumption timeout and retry count to avoid network storms; at the same time, it records the history of resumption requests, optimizes the resumption strategy, and reduces duplicate resumptions.
[0150] ⑥ Parameter optimization module: Collects terminal feedback data in real time (computing power, packet loss rate, synchronization deviation, noise reduction effect, etc.), performs full-link parameter optimization every 10 seconds, dynamically adjusts parameters such as sampling rate, interleaving matrix order, FEC redundancy, and synchronization period to form a closed-loop optimization; automatically triggers parameter adjustment commands to address issues such as insufficient terminal computing power, excessively high packet loss rate, and excessive synchronization deviation, ensuring stable system operation.
[0151] ⑦ Multi-broadcast Collaborative Management and Control Module: As the core management and control unit of the server, it realizes priority scheduling, mode switching, intercom and broadcast linkage, and area control for multiple broadcast types; maintains the broadcast request queue and processes them in order of fire > emergency > public > intercom; controls the switching of system working modes, and coordinates the adjustment of parameters of each module when the fire / emergency mode is triggered; manages the intercom and broadcast linkage logic, and realizes functions such as pausing intercom broadcast and converting intercom voice to broadcast; supports area group management, accurately sends different types of broadcasts to designated areas, and coordinates the adaptation of terminal parameters within the area.
[0152] 3. Supplementary Functions of Each Module in the IP Broadcast Terminal:
[0153] ① Clock Synchronization Module: Synchronizes with the server's NTP clock service, sends synchronization requests periodically, and reports its own computing power level; calculates link latency and clock deviation, and performs smoothing correction and long-term drift correction, with the following formulas respectively. , , It integrates a clock deviation early warning unit, which triggers secondary synchronization when the deviation of the fire / emergency terminal exceeds 1ms, ensuring that the synchronization error of multiple terminals is ≤2ms; it also synchronously adjusts the local sampling and decoding rate to keep it consistent with the sampling rate issued by the server.
[0154] ② Receive and Parse Module: Receives RTP service frames and FEC redundant frames sent by the server via UDP multicast, parses the broadcast type identifier, sampling rate identifier, frame group identifier, and frame sequence number, and buffers valid frames; filters invalid and erroneous frames, records the reception status, and provides a basis for packet loss detection; supports fast parsing of multiple types of identifiers, adapts to the processing requirements of different broadcast types, and ensures parsing efficiency and accuracy.
[0155] ③ Packet loss detection module: buffers the maximum sequence number of historically received frames. Each time a new frame is received ,pass Calculate the number of lost frames and generate a set of lost frame sequence numbers. It can determine the packet loss level (single frame, short continuous, long time) in real time and trigger the corresponding compensation process to ensure accurate and efficient packet loss detection and adapt to the computing power of embedded terminals.
[0156] ④ FEC Error Correction Module: Based on the broadcast type priority, FEC error correction is performed on lost frames of high-priority broadcasts first; when the number of lost frames is ≤1, it is corrected by... Restore lost frames, verify the restored audio quality, and trigger a resume request if restoration fails; support differentiated error correction for frames with different sampling rates, balancing error correction accuracy and computing power consumption.
[0157] ⑤ Resume Request Module: Initiates unicast resume requests according to broadcast priority, carrying the set of lost frame sequence numbers, frame group identifier, broadcast type identifier, and its own sampling rate parameters; sets the timeout and retry count for the corresponding broadcast type (fire / emergency). , Second; Public / Intercom , If no response is received within a timeout period, a retry will be performed. If the retry limit is reached, local compensation will be triggered. The server will receive a continuation frame, decode it, and then pass it into the subsequent processing flow.
[0158] ⑥ Local Compensation Module: Performs differentiated PLC layered compensation based on packet loss level, broadcast type, and sampling rate; single-frame packet loss uses linear / quadratic interpolation + fade-in / fade-out, short-term continuous packet loss uses pitch period + interpolation, and long-term packet loss uses gradual mute (low-end terminals) or third-order polynomial fitting (high-end terminals); long-term packet loss in fire / emergency situations triggers local alarms, and long-term packet loss in intercoms triggers interruption prompts; dynamically adjusts the compensation transition time to adapt to different sampling rate requirements and ensures compensation sound quality.
[0159] ⑦ Multi-level noise reduction module: Performs three-level lightweight noise reduction, with basic noise reduction using 8 32-point moving average filtering, adaptive noise reduction employs improved spectral subtraction combined with minimum statistical noise estimation, and post-noise reduction dynamically adjusts the gain based on VAD detection; noise reduction parameters (filter window, ...) are dynamically adjusted according to broadcast type, scene, and sampling rate. coefficient, (etc.); high-end terminals can enable sub-band noise reduction modules to further improve noise reduction effects and adapt to complex noise scenarios.
[0160] ⑧ Synchronous playback module: Restores the audio frame timing according to the inverse interleaving matrix and dynamically adjusts the adaptive buffer delay. It optimizes buffer parameters based on jitter values to achieve synchronized playback across multiple terminals; performs fade-in and fade-out processing when switching between multiple broadcast types to avoid popping and audio dropouts; controls audio playback gain, with fire / emergency broadcasts receiving an additional 1.2 times gain to ensure clear alarms; and supports audio playback status feedback, reporting playback quality to the server.
[0161] ⑨ Terminal parameter optimization module: Real-time monitoring of terminal computing power utilization, sampling rate switching status, noise reduction effect, etc., and reporting the running status to the server every 30 seconds; receiving parameter adjustment instructions from the server, smoothly adjusting sampling rate, noise reduction parameters, buffer delay, etc., to ensure that the terminal adapts to the system optimization strategy and runs stably.
[0162] 4. Additional functions for the loudspeaker intercom terminal:
[0163] The loudspeaker intercom terminal integrates all the core modules of the IP broadcast terminal, ensuring the ability to receive and process various types of broadcast audio, and achieve synchronous playback and parameter adaptation. In addition, four new dedicated intercom modules are added to achieve seamless linkage between two-way intercom and broadcasting:
[0164] ① Intercom audio acquisition module: integrates an adaptive sampling unit, synchronously matches the sampling rate sent by the server, and acquires two-way intercom voice; performs NLMS adaptive echo cancellation preprocessing to eliminate echo interference during intercom, with computing power consumption ≤30MIPS; preprocesses the acquired intercom voice to reduce noise, suppress background noise, and ensure clear intercom voice.
[0165] ② Intercom Transmission Module: The collected and processed intercom voice is encoded according to the corresponding encoding method (AAC-LD for public broadcast / intercom, and G.711 for emergency intercom), encapsulated into RTP packets, and sent to the server or designated intercom terminal via unicast; it carries intercom identifier and sampling rate identifier to ensure that the server and other terminals can accurately identify the intercom voice.
[0166] ③ Intercom receiving module: Receives intercom voice RTP packets forwarded by the server or sent by other intercom terminals, parses and decodes them, performs noise reduction processing, and plays them through the speaker; synchronously adjusts the playback gain to balance the clarity and volume of the call, avoiding excessively loud or soft volume from affecting the experience; supports real-time playback of intercom voice, with latency controlled within 100ms to meet the needs of two-way communication.
[0167] ④ Intercom Linkage Module: Receives linkage commands from the server to achieve coordinated control of intercom and broadcast; when an intercom is initiated, it automatically pauses the public broadcast in the corresponding area and synchronously matches the sampling rate and noise reduction parameters; during the intercom, if a fire / emergency broadcast is triggered, the intercom is automatically paused, and the fire / emergency voice is played first; it supports one-click conversion of intercom voice to public broadcast, sending conversion commands to the server to achieve seamless connection between intercom and broadcast; after the intercom ends, it automatically restores the public broadcast in the corresponding area to ensure continuous control.
[0168] 5. System Adaptability Design: The system supports low-end MCUs (STM32F1 series, computing power ≤100MIPS) and mid-range MCUs (STM32F4 series, computing power ≤100MIPS). It adapts to all types of embedded terminals, from 200MIPS to high-end ARM / Linux terminals (computing power ≥ 200MIPS); it ensures stable operation of low-end terminals through lightweight algorithm design (such as simplified noise reduction and compensation algorithms) and dynamic parameter adjustment (such as downgrading the sampling rate for low-computing-power terminals); it improves the sound quality and functional experience of high-end terminals by adding AI noise reduction, sub-band noise reduction and other modules; the system supports hot-swapping of terminals, and when a new terminal is connected, it automatically reports the computing power level and the server dynamically distributes the adaptation parameters, eliminating the need for manual configuration and improving the efficiency of engineering deployment.
[0169] Example 2
[0170] This embodiment uses an "industrial plant integrated broadcasting system" as the application scenario. This scenario includes low-end MCU terminals (STM32F103, computing power 80MIPS), mid-range MCU terminals (STM32F407, computing power 150MIPS), and high-end ARM terminals (RK3399, computing power 500MIPS), covering four major types of needs: fire broadcasting, emergency broadcasting, public broadcasting, and intercom. There is low-frequency noise from equipment and ambient noise in the plant area, and the network experiences sudden packet loss (packet loss rate ≤8%). The synchronization error of multiple terminals is required to be ≤2ms. Real-time performance of fire / emergency broadcasting is given priority, while public broadcasting / intercom takes into account both sound quality and latency.
[0171] 1. System Deployment: Deploy 1 broadcast server (equipped with chronyNTP clock service and multi-broadcast collaborative management module), 50 IP broadcast terminals (30 low-end, 15 mid-range, and 5 high-end), and 20 public address intercom terminals (10 mid-range and 10 high-end). All devices communicate via the factory's local area network based on the TCP / IP protocol. The server and terminals are all connected to the same NTP clock source to ensure consistent reference clocks. The factory area is divided into 5 management groups (production area, office area, storage area, fire exit, and central control room). The production area and fire exit mainly use low-end / mid-range terminals, while the office area and central control room mainly use high-end terminals.
[0172] 2. Specific implementation steps (corresponding to invention method S1) S6):
[0173] S1. Audio Preprocessing and Encoding:
[0174] The server collects various audio signals: fire broadcast (alarm voice), emergency broadcast (dispatch instructions), public broadcast (background music, notification), and intercom voice (two-way communication between terminals); when the terminal logs in, it reports the computing power level through the TCP control channel: STM32F103 (low-end), STM32F407 (mid-end), RK3399 (high-end), and the server establishes a terminal computing power table.
[0175] Sampling rate decision: In production areas and fire exits (complex noise scenarios), fire / emergency broadcasts use 32000Hz (mid-range terminals) and 16000Hz (low-end terminals); in office areas and central control rooms (normal scenarios), public broadcasts / intercoms use 48000Hz (high-end terminals) and 32000Hz (mid-range terminals); the server collects terminal feedback every 5 seconds. If the packet loss rate of terminals in the production area is greater than 5% for 3 consecutive seconds, the sampling rate is temporarily reduced to 16000Hz, and then increased to 32000Hz 2 seconds after the network is restored.
[0176] Encoding and Preprocessing: Fire / emergency broadcasts use G.711 encoding, and public broadcasts / intercoms use AAC-LD encoding; production area terminals implement 64-point median filtering (to suppress equipment impulse noise) and 300Hz... 3400Hz bandpass FIR filtering; high-end terminals integrate RNNoise lightweight AI noise reduction module (45MIPS computing power); the intercom terminal performs NLMS adaptive echo cancellation (28MIPS computing power) to ensure two-way communication without echo.
[0177] S2. Audio Interleaving and FEC Redundancy Generation:
[0178] Fire / emergency broadcasting uses an 8×8 interleaving matrix (M=8), while public broadcasting / intercom uses a 6×6 interleaving matrix (M=6); according to mapping relationships The audio frames are shuffled, encapsulated into RTP packets, and carry a sampling rate identifier (0=16k, 1=32k, 2=48k) and a broadcast type identifier.
[0179] FEC redundancy generation: Every 6 service frames form a group (matching the order of the interleaving matrix), with a redundancy of 20% for fire / emergency broadcast and 15% for public broadcast / intercom; FEC redundant frames are generated through XOR operation, carrying the same frame group identifier as the service frames, and are distributed to the corresponding regional terminals via UDP multicast.
[0180] S3. Multi-terminal clock synchronization:
[0181] The server deploys the chronyNTP clock service. The synchronization cycle for fire exits and fire / emergency terminals in the production area is 1 second, the synchronization cycle for low-end terminals is 2 seconds, and the synchronization cycle for mid-range / high-end terminals is 1 second. The terminal calculates the link delay d and clock deviation θ, performs smoothing correction (single correction amplitude ≤ 20ms), and performs long-term drift correction every 10 seconds.
[0182] The fire / emergency terminal now features a deviation warning. When the deviation exceeds 1ms, a secondary synchronization is triggered to ensure that the synchronization error of multiple terminals is ≤2ms. During the synchronization process, the terminal adjusts its local sampling and decoding rates in real time to keep them consistent with the sampling rate sent by the server, thus avoiding audio stuttering caused by synchronization deviation.
[0183] S4. Terminal packet loss detection and hierarchical compensation:
[0184] The terminal receives RTP packets and FEC redundant frames, parses the identification information, and caches the historical maximum frame sequence number. ,pass Calculate the number of lost frames and generate a set of lost frame sequence numbers.
[0185] FEC priority error correction: Prioritize fire protection > emergency > public > intercom; when the number of lost frames is ≤1, use... Restore lost frames; if restoration fails, trigger a resume request.
[0186] Resumption of service request: Fire / Emergency Broadcast , Next, public address / intercom , Next, unicast resume is used; if the retry limit is reached, local compensation is initiated.
[0187] Local compensation: For single-frame packet loss (L=1), low-end terminals use linear interpolation + fade-in / fade-out, and high-end terminals use quadratic interpolation + fade-in / fade-out, increasing the gain of fire / emergency frames by 1.2 times; for short-term continuous packet loss (2≤L≤3), the previous frame's pitch period is reused + interpolation; for long-term packet loss (L>3), low-end terminals gradually mute, high-end terminals use third-order polynomial fitting, fire / emergency terminals trigger local alarms, and intercom terminals trigger interruption prompts.
[0188] S5. Terminal audio noise reduction and synchronized playback:
[0189] Three-level lightweight noise reduction: 32-point moving average filtering is used in the production area terminals, and 16-point filtering is used in the office area / central control room; adaptive noise reduction is also used for fire / emergency broadcasting. Public Broadcasting intercom High sampling rate frames Reduced by 0.05; Rear noise reduction in progress, fire / emergency Public address / intercom High-end terminals utilize eight sub-band noise reduction modules.
[0190] Timing Restoration and Synchronized Playback: By Inverse Mapping Relationship Restore audio frame timing; dynamically adjust buffer delay: At that time, fire / emergency response time is 40ms, public / intercom response time is 60ms; 10ms ≤ If the buffer delay is ≤20ms, the buffer delay is increased by 20ms; When adding 30ms (not exceeding 120ms), the time limit is increased.
[0191] Multi-broadcast type switching: Execute 50 A 60ms fade-in / fade-out mechanism is used, with 60ms for high-sampling-rate frames and 50ms for low-sampling-rate frames. The sampling rate, noise reduction parameters, and buffer delay are adjusted synchronously to achieve seamless switching.
[0192] Table 1: Differentiated Parameter Fitting Table
[0193]
[0194] S6. Iterative optimization and collaborative management of multiple broadcast types:
[0195] Iterative optimization: The server collects terminal feedback every 10 seconds. If the computing power utilization rate of low-end terminals in the production area continues to exceed 80%, the sampling rate is reduced to 16000Hz. If the packet loss rate continues to exceed 5%, the redundancy of the fire / emergency broadcast FEC is increased to 22%. If the clock deviation continues to exceed 1ms, a second synchronization is triggered.
[0196] Collaborative Management and Control: Priority dispatch is as follows: fire > emergency > public > intercom. When a fire alarm is triggered, the system switches to fire mode, other broadcasts are paused, and the sampling rate is uniformly adjusted to 32000Hz (16000Hz for low-end systems). When an intercom terminal initiates an intercom, the public broadcast in the corresponding area is paused, and the sampling rate is matched synchronously. Broadcasts can be issued by area. When emergency instructions are issued in the production area, the public broadcast in the office area will continue to play normally without interference.
[0197] 3. Implementation effect verification:
[0198] Through the deployment of this embodiment, the system achieves seamless integration of four types of broadcasts, with a multi-terminal synchronization error of ≤1.5ms, meeting the needs of fire / emergency coordination; the packet loss compensation success rate is 96.3%, and there is no significant distortion in audio quality when three consecutive frames of packet loss occur; the computing power utilization rate of low-end terminals is ≤75%, and that of high-end terminals is ≤60%, adapting to terminals with different computing power; the noise suppression ratio is improved by 35dB after noise reduction, the fire / emergency alarm voice is clear, and the audio quality of public broadcasts / intercoms is good, with no music noise or echo interference; the system operates stably, and the parameters can be automatically adapted after hot-swapping of terminals, making deployment flexible and fully meeting the needs of multi-scenario integrated broadcasting in industrial plants.
[0199] All formulas in this invention specification The symbols have completely consistent meanings, with no differentiated interpretations or ambiguous scenarios. The unified definition throughout the text is as follows:
[0200] The binary bitwise XOR operation is the core operation method of the FEC forward error correction in this invention. The basic operation rules are as follows: , , , This operation has a self-cancellation property (the result of XORing any data with itself is 0), which can achieve accurate restoration of lost audio data in a single frame.
[0201] It should be noted that, in this document, relational terms such as "one" and "two" are used merely to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, the phrase "comprising an element defined as..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
[0202] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.
Claims
1. A method for processing IP broadcast audio in a converged broadcast system, characterized in that, The method achieves seamless integration of four major broadcast types: emergency broadcast, fire broadcast, public broadcast, and intercom, and is compatible with various embedded terminals ranging from low-end MCUs to high-end ARM / Linux. It includes the following steps: S1. Audio Preprocessing and Encoding: The broadcast server collects multiple types of audio signals and adopts a three-dimensional decision model of broadcast type + terminal computing power + scene noise to dynamically adjust the sampling rate of 16000Hz / 32000Hz / 48000Hz. The terminal completes seamless switching and feedback according to the sampling rate instructions issued by the server; the encoding method is dynamically switched according to the broadcast type, and preprocessing noise reduction and intercom-specific echo cancellation are performed. S2. Audio Interleaving and FEC Redundancy Generation: The server constructs an adjustable interleaving matrix of order 4-8, shuffles and rearranges continuous audio frames and encapsulates them into RTP packets. Every 4-8 consecutive RTP service frames are grouped together, and FEC redundancy frames are generated through XOR operation. Differentiated FEC redundancy is set according to broadcast type. The FEC redundancy frames and service frames are sent down via UDP multicast with the same identifier. S3. Multi-terminal clock synchronization: The server deploys the chronyNTP clock service. The terminal sends synchronization requests and reports its own computing power level according to a dynamic period. It calculates the link delay and clock deviation, performs smoothing correction and long-term drift correction, and triggers secondary synchronization when the deviation of the fire / emergency terminal exceeds 1ms. S4. Terminal packet loss detection and hierarchical compensation: A four-layer protection mechanism is adopted, which is interleaved to resist bursts → FEC priority error correction → unicast resume → local hierarchical compensation. Error correction and resume are processed according to the priority of broadcast type, and differentiated local compensation is performed according to the packet loss level, broadcast type and sampling rate. S5. Terminal audio noise reduction and synchronous playback: The terminal performs three-level lightweight noise reduction on various audio frames, dynamically adjusts the noise reduction parameters to adapt to broadcast type, scene and sampling rate; restores the timing according to the interleaved inverse matrix, dynamically adjusts the adaptive buffer delay, and realizes synchronous playback of multiple terminals and smooth switching of multiple broadcast types. S6. Iterative optimization and multi-broadcast type collaborative management: The server processes terminal resume requests according to priority, and the terminal executes the processing flow in a loop; it realizes priority scheduling of multiple broadcast types, mode switching, intercom and broadcast linkage and regional management, and dynamically optimizes the parameters of the entire link.
2. The method according to claim 1, characterized in that, The specific implementation of adaptive sampling in step S1 includes: during the NTP synchronization or login phase, the terminal reports its computing power level through the TCP control channel, and the server determines the target sampling rate according to the priority of broadcast type > terminal computing power > scene noise; fire / emergency broadcasts prioritize 32000Hz, low-end terminals reduce to 16000Hz, and high-end terminals maintain 32000Hz; public broadcasts / intercoms use 48000Hz for high-end terminals, 32000Hz for mid-range terminals, and 16000Hz for low-end terminals; in noisy scenarios, the sampling rate is increased by one level, and in normal scenarios, it is decreased by one level, and does not exceed the range of 16000Hz-48000Hz.
3. The method according to claim 2, characterized in that, The encoding method switching logic in step S1 is as follows: fire / emergency broadcasts use G.711 encoding to prioritize real-time performance; Public address / loudspeaker intercom uses AAC-LD encoding, balancing sound quality and latency; Preprocessing noise reduction includes 32-point median filtering and 300Hz-3400Hz bandpass FIR filtering. For industrial plant scenes, the median filtering window is increased to 64 points. The intercom terminal uses the NLMS adaptive echo cancellation algorithm, with a computing power consumption of ≤30MIPS; High-end terminals can optionally integrate the RNNoise lightweight AI noise reduction module, with computing power consumption ≤50MIPS.
4. The method according to claim 3, characterized in that, The adjustment logic of the interleaving matrix in step S2 is as follows: fire / emergency broadcasting uses an 8×8 interleaving matrix, and public broadcasting / intercom uses a 4-6×4-6 interleaving matrix; FEC redundancy is set as follows: 20% for fire / emergency broadcasting, and 10%-15% for public broadcasting / intercom. The formula for calculating FEC redundancy frames is: , in For k service frames in the same group, k is adjustable from 4 to 8 and matches the order of the interleaving matrix.
5. The method according to claim 4, characterized in that, The specific implementation of multi-terminal clock synchronization in step S3 includes: the synchronization period of fire / emergency terminals is forcibly set to 1 second, while that of low-end terminals can be adjusted to 2 seconds; the calculation formulas for link delay d and clock deviation θ are respectively... , ,in When the terminal sends a request, When the server receives a request, The time when the server sends a response. This is the moment when the terminal receives the response; The smoothing correction formula is Long-term drift correction is performed every 10 seconds, based on drift error. Adjust the local clock; in, : The original local clock time currently running on the terminal before a single smooth correction is executed; The updated local clock time of the terminal after amplitude limiting and smoothing correction provides a clock reference for subsequent audio timing alignment and multi-terminal synchronous playback; : Dedicated limiting and truncation function, used to constrain the amplitude of a single clock correction, avoiding audio stuttering, dropouts, and timing errors caused by large clock jumps; The system's preset maximum upper and lower limits for single clock correction are fixed engineering parameters. Function execution logic: When When, the function outputs ;when When, the function outputs ;when At that time, the function directly outputs the original offset. ; The global standard reference time output by the broadcast server chronyNTP clock service serves as a unified clock reference for all terminals across the network. : The terminal's local real-time clock before long-term drift correction is performed; The clock drift error caused by the long-term operation of the terminal hardware crystal oscillator is measured in milliseconds. The system performs drift correction every 10 seconds, and uses this error to fine-tune the local clock, suppressing long-term clock offset and ensuring stable synchronization accuracy in long-term broadcast scenarios.
6. The method according to claim 5, characterized in that, The specific implementation of packet loss detection and hierarchical compensation in step S4 includes: determining packet loss by frame sequence number difference and the number of lost frames. ,in The current frame number. The largest frame number in history; FEC priority correction is performed according to the priority order: fire > emergency > public > intercom. When the number of lost frames is ≤1, the FEC restoration formula is used. Restore the lost frames; The conditions for triggering a resume request are L≥2 and FEC restoration failure or keyframe packet loss, and the fire / emergency broadcast timeout period. Number of retries Next, public address / intercom , Secondly, unicast continuation is used to avoid network storms.
7. The method according to claim 1, characterized in that, The specific implementation of the three-level lightweight noise reduction in step S5 includes: Basic noise reduction: 8-32 point moving average filtering, 32 points are used for high sampling rate and noisy complex scenes, and 8-16 points are used for low sampling rate and normal scenes; Adaptive noise reduction: Improved spectral subtraction combined with minimum statistical noise estimation for fire / emergency broadcasting. Public Broadcasting intercom High sampling rate frames The value is slightly lower; Post-noise cancellation: VAD detection distinguishes between voice / silent segments, fire / emergency. Public address / intercom High-end terminals can add sub-band noise reduction for 8 sub-bands, and intercom terminals can add dynamic adjustment of noise threshold.
8. The method according to claim 1, characterized in that, The specific logic for multi-broadcast type collaborative management in step S6 includes: Priority dispatching: Broadcast requests are processed in the order of fire protection > emergency > public > intercom. Mode switching: When a fire / emergency alarm is triggered, the mode is automatically switched and other broadcasts are paused, and the sampling rate is adjusted to the appropriate range; Intercom and broadcast integration: When intercom is in progress, the corresponding area's public broadcast is paused, and the sampling rate is matched synchronously. Intercom voice can be converted to public broadcast with one click. Regional control: Accurately distribute different types of broadcasts to designated areas and match the sampling rate according to the computing power of the regional terminals.
9. An IP broadcast audio processing system for a converged broadcast system, used to implement the method according to any one of claims 1-8, characterized in that, It includes a broadcast server, at least one IP broadcast terminal, and at least one public address intercom terminal, which are connected via an IP network based on the TCP / IP protocol. The broadcast server includes an audio acquisition and encoding module, an interleaving encoding module, an FEC redundancy generation module, an NTP clock service module, a resume response module, a parameter optimization module, and a multi-broadcast collaborative management and control module. The IP broadcast terminal includes a clock synchronization module, a receiver parsing module, a packet loss detection module, an FEC error correction module, a resume request module, a local compensation module, a multi-level noise reduction module, a synchronous playback module, and a terminal parameter optimization module. The loudspeaker intercom terminal integrates the core functions of an IP broadcast terminal and adds an intercom audio acquisition module, an intercom transmission module, an intercom reception module, and an intercom linkage module to achieve seamless linkage between intercom and broadcast.
10. The system according to claim 9, characterized in that, The audio acquisition and encoding module of the broadcast server integrates an adaptive sampling unit, which can dynamically adjust the sampling rate to three levels: 16000Hz / 32000Hz / 48000Hz, dynamically switch the encoding mode, and integrate an echo cancellation module and an optional lightweight AI noise reduction module. The multi-broadcast collaborative management module is used to realize priority scheduling of multiple broadcast types, mode switching, intercom and broadcast linkage, and area management, and coordinate the linkage of parameters of each module. The clock synchronization module of the IP broadcast terminal can synchronize with the NTP server according to a dynamic period, report its own computing power level, perform clock smoothing correction and long-term drift suppression, and support early warning of clock deviation for fire / emergency terminals; the local compensation module can perform differentiated compensation according to packet loss level, broadcast type and sampling rate; the multi-level noise reduction module can dynamically adjust noise reduction parameters to adapt to different scenarios and broadcast type requirements.
Citation Information
Patent Citations
A cloud broadcasting system and method
CN102291244A
Broadcast intercom system based on IP network
CN104468146A