A multi-user concurrent access audio and video intercom dispatching system with intelligent noise reduction function
By employing a multi-dimensional concurrent noise suppression engine and distributed computing, the problems of noise accumulation and latency in multi-user concurrent access audio and video intercom scheduling systems are solved, achieving high-quality communication in extreme acoustic environments.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SHENZHEN ZHILIAN TECH CO LTD
- Filing Date
- 2026-03-28
- Publication Date
- 2026-07-07
AI Technical Summary
Existing technologies in multi-user concurrent audio and video intercom dispatch systems suffer from noise accumulation effects, reduced voice intelligibility, and processing delays, making it difficult to guarantee communication quality in extreme acoustic environments.
It employs a multi-dimensional concurrent noise suppression engine, including spectral slicing analysis, speech feature detail compensation, multi-path concurrent noise residual modeling, and nonlinear mixing weight control. Combined with edge-side preprocessing and central intelligent scheduling, it achieves global acoustic scene perception and collaborative intervention through distributed computing and hardware acceleration.
It effectively eliminates the noise accumulation effect, improves voice intelligibility and real-time performance, ensures clear communication in high-noise environments, and meets the communication needs under high-risk working conditions.
Smart Images

Figure CN122348993A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of communication technology, specifically a multi-user concurrent access audio and video intercom scheduling system with intelligent noise reduction function. Background Technology
[0002] Audio and video intercom dispatch systems are the core hubs in high-risk and complex working conditions, undertaking the functions of issuing instructions, providing information feedback, and coordinating multiple parties. They have evolved from single voice intercom to a multimodal audio and video fusion mode, supporting concurrent access of multiple heterogeneous terminals. The real-time communication performance and accurate voice transmission under extreme acoustic environments are directly related to scene safety. Misreading instructions may cause major accidents or delay rescue.
[0003] Existing technologies employ a single-channel noise reduction scheme based on statistical acoustics principles. By relying on algorithms such as spectral subtraction and Wiener filtering to remove target speech, it can suppress stable background noise and improve the signal-to-noise ratio in ideal scenarios such as single-point intercom and low-density access. On the terminal side, constant noise is attenuated by establishing a noise profile to ensure basic call quality.
[0004] As the demands for multi-user concurrency become more stringent, existing single-path processing solutions suffer from deep-seated technical contradictions. Residual noise converges, resulting in a noise accumulation effect. Aggressive noise reduction strategies impair speech intelligibility. The surge in computing power overhead during concurrent access causes processing delays. These problems are difficult to solve by simple parameter adjustments and constitute the core technical bottleneck of the system.
[0005] Therefore, the present invention provides a multi-user concurrent access audio and video intercom scheduling system with intelligent noise reduction function. Summary of the Invention
[0006] In order to overcome the shortcomings of the prior art, at least one technical problem raised in the background art is solved.
[0007] The technical solution adopted by this invention to solve its technical problem is as follows: A multi-user concurrent access audio and video intercom dispatching system with intelligent noise reduction function, comprising a front-end heterogeneous terminal access cluster, an edge-side preprocessing gateway, a central intelligent intercom dispatching core server, and a multi-dimensional concurrent noise suppression engine. The front-end heterogeneous terminal access cluster includes multiple distributed handheld intercom terminals, vehicle-mounted communication stations, and mobile intelligent access devices; each terminal establishes a bidirectional data connection with the edge-side preprocessing gateway through a wireless or wired communication link.
[0008] Each terminal in the heterogeneous terminal access cluster is equipped with a high-sensitivity audio acquisition array, which consists of at least two microelectromechanical system (MEMS) microphones with a spacing of 20mm to 50mm. The signal-to-noise ratio of the microphones is not less than 65dB, and the frequency response range covers 20Hz to 20kHz. The analog-to-digital converter integrated inside the terminal converts the acquired analog voice signal into a raw pulse code modulation (PCM) data stream with a sampling rate of 48kHz and a bit depth of 24 bits.
[0009] The edge-side preprocessing gateway is deployed on the base station side or in the regional center. Its hardware core consists of a high-performance field-programmable gate array (FPGA) and a multi-core digital signal processor (DSP). The gateway receives the raw audio data stream from the front-end terminal and performs streaming decapsulation and protocol conversion. The gateway is equipped with a preliminary noise profile estimation module, which performs real-time statistical feature extraction of environmental background noise in the non-speech segment of voice activation detection (VAD).
[0010] The central intelligent intercom dispatch core server is the logical control hub of the system, responsible for handling connection management, channel allocation, and synchronous distribution of audio and video access. The server is equipped with a media processing module based on a distributed computing architecture, which supports the simultaneous processing of no less than 128 concurrent audio streams. The server interacts with the multi-dimensional concurrent noise suppression engine through a high-speed Ethernet interface.
[0011] The multidimensional concurrent noise suppression engine is the core module for achieving intelligent noise reduction. The engine includes a spectrum slicing analysis unit, a speech feature detail compensation unit, a multi-channel concurrent noise residual modeling unit, and a nonlinear mixing weight controller.
[0012] The spectrum slicing analysis unit maps the received multiple digital audio streams to the frequency domain space and uses time-frequency overlap and addition technology to divide the signal into several narrow-band sub-bands. Within each sub-band, the analysis unit monitors the energy distribution and harmonic structure of the signal in real time, identifying stationary noise, transient impulse noise, and target speech components.
[0013] The speech feature detail compensation unit reversely adjusts for speech impairments during the noise reduction process. This unit includes a speech formant protector, which first locks the positions of the first three formants of the speech signal before the engine performs the main noise reduction operation. By setting a dynamic protection threshold, it ensures that the spectral energy of the formant region remains intact when attenuating background noise. At the same time, the detail compensation unit extracts the high-frequency harmonic components of the speech signal to compensate for transient details lost due to bandpass filtering or spectral subtraction operations in the denoised signal, thereby improving the intelligibility and naturalness of the speech.
[0014] The multi-channel concurrent noise residual modeling unit is specifically designed to address the noise accumulation problem during concurrent access by multiple users. This unit does not process single signals in isolation, but treats the residual noise of all access channels as a whole acoustic scene. The modeling unit extracts the features of coherent and incoherent noise that are not completely suppressed in each signal, constructs a global noise accumulation model, identifies common-mode noise components by calculating the cross-correlation coefficients of noise between channels, and uses spatial filtering techniques to cancel common-mode noise before physical mixing.
[0015] Preferably, the nonlinear mixing weight controller dynamically allocates mixing weights based on the real-time signal-to-noise ratio and voice activity intensity of each signal. When multiple channels are running concurrently, the controller does not perform a simple linear weighted summation, but adopts a nonlinear suppression strategy. When a channel is in a silent or pure noise state, the channel enters a low-gain suspension mode, and its weight coefficient is reduced to below a preset noise floor threshold, thereby cutting off the injection of its residual noise into the master mixing bus.
[0016] Preferably, the system further includes an audio-visual synchronization alignment module. When multiple users access the audio and video streams concurrently, the module uses nanosecond-level timestamps generated by each terminal to perform sequence calibration on the audio and video frames. In order to compensate for the computational latency caused by the noise reduction engine, the module is equipped with a dynamic cache queue. Based on the real-time processing time of the noise reduction engine, the rendering waiting time of the video frames is automatically adjusted to ensure that the audio-visual phase error output by the scheduling center is within 20ms.
[0017] Preferably, the handheld walkie-talkie terminal is equipped with a dedicated hardware noise reduction acceleration chip. This chip is connected to the MEMS microphone array via an I2S bus. The chip has embedded beamforming logic for near-field voice acquisition. By calculating the arrival time difference between the microphone arrays, a narrow beam pointing towards the speaker's mouth is formed in real time, thus achieving the first level of spatial noise reduction at the physical level.
[0018] Preferably, the edge-side preprocessing gateway has an abnormal sound signal recognition function. When it detects impact noise with ultra-large amplitude, such as mine blasting or electromagnetic arc discharge, the gateway immediately activates the fast automatic gain reduction (AGC) protection circuit to prevent nonlinear clipping distortion from occurring in the subsequent audio processing chain, and at the same time sends a synchronous sudden noise alarm flag to the core server.
[0019] The central intelligent intercom dispatch core server includes a load balancer. When the number of concurrent users surges and the single computing resources become saturated, the scheduler will allocate low-priority audio streams to the backup general-purpose processor cores for basic noise reduction according to preset priority weights, while keeping high-priority command and dispatch command streams in the high-performance computing pool of the multi-dimensional concurrent noise suppression engine to ensure the highest sound quality level of the core commands.
[0020] The execution logic of the multi-path concurrent noise residual modeling unit in the multi-dimensional concurrent noise suppression engine includes: A residual detection point is set at the mixer input to quantify the signal quality of each noise-reduced channel in real time. When the accumulated noise level caused by multiple concurrent channels exceeds the preset scheduling environment tolerance threshold, the engine automatically increases the convergence speed of each branch noise reduction algorithm and introduces a predictive background noise offset signal to achieve noise cancellation in the digital domain.
[0021] The operation process of the nonlinear mixing weight controller is as follows: First, the energy envelope of each input channel is detected; If the detected speech features are not significant and the energy fluctuations conform to the distribution of white noise or colored noise, then the mixing gain of the signal is set to a minimum value. If a strong command voice is detected, the gain is quickly restored to the standard level, and the signal is dynamically compressed to make it stand out in a noisy background.
[0022] Preferably, the system further includes a full-duplex echo cancellation module. During multi-user intercom, when the dispatcher turns on speaker monitoring, the module models and cancels the signal fed back from the speaker to the microphone through an adaptive filter. The module adopts a multi-level convergence strategy to automatically adjust the filter tap coefficients for the reverberation time of different physical environments, thereby eliminating voice overlap interference caused by echo.
[0023] The hardware implementation of the system of this invention involves specific circuit interconnection logic. The audio processing circuit of each access terminal includes a low-noise preamplifier circuit, an anti-aliasing filter, and a high-speed analog-to-digital conversion unit. The cutoff frequency of the anti-aliasing filter is set to 22kHz, and it has a stopband attenuation rate of not less than 80dB, which ensures the purity of the original sampled data. The data transmission link adopts the Secure Real-Time Transmission Protocol (SRTP) and uses a 128-bit encryption algorithm to ensure the privacy of audio and video data in the public network environment.
[0024] Preferably, at the software logic level, the multi-dimensional concurrent noise suppression engine adopts a pipelined processing architecture. Each frame of audio data undergoes DC component removal, pre-emphasis, frame-by-frame windowing (Hamming window), Fourier transform, noise masking calculation, gain smoothing, inverse transform, and deemphasis in the processing sequence. All processing steps are deeply optimized at the register level to minimize storage access latency.
[0025] To address the echo and noise issues when mobile apps connect, the system configures a software-defined audio processing stack for mobile terminals on the gateway side. This processing stack can automatically load the corresponding acoustic configuration parameters based on the hardware model of the mobile device to compensate for the non-linear distortion of the microphone frequency response of a specific terminal.
[0026] The central intelligent intercom dispatch core server is equipped with a redundancy backup mechanism. The system monitors the heartbeat status of the main server in real time. Once a hardware failure or network congestion is detected, the backup server can take over the current concurrent intercom session within 50ms and seamlessly inherit the current noise suppression profile parameters to ensure the continuity of dispatch tasks.
[0027] The system architecture conceived in this invention constructs a comprehensive, three-dimensional intelligent acoustic protection network by simultaneously focusing on four levels: physical acquisition, edge processing, central engine, and logical mixing. Its core lies not only in the optimization of single-point algorithms, but also in solving the problem of non-obvious physical noise accumulation in concurrent environments through the collaborative perception and interactive intervention of multiple signals from a systems engineering perspective. This system-level technological breakthrough provides extremely reliable communication support for command and dispatch under high-risk and extreme conditions.
[0028] Preferably, considering the severe sound reflections and reverberation interference caused by multipath effects in mine roadways, an adaptive de-reverberation unit is additionally integrated into the multidimensional concurrent noise suppression engine. This unit uses long short-term memory features to evaluate the environmental reverberation time and reduces the trailing effect of speech by adjusting the parameters of the inverse filter, thereby further improving the clarity of the boundaries of the command words.
[0029] In emergency command and dispatch, when faced with multiple commanders with different accents and speaking speeds, the system's speech feature detail compensation unit has adaptive equalization capabilities. It can automatically compensate for the gain in the low-frequency or high-frequency bands based on the speaker's spectral envelope characteristics, so that the timbre of each party in the final mixed output tends to be balanced and consistent, reducing the auditory fatigue of dispatchers during long-term duty.
[0030] In addition, the system's audio and video synchronization alignment module exhibits strong robustness when handling concurrent access in weak network environments. When network jitter occurs on a certain access link, the module repairs the damaged audio packets through forward error correction (FEC) and packet loss concealment (PLC) technologies; During the repair process, the system sends a temporary instruction to the noise reduction engine to appropriately reduce the dynamic processing depth of the signal path in exchange for lower reconstruction latency, ensuring that audible and recognizable voice commands can still be transmitted even under extremely poor network conditions.
[0031] Preferably, the multi-user concurrent access audio and video intercom scheduling system with intelligent noise reduction function of the present invention covers the entire chain of innovation from the bottom-level sensor array design, the middle-level protocol stack encapsulation, to the high-level multi-dimensional semantic-level noise reduction engine. It not only breaks through the isolation limitations of traditional noise reduction algorithms in terms of technical principles, but also solves the optimization problems of the three mutually restrictive core indicators of multi-concurrency, strong noise, and high real-time in engineering practice. By upgrading noise processing from "single-path repair" to "global control", the present invention establishes a brand-new technical paradigm for the field of audio and video scheduling, which has strong non-obviousness and significant technical progress.
[0032] In the implementation plan for power line inspection scenarios, considering that the electromagnetic induction noise generated by high-voltage cables often exhibits specific power frequency harmonic characteristics (such as 50Hz and its harmonics), the system's spectrum slicing analysis unit incorporates a notch filter array. This array can automatically scan and lock onto power frequency harmonic interference points based on the power grid environment, performing precise energy trapping with an extremely narrow bandwidth. This refined processing of specific industrial background noise, combined with subsequent detail compensation, enables inspection personnel to maintain clear communication quality while working inside substations.
[0033] The system's dynamic decision-making process is entirely based on real-time physical quantity feedback; For example, a nonlinear mixing weight controller performs actions based on the ratio of "voice energy / residual noise energy" (SNR_R) of the current access channel: When SNR_R is lower than the first preset threshold, complete mute is executed. When SNR_R is between the first and second thresholds, nonlinear suppression is performed; When SNR_R is higher than the second threshold, the original tone pass-through is executed and gain compensation is added. This control logic based on deterministic thresholds and dynamic feedback ensures the stability of system operation.
[0034] The connection logic between the modules in this invention is clear and the functions are well-defined. The electrical signal collected by the front-end microphone is amplified by the operational amplifier and then enters the ADC chip. The I2S signal output by the ADC enters the local MCU for protocol encapsulation and is then sent to the edge gateway through the Wi-Fi / 4G / 5G module. The edge gateway sends the aggregated RTP packets to the server in the data center through a fiber optic leased line. After the core server calls the multi-dimensional noise reduction engine to process the data, it mixes and re-encodes the various audio streams and pushes them to each monitoring terminal through the distribution node. The entire data flow is absolutely deterministic, and the electrical parameters and interface protocols (such as RTSP, RTMP, SIP, I2S, PCIE, etc.) between modules strictly follow international standards and are mutually matched, ensuring the overall compatibility and reliability of the system.
[0035] The system successfully eliminated the noise accumulation effect through multiple means, including multi-channel concurrent noise residual modeling, speech formant protection, nonlinear weight control, and audio-visual phase calibration. It maintained extremely high speech intelligibility in noisy scenarios and met the stringent requirements of complex scheduling environments for high-performance communication systems.
[0036] Preferably, in order to ensure system stability under extremely high concurrency (e.g., more than 500 terminals accessing at the same time), the central intelligent intercom dispatch core server adopts a resource prediction and scheduling algorithm, which monitors the processor utilization and video memory bandwidth usage of the multi-dimensional concurrent noise suppression engine in real time. When the pre-computational resources reach 90% of the threshold within the next three processing frame cycles, the system will automatically start the hierarchical processing logic: for secondary background audio streams accessed by ordinary mobile apps, the FFT points will be reduced from 1024 to 512 to reduce the amount of computation; while for intercom streams marked as "emergency command center", the maximum sampling rate and the highest complexity of noise reduction parameters will always be maintained, ensuring "zero damage" to core business from the perspective of resource allocation.
[0037] Preferably, the detail compensation unit inside the multi-dimensional concurrent noise suppression engine also integrates preset profiles for different languages and pronunciation habits; Since the scheduling scenario may involve personnel from different regions, the system automatically matches the corresponding acoustic correction curve by analyzing the distribution range of the fundamental frequency (F0) of the input audio in real time. For example, for pronunciation styles with rich high-frequency energy, the system will automatically fine-tune the corner frequency of the high-pass filter to retain more fricative details, thereby ensuring the accuracy of instruction transmission in cross-cultural and cross-regional collaboration.
[0038] Preferably, in terms of security design, in addition to encrypting audio data, the system also performs integrity verification on the processing logs of the noise reduction engine; Each noise-reduced and output mixing instruction contains a summary of the engine parameters that processed that frame in its metadata. This has significant legal evidentiary value during the accident retrospective phase, as it can prove that the voice recorded by the dispatch center was generated through deterministic and auditable algorithmic logic, ruling out the possibility of human intervention or random failures causing instruction distortion.
[0039] Preferably, the microphone of the handheld intercom terminal adopts hydrophobic nano-coating technology and is combined with an acoustic damping mesh to reduce turbulence noise (wind noise) generated when strong winds blow through the aperture. This physical windproof noise reduction, together with the backend digital beamforming and spectrum suppression, forms a three-tiered defense, enabling the system to extract clear and distinguishable human voices even in severe weather conditions with wind speeds reaching 15 m / s (level 7 wind). This has strong practicality in coastal power line inspection or open-pit mining operations.
[0040] Preferably, the audio-visual synchronization alignment module in the central intelligent intercom dispatch core server adopts an adaptive interleaving buffer strategy to address the packet loss and out-of-order problem caused by multi-path network transmission. The module dynamically adjusts the size of the receiving end buffer based on the real-time measured round-trip time (RTT) variance. In mobile network environments with extremely high packet loss rates (e.g., 10% to 20%), the module interacts with the noise reduction engine through packet loss recovery logic and feedback, and while reconstructing lost audio frames, it automatically reduces the gain smoothing coefficient of the noise reduction engine to suppress the instantaneous "mechanical" noise caused by packet repair, thereby ensuring the continuity of the listening experience.
[0041] The system of this invention also involves refined control of terminal power consumption. Since multi-channel acquisition and noise reduction have high requirements for computing resources, when the handheld terminal detects no local voice input (silent state) for more than 500ms, it automatically enters ultra-low power standby mode, keeping only the VAD monitoring circuit working. Once the speaker starts to speak, the circuit is instantly activated within 10ms, ensuring the complete acquisition of intercom commands, while increasing the standby time of the terminal by about 40%.
[0042] The beneficial effects of this invention are as follows: 1. The present invention provides a multi-user concurrent access audio and video intercom scheduling system with intelligent noise reduction function. By introducing a multi-channel concurrent noise residual modeling unit and global acoustic scene perception, the system can identify and cancel common-mode noise between channels, and dynamically cut off the noise injection path of idle channels using nonlinear mixing control technology.
[0043] 2. The multi-user concurrent access audio and video intercom dispatching system with intelligent noise reduction function described in this invention, through a voice feature detail compensation unit and a formant protection mechanism, accurately preserves and enhances the formant features and high-frequency details of the voice while suppressing strong noise. This improves the recognizability of dispatching instructions in high-noise environments and reduces the risk of misoperation.
[0044] 3. The multi-user concurrent access audio and video intercom scheduling system with intelligent noise reduction function described in this invention, through the hardware acceleration scheme of FPGA and DSP, and the dynamic cache alignment mechanism, the system can control the total end-to-end latency of the audio link within 150ms when executing multi-dimensional complex noise reduction algorithms, thus ensuring the immediacy of intercom scheduling.
[0045] 4. The multi-user concurrent access audio and video intercom dispatch system with intelligent noise reduction function described in this invention, through the design of heterogeneous terminal access cluster, enables the system to be compatible with communication equipment of various technical specifications. The distributed deployment scheme of edge preprocessing gateway reduces the computing power burden of the central server. The system can adaptively adjust the noise reduction profile according to the actual working conditions, and has extremely high engineering practical value. Attached Figure Description
[0046] The invention will now be further described with reference to the accompanying drawings.
[0047] Figure 1 This is a structural block diagram of a multi-user concurrent access audio and video intercom scheduling system with intelligent noise reduction function according to the present invention. Detailed Implementation
[0048] To make the technical means, creative features, objectives and effects of this invention easier to understand, the invention will be further described below in conjunction with specific embodiments.
[0049] like Figure 1 As shown in the embodiment of the present invention, a multi-user concurrent access audio and video intercom scheduling system with intelligent noise reduction function includes a front-end heterogeneous terminal access cluster, an edge-side preprocessing gateway, a central intelligent intercom scheduling core server, and a multi-dimensional concurrent noise suppression engine coupled through a high-speed industrial-grade communication link. At the hardware deployment level, each mobile handheld terminal in the front-end heterogeneous terminal access cluster adopts a structural design that meets IP68 protection standards, and its key internal audio pickup component is a high-sensitivity audio acquisition array. The array consists of two high-performance microelectromechanical systems (MEMS) microphones, with the physical distance between the two microphones strictly limited to between 20mm and 50mm. This specific physical span is set based on the beamforming algorithm requirements for near-field acquisition of human voices, aiming to obtain first-order directional gain through spatial differential effects. The MEMS microphone has a signal-to-noise ratio of no less than 65dB, and its frequency response curve remains highly flat in the range of 20Hz to 20kHz, ensuring the integrity of the original information during the audio acquisition stage.
[0050] Specifically, in the internal circuit design of the handheld terminal, the acquired analog voice signal first passes through an ultra-low noise preamplifier, and then enters a high dynamic range analog-to-digital converter. The converter oversamples the analog signal at a sampling rate of 48kHz and outputs a 24-bit deep raw pulse code modulation data stream. To reduce structural noise caused by chassis vibration or hand friction at the physical layer, the microphone assembly is encapsulated with an independent silicone suspension bracket; During the data transmission phase, the communication module inside the terminal encapsulates the PCM data stream in a secure real-time transmission protocol data packet and uses a 128-bit encryption algorithm to perform real-time encryption on the stream. Furthermore, the handheld walkie-talkie terminal integrates a hardware noise reduction and acceleration chip. This chip directly interfaces with the acquisition array via the I2S bus and uses the instruction set embedded in the chip to execute hardware-level beamforming logic. Based on the signal arrival time difference between the microphones, it calculates in real time and forms a narrow beam pointing towards the speaker's mouth area, thereby effectively isolating background noise from non-target areas from the acoustic front end.
[0051] The edge-side preprocessing gateway acts as a buffer and processing node between the access cluster and the central server. Its hardware core is composed of a high-performance field-programmable gate array (FPGA) and a multi-core digital signal processor (DSP). This heterogeneous computing architecture allows the gateway to perform real-time decapsulation and preliminary processing of massive audio streams with extremely low power consumption. The gateway internally deploys a preliminary noise profile estimation module, which integrates a highly sensitive voice activation detector. When the signal is detected to be in a non-speech segment, the module will perform feature extraction on the ambient noise based on statistical distribution, including calculating the energy variance, spectral centroid and zero-crossing rate of the noise, and generating real-time noise profile parameters. These parameters will be uploaded to the central core server along with the audio stream. For specific extreme scenarios, such as mining or power maintenance environments, the edge gateway is also equipped with an abnormal sound signal recognition circuit. When the instantaneous increase of the input signal exceeds the preset dynamic range threshold (such as blasting impact or high-voltage discharge sound), the fast automatic gain reduction circuit in the gateway can respond within microseconds, forcibly reducing the front-end input gain to prevent harmonic distortion caused by nonlinear clipping in the subsequent digital processing chain, and injecting a sudden noise alarm flag in the packet header to notify the dispatch center.
[0052] The central intelligent intercom dispatch core server adopts a distributed computing architecture, and its internal media processing module is responsible for managing the channel allocation and synchronous distribution of no less than 128 audio and video streams. The server motherboard interacts directly with the multi-dimensional concurrent noise suppression engine through a high-bandwidth PCIe interface to ensure the real-time data exchange during large-scale concurrent access. The load balancer inside the server is the key to resource allocation. It dynamically adjusts the allocation of computing resources according to preset business priority weights (such as command personnel flow, ordinary operation personnel flow, and background environment flow). When the system detects that the number of concurrent accesses is close to the peak computing power, the scheduler will automatically allocate the lower priority audio streams to the backup general-purpose processor cores to execute the basic spectrum suppression algorithm, while keeping the core command stream in the multi-dimensional concurrent noise reduction engine with hardware acceleration units, thereby ensuring the absolute clarity and audio quality priority of the core decision commands at the system level.
[0053] The multidimensional concurrent noise suppression engine serves as the technical hub of the system, and its core components include a spectrum slicing analysis unit, a speech feature detail compensation unit, a multi-channel concurrent noise residual modeling unit, and a nonlinear mixing weight controller. The spectrum slicing analysis unit uses time-frequency overlap and addition technology to map each input digital audio stream to the frequency domain space and divide it into several groups of narrowband sub-band sequences; Within each subband, the analysis unit monitors the energy distribution and harmonic structure of the signal in real time. Through fine-grained scanning of the spectral envelope, it identifies stationary noise, sudden transient impact noise, and reverberation components caused by multipath effects.
[0054] The speech feature detail compensation unit is the key to solving the noise reduction damage in this invention. The unit is equipped with a speech formant protector. Before the noise reduction engine performs a strong spectrum reduction operation, the protector will first lock the positions of the first three formants of the current speech frame through a spectrum energy peak tracking algorithm. By setting a dynamic protection threshold, the system artificially raises the lower limit of the gain in the formant region when calculating the noise suppression gain function, ensuring that the key characteristic frequencies of human voice are not mistakenly damaged. On this basis, the detail compensation unit extracts the high-frequency harmonic components of the speech signal and compensates for the transient details lost due to conventional filtering operations in the denoised signal. This processing method can significantly improve the naturalness of the denoised speech and eliminate the "mechanical sound" or "watery sound" commonly found in traditional denoising algorithms.
[0055] The multi-channel concurrent noise residual modeling unit provides a system-level solution to the problem of physical noise accumulation during multi-user access. This unit does not treat each access signal as an isolated entity, but rather extracts the residual noise features that are not fully suppressed from all access channels. It constructs a global acoustic scene accumulation model in the digital domain. By calculating the cross-correlation coefficient of residual noise between any two channels, the modeling unit can identify noise components with common-mode characteristics (such as the noise from mine exhaust fans simultaneously collected by various terminals). Utilizing spatial filtering techniques and phase cancellation logic, the system performs phase cancellation processing on these common-mode noises before physical mixing, fundamentally curbing the increase in noise floor level caused by multi-channel superposition.
[0056] The nonlinear mixing weight controller dynamically adjusts the gain ratio of each signal in the mixing bus based on the real-time signal-to-noise ratio and voice activity intensity feedback. This controller abandons the traditional equal-gain weighted summation method and adopts a nonlinear suppression strategy. When the signal-to-noise ratio of a certain channel is lower than the preset minimum threshold and the voice activation detection determines that there is no valid voice, the controller will switch the channel into a low-gain suspension mode to reduce the energy injected into the mixing bus to below the system noise floor. Once a strong command voice is detected, the controller will restore the standard gain within milliseconds and activate the dynamic range compression module to ensure that important commands are highly recognizable in complex backgrounds.
[0057] To achieve perfect alignment of audio and video, the central server also integrates an audio-visual synchronization alignment module. This module utilizes nanosecond-level high-precision timestamps generated by each terminal to perform sequence calibration on audio and video frames. Since the multi-dimensional noise reduction engine generates a certain algorithm latency, this module maintains a dynamic buffer queue and automatically adjusts the waiting time of video frames at the rendering end based on the real-time processing time feedback from the engine. This closed-loop adjustment mechanism ensures that the audio-visual phase error output by the scheduling center is always kept within 20ms, achieving zero-time-difference synchronization of human visual and auditory perception.
[0058] Looking further into the details of engineering implementation, the software logic of the multi-dimensional concurrent noise suppression engine runs on a deeply optimized instruction set architecture. Each frame of audio data needs to go through DC component removal, frame-by-frame windowing based on Hamming window, fast Fourier transform, noise masking calculation, gain smoothing based on prior signal-to-noise ratio estimation, inverse Fourier transform, and de-emphasis processing in sequence. All of these processing flows are deeply pipelined at the FPGA register level to minimize the latency caused by memory access. When dealing with special scenarios such as mines, the adaptive dereverberation unit, which is additionally integrated into the engine, uses long short-term memory features to evaluate the reflection characteristics of the tunnel. By dynamically adjusting the tap coefficients of the anti-filter, it reduces the speech trailing caused by multiple reflections and enhances the clarity of word boundaries.
[0059] In the implementation of this system in power inspection scenarios, the notch filter array built into the spectrum slicing analysis unit achieves precise removal of power frequency harmonic interference. The array can automatically scan for 50Hz and its harmonic interference points according to the ambient electromagnetic field strength and perform narrow-bandwidth energy suppression. Meanwhile, the system's full-duplex echo cancellation module models the signal fed back to the microphone from the dispatcher's speaker using an adaptive filter, and employs a multi-level convergence strategy to cancel the echo, ensuring clear voice during multi-party calls.
[0060] A comparative experiment was conducted under a simulated extreme industrial scenario (ambient noise floor of 105dB, simulating simultaneous access of 16 terminals). The experiment was divided into three groups: Comparative Example 1 uses a traditional single-channel noise reduction and linear mixing system based on spectral subtraction; Comparative Example 2 uses an existing noise reduction algorithm based on a simple neural network; Example 1 uses the multi-user concurrent access audio and video intercom scheduling system with intelligent noise reduction function described in this invention.
[0061] The experimental data are recorded in Table 1 below: Table 1: Comparison of Performance Indicators of Audio and Video Intercom Dispatch Systems Based on the quantitative data in Table 1, it can be clearly observed that Embodiment 1 of the present invention demonstrates a significant generational advantage in several core technical indicators. First, under extreme conditions with 16 concurrent access channels, this invention, through the collaboration of a multi-channel concurrent noise residual modeling unit and a nonlinear mixing controller, suppressed the total noise floor level to -58.5dBm, a reduction of more than 23dB compared to Comparative Example 1. This means that during multi-party calls, the background noise heard by the dispatcher is extremely faint, completely eliminating the negative effect of "multi-channel noise accumulation." In terms of speech quality metrics (PESQ score) and command recognition, this invention, with its speech feature detail compensation unit and formant protector, maintains a recognition rate of 96.8% even under strong interference. In contrast, although Comparative Example 2 uses neural network noise reduction, its recognition rate and real-time performance are significantly inferior due to the lack of system-level modeling for concurrent scenarios and speech impairment caused by computational resource limitations. Especially in end-to-end latency control, this invention, through hardware-level pipeline optimization of FPGA / DSP, controls the latency to 142ms, which is far superior to the 320ms of Comparative Example 2, meeting the stringent real-time requirements of emergency dispatch.
[0062] In the specific details of circuit connection and communication protocol implementation, the audio acquisition array of the front-end access terminal is connected to the ADC chip through a high-precision operational amplifier link. The digital output interface of the ADC adopts the I2S standard to transmit the digital voice signal to the local MCU. The MCU is responsible for encapsulating audio frames, nanosecond-level timestamps, and terminal state vectors into RTP protocol packets. The data transmission layer uses the Secure Real-Time Transport Protocol (SRTP) and ensures data security through a 128-bit AES encryption algorithm. The edge gateway sends the aggregated bitstream to the core server in the data center via a fiber optic leased line or a 5G slice network. The core server then pushes the various audio streams to the computing power pool of the multi-dimensional noise reduction engine in real time through a high-performance Ethernet switch.
[0063] For scenarios with extremely high concurrency (such as more than 500 connected terminals), the resource prediction and scheduling algorithm of the central server plays a core role. This algorithm monitors the occupancy rate of the internal registers and bus load of the engine in real time. When it predicts that the computing resources are about to reach the critical threshold of 90%, it will automatically start the hierarchical noise reduction logic. Specifically, for channels marked as "emergency command", the system always maintains a 1024-point FFT resolution and a full-parameter noise reduction profile; while for peripheral ordinary auxiliary channels, the system will smoothly downsample them to 512-point FFT processing to reduce the amount of computation by about 40%, thereby prioritizing the absolute clarity of the core command chain.
[0064] In terms of physical protection and structural noise reduction of the handheld terminal, the pickup hole adopts hydrophobic nano-coating technology and has a built-in multi-layer stainless steel acoustic damping mesh. This physical structure can not only block liquid from entering, but also effectively reduce the random turbulence noise (wind noise) generated when strong wind blows through the aperture. Combined with the back-end digital beamforming technology, the system can still extract clear human voices from both physical and digital dimensions in an environment of 15 m / s (level 7 gale). This has extremely high engineering value for coastal power facility inspection or open-pit mine dispatching.
[0065] Furthermore, this invention also relates to an optimized design for terminal power management. The audio processing circuit of the handheld terminal incorporates ultra-low power standby logic based on threshold triggering. When the VAD monitoring module detects no voice input for 500ms consecutively, it automatically shuts down the high-power ADC and part of the processing core, retaining only a weak monitoring current. Once the sound energy exceeds the preset trigger threshold, the circuit can achieve a full-function cold start within 10ms, ensuring complete capture of the first word of the intercom and increasing the average battery life of the terminal by more than 40%.
[0066] The system's security is not only reflected in data encryption, but also includes comprehensive log auditing and parameter backtracking functions; While outputting the mixing results, the multi-dimensional concurrent noise suppression engine generates a metadata summary for each frame of instructions, which includes the current engine processing parameters (such as noise reduction gain, formant offset, mixing weight factor, etc.). In the retrospective phase of a production safety accident, these records can serve as legal evidence to prove that the generation process of the dispatch recordings followed deterministic algorithmic logic, eliminating the possibility of random interference compromising the authenticity of the instructions.
[0067] In implementations for high humidity and high temperature environments, all chips, passive components and connectors selected for the system strictly adhere to industrial-grade standards (-40°C to +85°C) and have undergone rigorous salt spray corrosion and electromagnetic compatibility tests. The noise reduction processing link of the central server adopts SIMD (Single Instruction Multiple Data) instruction set optimization, which greatly improves the efficiency of parallel processing of audio data.
[0068] In summary, this invention provides a system-level optimal solution for noise suppression in multi-user concurrent access scenarios by reconstructing the entire process of the audio and video intercom dispatch system, from bottom-level sensor acquisition, edge preprocessing, central-level concurrent modeling to final nonlinear mixing. It not only achieves breakthroughs in suppressing the accumulation of physical noise, but also reaches extremely high industrial application standards in ensuring voice intelligibility and system real-time performance. It provides reliable technical support for communication dispatch in extreme and high-risk environments. Whether deep in mines, inside substations, or at emergency command sites, this system can ensure that every critical instruction is delivered to decision-making terminals at all levels in a clear, accurate, and timely manner under harsh conditions of high noise and high concurrency.
[0069] The foregoing has shown and described the basic principles, main features, and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited to the above embodiments. The embodiments and descriptions in the specification are merely illustrative of the principles of the invention. Various changes and modifications can be made to the invention without departing from its spirit and scope, and all such changes and modifications fall within the scope of the present invention as claimed. The scope of protection of the present invention is defined by the appended claims and their equivalents.
Claims
1. A multi-user concurrent access audio and video intercom dispatch system with intelligent noise reduction function, characterized in that, The system includes: The front-end heterogeneous terminal access cluster includes multiple distributed handheld walkie-talkies, vehicle-mounted communication stations, and mobile intelligent access devices, which are used to collect environmental acoustic signals and convert them into encrypted digital audio streams. The edge-side preprocessing gateway is connected to the front-end heterogeneous terminal access cluster via a communication link. Its hardware core consists of a field-programmable gate array (FPGA) and a multi-core digital signal processor (DSP), which is used to perform streaming decapsulation, protocol conversion, and preliminary noise profile estimation on the received raw audio data stream. The central intelligent intercom dispatch core server is connected to the edge-side preprocessing gateway and is responsible for channel management, media distribution and resource scheduling of multiple parallel audio and video streams; The multidimensional concurrent noise suppression engine interacts with the central intelligent intercom scheduling core server. The multidimensional concurrent noise suppression engine integrates a spectrum slicing analysis unit, a voice feature detail compensation unit, a multi-channel concurrent noise residual modeling unit, and a nonlinear mixing weight controller. The spectrum slicing analysis unit uses time-frequency overlap and addition technology to map multiple digital audio streams to narrowband sub-bands in the frequency domain, and identifies noise components and target speech components based on the energy distribution and harmonic structure within each sub-band. The speech feature detail compensation unit locks the position of the speech signal's formant by using a built-in speech formant protector, and performs high-frequency harmonic detail compensation on the noise-reduced signal based on a dynamic protection threshold. The multi-channel concurrent noise residual modeling unit extracts the residual noise features of each access channel and constructs a global noise accumulation model, using spatial filtering technology to cancel the common-mode noise components between channels; The nonlinear mixing weight controller performs a low-gain suspension operation on channels that are in a silent or pure noise state based on the real-time signal-to-noise ratio and voice activity intensity of each signal, and dynamically controls the injection weight of each audio stream into the mixing bus.
2. The multi-user concurrent access audio and video intercom dispatch system with intelligent noise reduction function according to claim 1, characterized in that: Each terminal in the heterogeneous terminal access cluster is equipped with a high-sensitivity audio acquisition array, which consists of at least two microelectromechanical system (MEMS) microphones. The analog-to-digital converters integrated within each terminal convert the analog voice signal into a raw pulse code modulation (PCM) data stream, and then encrypt it using the Secure Real-Time Transmission Protocol (SRTP) before uploading it to the edge-side preprocessing gateway.
3. A multi-user concurrent access audio and video intercom dispatching system with intelligent noise reduction function according to claim 1, characterized in that: The edge-side preprocessing gateway is equipped with a preliminary noise profile estimation module. The preliminary noise profile estimation module uses speech activation detection (VAD) to identify non-speech segments and extracts real-time statistical features of the environmental noise floor of non-speech segments, including energy variance, spectral centroid, and zero-crossing rate, to generate noise profile parameters. The edge-side preprocessing gateway also includes an abnormal sound signal recognition circuit. When an impact noise signal with an instantaneous increase exceeding a preset dynamic range threshold is detected, the edge-side preprocessing gateway activates the fast automatic gain control (AGC) protection circuit to perform amplitude limiting and suppression, and simultaneously sends a sudden noise alarm flag to the central intelligent intercom dispatch core server.
4. A multi-user concurrent access audio and video intercom dispatching system with intelligent noise reduction function according to claim 1, characterized in that: The central intelligent intercom dispatch core server is equipped with a media processing module based on a distributed computing architecture, which supports the processing of no less than 128 audio and video streams. The central intelligent intercom dispatch core server also includes a load balancer. The load balancer monitors the computing power resource utilization rate in real time. When the number of concurrent users exceeds a preset threshold, the load balancer allocates the low-priority audio stream to the backup general-purpose processor core for basic noise reduction processing according to the preset priority weight, while retaining the high-priority command and dispatch instruction stream in the high-performance computing power pool of the multi-dimensional concurrent noise suppression engine for deep noise reduction.
5. A multi-user concurrent access audio and video intercom dispatching system with intelligent noise reduction function according to claim 1, characterized in that: The execution logic of the speech formant protector in the speech feature detail compensation unit is as follows: Before the multidimensional concurrent noise suppression engine performs spectral reduction operation, the positions of the first three formants of the current speech frame are locked using the spectral energy peak tracking algorithm; By locking the spectral energy of the first three resonant peak regions through a dynamic protection threshold, it is ensured that the lower limit of the gain of the resonant peak regions is not lower than a preset value during the attenuation of background noise. By extracting the high-frequency harmonic components of the speech signal, the transient detail features lost due to bandpass filtering are superimposed onto the denoised audio signal to compensate for them.
6. A multi-user concurrent access audio and video intercom dispatching system with intelligent noise reduction function according to claim 1, characterized in that: The multi-channel concurrent noise residual modeling unit is equipped with multiple residual detection points at the mixer input, which are used to quantify the signal quality of each independent noise reduction channel in real time. The multi-channel concurrent noise residual modeling unit identifies the common-mode noise components with acoustic coherence collected by multiple terminals by calculating the cross-correlation coefficient of the unsuppressed residual noise between any two channels. When the accumulated noise level caused by multiple concurrent noises exceeds the preset scheduling environment tolerance threshold, the multi-concurrent noise residual modeling unit uses spatial filtering technology to perform anti-phase cancellation processing on the identified common-mode noise to suppress the multi-noise accumulation effect.
7. A multi-user concurrent access audio and video intercom dispatching system with intelligent noise reduction function according to claim 1, characterized in that: The nonlinear mixing weight controller executes the following control logic based on the energy envelope of each input channel and the ratio of speech energy to residual noise energy, SNR_R: If the ratio SNR_R is lower than the first preset threshold, the channel is determined to be in a silent or pure noise state. Then, the mixing weight coefficient of the signal is reduced to below the preset noise floor threshold, and the channel enters a low gain suspension mode. If the ratio SNR_R is between the first preset threshold and the second preset threshold, a nonlinear suppression strategy is executed to smooth the background noise. If the ratio SNR_R is higher than the second preset threshold, it is determined that there is strong command voice in the channel. Then, the mixing gain is quickly restored to the standard level, and the dynamic range compression module is activated to perform gain compensation on the signal, making it stand out in the mixing bus.
8. A multi-user concurrent access audio and video intercom dispatching system with intelligent noise reduction function according to claim 1, characterized in that: The central intelligent intercom dispatch core server also includes an audio-visual synchronization alignment module, which uses nanosecond-level timestamps generated by each terminal to perform sequence calibration on concurrently accessed audio and video frames. The audio-visual synchronization alignment module is equipped with a dynamic cache queue, which dynamically adjusts the rendering wait time of video frames based on the real-time processing time feedback from the multi-dimensional concurrent noise suppression engine. When network packet loss is detected, the audio-visual synchronization alignment module repairs the audio packet using forward error correction (FEC) and packet loss hiding (PLC) technology, and simultaneously instructs the multidimensional concurrent noise suppression engine to lower the gain smoothing coefficient.
9. A multi-user concurrent access audio and video intercom dispatching system with intelligent noise reduction function according to claim 2, characterized in that: The handheld walkie-talkie terminal is equipped with a dedicated hardware noise reduction and acceleration chip, which is connected to the audio acquisition array via an I2S bus. The hardware noise reduction acceleration chip has a beamforming logic embedded in it. By calculating the signal arrival time difference between each microphone in the audio acquisition array, it forms a narrow beam pointing towards the speaker's mouth in real time to perform first-level spatial noise reduction. The handheld walkie-talkie terminal's microphone hole is coated with a hydrophobic nano-coating and combined with a multi-layer stainless steel acoustic damping mesh to physically reduce turbulent wind noise.
10. A multi-user concurrent access audio and video intercom dispatching system with intelligent noise reduction function according to claim 1, characterized in that: The system also includes a full-duplex echo cancellation module, which models and cancels the signal fed back to the microphone from the dispatcher's side speaker through an adaptive filter, and adjusts the filter tap coefficients in real time according to the reverberation time of different physical environments. The multidimensional concurrent noise suppression engine also integrates an adaptive de-reverberation unit, which uses long short-time memory features to evaluate the environmental reverberation time and adjust the parameters of the inverse filter.