A software radio intelligent voice noise reduction system design method

CN121122297BActive Publication Date: 2026-09-08广东公信智能会议股份有限公司
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511267528.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-09-05
Publication Date
2026-09-08
Estimated Expiration
2045-09-05

AI Technical Summary

Technical Problem

基于通用处理器(CPU)或数字信号处理器(DSP)的传统软件无线电平台虽然具有灵活性优势,但在处理日益复杂的智能语音降噪算法(如深度神经网络)时面临显著瓶颈:其串行处理架构难以高效满足语音信号处理固有的高吞吐量和并行计算需求,导致处理延迟增大,难以达到实时语音交互所需的毫秒级响应要求,尤其在宽带或多通道场景下性能受限更为突出

Benefits of technology

本发明软件无线电智能语音降噪系统设计为采用Xilinx Zynq7020芯片与AD9361芯片的软件无线电智能语音降噪系统,通过创新的异构计算架构,实现了卓越的性能提升。Xilinx Zynq7020芯片集成的ARM处理器与FPGA紧密协同,充分发挥了硬件并行加速优势,使得复杂的DCNN智能降噪模型能够高效运行,显著提升了语音清晰度和信噪比。同时,软件无线电的核心特性——灵活性,在本系统中得到充分体现:AD9361射频前端支持宽频段可调,使系统能够无缝适配多种无线通信频段和标准;而基于FPGA的可重构能力,更支持DCNN降噪算法在不中断服务的情况下进行远程动态更新与优化,赋予了系统强大的环境适应性和持续进化能力。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121122297B_ABST
    Figure CN121122297B_ABST
Patent Text Reader

Abstract

The application discloses a software radio intelligent voice noise reduction system design method, through an analog / digital conversion module, radio frequency signals are collected, frequency conversion and digitization are carried out, input into a pre-processing module for anti-aliasing filtering and voice frame processing, then input into a DCNN hardware accelerator for convolution, activation and pooling operation, input into a DQPSK modulation and demodulation module for differential encoding, pulse shaping filtering and quadrature carrier modulation, to generate a baseband signal to be transmitted, and return to the analog / digital conversion module for frequency conversion and radio frequency transmission, when strong interference is detected in the current working frequency band, resulting in a decrease in communication quality, according to a preset strategy, a control instruction is automatically generated, the analog / digital conversion module is driven to quickly switch to a backup clean communication frequency band, and the communication link robustness is ensured. The application can meet the requirements of high real-time, high flexibility and high performance intelligent voice noise reduction in a complex noise environment, and has high practicability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of communication technology, and more specifically to a design method for a software-defined radio intelligent voice noise reduction system. Background Technology

[0002] Achieving high-quality, low-latency real-time voice communication under complex electromagnetic environments and strong noise interference is a key challenge in current wireless communication. While traditional software-defined radio platforms based on general-purpose processors (CPUs) or digital signal processors (DSPs) offer advantages in flexibility, they face significant bottlenecks when handling increasingly complex intelligent speech denoising algorithms (such as deep neural networks). Their serial processing architecture struggles to efficiently meet the inherent high throughput and parallel computing demands of speech signal processing, leading to increased processing latency and making it difficult to achieve the millisecond-level response requirements of real-time voice interaction. This performance limitation is particularly pronounced in broadband or multi-channel scenarios.

[0003] Therefore, there is an urgent need for a software radio intelligent voice noise reduction system design method based on field-programmable gate array (FPGA) to achieve high real-time performance, high flexibility, and high performance intelligent voice communication in complex noise environments. Summary of the Invention

[0004] The purpose of this invention is to overcome the shortcomings of the prior art and provide a software radio intelligent voice noise reduction system design method that can meet the requirements of high real-time performance, high flexibility, and high performance intelligent voice noise reduction in complex noise environments.

[0005] To achieve the above-mentioned objectives, the technical solution adopted by the present invention is as follows: This invention provides a software-defined radio intelligent voice noise reduction system design method, including an SDR board and a Xilinx Zynq7020 chip and a digital-to-analog converter / analog-to-digital converter module mounted on the SDR board. The Xilinx Zynq7020 chip runs a preprocessing module, a DCNN hardware accelerator, a DQPSK modulation and demodulation module, and a spectrum monitoring unit. The Xilinx Zynq7020 chip includes a PS terminal and a PL terminal, wherein the PS terminal runs the spectrum monitoring unit, and the PL terminal runs the preprocessing module, the DCNN hardware accelerator, and the DQPSK modulation and demodulation module. This design method includes the following steps: Step 1: The PS terminal runs an embedded operating system. The analog-to-digital converter / digital-to-analog converter module collects the voice radio frequency signal, performs down-conversion and digitization, and its I / Q data is directly sent to the PL terminal through a high-speed interface. The data first passes through the preprocessing module for anti-aliasing filtering and voice frame processing, and then is input to the depth-optimized DCNN hardware accelerator. Step 2: After the DCNN hardware accelerator performs convolution, activation, and pooling operations on the input data, it inputs the data into the DQPSK modulation and demodulation module for differential coding, pulse shaping filtering, and quadrature carrier modulation to generate the baseband signal to be transmitted. The signal is then sent back to the analog-to-digital converter / digital-to-analog converter module for up-conversion and radio frequency transmission, forming a fully hardware-accelerated closed-loop processing pipeline. Step 3: The PS terminal continuously monitors the quality of the received signal through the spectrum monitoring unit. When strong interference is detected in the current working frequency band, causing a decrease in communication quality, the PS terminal automatically generates control commands according to the preset strategy to drive the analog-to-digital converter / digital-to-analog converter module to quickly switch to the backup clean communication frequency band, ensuring the robustness of the communication link.

[0006] Preferably, the PS terminal is the processing part of the Xilinx Zynq7020 chip, which is an ARM-based processor subsystem responsible for global system control, communication protocol management, and dynamic parameter configuration of the RF front-end of the analog-to-digital converter / digital-to-analog converter module; the dynamic parameter settings include frequency band, gain, and bandwidth adjustment.

[0007] Preferably, the analog-to-digital converter / digital-to-analog converter module uses the AD9361 chip.

[0008] Preferably, the signal quality includes signal-to-noise ratio and bit error rate.

[0009] Preferably, the DQPSK modulation and demodulation module performs modulation and demodulation processing on the data. If the bit error rate of the demodulated data exceeds a preset threshold, it indicates that channel interference has affected data transmission, i.e., communication quality has deteriorated.

[0010] Preferably, when the spectrum monitoring unit detects that the signal-to-noise ratio of the received signal is lower than a preset threshold, it indicates that the noise or interference power significantly exceeds the useful signal, i.e., the communication quality deteriorates.

[0011] Preferably, the spectrum monitoring unit converts the time-domain signal into a frequency-domain signal using a Fourier transform formula to monitor the amplitude and phase of the spectrum; the Fourier transform formula is as follows; Where F(jω) is the spectral density function of the signal f(t), t is time, ω is the angular frequency, and j is the imaginary unit. It is a complex exponential function.

[0012] Preferably, the preset strategy is as follows: when the signal-to-noise ratio (SNR) is lower than a preset threshold for 100 consecutive cycles, a frequency band switch will be triggered; each cycle is 5ms, and the SNR is expressed as: SNR: 10*log10 (Ps / Pn) Where Ps represents the average power of the signal, and Pn represents the average power of the noise.

[0013] Preferably, the convolution is performed by sliding a filter through element-wise multiplication and summation on the input data; The activation uses ReLU as the activation function. For any input number x, the ReLU activation function is calculated as follows: , For input features The output feature map after ReLU activation is: , Where c represents the channel index of the input feature map X. Here, 'c' is the input channel index, and 'c' is an index variable used to iterate through all channels of the input feature map. The value of 'c' ranges from 0 to 1. -1; This represents the value of the output feature map at position (i,j) and channel c. This represents the value of the input feature map at the same position and channel; The pooling operation employs average pooling. By setting the pooling window size and stride, the average value within each pooling window is selected as the output to complete the pooling operation. The formula is as follows: , The pooling window size is The step size of the pooling window sliding on the input feature map is... , This represents the height of the pooling window and defines the size of the region where pooling operations are performed vertically. This represents the width of the pooling window and defines the size of the region where pooling operations are performed horizontally. This represents the step size of the pooling operation in the vertical direction, which is the number of pixels that the pooling window advances in the vertical direction each time it moves. This represents the step size of the pooling operation in the horizontal direction, which is the number of pixels the pooling window moves forward in the horizontal direction each time it moves.

[0014] Beneficial effects Compared with the prior art, the beneficial effects achieved by the present invention are as follows: This invention presents a software-defined radio (SDR) intelligent speech denoising system designed using the Xilinx Zynq7020 and AD9361 chips. Through an innovative heterogeneous computing architecture, it achieves significant performance improvements. The Xilinx Zynq7020 chip's integrated ARM processor and FPGA work closely together, fully leveraging the advantages of hardware parallel acceleration. This enables the complex DCNN intelligent denoising model to run efficiently, significantly improving speech clarity and signal-to-noise ratio. Simultaneously, the core characteristic of SDR—flexibility—is fully reflected in this system: the AD9361 RF front-end supports wide-band adjustable frequency, allowing the system to seamlessly adapt to various wireless communication frequency bands and standards; and the FPGA-based reconfigurability supports remote dynamic updates and optimizations of the DCNN denoising algorithm without service interruption, giving the system strong environmental adaptability and continuous evolution capabilities.

[0015] This invention not only boasts excellent intelligent noise reduction performance but also excels in communication reliability and overall energy efficiency. The adoption of DQPSK modulation technology effectively enhances the anti-interference capability and transmission robustness of the baseband signal under complex channel conditions, ensuring the quality of the voice communication link. The heterogeneous architecture of the Xilinx Zynq7020 chip, through reasonable hardware and software task partitioning and collaborative optimization, achieves high-efficiency operation of the entire system, significantly reducing power consumption and making it highly suitable for portable or embedded applications with long battery life requirements. Ultimately, this end-to-end solution, integrating high-performance intelligent noise reduction, full-band software radio flexibility, robust anti-interference communication, and low-power design, provides strong technical support for achieving clear, reliable, and real-time voice communication in harsh noisy environments. Attached Figure Description

[0016] Figure 1 This is a flowchart illustrating the design method of the software-defined radio intelligent voice noise reduction system of the present invention.

[0017] Figure 2 The diagram shows the DCNN intelligent noise reduction model of the software radio intelligent speech noise reduction system design method of the present invention. Detailed Implementation

[0018] To make the objectives, technical solutions, and advantages of the present invention clearer and more complete, the present invention will be further described in detail below with reference to the embodiments. Obviously, the embodiments described below are some embodiments of the present invention, but the scope of protection claimed by the present invention is not limited to the specific embodiments below.

[0019] like Figure 1As shown, a software-defined radio intelligent voice noise reduction system design method includes an SDR board and a Xilinx Zynq7020 chip and a digital-to-analog converter / analog-to-digital converter module mounted on the SDR board. The analog-to-digital converter / digital-to-analog converter module uses the AD9361 chip. The AD9361 chip is a high-performance, highly integrated RF transceiver that integrates the RF front-end and a flexible mixed-signal baseband section. It integrates a frequency synthesizer and has two independent direct-conversion receivers. Each receiver subsystem has independent automatic gain control, DC offset correction, quadrature correction, and digital filtering functions. At the same time, the transmitter of the AD9361 chip adopts a direct-conversion architecture, which can achieve high modulation accuracy and ultra-low noise.

[0020] The Xilinx ZYNQ7020 chip runs a preprocessing module, a DCNN hardware accelerator, a DQPSK modulation / demodulation module, and a spectrum monitoring unit. The Xilinx ZYNQ7020 chip includes a PS (Processing System) terminal and a PL (Processing Subsystem) terminal. The PS terminal runs the spectrum monitoring unit, while the PL terminal runs the preprocessing module, the DCNN hardware accelerator, and the DQPSK modulation / demodulation module. The PS terminal (Processing System (PS) terminal, which is the processing system portion of the Xilinx ZYNQ7020 chip, is an ARM-based processor subsystem) is responsible for global system control, communication protocol management, and dynamic parameter configuration of the AD9361 chip's RF front-end, such as frequency band, gain, and bandwidth adjustment. The Programmable Logic (PL) section (PL) is the programmable logic portion of the Zynq chip, based on an FPGA architecture. It allows users to implement custom hardware logic according to their needs. Through customized hardware logic design, it enables critical high-speed data processing. This architecture ensures seamless high-speed transition from RF signals to baseband processing, laying the foundation for low latency. The design method of this invention fully utilizes the hardware parallel processing capabilities of the FPGA to accelerate computationally intensive intelligent noise reduction algorithms, achieving extremely low latency. Simultaneously, combined with the high reconfigurability of the FPGA, it can flexibly adapt to different wireless communication waveforms and frequency bands (inheriting the flexibility of software-defined radio) and perform hardware-level optimization for specific noise reduction algorithms. like Figure 1 As shown, this design method includes the following steps: Step 1: The PS terminal runs an embedded operating system. The analog-to-digital converter / digital-to-analog converter module acquires the voice radio frequency signal. After down-conversion and digitization, its I / Q data is directly sent to the PL terminal through a high-speed interface. The data first passes through a preprocessing module for anti-aliasing filtering and voice framing (i.e., anti-aliasing filtering is performed through a digital FIR low-pass filter to filter out high-frequency noise interference, followed by noise reduction sampling, and then voice framing. Through a Hamming window sliding window, continuous voice is divided into frames of fixed length, such as 256 points). Finally, it is sent to a depth-optimized DCNN hardware accelerator.

[0021] Step 2: After the DCNN hardware accelerator performs convolution, activation, and pooling operations on the input data, it inputs the data into the DQPSK modulation and demodulation module for differential coding, pulse shaping filtering, and quadrature carrier modulation to generate the baseband signal to be transmitted. The signal is then sent back to the analog-to-digital converter / digital-to-analog converter module for up-conversion and radio frequency transmission, forming a fully hardware-accelerated closed-loop processing pipeline. Specifically, the convolution process is as follows: the convolutional layer processes the input data by applying a filter, and the filter is multiplied and summed element-wise with a local region of the input data to generate a feature map; that is, it is accomplished by sliding a filter on the input data to perform element-wise multiplication and summation. The activation uses ReLU as the activation function. For any input number x, the ReLU activation function is calculated as follows: , For input features The output feature map after ReLU activation is: , Where c represents the channel index of the input feature map X. Here, 'c' is the input channel index, and 'c' is an index variable used to iterate through all channels of the input feature map. The value of 'c' ranges from 0 to 1. -1; This represents the value of the output feature map at position (i,j) and channel c. This represents the value of the input feature map at the same position and channel; The pooling operation employs average pooling. By setting the pooling window size and stride, the average value within each pooling window is selected as the output to complete the pooling operation. The formula is as follows: , The pooling window size is The step size of the pooling window sliding on the input feature map is... , This represents the height of the pooling window and defines the size of the region where pooling operations are performed vertically. This represents the width of the pooling window and defines the size of the region where pooling operations are performed horizontally. This represents the step size of the pooling operation in the vertical direction, which is the number of pixels that the pooling window advances in the vertical direction each time it moves. This represents the step size of the pooling operation in the horizontal direction, which is the number of pixels the pooling window moves forward in the horizontal direction each time it moves.

[0022] like Figure 2 As shown, the DCNN (Deep Convolutional Neural Network) hardware accelerator uses a DCNN intelligent noise reduction model. This model is specifically trained for filtering external environmental noise. An example of the DCNN intelligent noise reduction model's workflow is as follows: The input feature map, with a data shape of [129, 8, 1], contains data information in terms of frequency, time, and number of channels. The input data is zero-padding in a ZeroPadding2D layer to ensure the output size of subsequent convolutional operations matches expectations. Then, it enters a repeated convolutional module for convolution operations, using ReLU activation and BatchNormalization to introduce non-linearity, accelerate training, and improve model stability. Skip connections are designed between the second and eleventh, and fifth and eighth convolutional modules. These skip connections add the output of earlier layers to the output of the current layer before continuing with subsequent convolutions, activations, and normalization operations, mitigating the vanishing gradient problem and enabling the network to learn more complex feature representations. After the output of the last convolutional module, it enters a SpatialDropout2D layer, randomly zeroing out a portion of the entire feature map to reduce overfitting and improve the model's generalization ability. The final convolutional layer transforms the feature map into the final data output with a shape of [129, 1, 1]. Step 3: To cope with the complex and ever-changing electromagnetic environment, this invention designs an intelligent spectrum management mechanism. The PS end continuously monitors the quality of the received signal (such as signal-to-noise ratio and bit error rate) through a spectrum monitoring unit. When the spectrum monitoring unit detects strong interference in the current operating frequency band that causes a decline in communication quality (i.e., when the spectrum monitoring unit detects that the signal-to-noise ratio of the received signal is lower than a preset threshold, indicating that the noise or interference power significantly exceeds the useful signal, or the bit error rate after DQPSK modulation and demodulation exceeds a preset range, indicating that channel interference has affected the reliability of data transmission), it automatically generates control commands according to a preset strategy to drive the AD9361 to quickly switch to a backup clean communication frequency band, effectively avoiding interference sources and ensuring the robustness of the communication link.

[0023] Specifically, the spectrum monitoring unit converts the time-domain signal into a frequency-domain signal using the Fourier transform formula to monitor the amplitude and phase of the spectrum. The Fourier transform formula is as follows; Where F(jω) is the spectral density function of the signal f(t), t is time, ω is the angular frequency, and j is the imaginary unit. It is a complex exponential function.

[0024] The preset strategy is as follows: when the signal-to-noise ratio (SNR) is lower than a preset threshold for 100 consecutive cycles (5ms per cycle), a frequency band switch is triggered; the SNR is expressed as: SNR: 10*log10 (Ps / Pn) Where Ps represents the average power of the signal, and Pn represents the average power of the noise.

[0025] This invention presents a software-defined radio (SDR) intelligent voice noise reduction system design method. Utilizing a SDR intelligent voice noise reduction system based on the Xilinx Zynq7020 and AD9361 chips, it achieves significant performance improvements through an innovative heterogeneous computing architecture. The Zynq7020 chip's integrated ARM processor and FPGA work closely together, fully leveraging the advantages of hardware parallel acceleration. This enables the complex DCNN intelligent noise reduction model to run efficiently, significantly improving voice clarity and signal-to-noise ratio. Simultaneously, the core characteristic of SDR—flexibility—is fully reflected in this system: the AD9361 chip's RF front-end supports wide-band adjustable frequency, allowing the system to seamlessly adapt to various wireless communication frequency bands and standards; while the FPGA's reconfigurability supports remote dynamic updates and optimizations of the DCNN noise reduction algorithm without service interruption, endowing the system with strong environmental adaptability and continuous evolution capabilities. This design, combining SDR spectrum flexibility with heterogeneous computing hardware acceleration, enables the system to continuously provide clear, reliable, and real-time voice communication capabilities in various high-noise and strong-interference real-world application scenarios, with all core processing efficiently completed within a single-chip system.

[0026] This invention not only boasts excellent intelligent noise reduction performance but also excels in communication reliability and overall energy efficiency. The adoption of DQPSK modulation technology effectively enhances the anti-interference capability and transmission robustness of the baseband signal under complex channel conditions, ensuring the quality of the voice communication link. The heterogeneous architecture of the Xilinx Zynq7020 chip, through reasonable hardware and software task partitioning and collaborative optimization, achieves high-efficiency operation of the entire system, significantly reducing power consumption and making it highly suitable for portable or embedded applications with long battery life requirements. Ultimately, this end-to-end solution, integrating high-performance intelligent noise reduction, full-band software radio flexibility, robust anti-interference communication, and low-power design, provides strong technical support for achieving clear, reliable, and real-time voice communication in harsh noisy environments.

[0027] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any changes or substitutions within the technical scope disclosed in this application should be included within the scope of protection of this application. Based on the disclosure and teachings of the above specification, those skilled in the art can also make changes and modifications to the above embodiments. Therefore, this invention is not limited to the specific embodiments disclosed and described above, and some modifications and changes to the invention should also fall within the scope of protection of the claims of this invention. Furthermore, although some specific terms are used in this specification, these terms are only for convenience of explanation and do not constitute any limitation on this invention.

Claims

1. A design method for a software-defined radio intelligent voice noise reduction system, characterized in that, The device includes an SDR board and a Xilinx Zynq7020 chip and a digital-to-analog converter / analog-to-digital converter module mounted on the SDR board. The Xilinx Zynq7020 chip runs a preprocessing module, a DCNN hardware accelerator, a DQPSK modulation and demodulation module, and a spectrum monitoring unit. The Xilinx Zynq7020 chip includes a PS end and a PL end, wherein the PS end runs the spectrum monitoring unit, and the PL end runs the preprocessing module, the DCNN hardware accelerator, and the DQPSK modulation and demodulation module. This design method includes the following steps: Step 1: The PS terminal runs an embedded operating system. The analog-to-digital converter / digital-to-analog converter module collects the voice radio frequency signal, performs down-conversion and digitization, and its I / Q data is directly sent to the PL terminal through a high-speed interface. The data first passes through the preprocessing module for anti-aliasing filtering and voice frame processing, and then is input to the depth-optimized DCNN hardware accelerator. Step 2: After the DCNN hardware accelerator performs convolution, activation, and pooling operations on the input data, it inputs the data into the DQPSK modulation and demodulation module for differential coding, pulse shaping filtering, and quadrature carrier modulation to generate the baseband signal to be transmitted. The signal is then sent back to the analog-to-digital converter / digital-to-analog converter module for up-conversion and radio frequency transmission, forming a fully hardware-accelerated closed-loop processing pipeline. Step 3: The PS terminal continuously monitors the quality of the received signal through the spectrum monitoring unit. When strong interference is detected in the current working frequency band, causing a decrease in communication quality, the PS terminal automatically generates control commands according to the preset strategy to drive the analog-to-digital converter / digital-to-analog converter module to quickly switch to the backup clean communication frequency band to ensure the robustness of the communication link. The PS terminal is the processing section of the Xilinx Zynq7020 chip, which is an ARM-based processor subsystem responsible for global system control, communication protocol management, and dynamic parameter configuration of the RF front-end of the analog-to-digital converter / digital-to-analog converter module; the dynamic parameter settings include frequency band, gain, and bandwidth adjustment; The preset strategy is as follows: when the signal-to-noise ratio (SNR) is lower than a preset threshold for 100 consecutive cycles, a frequency band switch will be triggered; each cycle is 5ms, and the SNR is expressed as: SNR: 10*log10(Ps / Pn) Where Ps represents the average power of the signal and Pn represents the average power of the noise.

2. The design method of a software-defined radio intelligent voice noise reduction system according to claim 1, characterized in that: The analog-to-digital converter / digital-to-analog converter module uses the AD9361 chip.

3. The design method for a software-defined radio intelligent voice noise reduction system according to claim 1, characterized in that: The signal quality includes the signal-to-noise ratio and the bit error rate.

4. The design method of a software-defined radio intelligent voice noise reduction system according to claim 3, characterized in that: The DQPSK modulation and demodulation module performs modulation and demodulation processing on the data. If the bit error rate of the demodulated data exceeds the preset threshold, it indicates that channel interference has affected data transmission, i.e., the communication quality has deteriorated.

5. The design method of a software-defined radio intelligent voice noise reduction system according to claim 4, characterized in that: When the spectrum monitoring unit detects that the signal-to-noise ratio of the received signal is lower than a preset threshold, it indicates that the noise or interference power significantly exceeds the useful signal, i.e., the communication quality deteriorates.

6. The design method of a software-defined radio intelligent voice noise reduction system according to claim 5, characterized in that: The spectrum monitoring unit uses the Fourier transform formula to convert the time-domain signal into a frequency-domain signal to monitor the amplitude and phase of the spectrum. The Fourier transform formula is as follows: Where F(jω) is the spectral density function of the signal f(t), t is time, ω is the angular frequency, and j is the imaginary unit. It is a complex exponential function.

7. The design method of a software-defined radio intelligent voice noise reduction system according to claim 1, characterized in that: The convolution operation is performed by sliding a filter through element-wise multiplication and summation on the input data; the activation uses ReLU as the activation function, and the formula for calculating the ReLU activation function for any input number x is as follows: , For the input feature map The output feature map after ReLU activation is: , Where c represents the channel index of the input feature map X. Here, 'c' is the input channel index, and 'c' is an index variable used to iterate through all channels of the input feature map. The value of 'c' ranges from 0 to 1. -1; This represents the value of the output feature map at position (i,j) and channel c; This represents the value of the input feature map at the same position and channel; The pooling operation employs average pooling. By setting the pooling window size and stride, the average value within each pooling window is selected as the output to complete the pooling operation. The formula is as follows: , The pooling window size is The step size of the pooling window sliding on the input feature map is... , This represents the height of the pooling window and defines the size of the region where pooling operations are performed vertically. This represents the width of the pooling window and defines the size of the region where pooling operations are performed horizontally. This represents the step size of the pooling operation in the vertical direction, which is the number of pixels that the pooling window advances in the vertical direction each time it moves. This represents the step size of the pooling operation in the horizontal direction, which is the number of pixels the pooling window moves forward in the horizontal direction each time it moves.

Citation Information

Patent Citations

  • Voice processing method and device, computer equipment and storage medium

    CN117894306A

  • Speech enhancement method and device using fast fourier convolution

    WO2023182765A1