Audio noise reduction method and system based on Polar code transmission

Through the audio noise reduction method based on Polar code transmission, audio feature vectors are extracted and feature domain noise suppression is performed. Combined with Polar encoding, the problem of insufficient clarity of audio signals in complex environments is solved, and high-quality audio transmission and data reliability are achieved.

CN120498595APending Publication Date: 2025-08-15WUHAN PANSHENG DINGCHENG TECH CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510722604.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-30
Publication Date
2025-08-15

AI Technical Summary

Technical Problem

The prior art is difficult to effectively improve the clarity of audio signals in complex environments, especially in scenarios such as mobile communications and remote meetings. Traditional audio noise reduction methods are difficult to meet the needs of high-quality audio transmission.

Method used

The audio noise reduction method based on Polar code transmission is adopted, and the original audio signal is obtained, the audio feature vector sequence is extracted, and the feature domain acoustic noise suppression is performed, and the anti-interference ability is improved by using Polar code, ultimately achieving high-quality wireless transmission.

Benefits of technology

It improves the audio transmission quality and data transmission reliability, and enhances the noise resistance of audio signals in harsh wireless environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120498595A_ABST
    Figure CN120498595A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of audio noise reduction, and discloses an audio noise reduction method and system based on Polar code transmission, and the method comprises the steps: firstly obtaining an original audio signal, extracting an audio feature vector of the original audio signal, carrying out the suppression of a noise component in a feature domain, and obtaining an audio feature vector sequence after noise reduction; and feature coding and quantization are carried out on the feature sequence, and the feature sequence is converted into a bit stream with a compact structure so as to adapt to the narrow bandwidth and efficient transmission requirements of a wireless channel. In the channel coding stage, the high error correction capability and the polarization characteristic of the Polar code are fully utilized, the characteristic bit stream is effectively coded, and the anti-interference performance and the transmission robustness of the audio data in the severe wireless environment are greatly improved. And finally, mapping the bit stream after Polar coding into a modulation symbol, completing conversion to an analog radio frequency signal, and realizing high-quality wireless noise reduction transmission of an audio signal. Therefore, not only is the audio transmission quality improved, but also the reliability of data transmission is ensured.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of digital information transmission, and more specifically, to an audio noise reduction method and system based on Polar code transmission. Background Art

[0002] With the rapid development of wireless communication technology, audio signal transmission faces numerous challenges, one of the most prominent being noise interference. In particular, in applications such as mobile communications, teleconferencing, and broadcast systems, the demand for high-quality audio transmission is growing. This not only requires low latency and high efficiency during transmission, but also places higher standards on audio signal quality. Traditional audio noise reduction methods often struggle to meet these demands, especially in complex environmental noise environments. Effectively improving audio signal clarity has become a pressing technical challenge.

[0003] In this context, Polar codes, as a new channel coding technology, demonstrate tremendous potential for application in 5G and future communications due to their superior error correction performance and efficient encoding and decoding algorithms. Applying Polar codes to audio signal transmission not only enhances data transmission reliability but also improves the audio signal's noise immunity to a certain extent. However, relying solely on Polar codes cannot fully resolve all issues encountered during audio signal acquisition and transmission, particularly with regard to acoustic noise suppression. Summary of the Invention

[0004] To address the above technical issues, this application is proposed by combining advanced audio processing technology with the advantages of Polar codes. The embodiments of this application propose an audio noise reduction method and system based on Polar code transmission, which aims to solve the problem of audio signal quality degradation caused by noise interference during wireless transmission.

[0005] According to one aspect of the present application, a method and system for audio noise reduction based on Polar code transmission are provided, including: obtaining an original audio signal; extracting an audio feature vector sequence from the original audio signal; performing feature domain acoustic noise suppression on the audio feature vector sequence to obtain a noise-reduced audio feature vector sequence; performing feature encoding and quantization on the noise-reduced audio feature vector sequence to obtain a quantized feature bit stream; performing Polar encoding on the quantized feature bit stream to obtain a Polar-encoded bit stream; and mapping the Polar-encoded bit stream into modulation symbols to obtain an analog radio frequency signal for transmission via a wireless signal.

[0006] In one possible implementation, extracting an audio feature vector sequence from the original audio signal includes: performing frame processing on the original audio signal to obtain an audio frame sequence; performing feature extraction based on short-time Fourier transform on each audio frame in the audio frame sequence to obtain an amplitude spectrum vector sequence and a phase spectrum vector sequence, wherein the amplitude spectrum vector sequence is the audio feature vector sequence.

[0007] In a possible implementation, in the process of performing frame processing on the original audio signal to obtain an audio frame sequence, the frame length is set to 32 ms and the frame shift is set to 16 ms.

[0008] In one possible implementation, feature extraction based on short-time Fourier transform is performed on each audio frame in the audio frame sequence to obtain an amplitude spectrum vector sequence and a phase spectrum vector sequence, including: applying a Hanning window to each audio frame in the audio frame sequence to obtain a windowed audio frame sequence; performing short-time Fourier transform on each windowed audio frame in the windowed audio frame sequence and extracting the amplitude and phase of each frequency component to obtain an amplitude spectrum vector and a phase spectrum vector.

[0009] In one possible implementation, feature domain acoustic noise suppression is performed on the audio feature vector sequence to obtain a noise-reduced audio feature vector sequence, including: inputting the audio feature vector sequence into a trained deep learning model based on a gated recurrent unit to obtain an audio feature time-series context-associated feature vector sequence; and feature decoding is performed on the audio feature time-series context-associated feature vector sequence to obtain a noise-reduced audio feature vector sequence.

[0010] In one possible implementation, the gated recurrent unit-based deep learning model is trained on noisy speech data pairs generated by mixing clean speech and various background noises.

[0011] In one possible implementation, the denoised audio feature vector sequence is feature encoded and quantized to obtain a quantized feature bit stream, including: quantizing each denoised audio feature vector in the denoised audio feature vector sequence to obtain an amplitude spectrum quantization proportional stream; quantizing each phase spectrum vector in the phase spectrum vector sequence to obtain a phase spectrum quantization bit stream; and merging the amplitude spectrum quantization proportional stream and the phase spectrum quantization bit stream to obtain the quantized feature bit stream.

[0012] In one possible implementation, the denoised audio feature vector sequence is feature encoded and quantized to obtain a quantized feature bit stream, including: quantizing each denoised audio feature vector in the denoised audio feature vector sequence to obtain an amplitude spectrum quantization proportional stream; quantizing each phase spectrum vector in the phase spectrum vector sequence to obtain a phase spectrum quantization bit stream; optimizing the amplitude spectrum quantization proportional stream and the phase spectrum quantization bit stream to obtain an optimized amplitude spectrum quantization proportional stream and an optimized phase spectrum quantization bit stream; and merging the optimized amplitude spectrum quantization proportional stream and the optimized phase spectrum quantization bit stream to obtain the quantized feature bit stream.

[0013] In one possible implementation, the amplitude spectrum quantization proportional stream and the phase spectrum quantization bit stream are optimized to obtain an optimized amplitude spectrum quantization proportional stream and an optimized phase spectrum quantization bit stream, including: calculating the transient response modulation factor and the transient prediction offset compensation factor between each amplitude spectrum quantization feature in the amplitude spectrum quantization proportional stream and each phase spectrum quantization feature in the phase spectrum quantization bit stream; performing a transient-steady-state adaptive transition by calculating the integral of the transient prediction offset compensation factor to obtain a transition adjustment factor; based on the quadratic curvature, in combination with the transient prediction offset compensation factor and the transition adjustment factor, updating each amplitude spectrum quantization feature in the amplitude spectrum quantization proportional stream and each phase spectrum quantization feature in the phase spectrum quantization bit stream respectively to obtain the optimized amplitude spectrum quantization proportional stream and the optimized phase spectrum quantization bit stream.

[0014] According to another aspect of the present application, an audio noise reduction system based on Polar code transmission is provided, including: an audio signal acquisition module for acquiring an original audio signal; an audio feature extraction module for extracting an audio feature vector sequence from the original audio signal; an acoustic noise suppression module for performing feature domain acoustic noise suppression on the audio feature vector sequence to obtain a noise-reduced audio feature vector sequence; a feature encoding and quantization module for performing feature encoding and quantization on the noise-reduced audio feature vector sequence to obtain a quantized feature bit stream; a Polar encoding module for performing Polar encoding on the quantized feature bit stream to obtain a Polar-encoded bit stream; and a Polar bit stream modulation module for mapping the Polar-encoded bit stream into modulation symbols to obtain an analog radio frequency signal for transmission via wireless signals.

[0015] Compared with the prior art, the audio noise reduction method and system based on Polar code transmission provided by the present application first obtains the original audio signal and extracts its audio feature vector, and suppresses the noise component in the feature domain to obtain the audio feature vector sequence after noise reduction. The feature sequence is then feature encoded and quantized, and converted into a compact bit stream to adapt to the narrow bandwidth and efficient transmission requirements of the wireless channel. In the channel coding stage, the high error correction capability and polarization characteristics of the Polar code are fully utilized to effectively encode the feature bit stream, which greatly improves the anti-interference and transmission robustness of audio data in harsh wireless environments. Finally, the Polar-coded bit stream is mapped to modulation symbols to complete the conversion to an analog RF signal, thereby realizing high-quality wireless noise reduction transmission of the audio signal. In this way, not only the audio transmission quality is improved, but also the reliability of data transmission is ensured. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] The above and other purposes, features, and advantages of the present application will become more apparent through a more detailed description of the embodiments of the present application in conjunction with the accompanying drawings. The accompanying drawings are intended to provide a further understanding of the embodiments of the present application and constitute a part of the specification. Together with the embodiments of the present application, they are used to explain the present application and do not constitute a limitation of the present application. In the drawings, the same reference numerals generally represent the same components or steps.

[0017] Figure 1 The figure shows a schematic flowchart of an audio noise reduction method based on Polar code transmission according to an embodiment of the present application.

[0018] Figure 2 The figure shows a schematic flowchart of step S2 in the audio noise reduction method based on Polar code transmission according to an embodiment of the present application.

[0019] Figure 3 The figure illustrates a schematic flowchart of step S21 in the audio noise reduction method based on Polar code transmission according to an embodiment of the present application.

[0020] Figure 4 The figure shows a schematic flowchart of step S3 in the audio noise reduction method based on Polar code transmission according to an embodiment of the present application.

[0021] Figure 5 The figure shows a schematic flowchart of step S4 in the audio noise reduction method based on Polar code transmission according to an embodiment of the present application.

[0022] Figure 6 The figure shows a schematic block diagram of an audio noise reduction system based on Polar code transmission according to an embodiment of the present application. DETAILED DESCRIPTION

[0023] Below, the exemplary embodiments according to the present application will be described in detail with reference to the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all the embodiments of the present application, and it should be understood that the present application is not limited to the exemplary embodiments described herein.

[0024] In response to the above technical problems, in the technical solution of this application, an audio noise reduction method based on Polar code transmission is proposed. Figure 1 The figure shows a schematic flowchart of an audio noise reduction method based on Polar code transmission according to an embodiment of the present application.

[0025] like Figure 1 As shown, the audio noise reduction method based on Polar code transmission proposed in this application includes:

[0026] S1, obtain the original audio signal;

[0027] S2, extracting an audio feature vector sequence from the original audio signal;

[0028] S3, performing feature domain acoustic noise suppression on the audio feature vector sequence to obtain a noise-reduced audio feature vector sequence;

[0029] S4, performing feature encoding and quantization on the denoised audio feature vector sequence to obtain a quantized feature bit stream;

[0030] S5, performing Polar encoding on the quantized feature bit stream to obtain a Polar encoded bit stream;

[0031] S6, mapping the Polar-encoded bit stream into modulation symbols to obtain an analog radio frequency signal for wireless signal transmission.

[0032] Specifically, in step S1, an original audio signal is acquired. It should be understood that the original audio signal contains all the information that needs to be preserved, such as speech content and musical details. It may also contain various types of noise, which may come from the ambient background, the device itself, or other interference sources. To achieve high-quality audio transmission, this information must first be accurately captured so that it can be processed in subsequent steps.

[0033] In one specific embodiment, acquiring the original audio signal is accomplished using a microphone array. These microphones are designed to capture sound from different directions, effectively separating the primary voice signal from other ambient noise. Modern microphone technology not only records sound with high fidelity but also, to a certain extent, reduces the impact of mechanical vibration and electromagnetic interference on recording quality. Once the audio signal is captured, it exists in the form of an analog signal, which then needs to be converted to a digital signal through an analog-to-digital converter (ADC).

[0034] In another specific embodiment, in the process of acquiring the original audio signal in a smart speaker or personal assistant device. Here, the micro-microphone inside the device continuously monitors the surrounding acoustic environment and converts the captured sound into an electrical signal. Considering that such devices are usually placed in a home or office environment, the challenge they face is how to accurately recognize the user's voice commands in a complex and changing acoustic background. Therefore, in addition to the basic microphone hardware, a series of advanced algorithms are also used to enhance the quality of signal acquisition. For example, beamforming technology can focus on capturing sounds in a specific direction, while echo cancellation algorithms can remove echoes generated by the sounds played by the device itself.

[0035] Specifically, in step S2, an audio feature vector sequence is extracted from the original audio signal. It should be understood that the original audio signal is converted into a series of discrete digital sampling points after a continuous time sound wave passes through a microphone or other sound pickup device and undergoes analog-to-digital conversion. Although this time domain signal fully reflects the changing process of physical sound, the information it contains is extremely complex. It contains a large amount of structural information related to actual hearing and speech meaning, and it also inevitably contains a large amount of redundant or irrelevant data to the target task. In addition, in real sound scenes, audio signals are often mixed with background noise, environmental interference, etc., and it is difficult to directly distinguish and efficiently model factors such as sound source structure, speech content, and noise energy using the original time domain data. If the original audio is directly denoised and transmitted in the time domain, not only will the data volume be too large, the processing and storage pressure will be huge, and it will also be difficult to effectively extract and utilize signal features, resulting in poor results in noise reduction, encoding, and subsequent recognition. Based on this, in the technical solution of the present application, an audio feature vector sequence is extracted from the original audio signal.

[0036] In one embodiment, Figure 2 As shown, extracting an audio feature vector sequence from the original audio signal includes:

[0037] S21, performing frame processing on the original audio signal to obtain an audio frame sequence;

[0038] S22, performing short-time Fourier transform-based feature extraction on each audio frame in the audio frame sequence to obtain an amplitude spectrum vector sequence and a phase spectrum vector sequence, wherein the amplitude spectrum vector sequence is the audio feature vector sequence.

[0039] In one embodiment, Figure 3 As shown, performing feature extraction based on short-time Fourier transform on each audio frame in the audio frame sequence to obtain an amplitude spectrum vector sequence and a phase spectrum vector sequence, including:

[0040] S221, applying a Hanning window to each audio frame in the audio frame sequence to obtain a windowed audio frame sequence;

[0041] S222: Perform short-time Fourier transform on each windowed audio frame in the windowed audio frame sequence and extract the amplitude and phase of each frequency component to obtain an amplitude spectrum vector and a phase spectrum vector.

[0042] Specifically, first, the continuous audio signal is divided into a series of overlapping short segments. In one embodiment, in the process of framing the original audio signal to obtain an audio frame sequence, the frame length is set to 32ms and the frame shift is set to 16ms. It should be understood that framing can ensure that even subtle changes in rapidly changing audio signals can be captured. Next, a Hanning window function is applied to each frame to reduce spectral leakage, that is, to avoid the energy caused by the truncation effect from spreading to irrelevant frequency components. The Hanning window is a common weighted window function, which has the shape of a cosine square waveform within one period, with both ends close to zero and reaching a maximum value in the middle.

[0043] After the windowing process is completed, the short-time Fourier transform is then performed on the windowed audio frame. The basic idea of STFT is to perform localized analysis of the signal in the time domain, so that the changes in the signal can be observed simultaneously in both time and frequency dimensions. For each audio frame, its discrete Fourier transform is calculated, and the result is a complex array representing the amplitude and phase information of the frame at each frequency. Specifically, the audio signal is regarded as a real-valued function x[n], where n represents the sampling point index. For the k-th audio frame, its corresponding short-time Fourier transform X[k,m] can be calculated by the following formula:

[0044]

[0045] Where N is the frame length (e.g., 512 samples), H is the frame shift (e.g., 256 samples), ω[n] is the Hanning window function, j is the imaginary unit, and m is an integer from 0 to N-1, representing different frequency components. After a short-time Fourier transform, the amplitude and phase spectra of each audio frame are obtained. The amplitude spectrum reflects the intensity distribution of different frequency components and is an important basis for subsequent noise suppression. While the phase spectrum also contains important information, it may not be as critical as the amplitude spectrum in some applications.

[0046] Specifically, in step S3, feature-domain acoustic noise suppression is performed on the audio feature vector sequence to obtain a denoised audio feature vector sequence. It should be understood that the feature domain is the space where signal structure and noise energy distribution are most intuitive, easiest to separate, and most easily modeled. The human ear's ability to separate speech and noise relies heavily on time-frequency distribution and dynamic characteristics. In direct time-domain signals, the characteristics of noise and speech are highly overlapping, making them difficult to distinguish. However, by mapping the time-domain signal to the time-frequency domain through steps such as the short-time Fourier transform, audio feature vectors such as the amplitude spectrum and phase spectrum can clearly demonstrate the differences in energy bands and modulation structures between the speech main stem and background noise. Feature-domain denoising not only allows for the suppression of non-speech noise components at fine spectral resolution, but also reduces the degradation of speech information details. Compared to traditional methods such as simple time-domain filtering, traditional spectral subtraction, and multi-band statistical processing, deep learning in the feature domain can fully utilize spectral information and sequence context dependencies, improving the nonlinear modeling capabilities and generalization robustness of denoising.

[0047] In one embodiment, Figure 4 As shown, performing feature domain acoustic noise suppression on the audio feature vector sequence to obtain a noise-reduced audio feature vector sequence includes:

[0048] S31, inputting the audio feature vector sequence into a trained deep learning model based on a gated recurrent unit to obtain an audio feature temporal context-related feature vector sequence;

[0049] S32: Perform feature decoding on the audio feature temporal context associated feature vector sequence to obtain a noise-reduced audio feature vector sequence.

[0050] Specifically, an audio feature vector sequence is essentially a set of multidimensional vectors with a distinct temporal structure. Speech coherence, syllable transitions, formants, harmonic structure, speaking rate, pauses, and other characteristics are all reflected in these temporal features. Deep learning models based on gated recurrent units (GRUs) can adaptively adjust the historical state of the input sequence and its dependency on current input features through internal update and reset gates. Even in extreme audio environments, they can automatically forget useless noise information, efficiently capturing the dynamic contextual evolution of the main speech components, and thus achieving multi-layer semantic fusion in both spatial and temporal dimensions.

[0051] In a specific embodiment, the deep learning model based on the gated recurrent unit consists of an input layer, several GRU recursive layers and an output layer. The input layer is responsible for receiving a sequence of audio feature vectors. The feature input is then recursively calculated through a series of GRU layers. Each GRU unit contains a reset gate, an update gate and a candidate layer. The reset gate controls the information reset ratio, and the update gate determines the fusion ratio of the old state and the new candidate state. If multiple layers of GRU are stacked, the output of each layer can be used as the input of the next layer of GRU to achieve richer sequence feature abstraction, and the layer depth (for example, 2-4 layers) can be tuned according to the actual scenario and model capacity. And in the last layer, that is, the output layer, an audio feature temporal context-related feature vector sequence is generated.

[0052] Next, the audio feature temporal context-associated feature vector sequence is feature decoded to obtain the denoised audio feature vector sequence. It should be understood that feature decoding can use full connection or convolution decoding to map the GRU hidden state to the target denoised amplitude spectrum feature. Feature decoding can achieve the restoration of the speech backbone structure and the minimization of noise residues. Therefore, a regression output is generally used with the target clean spectrum minimum mean square error (MSE) or weighted L1 / L2 loss as the optimization target. For fine-grained feature recovery, a mask-based decoding strategy can also be used, that is, the network does not directly output the amplitude spectrum, but outputs a multiplier mask (value range [0,1]), and performs element-by-element point multiplication on the input amplitude spectrum frame by frame to achieve spectrum soft gating selection and enhance speech retention and noise suppression capabilities.

[0053] In a specific embodiment, the decoding layer uses a fully connected layer to completely restore the hidden state of each frame to the amplitude spectrum feature points of the same dimension as the input, thereby realizing fixed-point noise reduction and directly outputting the noise-reduced amplitude spectrum itself. In a specific embodiment, the decoding layer can adopt a mask-based strategy to output the result of multiplying the amplitude spectrum by the mask. Specifically, the output mask vector is then multiplied with the original amplitude spectrum frame by frame and point by point to dynamically suppress the noise band and retain only the structured speech signal. These two modes can be adaptively switched according to actual application requirements. For high-noise scenarios, mask-based decoding performs better in subjective perception because it can more finely adjust the relative energy distribution of the target frequency band and the interference signal, reduce traditional noise reduction artifacts such as pumping sound and speech fragmentation, and improve subsequent audio synthesis and feature restoration effects.

[0054] In one embodiment, the deep learning model based on the gated recurrent unit is trained on a pair of noisy speech data generated by mixing clean speech and various background noises. Specifically, a large-scale training data set is first constructed, including a large number of clean speech samples collected or imported from a voice library, and then mixed with various real-world background noises (such as traffic, restaurants, fans, rain, office chats, etc.) in different proportions to generate noisy speech. It is usually necessary to cover all subsets of multiple signal-to-noise ratios, multiple speakers, multiple dialects, different ages and speeds, while ensuring that the test set contains new noise source types and scenes that have not been seen. For each training sample, first frame it, then perform a short-time Fourier transform, and calculate the frame-level amplitude spectrum and phase spectrum. The amplitude spectrum is directly used as the model input, and the target output is the corresponding clean speech amplitude spectrum. The noisy spectrum constitutes the training input feature, the clean spectrum constitutes the label target, and the training is batched into the model in the original order of frame alignment.

[0055] During the model training phase, the loss function is set to the mean square error between the predicted output spectrum and the target spectrum; if a mask-based solution is adopted, it can also be the cross entropy or weighted MSE between the predicted mask and the ideal amplitude spectrum mask. In order to improve the final subjective listening performance and avoid the superposition of sound quality artifacts, in other embodiments of the present application, the training process can also be combined with perceptual loss indicators (such as SDR signal-to-noise distortion rate, PESQ, STOI, etc.) for auxiliary optimization. The optimization algorithm includes adaptive learning rate strategies such as Adam and RMSProp, which continuously update the weights of the entire network in small batches (batch size is usually 4-64). In order to prevent overfitting and improve generalization ability, it can also be equipped with dropout, layer normalization, data enhancement (time shift, channel reverberation, gain perturbation, etc.), and adopt strategies such as Early-stopping or dynamic learning rate decay to monitor and guide model performance.

[0056] Specifically, in step S4, the denoised audio feature vector sequence is feature encoded and quantized to obtain a quantized feature bitstream. It should be understood that the denoised audio feature vector sequence is still expressed in the form of continuous real numbers. Although this type of data removes most noise components and more effectively and effectively expresses the original speech content, it still belongs to a high-dimensional real number space. On the one hand, real-number feature vectors describe information with extreme precision, containing many subtle linear and nonlinear changes in the signal; on the other hand, the storage, processing, and communication costs of such raw features are also extremely high, making it difficult to adapt to practical requirements such as bandwidth limitations of wireless channels and limited computing resources of edge inference terminals. In order to achieve efficient transmission of high-dimensional real-number features while effectively preserving the main structure and perceptual quality of the signal, they must be further encoded into digital, discretized low-bitstream data. In other words, the denoised audio feature vector sequence is feature encoded and quantized to obtain a quantized feature bitstream.

[0057] In one embodiment, Figure 5 As shown, feature encoding and quantization are performed on the denoised audio feature vector sequence to obtain a quantized feature bit stream, including:

[0058] S41, quantizing each denoised audio feature vector in the denoised audio feature vector sequence to obtain an amplitude spectrum quantization ratio stream;

[0059] S42, quantizing each phase spectrum vector in the phase spectrum vector sequence to obtain a phase spectrum quantization bit stream;

[0060] S43: Merge the amplitude spectrum quantization ratio stream and the phase spectrum quantization bit stream to obtain the quantized feature bit stream.

[0061] Specifically, feature quantization revolves around two dimensions of the signal: the quantization of the amplitude spectrum vector after noise reduction, and the quantization of the phase spectrum vector. The amplitude spectrum and phase spectrum are the primary characteristic expressions of audio signals in the short-time Fourier transform (STFT) domain. The amplitude spectrum carries the main energy distribution and speech structure information, while the phase spectrum is particularly critical for the precise restoration of time-domain signals. The actual data processing process typically involves quantizing the amplitude spectrum and phase spectrum of each frame after noise reduction, mapping their continuous values to discrete levels (i.e., segmentation and binning), and gradually encoding them into fixed-length or variable-length bit data. Finally, the amplitude and phase sub-streams are integrated to generate an overall efficient bit feature stream, preparing for subsequent polar code channel coding and wireless modulation.

[0062] In a specific embodiment, the amplitude spectrum quantization can adopt strategies such as linear quantization or logarithmic quantization. Linear quantization is to evenly divide the value range of the amplitude spectrum into several levels. For example, 8-bit linear quantization is to divide the value range of each amplitude spectrum feature into 256 parts, each of which is represented by an 8-bit unsigned integer. In this way, the original floating-point or fixed-point amplitude spectrum data points are directly mapped to one or a group of integers, and the storage requirements of a single data point are greatly reduced. The advantage of linear quantization is that it is simple to implement, the quantization error is uniformly distributed overall, and it is convenient for hardware acceleration. However, when the spectrum value distribution is concentrated in a small range and a small number of points occasionally increase, the performance of linear quantization will be limited. Therefore, in another embodiment of the present application, a logarithmic quantization method (such as logarithmic compression) can also be adopted, and the amplitude spectrum is first log transformed and then linearly quantized. Logarithmic quantization in dB units not only conforms to the nonlinear relationship between speech energy and human ear perception, but also effectively protects weak signals in low-energy frequency bands, improving the subjective quality and restoration resolution of the overall signal. For example, the amplitude spectrum can be normalized to between 0 and 1, then its logarithm is taken. The logarithmic features are then quantized using N-bit unsigned integers. Combining logarithmic quantization with segmented coding can further reduce redundant information in high-energy regions, thereby improving bit efficiency.

[0063] The quantization processing of the phase spectrum vector is more special than that of the amplitude spectrum. The phase itself is a periodic feature, usually distributed between -π and π. For its linear quantization, the interval can be divided into several equal parts (such as 16 levels, 32 levels, 64 levels, etc.), and each level is represented by a finite bit code. Since phase information plays a decisive role in restoring the waveform, but has limited impact on the subjective perception of speech, the number of phase quantization levels can be appropriately reduced in the actual system to save bandwidth and computing resources. In some scenarios where the robustness of the reconstructed waveform is not high, phase approximation or finite-level sampling can even be used to greatly reduce the bit requirement. The specific processing flow is to normalize the phase spectrum of each frame, fold the interval, map it to a fixed-length quantization code, and then convert it into an integer stream. With the synchronization clock and frame number, it is restored sequentially at the decoding end.

[0064] Throughout the quantization and encoding process, issues such as quantization error accumulation, distortion floor, and bandwidth utilization must be fully considered. Generally speaking, the quantization of amplitude and phase spectra requires a trade-off between the subjective quality of actual business applications and system capacity. A quantization accuracy of 3 to 8 bits per feature point is generally adopted, and the data volume of a single-frame feature stream is significantly smaller than the original features. For narrowband wireless transmission with extremely limited bandwidth, the quantization accuracy and encoding strategy can be dynamically adjusted to prioritize the restoration of key signal components, thereby achieving end-to-end progressive optimization of the communication system. Quantized data is typically encapsulated in a frame / packet format. Each frame corresponds to the complete audio features of a time period and is stored in an ordered buffer queue, providing near-real-time input for subsequent processes such as Polar encoding.

[0065] In addition to basic uniform quantization, in order to further compress data redundancy and improve bit efficiency, in other examples of this application, auxiliary algorithms such as entropy coding, differential coding, vector quantization, and predictive coding can also be used. For example, entropy coding (such as Huffman coding, arithmetic coding, etc.) can dynamically adjust the code length based on the statistical distribution of feature points, with common values represented by short codes and rare values by long codes, effectively improving the average transmission efficiency. Vector quantization aggregates a set of continuous features into vectors, and uses a preset codebook or self-learning codebook (such as the LBG algorithm) to directly map high-dimensional features into a small number of category labels, each of which is represented by a fixed-length bit. This strategy is particularly suitable for spectral feature compression of speech signals, and can save more codeword length than scalar quantization at the same quality. Differential coding and predictive coding encode the differences between adjacent frames based on the temporal correlation of the feature sequence, rather than directly encoding the absolute value. Such technologies often achieve good compression gains in speech or music signals.

[0066] The quantized amplitude spectrum quantization bit stream and phase spectrum quantization bit stream need to be merged and output eventually. Specifically, the quantized amplitude spectrum bit stream and all quantized phase spectrum bit streams of each frame of audio in unit time are merged together to form the final quantized feature bit stream of the frame or a series of continuous frames. In a specific embodiment, it can be implemented by adding frame header tags, alternating multiplexing, etc., to ensure efficient recovery and separation before subsequent Polar channel encoder processing, which is convenient for the design of encoding and decoding algorithms. The use of packet length check, multi-frame synchronization and other protocols for data streams can greatly reduce the risk of packet loss and bit misalignment, and improve the data robustness and self-recovery capability under wireless channels.

[0067] Here, considering the amplitude spectrum quantization characteristics a in the amplitude spectrum quantization ratio flow i is obtained by quantizing the amplitude spectrum vector sequence after noise reduction, and each phase spectrum quantization feature b in the phase spectrum quantization bit stream is obtained by quantizing the amplitude spectrum vector sequence after noise reduction, and the phase spectrum quantization feature b in the phase spectrum quantization bit stream is obtained by quantizing the amplitude spectrum vector sequence after noise reduction. iIt is obtained by directly quantizing the phase spectrum vector sequence, so a i and b i Obviously there is a certain spectrum sequence connection analysis bias, that is, if we want to take into account a i and b i In terms of transient and steady-state correlation from the perspective of overall spectrum analysis, the amplitude spectrum quantization ratio stream and the phase spectrum quantization bit stream are preferably optimized before being merged.

[0068] Based on this, in another embodiment, the denoised audio feature vector sequence is feature encoded and quantized to obtain a quantized feature bit stream, including: quantizing each denoised audio feature vector in the denoised audio feature vector sequence to obtain an amplitude spectrum quantization proportional stream; quantizing each phase spectrum vector in the phase spectrum vector sequence to obtain a phase spectrum quantization bit stream; optimizing the amplitude spectrum quantization proportional stream and the phase spectrum quantization bit stream to obtain an optimized amplitude spectrum quantization proportional stream and an optimized phase spectrum quantization bit stream; merging the optimized amplitude spectrum quantization proportional stream and the optimized phase spectrum quantization bit stream to obtain the quantized feature bit stream.

[0069] Specifically, optimizing the amplitude spectrum quantization proportional stream and the phase spectrum quantization bit stream to obtain an optimized amplitude spectrum quantization proportional stream and an optimized phase spectrum quantization bit stream includes:

[0070] First, based on the transient correlation perspective, through a i and b i The transient adaptive response modulation between the two is used to suppress the nonlinear load transient prediction offset caused by the difference in nonlinear noise reduction processing, that is, the transient response modulation factor and the transient prediction offset compensation factor between each amplitude spectrum quantization feature in the amplitude spectrum quantization scale stream and each phase spectrum quantization feature in the phase spectrum quantization bit stream are calculated, which can be expressed as:

[0071]

[0072] s i =lnr i

[0073] Among them, a i represents each amplitude spectrum quantization feature in the amplitude spectrum quantization ratio stream, b i represents each phase spectrum quantization feature in the phase spectrum quantization bit stream, e represents a natural constant, ln represents a natural logarithm, r i represents the transient response modulation factor, s i Represents the transient prediction offset compensation factor.

[0074] That is, first a iand b i Perform transient linkage coordination to construct the transient response modulation factor r i , then, a is compensated by dynamically adapting the natural constant exponent of the transient response to a logarithmic function with the natural constant as the base. i and b i The transient load prediction offset between

[0075] At the same time, based on the steady-state correlation of amplitude and phase under global conditions, the transient-steady-state adaptive transition is performed by calculating the integral of the transient prediction offset compensation factor to obtain the transition adjustment factor, that is:

[0076] q i =∫s i dr i =r i lnr i ―r i

[0077] That is, q i As a transition adjustment factor, it can be used to quantify the local-global tolerance of the correlated transition of parameters from local transient to global stable.

[0078] Finally, based on the quadratic curvature, combined with the transient prediction offset compensation factor and the transition adjustment factor, each amplitude spectrum quantization feature in the amplitude spectrum quantization proportional stream and each phase spectrum quantization feature in the phase spectrum quantization bit stream are updated respectively to obtain the optimized amplitude spectrum quantization proportional stream and the optimized phase spectrum quantization bit stream, that is:

[0079] a ′i =s i 2 a i +q i a i

[0080] b ′i =s i 2 b i +q i b i

[0081] Among them, a ′i represents the optimized amplitude spectrum quantization features in the optimized amplitude spectrum quantization ratio stream, b ′i Represents each optimized phase spectrum quantization feature in the optimized phase spectrum quantization bit stream.

[0082] Based on the transient and steady-state correlation from the perspective of overall spectrum analysis, effective connection admissible synergy can be established by suppressing the spectrum sequence connection analysis offset, thereby realizing the amplitude spectrum quantization feature a iand phase spectrum quantization feature b i Correlation adaptation and unification based on transient spectrum analysis and steady-state spectrum analysis.

[0083] Specifically, in step S5, the quantized feature bit stream is polar-encoded to obtain a polar-encoded bit stream. It should be understood that, faced with complex interference, channel noise, multipath fading, and occasional burst errors in real-world environments, relying solely on data quantization and digital processing still cannot fundamentally guarantee high-quality recovery of the original content at the receiving end. Especially in situations where bandwidth is limited, baud rate is restricted, and transmission distances are long, signals are highly susceptible to various communication noises and interferences during airborne propagation, resulting in bit errors and even the continuous generation of large numbers of erroneous bits. If the quantized feature bit stream is not efficiently channel-encoded, minor transmission distortions and errors will directly lead to the failure of the audio noise reduction process, and in severe cases, even affect the integrity of the entire audio communication task. Therefore, in situations where high-reliability wireless audio transmission is required, advanced channel coding mechanisms are further introduced. That is, the quantized feature bit stream is polar-encoded to obtain a polar-encoded bit stream.

[0084] Compared to conventional general-purpose channel coding technologies such as convolutional codes, BCH codes, and LDPC codes, Polar codes not only have clear polarization theory support and excellent fault tolerance, but also boast encoding and decoding complexity of O(NlogN), making them ideally suited for modern high-efficiency, low-power data communication scenarios. Furthermore, Polar codes have been adopted as a fundamental control channel coding technology in next-generation wireless communication standards such as 3GPP 5G NR, demonstrating strong engineering feasibility. For audio feature bitstreams, Polar codes significantly improve channel robustness with minimal code length, making them one of the optimal channel coding options for wireless transmission of quantized bitstreams.

[0085] Specifically, the Polar coded bitstream is highly resistant to both random and burst errors. Even in complex wireless channel environments, it can minimize the net bit error rate during transmission, significantly improving the availability and fidelity of audio content reaching the end user. Furthermore, the code rate, code length, and polarization structure of Polar codes are flexibly configurable on demand, supporting adaptive rate adjustment at any length. This makes it ideally suited for modern audio systems that face dynamic changes in wireless bandwidth and demand real-time, adaptive communication technologies. Furthermore, Polar codes' unique polarization construction allows some bits in the same data batch to reside on highly reliable channels while others reside on less reliable channels. By freezing unreliable bits and transmitting only the true information on the polarized, reliable bits, Polar codes naturally separate error correction from the core information, significantly improving decoding efficiency and error recovery. In particular, in audio systems that utilize frame-by-frame transmission, packet switching, and end-to-end acknowledgment, Polar codes can be further combined with various retransmission and interleaving mechanisms to achieve new levels of system-wide data integrity and latency optimization, enabling practical large-scale deployment of engineering systems.

[0086] Specifically, the quantized characteristic bit stream must first be grouped into codewords. Based on the target channel environment, required error performance, system bandwidth upper limit, and real-time delay requirements, appropriate Polar code parameters are selected, including the code length N (usually selected as an integer power of 2, such as 1024, 2048, or 4096), the number of information bits K (actual data bit length), the location of the frozen bits, and the specific polarization order. In audio communication scenarios, due to many factors such as real-time communication and channel changes, and frame length limitations, the Polar parameters need to be adaptively adjusted according to the actual environment, and switching can be completed in real time at the frame / packet level. The core idea of Polar coding is to decompose the physical channel into a group of equivalent sub-channels through a series of channel polarization operations, and then transmit real data information on the channel, filling in the predetermined frozen bits in the failed channel.

[0087] In one embodiment, performing Polar encoding on the quantized feature bitstream to obtain a Polar-encoded bitstream includes the following steps:

[0088] First, a bit channel sorting is generated based on a pre-designed polarization sequence, the K sub-channels with the highest reliability are selected, and the K bits of information are sequentially filled in.

[0089] Then, for the NK unreliable bits, all are filled with 0 or other agreed freezing conditions (such as CRC or error correction flag, etc.);

[0090] Then, it enters the Polar code encoder (actually a series of binary operation networks) to form a codeword of length N.

[0091] The encoding operation is essentially a recursive matrix multiplication performed through XNOR gate-level arrangements (Kronecker products can be used to construct the polarization generator matrix G_N). This entire process has extremely low latency in a typical Polar hardware encoder or software algorithm implementation, making it easy to pipeline and parallelize the process.

[0092] In a specific embodiment, using a simple example with N = 8, assume the input information bit stream is K = 5 bits. After polarization sorting, subchannels 0, 2, 3, 5, and 6 are selected as information bits, and the remaining bits are frozen. Once the input information is entered, the recursive layered polarization matrix encoding completes the codeword expansion output through an exclusive-OR operation.

[0093] The core matrix of Polar code can be expressed as in is the fundamental matrix, represents the Kronecker product, and n is the number of polarization levels. The encoding can be directly parallelized using polarization circuit networks and efficiently recursively implemented through cache arrays.

[0094] In this specific embodiment, the polar encoder takes as input a K-length information bit sequence (i.e., a quantized feature bit stream) and outputs a codeword sequence of length N. These codewords contain not only the actual data but also polarization parity information (frozen bits), providing strong error correction and self-recovery capabilities. All encoded codewords are then output as a complete, structured polar coded bit stream, which then enters signal modulation.

[0095] Specifically, in step S6, the Polar-encoded bit stream is mapped to modulation symbols to obtain an analog RF signal for wireless signal transmission. It should be understood that there is an interface gap between digital communications and the real wireless physical world. Only by using modulation technology to convert the digital bit stream into an analog signal suitable for propagation over a RF medium can true long-distance wireless transmission of information be achieved. Furthermore, the implementation of this process directly determines the system's bandwidth utilization efficiency, robustness, real-time performance, and anti-interference capabilities, and has a decisive impact on the end user's voice and audio experience and system performance.

[0096] Specifically, digital signals themselves are only discrete values in a logical sense, but electromagnetic waves and radio frequency channels can only carry continuous changes in voltage / current. Wireless signal transmission in audio applications requires converting the encoded digital bit stream into a continuous electromagnetic wave—an analog radio frequency signal—that can be transmitted through an antenna and propagated through the air. The modulation process is used to achieve reliable mapping of bits to physical signals and efficiently utilize limited physical channel bandwidth and power resources. Without modulation, even the highly error-resistant Polar encoded bit stream would not be able to cover all aspects of wireless transmission. Modulation technology is the core backbone of wireless audio communications, acting as a bridge for transcoding between data and physical media.

[0097] At the same time, modulation, as the intermediate layer connecting channel coding and physical signals, its type, architecture, and mapping strategy all affect the tuning performance, perceived quality, and energy efficiency of the entire audio wireless system. Modern wireless communications employ a variety of digital modulation schemes, including binary bit shift keying (BPSK), quadrature amplitude modulation (QAM), and orthogonal frequency division multiplexing (OFDM). Each modulation scheme arranges the relationship between the polar code output bit stream and the RF signal in a different way. For example, BPSK maps a single bit to positive and negative carrier phases, QPSK maps two bits together to four signal phases, and 16QAM / 64QAM / 256QAM simultaneously maps four to eight bits to different points on the two-dimensional I / Q plane. With the ever-increasing demand for bandwidth and data rates, choosing higher-order, more robust modulation schemes is particularly important.

[0098] Specifically, after obtaining the Polar coded bit stream, an appropriate modulation scheme (such as BPSK, QPSK, or QAM) is first selected based on the actual system and channel conditions. The mapping process involves bit grouping, bit-to-symbol mapping, pulse shaping, upconversion, and signal transmission. Bit grouping involves dividing the Polar coded bit stream according to the number of bits per symbol used by the modulation scheme. For example, QPSK carries 2 bits per symbol, while 16QAM carries 4 bits per symbol.

[0099] Therefore, in QPSK, the bit stream is split into groups of two, each mapped to four different phases. In 16QAM, four bits are grouped together, mapped to 16 complex points (that is, 16 signal patterns consisting of four amplitude combinations of the I (in-phase) and Q (quadrature) components). This mapping process is usually accomplished by table lookup, where each group of bits is searched in a pre-defined constellation table to obtain the corresponding complex coordinates, or modulation symbols.

[0100] After the entire bit stream is grouped and converted into modulation symbols through a table lookup, a symbol sequence is generated. These symbols form a complex sequence representing the physical signal to be transmitted on the channel at a specific sampling time. Pulse shaping (such as root-raised cosine filtering) is then performed to prevent band leakage, reduce intersymbol interference, and effectively control the signal's spectral width. After signal shaping, digital-to-analog conversion is performed before entering the RF section. Up-conversion then occurs, bringing the low-frequency baseband signal to the carrier frequency specified by the wireless channel. After power amplification and RF filtering, the signal is ultimately transmitted by the antenna, generating a simulated RF signal that can be received in a real-world environment.

[0101] In summary, the audio noise reduction method and system based on Polar code transmission provided by the present application first obtains the original audio signal and extracts its audio feature vector, and suppresses the noise component in the feature domain to obtain the audio feature vector sequence after noise reduction. The feature sequence is then feature encoded and quantized, and converted into a compact bit stream to adapt to the narrow bandwidth and efficient transmission requirements of the wireless channel. In the channel coding stage, the high error correction capability and polarization characteristics of the Polar code are fully utilized to effectively encode the feature bit stream, which greatly improves the anti-interference and transmission robustness of audio data in harsh wireless environments. Finally, the Polar encoded bit stream is mapped to modulation symbols to complete the conversion to analog RF signals, thereby realizing high-quality wireless noise reduction transmission of audio signals. In this way, not only the audio transmission quality is improved, but also the reliability of data transmission is ensured.

[0102] The present application also provides an audio noise reduction system based on Polar code transmission. Figure 6 FIG shows a schematic block diagram of an audio noise reduction system based on Polar code transmission according to an embodiment of the present application, as shown in FIG. Figure 6 As shown, the audio noise reduction system 600 based on Polar code transmission includes:

[0103] The audio signal acquisition module 610 is used to acquire the original audio signal;

[0104] An audio feature extraction module 620 is configured to extract an audio feature vector sequence from the original audio signal;

[0105] an acoustic noise suppression module 630, configured to perform feature-domain acoustic noise suppression on the audio feature vector sequence to obtain a noise-reduced audio feature vector sequence;

[0106] A feature coding and quantization module 640 is configured to perform feature coding and quantization on the denoised audio feature vector sequence to obtain a quantized feature bit stream;

[0107] Polar encoding module 650, configured to perform Polar encoding on the quantized feature bit stream to obtain a Polar encoded bit stream;

[0108] The Polar bit stream modulation module 660 is configured to map the Polar-encoded bit stream into modulation symbols to obtain an analog radio frequency signal for wireless signal transmission.

[0109] The basic principles of this application have been described above in conjunction with specific embodiments. However, it should be noted that the advantages, strengths, and effects mentioned in this application are merely illustrative and not restrictive, and it should not be assumed that these advantages, strengths, and effects are required of each embodiment of this application. In addition, the specific details disclosed above are merely illustrative and facilitating understanding, and are not restrictive. The above details do not limit this application to being implemented using the above specific details.

[0110] The above description of the disclosed aspects is provided to enable any person skilled in the art to make or use the present application. Various modifications to these aspects will be readily apparent to those skilled in the art, and the general principles defined herein may be applied to other aspects without departing from the scope of the present application. Therefore, the present application is not intended to be limited to the aspects shown herein, but rather to be accorded the widest scope consistent with the principles and novel features disclosed herein.

[0111] The above description has been provided for the purpose of illustration and description. Furthermore, this description is not intended to limit the embodiments of the present application to the forms disclosed herein. Although a number of example aspects and embodiments have been discussed above, those skilled in the art will recognize certain variations, modifications, alterations, additions, and sub-combinations thereof.

Claims

1. An audio noise reduction method based on Polar code transmission, characterized in that: include: Get the original audio signal; Extracting an audio feature vector sequence from the original audio signal; Performing feature-domain acoustic noise suppression on the audio feature vector sequence to obtain a noise-reduced audio feature vector sequence; Performing feature encoding and quantization on the denoised audio feature vector sequence to obtain a quantized feature bit stream; Performing Polar encoding on the quantized feature bit stream to obtain a Polar encoded bit stream; The polar-encoded bit stream is mapped into modulation symbols to obtain an analog radio frequency signal for transmission via a wireless signal.

2. The audio noise reduction method based on Polar code transmission according to claim 1, characterized in that: Extracting an audio feature vector sequence from the original audio signal, comprising: Performing frame processing on the original audio signal to obtain an audio frame sequence; Perform feature extraction based on short-time Fourier transform on each audio frame in the audio frame sequence to obtain an amplitude spectrum vector sequence and a phase spectrum vector sequence, wherein the amplitude spectrum vector sequence is the audio feature vector sequence.

3. The audio noise reduction method based on Polar code transmission according to claim 2, characterized in that: In the process of performing frame processing on the original audio signal to obtain an audio frame sequence, the frame length is set to 32 ms and the frame shift is set to 16 ms.

4. The audio noise reduction method based on Polar code transmission according to claim 2, characterized in that: Performing feature extraction based on short-time Fourier transform on each audio frame in the audio frame sequence to obtain an amplitude spectrum vector sequence and a phase spectrum vector sequence, including: Applying a Hanning window to each audio frame in the audio frame sequence to obtain a windowed audio frame sequence; Performing short-time Fourier transform on each windowed audio frame in the windowed audio frame sequence and extracting the amplitude and phase of each frequency component to obtain an amplitude spectrum vector and a phase spectrum vector.

5. The audio noise reduction method based on Polar code transmission according to claim 1, characterized in that: Performing feature-domain acoustic noise suppression on the audio feature vector sequence to obtain a noise-reduced audio feature vector sequence, comprising: Inputting the audio feature vector sequence into a trained deep learning model based on a gated recurrent unit to obtain an audio feature temporal context-related feature vector sequence; Feature decoding is performed on the audio feature temporal context associated feature vector sequence to obtain a noise-reduced audio feature vector sequence.

6. The audio noise reduction method based on Polar code transmission according to claim 5, characterized in that: The gated recurrent unit-based deep learning model is trained on noisy speech data pairs generated by mixing clean speech and various background noises.

7. The audio noise reduction method based on Polar code transmission according to claim 2, characterized in that: Feature encoding and quantization are performed on the noise-reduced audio feature vector sequence to obtain a quantized feature bit stream, including: quantizing each denoised audio feature vector in the denoised audio feature vector sequence to obtain an amplitude spectrum quantization ratio stream; quantizing each phase spectrum vector in the phase spectrum vector sequence to obtain a phase spectrum quantization bit stream; The amplitude spectrum quantization scale stream and the phase spectrum quantization bit stream are combined to obtain the quantized feature bit stream.

8. The audio noise reduction method based on Polar code transmission according to claim 2, characterized in that: Feature encoding and quantization are performed on the noise-reduced audio feature vector sequence to obtain a quantized feature bit stream, including: quantizing each denoised audio feature vector in the denoised audio feature vector sequence to obtain an amplitude spectrum quantization ratio stream; quantizing each phase spectrum vector in the phase spectrum vector sequence to obtain a phase spectrum quantization bit stream; Optimizing the amplitude spectrum quantization proportional stream and the phase spectrum quantization bit stream to obtain an optimized amplitude spectrum quantization proportional stream and an optimized phase spectrum quantization bit stream; The optimized amplitude spectrum quantization scale stream and the optimized phase spectrum quantization bit stream are combined to obtain the quantized feature bit stream.

9. The audio noise reduction method based on Polar code transmission according to claim 8, characterized in that: Optimizing the amplitude spectrum quantization proportional stream and the phase spectrum quantization bit stream to obtain an optimized amplitude spectrum quantization proportional stream and an optimized phase spectrum quantization bit stream, including: Calculating a transient response modulation factor and a transient prediction offset compensation factor between each amplitude spectrum quantization feature in the amplitude spectrum quantization scale stream and each phase spectrum quantization feature in the phase spectrum quantization bit stream; Performing transient-steady-state adaptive transition by calculating the integral of the transient prediction offset compensation factor to obtain a transition adjustment factor; Based on the quadratic curvature, combined with the transient prediction offset compensation factor and the transition adjustment factor, the amplitude spectrum quantization features in the amplitude spectrum quantization proportional stream and the phase spectrum quantization features in the phase spectrum quantization bit stream are updated respectively to obtain the optimized amplitude spectrum quantization proportional stream and the optimized phase spectrum quantization bit stream.

10. An audio noise reduction method based on Polar code transmission, characterized in that: include: An audio signal acquisition module is used to acquire the original audio signal; An audio feature extraction module, configured to extract an audio feature vector sequence from the original audio signal; an acoustic noise suppression module, configured to perform feature-domain acoustic noise suppression on the audio feature vector sequence to obtain a noise-reduced audio feature vector sequence; A feature coding and quantization module, configured to perform feature coding and quantization on the denoised audio feature vector sequence to obtain a quantized feature bit stream; A Polar encoding module, configured to perform Polar encoding on the quantized feature bit stream to obtain a Polar encoded bit stream; The Polar bit stream modulation module is used to map the Polar encoded bit stream into modulation symbols to obtain an analog radio frequency signal for transmission via a wireless signal.

Citation Information

Cited By

  • Selection method and system of R16QAM structured bit selector with high spectral efficiency

    CN120934695A