Analog-digital hybrid SRAM memory audio noise reduction method and system
Through the analog-to-digital hybrid SRAM storage and computing audio noise reduction method, the SRAM storage and computing chip is used for differential sampling quantization and frequency domain feature extraction to identify and suppress noise, solving the problems of high power consumption and cost in traditional methods and achieving low-power, high-efficiency audio noise reduction effects.
Patent Information
- Application Number
- CN202511070763.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-31
- Publication Date
- 2025-09-26
- Estimated Expiration
- 2045-07-31
AI Technical Summary
Traditional digital signal processing methods rely on high-precision analog-to-digital converters and high-performance processors, resulting in a significant increase in power consumption and hardware costs, making it difficult to meet the low power consumption and small size requirements of edge devices.
A hybrid analog-digital SRAM storage and computing audio noise reduction method is adopted. A multi-bit-width digital audio code stream is generated through differential sampling and quantization processing. The SRAM storage and computing chip is used for mixed mode writing, frequency domain feature extraction and noise identification are performed, and a time-varying noise suppression coefficient is generated to finally restore the noise-reduced audio signal.
It achieves low-power, high-real-time audio noise reduction, improves audio restoration quality and listening experience, and reduces hardware costs.
Smart Images

Figure CN120708645A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of audio noise reduction, and in particular to an analog-to-digital hybrid SRAM storage and calculation audio noise reduction method and system. Background Art
[0002] In recent years, with the rapid development of artificial intelligence and the Internet of Things (IoT) technologies, audio signal processing has played an increasingly important role in fields such as intelligent voice assistants, wearable devices, and communication systems. Traditional digital signal processing methods typically rely on high-precision analog-to-digital converters (ADCs) and high-performance processors, resulting in a significant increase in power consumption and hardware costs, making it difficult to meet the low power and small size requirements of edge devices. In addition, existing noise reduction algorithms are mostly based on complex software models, such as deep learning networks, whose computationally intensive nature further exacerbates real-time and energy efficiency bottlenecks. Therefore, there is an urgent need for a method that can efficiently implement audio noise reduction at the hardware level to overcome the energy efficiency and real-time limitations of traditional architectures. Summary of the Invention
[0003] The main purpose of the present invention is to provide an analog-to-digital hybrid SRAM storage and calculation audio noise reduction method, which solves the technical problem that traditional digital signal processing methods usually rely on high-precision analog-to-digital converters and high-performance processors, resulting in a significant increase in power consumption and hardware costs.
[0004] To achieve the above object, the present invention provides an analog-digital hybrid SRAM storage and calculation audio noise reduction method, comprising the following steps: Perform differential sampling and quantization processing on the input audio signal to obtain a multi-bit width digital audio code stream; Inputting the multi-bit width digital audio code stream into a preset SRAM storage chip to perform a mixed mode write operation to obtain an analog-digital hybrid storage matrix; Performing frequency domain feature extraction on the analog-digital hybrid storage matrix to obtain frequency domain extraction features; Performing noise identification and suppression on the audio signal based on the frequency domain extraction features to obtain a time-varying noise suppression coefficient; An audio signal restoration process is performed on the time-varying noise suppression coefficient to obtain a noise-reduced output audio signal.
[0005] Preferably, performing differential sampling and quantization processing on the input audio signal to obtain a multi-bit width digital audio code stream includes: Performing time-domain differential sampling processing on the audio signal to obtain an audio sampling sequence, and extracting an audio amplitude reference of the audio sampling sequence; Performing dynamic threshold analysis on the audio amplitude reference to obtain an acoustic quantization sequence, and performing amplitude mapping calculation based on the acoustic quantization sequence to obtain an audio quantization parameter; An audio bit width characteristic sequence in the audio signal is mapped based on the audio quantization parameter, and the audio bit width characteristic sequence is converted into a multi-bit width code to obtain a multi-bit width digital audio code stream.
[0006] Preferably, the step of inputting the multi-bit width digital audio code stream into a preset SRAM storage chip to perform a mixed mode write operation to obtain an analog-digital hybrid storage matrix includes: Performing bit width segmentation mapping on the multi-bit width digital audio code stream to obtain multi-level audio storage mapping parameters, and performing analog voltage magnitude conversion on the multi-level audio storage mapping parameters to form an analog voltage storage sequence; Inputting the analog voltage storage sequence into a preset SRAM storage chip for mixed mode coding and writing to obtain a mixed mode storage cell state, and performing voltage-current domain conversion on the mixed mode storage cell state to extract an analog current feature sequence; The analog current feature sequence is adaptively merged in bit width by a current domain parallel computing mechanism to generate an analog calculation feature vector, and digital-analog hybrid storage mapping is performed based on the analog calculation feature vector to obtain an analog-digital hybrid storage matrix.
[0007] Preferably, performing frequency domain feature extraction on the analog-digital hybrid storage matrix to obtain frequency domain extracted features includes: Performing a block Fourier transform operation on the analog-digital hybrid storage matrix to generate a time-frequency domain energy distribution, and performing nonlinear spectrum peak detection on the time-frequency domain energy distribution to obtain an audio spectrum feature sequence; Adaptively dividing the time-frequency domain energy distribution into frequency bands based on the audio spectrum feature sequence to obtain multi-scale frequency band energy parameters, and analyzing the time-domain correlation of the multi-scale frequency band energy parameters to obtain an inter-frequency band energy flow feature map; Multi-resolution analysis is performed on the energy flow characteristic graph between frequency bands through wavelet packet decomposition to obtain a multi-scale representation of frequency domain features, and the frequency domain feature space of the multi-scale representation of frequency domain features is reconstructed to output frequency domain extraction features.
[0008] Preferably, the adaptively dividing the time-frequency domain energy distribution into frequency bands based on the audio spectrum feature sequence to obtain multi-scale frequency band energy parameters includes: Calculating the perceptual importance factor in the audio spectrum feature sequence to obtain a frequency band sensitivity distribution vector, and performing weight analysis on the frequency band sensitivity distribution vector to obtain a dynamic frequency band weight; Dividing the non-uniform frequency bands in the time-frequency domain energy distribution based on the dynamic frequency band weights, and performing isolation and suppression processing on the non-uniform frequency bands to obtain an isolation band energy matrix; The isolation band energy matrix is decomposed at multiple scales to generate hierarchical band energy groups, and entropy features of the hierarchical band energy groups are extracted to obtain multi-scale band energy parameters.
[0009] Preferably, the performing noise identification and suppression on the audio signal based on the frequency domain extraction feature to obtain a time-varying noise suppression coefficient includes: Performing subspace decomposition on the frequency domain extracted features to obtain a multidimensional feature subspace group, and performing spectrum peak trajectory identification and tracking on the multidimensional feature subspace group to output a spectrum peak feature sequence; Decoupling the signal-to-noise components of the spectrum peak characteristic sequence by a singular spectrum analysis method to obtain a noise characteristic component matrix, and determining noise distribution characteristics based on the noise characteristic component matrix; Calculating a suppression coefficient based on the noise distribution characteristics to obtain a frequency band suppression coefficient group, and performing dimension expansion on the frequency band suppression coefficient group to obtain an initial suppression coefficient tensor; The initial suppression coefficient tensor is subjected to phase preservation optimization to obtain a phase compensation characteristic sequence, and the initial suppression coefficient tensor is dynamically corrected based on the phase compensation characteristic sequence to obtain a time-varying noise suppression coefficient.
[0010] Preferably, performing audio signal restoration processing on the time-varying noise suppression coefficient to obtain a noise-reduced output audio signal includes: performing bit width conversion processing on the time-varying noise suppression coefficient to obtain a suppression parameter sequence, and performing dynamic gain control on the audio signal based on the suppression parameter sequence to obtain a preliminary noise reduction signal sequence; Performing frequency-domain-time-domain conversion on the preliminary noise reduction signal sequence by discrete cosine transform to form a time-domain reconstructed signal, and compensating for an error in the time-domain reconstructed signal to generate a compensated audio signal; Filtering the compensated audio signal to obtain a de-artifacted audio signal, and performing dynamic range adjustment and enhancement on the de-artifacted audio signal to obtain an enhanced audio signal; The enhanced audio signal is subjected to multi-channel fusion recovery processing to output an audio signal, and the audio signal is format-converted to obtain a noise-reduced output audio signal.
[0011] The present invention also provides an analog-digital hybrid SRAM storage and calculation audio noise reduction system, comprising: The quantization processing module is used to perform differential sampling and quantization processing on the input audio signal to obtain a multi-bit width digital audio code stream; An operation module is used to input the multi-bit width digital audio code stream into a preset SRAM storage and computing chip to perform a mixed mode write operation to obtain an analog-digital hybrid storage matrix; An extraction processing module, configured to perform frequency domain feature extraction on the analog-digital hybrid storage matrix to obtain frequency domain extraction features; a suppression processing module, configured to identify and suppress noise on the audio signal based on the frequency domain extraction features, and obtain a time-varying noise suppression coefficient; The restoration processing module is used to perform audio signal restoration processing on the time-varying noise suppression coefficient to obtain a noise-reduced output audio signal.
[0012] The present invention also provides a computer device, comprising a memory and a processor, wherein the memory stores a computer program, and the processor implements the steps of any one of the above methods when executing the computer program.
[0013] The present invention also provides a computer-readable storage medium having a computer program stored thereon, and when the computer program is executed by a processor, the steps of any of the above methods are implemented.
[0014] The present invention provides an analog-to-digital hybrid SRAM storage and computing audio noise reduction method, comprising the following steps: performing differential sampling and quantization processing on an input audio signal to obtain a multi-bit width digital audio code stream; inputting the multi-bit width digital audio code stream into a preset SRAM storage and computing chip to perform a mixed mode write operation to obtain an analog-digital hybrid storage matrix; performing frequency domain feature extraction on the analog-digital hybrid storage matrix to obtain frequency domain extraction features; performing noise identification and suppression on the audio signal based on the frequency domain extraction features to obtain a time-varying noise suppression coefficient; performing audio signal recovery processing on the time-varying noise suppression coefficient to obtain a noise-reduced output audio signal. The method solves the technical problem that traditional digital signal processing methods usually rely on high-precision analog-to-digital converters and high-performance processors, resulting in a significant increase in power consumption and hardware costs. The method realizes effective identification and adaptive suppression of time-varying noise through intelligent analysis of frequency domain feature tensors, and generates a dynamic noise suppression coefficient matrix, thereby maintaining good audio recovery quality in complex noise environments and improving the technical effect of improving the auditory experience. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0016] Figure 1 1 is a schematic diagram of the steps of an analog-digital hybrid SRAM storage and calculation method for audio noise reduction in one embodiment of the present invention; Figure 2This is a structural block diagram of an analog-to-digital hybrid SRAM storage and calculation audio noise reduction system in one embodiment of the present invention; Figure 3 It is a schematic block diagram of the structure of a computer device according to an embodiment of the present invention.
[0017] The purpose, features and advantages of the present invention will be further described with reference to the accompanying drawings and in conjunction with the embodiments. DETAILED DESCRIPTION
[0018] In order to make the purpose, technical solutions and advantages of the present invention more clearly understood, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.
[0019] like Figure 1 As shown, Figure 1 In one embodiment of the present invention, a method for audio noise reduction using a hybrid analog-digital SRAM storage and calculation method includes the following steps: Step S1 : performing differential sampling and quantization processing on the input audio signal to obtain a multi-bit width digital audio code stream.
[0020] Specifically, in this method, the input audio signal is first subjected to differential sampling and quantization processing to obtain a multi-bit width digital audio code stream. This process is the starting point of the entire audio noise reduction process. Its core lies in reducing the redundant information between adjacent sampling points through differential sampling, thereby reducing the complexity of subsequent data processing and the transmission bandwidth requirements. Specifically, in the differential sampling process, the system does not directly collect the absolute amplitude of the original audio signal, but records the difference between the current sampling point and the sampling point at the previous moment, thereby forming a differential signal; the differential signal is then quantized, that is, it is mapped into a digital representation with a limited bit width, such as quantization with 8-bit or 16-bit precision, thereby obtaining a multi-bit width digital audio code stream. This processing method can not only effectively compress the amount of data, but also improve the matching and processing efficiency of the signal when mixed mode writing is carried out in the SRAM storage chip. For example, in an intelligent voice assistant, when a user issues a voice command, the analog audio signal collected by the microphone first enters this step for preprocessing. After differential sampling and quantization, a compact digital audio code stream is generated, which can then be efficiently sent to the SRAM storage and computing chip for mixed storage and calculation, thereby achieving low-power, high-real-time audio noise reduction function.
[0021] Step S2: input the multi-bit width digital audio code stream into a preset SRAM storage chip to perform a mixed mode write operation to obtain an analog-digital hybrid storage matrix.
[0022] Specifically, in the process of inputting the multi-bit width digital audio code stream into the preset SRAM storage and computing chip for mixed mode write operation, the core goal is to utilize the unique characteristics of the SRAM storage and computing integrated architecture to achieve efficient mapping of digital signals to analog-digital hybrid storage matrix. In this step, the multi-bit width digital audio code stream is not stored in the SRAM unit as pure digital information in the traditional way, but is converted into a data form that can carry both analog and digital signal characteristics in the SRAM unit through a specific encoding and voltage mapping strategy, thereby completing the mixed mode write operation. Specifically, inside the SRAM storage and computing chip, each storage unit is not only used to store digital bits, but also can control the amplitude or duration of the write voltage to make the unit work in different level states, forming a storage structure with analog characteristics, and finally constructing a two-dimensional or multi-dimensional hybrid storage matrix containing analog and digital information. This design allows subsequent computational tasks such as frequency domain feature extraction to be performed directly in the analog domain or hybrid domain, reducing the overhead of frequent analog-to-digital conversion and improving overall processing efficiency. For example, in an intelligent voice assistant, the voice signal after differential sampling and quantization processing is sent to the SRAM storage chip, and an analog-digital hybrid storage matrix is constructed through a mixed mode write mechanism, thereby supporting subsequent low-power, high-real-time noise recognition and suppression operations, significantly improving the energy efficiency and response speed of the audio noise reduction system.
[0023] Step S3: extract frequency domain features from the analog-digital hybrid storage matrix to obtain frequency domain extracted features.
[0024] Specifically, frequency domain feature extraction from the analog-digital hybrid storage matrix to obtain extracted frequency domain features is a key prerequisite for noise identification and suppression in this audio noise reduction method. Its purpose is to exploit the frequency domain distribution characteristics of the audio signal within the hybrid storage structure, providing high-dimensional data support for subsequent time-varying noise modeling. This process leverages the analog computing power within the SRAM memory chip. Without transferring the data to an external processor, it directly performs a Fast Fourier Transform (FFT) or similar frequency domain transformation on the audio data stored in the hybrid analog-digital state, thereby obtaining the energy distribution of the audio signal across different frequency components. Because the matrix contains both analog and digital information, frequency domain feature extraction not only involves the application of traditional digital signal processing algorithms but also requires the integration of analog domain parallel computing techniques to improve computational efficiency and reduce power consumption. The resulting frequency domain extracted features are a multidimensional array that reflects the changing trends of the audio signal across multiple dimensions, such as time, frequency, and amplitude, providing a refined data foundation for subsequent noise identification based on this tensor. For example, in an intelligent voice assistant, when a user issues the command "Hey, Xiaozhi, open navigation" in a noisy environment, the system can accurately capture the frequency band of the keyword in the voice through this step, and effectively distinguish the fan sound or traffic noise in the background, thereby providing accurate basis for the next step of noise suppression.
[0025] Step S4: performing noise identification and suppression on the audio signal based on the frequency domain extraction feature to obtain a time-varying noise suppression coefficient.
[0026] Specifically, the method uses the frequency domain extraction feature to identify and suppress noise in the audio signal, obtaining a time-varying noise suppression coefficient. This is the core computational step in achieving intelligent audio noise reduction. The key lies in utilizing the multidimensional frequency information contained in the frequency domain extraction feature to dynamically analyze the noise components in the audio signal and generate a noise suppression coefficient matrix that adapts to time variations. In the specific implementation process, the system first identifies frequency bands that may be background noise within a specific time period, such as low-frequency humming or high-frequency environmental noise, based on the energy distribution of different frequency components in the frequency domain extraction feature. Then, combining a preset noise model and an adaptive learning mechanism, the system uses an analog-digital hybrid calculation method within the SRAM memory chip to adjust the suppression weights of each frequency component in real time, thereby constructing a noise suppression coefficient matrix that evolves over time. This matrix can reflect the noise reduction intensity required at each moment and in each frequency band, providing precise control parameters for subsequent audio signal recovery. For example, in the application scenario of an intelligent voice assistant, when a user issues the command "play my favorite music" at a subway station, the system can identify the train arrival announcement and crowd noise as the main interference sources through this step, and dynamically generate the corresponding time-varying noise suppression coefficient, thereby effectively reducing the impact of background noise without losing voice clarity and improving voice recognition accuracy.
[0027] Step S5: performing audio signal restoration processing on the time-varying noise suppression coefficient to obtain a noise-reduced output audio signal.
[0028] Specifically, performing audio signal restoration processing on the time-varying noise suppression coefficients to obtain a noise-reduced output audio signal is a key step in this method for ultimately improving audio quality and ensuring speech intelligibility. Its core lies in adaptively fusing the dynamic noise suppression parameters generated in the previous stage with the original audio signal in the frequency or time domain to reconstruct clearer and cleaner speech content. Specifically, during the audio signal restoration process, the system utilizes the time-frequency suppression weights contained in the time-varying noise suppression coefficients to perform frame-by-frame weighted correction on the audio data after frequency domain feature extraction. This dynamically adjusts the energy distribution of each component based on the noise intensity at different time points and frequency bands. The corrected frequency domain signal is then restored to a time domain audio waveform using methods such as the inverse Fourier transform (IFFT), resulting in the noise-reduced output audio signal. This process fully utilizes the hybrid computing power of the SRAM memory and computing chip to perform signal reconstruction directly on-chip, avoiding the latency and power consumption issues associated with frequent data movement in traditional architectures. For example, in the application scenario of an intelligent voice assistant, when a user issues the command "Hey, Xiaozhi, call Zhang San" in a noisy environment, the system can accurately suppress the air-conditioning noise and the sound of people talking in the background according to the time-varying noise suppression coefficient through this step, and finally output a high-quality voice signal, ensuring that the voice recognition module accurately interprets the user's intention and completes the corresponding operation.
[0029] In some embodiments, performing differential sampling and quantization processing on the input audio signal to obtain a multi-bit width digital audio code stream includes: Performing time-domain differential sampling processing on the audio signal to obtain an audio sampling sequence, and extracting an audio amplitude reference of the audio sampling sequence; Performing dynamic threshold analysis on the audio amplitude reference to obtain an acoustic quantization sequence, and performing amplitude mapping calculation based on the acoustic quantization sequence to obtain an audio quantization parameter; An audio bit width characteristic sequence in the audio signal is mapped based on the audio quantization parameter, and the audio bit width characteristic sequence is converted into a multi-bit width code to obtain a multi-bit width digital audio code stream.
[0030] Specifically, the process of performing differential sampling and quantization processing on the input audio signal to obtain a multi-bit width digital audio code stream is a key preprocessing link in the entire analog-to-digital hybrid SRAM storage and computing audio noise reduction method. Its core lies in efficiently converting the original analog audio signal into a digital code stream suitable for subsequent SRAM storage and computing chip mixed mode writing through a series of structured signal analysis and transformation operations. The process specifically includes three sub-steps: first, time domain differential sampling processing is performed on the audio signal to obtain an audio sampling sequence, and an audio amplitude benchmark is extracted from it to generate an audio amplitude benchmark; second, dynamic threshold analysis is performed based on the audio amplitude benchmark to obtain an acoustic quantization sequence, and amplitude mapping calculation is performed on this basis to generate audio quantization parameters; finally, the bit width features in the audio signal are mapped according to the audio quantization parameters, and these feature sequences are converted into multi-bit width digital audio code streams as data input for subsequent SRAM storage and computing chip processing. During implementation, the system first performs time-domain differential sampling on the received analog audio signal. The core idea behind this operation is to not directly capture the absolute amplitude at each time point, but instead to record the amplitude difference between the current and previous moments, thereby constructing an audio sample sequence that reflects the local dynamic characteristics of the audio signal. This differential mechanism effectively reduces data redundancy and improves subsequent processing efficiency. The system then extracts an audio amplitude benchmark from this audio sample sequence—a reference value that represents the overall energy distribution trend of the audio signal—and organizes this into an audio amplitude benchmark for subsequent quantitative analysis. For example, in an intelligent voice assistant, when a user says, "Hey, Xiaozhi, what's the weather like tomorrow?" the continuous analog audio signal picked up by the microphone first enters this processing stage, forming a differential audio sample sequence and generating a corresponding audio amplitude benchmark. Next, the system performs dynamic threshold analysis based on this audio amplitude benchmark. This process aims to adaptively divide the quantization interval based on the instantaneous energy characteristics of the audio signal, thereby generating an environmentally adaptive acoustic quantization sequence. Unlike the traditional fixed-bit-width quantization method, this method introduces a dynamic adjustment mechanism, which enables the quantization accuracy to be automatically adjusted as the audio content changes, thereby reducing unnecessary information redundancy while ensuring sound quality. Afterwards, the system further performs amplitude mapping calculations based on the acoustic quantization sequence, that is, mapping each audio sample to a specific quantization level according to its degree of offset relative to the reference value, and thereby constructing audio quantization parameters. This tensor not only contains the quantization results of the audio signal, but also integrates multi-dimensional information such as time, frequency and amplitude, providing sufficient data support for subsequent bit-width encoding. After completing the construction of the audio quantization parameters, the system enters the final multi-bit-width encoding conversion stage, whose goal is to map the bit-width feature sequence contained in the audio signal into a multi-bit-width digital audio code stream suitable for processing by the SRAM storage and computing chip.The so-called "bit-width feature sequence" refers to the required accuracy for representing the audio signal in different time segments. For example, a lower bit-width can be used during periods of speech silence to conserve resources, while a higher bit-width can be used during periods of speech bursts to preserve detail. The system encodes these bit-width features based on the mapping rules provided by the audio quantization parameters, generating a unified, multi-bit-width digital audio stream. This stream offers excellent compressibility and compatibility, meeting the input data format requirements of the SRAM storage and computing chip while balancing audio signal quality and energy efficiency. For example, in an intelligent voice assistant application scenario, when a user issues the command "Play my favorite music" at a subway station, the system must respond quickly and accurately recognize the speech content due to the complex and constantly changing background noise. By applying the aforementioned differential sampling and quantization process to the input audio signal, the system significantly reduces data volume and processing latency while maintaining speech clarity. This provides an efficient and stable data foundation for subsequent mixed-mode writing and frequency-domain feature extraction in the SRAM storage and computing chip, ultimately achieving low-power, high-precision audio noise reduction.
[0031] In some embodiments, the step of inputting the multi-bit-width digital audio code stream into a preset SRAM storage chip for performing a mixed-mode write operation to obtain an analog-digital hybrid storage matrix includes: Performing bit width segmentation mapping on the multi-bit width digital audio code stream to obtain multi-level audio storage mapping parameters, and performing analog voltage magnitude conversion on the multi-level audio storage mapping parameters to form an analog voltage storage sequence; Inputting the analog voltage storage sequence into a preset SRAM storage chip for mixed mode coding and writing to obtain a mixed mode storage cell state, and performing voltage-current domain conversion on the mixed mode storage cell state to extract an analog current feature sequence; The analog current feature sequence is adaptively merged in bit width by a current domain parallel computing mechanism to generate an analog calculation feature vector, and digital-analog hybrid storage mapping is performed based on the analog calculation feature vector to obtain an analog-digital hybrid storage matrix.
[0032] Specifically, the process of inputting the multi-bit width digital audio code stream into a preset SRAM storage chip for mixed mode write operation to obtain an analog-digital hybrid storage matrix is one of the key technical paths for realizing analog-digital collaborative computing and low-power processing in the audio noise reduction method. Its core lies in converting the SRAM unit that is traditionally used only for storage into a composite storage structure that can carry analog signal characteristics and perform some computing functions through steps such as bit width segmented mapping, voltage magnitude conversion, mixed mode coding writing, and current domain parallel computing. Specifically, the system first performs bit width segmented mapping on the multi-bit width digital audio code stream generated after differential sampling and quantization processing, that is, according to the bit width requirements of the audio data in different time segments, it is split into multiple audio data subsets with hierarchical characteristics, and multi-level audio storage mapping parameters are constructed accordingly. This tensor not only retains the timing characteristics of the original audio signal, but also provides a structured basis for subsequent analog voltage conversion. The system then converts these multi-level audio storage mapping parameters into an analog voltage storage sequence. This step utilizes a digital-to-analog converter (DAC) mechanism to map digital audio data into analog voltage signals of specific amplitudes based on their bit width. For example, with 8-bit precision, each audio sample is mapped to one of 256 possible voltage levels; 16-bit precision allows for even higher voltage resolution. This voltage mapping method enables audio data to be stored and accessed as level states in SRAM cells, breaking the binary limitations of traditional digital storage. The system then inputs these analog voltage signals into a pre-defined SRAM memory chip, performing a mixed-mode encoding write operation. This involves writing some data in standard digital form while loading the other data as analog voltages into the drain or wordline of the SRAM cell, thereby forming a mixed-mode storage cell state. This matrix reflects the SRAM cell's ability to dynamically switch between analog and digital states, laying the foundation for subsequent parallel computing in the analog domain. After completing the mixed mode write, the system further performs a voltage-current domain conversion on the mixed mode storage cell state, that is, through the transconductance characteristics of the transistor inside the SRAM cell, the voltage signal is converted into a corresponding current response, and an analog current feature sequence is extracted. Since the current signal has natural superposition and parallelism, the analog current feature sequence is very suitable for performing basic operations such as addition and multiplication without the need to move the data to an external processor for digital domain processing. In order to further improve the flexibility and adaptability of the system, the system adopts a current domain parallel computing mechanism to perform bit-width adaptive merging of these analog current feature sequences, that is, dynamically adjust the number of current channels participating in the parallel calculation according to the complexity of the current audio signal and the noise environment, and finally generate an analog calculation feature vector. This vector is essentially a high-level abstract representation of the audio signal in the analog domain, which contains rich spectrum and energy distribution information.On this basis, the system further performs digital-analog hybrid storage mapping based on the analog calculation feature vector, that is, establishes a data structure in the SRAM array that contains both digital control information and analog calculation results, thereby ultimately generating an analog-digital hybrid storage matrix. This matrix can not only be used for subsequent frequency domain feature extraction and noise recognition, but also as an input source for the internal computing resources of the SRAM storage chip to achieve high-efficiency audio signal processing. For example, in the application scenario of an intelligent voice assistant, when a user issues the command "Hey, Xiaozhi, order me a pizza" in a noisy environment, the system uses the above-mentioned mixed mode writing process to efficiently map the voice signal into an analog-digital hybrid data structure that can be directly processed in the SRAM, significantly reducing the delay and energy consumption caused by data transfer, thereby achieving real-time, low-power audio noise reduction and speech recognition functions, and providing strong technical support for edge voice interaction devices.
[0033] In some embodiments, performing frequency domain feature extraction on the analog-digital hybrid storage matrix to obtain frequency domain extracted features includes: Performing a block Fourier transform operation on the analog-digital hybrid storage matrix to generate a time-frequency domain energy distribution, and performing nonlinear spectrum peak detection on the time-frequency domain energy distribution to obtain an audio spectrum feature sequence; Adaptively dividing the time-frequency domain energy distribution into frequency bands based on the audio spectrum feature sequence to obtain multi-scale frequency band energy parameters, and analyzing the time-domain correlation of the multi-scale frequency band energy parameters to obtain an inter-frequency band energy flow feature map; Multi-resolution analysis is performed on the energy flow characteristic graph between frequency bands through wavelet packet decomposition to obtain a multi-scale representation of frequency domain features, and the frequency domain feature space of the multi-scale representation of frequency domain features is reconstructed to output frequency domain extraction features.
[0034] Specifically, the process of extracting frequency-domain features from the analog-digital hybrid storage matrix to obtain extracted frequency-domain features is the fundamental supporting step for noise identification and suppression in this audio noise reduction method. Its core goal is to exploit the frequency-domain distribution characteristics of the audio signal from the analog-digital hybrid data structure constructed within the SRAM memory chip, providing a high-precision, multi-dimensional feature basis for the subsequent generation of time-varying noise suppression coefficients. Specifically, this step includes multiple technical sub-processes, including block Fourier transform operations, nonlinear spectral peak detection, adaptive frequency band partitioning, time-domain correlation analysis, wavelet packet decomposition, and frequency-domain feature space reconstruction. These closely interconnected steps collectively achieve an efficient mapping from the hybrid storage structure to the frequency-domain feature tensor. First, the system performs a block Fourier transform on the analog-digital hybrid storage matrix to generate a time-frequency energy distribution. The core of this step is to perform a fast Fourier transform (FFT) on the audio data segments in the hybrid storage matrix using local time windows to obtain the frequency energy distribution within each time window. This block-based processing approach not only reduces computational complexity but also effectively captures the dynamic time-frequency characteristics of the audio signal. The system then performs nonlinear peak detection on the obtained time-frequency energy distribution. This involves identifying spectral components with significant speech energy by setting a dynamic threshold or using a peak search algorithm, thereby extracting an audio spectral feature sequence. This sequence reflects the frequency ranges of key phonemes in the speech content, facilitating the precise location of subsequent noise modeling. Furthermore, based on this audio spectral feature sequence, the system performs adaptive frequency band segmentation on the time-frequency energy distribution. This involves dividing the entire frequency range into multiple sub-bands of varying scales based on the concentration of speech energy distribution. Multi-scale frequency band energy parameters are then constructed based on these parameters. This segmentation strategy automatically adjusts the frequency band width based on changes in audio content, using narrower bands in areas of active speech to improve resolution and wider bands in areas of background noise to reduce redundant computation. The system then performs time-domain correlation analysis on these multi-scale frequency band energy parameters, calculating the correlation between energy changes in different frequency bands over time. This allows for the extraction of a characteristic map of inter-band energy flow. This map reveals the energy transfer patterns between different frequency channels in the audio signal and is crucial for identifying periodic or non-stationary characteristics of ambient noise. To further enhance the representation of frequency domain features, the system introduces wavelet packet decomposition technology to perform multi-resolution analysis of the inter-band energy flow characteristic maps, thereby obtaining a multi-scale representation of frequency domain features. The advantage of wavelet packet decomposition is its ability to meticulously characterize signals at different scales, making it particularly suitable for analyzing transient components and local features in non-stationary audio signals. Through this process, the system can obtain multi-level frequency domain information, including high-frequency details and low-frequency trends, forming a richer and more robust feature set.Ultimately, the system reconstructs these multi-scale representations into a frequency-domain feature space, mapping the frequency-domain features at different scales into a unified high-dimensional tensor structure, thereby outputting frequency-domain extracted features. This tensor not only contains complete information about the audio signal across the three dimensions of time, frequency, and amplitude, but is also highly interpretable and actionable, providing a solid data foundation for subsequent noise recognition and suppression. For example, in an intelligent voice assistant application scenario, when a user utters the command "Hey, Xiaozhi, order me a pizza" in a noisy environment, the system, through the aforementioned frequency-domain feature extraction process, can accurately identify the frequency bands containing keywords such as "order" and "pizza" in the voice, and effectively distinguish between background air conditioning noise and the sounds of people chatting. The resulting frequency-domain extracted features are then used to drive the subsequent noise suppression module, ensuring that interference is removed while preserving speech clarity, thereby improving speech recognition accuracy and user experience. Throughout this process, all frequency-domain analysis is performed internally within the SRAM memory chip using a hybrid analog-digital approach, eliminating the need for frequent access to an external processor, significantly reducing system power consumption and latency, and fully demonstrating the technical advantages of this approach for low-power audio processing at the edge.
[0035] In some embodiments, the adaptively dividing the time-frequency domain energy distribution into frequency bands based on the audio spectrum feature sequence to obtain multi-scale frequency band energy parameters includes: Calculating the perceptual importance factor in the audio spectrum feature sequence to obtain a frequency band sensitivity distribution vector, and performing weight analysis on the frequency band sensitivity distribution vector to obtain a dynamic frequency band weight; Dividing the non-uniform frequency bands in the time-frequency domain energy distribution based on the dynamic frequency band weights, and performing isolation and suppression processing on the non-uniform frequency bands to obtain an isolation band energy matrix; The isolation band energy matrix is decomposed at multiple scales to generate hierarchical band energy groups, and entropy features of the hierarchical band energy groups are extracted to obtain multi-scale band energy parameters.
[0036] Specifically, the process of adaptively dividing the time-frequency domain energy distribution into frequency bands based on the audio spectrum feature sequence to obtain multi-scale frequency band energy parameters is an important technical link in the audio noise reduction method for achieving refined frequency domain processing. The core of the process is to construct a multi-level energy structure that can reflect the local spectral characteristics of the audio signal through the calculation of perceptual importance factors, the generation of dynamic frequency band weights, the isolation and suppression of non-uniform frequency bands, and multi-scale decomposition and entropy analysis. Specifically, in the frequency domain feature extraction process, the system first extracts the perceptual importance factors from the audio spectrum feature sequence, that is, according to the sensitivity of the human ear's auditory characteristics or the speech recognition model to different frequency components, assigns corresponding weight values to each frequency band, and generates a frequency band sensitivity distribution vector accordingly. This vector reflects the critical distribution of the audio signal in the frequency domain for auditory perception or speech recognition, and provides a cognitive reference for subsequent frequency band division. The system then performs a weighted analysis on the band sensitivity distribution vector to obtain dynamic band weights. The key to this step is to dynamically adjust the weighting coefficients of each frequency band based on the temporal characteristics of the current audio content, such as the intensity of speech activity and the type of background noise. This allows the resulting band division to adapt automatically to changes in the audio environment. For example, the band resolution is increased in areas with clear speech, while the bandwidth is appropriately widened in areas dominated by noise, thereby achieving efficient feature extraction. Furthermore, the system uses these dynamic band weights to perform a non-uniform division of the frequency bands in the time-frequency energy distribution. This division, in turn, divides the entire spectrum into several sub-bands of varying widths and positions, creating a non-uniform band structure. This division is more adaptable to the non-stationary and nonlinear characteristics of speech signals than traditional equal-width band division. Furthermore, the system performs isolation suppression on these non-uniform bands, identifying and masking low-energy or noise-dominated bands while retaining those containing valid speech information. This results in an isolated band energy matrix. This matrix contains only the filtered critical band information, reducing the impact of redundant data on subsequent processing and improving the targeted and accurate feature extraction. On this basis, the system employs a multi-scale decomposition strategy to hierarchically process the isolated band energy matrix. This involves further subdividing each band into multiple subbands through filter banks or multi-resolution analysis methods, thereby generating hierarchical band energy groups. This hierarchical mechanism not only enhances the ability to capture speech details but also improves the system's adaptability to complex noise environments. Finally, the system extracts entropy features from these hierarchical band energy groups, measuring the degree of uncertainty in the energy distribution within each band, and constructs multi-scale band energy parameters. This matrix integrates information from multiple dimensions, including time, frequency, energy, and uncertainty, offering high expressiveness and robustness, providing precise data support for subsequent noise identification and suppression.For example, in the application scenario of an intelligent voice assistant, when a user issues the command "Hey, Xiaozhi, order me a pizza" at a subway station, the system, through the adaptive frequency band division process described above, can accurately identify the frequency bands containing key words such as "order" and "pizza" in the voice, and effectively distinguish between air conditioning noise and the sound of people talking in the background, thereby providing a high-precision frequency domain feature foundation for the next step of noise modeling and suppression, ensuring that the final output voice signal is both clean and undistorted, significantly improving the stability of the voice recognition system and user experience. Throughout the entire process, all operations are completed based on the analog-digital hybrid computing architecture within the SRAM memory chip, fully leveraging its advantages of low power consumption and high parallelism, providing strong technical support for real-time audio processing at the edge.
[0037] In some embodiments, the performing noise identification and suppression on the audio signal based on the frequency domain extracted features to obtain a time-varying noise suppression coefficient includes: Performing subspace decomposition on the frequency domain extracted features to obtain a multidimensional feature subspace group, and performing spectrum peak trajectory identification and tracking on the multidimensional feature subspace group to output a spectrum peak feature sequence; Decoupling the signal-to-noise components of the spectrum peak characteristic sequence by a singular spectrum analysis method to obtain a noise characteristic component matrix, and determining noise distribution characteristics based on the noise characteristic component matrix; Calculating a suppression coefficient based on the noise distribution characteristics to obtain a frequency band suppression coefficient group, and performing dimension expansion on the frequency band suppression coefficient group to obtain an initial suppression coefficient tensor; The initial suppression coefficient tensor is subjected to phase preservation optimization to obtain a phase compensation characteristic sequence, and the initial suppression coefficient tensor is dynamically corrected based on the phase compensation characteristic sequence to obtain a time-varying noise suppression coefficient.
[0038] Specifically, the process of identifying and suppressing noise in the audio signal based on the frequency-domain extracted features to obtain time-varying noise suppression coefficients is a key technical step in achieving intelligent and dynamic noise modeling and control in this hybrid analog-to-digital SRAM audio noise reduction method. Its core lies in constructing a noise suppression parameter system that can adaptively adjust over time by performing subspace decomposition, peak tracking, signal-to-noise component decoupling, and phase optimization on the multidimensional spectral information contained in the frequency-domain extracted features, thereby providing a precise control basis for subsequent audio signal recovery. Specifically, the system first performs subspace decomposition on the frequency-domain extracted features to obtain a multidimensional feature subspace group. This process utilizes mathematical tools such as high-order tensor decomposition or principal component analysis (PCA) to map the spectral energy distribution in the original tensor into multiple physically meaningful feature subspaces. Each subspace corresponds to a region of energy concentration in the audio signal at a different dimension, such as a speech formant, background noise band, or transient interference source. The system then further identifies and tracks the spectral peak trajectories of these multidimensional feature subspace groups. By detecting the changing paths of energy peaks in each subspace, it extracts the stable frequency components in the speech signal and their evolution patterns, ultimately outputting a spectral peak feature sequence. This sequence reflects the spectral trajectory of key phonemes in the speech content and plays an important role in distinguishing speech from noise. Furthermore, the system uses singular spectrum analysis (SSA) to decouple the signal-to-noise components of the spectral peak feature sequence, separating the energy components into a speech-dominated "signal component" and a noise-dominated "noise component." This matrix then generates a noise feature component matrix. This matrix comprehensively captures the frequency domain distribution characteristics of the background noise, such as its periodicity and its concentration in specific frequency bands. Based on this noise feature component matrix, the system further analyzes its overall distribution pattern, extracting key characteristics such as local noise intensity, duration, and spectral correlation. This determines the noise distribution characteristics and provides a quantitative basis for subsequent noise suppression strategies. Subsequently, the system calculates the suppression coefficient corresponding to each frequency band based on the noise distribution characteristics, that is, sets a weight value for each frequency band that reflects the degree to which it is affected by noise, and thereby forms a frequency band suppression coefficient group. In order to adapt it to the multidimensional data structure required for the subsequent audio signal recovery stage, the system further performs a dimensional expansion operation on the frequency band suppression coefficient group, upgrading it from a one-dimensional vector to a three-dimensional tensor form that matches the frequency domain feature tensor, and generates an initial suppression coefficient tensor. This tensor not only contains the suppression strength of each frequency band, but also integrates information in dimensions such as time and amplitude, and has good operability and generalization capabilities. However, in the traditional noise suppression process, simply adjusting the frequency domain amplitude may cause phase distortion of the audio signal in the time domain, thereby affecting the naturalness and intelligibility of the speech.To this end, the system further performs a phase-preserving optimization operation on the initial suppression coefficient tensor. This involves analyzing the phase information of the original audio signal in the frequency domain to extract a phase-compensated feature sequence. This sequence is then used to dynamically correct the initial suppression coefficient tensor, resulting in a final noise suppression coefficient that not only effectively suppresses noise in amplitude but also maintains the coherence and integrity of the speech signal in phase. Ultimately, the system outputs a time-varying noise suppression coefficient matrix. This matrix, a multidimensional structure that evolves over time, accurately reflects the required noise suppression strength for each frequency band over time, providing high-quality control parameters for subsequent audio signal recovery. For example, in an intelligent voice assistant application scenario, when a user issues the command "Hey, Xiaozhi, order me a pizza," the ambient noise is complex and constantly changing. Through the aforementioned noise identification and suppression process, the system can identify key noise sources, such as train arrival announcements and crowd conversations, in real time and dynamically generate corresponding time-varying noise suppression coefficients. This matrix automatically adjusts the suppression strength based on the noise intensity of different frequency bands, effectively reducing background noise interference while ensuring speech clarity, thereby significantly improving the accuracy and stability of the speech recognition system. The entire process is completed based on the analog-digital hybrid computing architecture inside the SRAM storage chip, giving full play to its advantages of low power consumption and high parallelism, and providing strong technical support for real-time audio processing at the edge.
[0039] In some embodiments, performing audio signal restoration processing on the time-varying noise suppression coefficient to obtain a noise-reduced output audio signal includes: performing bit width conversion processing on the time-varying noise suppression coefficient to obtain a suppression parameter sequence, and performing dynamic gain control on the audio signal based on the suppression parameter sequence to obtain a preliminary noise reduction signal sequence; Performing frequency-domain-time-domain conversion on the preliminary noise reduction signal sequence by discrete cosine transform to form a time-domain reconstructed signal, and compensating for an error in the time-domain reconstructed signal to generate a compensated audio signal; Filtering the compensated audio signal to obtain a de-artifacted audio signal, and performing dynamic range adjustment and enhancement on the de-artifacted audio signal to obtain an enhanced audio signal; The enhanced audio signal is subjected to multi-channel fusion recovery processing to output an audio signal, and the audio signal is format-converted to obtain a noise-reduced output audio signal.
[0040] Specifically, the process of performing audio signal recovery processing on the time-varying noise suppression coefficients to obtain the noise-reduced output audio signal is the core step in achieving final speech reconstruction and quality assurance in this analog-to-digital hybrid SRAM storage and computing audio noise reduction method. Its core goal is to efficiently map the dynamic noise suppression parameters generated in the previous stage back to the original audio signal structure to reconstruct clear, natural, and highly intelligible audio content. Specifically, this step includes multiple sub-processes such as bitwidth conversion, dynamic gain control, frequency-domain-time domain conversion, error compensation, filtering artifact removal, dynamic range adjustment and enhancement, multi-channel fusion recovery, and format conversion. These interrelated steps collectively complete the complete reconstruction process from suppression parameters to high-quality audio output. At the initial stage of audio signal recovery, the system first performs bitwidth conversion on the time-varying noise suppression coefficients. This process converts the multidimensional tensor data originally used to represent the noise suppression strength into a low-dimensional digital sequence suitable for use in subsequent audio processing modules, according to the data accuracy requirements supported by the SRAM storage and computing chip. This generates a suppression parameter sequence. This sequence is essentially a set of time-varying noise suppression weights that guide the energy adjustment strategy for each frequency band in the subsequent audio signal reconstruction process. The system then applies dynamic gain control to the original audio signal based on this suppression parameter sequence. This automatically adjusts the gain coefficients for corresponding frequency bands based on the noise intensity of different frequency components within each time segment, effectively suppressing background noise while avoiding speech distortion or energy attenuation. This process effectively demonstrates the closed-loop feedback mechanism between noise modeling and signal reconstruction. Next, the system performs a frequency-to-time domain conversion on the initial noise-reduced signal sequence using the discrete cosine transform (DCT) method to generate a time-domain reconstructed signal. The core of this step is to restore the audio data after frequency-domain noise suppression to a continuous time series while preserving the temporal continuity and rhythmic characteristics of the speech signal. Because frequency-domain processing may introduce some information loss or phase distortion, the system further performs error compensation on the time-domain reconstructed signal. This locally corrects the residual information between the original audio samples and the reconstructed signal to generate a compensated audio signal. This resulting signal sounds more natural and effectively reduces the "metallic" or "stuttering" effect caused by noise reduction. To further improve audio quality, the system filters the compensated audio signal, focusing on removing high-frequency artifacts or nonlinear distortion components introduced by frequency domain processing or gain control, thereby obtaining a de-artifacted audio signal. This signal has a smoother spectral distribution and significantly improves speech clarity. On this basis, the system also performs dynamic range adjustment and enhancement on the de-artifacted audio signal. That is, based on the background noise level in the current environment and the user's auditory preferences, the system adaptively expands or compresses the overall dynamic range of the audio signal, so that the speech can maintain good audibility and comfort in different scenarios, thereby generating an enhanced audio signal.To adapt to multi-channel audio interaction scenarios such as intelligent voice assistants, the system further performs multi-channel fusion recovery processing on the enhanced audio signal, that is, synthesizing and optimizing the audio data of multiple channels according to the spatial position relationship, so that the output voice has better spatial sense and directionality, which is particularly suitable for wearable devices or in-vehicle voice systems. Finally, the system performs a format conversion operation on the fused audio signal, unifying it into a standard audio coding format such as PCM, AAC or MP3, so as to facilitate subsequent playback, transmission or storage, thereby obtaining a noise-reduced output audio signal. For example, in the application scenario of an intelligent voice assistant, when a user issues the command "Hey, Xiaozhi, help me order a pizza" at the subway station, the system can accurately reconstruct the clear voice content through the above audio signal recovery process. Even under the interference of the train arrival announcement and the crowd noise, it can still ensure the voice integrity of keywords such as "order" and "pizza", thereby improving the recognition accuracy of the voice recognition module and the user's interactive experience. Throughout the entire process, all audio reconstruction operations are completed based on the analog-digital hybrid computing architecture within the SRAM storage chip, fully leveraging its advantages of low power consumption and high parallelism, providing strong technical support for real-time audio processing at the edge.
[0041] In some embodiments, performing dynamic gain control on the audio signal based on the suppression parameter sequence to obtain a preliminary noise reduction signal sequence includes: performing sub-band decomposition on the suppression parameter sequence to obtain a multi-level suppression parameter group, and performing gain threshold calculation on the multi-level suppression parameter group to obtain a gain threshold matrix; performing segmented envelope detection on the audio signal based on the gain threshold matrix to obtain an audio envelope sequence, and performing dynamic range compression on the audio envelope sequence to obtain a compression coefficient vector; Performing time-frequency joint mapping on the compression coefficient vector to obtain a gain control feature group, and performing nonlinear compensation on the gain control feature group to obtain a compensation gain sequence; The audio signal is gain modulated based on the compensation gain sequence to obtain a modulated audio sequence, and the modulated audio sequence is subjected to noise threshold control to obtain a preliminary noise reduction signal sequence.
[0042] Specifically, the process of dynamically controlling the gain of the audio signal based on the suppression parameter sequence to obtain a preliminary noise reduction signal sequence is the key link in realizing speech signal energy reconstruction and noise suppression in the analog-to-digital hybrid SRAM storage audio noise reduction method. Its core lies in constructing a gain control system that can be adaptively adjusted with the audio content through steps such as sub-band decomposition, gain threshold calculation, envelope detection, dynamic range compression, time-frequency joint mapping, nonlinear compensation, and noise threshold control, thereby effectively reducing the impact of background noise while preserving speech clarity. Specifically, in the dynamic gain control process, the system first performs a sub-band decomposition operation on the suppression parameter sequence, that is, the originally unified noise suppression parameters are divided into multiple sub-band levels with different frequency ranges according to the spectral distribution characteristics, and a multi-level suppression parameter group is generated accordingly. This process enables each frequency band to independently set a gain strategy based on its own noise intensity and speech energy characteristics, avoiding the speech distortion problem caused by traditional single global gain control. The system then calculates gain thresholds based on this multi-level suppression parameter set. This involves setting gain adjustment thresholds based on the noise energy levels within different frequency bands to prevent over-suppression or over-amplification of low-energy speech components. This ultimately forms a gain threshold matrix. This matrix provides a precise control basis for subsequent audio signal processing. Furthermore, the system applies this gain threshold matrix to the original audio signal and performs segmented envelope detection. This involves extracting the amplitude trends of the audio signal over various time periods using a sliding time window, thereby generating an audio envelope sequence. This sequence reflects the energy fluctuations of the speech signal in the time domain and is particularly useful for identifying energy differences between speech bursts and silences. To further improve speech intelligibility and mitigate the energy compression effect caused by noise reduction, the system applies dynamic range compression to this audio envelope sequence. This process automatically adjusts the overall gain curve based on the current speech activity, ensuring that the speech remains comfortable to listen to under varying signal-to-noise ratios. This process generates a compression coefficient vector, which contains the desired compression ratio for each time segment and guides subsequent gain modulation. In order to make the gain control more precise and take into account the time-frequency structure of the speech signal, the system further performs a joint time-frequency mapping operation on the compression coefficient vector, that is, mapping the compressed gain parameters to the time-frequency two-dimensional space to form a gain control feature group. This feature group not only includes the gain requirements of each frequency band at different times, but also integrates high-order speech features such as speech resonance peaks, speaking speed and rhythm, thereby improving the intelligence and robustness of gain control. Since the audio signal may introduce certain nonlinear distortion during frequency domain processing, the system also needs to perform nonlinear compensation processing on the gain control feature group, that is, based on the local energy distribution and phase continuity of the speech signal, the gain curve is locally corrected to eliminate the "metallic sound" or "intermittent feeling" of the speech caused by the frequency domain transformation, and finally generate a compensation gain sequence.Finally, the system performs gain modulation on the audio signal based on the compensation gain sequence. This involves multiplying the audio samples within each time segment by their corresponding gain coefficients, thereby dynamically enhancing or suppressing the speech signal and generating a modulated audio sequence. To ensure the quality of the output signal, the system further applies noise thresholding to the modulated audio sequence. This involves setting a minimum energy threshold, below which signal components are treated as residual noise and attenuated. This results in a preliminary noise-reduced signal sequence. This sequence, serving as the basis for the subsequent audio restoration process, offers high speech clarity and low background noise interference, significantly improving the recognition accuracy of the speech recognition module. For example, in an intelligent voice assistant application scenario, when a user utters the command "Hey, Xiaozhi, order me a pizza" at a subway station, the system, through the aforementioned dynamic gain control process, can accurately identify the frequency bands containing key words such as "order" and "pizza." The system then dynamically adjusts the gain intensity of each frequency band based on the background announcement of train arrivals and the sounds of people chatting. This approach suppresses ambient noise while preserving speech details, ensuring the naturalness and intelligibility of the output speech. The entire process is completed based on the analog-digital hybrid computing architecture inside the SRAM storage chip, giving full play to its advantages of low power consumption and high parallelism, and providing strong technical support for real-time audio processing at the edge.
[0043] The above describes the analog-digital hybrid SRAM storage and computing audio noise reduction method in the embodiment of the present invention. The following describes the analog-digital hybrid SRAM storage and computing audio noise reduction system in the embodiment of the present invention. Figure 2 In one embodiment of the present invention, an analog-digital hybrid SRAM storage and computing audio noise reduction system includes: The quantization processing module 21 is used to perform differential sampling quantization processing on the input audio signal to obtain a multi-bit width digital audio code stream; An operation module 22 is configured to input the multi-bit width digital audio code stream into a preset SRAM storage and computing chip to perform a mixed mode write operation to obtain an analog-digital hybrid storage matrix; An extraction module 23 is configured to perform frequency domain feature extraction on the analog-digital hybrid storage matrix to obtain frequency domain extracted features; a suppression module 24, configured to identify and suppress noise on the audio signal based on the frequency domain extraction feature, and obtain a time-varying noise suppression coefficient; The restoration module 25 is configured to perform audio signal restoration processing on the time-varying noise suppression coefficient to obtain a noise-reduced output audio signal.
[0044] In this embodiment, for the specific implementation of each unit in the above system embodiment, please refer to the above method embodiment, which will not be repeated here.
[0045] Reference Figure 3In an embodiment of the present invention, a computer device is also provided, wherein the internal structure of the computer device can be as follows: Figure 3 As shown. The computer device includes a processor, memory, display screen, input device, network interface and database connected via a system bus. The processor of the computer design is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and computer program in the non-volatile storage medium. The database of the computer device is used to store the corresponding data in this embodiment. The network interface of the computer device is used to communicate with an external terminal via a network connection. When the computer program is executed by the processor, the above method is implemented.
[0046] Those skilled in the art will understand that Figure 3 The structure shown in the figure is merely a block diagram of a portion of the structure related to the solution of the present invention and does not constitute a limitation on the computer device to which the solution of the present invention is applied.
[0047] An embodiment of the present invention further provides a computer-readable storage medium having a computer program stored thereon, which implements the above-described method when executed by a processor. It is understood that the computer-readable storage medium in this embodiment can be a volatile readable storage medium or a non-volatile readable storage medium.
[0048] Those skilled in the art will appreciate that all or part of the processes in the above-described method embodiments can be implemented by instructing the relevant hardware using a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the above-described method embodiments. Any reference to memory, storage, database, or other media provided herein and used in the embodiments may include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double-speed SDRAM (SSRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), RAMbus direct RAM (RDRAM), direct RAMbus dynamic RAM (DRDRAM), and RAMbus dynamic RAM.
[0049] It should be noted that, in this document, the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, apparatus, article, or method comprising a series of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, apparatus, article, or method. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of other identical elements in the process, apparatus, article, or method comprising the element.
[0050] The above description is only a preferred embodiment of the present invention and does not limit the patent scope of the present invention. Any equivalent structure or equivalent process transformation made by using the contents of the present invention description and drawings, or directly or indirectly applied in other related technical fields, are also included in the patent protection scope of the present invention.
Claims
1. A method for audio noise reduction using analog-digital hybrid SRAM storage and calculation, characterized in that: The following steps are involved: Perform differential sampling and quantization processing on the input audio signal to obtain a multi-bit width digital audio code stream; Inputting the multi-bit width digital audio code stream into a preset SRAM storage chip to perform a mixed mode write operation to obtain an analog-digital hybrid storage matrix; Performing frequency domain feature extraction on the analog-digital hybrid storage matrix to obtain frequency domain extraction features; Performing noise identification and suppression on the audio signal based on the frequency domain extraction features to obtain a time-varying noise suppression coefficient; An audio signal restoration process is performed on the time-varying noise suppression coefficient to obtain a noise-reduced output audio signal.
2. The analog-digital hybrid SRAM storage and calculation audio noise reduction method according to claim 1, characterized in that: The differential sampling and quantization processing of the input audio signal to obtain a multi-bit width digital audio code stream includes: Performing time-domain differential sampling processing on the audio signal to obtain an audio sampling sequence, and extracting an audio amplitude reference of the audio sampling sequence; Performing dynamic threshold analysis on the audio amplitude reference to obtain an acoustic quantization sequence, and performing amplitude mapping calculation based on the acoustic quantization sequence to obtain an audio quantization parameter; An audio bit width characteristic sequence in the audio signal is mapped based on the audio quantization parameter, and the audio bit width characteristic sequence is converted into a multi-bit width code to obtain a multi-bit width digital audio code stream.
3. The analog-digital hybrid SRAM storage and calculation audio noise reduction method according to claim 1, characterized in that: The multi-bit width digital audio code stream is input into a preset SRAM storage chip to perform a mixed mode write operation to obtain an analog-digital hybrid storage matrix, including: Performing bit width segmentation mapping on the multi-bit width digital audio code stream to obtain multi-level audio storage mapping parameters, and performing analog voltage magnitude conversion on the multi-level audio storage mapping parameters to form an analog voltage storage sequence; Inputting the analog voltage storage sequence into a preset SRAM storage chip for mixed mode coding and writing to obtain a mixed mode storage cell state, and performing voltage-current domain conversion on the mixed mode storage cell state to extract an analog current feature sequence; The analog current feature sequence is adaptively merged in bit width by a current domain parallel computing mechanism to generate an analog calculation feature vector, and digital-analog hybrid storage mapping is performed based on the analog calculation feature vector to obtain an analog-digital hybrid storage matrix.
4. The analog-digital hybrid SRAM storage and calculation audio noise reduction method according to claim 1, characterized in that: The performing frequency domain feature extraction on the analog-digital hybrid storage matrix to obtain frequency domain extraction features includes: Performing a block Fourier transform operation on the analog-digital hybrid storage matrix to generate a time-frequency domain energy distribution, and performing nonlinear spectrum peak detection on the time-frequency domain energy distribution to obtain an audio spectrum feature sequence; Adaptively dividing the time-frequency domain energy distribution into frequency bands based on the audio spectrum feature sequence to obtain multi-scale frequency band energy parameters, and analyzing the time-domain correlation of the multi-scale frequency band energy parameters to obtain an inter-frequency band energy flow feature map; Multi-resolution analysis is performed on the energy flow characteristic graph between frequency bands through wavelet packet decomposition to obtain a multi-scale representation of frequency domain features, and the frequency domain feature space of the multi-scale representation of frequency domain features is reconstructed to output frequency domain extraction features.
5. The analog-digital hybrid SRAM storage and calculation audio noise reduction method according to claim 4, characterized in that: The adaptively dividing the time-frequency domain energy distribution into frequency bands based on the audio spectrum feature sequence to obtain multi-scale frequency band energy parameters includes: Calculating the perceptual importance factor in the audio spectrum feature sequence to obtain a frequency band sensitivity distribution vector, and performing weight analysis on the frequency band sensitivity distribution vector to obtain a dynamic frequency band weight; Dividing the non-uniform frequency bands in the time-frequency domain energy distribution based on the dynamic frequency band weights, and performing isolation and suppression processing on the non-uniform frequency bands to obtain an isolation band energy matrix; The isolation band energy matrix is decomposed at multiple scales to generate hierarchical band energy groups, and entropy features of the hierarchical band energy groups are extracted to obtain multi-scale band energy parameters.
6. The analog-digital hybrid SRAM storage and calculation audio noise reduction method according to claim 1, characterized in that: The performing noise identification and suppression on the audio signal based on the frequency domain extraction feature to obtain a time-varying noise suppression coefficient includes: Performing subspace decomposition on the frequency domain extracted features to obtain a multidimensional feature subspace group, and performing spectrum peak trajectory identification and tracking on the multidimensional feature subspace group to output a spectrum peak feature sequence; Decoupling the signal-to-noise components of the spectrum peak characteristic sequence by a singular spectrum analysis method to obtain a noise characteristic component matrix, and determining noise distribution characteristics based on the noise characteristic component matrix; Calculating a suppression coefficient based on the noise distribution characteristics to obtain a frequency band suppression coefficient group, and performing dimension expansion on the frequency band suppression coefficient group to obtain an initial suppression coefficient tensor; The initial suppression coefficient tensor is subjected to phase preservation optimization to obtain a phase compensation characteristic sequence, and the initial suppression coefficient tensor is dynamically corrected based on the phase compensation characteristic sequence to obtain a time-varying noise suppression coefficient.
7. The analog-digital hybrid SRAM storage and calculation audio noise reduction method according to claim 1, characterized in that: The performing audio signal restoration processing on the time-varying noise suppression coefficient to obtain a noise-reduced output audio signal includes: performing bit width conversion processing on the time-varying noise suppression coefficient to obtain a suppression parameter sequence, and performing dynamic gain control on the audio signal based on the suppression parameter sequence to obtain a preliminary noise reduction signal sequence; Performing frequency-domain-time-domain conversion on the preliminary noise reduction signal sequence by discrete cosine transform to form a time-domain reconstructed signal, and compensating for an error in the time-domain reconstructed signal to generate a compensated audio signal; Filtering the compensated audio signal to obtain a de-artifacted audio signal, and performing dynamic range adjustment and enhancement on the de-artifacted audio signal to obtain an enhanced audio signal; The enhanced audio signal is subjected to multi-channel fusion recovery processing to output an audio signal, and the audio signal is format-converted to obtain a noise-reduced output audio signal.
8. An analog-digital hybrid SRAM storage and calculation audio noise reduction system, characterized in that: include: The quantization processing module is used to perform differential sampling and quantization processing on the input audio signal to obtain a multi-bit width digital audio code stream; An operation module is used to input the multi-bit width digital audio code stream into a preset SRAM storage and computing chip to perform a mixed mode write operation to obtain an analog-digital hybrid storage matrix; An extraction processing module, configured to perform frequency domain feature extraction on the analog-digital hybrid storage matrix to obtain frequency domain extraction features; a suppression processing module, configured to identify and suppress noise on the audio signal based on the frequency domain extraction features, and obtain a time-varying noise suppression coefficient; The restoration processing module is used to perform audio signal restoration processing on the time-varying noise suppression coefficient to obtain a noise-reduced output audio signal.
9. A computer device comprising a memory and a processor, wherein a computer program is stored in the memory, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 7 are implemented.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.
Citation Information
Patent Citations
Audio denoising system and method of vehicular radio
CN104252863A
Audio noise reduction method, device and system based on deep learning
CN120148537A
Noise reducer, noise reducing method, and recording medium
US20070156399A1
Cited By
Quantitative compression and computing power adaptive optimization method and system of neural network
CN120996129A