Multifunctional audio processing method, device and system

By using neural network accelerators to perform audio signal spectrum analysis and operator optimization on end-side devices, the computing and storage challenges of multifunctional audio processing on resource-constrained devices are solved, and efficient multifunctional audio processing is achieved.

CN120375828APending Publication Date: 2025-07-25SUN YAT SEN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510453342.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-11
Publication Date
2025-07-25

AI Technical Summary

Technical Problem

When deploying multiple audio processing neural network models on end-side devices with resource-constrained, computing and storage requirements are high, making it difficult to achieve multi-functional audio processing.

Method used

The multi-functional audio processing method is adopted to analyze the spectrum of audio signal through the neural network accelerator, decompose and optimize the operators in the neural network model, support multiple audio processing functions, and manage data and parameters through the cache module.

Benefits of technology

It improves hardware reuse rate, saves hardware resource consumption, supports multiple audio processing tasks, and realizes flexible multi-function audio processing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120375828A_ABST
    Figure CN120375828A_ABST
Patent Text Reader

Abstract

The invention discloses a multifunctional audio processing method, device and system, and relates to the technical field of audio processing, and the method comprises the steps: calculating spectrum amplitude data from a spectrogram of an audio signal; selecting an audio processing function needing to be executed, and outputting control signals and instructions required by each module for realizing the audio processing function; inputting the audio signal or the spectrum amplitude data into a neural network accelerator, and transmitting the weight and bias parameters of a pre-trained neural network model corresponding to the audio processing function to a multiply-accumulate array and each storage module in the neural network accelerator; and decomposing, equivalently and optimizing various operators in the neural network model according to the instruction, so that the various operators are converted into calculations which can be executed by the neural network accelerator, and then the various executable operators are utilized to calculate the audio signal or the spectrum amplitude data according to the weight and the bias parameter to obtain a processing result of the audio processing function. The audio processing efficiency can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of audio processing, and in particular, to a multifunctional audio processing method, apparatus, and system. Background Art

[0002] Audio processing technology has a wide range of applications in daily life, including speech keyword recognition, speech noise reduction, speech separation, audio classification, speaker recognition, etc. In recent years, deep learning has been widely used in various audio processing-related tasks, such as speech keyword recognition, speech noise reduction, speaker recognition, etc. Deploying neural network models for speech processing on edge devices requires these devices to have low power consumption and real-time performance. To meet these requirements, many neural network accelerators and audio processing systems specifically designed for speech processing have been proposed. With the continuous growth of the demand for intelligence in wearable devices and smart home appliances, the trend for edge devices to support multifunctional audio processing applications is increasing.

[0003] However, deploying multiple audio processing neural networks in resource-constrained edge scenarios is challenging. Deploying multiple audio processing applications in an edge hardware leads to problems of large computing and storage requirements, and it is difficult to be deployed on edge devices with limited power consumption and computing power. Most audio systems are limited to running a single neural network model, and thus lack the flexibility to process multiple neural network models. Summary of the Invention

[0004] The main purpose of the embodiments of the present application is to propose a multifunctional audio processing method, apparatus, and system to achieve multifunctional audio processing in resource-constrained devices.

[0005] To achieve the above object, on the one hand, an embodiment of the present application proposes a multifunctional audio processing method, and the method includes the following steps:

[0006] Collect an audio signal;

[0007] Obtain a spectrogram according to the audio signal, calculate spectral amplitude data from the spectrogram, and store the spectral amplitude data in a corresponding video memory area in a cache module;

[0008] Select an audio processing function to be executed, and output control signals and instructions required for each module to implement the audio processing function;

[0009] Input the audio signal or the spectral amplitude data into a neural network accelerator, configure the working state of the neural network accelerator according to the control signals and the instructions, and transfer the weights and bias parameters of a pre-trained neural network model corresponding to the audio processing function to a multiply-accumulate array and each storage module in the neural network accelerator;

[0010] Decompose, equivalentize, and optimize various operators in the neural network model according to the instructions, so that various operators are transformed into computations executable by the neural network accelerator. Furthermore, use the executable various operators to calculate the processing result of the audio processing function based on the weight and the bias parameter for the audio signal or the spectral amplitude data.

[0011] In some embodiments, the method further includes the following steps:

[0012] Dynamically display the spectral amplitude data and the spectrum corresponding to the processing result;

[0013] Perform interactive operations dynamically through a host computer, where the interactive operations include switching the audio processing function, outputting the processing result, and starting or ending the processing of the audio signal.

[0014] To achieve the above object, on the other hand, an embodiment of the present application proposes a multifunctional audio processing device, where the device includes:

[0015] An audio acquisition unit, configured to acquire an audio signal;

[0016] A spectrum generation unit, configured to obtain a spectrogram according to the audio signal, calculate spectral amplitude data from the spectrogram, and store the spectral amplitude data in a corresponding video memory area in a cache module;

[0017] A control unit, configured to select an audio processing function to be executed, and output control signals and instructions required for each module to implement the audio processing function;

[0018] A multi-model selection unit, configured to input the audio signal or the spectral amplitude data into a neural network accelerator, configure the working state of the neural network accelerator according to the control signal and the instruction, and transfer the weights and bias parameters of the neural network model corresponding to the audio processing function and pre-trained to a multiply-accumulate array and each storage module in the neural network accelerator;

[0019] An audio processing unit, configured to decompose, equivalentize, and optimize various operators in the neural network model according to the instruction, so that various operators are transformed into computations executable by the neural network accelerator. Furthermore, use the executable various operators to calculate the processing result of the audio processing function based on the weight and the bias parameter for the audio signal or the spectral amplitude data.

[0020] To achieve the above object, on the other hand, an embodiment of the present application proposes that the system includes:

[0021] An audio acquisition module for acquiring audio signals;

[0022] A spectrum generation module for obtaining a spectrogram according to the audio signal, calculating spectrum amplitude data from the spectrogram, and storing the spectrum amplitude data in a corresponding video memory area in the cache module;

[0023] A control module for selecting an audio processing function to be executed and outputting control signals and instructions required for each module to implement the audio processing function;

[0024] A multi-model selection module for inputting the audio signal or the spectrum amplitude data into a neural network accelerator, configuring the working state of the neural network accelerator according to the control signal and the instruction, and transferring the weights and bias parameters of the pre-trained neural network model corresponding to the audio processing function to the multiply-accumulate array and each storage module in the neural network accelerator;

[0025] An audio processing module for decomposing, equivalentizing, and optimizing various operators in the neural network model according to the instruction, transforming various operators into calculations executable by the neural network accelerator, and then using the executable operators to calculate a processing result of the audio processing function according to the weights and the bias parameters for the audio signal or the spectrum amplitude data;

[0026] A cache module for controlling a DDR memory to implement the saving and reading / writing of the spectrum amplitude data, the weights of each neural network model, the bias parameters, and the instructions; wherein, the DDR memory includes multiple AXI interfaces to complete the connection between the user AXI interface and the actual AXI interface, and complete the arbitration for multiple modules to read / write the AXI interface;

[0027] A display driver module for displaying a function menu bar of the system, displaying the spectrum corresponding to the spectrum amplitude data and the processing result, and constructing an output data stream of the required display content;

[0028] An Ethernet module for completing data interaction with a host computer, dynamically sending the audio signal or the spectrum amplitude data to the host computer, and initializing the DDR memory in the cache module according to the data sent by the host computer;

[0029] A host computer interaction module for dynamically performing interactive operations on the system, where the interactive operations include switching the audio processing function, outputting the processing result, and starting or ending the processing of the audio signal.

[0030] In some embodiments, the neural network accelerator includes a logic control unit, the multiply-accumulate array, the storage module, and a post-processing unit;

[0031] The logic control unit includes a finite state machine module, a decoder module, and an address generation unit; the finite state machine module is used to switch between different applications according to different user requirements; the decoder module is used to decode the instructions configured for the neural network accelerator and generate the control signals; the address generation unit is used to perform the address mapping calculation required for the decomposition and transformation of the corresponding operator according to the decoding result of the instructions configured for the neural network accelerator, generate the read and write addresses of the storage module, control the input data and output data of the multiply-accumulate array according to the read and write addresses, and read the weights and bias parameters of the neural network model corresponding to each application to the multiply-accumulate array for calculation;

[0032] The multiply-accumulate array includes a plurality of multiply-add units, and each multiply-add unit obtains the input accumulation sum from the adjacent previous multiply-add unit, obtains the weights and input data from the storage module, performs multiply-add calculation to obtain the output accumulation sum, and outputs it to the adjacent next multiply-add unit;

[0033] The storage module includes a feature map storage module, an instruction storage module, a weight storage module, a bias storage module, an intermediate result cache module, an audio input storage module, and an audio output storage module; among them, the weight storage module, the bias storage module, and the instruction storage module adopt a ping-pong cache mode; the number, size, storage data format, and read-write methods of different modules in the storage module are different;

[0034] The post-processing unit includes a bias and activation unit, a maximum value unit, a pooling unit, a mean and variance calculation unit, and a layer normalization unit; the bias and activation unit is used to implement the calculation of the non-linear activation function and the addition of bias parameters in the neural network model; the maximum value unit is used to calculate the maximum value of the output feature map; the pooling unit is used to receive the maximum value of the output feature map and implement the calculations of maximum pooling, average pooling, and statistical value pooling; the mean and variance calculation unit is used to implement the calculation of the mean and variance of the output feature map, and transmit the mean and variance to the pooling unit and the layer normalization unit; the layer normalization unit is used to receive the mean and variance of the output feature map and implement layer normalization calculation.

[0035] In some embodiments, the multiply-accumulate unit is configured to truncate the input quantized fixed-point data according to the truncation bit number calculated by the post-processing unit, then multiply it by the quantized fixed-point weight provided by the storage module and add it to the calculation result output by the previous adjacent multiply-accumulate unit, save the obtained output result in the output register, and determine whether to output the output result in the next cycle or output the input data to the next adjacent multiply-accumulate unit by a skip connection selection logic.

[0036] In some embodiments, the display driving module is configured to implement the line buffer refresh operation and the synchronization signal output control of the display data based on the line scan coordinate dynamic management mechanism, download the corresponding display data to the line buffer unit according to the current line scan coordinate, and then complete the timing matching output of the data and the line synchronization signal.

[0037] In some embodiments, the input data of the line buffer unit includes the graphic interface rendering data and the spectrum analysis data; the line buffer unit is configured to implement the binary pixel mapping mechanism for the graphic interface rendering data, generate a pixel array conforming to the RGB888 standard through the two-color decoding process of the preset background color and the highlight color in the timing construction stage; execute the dynamic length decoding algorithm for the spectrum analysis data, convert the spectrum feature parameters uploaded by the spectrum generation module into the pixel lighting length of the longitudinal spectrum histogram by parsing, and the number of continuously activated pixels in each row has a linear mapping relationship with the amplitude value of the corresponding frequency point.

[0038] In some embodiments, the host computer interaction module is implemented by a Python script, and the Python script includes a main process, an audio acquisition process, and an inference calculation process;

[0039] The main process is configured to maintain the GUI interface, ensure the real-time refresh of the interface through the event-driven mechanism, and at the same time act as a cross-process communication hub to implement instruction distribution and result aggregation; the main process is also configured to write neural network parameters to the cache module of the neural network accelerator through the network port;

[0040] The audio acquisition process is configured to execute the audio stream circular buffer management and implement the data life cycle control policy, and the data life cycle control policy includes the real-time audio stream writing operation and the first-in-first-out replacement policy for expired data;

[0041] The inference calculation process is configured to build an asynchronous task scheduling mechanism, dynamically retrieve the data in the corresponding section of the audio buffer by parsing the operation instructions issued by the main control process, complete the algorithm inference calculation of the specified function, and feedback the calculation result to the main control process according to the preset inference result feedback protocol.

[0042] The embodiments of the present application at least include the following beneficial effects:

[0043] 1. Since the neural network accelerator of the present application supports multiple neural network operators, by decomposing and optimizing the designs of multiple neural operators and equating multiple neural network operators to the computations supported by the neural network accelerator, it brings the advantages of high hardware reuse rate and saving hardware resource consumption.

[0044] 2. Since the neural network accelerator of the present application uses the original audio signal or the spectral amplitude data of the audio signal as the input when computing various audio processing tasks, without the introduction of feature extraction modules such as Mel Frequency Cepstral Coefficients, it brings the advantage of saving the computation and resource consumption of additional feature extraction modules. BRIEF DESCRIPTION OF THE DRAWINGS

[0045] To more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0046] Figure 1 It is a schematic flowchart of a multifunctional audio processing method provided by an embodiment of the present application;

[0047] Figure 2 It is an exemplary structural diagram of a multifunctional audio processing system provided by an embodiment of the present application;

[0048] Figure 3 It is an exemplary structural diagram of a display driving module provided by an embodiment of the present application;

[0049] Figure 4 It is an exemplary structural diagram of a neural network accelerator provided by an embodiment of the present application;

[0050] Figure 5 It is an exemplary circuit structural diagram of a multiply-accumulate array provided by an embodiment of the present application;

[0051] Figure 6 It is an exemplary circuit structural diagram of a multiply-add unit provided by an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0052] To make the objectives, technical solutions, and advantages of this application more clearly understood, the following further elaborates on this application in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely for explaining this application and are not intended to limit this application. When the following description involves the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The implementation manners described in the following exemplary embodiments do not represent all implementation manners consistent with the embodiments of this application. They are merely examples of devices and methods consistent with some aspects of the embodiments of this application as detailed in the appended claims.

[0053] It can be understood that the terms "first", "second", etc. used in this application may be used herein to describe various concepts, but unless otherwise specified, these concepts are not limited by these terms. These terms are only used to distinguish one concept from another. For example, without departing from the scope of the embodiments of this application, the first information may also be referred to as the second information. Similarly, the second information may also be referred to as the first information. Depending on the context, as used herein, the words "if", "when" can be interpreted as "when...", "while...", or "in response to determining".

[0054] The terms "at least one", "multiple", "each", "any one", etc. used in this application, at least one includes one, two, or more than two, multiple includes two or more than two, each refers to each of the corresponding multiple, and any one refers to any one of the multiple.

[0055] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which this application belongs. The terms used herein are only for the purpose of describing the embodiments of this application and are not intended to limit this application.

[0056] Before elaborating on the embodiments of this application in detail, first, some related technologies involved in the embodiments of this application are described as follows:

[0057] Deploying multiple neural networks for audio processing in a resource-constrained edge scenario is challenging. Deploying multiple audio processing applications in an edge hardware results in problems of high computing and storage requirements, and it is difficult to be deployed on edge devices with limited power consumption and computing power. Most existing audio systems are limited to running a single neural network model, so they lack the flexibility to process multiple neural network models and cannot achieve multifunctional audio processing based on neural networks.

[0058] Refer to Figure 1 , the embodiments of this application provide a multifunctional audio processing method, which may include but is not limited to S100 to S140, specifically as follows:

[0059] S100: Collect an audio signal;

[0060] S110: Obtain a spectrogram based on the audio signal, calculate spectral amplitude data from the spectrogram, and store the spectral amplitude data in a corresponding video memory area in a cache module;

[0061] S120: Select an audio processing function to be executed, and output control signals and instructions required for each module to implement the audio processing function;

[0062] S130: Input the audio signal or the spectral amplitude data into a neural network accelerator, configure the working state of the neural network accelerator according to the control signals and the instructions, and transfer weights and bias parameters of a pre-trained neural network model corresponding to the audio processing function to a multiply-accumulate array and each storage module in the neural network accelerator;

[0063] S140: Decompose, equivalentize, and optimize various operators in the neural network model according to the instructions, transform various operators into calculations executable by the neural network accelerator, and then use the executable operators to calculate a processing result of the audio processing function based on the weights and the bias parameters for the audio signal or the spectral amplitude data.

[0064] Exemplarily, the processing result of the embodiment of the present application includes at least one of an audio classification result, a keyword recognition result, a voice conversion result, a noise reduction result, a speech recognition result, or a voiceprint recognition result corresponding to the audio signal.

[0065] Optionally, the method further includes the following steps:

[0066] Dynamically display the spectral amplitude data and the spectrum corresponding to the processing result;

[0067] Perform dynamic interaction operations through a host computer, where the interaction operations include switching the audio processing function, outputting the processing result, and starting or ending the processing of the audio signal.

[0068] To solve the problems existing in the prior art, an embodiment of the present application further provides a multifunctional audio processing device, and the device includes:

[0069] An audio acquisition unit, configured to collect an audio signal;

[0070] A spectrum generation unit, configured to obtain a spectrogram based on the audio signal, calculate spectral amplitude data from the spectrogram, and store the spectral amplitude data in a corresponding video memory area in a cache module;

[0071] A control unit for selecting an audio processing function to be executed and outputting control signals and instructions required for each module to implement the audio processing function;

[0072] A multi-model selection unit for inputting the audio signal or the spectral amplitude data into a neural network accelerator, configuring the working state of the neural network accelerator according to the control signal and the instruction, and transferring the weights and bias parameters of the pre-trained neural network model corresponding to the audio processing function to the multiply-accumulate array and each storage module in the neural network accelerator;

[0073] An audio processing unit for decomposing, equivalentizing, and optimizing various operators in the neural network model according to the instruction, transforming various operators into calculations executable by the neural network accelerator, and then calculating the processing result of the audio processing function from the audio signal or the spectral amplitude data according to the weights and the bias parameters by using the executable operators.

[0074] To solve the problems existing in the prior art, an embodiment of the present application further proposes that the system includes:

[0075] An audio acquisition module for acquiring an audio signal;

[0076] A spectrum generation module for obtaining a spectrogram from the audio signal, calculating spectral amplitude data from the spectrogram, and storing the spectral amplitude data in a corresponding video memory area in a cache module;

[0077] A control module for selecting an audio processing function to be executed and outputting control signals and instructions required for each module to implement the audio processing function;

[0078] A multi-model selection module for inputting the audio signal or the spectral amplitude data into a neural network accelerator, configuring the working state of the neural network accelerator according to the control signal and the instruction, and transferring the weights and bias parameters of the pre-trained neural network model corresponding to the audio processing function to the multiply-accumulate array and each storage module in the neural network accelerator;

[0079] An audio processing module for decomposing, equivalentizing, and optimizing various operators in the neural network model according to the instruction, transforming various operators into calculations executable by the neural network accelerator, and then calculating the processing result of the audio processing function from the audio signal or the spectral amplitude data according to the weights and the bias parameters by using the executable operators;

[0080] A cache module, which is used to control a DDR memory to implement the storage and read / write of the spectral amplitude data, the weights of each neural network model, the bias parameters, and the instructions; wherein, the DDR memory includes multiple AXI interfaces to complete the connection between the user AXI interface and the actual AXI interface, and complete the arbitration of multiple modules for reading and writing the AXI interfaces.

[0081] A display driver module, which is used to display the function menu bar of the system, display the spectral amplitude data and the spectrum corresponding to the processing result, and construct an output data stream of the content to be displayed.

[0082] An Ethernet module, which is used to complete data interaction with a host computer, dynamically send the audio signal or the spectral amplitude data to the host computer, and initialize the DDR memory in the cache module according to the data sent by the host computer.

[0083] A host computer interaction module, which is used to dynamically perform interaction operations on the system, and the interaction operations include switching the audio processing function, outputting the processing result, and starting or ending the processing of the audio signal.

[0084] Optionally, the neural network accelerator includes a logic control unit, the multiply-accumulate array, the storage module, and a post-processing unit.

[0085] The logic control unit includes a finite state machine module, a decoder module, and an address generation unit; the finite state machine module is used to switch between different applications according to different user requirements; the decoder module is used to decode the instructions configured for the neural network accelerator and generate the control signals; the address generation unit is used to perform address mapping calculations required for the decomposition and transformation of the corresponding operator according to the decoding result of the instructions configured for the neural network accelerator, generate the read / write addresses of the storage module, control the input data and output data of the multiply-accumulate array according to the read / write addresses, and read the weights and bias parameters of the neural network models corresponding to each application to the multiply-accumulate array for calculation.

[0086] The multiply-accumulate array includes multiple multiply-add units, and each multiply-add unit obtains the input accumulation sum from the adjacent previous multiply-add unit, obtains the weight and input data from the storage module, performs a multiply-add calculation to obtain the output accumulation sum, and outputs it to the adjacent next multiply-add unit.

[0087] The storage module includes a feature map storage module, an instruction storage module, a weight storage module, a bias storage module, an intermediate result cache module, an audio input storage module, and an audio output storage module; among them, the weight storage module, the bias storage module, and the instruction storage module adopt a ping-pong cache mode; the number, size, stored data format, and read / write methods of different modules in the storage module are different;

[0088] The post-processing unit includes a bias and activation unit, a maximum value unit, a pooling unit, a mean and variance calculation unit, and a layer normalization unit; the bias and activation unit is used to implement the calculation of the non-linear activation function and the addition of bias parameters in the neural network model; the maximum value unit is used to calculate the maximum value of the output feature map; the pooling unit is used to receive the maximum value of the output feature map and implement the calculations of max pooling, average pooling, and statistical value pooling; the mean and variance calculation unit is used to implement the calculation of the mean and variance of the output feature map, and transmit the mean and variance to the pooling unit and the layer normalization unit; the layer normalization unit is used to receive the mean and variance of the output feature map and implement layer normalization calculation.

[0089] In some embodiments, the multiply-accumulate unit is used to truncate the input fixed-point quantized data according to the truncation bit number calculated by the post-processing unit, then multiply it with the fixed-point quantized weight provided by the storage module and add it to the calculation result output by the previous adjacent multiply-accumulate unit, and the obtained output result is saved in the output register, and a skip connection selection logic determines whether to output the output result in the next cycle or output the input data to the next adjacent multiply-accumulate unit.

[0090] Optionally, the display driving module is used to implement the row buffer refresh operation and synchronous signal output control of the display data based on the row scan coordinate dynamic management mechanism, download the corresponding display data to the row buffer unit according to the current row scan coordinate, and then complete the timing matching output of the data and the row synchronous signal.

[0091] Optionally, the input data of the row buffer unit includes graphic interface rendering data and spectrum analysis data; the row buffer unit is used to implement a binary pixel mapping mechanism for the graphic interface rendering data, generate a pixel array conforming to the RGB888 standard through two-color decoding processing of a preset background color and a highlight color in the timing construction stage; execute a dynamic length decoding algorithm on the spectrum analysis data, convert it into the pixel lighting length of a longitudinal spectrum histogram by parsing the spectrum feature parameters uploaded by the spectrum generation module, where the number of continuously activated pixels in each row has a linear mapping relationship with the amplitude value of the corresponding frequency point.

[0092] Optionally, the host computer interaction module is implemented by a Python script, and the Python script includes a main process, an audio acquisition process, and an inference calculation process;

[0093] The main process is used to maintain the GUI interface, ensure real-time interface refresh through the event-driven mechanism, and at the same time serve as a cross-process communication hub to implement instruction distribution and result aggregation; the main process is also used to write neural network parameters to the cache module of the neural network accelerator through the network port;

[0094] The audio acquisition process is used to perform audio stream circular buffer management and implement a data life cycle control policy, and the data life cycle control policy includes real-time audio stream writing operations and a first-in, first-out replacement policy for expired data;

[0095] The inference calculation process is used to build an asynchronous task scheduling mechanism, dynamically retrieve data in the corresponding section of the audio buffer by parsing the operation instructions issued by the main control process, complete algorithm inference calculations for specified functions, and feedback the calculation results to the main control process according to the preset inference result feedback protocol.

[0096] Next, specific application examples will be combined to introduce and explain the solution of the embodiment of the present application in detail.

[0097] Refer to Figure 2 , the embodiment of the present application provides a multifunctional audio processing system, specifically including:

[0098] An audio acquisition module, which is used to acquire audio, convert the audio into a digital signal, and output the audio to the neural network accelerator module and the spectrum generation module.

[0099] A control module, which is used to control the states of each module in the system and switch the functions of the multifunctional audio processing system.

[0100] A neural network accelerator module, which is used to implement and accelerate the inference calculations of neural network models corresponding to multiple audio processing functions, and output processed audio results or recognition results.

[0101] A spectrum generation module, which is used to implement FFT calculations to convert the input audio or output audio into frequency domain data, convert the complex frequency domain data into spectrum amplitude data, and store the frequency domain data in the corresponding video memory area in the DDR.

[0102] A cache module, which is used to control a DDR to implement the storage and reading and writing of spectrum data, weights, biases, and instructions of each neural network model. The control logic therein expands a single AXI interface of the DDR into multiple interfaces, completes the connection between the user AXI interface and the actual AXI interface, and completes the arbitration of multiple user read and write AXI interfaces.

[0103] An Ethernet module is used to complete data interaction with the host computer, send audio data to the host computer in real time, and initialize the DDR in the cache module according to the data sent by the host computer.

[0104] Refer to Figure 3 , the display driver module is used to download data from the cache module according to the display timing to refresh the content of the line buffer, and then perform specific data decoding according to different data types, and combine the field synchronization signal and the line synchronization signal to generate a display output data stream.

[0105] The host computer interaction module is used to initialize the neural network parameters of the cache module and implement the extended audio processing function.

[0106] Next, the neural network accelerator of the embodiment of the present application will be described. Refer to Figure 4 , the embodiment of the present application provides an example structure diagram of a neural network accelerator.

[0107] The neural network accelerator includes a logic control unit, a multiply-accumulate array, a storage module, and a post-processing unit.

[0108] The logic control unit includes a finite state machine module, a decoder module, and an address generation unit. The finite state machine module is used to switch between different applications according to different user requirements. The decoder module is used to decode the instructions configured for the neural network accelerator and generate control signals. The address generation unit is used to perform address mapping calculations required for the decomposition and transformation of the corresponding operators according to the decoding results of the instructions configured for the neural network accelerator, generate read and write addresses of the storage unit, control the input data and output data of the multiply-accumulate array according to the read and write addresses, and read the weights and bias parameters of the neural network model corresponding to each application to the multiply-accumulate array for calculation.

[0109] The multiply-accumulate array is an array composed of multiple multiply-accumulate units. Each multiply-accumulate unit obtains the input accumulation sum from the adjacent previous multiply-accumulate unit, obtains the weight and input data from the storage module, performs multiply-accumulate calculation to obtain the output accumulation sum, and outputs it to the adjacent next multiply-accumulate unit.

[0110] Furthermore, the multiply-accumulate array of the embodiment of the present application will be described. Refer to Figure 5 , the embodiment of the present application provides an example circuit structure diagram of a multiply-accumulate array.

[0111] Specifically, the multiply-accumulate array of the embodiments of the present application may be a two-dimensional rectangular array composed of multiply-add units. The multiply-add units in each row calculate and transmit the partial sum from left to right in sequence. The input partial sum of the multiply-add unit located on the leftmost side may come from the intermediate result cache module or the feature map storage module to implement the bypass addition operation of the residual block. The output partial sum of the multiply-add unit located on the rightmost side will be stored in the intermediate result cache module. The input partial sum of the multiply-add unit located in the middle is the calculation result of the adjacent multiply-add unit on its left in the previous cycle, and its calculation result will also be transmitted to the adjacent multiply-add unit on its right in the next cycle as its input partial sum. By designing the data flow to batch process the input feature map and the convolution kernel, the calculation of one-dimensional pointwise convolution or one-dimensional layer-by-layer convolution is realized. Then, according to the configuration of the instructions, various operators in the neural network model are decomposed, equivalent, and optimized, so that the calculations of various neural operators in the neural network model are transformed into the calculations of one-dimensional pointwise convolution or one-dimensional layer-by-layer convolution in the multiply-accumulate array, realizing the calculations of various neural network operators, including: one-dimensional pointwise convolution, one-dimensional layer-by-layer convolution, one-dimensional conventional convolution, fully connected layer, one-dimensional transposed convolution, one-dimensional layer-by-layer dilated convolution.

[0112] Further, the multiply-add unit in the multiply-accumulate array of the embodiments of the present application is described. Referring to Figure 6 , an example circuit structure diagram of a multiply-add unit is provided in the embodiments of the present application.

[0113] Specifically, the working process of the multiply-add unit may include: the input fixed-point quantized data is first truncated according to the truncation bit number calculated by the post-processing unit, and then multiplied by the fixed-point quantized weight data provided by the storage module and added to the calculation result output by the adjacent previous unit. The obtained output result is stored in an output register, and a skip connection selection logic determines whether to output the output result in the next cycle or directly output the input data to the next adjacent unit.

[0114] The storage module includes five parts: a feature map storage module, an instruction storage module, a weight storage module, a bias storage module, an intermediate result cache module, an audio input storage module, and an audio output storage module. Among them, the weight storage module, the bias storage module, and the instruction storage module adopt the ping-pong cache mode. The number, size, stored data format, and read-write methods of different modules are different.

[0115] The post - processing unit includes a bias and activation unit, a maximum value unit, a pooling unit, a mean and variance calculation unit, and a layer normalization unit. The bias and activation unit is used to implement the calculation of non - linear activation functions that may be involved in the neural network model and the addition of biases. The maximum value unit is used to calculate the maximum value of the output feature map. The pooling unit is used to receive the maximum value of the output feature map and implement the calculations of max - pooling, average - pooling, and statistical - value pooling. The mean and variance calculation unit is used to calculate the mean and variance of the output feature map and transmit the mean and variance to the pooling unit and the layer normalization unit. The layer normalization unit is used to receive the mean and variance of the output feature map and implement layer normalization calculations.

[0116] The beneficial effects of the embodiments of this application at least include:

[0117] 1. Since a neural network accelerator module that supports multiple neural network operators is adopted, by decomposing and optimizing the designs of multiple neural operators and equating multiple neural network operators to the calculations supported by the neural network accelerator, it brings the advantages of high hardware reuse rate and saving hardware resource consumption.

[0118] 2. Since the neural network accelerator uses the original audio input or the audio spectrum as the input when calculating various audio processing tasks, without the introduction of feature extraction modules such as Mel - Frequency Cepstral Coefficients, it brings the advantages of saving the calculations and resource consumption of additional feature extraction modules.

[0119] 3. Since it supports modular host - computer interaction functions, while accelerating the conventional audio processing process, the present invention can quickly deploy customized function modules and can perform multi - dimensional data analysis and processing on the collected audio, bringing the dual advantages of both system scalability and engineering practicality.

[0120] The embodiments described in the embodiments of this application are for more clearly illustrating the technical solutions of the embodiments of this application, and do not constitute a limitation on the technical solutions provided by the embodiments of this application. Those skilled in the art know that with the evolution of technology and the emergence of new application scenarios, the technical solutions provided by the embodiments of this application are equally applicable to similar technical problems.

[0121] Those skilled in the art can understand that the technical solutions shown in the figures do not constitute a limitation on the embodiments of this application, and may include more or fewer technical solutions than those shown in the figures, or combine some technical solutions, or different technical solutions.

[0122] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0123] Those of ordinary skill in the art will understand that all or some of the steps in the methods disclosed above, and the functional modules / units in systems and devices, can be implemented as software, firmware, hardware, or a suitable combination thereof.

[0124] As used in the description of this application and the above drawings, the terms "first", "second", "third", "fourth", etc. (if any) are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances so that the embodiments of this application described here can be implemented in an order different from those illustrated or described here. In addition, the terms "comprising" and "having", and any variations thereof, are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that comprises a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products, or devices.

[0125] It should be understood that in this application, "at least one (item)" means one or more, and "a plurality" means two or more. "And / or" is used to describe the association relationship of associated objects and indicates that three relationships may exist. For example, "A and / or B" may mean: only A exists, only B exists, and both A and B exist at the same time. Among them, A and B can be singular or plural. The character " / " generally means that the associated objects before and after are in an "or" relationship. "At least one (one) of the following" or a similar expression means any combination of these items, including any combination of single items (ones) or plural items (ones). For example, at least one (one) of a, b, or c may mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, and c can be single or multiple.

[0126] In several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the above division of units is only a logical function division, and there may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection to each other can be through some interfaces, and the indirect coupling or communication connection of devices or units can be in electrical, mechanical, or other forms.

[0127] The units described above as separate components may or may not be physically separated. The components shown as units may or may not be physical units, that is, they may be located in one place or distributed over multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0128] In addition, each functional unit in various embodiments of the present application may be integrated in a processing unit, may exist separately as individual physical units, or two or more units may be integrated in one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional units.

[0129] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes multiple instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods in various embodiments of the present application. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs that can store programs.

[0130] The preferred embodiments of the embodiments of the present application have been described above with reference to the accompanying drawings, but this does not limit the scope of the rights of the embodiments of the present application. Any modifications, equivalent replacements, and improvements made by those skilled in the art without departing from the scope and essence of the embodiments of the present application shall be within the scope of the rights of the embodiments of the present application.

Claims

1. A multifunctional audio processing method, characterized in that The method includes the following steps: Collect an audio signal; Obtain a spectrogram according to the audio signal, calculate spectral amplitude data from the spectrogram, and store the spectral amplitude data in a corresponding video memory area in a cache module; Select an audio processing function to be executed, and output control signals and instructions required for each module to implement the audio processing function; Input the audio signal or the spectral amplitude data into a neural network accelerator, configure the working state of the neural network accelerator according to the control signals and the instructions, and transfer the weights and bias parameters of a pre-trained neural network model corresponding to the audio processing function to a multiply-accumulate array and each storage module in the neural network accelerator; Decompose, equivalentize, and optimize various operators in the neural network model according to the instructions, transform various operators into computations executable by the neural network accelerator, and then use the executable operators to calculate a processing result of the audio processing function according to the weights and the bias parameters for the audio signal or the spectral amplitude data.

2. The multifunctional audio processing method according to claim 1, characterized in that The method further includes the following steps: Dynamically display the spectral amplitude data and the spectrum corresponding to the processing result; Perform dynamic interaction operations through a host computer, where the interaction operations include switching the audio processing function, outputting the processing result, and starting or ending the processing of the audio signal.

3. A multifunctional audio processing device, characterized in that, The apparatus includes: An audio acquisition unit for collecting an audio signal; A spectrum generation unit for obtaining a spectrogram according to the audio signal, calculating spectral amplitude data from the spectrogram, and storing the spectral amplitude data in a corresponding video memory area in a cache module; A control unit for selecting an audio processing function to be executed and outputting control signals and instructions required for each module to implement the audio processing function; A multi-model selection unit for inputting the audio signal or the spectral amplitude data into a neural network accelerator, configuring the working state of the neural network accelerator according to the control signals and the instructions, and transferring the weights and bias parameters of a pre-trained neural network model corresponding to the audio processing function to a multiply-accumulate array and each storage module in the neural network accelerator; An audio processing unit for decomposing, equivalentizing, and optimizing various operators in the neural network model according to the instructions, transforming various operators into computations executable by the neural network accelerator, and then using the executable operators to calculate a processing result of the audio processing function according to the weights and the bias parameters for the audio signal or the spectral amplitude data.

4. A multifunctional audio processing system, characterized in that, The system includes: An audio acquisition module for collecting an audio signal; A spectrum generation module for obtaining a spectrogram according to the audio signal, calculating spectral amplitude data from the spectrogram, and storing the spectral amplitude data in a corresponding video memory area in a cache module; A control module for selecting an audio processing function to be executed and outputting control signals and instructions required for each module to implement the audio processing function; A multi-model selection module for inputting the audio signal or the spectral amplitude data into a neural network accelerator, configuring the working state of the neural network accelerator according to the control signal and the instruction, and transferring the weights and bias parameters of the pre-trained neural network model corresponding to the audio processing function to the multiply-accumulate array and each storage module in the neural network accelerator; An audio processing module for decomposing, equivalentizing, and optimizing various operators in the neural network model according to the instruction, transforming various operators into calculations executable by the neural network accelerator, and then calculating the processing result of the audio processing function for the audio signal or the spectral amplitude data according to the weights and the bias parameters by using the executable various operators; A cache module for controlling a DDR memory to implement the saving, reading, and writing of the spectral amplitude data, the weights of each neural network model, the bias parameters, and the instruction; wherein, the DDR memory includes multiple AXI interfaces to complete the connection between the user AXI interface and the actual AXI interface, and complete the arbitration of multiple modules for reading and writing the AXI interface; A display driving module for displaying the function menu bar of the system, displaying the spectrum corresponding to the spectral amplitude data and the processing result, and constructing an output data stream of the required display content; An Ethernet module for completing data interaction with a host computer, dynamically sending the audio signal or the spectral amplitude data to the host computer, and initializing the DDR memory in the cache module according to the data sent by the host computer; A host computer interaction module for dynamically performing interactive operations on the system, and the interactive operations include switching the audio processing function, outputting the processing result, and starting or ending the processing of the audio signal.

5. A multifunctional audio processing system according to claim 4, wherein The neural network accelerator includes a logic control unit, the multiply-accumulate array, the storage module, and a post-processing unit; The logic control unit includes a finite state machine module, a decoder module, and an address generation unit; the finite state machine module is used for switching between different applications according to different user requirements; the decoder module is used for decoding the instructions configured for the neural network accelerator and generating the control signal; the address generation unit is used for performing address mapping calculations required for the decomposition and transformation of the corresponding operators according to the decoding result of the instructions configured for the neural network accelerator, generating the read and write addresses of the storage module, controlling the input data and output data of the multiply-accumulate array according to the read and write addresses, and reading the weights and bias parameters of the corresponding neural network models of each application to the multiply-accumulate array for calculation; The multiply-accumulate array includes multiple multiply-add units, and each multiply-add unit obtains the input accumulation sum from the adjacent previous multiply-add unit, obtains the weight and the input data from the storage module, performs multiply-add calculation to obtain the output accumulation sum, and outputs it to the adjacent next multiply-add unit; The storage module includes a feature map storage module, an instruction storage module, a weight storage module, a bias storage module, an intermediate result cache module, an audio input storage module, and an audio output storage module; among them, the weight storage module, the bias storage module, and the instruction storage module adopt a ping-pong cache mode; the number, size, stored data format, and read-write methods of different modules in the storage module are different; The post-processing unit includes a bias and activation unit, a maximum value unit, a pooling unit, a mean and variance calculation unit, and a layer normalization unit; the bias and activation unit is used to implement the calculation of the non-linear activation function and the addition of bias parameters in the neural network model; the maximum value unit is used to calculate the maximum value of the output feature map; the pooling unit is used to receive the maximum value of the output feature map and implement the calculations of maximum value pooling, average value pooling, and statistical value pooling; the mean and variance calculation unit is used to implement the calculation of the mean and variance of the output feature map, and transmit the mean and variance to the pooling unit and the layer normalization unit; the layer normalization unit is used to receive the mean and variance of the output feature map and implement layer normalization calculation.

6. A multifunctional audio processing system according to claim 5, characterized in that, The multiply-accumulate unit is used to truncate the input fixed-point quantized data according to the truncation bit number calculated by the post-processing unit, then multiply it by the fixed-point quantized weight provided by the storage module and add it to the calculation result output by the previous adjacent multiply-accumulate unit. The obtained output result is stored in the output register, and a skip connection selection logic determines whether to output the output result in the next cycle or output the input data to the next adjacent multiply-accumulate unit.

7. A multifunctional audio processing system according to claim 4, wherein, The display driving module is used to implement the row buffer refresh operation and synchronous signal output control of the display data based on the row scan coordinate dynamic management mechanism, download the corresponding display data to the row buffer unit according to the current row scan coordinate, and then complete the timing matching output of the data and the row synchronous signal.

8. A multifunctional audio processing system according to claim 7, characterized in that, The input data of the row buffer unit includes graphic interface rendering data and spectrum analysis data; the row buffer unit is used to implement a binary pixel mapping mechanism for the graphic interface rendering data, generate a pixel array conforming to the RGB888 standard through two-color decoding processing of a preset background color and a highlight color in the timing construction stage; execute a dynamic length decoding algorithm for the spectrum analysis data, convert the spectrum feature parameters uploaded by the spectrum generation module by parsing the spectrum into the pixel lighting length of a longitudinal spectrum histogram, where the number of continuously activated pixels in each row has a linear mapping relationship with the amplitude value of the corresponding frequency point.

9. A multifunctional audio processing system according to claim 4, characterized in that, The host computer interaction module is implemented with a Python script, and the Python script includes a main process, an audio acquisition process, and an inference calculation process; The main process is used to maintain the GUI interface, ensure real-time interface refresh through an event-driven mechanism, and at the same time act as a cross-process communication hub to implement instruction distribution and result aggregation; the main process is also used to write neural network parameters to the cache module of the neural network accelerator through a network port; The audio acquisition process is used to perform audio stream circular buffer management and implement a data life cycle control policy, and the data life cycle control policy includes real-time audio stream writing operations and a first-in, first-out replacement policy for expired data; The inference calculation process is used to construct an asynchronous task scheduling mechanism. By parsing the operation instructions issued by the main control process, it dynamically retrieves the data in the corresponding section of the audio buffer, completes the algorithm inference calculation of the specified function, and feeds back the calculation results to the main control process according to the preset inference result feedback protocol.